Media content processing method and apparatus, device and storage medium

By generating and combining component information and initializing it with asynchronous task management optimization algorithm, the lag problem caused by the initialization of multiple algorithms in video effects processing is solved, and the smoothness of video playback and user experience is improved.

WO2025175881A1PCT designated stage Publication Date: 2025-08-28BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/138061
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-19
Filing Date
2024-12-10
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

In the prior art, in video special effects processing, the initialization of multiple algorithms causes video playback to be stuttered and affects the user experience.

Method used

By obtaining special-effect component information, generating merged component information, and performing preprocessing on the target component, the initialization process is utilized by asynchronous task management and resource pool optimization algorithm.

Benefits of technology

Improves the fluency of media content, improves user experience, and avoids lag problems caused by synchronization waiting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024138061_28082025_PF_FP_ABST
    Figure CN2024138061_28082025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a media content processing method and apparatus, a device and a storage medium. The method comprises: acquiring corresponding special effect component information of one or more special effects, wherein the one or more special effects will be applied within target time of media content, and the special effect component information of each special effect indicates an association relationship between components used by the special effect; on the basis of the corresponding special effect component information of the one or more special effects, generating merged component information related to a group of target components to be applied within the target time, the merged component information at least indicating an association relationship between the group of target components; and, on the basis of the merged component information, executing preprocessing of one or more target components among the group of target components. Therefore, the smoothness of media content display and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and storage medium for media content processing

[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, equipment and storage media for media content processing” and application number 202410185843.7, filed on February 19, 2024. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for media content processing. Background Art

[0003] With the rapid development of computer technology, a wide variety of short and long video content is becoming increasingly abundant. People's aesthetic standards for videos are also becoming increasingly sophisticated. Users are increasingly inclined to add rich and interesting special effects and gameplay to their videos. These special effects and gameplay are inseparable from the execution of algorithms. Therefore, we look forward to the efficient execution of algorithms to ensure smooth playback of media content and provide users with a better experience. Summary of the Invention

[0004] In a first aspect of the present disclosure, a method for media content processing is provided. The method includes: obtaining corresponding special effect component information for one or more special effects, wherein the one or more special effects are to be applied within a target time of the media content, the special effect component information for each special effect indicating an association relationship between components used by the special effect; generating, based on the corresponding special effect component information for the one or more special effects, merged component information related to a set of target components to be used at the target time, the merged component information at least indicating an association relationship between the set of target components; and performing preprocessing on one or more target components in the set of target components based on the merged component information.

[0005] In a second aspect of the present disclosure, a device for media content processing is provided. The device includes: an information acquisition module configured to acquire corresponding special effect component information of one or more special effects, wherein the one or more special effects are to be applied within a target time of the media content, and the special effect component information of each special effect indicates the association relationship between the components used by the special effect; an information generation module configured to generate, based on the corresponding special effect component information of the one or more special effects, merged component information related to a group of target components to be used at the target time, the merged component information at least indicating the association relationship between the group of target components; and an execution module configured to perform preprocessing on one or more target components in the group of target components based on the merged component information.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0009] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0011] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] 2A and 2B are schematic diagrams showing applicable scenarios according to some embodiments of the present disclosure;

[0013] FIG3 shows a schematic diagram of media content processing according to some embodiments of the present disclosure;

[0014] FIG4 shows a flowchart for performing pre-processing on a target component according to some embodiments of the present disclosure;

[0015] FIG5 illustrates a schematic diagram of an example architecture for component and rendering synchronization according to some embodiments of the present disclosure;

[0016] FIG6 shows a schematic diagram of an example architecture for media content processing according to some embodiments of the present disclosure;

[0017] FIG7 illustrates a flow chart of a process for media content processing according to some embodiments of the present disclosure;

[0018] FIG8 shows a schematic structural block diagram of an apparatus for media content processing according to certain embodiments of the present disclosure; and

[0019] FIG9 illustrates a block diagram showing an electronic device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0020] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0021] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0022] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0023] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0024] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0025] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0026] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.

[0027] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0028] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.

[0029] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. In this article, "model" may also be referred to as "machine learning model", "machine learning network" or "network", and these terms are used interchangeably in this article. A model can also include different types of processing units or networks.

[0030] As used herein, the term "component" may refer to any suitable model, module, unit, etc. used to implement a special effect. Such a component may provide a corresponding output based on the input provided and may include any suitable operation, calculation, etc. An example of a component is an algorithm. Below, some embodiments of the present disclosure will be primarily described with reference to algorithms, but it should be understood that such embodiments are also applicable to other types of components.

[0031] As used herein, the term "preprocessing" may include performing appropriate processing on a component prior to execution to place the component in a state ready for execution. As an example, preprocessing may include initialization, such as initializing parameters of the component (such as parameters of a model). Alternatively or additionally, preprocessing may include recording the component in a memory so that it is ready for execution. Preprocessing may also include other suitable processes or operations.

[0032] As briefly mentioned above, with the increasing richness of various short and long video content, people's aesthetic standards for videos are also getting higher and higher, and users are gradually tending to add rich and interesting special effects and gameplay to their videos. These special effects and gameplay are inseparable from algorithm detection, such as face detection (for example, beauty), human key point detection, etc. The first frame in which these special effects take effect will include the initialization of all related algorithms. According to effective statistics, the initialization of a single algorithm can reach tens or even hundreds of milliseconds. The superposition of multiple algorithms will inevitably cause the originally smoothly playing video to suddenly drop frames, causing obvious physical lag to the user.

[0033] Conventionally, when rendering the first frame of an effect, the system synchronously waits for all algorithms to initialize and complete inference before consuming the algorithm results. However, this approach can cause noticeable first-frame lag, resulting in a poor user experience. Another approach is to perform algorithm initialization asynchronously, but still rely on the algorithm results for rendering. Without proper waiting logic, rendering logic can fail to consume the algorithm results, potentially leading to script and program confusion, resulting in effects not taking effect or even program crashes.

[0034] In view of this, the embodiments of the present disclosure propose an improved solution for media content processing. According to various embodiments of the present disclosure, corresponding special effect component information of one or more special effects is obtained, and one or more special effects are to be applied within the target time of the media content. The special effect component information of each special effect indicates the association relationship between the components used by the special effect. Based on the corresponding special effect component information of one or more special effects, merged component information related to a group of target components to be used at the target time is generated, and the merged component information at least indicates the association relationship between a group of target components. Then, based on the merged component information, preprocessing is performed on one or more target components in a group of target components. Thereby, the smoothness of the media content display is improved, and the user experience is improved.

[0035] Sample Environment

[0036] 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the example environment 100, an application 120 is installed in a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or a device attached to the terminal device 110.

[0037] In some embodiments, the application 120 may be a content sharing application, a content editing application, a content creation application, etc. The application 120 can provide various services related to media content (also referred to as media content items, content items, media items, etc.) to the user 140, including browsing, commenting, forwarding, creating (e.g., shooting and / or editing), and publishing of media content.

[0038] In the environment 100 of FIG1 , if the application 120 is active, the terminal device 110 may present an interface 150 of the application 120. The interface 150 may include various interfaces provided by the application 120, such as a media content presentation interface, a media content creation interface, a media content publishing interface, etc. The application 120 may provide a media content editing function (for example, the application 120 may be an editing application) to support editing (e.g., editing) of media content in the application 120.

[0039] In some embodiments, the terminal device 110 communicates with the server 130 to enable the supply of services to the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.). The server 130 can be various types of computing systems / servers that can provide computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.

[0040] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.

[0041] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0042] The following description will use Figures 2A and 2B to illustrate scenarios applicable to some example embodiments of the present disclosure. However, this is merely exemplary and is not intended to limit the present disclosure. Figures 2A and 2B illustrate schematic diagrams of applicable scenarios according to some embodiments of the present disclosure. For ease of discussion, the following description will be from the perspective of terminal device 110.

[0043] As shown in FIG2A , the terminal device 110 receives selections of “special effect A” 211 - 1 and “special effect B” 211 - 2 and applies “special effect A” and “special effect B” to the video 212 according to an algorithm to provide services to the user.

[0044] As shown in Figure 2B, the terminal device 110 receives the selections of "special effect A" 211-1, "special effect B" 211-2, "special effect C" 221-1, and "special effect E" 221-2, and the terminal device 110 applies "special effect A", "special effect B", "special effect C", and "special effect E" to the video 212 according to the algorithm to provide services to the user.

[0045] However, the first frame in which these special effects take effect includes the initialization of all related algorithms. Initialization of a single algorithm for each special effect can take tens or even hundreds of milliseconds. The superposition of multiple algorithms corresponding to multiple special effects will inevitably cause a video that was previously playing smoothly to suddenly drop frames, causing noticeable lag for the user. Therefore, the disclosed embodiments propose an improved solution for media content processing to address the aforementioned issues.

[0046] The following describes a solution for media content processing according to the present disclosure with reference to Figure 3. Figure 3 shows a schematic diagram of media content processing according to some embodiments of the present disclosure. For ease of discussion, the following description will be from the perspective of the terminal device 110.

[0047] In some embodiments, the terminal device 110 obtains corresponding special effect component information for one or more special effects. In some examples, the one or more special effects obtained by the terminal device 110 can be visual effects, sound effects, and the like. In some embodiments, the special effect component information is represented by a component graph, which includes nodes representing components and edges representing the relationships between components. For example, the corresponding special effect component information for one or more special effects can be represented by an algorithm graph. For ease of discussion, the description in the accompanying drawings or below will be based on an algorithm graph. This is merely exemplary.

[0048] As shown in FIG3 , terminal device 110 obtains "Special Effect A" 331, "Special Effect B" 332, and "Special Effect C" 333. The special effect component information corresponding to "Special Effect A" 331 is "Algorithm Graph 1" 331-1. The special effect component information corresponding to "Special Effect B" 332 is "Algorithm Graph 2" 332-1. The special effect component information corresponding to "Special Effect C" 333 is "Algorithm Graph 3" 333-1.

[0049] In some examples, the corresponding special effect components of one or more special effects acquired by the terminal device 110 can be used to indicate an algorithm, such as a face detection algorithm, a key point detection algorithm, or a skeleton detection algorithm. The algorithm can be implemented through a machine learning model.

[0050] In some embodiments, the one or more special effects obtained by the terminal device 110 are applied to the media content within the target time. The special effect component information of each special effect is used to indicate the association relationship between the components used by the special effect. In some embodiments, the target time can be determined by a time window with a predetermined length. For example, as shown in Figures 2A and 2B, "special effect A" and "special effect B" are applied to the media content within a time window, and "special effect A", "special effect B", "special effect C", and "special effect E" are applied to the media content within another time window.

[0051] In some embodiments, terminal device 110 generates merged component information related to a set of target components to be used within a target time based on the corresponding special effect component information of one or more special effects. The merged component information at least indicates the association relationship between the set of target components. In some examples, the association relationship between the set of target components may also indicate the execution status or preprocessing status of the components. Based on the merged component information, terminal device 110 performs preprocessing on one or more target components in the set of target components.

[0052] In some embodiments, merged component information is also represented by a component graph, which includes nodes representing components and edges representing associations between components. For example, the merged component information generated by the terminal device 110 based on the corresponding special effect component information of one or more special effects can be represented by a merge algorithm graph. For ease of discussion, the description in the accompanying drawings and below will be based on the merge algorithm graph. This is merely exemplary.

[0053] As shown in Figure 3, the terminal device 110 merges "Algorithm Diagram 1", "Algorithm Diagram 2" and "Algorithm Diagram 3" into "Merged Algorithm Diagram" 334 in the preloading interface 320 according to the "Algorithm Diagram 1", "Algorithm Diagram 2" and "Algorithm Diagram 3" corresponding to "Special Effect A", "Special Effect B" and "Special Effect C" respectively.

[0054] Additionally, the terminal device 110 may also obtain current component information indicating the association relationship between components currently in execution. The terminal device 110 determines a set of target components based on the components used by one or more special effects and the components currently in execution. Then, based on the set of target components, the terminal device 110 merges the corresponding special component information and the current component information into merged component information. In this way, the currently executing component information does not need to be initialized again, thereby reducing lag and providing a better user experience.

[0055] In some examples, current component information is also represented by a component graph, which includes nodes representing components and edges representing relationships between components. For example, current component information can be represented by an execution graph. For ease of discussion, the following description will be described using execution graphs. However, this is for illustrative purposes only.

[0056] As shown in FIG3 , terminal device 110 obtains the currently executing "execution graph" 335 and uses the union of "algorithm graph 1," "algorithm graph 2," "algorithm graph 3," and "execution graph" 335 as a set of target components. Terminal device 110 merges "algorithm graph 1," "algorithm graph 2," "algorithm graph 3," and "execution graph" into "merged algorithm graph" 334 via preloading interface 320 .

[0057] In some embodiments, the terminal device 110 determines a target component from a set of target components that is included in the special effect component information but not in the current component information. The terminal device 110 then sets the preprocessing status of the determined target component to pending preprocessing in the merged component information. In some embodiments, the execution status of the target component to be preprocessed can be set to non-executable.

[0058] In some examples, the current component information included in a set of target components has already been preprocessed, while the other components (i.e., target components) have not been preprocessed. Therefore, the target components are marked as pending preprocessing. In some examples, each algorithm corresponding to a special effect has a flag to indicate whether it is pending preprocessing or preprocessing.

[0059] In the embodiment of the present disclosure, first of all, all calls to the algorithm are organized in the form of a graph. For example, for special effects scenes, the upper layer configures all the algorithms needed for a special effect (for example, face, body or segmentation, etc.) as vertices in the graph, and agrees on the corresponding dependencies. The algorithm module merges the preloaded graph and the current execution graph. The status of all vertices in the preload (for example, corresponding to the algorithm) is set to be only initializable and not executable, and is restored to its original state only when waiting for the special effects to be actually executed.

[0060] In some embodiments, the merged component information may further indicate the corresponding pre-processing status of the set of target components. For example, whether the set of target components has been initialized, whether the set of target components is being initialized, etc. The terminal device 110 may determine one or more target components to be pre-processed from the set of target components based on the corresponding pre-processing status of the set of target components. In some examples, the terminal device 110 may determine one or more target components from the set of target components that have not been initialized or are not being initialized based on the corresponding pre-processing status of the set of target components.

[0061] The terminal device 110 then performs preprocessing on one or more target components. In some embodiments, the terminal device 110 encapsulates the preprocessing task for a given target component among the one or more target components as an asynchronous task. The terminal device 110 then utilizes an asynchronous resource pool (e.g., an asynchronous thread pool) to process the asynchronous task. In some embodiments, if the asynchronous task is completed, the terminal device 110 updates the preprocessing status of the given target component to preprocessed.

[0062] The following describes a process 400 for performing preprocessing on one or more target components with reference to FIG4 . FIG4 illustrates a flowchart 400 for performing preprocessing on target components according to some embodiments of the present disclosure. In some examples, process 400 is also referred to as an asynchronous task management process. For ease of description, reference will be made to FIG1 and FIG3 .

[0063] In some examples, each algorithm corresponding to a special effect has a flag to determine whether it has been initialized. After the terminal device 110 merges the execution graph 3 with the preload graph (e.g., "Algorithm Graph 1," "Algorithm Graph 2," and "Algorithm Graph 3") in the preload interface 320, it determines the merged algorithm graph 334.

[0064] In block 411, the terminal device 110 traverses each algorithm instance in the final merged algorithm graph 334. In block 412, the terminal device 110 determines whether the merged algorithm graph 334 is an algorithm graph that has been initialized or is being initialized.

[0065] In block 413, if the terminal device 110 determines that each algorithm instance in the merged algorithm graph 334 has not been initialized or is being initialized, the preload task (e.g., including initialization) is encapsulated into an asynchronous task. In block 414, the algorithm encapsulated into the asynchronous task is placed into an asynchronous thread pool (e.g., an asynchronous task management object). The asynchronous task management object then determines, based on the address of each algorithm, whether it is currently being initialized or has already been initialized. If so, it is skipped; otherwise, it is placed into the asynchronous thread pool for asynchronous loading.

[0066] For ease of understanding, an example architecture for synchronizing components (e.g., algorithms) and rendering in a media content processing process will be described below with reference to Figure 5. Figure 5 shows a schematic diagram of an example architecture 500 for synchronizing components (e.g., algorithms) and rendering according to some embodiments of the present disclosure.

[0067] As shown in Figure 5, the upper-level caller 511 interacts with the algorithm graph set 310 in the form of a preloading interface 320. When designing the interface, in addition to allowing the upper layer to pass in the algorithm graph that needs to be preloaded, it also supports the upper layer to set an asynchronous callback 512. In some embodiments, the terminal device 110 provides a prompt for the completion of the preprocessing task in response to determining that a group of target components have been preprocessed. In some examples, when all asynchronous initialization tasks added to the asynchronous thread pool 513 are executed, the upper-level caller 511 is aware of the completion of the task and whether the task is successful through the asynchronous callback 512. In this way, components (for example, algorithms) can be synchronized with rendering.

[0068] In other examples, when executing an algorithm, the normal link has timing guarantees (for example, the algorithm must be executed before rendering in the same frame). During the algorithm execution, the terminal device 110 determines whether the current algorithm has an asynchronous initialization task. If so, it waits for the task. This also allows components (for example, algorithms) to be synchronized with rendering.

[0069] The disclosed embodiments, through both of the above approaches, can address issues such as asynchronous algorithm initialization and rendering dependency on algorithm results. Without proper waiting logic, rendering logic may not be able to consume the algorithm results, potentially leading to script and program level confusion, resulting in ineffective special effects and even program crashes. This ensures that rendering correctly consumes the corresponding algorithm results.

[0070] To better understand the present disclosure, the following description of the present disclosure for media content processing is provided with reference to FIG6 . FIG6 shows a schematic diagram of an example architecture 600 for media content processing according to some embodiments of the present disclosure. This is merely exemplary, and for ease of discussion, the following description will be provided with reference to FIG1 .

[0071] In the scenario where the client algorithm capabilities needed for editing are required: the upper-level software development kit (SDK) creates the fragment information in advance, calls the preloading capability at the starting position of the special effect, organizes the algorithm graph, and adds the node graph in sequence. As shown in Figure 6, the upper-level software development kit (SDK) calls the algorithm capability for "Special Effect 0" 611 at time T0 (for example, the current time). Then, the preloading capability is called for "Special Effect 1" 612 at time T1, and the preloading capability is called for "Special Effect 2" 613 at time T2. In the component graph 0 corresponding to "Special Effect 0", component D (for example, an algorithm instance) and component C (for example, an algorithm instance) depend on component A. For example, the output of component A serves as the input of component D and component C. In the component graph 1 corresponding to "Special Effect 1", component B and component C depend on component A. In the component graph 2 corresponding to "Special Effect 2", component E depends on component A.

[0072] The preloaded algorithm nodes are not enabled by default. During the seekframe function used for precise positioning in the user interface development toolkit (swing), only the algorithm nodes of the current segment are enabled.

[0073] Then, the terminal device 110 parses "Special Effect 0", "Special Effect 1" and "Special Effect 2" through the special effect parser (Effect Parser) 614. The terminal device 110 uses the component manager (AigorithManager) to merge the component information (for example, component graph) corresponding to "Special Effect 0", "Special Effect 1" and "Special Effect 2" to obtain merged component information (for example, merged graph) 616. Since the components corresponding to "Special Effect 0" include component A, component D and component C, the components corresponding to "Special Effect 1" include component A, component B and component C, and the component information corresponding to "Special Effect 2" includes component A and component E. Therefore, the merged component information (for example, merged graph) 616 includes component A, component D, component C, component B and component E, wherein component A, component D and component C are marked as components that have been initialized, and component B and component E are marked as components to be initialized.

[0074] Finally, if the terminal device 110 determines that each algorithm instance in the merge graph 616 has not been initialized or is in the process of initialization, it encapsulates its preloaded task (e.g., initialization) into an asynchronous task. The algorithm encapsulated as an asynchronous task is then placed into the asynchronous thread pool. When all asynchronous initialization tasks added to the asynchronous thread pool are completed, the upper-level caller 511 is notified of the task completion and success via asynchronous callback 512.

[0075] Correspondingly, in other examples, in scenarios where client algorithm capabilities are needed for shooting (for example, XX application). Most camera applications will turn on the beauty function by default, and beauty will use the algorithm capabilities corresponding to the face algorithm. At this time, the corresponding algorithm graph can be organized as soon as the shooting begins by merging the corresponding special effect component information of one or more special effects into the corresponding asynchronous initialization capability of the merged component information. At the same time, start the face algorithm initialization task, and there is no need to block and wait until the algorithm is executed.

[0076] Therefore, when the user clicks on a special effect, the algorithm to be executed has been determined, so the preloading of the corresponding algorithm diagram can be started at the first time.

[0077] In summary, by merging the current component information with the corresponding special effect component information of one or more special effects, the first frame lag problem caused by synchronously waiting for all algorithms to initialize and complete can be solved. Furthermore, if the algorithm is initialized asynchronously, the embodiment of the present disclosure can achieve synchronization between the algorithm and rendering. As a result, the embodiment of the present disclosure can improve the smoothness of media playback and enhance the user experience.

[0078] Example Process

[0079] 7 shows a flow chart of a process 700 for media content processing according to some embodiments of the present disclosure. The process 700 may be implemented at the terminal device 110. The process 700 is described below with reference to FIG1.

[0080] In box 710, the terminal device 110 obtains corresponding special effect component information of one or more special effects, where the one or more special effects will be applied within the target time of the media content. The special effect component information of each special effect indicates the association relationship between the components used by the special effect.

[0081] In block 720 , the terminal device 110 generates merged component information related to a set of target components to be used at a target time based on corresponding special effect component information of one or more special effects. The merged component information at least indicates an association relationship between the set of target components.

[0082] In block 730 , the terminal device 110 performs pre-processing on one or more target components in the set of target components based on the merged component information.

[0083] In some embodiments, generating merged component information includes: obtaining current component information, the current component information indicating the association relationship between components currently in execution; determining a set of target components based on components used by one or more special effects and components currently in execution; and merging corresponding special effect component information and current component information into merged component information based on a set of target components.

[0084] In some embodiments, process 700 further includes: determining a target component included in the special effect component information but not included in the current component information from a set of target components; and setting the preprocessing status of the determined target component to be preprocessed in the merged component information.

[0085] In some embodiments, the merged component information also indicates corresponding preprocessing statuses of a set of target components, and performing preprocessing on one or more target components in a set of target components includes: determining one or more target components to be preprocessed from a set of target components based on the corresponding preprocessing statuses of the set of target components; and performing preprocessing on the one or more target components.

[0086] In some embodiments, performing preprocessing on one or more target components includes: for a given target component among the one or more target components, encapsulating a preprocessing task for the given target component as an asynchronous task; and processing the asynchronous task using an asynchronous resource pool.

[0087] In some embodiments, process 700 further includes updating a pre-processing status of the given target component to pre-processed in response to completion of the asynchronous task.

[0088] In some embodiments, process 700 further includes providing an indication of completion of the pre-processing task in response to determining that the set of target components are all pre-processed.

[0089] In some embodiments, the special effect component information and the merged component information are respectively represented by a component graph, which includes nodes representing components and edges representing association relationships between components.

[0090] In some embodiments, the target time is determined by a time window having a predetermined length.

[0091] Example devices and equipment

[0092] 8 shows a schematic structural block diagram of an apparatus 800 for media content processing according to certain embodiments of the present disclosure. Apparatus 800 may be implemented as or included in terminal device 110. Each module / component in apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.

[0093] As shown in the figure, the device 800 includes an information acquisition module 810, which is configured to obtain corresponding special effect component information of one or more special effects, where one or more special effects are to be applied within the target time of the media content, and the special effect component information of each special effect indicates the association relationship between the components used by the special effect.

[0094] The device 800 also includes an information generation module 820, which is configured to generate merged component information related to a group of target components to be used at a target time based on corresponding special effect component information of one or more special effects, and the merged component information at least indicates the association relationship between a group of target components.

[0095] The apparatus 800 further includes an execution module 830 configured to perform pre-processing on one or more target components in a group of target components based on the merged component information.

[0096] In some embodiments, the information generation module 820 is also configured to obtain current component information, which indicates the association relationship between components currently in execution; determine a set of target components based on the components used by one or more special effects and the components currently in execution; and merge the corresponding special effect component information and the current component information into merged component information based on the set of target components.

[0097] In some embodiments, the device module 800 also includes a component determination module, which is configured to determine a target component included in the special effect component information but not included in the current component information from a group of target components; and in the merged component information, set the preprocessing status of the determined target component to be preprocessed.

[0098] In some embodiments, the merged component information also indicates a corresponding preprocessing status of a set of target components, and the execution module 830 is further configured to determine one or more target components to be preprocessed from a set of target components based on the corresponding preprocessing status of the set of target components; and perform preprocessing of the one or more target components.

[0099] In some embodiments, the execution module 830 is further configured to encapsulate the pre-processing task for a given target component among the one or more target components into an asynchronous task; and utilize an asynchronous resource pool to process the asynchronous task.

[0100] In some embodiments, the apparatus 800 further includes a status updating module configured to update the pre-processing status of a given target component to pre-processed in response to completion of the asynchronous task.

[0101] In some embodiments, the apparatus 800 further includes a providing module configured to provide an indication of completion of the pre-processing task in response to determining that a group of target components have all been pre-processed.

[0102] In some embodiments, the special effect component information and the merged component information are respectively represented by a component graph, which includes nodes representing components and edges representing association relationships between components.

[0103] In some embodiments, the target time is determined by a time window having a predetermined length.

[0104] FIG9 shows a block diagram of an electronic device 900 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 900 shown in FIG9 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 900 shown in FIG9 can be used to implement the terminal device 110 of FIG1 .

[0105] As shown in FIG9 , electronic device 900 is a general-purpose electronic device. Components of electronic device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. Processing unit 910 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 920. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 900.

[0106] The electronic device 900 typically includes a plurality of computer storage media. Such media can be any accessible media that the electronic device 900 can access, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 920 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 930 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 900.

[0107] The electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 9 , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 920 may include a computer program product 925 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0108] The communication unit 940 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 900 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 900 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0109] The input device 950 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 960 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 900 may also communicate with one or more external devices (not shown) through the communication unit 940 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 900, or with any device that allows the electronic device 900 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0110] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0111] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0112] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0113] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0114] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0115] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A media content processing method, comprising: Obtaining corresponding special effect component information of one or more special effects, wherein the one or more special effects are to be applied within a target time of the media content, the special effect component information of each special effect indicating an association relationship between components used by the special effect; generating, based on corresponding special effect component information of the one or more special effects, merged component information related to a group of target components to be used at the target time, the merged component information at least indicating an association relationship between the group of target components; as well as Pre-processing is performed on one or more target components in the set of target components based on the merged component information.

2. The method according to claim 1, wherein generating the merge component information comprises: Acquire current component information, where the current component information indicates association relationships between components currently in execution state; Determining the set of target components based on the components used by the one or more special effects and the components currently being executed; as well as According to the group of target components, the corresponding special effect component information and the current component information are merged into the merged component information.

3. The method according to claim 2, further comprising: Determine, from the set of target components, a target component that is included in the special effect component information but is not included in the current component information; as well as In the merged component information, the preprocessing status of the determined target component is set to be preprocessed.

4. The method according to any one of claims 1 to 3, wherein the merged component information further indicates a corresponding pre-processing status of the set of target components, and performing pre-processing on one or more target components in the set of target components comprises: determining the one or more target components to be preprocessed from the set of target components based on respective preprocessing states of the set of target components; as well as Pre-processing of the one or more target components is performed.

5. The method of claim 4, wherein performing pre-processing on the one or more target components comprises: For a given target component among the one or more target components, encapsulate a pre-processing task for the given target component as an asynchronous task; as well as An asynchronous resource pool is used to process the asynchronous task.

6. The method according to claim 5, further comprising: In response to completion of the asynchronous task, the pre-processing status of the given target component is updated to pre-processed.

7. The method according to any one of claims 1 to 3, further comprising: In response to determining that the set of target components are all pre-processed, an indication of completion of the pre-processing task is provided.

8. The method according to any one of claims 1 to 3, wherein the special effect component information and the merged component information are respectively represented by a component graph, and the component graph includes nodes representing components and edges representing association relationships between components.

9. The method according to any one of claims 1 to 3, wherein the target time is determined by a time window having a predetermined length.

10. A device for media content processing, comprising: an information acquisition module configured to acquire corresponding special effect component information of one or more special effects to be applied within a target time of the media content, wherein the special effect component information of each special effect indicates an association relationship between components used by the special effect; an information generating module configured to generate, based on corresponding special effect component information of the one or more special effects, merged component information related to a group of target components to be used at the target time, the merged component information at least indicating an association relationship between the group of target components; as well as The execution module is configured to perform pre-processing on one or more target components in the group of target components based on the merged component information.

11. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processing unit.

12. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product comprising a computer program, wherein the computer program implements the method according to any one of claims 1 to 9 when executed by a processor.

Citation Information

Patent Citations

  • Media content processing method and device, equipment and storage medium

    CN118175368A

  • Special effect merging method and device, electronic equipment and computer readable storage medium

    CN110221822A

  • Special effect processing method and device, electronic equipment and storage medium

    CN110674341A

  • Special effect rendering method and device, electronic equipment and computer readable storage medium

    CN116204163A

  • Method and apparatus for creating and managing high impact special effects

    US20070296734A1