Media content processing method and apparatus, device, and storage medium

By generating and displaying media content on the video playback page, the problem of single functions in the existing technology is solved, rich media content interaction functions are realized, and user experience is improved.

WO2025167287A1PCT designated stage Publication Date: 2025-08-14BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/136182
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-05
Filing Date
2024-12-02
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing media content processing applications have a single function and require switching back and forth between multiple applications to complete processing, resulting in poor user experience.

Method used

The pause frame screen is obtained by pausing video playback, and the media content is generated and displayed using the preset biographic model, which supports rich interactive functions on the video playback page, such as content tags, editing, publishing, etc.

Benefits of technology

It realizes the function of generating and displaying media content on the video playback page, enriches the interactive functions related to media content, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136182_14082025_PF_FP_ABST
    Figure CN2024136182_14082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a media content processing method and apparatus, a device, and a storage medium. The method comprises: in response to a playback pause operation on a playback page of a target video, pausing playback of the target video, and acquiring a pause frame picture corresponding to the target video; and generating first target media content on the basis of the pause frame picture, and displaying the first target media content on the playback page of the target video, wherein the first target media content is generated on the basis of part or all of the content on the pause frame picture and by using a preset image generation model.
Need to check novelty before this filing date? Find Prior Art

Description

A media content processing method, device, equipment and storage medium

[0001] This application claims priority to the Chinese patent application filed on February 5, 2024, with application number 202410167955.X and invention name “A media content processing method, device, equipment and storage medium”. The entire contents of the application are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of data processing, and in particular to a method, apparatus, device, and storage medium for processing media content. Background Art

[0003] With the continuous development of computer technology, applications related to media content processing have emerged in an endless stream. For example, one application has a photo editing function, while another application has a video editing function. Summary of the Invention

[0004] In order to solve the above technical problems, the embodiments of the present disclosure provide a media content processing method, apparatus, device, and storage medium.

[0005] In a first aspect, the present disclosure provides a method for processing media content, the method comprising:

[0006] In response to a pause operation on a play page of a target video, pausing the target video and acquiring a pause frame corresponding to the target video;

[0007] Generate first target media content based on the pause frame image, and display the first target media content on the playback page of the target video; wherein, the first target media content is generated based on part or all of the content on the pause frame image using a preset image generation model.

[0008] In an optional implementation manner, before generating the first target media content based on the pause frame image and after pausing the target video, the method further includes:

[0009] Displaying at least one content tag on the pause frame of the target video; wherein the content tag is obtained by analyzing the pause frame, and the content tag is used to identify part or all of the content on the pause frame;

[0010] Accordingly, generating the first target media content based on the pause frame image includes:

[0011] In response to a selection operation on a target content tag among the at least one content tag, first target media content is generated based on the target content tag and the pause frame image.

[0012] In an optional implementation, the method further includes:

[0013] In response to a regeneration operation triggered for the first target media content, generating second target media content based on the pause frame image using the preset image generation model;

[0014] The first target media content is replaced by the second target media content and displayed on the playback page of the target video.

[0015] In an optional implementation, the method further includes:

[0016] In response to an edit triggering operation on the first target media content, displaying a media content editing page;

[0017] In response to a first editing operation triggered on the media content editing page, the first target media content is updated and presented based on the first editing operation.

[0018] In an optional embodiment, in response to a first editing operation triggered on the media content editing page, before updating and displaying the first target media content based on the first editing operation, the method further includes:

[0019] Displaying an editing session window on the media content editing page;

[0020] Accordingly, in response to the first editing operation triggered on the media content editing page, updating and displaying the first target media content based on the first editing operation includes:

[0021] In response to an input operation on a content editing text in the editing session window, generating third target media content based on a first editing operation indicated by the content editing text and the first target media content;

[0022] In response to a preset application operation on the third target media content, the first target media content is replaced by the third target media content for presentation.

[0023] In an optional implementation, in response to a first editing operation triggered on the media content editing page, updating and displaying the first target media content based on the first editing operation includes:

[0024] In response to an editing operation triggered on the media content editing page for a target object in the first target media content, displaying a plurality of candidate objects corresponding to the target object on the media content editing page;

[0025] In response to a selection operation on a target candidate object among the multiple candidate objects, the target object in the first target media content is replaced with the target candidate object; wherein the target candidate object is used to update and present the first target media content.

[0026] In an optional implementation, the method further includes:

[0027] In response to a work publishing operation for the first target media content, a multimedia work is generated and published based on the first target media content.

[0028] In an optional implementation manner, before generating the first target media content based on the pause frame image, the method further includes:

[0029] Acquire a target digital model; wherein the target digital model is constructed based on pre-input modeling parameter information;

[0030] Accordingly, generating the first target media content based on the pause frame image includes:

[0031] First target media content is generated based on the pause frame image and the target digital model.

[0032] In an optional implementation, the method further includes:

[0033] In response to a comment triggering operation on the target video, a comment panel is displayed, and recommended comment content is displayed for the comment input box on the comment panel; wherein the recommended comment content includes the first target media content.

[0034] In an optional implementation, the method further includes:

[0035] In response to a comment trigger operation for the target video, a comment panel is displayed; wherein, the comment panel displays first target comment content including fourth target media content, and the first target comment content is provided with a first target control, which is used to trigger the generation of fifth target media content based on the fourth target media content.

[0036] In an optional implementation, the method further includes:

[0037] In response to a display trigger operation for the sixth target media content on the comment panel, the sixth target media content is displayed based on the image comment information flow corresponding to the target video; wherein, the sixth target media content belongs to the second target comment content for the target video, and a second target control is set on the display page of the sixth target media content, and the second target control is used to trigger the generation of the seventh target media content based on the sixth target media content.

[0038] In a second aspect, the present disclosure further provides a media content processing device, the device comprising:

[0039] A first pause module, configured to pause the target video in response to a pause operation performed on a play page of the target video, and obtain a pause frame corresponding to the target video;

[0040] The first generation module is used to generate first target media content based on the pause frame image, and display the first target media content on the playback page of the target video; wherein, the first target media content is generated based on part or all of the content on the pause frame image using a preset raw image model.

[0041] In a third aspect, the present disclosure provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium. When the instructions are executed on a terminal device, the terminal device implements the above method.

[0042] In a fourth aspect, the present disclosure provides a media content processing device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method when executing the computer program.

[0043] In a fifth aspect, the present disclosure provides a computer program product, which includes a computer program / instructions, and the computer program / instructions implement the above method when executed by a processor.

[0044] An embodiment of the present disclosure provides a media content processing method. Specifically, in response to a pause playback operation on a playback page of a target video, the target video is paused, and a pause frame image corresponding to the target video is obtained. First target media content is generated based on the pause frame image, and the first target media content is displayed on the playback page of the target video, wherein the first target media content is generated based on part or all of the content on the pause frame image using a preset image generation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0046] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] FIG1 is a flow chart of a method for processing media content provided by an embodiment of the present disclosure;

[0048] FIG2 is a schematic diagram of a target video playback page provided by an embodiment of the present disclosure;

[0049] FIG3 is a schematic diagram of another target video playback page provided by an embodiment of the present disclosure;

[0050] FIG4 is a schematic diagram of another target video playback page provided by an embodiment of the present disclosure;

[0051] FIG5 is a schematic diagram of another target video playback page provided by an embodiment of the present disclosure;

[0052] FIG6 is a schematic diagram of a media content editing page provided by an embodiment of the present disclosure;

[0053] FIG7 is a schematic diagram of another media content editing page provided by an embodiment of the present disclosure;

[0054] FIG8 is a schematic diagram of another media content editing page provided by an embodiment of the present disclosure;

[0055] FIG9 is a schematic diagram of another media content editing page provided by an embodiment of the present disclosure;

[0056] FIG10 is a schematic diagram of a target video playback page provided by an embodiment of the present disclosure;

[0057] FIG11 is a schematic diagram of another target video playback page provided by an embodiment of the present disclosure;

[0058] FIG12 is a schematic diagram of a media content display page provided by an embodiment of the present disclosure;

[0059] FIG13 is a schematic structural diagram of a media content processing device provided by an embodiment of the present disclosure;

[0060] FIG14 is a schematic structural diagram of a media content processing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0061] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.

[0062] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0063] Currently, applications related to media content processing offer relatively limited functionality. For example, an application might only offer photo editing or editing. Consequently, processing media content often requires switching between multiple applications. Therefore, enriching interactive features related to media content and improving user experience is a pressing technical challenge.

[0064] To this end, an embodiment of the present disclosure provides a media content processing method. Specifically, in response to a pause playback operation on the playback page of a target video, the target video is paused, and a pause frame image corresponding to the target video is obtained. First target media content is generated based on the pause frame image, and the first target media content is displayed on the playback page of the target video, wherein the first target media content is generated based on part or all of the content on the pause frame image using a preset image generation model.

[0065] The technical solution provided by the embodiments of the present disclosure has at least the following advantages compared with the prior art:

[0066] The disclosed embodiments support the function of generating the first target media content based on part or all of the content on the pause frame of the target video and displaying it on the video playback page by triggering the pause of the video, thereby enriching the media content-related interactive functions and improving the user experience.

[0067] Specifically, an embodiment of the present disclosure provides a media content processing method. Referring to FIG1 , which is a flow chart of a media content processing method provided by an embodiment of the present disclosure, the method specifically includes:

[0068] S101: In response to a pause operation on a play page of a target video, the target video is paused, and a pause frame corresponding to the target video is obtained.

[0069] The media content processing method provided by the embodiments of the present disclosure can be applied to a client. Specifically, the client can be deployed on a terminal such as a smart phone, a tablet computer, or a desktop computer.

[0070] In the disclosed embodiments, the target video can be any video. During playback of the target video, when a pause operation is received on the target video's playback page, the target video is paused. At this time, a pause frame corresponding to the target video is obtained. Specifically, content such as character images, background images, clothing, and accessories in the pause frame can be obtained. The pause frame is the frame corresponding to the target video at the moment of pause.

[0071] The operation of acquiring the pause frame image may include acquiring the content of the pause frame image using techniques such as image recognition. The embodiment of the present disclosure does not impose any limitation on the specific method of acquiring the pause frame image.

[0072] S102: Generate first target media content based on the pause frame image, and display the first target media content on the playback page of the target video.

[0073] The first target media content is generated based on part or all of the content on the pause frame screen using a preset image generation model.

[0074] In the embodiment of the present disclosure, based on part or all of the content on the acquired pause frame image, a preset image generation model is used to generate the first target media content, and the content is displayed on the playback page of the target video.

[0075] Specifically, part or all of the content on the pause frame is input into a preset raw image model. The preset raw image model generates first target media content based on part or all of the content on the pause frame and displays it on the playback page of the target video. In this case, the generated first target media content is related to the pause frame.

[0076] In actual applications, the first target media content may be generated based on the target content tag and the pause frame image.

[0077] In an optional embodiment, when a pause frame playback operation is received on the playback page of a target video, the target video is paused, the pause frame image is obtained, and a content label is obtained by analyzing part or all of the content on the pause frame image and displayed on the corresponding content of the pause frame image. The content label is used to describe part or all of the content on the pause frame image.

[0078] Analyzing part or all of the content on the pause frame may include analyzing the pause frame using techniques such as image analysis to obtain a content tag, which is not limited in any way by the present disclosure.

[0079] As shown in Figure 2, a schematic diagram of a target video playback page provided by an embodiment of the present disclosure is shown on the target video playback page. The target video playback page displays a pause frame image corresponding to the target video, as well as content tags 201 and 202.

[0080] Based on displaying a pause frame image and content tags corresponding to the target video on the target video's playback page, when a selection operation is received for a target content tag among at least one of the content tags, first target media content is generated using a preset image generation model based on the target content tag and the pause frame image, and displayed on the target video's playback page. The first target media content includes media content related to the target content tag.

[0081] In one application scenario, when a target content tag is selected, a media content display panel is displayed on the playback page of the target video. The media content display panel displays the progress information of generating the first target media content. Part or all of the content on the pause frame screen and the target content tag are input into a preset raw image model. The preset raw image model generates the first target media content based on part or all of the content on the pause frame screen and the target content tag, and displays it on the media content display panel.

[0082] As shown in Figure 3, a schematic diagram of another target video playback page provided by an embodiment of the present disclosure is shown. The target video playback page displays a media content display panel 301. During the process of generating the first target media content based on the pause frame corresponding to the target video, the media content display panel 301 can display the first target media content generation progress information, such as the generation progress percentage information or the "generating" progress prompt information.

[0083] 4 is a schematic diagram of another target video playback page provided by an embodiment of the present disclosure, wherein a media content display panel 401 is displayed on the target video playback page, and a first target media content 402 is displayed on the media content display panel.

[0084] In actual applications, to support switching target content tags to meet user needs, in the disclosed embodiments, during the generation of the first target media content, content corresponding to the content tag in the pause frame image can be displayed on the target video's playback page. By triggering the content corresponding to the content tag, the first target media content can be updated and displayed. For example, FIG4 shows content 403 and 404 corresponding to the content tag displayed on the target video's playback page.

[0085] In actual applications, the first target media content may also be generated based on the target digital model and the pause frame image.

[0086] In an optional embodiment, when a pause frame playback operation is received on the playback page of the target video, the target video is paused, the pause frame image and the target digital model are obtained, and based on the pause frame image and the target digital model, the first target media content is generated using a preset raw image model.

[0087] The target digital model may be constructed using pre-input modeling parameter information, where the modeling parameter information is information required to construct the target digital model.

[0088] Based on the above embodiment, more preferably, the first target media content may be generated based on the pause frame image, the target content tag, and the target digital model using a preset raw image model.

[0089] On the basis of the above content, based on the generated first target media content, a work of the first target media content may also be published.

[0090] Specifically, when the generated first target media content is displayed on the playback page of the target video, when a work publishing operation for the first target media content is received, a multimedia work is generated based on the first target media content and published.

[0091] The work publishing operation for the first target media content can be implemented by triggering a publishing control to achieve the work publishing of the first target media content. The embodiment of the present disclosure does not impose any limitation on the work publishing operation for the first target media content.

[0092] In the embodiment of the present disclosure, by generating a multimedia work for the generated first target media content and publishing the multimedia work, rapid publishing of the first target media content is supported.

[0093] In the media content processing method provided by the embodiments of the present disclosure, specifically, in response to a pause playback operation on the playback page of the target video, the target video is paused, and a pause frame image corresponding to the target video is obtained, first target media content is generated based on the pause frame image, and the first target media content is displayed on the playback page of the target video, wherein the first target media content is generated based on part or all of the content on the pause frame image using a preset image generation model.

[0094] The disclosed embodiments support the function of generating the first target media content based on part or all of the content on the pause frame of the target video and displaying it on the video playback page by triggering the pause of the video, thereby enriching the media content-related interactive functions and improving the user experience.

[0095] On the basis of the above content, in order to further enrich the function of media content processing, the first target internal content displayed on the playback page of the target video can also be regenerated, that is, the second target media content can be regenerated based on the pause frame picture.

[0096] Specifically, based on the first target media content displayed on the playback page of the target video, when a regeneration operation triggered for the first target media content is received, a preset raw image model is used to generate a second target media content based on the pause frame image, and the second target media content is used to replace the first target media content and displayed on the playback page of the target video.

[0097] The regeneration operation may include displaying a regeneration control on the first target media content, and triggering the regeneration control to facilitate the preset image generation model to generate the second target media content based on the pause frame image.

[0098] As shown in Figure 5, a schematic diagram of another target video playback page provided by an embodiment of the present disclosure is provided. The first target media content is displayed on the target video playback page, and a regeneration control 501 is displayed on the first target media content. When the regeneration control 501 is triggered, a second target media content is generated based on the pause frame image using a preset image generation model, and the first target media content is replaced with the second target media content. At this time, the second target media content is displayed on the target video playback page, and the second target media content is related to the pause frame image.

[0099] In actual applications, after the first target media content is generated and displayed on the playback page of the target video, the first target media content may be edited to optimize the first target media content and meet user needs.

[0100] Specifically, when an edit trigger operation for the first target media content is received, the media content edit page is displayed. When a first edit operation triggered on the media content edit page is received, the first target media content is updated and displayed based on the first edit operation. The edit trigger operation for the first target media content may include a trigger operation for an edit control. Specifically, an edit control may be displayed on the first target media content, and the media content edit page is displayed by triggering the edit control. For example, the edit control 502 in the play page of the target video in FIG5 is shown above. The first edit operation is any edit operation.

[0101] In the embodiment of the present disclosure, after the media content editing page is displayed, a first editing operation may be triggered based on a target object of the first target media content.

[0102] In an optional embodiment, after receiving an editing trigger operation for the first target media content and displaying the media content editing page, when receiving an editing operation triggered on the media content editing page for a target object in the first target media content, multiple candidate objects corresponding to the target object are displayed on the media content editing page.

[0103] The target object is one or more components of the first target media content. For example, if the first target media content includes background environment, clothing, shoes, hats, etc., the target object can be clothing. The multiple candidate objects are objects with the same or similar attributes as the target object. For example, if the target object is clothing, the multiple candidate objects can be clothing of the same brand or style as the target object (clothing).

[0104] In an embodiment of the present disclosure, the editing operation triggered for the target object in the first media content may specifically include a triggering operation for a label control of the target object. In one application scenario, when an editing triggering operation for the first target media content is received, a media content editing page is displayed, on which the first target media content is displayed, and a label control for displaying the target object on the first target media content is displayed. When the label control for the target object is triggered, multiple candidate objects corresponding to the target object are displayed on the media content editing page. The label control is used to trigger the display of candidate objects corresponding to the target object.

[0105] As shown in Figure 6, a schematic diagram of a media content editing page provided in an embodiment of the present disclosure is shown on the media content editing page. A first target media content 601 is displayed on the media content editing page, and label controls for multiple target objects are displayed on the first target media content, including a label control 602 and a label control 603. When the label control 602 of one of the target objects is triggered, multiple candidate objects corresponding to the target object are displayed on the media content editing page. More candidate objects can be browsed by sliding the media content editing page.

[0106] 7 is a schematic diagram of another media content editing page provided by an embodiment of the present disclosure, wherein the media content editing page displays first target media content and a plurality of candidate objects corresponding to the target object of the first target media content, including candidate object 701 .

[0107] In the disclosed embodiment, after displaying multiple candidate objects corresponding to a target object on a media content editing page, when a selection operation is received for a target candidate object among the multiple candidate objects, the target candidate object is used to replace the target object in the first target media content. The target candidate object is used to update the displayed first target media content, specifically updating the target object on the displayed first target media content. In other words, based on the target candidate object, the first target media content containing the target candidate object is generated. The target candidate object is any one of the candidate objects.

[0108] 7 , when candidate object 701 is selected, candidate object 701 is the target candidate object, and the target candidate object is used to replace the target object (clothing) in the first target media content. At this time, the first target media content displayed on the media content editing page includes the candidate target object.

[0109] In actual applications, when the first target media content is displayed on a media content editing page, and when the first target media content includes a candidate target object, an object information card for the candidate target object can be displayed on the media content editing page. The object information card can specifically display image information, price information, etc. of the candidate target object. By triggering the object information card, detailed information about the candidate target object can be displayed.

[0110] Specifically, when a trigger action is received for an object information card, the media content editing page switches to the object details page displaying the target candidate object. The object details page displays detailed information about the target candidate object and supports target interaction operations for the target candidate object, such as placing a purchase order.

[0111] As shown in FIG. 7 , the media content editing page displays an object information card 702 of a target candidate object. By triggering the object information card, the media content editing page switches to an object details page displaying the target candidate object.

[0112] In the embodiment of the present disclosure, after the media content editing page is displayed, an editing session window may be further displayed on the media content editing page to trigger the first editing operation.

[0113] In an optional embodiment, when an editing trigger operation for the first target media content is received, a media content editing page is displayed, and an editing session window is displayed on the media content editing page. When an input operation for the content editing text is received in the editing session window, a third target media content is generated based on the first editing operation indicated by the content editing text and the first target media content using a preset raw image model.

[0114] In the disclosed embodiment, triggering the display of the editing session window may include triggering the display of the editing session window control. Specifically, the editing session window control is displayed on the media content editing page. When the editing session window control is triggered, the editing session window is displayed on the media content editing page.

[0115] As shown in FIG6 , the media content editing page displays an editing session window control 604 . When the editing session window control 604 is triggered, an editing session window is displayed on the media content editing page.

[0116] In addition, one or more preset content edit texts may be displayed at the edit session window control to prompt the user to trigger the edit session window control. By triggering the preset content edit texts, the edit session window may also be triggered to be displayed on the media content edit page. For example, the preset content edit text 605 in the media content edit page shown in FIG. 6 is shown above.

[0117] FIG8 is a schematic diagram of another media content editing page provided by an embodiment of the present disclosure, which displays an editing session window 801. When content editing text is entered and sent, third target media content 802 is generated based on the content editing text and first target media content using a preset image generation model and displayed in the editing session window.

[0118] Based on the third target media content displayed in the editing session window, the first target media content may be replaced by the third target media content to be displayed on the media content editing page.

[0119] Specifically, when a preset application operation for the third target media content is received, the first target media content is replaced by the third target media content and displayed on the media content editing page.

[0120] Among them, the preset application operation for the third target media content may include setting a usage control within the display area of ​​the third target media content, such as the usage control 803 displayed on the media content editing page shown in Figure 8 above. By triggering the usage control 803, the third target media content replaces the first target media content. At this time, the media content editing page is displayed, and the third target media content is displayed on the media content editing page.

[0121] 9 is a schematic diagram of another media content editing page provided by an embodiment of the present disclosure, wherein the media content editing page displays third target media content 802 .

[0122] In actual applications, after generating the first target media content based on the pause frame image corresponding to the target video, recommended comment content for the target video can be set based on the first target media content, wherein the recommended comment content includes the first target media content, that is, there is an association between the recommended comment content and the commented target video.

[0123] Specifically, first, after generating the first target media content based on the pause frame image corresponding to the target video, the first target media content is saved locally by triggering the save control. Then, when a comment trigger operation for the target video is received, the comment panel is displayed, and the recommended comment content is displayed for the comment input box on the comment panel, that is, when the comment input box on the comment panel is triggered, the recommended comment content is displayed, and the recommended comment content includes the first target media content.

[0124] As shown in Figure 10, a schematic diagram of a target video playback page provided by an embodiment of the present disclosure is shown. A comment panel 1001 is displayed on the target video playback page, and when the comment input box on the comment panel is triggered, recommended comment content 1002 is displayed.

[0125] In actual applications, while a target video is playing on a playback page, a user can trigger the generation of fifth target media content based on a comment on the target video. Specifically, when a comment trigger operation is received for the target video, a comment panel is displayed, which displays first target comment content, which is a comment content containing fourth target media content, and a first target control is set on the first target comment content, which is used to trigger the generation of fifth target media content based on the fourth target media content. The fourth target media content is generated by other users based on the pause frame of the target video.

[0126] As shown in Figure 11, a schematic diagram of another target video playback page provided by an embodiment of the present disclosure is shown. The target video playback page displays a comment panel 1101, which displays first target comment content 1102. The first target comment content includes comment content for fourth target media content 1103, and the first target comment content is provided with a first target control 1104. By triggering the first target control 1104, fifth target media content is generated based on the fourth target media content using a preset raw image model.

[0127] In actual applications, after a comment panel is displayed on the target video's playback page, the seventh target media content can be generated based on the sixth target media content displayed in the image comment information stream corresponding to the target video. The sixth target media content is the second target comment content for the target video.

[0128] Specifically, when a display trigger operation for the sixth target media content on the comment panel is received, the sixth target media content is displayed based on the image review information flow corresponding to the target video. A second target control is set on the display page of the sixth target media content. By triggering the second target control, the seventh target media content can be generated based on the sixth target media content using a preset raw image model.

[0129] FIG12 is a schematic diagram of a media content display page provided by an embodiment of the present disclosure. The media content display page displays sixth target media content and a second target control 1201. By triggering the second target control, seventh target media content is generated based on the sixth media content using a preset image generation model.

[0130] Based on the above method embodiment, the present disclosure further provides a media content processing device. Referring to FIG13 , which is a schematic structural diagram of a media content processing device provided by an embodiment of the present disclosure, the device includes:

[0131] A first pause module 1301 is configured to pause the target video in response to a pause operation performed on the target video's playback page, and obtain a pause frame corresponding to the target video;

[0132] The first generation module 1302 is used to generate first target media content based on the pause frame image, and display the first target media content on the playback page of the target video; wherein, the first target media content is generated based on part or all of the content on the pause frame image using a preset raw image model.

[0133] In an optional embodiment, the device further includes:

[0134] a first display module, configured to display at least one content tag on the pause frame of the target video; wherein the content tag is obtained by analyzing the pause frame, and is used to identify part or all of the content on the pause frame;

[0135] Accordingly, the first generating module is specifically configured to:

[0136] In response to a selection operation on a target content tag among the at least one content tag, first target media content is generated based on the target content tag and the pause frame image.

[0137] In an optional embodiment, the device further includes:

[0138] A second generating module is configured to generate second target media content based on the pause frame image using the preset image generation model in response to a regeneration operation triggered for the first target media content;

[0139] The first replacement module is configured to replace the first target media content with the second target media content for display on the playback page of the target video.

[0140] In an optional embodiment, the device further includes:

[0141] a second display module, configured to display a media content editing page in response to an editing trigger operation on the first target media content;

[0142] The update display module is configured to respond to a first editing operation triggered on the media content editing page and update and display the first target media content based on the first editing operation.

[0143] In an optional embodiment, the device further includes:

[0144] a third display module, configured to display an editing session window on the media content editing page;

[0145] Accordingly, the update display module includes:

[0146] a third generating module, configured to generate, in response to an input operation on a content editing text in the editing session window, third target media content based on a first editing operation indicated by the content editing text and the first target media content;

[0147] The second replacement module is configured to, in response to a preset application operation on the third target media content, replace the first target media content with the third target media content for display.

[0148] In an optional implementation, the update display module includes:

[0149] a first display module, configured to display, on the media content editing page, a plurality of candidate objects corresponding to the target object in the first target media content in response to an editing operation triggered on the media content editing page for the target object;

[0150] A third replacement module is configured to, in response to a selection operation on a target candidate object among the multiple candidate objects, replace the target object in the first target media content with the target candidate object; wherein the target candidate object is used to update and display the first target media content.

[0151] In an optional embodiment, the device further includes:

[0152] The publishing module is configured to generate and publish a multimedia work based on the first target media content in response to a work publishing operation for the first target media content.

[0153] In an optional embodiment, the device further includes:

[0154] An acquisition module is used to acquire a target digital model; wherein the target digital model is constructed based on pre-input modeling parameter information;

[0155] Accordingly, the first generating module is specifically configured to:

[0156] First target media content is generated based on the pause frame image and the target digital model.

[0157] In an optional embodiment, the device further includes:

[0158] A fourth display module is configured to display a comment panel in response to a comment triggering operation on the target video, and to display recommended comment content for the comment input box on the comment panel; wherein the recommended comment content includes the first target media content.

[0159] In an optional embodiment, the device further includes:

[0160] A fifth display module is used to display a comment panel in response to a comment trigger operation for the target video; wherein, the comment panel displays first target comment content including fourth target media content, and the first target comment content is provided with a first target control, which is used to trigger the generation of fifth target media content based on the fourth target media content.

[0161] In an optional embodiment, the device further includes:

[0162] The second display module is used to respond to the display trigger operation of the sixth target media content on the comment panel, and display the sixth target media content based on the picture comment information flow corresponding to the target video; wherein, the sixth target media content belongs to the second target comment content for the target video, and a second target control is set on the display page of the sixth target media content, and the second target control is used to trigger the generation of the seventh target media content based on the sixth target media content.

[0163] In the media content processing device provided by the embodiment of the present disclosure, in response to a pause playback operation performed on the playback page of the target video, the target video is paused, and a pause frame image corresponding to the target video is obtained, first target media content is generated based on the pause frame image, and the first target media content is displayed on the playback page of the target video, wherein the first target media content is generated based on part or all of the content on the pause frame image using a preset image generation model.

[0164] The disclosed embodiments support the function of generating the first target media content based on part or all of the content on the pause frame of the target video and displaying it on the video playback page by triggering the pause of the video, thereby enriching the media content-related interactive functions and improving the user experience.

[0165] In addition to the above-mentioned method and apparatus, the embodiments of the present disclosure further provide a computer-readable storage medium, which stores instructions. When the instructions are executed on a terminal device, the terminal device implements the media content processing method described in the embodiments of the present disclosure.

[0166] The embodiments of the present disclosure further provide a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, the media content processing method described in the embodiments of the present disclosure is implemented.

[0167] In addition, the embodiment of the present disclosure further provides a media content processing device, as shown in FIG14 , which may include:

[0168] Processor 1401, memory 1402, input device 1403, and output device 1404. The page interaction device may include one or more processors 1401, with one processor being used as an example in FIG14 . In some embodiments of the present disclosure, processor 1401, memory 1402, input device 1403, and output device 1404 may be connected via a bus or other means, with FIG14 using a bus as an example.

[0169] The memory 1402 can be used to store software programs and modules. The processor 1401 executes various functional applications and data processing of the page interaction device by running the software programs and modules stored in the memory 1402. The memory 1402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc. In addition, the memory 1402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. The input device 1403 can be used to receive input digital or character information, and to generate signal input related to user settings and function control of the page interaction device.

[0170] Specifically in this embodiment, the processor 1401 will load the executable files corresponding to the processes of one or more applications into the memory 1402 according to the following instructions, and the processor 1401 will run the applications stored in the memory 1402, thereby realizing the various functions of the above-mentioned page interaction device.

[0171] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0172] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.

Claims

1. A method for processing media content, wherein: The method comprises: In response to a pause operation on a play page of a target video, pausing the target video and acquiring a pause frame corresponding to the target video; Generate first target media content based on the pause frame image, and display the first target media content on the playback page of the target video; wherein, the first target media content is generated based on part or all of the content on the pause frame image using a preset image generation model.

2. The method according to claim 1, wherein Before generating the first target media content based on the pause frame image and after pausing the target video, the method further includes: Displaying at least one content tag on the pause frame of the target video; wherein the content tag is obtained by analyzing the pause frame, and the content tag is used to identify part or all of the content on the pause frame; Accordingly, generating the first target media content based on the pause frame image includes: In response to a selection operation on a target content tag among the at least one content tag, first target media content is generated based on the target content tag and the pause frame image.

3. The method according to claim 1, wherein The method further comprises: In response to a regeneration operation triggered for the first target media content, generating second target media content based on the pause frame image using the preset image generation model; The first target media content is replaced by the second target media content and displayed on the playback page of the target video.

4. The method according to claim 1, wherein The method further comprises: In response to an edit triggering operation on the first target media content, displaying a media content editing page; In response to a first editing operation triggered on the media content editing page, the first target media content is updated and presented based on the first editing operation.

5. The method according to claim 4, wherein The method further includes: responding to a first editing operation triggered on the media content editing page, and updating and displaying the first target media content based on the first editing operation; Displaying an editing session window on the media content editing page; Accordingly, in response to the first editing operation triggered on the media content editing page, updating and displaying the first target media content based on the first editing operation includes: In response to an input operation on a content editing text in the editing session window, generating third target media content based on a first editing operation indicated by the content editing text and the first target media content; In response to a preset application operation on the third target media content, the first target media content is replaced by the third target media content for presentation.

6. The method according to claim 4, wherein: In response to a first editing operation triggered on the media content editing page, updating and presenting the first target media content based on the first editing operation includes: In response to an editing operation triggered on the media content editing page for a target object in the first target media content, displaying a plurality of candidate objects corresponding to the target object on the media content editing page; In response to a selection operation on a target candidate object among the multiple candidate objects, the target object in the first target media content is replaced with the target candidate object; wherein the target candidate object is used to update and present the first target media content.

7. The method according to claim 1, wherein The method further comprises: In response to a work publishing operation for the first target media content, a multimedia work is generated and published based on the first target media content.

8. The method according to claim 1, wherein Before generating the first target media content based on the pause frame image, the method further includes: Acquire a target digital model; wherein the target digital model is constructed based on pre-input modeling parameter information; Accordingly, generating the first target media content based on the pause frame image includes: First target media content is generated based on the pause frame image and the target digital model.

9. The method according to claim 1, wherein: The method further comprises: In response to a comment triggering operation on the target video, a comment panel is displayed, and recommended comment content is displayed for the comment input box on the comment panel; wherein the recommended comment content includes the first target media content.

10. The method according to claim 1, wherein The method further comprises: In response to a comment trigger operation for the target video, a comment panel is displayed; wherein, the comment panel displays first target comment content including fourth target media content, and the first target comment content is provided with a first target control, which is used to trigger the generation of fifth target media content based on the fourth target media content.

11. The method according to claim 1, wherein The method further comprises: In response to a display trigger operation for the sixth target media content on the comment panel, the sixth target media content is displayed based on the image comment information flow corresponding to the target video; wherein, the sixth target media content belongs to the second target comment content for the target video, and a second target control is set on the display page of the sixth target media content, and the second target control is used to trigger the generation of the seventh target media content based on the sixth target media content.

12. A media content processing device, wherein: The device comprises: A first pause module, configured to pause the target video in response to a pause operation performed on a play page of the target video, and obtain a pause frame corresponding to the target video; The first generation module is used to generate first target media content based on the pause frame image, and display the first target media content on the playback page of the target video; wherein, the first target media content is generated based on part or all of the content on the pause frame image using a preset raw image model.

13. A computer-readable storage medium, wherein: The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device implements the method according to any one of claims 1 to 11.

14. A media content processing device, wherein: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Video screen editing method and device

    CN104023272A

  • Image processing method and device, computer equipment and storage medium

    CN114285961A

  • Video processing method and device, terminal and storage medium

    CN116257165A

  • System and method for providing and interacting with coordinated presentations

    WO2016029209A1