Media content processing method and device, and storage medium and program product

By recognizing users' editing intentions through human-computer dialogue, intelligent editing of media content is achieved, solving the problem that users need to be familiar with the location of functional controls and the complexity of operations, thus improving editing efficiency and intelligence.

WO2025036409A9PCT designated stage expired Publication Date: 2026-01-29BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/112089
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-14
Filing Date
2024-08-14
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

In media content editing software, users need to familiarize themselves with the location and operation of functional controls, resulting in high learning costs, low levels of intelligence, and complex operation.

Method used

The system uses a human-computer dialogue approach to determine the user's editing intent, and then edits the media content to be processed based on the dialogue content. This includes displaying the dialogue content and editing track on the interface, previewing the editing effect in real time, and using natural language processing to identify the editing intent and perform editing operations.

Benefits of technology

It reduces the learning cost for users, improves the intelligence level of media content generation, and is simple and convenient to operate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024112089_29012026_PF_FP_ABST
    Figure CN2024112089_29012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a media content processing method and device, and a storage medium and a program product. The media content processing method comprises: displaying dialogue content in a dialog window of a first interface, wherein the dialogue content is used for indicating an editing intent of a user for media content to be processed, and the first interface comprises an editing track, an editing effect preview area and the dialog window, the editing track being configured to present a video clip formed on the basis of said media content, and the editing effect preview area being configured to play the video clip formed on the basis of said media content; and on the basis of the dialogue content, editing said media content, so as to obtain target media content, wherein the target media content comprises said media content and editing information. The embodiments of the present disclosure reduce the learning cost of a target user, have simple operations, and increase the level of intelligence of media content generation.
Need to check novelty before this filing date? Find Prior Art

Description

Media content processing methods, equipment, storage media and software products

[0001] This application claims priority to Chinese Patent Application No. 202311024778.1, filed on August 14, 2023, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0002] Embodiments of this disclosure relate to a media content processing method, apparatus, storage medium, and program product. Background Technology

[0003] In media content editing software, functional controls are provided in fixed locations on the editing page for editing media content. Users need to be familiar with the location of these functional controls and master their operation methods during the editing process. The operation is complicated, the learning cost is high, the level of intelligence is low, and it is not conducive to user operation.

[0004] Summary of the Invention

[0005] This disclosure provides a media content processing method, apparatus, storage medium, and program product.

[0006] In a first aspect, embodiments of this disclosure provide a media content processing method, including:

[0007] The dialog window of the first interface displays the dialog content; the dialog content is used to indicate the user's editing intention for the media content to be processed; the first interface includes an editing track, an editing effect preview area and the dialog window, the editing track is used to present video clips formed based on the media content to be processed, and the editing effect preview area is used to play the video clips formed based on the media content to be processed;

[0008] Based on the dialogue content, the media content to be processed is edited to obtain target media content; the target media content includes the media content to be processed and editing information; the editing information is used to indicate the editing operation to be performed on the media content to be processed;

[0009] Video clips based on the target media content are presented on the editing track.

[0010] Secondly, embodiments of this disclosure provide a media content processing device, including:

[0011] The display module is configured to display dialogue content in a dialog window of the first interface; the dialogue content is used to indicate the user's editing intention for the media content to be processed; the first interface includes an editing track, an editing effect preview area and the dialog window, the editing track is used to present video clips formed based on the media content to be processed, and the editing effect preview area is used to play the video clips formed based on the media content to be processed;

[0012] An editing module is configured to edit the media content to be processed based on the dialogue content to obtain target media content; the target media content includes the media content to be processed and editing information; the editing information is used to indicate the editing operation to be performed on the media content to be processed.

[0013] The display module is also configured to present video clips based on the target media content on the editing track.

[0014] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;

[0015] The memory stores computer-executed instructions;

[0016] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the media content processing method as described in the first aspect and various possible designs of the first aspect.

[0017] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the media content processing method described in the first aspect and various possible designs of the first aspect.

[0018] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the media content processing method described in the first aspect and various possible designs of the first aspect. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 is a schematic flowchart of the media content processing method provided in an embodiment of this disclosure;

[0021] Figure 2 is an interactive schematic diagram of the media content processing method provided in the embodiments of this disclosure;

[0022] Figures 3a to 3d are schematic diagrams of the interface for the dialog content of the music editing process provided in the embodiments of this disclosure;

[0023] Figures 4a to 4d are schematic diagrams of the interface for background editing processing provided in the embodiments of this disclosure;

[0024] Figure 5 is a structural block diagram of the media content processing device provided in an embodiment of this disclosure; and

[0025] Figure 6 is a schematic diagram of the hardware structure of the media content processing device provided in an embodiment of this disclosure. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0027] To address the issues of high operational difficulty and low intelligence level in the aforementioned media content editing methods, the inventors of this disclosure have discovered that a human-computer dialogue method can be used to determine the user's editing intent, and then the media content to be processed can be edited based on this intent to obtain the target media content. Based on this, embodiments of this disclosure provide a media content processing method.

[0028] The media content processing method provided in this disclosure can be implemented through electronic devices, or applications (APPs), web pages, etc., within electronic devices. Electronic devices may include mobile phones, tablets, wearable electronic devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), smart TVs, smart screens, high-definition TVs, 4K TVs, smart speakers, smart projectors, and other smart home devices. This disclosure does not limit the specific type of electronic device. This disclosure also does not limit the type of operating system used by the electronic device; for example, Android, Linux, Windows, iOS, etc.

[0029] Referring to Figure 1, which is a schematic flowchart of a media content processing method provided in an embodiment of this disclosure, the media content processing method includes:

[0030] 101. Display the dialogue content in the dialog window of the first interface; the dialogue content is used to indicate the user's editing intention for the media content to be processed; the first interface includes an editing track, an editing effect preview area and a dialog window, the editing track is used to present video clips formed based on the media content to be processed, and the editing effect preview area is used to play video clips formed based on the media content to be processed.

[0031] In the specific implementation process, the media content to be processed can be imported into the editing track of the first interface, and then the initial editing effect of the media content to be processed can be viewed in the editing preview area. In the subsequent editing process of the media content to be processed based on the dialogue content, the edited media content to be processed can be displayed on the editing track in real time, so that the editing effect of the edited media content to be processed can be viewed in real time through the editing preview area.

[0032] For example, after changing the background of the media content to be processed based on the dialogue content, the user can view the editing effect of the media content to be processed after changing the background in the editing effect preview area. Then, after changing the music, the user can view the editing effect of the media content to be processed after changing the music in real time. This makes it easier for the user to input dialogue information for the next editing operation based on the editing effect presented in the editing effect preview area.

[0033] The media content to be processed can be original materials (such as images, video materials, etc.). Correspondingly, the video clip formed based on the media content to be processed can contain one or more frames of images, can refer to video, or can be a video draft generated in advance based on the original materials (such as images, video materials, etc.). The video draft includes the original materials and editing information, which is used to indicate the editing operations to be performed on the original materials. In one embodiment of this disclosure, the content to be processed is video material; before displaying the dialogue content in the dialog window of the first interface, the method further includes: acquiring video material; presenting a video clip formed based on the video material on the editing track; and editing the media content to be processed according to the dialogue content to obtain the target media content, including: editing the video material according to the dialogue content to obtain a video draft; the video draft includes video material and editing information; the editing information is used to indicate the editing operations to be performed on the video material.

[0034] In addition, the dialogue content can include user dialogue information from users, as well as system dialogue information from the system. Dialogue information can include data of various types such as text, numbers, symbols, links, emoticons, images, audio, and video clips, without limitation; there are no limitations on display parameters such as the size, color, display position, and transparency of the editing effect preview area, editing track, and dialogue window.

[0035] There are several ways for users to input dialogue information.

[0036] In one possible implementation, to facilitate guidance for users with limited editing experience, operation controls can be displayed on the first interface, allowing users to express their editing intentions by triggering these controls. Specifically, the dialogue content may include first user dialogue information, and the media content processing method may further include: displaying a target function control on the first interface; and generating first user dialogue information corresponding to the target function control in response to a touch operation applied to the target function control.

[0037] The target function controls can be controls for changing aspect ratios, changing music, intelligently adding text, replacing materials, adjusting filters, adding special effects, etc. Target function controls can be displayed in the form of icons, text, symbols, images, or a combination of these forms; this disclosure does not impose any limitations on this.

[0038] In another possible implementation, to improve the flexibility of user input of user dialogue information, an input area can be set in the first interface, allowing the user to input user dialogue information within the input area. Specifically, the dialogue content may include second user dialogue information, the first interface may include a dialogue input area, and the media content processing method may further include: responding to an input operation applied to the dialogue input area and acquiring the second user dialogue information. In this embodiment, the display parameters such as the size, color, display position, and display transparency of the input area are not limited.

[0039] In another possible implementation, to broaden the user's editing ideas, editing suggestions can be presented in a dialog window on the first interface, allowing the user to input their own dialogue based on these suggestions. These suggestions can be determined based on the media content to be processed; for example, if the media content includes background music, suggestions to change the music could be provided.

[0040] It should be noted that the above-mentioned various methods of inputting user dialogue information can be implemented individually or in any combination, and this disclosure does not limit this.

[0041] Dialogue information from different parties in the dialogue content can be set on the same side or on different sides. For example, user dialogue information from the user can be set on the right side of the dialogue window and right-aligned, while system dialogue information from the system can be set on the left side of the dialogue window and left-aligned. Alternatively, the dialogue information from both parties can be left-aligned or right-aligned. This embodiment of the disclosure does not limit this.

[0042] 102. Based on the dialogue content, edit the media content to be processed to obtain the target media content; the target media content includes the media content to be processed and the editing information; the editing information is used to instruct the editing operations to be performed on the media content to be processed.

[0043] Specifically, after obtaining the dialogue content, the editing intent of the dialogue content can be identified and processed, and the media content to be processed can be edited based on the editing intent to obtain the target media content.

[0044] For example, the above-mentioned identification process may involve performing natural language processing on the dialogue content to obtain multiple tags representing creative intent. Based on these tags, corresponding operation controls (such as changing the background, changing the music, adding special effects, etc.) and parameters to be edited (such as background style, color, composition, etc.; music author, lyrics, album art, etc.; candidate styles of special effects, position of adding special effects, etc.) are determined. Then, based on the operation controls, the parameters are adjusted to obtain the target media content. The target media content can be data of types such as video, text, and images.

[0045] For example, a user can input the dialogue message "Help me change the background to blue sky and white clouds". The system can perform natural language processing on this dialogue message to determine the user's editing intent. Then, it can call two relevant components, such as a keying component and a background replacement component. The keying component is used to cut out objects such as people and scenery from the content to be edited, excluding the background. The background replacement component is used to provide the user with a card (control) consisting of multiple candidate backgrounds to be replaced in the dialogue window. Then, it receives the user's selection operation of the candidate background, and then merges the selected candidate background with the keying result to realize the editing intent of replacing the blue sky and white clouds background and obtain the target media content.

[0046] When there are multiple components (cards), there are several ways to determine the execution order of the components. For example, the order of different cards can be set in advance and then executed one by one. Specifically, parameters can be set for each card. During the execution of a card, the parameter is the first value. After the card is finished executing, the parameter is the second value. At this time, the component cannot execute any logic or output any renderable content.

[0047] 103. Present video clips based on target media content on the editing track.

[0048] Specifically, after generating the target media content, in order to facilitate viewing the editing effect of the target media content, video clips formed based on the target media content can be displayed on the editing track.

[0049] As can be seen from the above description, the media content processing method provided in this disclosure realizes information interaction between users and the system by adopting a dialogue method. Then, based on the obtained dialogue content representing the user's editing intention, the content to be processed is edited to obtain the target content. This reduces the learning cost for the target user, is simple to operate, and improves the level of intelligence in media content generation.

[0050] In one embodiment of this disclosure, based on the embodiment of FIG1 above, the media content processing method further includes: presenting the media content to be processed on the editing track; obtaining the current position of the preview pointer on the editing track; using the preview pointer to move along the editing track; determining the editing position of the media content to be processed according to the current position; and editing the media content to be processed according to the dialogue content, including: editing the material corresponding to the editing position in the media content to be processed according to the dialogue content.

[0051] Specifically, the application's creation page can display multiple video draft covers. Clicking a cover leads to the first interface, where the corresponding video draft is imported into the editing track. Users can then preview the edited video clips created from the drafts in the editing effect preview area. The preview pointer on the editing track can be moved to position the image frame corresponding to that pointer's location. This allows for human-computer interaction through a dialog window, enabling editing of that image frame.

[0052] In one embodiment of this disclosure, based on the embodiment of FIG1 above, editing the media content to be processed according to the dialogue content may include: determining the editing intent according to the dialogue content; determining the editing tools and editing parameters according to the editing intent; and editing the media content to be processed using the editing tools based on the editing parameters.

[0053] Specifically, natural language processing can be used to analyze the dialogue content, identify the user's editing intent, and determine the corresponding editing tools and parameters based on this intent. Then, these parameters are adjusted using the control panel to obtain the target media content. The target media content can be data types such as video, text, and images.

[0054] In this context, "editing tool" refers to a UI component, such as a React component, which is a UI part encapsulated with independent functionality. Examples include editing controls for changing backgrounds, music, and adding effects. "Editing parameters" refers to the editing parameters used by the editing tool during the editing process, such as background style, color, and composition; music author, lyrics, and album art; and effect candidate styles and effect placement.

[0055] In one embodiment of this disclosure, the dialogue content may include user dialogue information and system dialogue information. The source of the editing intent is twofold: firstly, the intent recognition of the user dialogue information, and secondly, the parsing of the user's touch operations on the operation controls provided in the system dialogue information. Specifically, determining the editing intent based on the dialogue content may include: determining a first editing intent based on the fourth user dialogue information in the dialogue content; determining the operation controls and their corresponding editing parameters based on the first editing intent; and correspondingly, editing the media content to be processed using an editing tool based on the editing parameters, which may include: displaying the operation controls in the dialogue window of the first interface; adjusting the editing parameters in response to touch operations applied to the operation controls to obtain target editing parameters; and editing the media content to be processed using an editing tool based on the target editing parameters.

[0056] In one embodiment of this disclosure, to save computing resources of the terminal device, the intent recognition step can be performed by the server. Specifically, determining the first editing intent based on the fourth user dialogue information in the dialogue content can include: sending the fourth user dialogue information to the server so that the server performs natural language processing on the fourth user dialogue information to obtain the corresponding editing intent; assembling and generating a response script based on the editing intent; and sending the response script to the client; the response script includes functions and corresponding parameter keywords; determining the operation controls and corresponding editing parameters based on the first editing intent can include: receiving the response script; determining the operation controls corresponding to the functions based on the response script; and determining the corresponding editing parameters based on the parameter keywords.

[0057] Specifically, the parameter keywords are the basic instruction keywords. Obtaining these parameters requires the client to request the API, assemble, and render them manually. For example, with music cards, the response script will only pass the keyword parameter (`keyword`), not information such as the tag icon, theme title, or music ID. The client needs to request the API, assemble, and render these parameters based on the instruction keywords.

[0058] For example, as shown in Figure 2, taking a video draft as the content to be edited, the user enters a statement (user dialogue information) in the dialog window of the client's first interface. Based on this input statement, the client sends an UpsertSessionReq to the server. After receiving the request, the server creates a session context, performs Natural Language Processing (NLP) on the input statement carried in the UpsertSessionReq, identifies the editing intent, determines the editing tool (UI component) based on the editing intent, and then assembles a response script (AnswerScript). Based on this response script, the server sends an UpsertSessionRsp to the client, which includes information about the response script. The client interprets the response script, injects editing parameters (Params), and then executes the script. During the execution of the script, a card of a component contained in the script is displayed in the dialog window, receiving the user's touch operation (click or long press). In response to the touch operation, the editing parameters are adjusted, the video draft is edited based on the adjusted editing parameters, and it is determined whether the preset conditions are met. If they are met (e.g., the previous card has been executed), the next component card is executed. After multiple sessions and once the editing effect meets the user's needs, the user can trigger a close operation on the client's first interface to close the draft editor, i.e., the first interface. Based on this close command, the client sends a DestroySessionReq request to the server.

[0059] In one embodiment of this disclosure, to better guide the user in editing operations, after determining the editing intent based on user dialogue information, the operation controls for the current editing operation and the recommended information for the next editing operation can be determined simultaneously based on the editing intent. Specifically, after determining the first editing intent based on the user dialogue information in the dialogue content, the process may further include: determining the recommended information for the next editing operation based on the first editing intent; displaying the recommended information for the next editing operation in the dialogue window of the first interface, so that the user can input fifth user dialogue information based on the recommended information for the next editing operation; the fifth user dialogue information is used to indicate the second editing intent in the editing intent.

[0060] Specifically, the dialog window on the first interface can display multiple user-inputted dialog messages, each of which can identify an editing intent. These editing intents collectively constitute the editing intent corresponding to the user dialog message. Of course, if the dialog content also includes system dialog information, the system dialog information can include the operation controls and editing parameters determined based on the editing intent based on the user dialog information. In this case, touch operations on the operation controls of the system dialog information can obtain the user editing intent determined based on the system dialog information.

[0061] The following examples, with reference to Figures 3a-3d, illustrate the process of editing music replacement.

[0062] After the application (media content editing application) is launched, the user interface 31 shown in Figure 3a can be displayed on the mobile phone. The user interface 31 is a conversational editing page (i.e., the first interface). The conversational editing page is mainly used to present and obtain the conversation content between the user and the system, and to edit and obtain the target media content based on the conversation content.

[0063] Referring to Figure 3a, the user interface 31 includes a dialog window 301, a user preview area 302, an editing track 303, and a dialog input area 304. The user preview area 302 displays the video editing frame of the media content to be processed, indicated by the preview pointer. The editing track 303 includes a preview pointer used to indicate the editing position. The dialog window 301 displays operation suggestions for the media content to be processed, such as "Change Music," "Add Packaging," "Smart Add Text," and "Replace Material." The dialog input area 304 is used to input user dialogue information. Users can input user dialogue information in the dialog input area 304 based on the suggestions to indicate their editing intentions. For example, suppose the user wants to perform the "Change Music" editing operation.

[0064] Referring to Figure 3b, the user dialogue information "Change music" entered by the user is displayed in the dialogue input area 304 of the user interface 32. The system obtains the user dialogue information and performs intent recognition on the user dialogue information. Based on the recognized editing intent, the system determines the operation controls and editing parameters of the current operation so that the operation controls can be displayed in the dialogue window. The editing parameters can be adjusted by the user's touch operation on the operation controls.

[0065] Referring again to Figure 3c, the dialog input area 304 of the user interface 33 displays the user dialog message "Change music" based on the user's input, as well as system dialog information. This system dialog information includes operation controls for the current operation determined based on the user's editing intent, as shown in Figure 3c. The operation controls include multiple candidate music tracks (original, music 1, music 2, music 3) and other editing parameters for the user to select from, as well as adjustment controls and their corresponding editing parameters. Furthermore, below the system dialog information, suggested information for the next operation is displayed: "Upbeat music" and "Recommended popular music." This suggested information can be determined by the system based on the user's editing intent of "Change music." In other words, the system determines both the operation controls for the current operation and the recommended information for the next operation based on the editing intent. This makes it easier for users to perform subsequent editing processes efficiently.

[0066] Referring again to Figure 3d, the user's touch operation (click or long press) on the adjustment control is received. In response to the touch operation, further operation controls are displayed in the dialog input area 304 of the user interface 32, including cropping and volume control for selecting segments and adjusting the volume.

[0067] Figures 3b to 3d also include a user preview area 302, an editing track 303, and a dialog input area 304. The description of the areas can be found in the description of Figure 3a, and will not be repeated here.

[0068] It should be noted that the embodiments shown in Figures 3a to 3d are merely examples of the media content processing method provided in this disclosure. In actual scenarios, the display method, presentation format, related control settings, and the layout, display parameters, and implementation methods of some pages and controls and entry points can be flexibly set according to requirements.

[0069] The following examples, with reference to Figures 4a-4d, illustrate the background changing editing process.

[0070] After the application (media content editing application) is launched, the user interface 41 shown in Figure 4a can be displayed on the mobile phone. The user interface 41 is a conversational editing page (i.e., the first interface). The conversational editing page is mainly used to present and obtain the conversation content between the user and the system, and to edit and obtain the target media content based on the conversation content.

[0071] Referring to Figure 4a, the user interface 41 includes a dialog window 301, a user preview area 302, an editing track 304, and a dialog input area 304. The user preview area 302 displays the video editing frame of the media content to be processed, indicated by the preview pointer. The editing track 304 includes a preview pointer used to indicate the editing position. The dialog window 301 displays operation suggestions for the media content to be processed, such as "Change scale," "Add packaging," "Smart add text," and "Replace material." The dialog input area 304 is used to input user dialogue information. Users can input user dialogue information in the dialog input area 304 based on the suggestions to indicate their editing intentions. As shown in Figure 4a, assuming the user wants to perform "Change scale" editing, the dialog input area 304 of the user interface 41 displays the user dialogue information "Change scale."

[0072] Referring to Figure 4b, the system acquires the user's dialogue information "change aspect ratio" and performs intent recognition on it. Based on the identified editing intent, it determines the operation controls and editing parameters corresponding to the current operation of "change aspect ratio," displaying the operation controls in the dialogue window. The editing parameters can be adjusted by the user's touch operation on the operation controls. As shown in Figure 4b, the dialogue input area 304 of the user interface 42 displays the system dialogue information based on the user's input "change aspect ratio," including the operation controls for the current operation determined based on the user's editing intent. The operation controls include multiple candidate aspect ratios (original, 9:16, 16:9, 1:1, 4:3) for the user to select. Upon receiving the user's touch operation on 9:16, and combining Figures 4a and 4b, the image frame in the editing effect preview area switches from the original aspect ratio to a 9:16 aspect ratio.

[0073] Furthermore, below the system dialogue information, suggested information for the next step is displayed, such as "Smart Packaging," "Adjust Filters," and "Add Effects." This suggestion information is determined by the system based on the user's editing intent, such as "Change Scale," meaning that the system determines both the current operation controls and the recommended information for the next step based on the editing intent. This makes it easier for users to perform subsequent editing processes more efficiently.

[0074] Referring to Figure 4c, in order to facilitate users to edit image frames, an editing function can be provided in the editing effect preview area. As shown in Figure 4c, the user's touch operation (click or long press) on the image area in the editing effect preview area 302 of the user interface 43 can be received. In response to the touch operation, the image area can be switched to the selected state.

[0075] Referring again to Figure 4d, the system receives user touch operations (clicks or drags) on the selected box. In response to these touch operations, the image within the selected box is magnified based on its magnification factor. For example, the magnification factor can be vaguely determined based on the magnification factor of the selected box. For instance, in one case, if the magnification factor of the selected box is greater than 1, the image is magnified by a preset factor, such as 1.1. In another case, the magnification factor of the selected box is divided into multiple intervals, with different intervals corresponding to different image magnification factors. For example, magnification factors of 1-1.2 and 1.2-1.3 correspond to image magnification factors of 1.1 and 1.2, respectively.

[0076] Figures 4b to 4d also include a user preview area 302, an editing track 304, and a dialog input area 304. The description of the areas can be found in the description of Figure 4a, and will not be repeated here.

[0077] It should be noted that the embodiments shown in Figures 4a to 4d are merely examples of the media content processing method provided in this disclosure. In actual scenarios, the display method, presentation format, related control settings, and the layout, display parameters, and implementation methods of some pages and controls and entry points can be flexibly set according to requirements.

[0078] Corresponding to the media content processing method in the above embodiments, FIG5 is a structural block diagram of a media content processing device provided in an embodiment of this disclosure. For ease of explanation, only the parts related to the embodiments of this disclosure are shown. Referring to FIG5, the media content processing device 50 includes: a display module 501 and an editing module 502.

[0079] The display module 501 is configured to display dialogue content in a dialog window of the first interface; the dialogue content is used to indicate the user's editing intentions for the media content to be processed; the first interface includes an editing track, an editing effect preview area, and a dialog window, the editing track is used to present video clips formed based on the media content to be processed, and the editing effect preview area is used to play video clips formed based on the media content to be processed;

[0080] The generation module 502 is configured to edit the media content to be processed based on the dialogue content to obtain the target media content; the target media content includes the media content to be processed and editing information; the editing information is used to indicate the editing operations to be performed on the media content to be processed;

[0081] Display module 501 is also configured to present video clips formed based on target media content on the editing track.

[0082] In one embodiment of this disclosure, the display module 501 is further configured to: display a target function control in a first interface; and generate first user dialogue information corresponding to the target function control in response to a touch operation applied to the target function control.

[0083] In one embodiment of this disclosure, the dialogue content includes second user dialogue information, the first interface includes a dialogue input area, and the display module 501 is further configured to: respond to an input operation applied to the dialogue input area and obtain the second user dialogue information.

[0084] In one embodiment of this disclosure, the dialogue content includes third user dialogue information, and the display module 501 is further configured to: display at least one editing operation recommendation information in a first interface; the editing operation recommendation information is determined based on the media content to be processed; and in response to an input operation triggered by at least one prompt information acting on the dialogue input area, obtain the third user dialogue information.

[0085] In one embodiment of this disclosure, the generation module 502 is specifically configured to: determine the editing intent based on the dialogue content; determine the editing tools and editing parameters based on the editing intent; and edit the media content to be processed using the editing tools based on the editing parameters.

[0086] In one embodiment of this disclosure, the generation module 502 is specifically configured to: determine a first editing intent based on the fourth user dialogue information in the dialogue content; determine an operation control and the corresponding editing parameters based on the first editing intent; display the operation control in the dialogue window of the first interface; adjust the editing parameters in response to a touch operation applied to the operation control to obtain target editing parameters; and edit the media content to be processed using an editing tool based on the target editing parameters.

[0087] In one embodiment of this disclosure, the generation module 502 is specifically configured to: send the fourth user dialogue information to the server so that the server performs natural language processing on the fourth user dialogue information to obtain the corresponding editing intent, assemble and generate a response script according to the editing intent, and send the response script to the client; the response script includes functions and corresponding parameter keywords; receive the response script; determine the operation controls corresponding to the functions according to the response script; and determine the corresponding editing parameters according to the parameter keywords.

[0088] In one embodiment of this disclosure, the generation module 502 is further configured to: determine recommended information for the next editing operation based on the first editing intent; display the recommended information for the next editing operation in the dialog window of the first interface, so that the user can input fifth user dialog information based on the recommended information for the next editing operation; the fifth user dialog information is used to indicate the second editing intent in the editing intent.

[0089] In one embodiment of this disclosure, the generation module 502 is further configured to: present the media content to be processed on the editing track; obtain the current position of the preview pointer on the editing track; use the preview pointer to move along the editing track; determine the editing position of the media content to be processed based on the current position; and edit the material corresponding to the editing position in the media content to be processed based on the dialogue content.

[0090] In one embodiment of this disclosure, the generation module 502 is further configured to: acquire video footage; present a video clip formed based on the video footage on an editing track; and edit the video footage according to the dialogue content to obtain a video draft; the video draft includes video footage and editing information; the editing information is used to indicate the editing operations performed on the video footage.

[0091] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0092] To implement the above embodiments, this disclosure also provides an electronic device.

[0093] Referring to Figure 6, a schematic diagram of the structure of an electronic device 900 suitable for implementing embodiments of the present disclosure is shown. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 6 is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present disclosure.

[0094] As shown in Figure 6, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0095] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 shows electronic device 900 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0096] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.

[0097] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0098] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0099] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0100] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0102] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0103] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0104] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0105] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0106] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0107] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A media content processing method, comprising: displaying conversation content in a conversation window of a first interface; wherein the conversation content is used to indicate a user's editing intention for to-be-processed media content; the first interface comprises an editing track, an editing effect preview area, and the conversation window, the editing track is used to present a video clip formed based on the to-be-processed media content, and the editing effect preview area is used to play the video clip formed based on the to-be-processed media content; performing editing processing on the to-be-processed media content according to the conversation content to obtain target media content; wherein the target media content comprises to-be-processed media content and editing information; the editing information is used to indicate an editing operation performed on the to-be-processed media content; presenting a video clip formed based on the target media content on the editing track.

2. The method of claim 1, wherein, The conversation content comprises first user conversation information, and the method further comprises: displaying a target function control in the first interface; generating the first user conversation information corresponding to the target function control in response to a touch operation acting on the target function control.

3. The method of claim 1, wherein, The conversation content comprises second user conversation information, and the first interface comprises a conversation input area, and the method further comprises: obtaining the second user conversation information in response to an input operation acting on the conversation input area.

4. The method of claim 3, wherein, The conversation content comprises third user conversation information, and the method further comprises: displaying at least one editing operation recommendation information in the first interface; wherein the editing operation recommendation information is determined according to the to-be-processed media content; obtaining the third user conversation information in response to an input operation acting on the conversation input area and triggered according to the at least one prompt information.

5. The method according to any one of claims 1 to 4, wherein, The editing processing on the to-be-processed media content according to the conversation content comprises: determining an editing intention according to the conversation content; determining an editing tool and an editing parameter according to the editing intention; performing editing processing on the to-be-processed media content by the editing tool based on the editing parameter.

6. The method of claim 5, wherein, The editing intention comprises a first editing intention, and the determining of the editing intention according to the conversation content comprises: determining the first editing intention in the editing intention according to fourth user conversation information in the conversation content; determining an operation control and an editing parameter corresponding to the operation control according to the first editing intention; The performing of the editing processing on the to-be-processed media content by the editing tool based on the editing parameter comprises: displaying the operation control in the conversation window of the first interface; adjusting the editing parameter in response to a touch operation acting on the operation control to obtain a target editing parameter, and performing editing processing on the to-be-processed media content by the editing tool based on the target editing parameter.

7. The method of claim 6, wherein, The determining of the first editing intention in the editing intention according to the fourth user conversation information in the conversation content comprises: send the fourth user conversation information to a server, so that the server performs natural language processing on the fourth user conversation information to obtain a corresponding editing intention, assemble and generate a reply script according to the editing intention, and send the reply script to a client; wherein the reply script includes a function and a parameter keyword corresponding to the function; the operation control and the editing parameter corresponding to the operation control are determined according to the first editing intention, including: receiving the reply script; determining the operation control corresponding to the function according to the reply script; determining the corresponding editing parameter according to the parameter keyword.

8. The method of claim 6, wherein, The editing intention further includes a second editing intention, and after determining the first editing intention in the editing intention according to the user conversation information in the conversation content, further including: determining recommendation information of a next editing operation according to the first editing intention; displaying the recommendation information of the next editing operation in the conversation window of the first interface, so that the user inputs fifth user conversation information according to the recommendation information of the next editing operation; wherein the fifth user conversation information is used to indicate the second editing intention in the editing intention.

9. The method of any one of claims 1-4, further comprising: presenting the to-be-processed media content on the editing track; obtaining a current position of a preview pointer on the editing track; wherein the preview pointer is used to move along the editing track; determining an editing position of the to-be-processed media content according to the current position; the editing processing of the to-be-processed media content according to the conversation content, including: editing processing of a material corresponding to the editing position in the to-be-processed media content according to the conversation content.

10. The method according to any one of claims 1-4, wherein, The to-be-processed content is a video material; before displaying the conversation content in the conversation window of the first interface, further including: obtaining a video material; presenting a video clip formed based on the video material on the editing track; the editing processing of the to-be-processed media content according to the conversation content to obtain a target media content, including: editing processing of the video material according to the conversation content to obtain a video draft; wherein the video draft includes a video material and editing information; the editing information is used to indicate an editing operation performed on the video material.

11. A video processing device, comprising: a display module configured to display conversation content in a conversation window of a first interface; wherein the conversation content is used to indicate an editing intention of a user on to-be-processed media content; the first interface includes an editing track, an editing effect preview area, and the conversation window, the editing track is used to present a video clip formed based on the to-be-processed media content, and the editing effect preview area is used to play a video clip formed based on the to-be-processed media content; An editing module configured to edit the to-be-processed media content according to the dialogue content, to obtain target media content; wherein the target media content comprises the to-be-processed media content and editing information; the editing information is used to indicate an editing operation performed on the to-be-processed media content; and The display module is further configured to present a video clip formed based on the target media content on the editing track.

12. An electronic device, comprising: Comprise: A processor and a memory; Wherein the memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the media content processing method as claimed in any one of claims 1 to 9.

13. A computer readable storage medium, wherein, The computer readable storage medium stores computer execution instructions, and when the processor executes the computer execution instructions, the media content processing method as claimed in any one of claims 1 to 9 is realized.

14. A computer program product comprising a computer program, wherein, The computer program is executed by the processor to realize the media content processing method as claimed in any one of claims 1 to 9.