Interaction method and device, equipment and storage medium
By receiving natural language editing requests in social applications and presenting optional editing modes, users solve the problem of cumbersome operations when editing media content, and achieve efficient and personalized media content editing.
Patent Information
- Application Number
- CN202311714195.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-13
AI Technical Summary
When users edit published media content in social applications, they need to manually select editing controls and constantly try, resulting in wasting time and energy, and it is difficult to select satisfactory materials.
An interactive method is provided to present at least one optional editing mode by receiving a media content editing request in natural language form at a first interface of a user device, and to apply the editing effect after the editing mode. If the user confirms the target editing mode, the edited media content after the target editing effect is presented in the second interface.
It significantly improves the editing efficiency of media content, simplifies the editing process, provides a more personalized and in line with user intentions, and improves user satisfaction.
Smart Images

Figure CN120144007A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, devices, equipment, and computer-readable storage media for interaction. Background Art
[0002] With the development of social applications, more and more users use social applications to post pictures or videos showing their daily lives or browse the content posted by other publishers in the application. Thus, users' personalized editing needs for the media content to be posted are also increasing day by day. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for interaction is provided. The method includes receiving an editing request for media content in a first interface, the editing request including a request in natural language form, the editing request indicating overlaying media materials and / or adding special effects to the media content; based on the editing request, presenting, in the first interface, at least one optional editing mode of the media content and a first editing effect after applying a first optional editing mode among the at least one optional editing modes to the media content; and in response to receiving a confirmation for a target editing mode among the at least one optional editing modes, presenting, in a second interface, the media content edited with a target editing effect corresponding to the target editing mode.
[0004] In a second aspect of the present disclosure, a device for interaction is provided. The device includes a receiving module configured to receive an editing request for media content in a first interface, the editing request including a request in natural language form, the editing request indicating overlaying media materials and / or adding special effects to the media content; a first presenting module configured to present, in the first interface, at least one optional editing mode of the media content and a first editing effect after applying a first optional editing mode among the at least one optional editing modes to the media content based on the editing request; and a second presenting module configured to present, in a second interface, the media content edited with a target editing effect corresponding to the target editing mode in response to receiving a confirmation for a target editing mode among the at least one optional editing modes.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to execute the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, and when the program is executed by a processor, the method of the first aspect is implemented.
[0007] It should be understood that the content described in the present invention content section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0009] Figure 1 A schematic diagram showing an exemplary environment in which the embodiments of the present disclosure can be implemented;
[0010] Figures 2A to 2D A schematic diagram showing an exemplary interface of an interaction process according to some embodiments of the present disclosure;
[0011] Figure 3 A block diagram showing an interaction process according to some embodiments of the present disclosure;
[0012] Figure 4 A flowchart showing a process for interaction according to some embodiments of the present disclosure;
[0013] Figure 5 A block diagram showing a device for interaction according to some embodiments of the present disclosure; and
[0014] Figure 6 A block diagram showing a device capable of implementing multiple embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0016] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.
[0017] It can be understood that the image data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage, or deletion of the data) should comply with the requirements of the corresponding laws, regulations, and related provisions.
[0018] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the data involved in the present disclosure should be informed to the relevant users and the authorization of the relevant users should be obtained through appropriate means according to the relevant laws and regulations, where the relevant users may include any type of right subject, such as an individual, an enterprise, or a group.
[0019] As used herein, the term "model" can learn the corresponding association relationship between the input and the output from the training data, so that after the training is completed, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes the input and provides the corresponding output by using multiple layers of processing units. The neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network", or "learning network", and these terms can be used interchangeably in this article.
[0020] As described above, during the use of a social application, a user may expect to timely learn about the content published by the publishers with a high social intimacy with the user in the social application. Especially for some content with a predetermined timeliness, because such content can often only be accessed within a limited time period after it is published.
[0021] When a user edits the media content to be published, the user often needs to manually select editing controls and keep trying, and then add materials, special effects, etc. to the media content to be published. For example, when a user wants to add background music to the media content to be published, the user needs to enter the music library and try them one by one randomly or in the order of the music list in the music library to determine a satisfactory background music. This process takes a lot of operation time and energy of the user, and the user may not necessarily select satisfactory materials.
[0022] Currently, attempts have been made to apply machine learning models to analyze and understand user requests in order to perform corresponding operations. Therefore, it is desirable to achieve personalized editing of media content based on media content editing requests in the form of natural language from users.
[0023] The technical solution of the present disclosure provides an interaction solution. According to the interaction solution of the present disclosure, an editing request for media content can be received on a first interface of a user device. The editing request is a request in the form of natural language and the editing request indicates adding media materials and / or adding special effects to the media content. In this first interface, while presenting at least one optional editing mode of the media content based on the editing request, a first editing effect after applying a first optional editing mode among the at least one optional editing modes to the media content is presented. If a confirmation for a target editing mode among the at least one optional editing modes is received, the media content edited with a target editing effect corresponding to the target editing mode is presented on a second interface.
[0024] In this way, it is possible to present recommendations for at least one optional editing mode based on the parsing of an editing request in the form of natural language, thereby significantly improving the editing efficiency of media content and simplifying the editing process. At the same time, it is possible to provide a more personalized and user-intention-compliant solution in media content editing, improving user satisfaction.
[0025] Example environment
[0026] First, refer to Figure 1 , which schematically shows a schematic diagram of an example environment 100 in which an exemplary implementation manner according to the present disclosure can be implemented.
[0027] Figure 1 A schematic diagram of an example environment 100 in which an embodiment of the present disclosure can be implemented is shown. In this example environment 100, a target application 120 is installed in an electronic device 110. A user 102 can interact with the target application 120 via the electronic device 110 and / or an attached device of the electronic device 110. The target application 120 can be a content sharing application, capable of providing services related to media content consumption to the user 102, including browsing, commenting, forwarding, creating (e.g., shooting and / or editing), publishing, etc. of media content. Herein, "media content" can be content in various forms, including videos, audios, images, image sets, texts, etc.
[0028] In Figure 1In an environment 100, if the target application 120 is active, the electronic device 110 can present a page 140 of the target application 120 to the user 102. The page 140 can be various types of pages that the target application 120 can provide, such as a page for presenting media content, a content creation page, a content editing page, a message page, a personal page, a shopping mall page, and so on.
[0029] In some embodiments, the target application 120 can include or be implemented as a digital assistant. The digital assistant can be configured to have the ability of intelligent conversation. In some embodiments, the digital assistant can be integrated within the target application 120 and act as a part of the target application 120 to assist in performing task processing within the target application 120. In some other embodiments, the digital assistant can be configured as an independently running application, such as a web application or other types of applications. In such an example, the digital assistant and the target application 120 can be regarded as the same application. The digital assistant is provided to assist users with various task processing requirements in different applications and scenarios. During the interaction with the digital assistant, the user inputs an interaction message, and the digital assistant provides a reply message in response to the user input. Generally, the digital assistant can support the user to input questions in a natural language manner and perform tasks and provide replies based on the understanding of the natural language input and logical reasoning ability.
[0030] In some embodiments, the digital assistant can interact with the user 102 as a contact of the user 102. For example, the digital assistant can be implemented in an instant messaging (IM) application. For example, the digital assistant can interact with the user 102 in a one-on-one chat session with the user 102. In some other embodiments, the digital assistant can also interact with multiple users in a group chat session including multiple users.
[0031] For the user 102, the client of the application running part (such as the electronic device 110 running the target application 120) can present an interaction window of the target application 120 or the digital assistant, such as a session window with the digital assistant, in the client interface (such as the page 140). The user 102 can input a session message in the session window, and the target application 120 can determine the reply message of the digital assistant 122 and present it to the user 102 in the interaction window. In some embodiments, depending on the configuration of the target application 120, the interaction messages with the target application 120 can include messages in multimodal forms, such as text messages (e.g., natural language text), voice messages, image messages, video messages, and so on.
[0032] The application running part can be deployed locally on the terminal device (such as the electronic device 110) of each user 102, and / or can be supported by the server device (such as the server 130). For example, the terminal device of the user 102 can run a client of the application running part, and this client can support the interaction between the user and the application running part provided by the server. When the application running part runs locally on the user's terminal device, the user 102 can directly interact with the local application running part using the terminal device. When the application running part runs on the server device, the server device can, based on the communication connection with the terminal device, realize the service supply to the client running on the terminal device. The application running part can present corresponding application pages to the user 102 based on the operations of the user 102, so as to output and / or receive information related to application usage from the user 102.
[0033] In some embodiments, the implementation of at least some functions of the target application 120, and / or the implementation of at least some functions of the digital assistant in the target application 120 can be based on the model 131. The model 131 can be deployed in the server 130, for example. During the running process of the target application 120, the capabilities of one or more models (such as the model 131) can be invoked. In the target application 120, the digital assistant can use the model 131 to understand the user input and provide a reply to the user 102 based on the output of the model 131.
[0034] In some embodiments, the model 131 can be a machine learning model, a deep learning model, a learning model, a neural network, etc. In some embodiments, the model 131 can be based on a language model (LM). By learning from a large amount of corpus, the language model can possess the ability to answer questions. The model 131 can also be based on other suitable models.
[0035] In some embodiments, the electronic device 110 communicates with the server 130 to implement the supply of services for the target application 120. The electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 is also capable of supporting any type of user interface (such as a "wearable" circuit, etc.). The server 130 is various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, and the like.
[0036] It should be understood that the structure and functions of the environment 100 are described only for exemplary purposes, without implying any limitation on the scope of the present disclosure.
[0037] Interaction process
[0038] Figures 2A to 2D A schematic diagram of an exemplary interface for interaction according to some embodiments of the present disclosure is shown. The following describes the interaction process of the embodiments of the present disclosure in conjunction with Figures 2A to 2D The interaction process of the embodiments of the present disclosure is described. This process can be implemented at least in part at Figure 1 the electronic device 110 shown. For ease of discussion, this process will be described with reference to Figure 1 the environment 100.
[0039] The interface of the electronic device 110 (hereinafter also referred to as the "third interface") can present unedited media content to be published. As Figure 2A shown, the interface 200A presents unedited media content to be published. In some embodiments, the third interface can also provide guidance for obtaining an editing request. For example, if the electronic device 110 detects a trigger operation by the user on the presented control 201 in the interface 200A, an interface (hereinafter also referred to as the "first interface") that can receive an editing request from the user for the media content to be published is presented.
[0040] For example, the control 201 of the interface 200A can be understood as an entry for enabling the user 102 to interact with the target application 120. Once a trigger operation on the control 201 is detected, the electronic device 110 can present an interaction window of the target application 120 in the first interface, such as a session window with a media editing assistant. AsFigure 2B As shown, the interface 200B presents a session window 202 with the media editing assistant. The user 102 can enter an editing request for the media content to be published in this session window 202.
[0041] It should be understood that in addition to the triggering operation on the control 201, in some other embodiments, the user 102 can also cause the electronic device 110 to present the session window 202 with the media editing assistant in other ways, such as by voice triggering and the like.
[0042] For example, if the user 102 enters an editing request "Help me find a material template suitable for spring" for the media content to be published through the session window 202 and sends it, then on the first interface, the media editing assistant's reply to this editing request can be presented. As Figure 2C shown, on the interface 200B, for the editing request 211 "Help me find a material template suitable for spring" entered by the user, the electronic device 110 can present the media editing assistant's reply 212 "Try these templates with spring elements" and selection controls 221 to 224 corresponding to multiple recommended optional editing modes respectively. While presenting the selection controls 221 to 224 corresponding to multiple recommended optional editing modes, on the interface 200B, the editing effect corresponding to applying one of the multiple recommended optional editing modes to the media content to be published can also be presented.
[0043] In some embodiments, the applied optional editing mode can be the first optional editing mode among the multiple recommended optional editing modes. In some other embodiments, the applied optional editing mode can also be any other optional editing mode among the multiple recommended optional editing modes.
[0044] It should be understood that any number of optional editing modes can be recommended based on the user's editing request, such as one or more recommended optional editing modes. The number of the recommended optional editing modes is not limited herein.
[0045] If the user is satisfied with the editing effect corresponding to the optional editing mode currently applied to the media content to be published, then the confirmation for the optional editing mode currently applied to the media content to be published can be triggered. For example, by triggering the control 214 on the interface 200B to exit the conversation with the media editing assistant to indicate the confirmation of the optional editing mode currently applied to the media content to be published.
[0046] If a user wants to try other optional editing modes than the one currently applied to the media content to be published, the selection control corresponding to the other optional editing mode can be triggered. For example, if the optional editing mode currently applied to the media content to be published is the optional editing mode corresponding to control 221, user 102 can trigger control 222 to apply the optional editing mode corresponding to control 222 to the media content to be published.
[0047] If the user is satisfied with the editing effect corresponding to the optional editing mode tried this time, the confirmation for the optional editing mode tried this time can be triggered. For example, by triggering control 214 on interface 200B to exit the conversation with the media editing assistant to indicate confirmation of the optional editing mode currently applied to the media content to be published.
[0048] If it is detected that the user confirms the optional editing mode, the optional editing mode is determined as the target editing mode for editing the media content to be published. Further, electronic device 110 can present the editing effect of applying the target editing mode to the media content to be published on another interface (hereinafter also referred to as the "second interface"). As Figure 2D shown, the editing effect corresponding to the selected target editing mode can be presented in interface 200C.
[0049] Edit request understanding and execution process
[0050] In the example described above, the user's editing request for the media content to be published can be associated with one or more functions (which can also be referred to as "tools") that electronic device 110 can implement. For example, for the editing request "Help me find a material template suitable for spring" input by the user shown in Figure 2C , electronic device 110 may need to perform special effect processing on the media content to be published, such as adjusting the color saturation and brightness and / or adding stickers / text, etc. In addition, electronic device 110 may also need to perform the operation of adding a relaxing background music to the media content to be published. This requires analyzing and understanding the editing request in the form of natural language input by the user.
[0051] As already mentioned above, the implementation of at least some functions of target application 120 can be based on a model. The model can be deployed in server 130. During the operation of target application 120, the capabilities of one or more models can be called. In target application 120, the media editing assistant can use the model to understand the user input and provide a response to the editing request of user 102 based on the output of the model. Based on this, Figure 3 shows a block diagram of an interaction process according to some embodiments of the present disclosure.
[0052] As Figure 3 shown, after the target application 120 deployed on the electronic device 110, such as the electronic device 110, receives an edit request 211 from the user for the media content to be published, the electronic device 110 provides the edit request 211 to the execution system 302 at the server 130.
[0053] In some embodiments, the execution system 302 may obtain registration information during the process of the electronic device 110 registering to the server 130. For example, the registration information may include available function information supported by the electronic device 110 and / or the target application 120. For example, the available function information may indicate that the electronic device 110 and / or the target application 120 have functions such as adding music, adding special effects, and overlaying materials for media content editing. In addition, the available function information may include, for example, a music list of the music library, a special effect template list of the special effect library, and a material list of the material library.
[0054] The received available function information (such as editing ability information) of the electronic device 110 and / or the target application 120 can be parsed into independent function information in the execution system 302. These function information have a predetermined information format. The function information may indicate, for example, function name, function description, function classification, and descriptions of model input and model output parameters associated with the available function. The following shows an example of the information format of the example function information.
[0055] Table 1
[0056]
[0057]
[0058] Based on the obtained available function information of the electronic device 110 and / or the target application 120 and the received edit request for the media content, the execution system 302 may provide at least one prompt word associated with the edit request and the available function information to the model 131. The model 131 may be, for example, a natural language model.
[0059] For example, if the edit request is "add interesting music", the execution system 302 determines that the function required to execute the edit request is adding music, and the prompt word is interesting music. In this case, the input provided to the model 131 may be represented as: function name: adding music; input parameter associated with the function (i.e., the prompt word): {keyword: "interesting music"}. The output of the model 131 may be, for example, the name of the music required for editing and the identifier (ID) of the music.
[0060] It should be understood that the above only gives an example of using Model 131 to determine the functions associated with the editing request, so as to understand the input and output of Mode 131. The input parameters and output parameters associated with the functions may also include other parameters shown in Table 1. For example, when the function is to add music, depending on the editing request, the prompt words may also include music types, music name keywords, etc.
[0061] The output of Model 131 is provided to Execution System 302, and Execution System 302 provides the output as the functions required for editing media content to Electronic Device 110, so that Electronic Device 110 presents at least one optional editing mode corresponding to the function through the interface of Target Application 120.
[0062] In some embodiments, the user's editing request for media content may be associated with multiple functions. After receiving an indication of the functions associated with the editing request provided by Server 130, Electronic Device 110 may determine whether the functions provided by Server 130 can meet the editing request.
[0063] For example, for the editing request "Help me find a material template suitable for spring" entered by the user shown in Figure 2C In addition to adding a relaxing background music, Electronic Device 110 may need to perform special effect processing on the media content to be published, such as adjusting the color saturation and brightness and / or adding stickers / text, etc. In this case, the editing request cannot be well implemented only based on the function of adding music.
[0064] Therefore, Electronic Device 110 can then request Server 130 again to obtain additional functions. Execution System 302 can obtain the output from Model 131 again based on other functions associated with the editing request, such as adding templates and the prompt words associated with the templates to be added, such as the recommended optional template names and template identifiers, etc., and provide the model output to Electronic Device 110.
[0065] If Electronic Device 110 determines that the functions obtained from the Server 130 side can implement the operations associated with the editing request, at least one optional editing mode corresponding to the obtained functions is presented on the interface of Target Application 120, such as the controls 221 to 224 corresponding to multiple optional editing modes shown in Figure 2C .
[0066] In addition, Electronic Device 110 can also send the result of the editing performed at Electronic Device 110 based on the functions provided by the server to Server 130.
[0067] In this way, it is possible to present at least one recommended optional editing mode to the user based on the parsing of the editing request in the form of natural language, thereby significantly improving the editing efficiency of media content and simplifying the editing process. At the same time, it is possible to provide a more personalized and user-intention-compliant solution in media content editing, improving user satisfaction.
[0068] Example process
[0069] Figure 4 FIG. 400 shows a flowchart of an interaction process 400 according to some embodiments of the present disclosure. The process 400 may be implemented at the electronic device 110. It should be understood that the process 400 may also be implemented at the server 130.
[0070] At block 410, the electronic device 110 receives an editing request for media content in a first interface, the editing request including a request in the form of natural language, the editing request indicating superimposing media materials and / or adding special effects to the media content.
[0071] At block 420, based on the editing request, the electronic device 110 simultaneously presents at least one optional editing mode of the media content and a first editing effect after applying a first optional editing mode among the at least one optional editing mode to the media content in the first interface.
[0072] At block 430, if the electronic device 110 receives a confirmation for a target editing mode among the at least one optional editing mode, then at block 440, the electronic device 110 presents the media content edited with a target editing effect corresponding to the target editing mode in a second interface.
[0073] In some embodiments, the electronic device 110 may provide a guide for obtaining the editing request in a third interface presenting the unedited media content; and in response to receiving a trigger operation for the guide, present the first interface.
[0074] In some embodiments, the electronic device 110 may, in response to receiving a confirmation for the first optional editing mode, present the media content with the first editing effect in the second interface.
[0075] In some embodiments, the electronic device 110 may, in response to receiving a selection of a second optional editing mode among the at least one optional editing mode, present a second editing effect after applying the second optional editing mode to the media content in the first interface; and in response to receiving a confirmation for the second optional editing mode, present the media content with the second editing effect in the second interface.
[0076] In some embodiments, the electronic device 110 may determine at least one target function corresponding to the editing request, where the at least one target function is determined by: providing at least one prompt word to a target model and obtaining a model output from the target model, the at least one prompt word being generated based on the editing request and information on a plurality of available functions associated with the editing of the media content; and presenting the at least one optional editing mode based on the at least one target function.
[0077] In some embodiments, the information on the plurality of available functions includes at least one of the following: function name; function description; function classification; and descriptions of model input and model output parameters associated with the available functions.
[0078] The electronic device 110 may, in response to determining a first target function corresponding to the editing request, determine whether the at least one optional editing mode can be presented based on the first target function; in response to determining that the at least one optional editing mode cannot be presented based on the first target function, determine a second target function corresponding to the editing request; and present the at least one optional editing mode based on the first target function and the second target function.
[0079] Example devices and equipment
[0080] Embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 5 A schematic structural block diagram of a device 500 for interaction according to some embodiments of the present disclosure is shown.
[0081] As Figure 5 shown, the device 500 may include a receiving module 510 configured to receive an editing request for the media content at a first interface, the editing request including a request in natural language form, the editing request indicating superimposing media materials and / or adding special effects to the media content. The device 500 may also include a first presentation module 520 configured to, based on the editing request, simultaneously present at least one optional editing mode of the media content and a first editing effect after applying a first optional editing mode among the at least one optional editing mode to the media content in the first interface. In addition, the device 500 may further include a second presentation module 530 configured to, in response to receiving a confirmation of a target editing mode among the at least one optional editing mode, present the media content edited with a target editing effect corresponding to the target editing mode in a second interface.
[0082] In some embodiments, the device 500 may further include a guidance providing module configured to: provide guidance for obtaining the editing request in a third interface that presents the unedited media content; and a third presentation module configured to: in response to receiving a trigger operation for the guidance, present the first interface.
[0083] In some embodiments, the device 500 may further include a first editing effect presentation module configured to: in response to receiving confirmation for the first optional editing mode, present the media content with the first editing effect in the second interface.
[0084] In some embodiments, the device 500 may further include a second editing effect presentation module configured to: in response to receiving a selection of a second optional editing mode from the at least one optional editing mode, present a second editing effect after applying the second optional editing mode to the media content in the first interface; and in response to receiving confirmation for the second optional editing mode, present the media content with the second editing effect in the second interface.
[0085] In some embodiments, the first presentation module 520 may further be configured to: determine at least one target function corresponding to the editing request, where the at least one target function is determined by: providing at least one prompt to a target model and obtaining a model output from the target model, the at least one prompt being generated based on the editing request and information about a plurality of available functions associated with the editing of the media content; and presenting the at least one optional editing mode based on the at least one target function.
[0086] In some embodiments, the information about the plurality of available functions includes at least one of the following: function name; function description; function classification; and descriptions of model input and model output parameters associated with the available function.
[0087] In some embodiments, the first presentation module 520 may further be configured to: in response to determining a first target function corresponding to the editing request, determine whether the at least one optional editing mode can be presented based on the first target function; in response to determining that the at least one optional editing mode cannot be presented based on the first target function, determine a second target function corresponding to the editing request; and present the at least one optional editing mode based on the first target function and the second target function.
[0088] The units included in apparatus 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units in apparatus 500 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0089] As Figure 6 shown, computing device / server 600 is in the form of a general-purpose computing device. The components of computing device / server 600 can include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 can be an actual or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capability of computing device / server 600.
[0090] Computing device / server 600 generally includes multiple computer storage media. Such media can be any available media accessible to computing device / server 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, magnetic disks, or any other media that can be used to store information and / or data (such as training data for training) and can be accessed within computing device / server 600.
[0091] Computing device / server 600 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 6As shown, a disk drive for reading from and writing to a removable, non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading from and writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. Memory 620 may include a computer program product 625 having one or more program modules configured to execute the various methods or acts of the various embodiments of the present disclosure.
[0092] Communication unit 640 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of computing device / server 600 can be implemented in a single computing cluster or multiple computer machines capable of communicating via a communication connection. Thus, computing device / server 600 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0093] Input device 650 can be one or more input devices such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices such as a display, speaker, printer, etc. Computing device / server 600 can also communicate with one or more external devices (not shown) as needed via communication unit 640, external devices such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with computing device / server 600, or communicate with any device that enables computing device / server 600 to communicate with one or more other computing devices (e.g., a network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0094] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which one or more computer instructions are stored, and the one or more computer instructions are executed by a processor to implement the method described above.
[0095] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0096] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0097] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other devices to implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0098] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.
[0099] The implementations of the present disclosure have been described above. The description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skilled persons in the art to understand the implementations disclosed herein.
Claims
1. A method for interaction, comprising: receiving, on a first interface, an editing request for media content, the editing request including a request in natural language form, the editing request indicating superimposing media materials and / or adding special effects to the media content; based on the editing request, presenting, on the first interface, at least one optional editing mode of the media content and a first editing effect after applying a first optional editing mode among the at least one optional editing mode to the media content; and in response to receiving confirmation of a target editing mode among the at least one optional editing mode, presenting, on a second interface, the media content edited with a target editing effect corresponding to the target editing mode.
2. The method according to claim 1, further comprising: providing, on a third interface presenting the unedited media content, a guide for obtaining the editing request; in response to receiving a trigger operation for the guide, presenting the first interface.
3. The method according to claim 1 or 2, further comprising: in response to receiving confirmation of the first optional editing mode, presenting, on the second interface, the media content with the first editing effect.
4. The method according to claim 1 or 2, further comprising: in response to receiving a selection of a second optional editing mode among the at least one optional editing mode, presenting, on the first interface, a second editing effect after applying the second optional editing mode to the media content; and in response to receiving confirmation of the second optional editing mode, presenting, on the second interface, the media content with the second editing effect.
5. The method according to claim 1, wherein presenting the at least one optional editing mode comprises: determining at least one target function corresponding to the editing request, wherein the at least one target function is determined by: providing at least one prompt word to a target model and obtaining a model output from the target model, the at least one prompt word being generated based on the editing request and information on a plurality of available functions associated with the editing of the media content; and presenting the at least one optional editing mode based on the at least one target function.
6. The method according to claim 5, wherein the information on the plurality of available functions includes at least one of the following: function name; function description; function classification; and description of model input and model output parameters associated with the available function.
7. The method according to claim 5, wherein presenting the at least one optional editing mode based on the at least one target function comprises: in response to determining a first target function corresponding to the editing request, determining whether the at least one optional editing mode can be presented based on the first target function; in response to determining that the at least one optional editing mode cannot be presented based on the first target function, determining a second target function corresponding to the editing request; and presenting the at least one optional editing mode based on the first target function and the second target function.
8. An apparatus for interaction, comprising: A receiving module, configured to receive an editing request for media content in a first interface, where the editing request includes a request in natural language form, and the editing request indicates superimposing media materials and / or adding special effects to the media content; A first presenting module, configured to, based on the editing request, simultaneously present at least one optional editing mode of the media content and a first editing effect after applying a first optional editing mode among the at least one optional editing modes to the media content in the first interface; And A second presenting module, configured to, in response to receiving confirmation of a target editing mode among the at least one optional editing modes, present the media content edited with a target editing effect corresponding to the target editing mode in a second interface.
9. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, and the instructions, when executed by the at least one processing unit, cause the electronic device to execute the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, having stored thereon a computer program, where the program, when executed by a processor, implements the method according to any one of claims 1 to 8.