Method and apparatus for interaction, device, and storage medium
By automatically sending multimedia content, prompts, and model parameters to the machine learning model through the interactive interface, the problems of user operation complexity and high cost are solved, and the process of evaluating the capabilities of the machine learning model is simplified.
Patent Information
- Application Number
- PCT/CN2024/108301
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-05
AI Technical Summary
When evaluating the multimedia content understanding ability of machine learning models, users need to upload multimedia content themselves and specify prompt words and model parameters, which increases the complexity and cost of operation, and unreasonable settings affect the model's parsing effect.
It provides an interactive interface containing multimedia content, prompts, and model parameters. After the user selects, the information is automatically sent to the target machine learning model and the execution results are displayed, simplifying the user operation.
Users can intuitively understand the capabilities of machine learning models by simply selecting multimedia content and confirming delivery, reducing operational costs and complexity.
Smart Images

Figure CN2024108301_05022026_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices, and storage media for interaction Technical Field
[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and more particularly to methods, apparatuses, devices, and computer-readable storage media for interaction. Background Technology
[0002] Evaluating machine learning models, such as their ability to understand multimedia content, requires users to upload multimedia content and specify prompt words and various model parameters. Since prompt words and model parameters play a crucial role in the model's understanding ability, inappropriate settings can directly impact the model's parsing performance. This makes it difficult for users to assess the model's comprehension capabilities, and setting prompt words and model parameters also increases the complexity and cost of user operations.
[0003] Summary of the Invention
[0004] In a first aspect of this disclosure, a method for interaction is provided. The method includes: presenting an interactive interface, the interactive interface including at least one first interactive object, each of the at least one first interactive object being associated with multimedia content, prompt information, and model parameters, the prompt information describing a target task that a target machine learning model needs to perform on the multimedia content, and the model parameters indicating parameters used by the target machine learning model in performing the target task; in response to detecting a selection of a first interactive object among the at least one first interactive object, presenting the selected first interactive object, the prompt information associated with the selected first interactive object, and the model parameters in the interactive interface; in response to detecting a sending instruction for the selected first interactive object, sending the multimedia content associated with the selected first interactive object, the prompt information, and the model parameters to the target machine learning model; and in response to receiving an execution result for the target task from the target machine learning model, presenting the execution result in the interactive interface.
[0005] In a second aspect of this disclosure, an apparatus for interaction is provided. The apparatus includes: a first presentation module configured to present an interactive interface, the interactive interface including at least one first interactive object, each of the at least one first interactive object being associated with multimedia content, prompt information, and model parameters, the prompt information describing a target task that a target machine learning model needs to perform on the multimedia content, and the model parameters indicating parameters used by the target machine learning model in performing the target task; a second presentation module configured to, in response to detecting a selection of a first interactive object among the at least one first interactive object, present the selected first interactive object, the prompt information associated with the selected first interactive object, and the model parameters in the interactive interface; an information sending module configured to, in response to detecting a sending instruction for the selected first interactive object, send the multimedia content associated with the selected first interactive object, the prompt information, and the model parameters to the target machine learning model; and a third presentation module configured to, in response to receiving an execution result for the target task from the target machine learning model, present the execution result in the interactive interface.
[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0008] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method of the first aspect.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] Figure 2 illustrates a flowchart of a user interaction process according to some embodiments of the present disclosure;
[0013] Figures 3A to 3C show schematic diagrams of example interfaces according to some embodiments of the present disclosure;
[0014] Figure 4 shows a block diagram of a user interaction apparatus according to some embodiments of the present disclosure; and
[0015] Figure 5 shows a block diagram of an apparatus capable of implementing several embodiments of the present disclosure. Detailed Implementation
[0016] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0017] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0018] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0019] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0021] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0022] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0023] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.
[0024] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0025] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably. A model can also include different types of processing units or networks.
[0026] As briefly mentioned earlier, evaluating machine learning models, such as their ability to understand multimedia content, requires users to upload multimedia content and specify prompts and various model parameters. Uploading multimedia content involves editing and other operations, increasing user complexity and cost. Furthermore, since prompts and model parameters play a crucial role in the model's understanding, inappropriate settings directly impact the model's parsing performance. Users lacking specialized knowledge may specify inappropriate prompts or parameters, directly affecting the model's comprehension of multimedia content. This makes it difficult for users to assess the model's understanding capabilities, and setting prompts and parameters further increases user complexity and cost.
[0027] Embodiments of this disclosure propose a scheme for interaction. According to various embodiments of this disclosure, an interactive interface is presented, comprising at least one first interactive object, each of which is associated with multimedia content, prompt information, and model parameters. In response to detecting a selection of one of the at least one first interactive objects, the selected first interactive object, its associated prompt information, and model parameters are presented in the interactive interface. In response to detecting a sending command for the selected first interactive object, the multimedia content associated with the selected first interactive object, its prompt information, and model parameters are sent to a target machine learning model. And in response to receiving an execution result for a target task from the target machine learning model, the execution result is presented in the interactive interface. In this manner, when a user evaluates the ability of a machine learning model to understand multimedia content, they only need to select the corresponding multimedia content and confirm sending to send the multimedia content, its associated prompt information, and model parameters to the target machine learning model, and the execution result of the target machine learning model is returned to the user. This allows the user to intuitively understand the capabilities of the machine learning model with minimal operational cost.
[0028] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In environment 100, an application 130 for evaluating machine learning models, such as a web application or other types of applications, is deployed on a terminal device 110. In some embodiments, application 130 is provided to assist users with various task processing needs in different applications and scenarios (e.g., evaluating the ability of machine learning models to understand multimedia content). During interaction with application 130, the user inputs interactive messages, and application 130 responds to the user's input by providing reply messages. Typically, application 130 is able to support users inputting questions in natural language and perform tasks and provide replies based on the understanding and logical reasoning ability of the natural language input.
[0029] In some embodiments, application 130 can interact with end user 150 as a contact. For example, application 130 can be implemented in an instant messaging (IM) application. Application 130 can interact with end user 150 in a one-on-one chat session. In some embodiments, application 130 can interact with multiple users in a group chat session that includes multiple users.
[0030] In some embodiments, interaction messages with application 130 may include multimodal messages, such as text messages (e.g., natural language text), voice messages, image messages, video messages, and so on.
[0031] In environment 100, machine learning models 140-1, 140-2, ..., 140-N are deployed on server 120. For ease of discussion, machine learning models 140-1, 140-2, ..., 140-N can be collectively referred to as machine learning model 140 or referred to individually as machine learning model 140.
[0032] In some embodiments, the machine learning model 140 may be, for example, a model that processes image data or a model that processes video data.
[0033] In some embodiments, the terminal device 110 may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the client device may also support any type of user-facing interface (such as "wearable" circuitry).
[0034] In some embodiments, server 120 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 120 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc.
[0035] A communication connection can be established between server 120 and terminal device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 120 and terminal device 110 can achieve signaling interaction through the communication connection between them.
[0036] It should be understood that the structure and function of environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0037] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0038] Figure 2 shows a flowchart of an interaction process 200 according to some embodiments of the present disclosure. The following description uses the execution of the interaction process 200 by a terminal device as an example. The interaction scheme provided by the present disclosure will then be described in detail with reference to Figure 1.
[0039] In frame 210, application 130 in terminal device 110 presents an interactive interface.
[0040] The interactive interface includes at least one first interactive object, each of which is associated with multimedia content, prompt words, and model parameters. The prompt words describe the target task that the target machine learning model needs to perform on the multimedia content. The model parameters indicate the parameters used by the target machine learning model in performing the target task.
[0041] In some embodiments, the multimedia content, prompt information, and model parameters associated with the first interactive object can be stored locally on the terminal device 110. In some embodiments, the multimedia content, prompt information, and model parameters associated with the first interactive object can be stored on the server 120.
[0042] In this document, "interactive object" refers to elements, components, etc., that can be interacted with by the user in an interactive interface. In some embodiments, the first interactive object may include, for example, images, icons, cards, logos, etc., used to display multimedia content corresponding to the first interactive object. This disclosure does not limit the specific presentation method of the first interactive object.
[0043] In some embodiments, multimedia content may include, but is not limited to, audio, video, or images. For example, if the multimedia content is a video, the corresponding prompt information could be something like "describe the details of the video," and the target task is to use a corresponding machine learning model (e.g., a video understanding model) to analyze the video to "describe the details of the video." As another example, if the multimedia content is an image, the corresponding prompt information could be something like "describe the content in the image," and the target task is to use a corresponding machine learning model (e.g., an image recognition model) to analyze the image to "describe the content in the image."
[0044] It should be understood that the examples of machine learning models and their tasks given herein are for illustrative purposes only. Embodiments of this disclosure can be applied to the capability evaluation of various machine learning models involved in multimedia content processing.
[0045] Taking the first interactive object as an image displaying multimedia content as an example. As shown in Figure 3A, the interface 300A that the terminal device 110 can present displays images 310-1, 310-2, ..., 310-N. For ease of discussion, images 310-1, 310-2, ..., 310-N can be collectively referred to as images 310 or individually as images 310.
[0046] In interface 300A, image 310 is configured to be selectable. As an example, a selection button can be displayed on image 310. After image 310 is selected, as shown in the figure, for example, the border of image 310 can be thickened.
[0047] In box 220, in response to detecting the selection of a first interactive object among at least one first interactive object, the terminal device 110 presents the selected first interactive object, the prompt word information associated with the selected first interactive object, and model parameters in the interactive interface.
[0048] In some embodiments, in response to detecting the selection of a first interactive object among at least one first interactive object, the terminal device 110 may obtain the multimedia content, prompt word information, and model parameters associated with the selected first interactive object locally on the terminal device 110 or from the server 120.
[0049] After obtaining the multimedia content, prompt word information, and model parameters associated with the selected first interactive object, the terminal device 110 can present the selected first interactive object, the prompt word information associated with the selected first interactive object, and the model parameters in the interactive interface.
[0050] As shown in Figure 3B, the interface 300B presents an input area 320 and a parameter configuration area 330. After the terminal device 110 detects that the first interactive object is selected, it can present the selected first interactive object and the corresponding prompt information 340 of the selected first interactive object in the input area 320, and present the model parameters 350 corresponding to the selected first interactive object in the parameter configuration area 330.
[0051] It should be understood that the selected first interactive object and the corresponding prompt information 340 of the selected first interactive object presented in the input area 320 can be presented in any suitable manner, and the model parameters 350 corresponding to the selected first interactive object presented in the parameter configuration area 330 can also be presented in any suitable manner. This embodiment of the present disclosure does not impose any restrictions.
[0052] In some embodiments, when the first interactive object is presented in the input area 320, the selected image 310 may be presented directly, or other icons, text, etc., that can indicate the selected image 310 may be presented. The embodiments disclosed herein are not limited to these embodiments.
[0053] In some embodiments, the prompt information presented in the interactive interface can be edited. Referring again to FIG3B, the prompt information 340 presented in the input area 320 can be, for example, default prompt information, that is, prompt information obtained by the terminal device 110. The terminal user 150 can edit the prompt information in the input area 320 according to actual needs.
[0054] In some embodiments, in response to detecting that the prompt word information 340 in the input area 320 has been edited, the terminal device 110 may present the edited prompt word information in the interface 300B. As an example, the terminal device 110 may present the user's editing process while the terminal user 150 is editing the prompt word information; or it may present the final edited result after the user has completed the editing process.
[0055] In some embodiments, the model parameters in the interactive interface can be edited. Referring again to Figure 3B, the model parameters 350 presented in the parameter configuration area 330 can be, for example, default model parameters, i.e., model parameters obtained by the terminal device 110. The terminal user 150 can edit the model parameters in the parameter configuration area 330 according to actual needs.
[0056] In some embodiments, in response to detecting that the model parameter 350 in the parameter configuration area 330 has been edited, the terminal device 110 can display the edited model parameter in the interface 300B.
[0057] In some embodiments, model parameters may include at least one of the following: sampling parameters, length of the execution result text, number of frames in the multimedia content, resolution used for capturing the multimedia content, frame rate used for capturing the multimedia content, method of capturing frames in the multimedia content, or region of capturing the multimedia content.
[0058] Sampling parameters indicate the randomness of the machine learning model's sampling of frames from multimedia content. For example, a higher sampling parameter results in greater randomness in the sampling of frames from multimedia content, and vice versa. The execution result text length indicates the maximum length of the text output by the machine learning model. The number of frames sampled from the multimedia content indicates the number of frames sampled from the multimedia content. The method of sampling frames from the multimedia content can indicate, for example, the method of sampling frames uniformly, or sampling one or more keyframes. The area for sampling multimedia content can include, for example, sampling a region of a predetermined shape from the frames of the multimedia content, generating a predetermined shape from the sampled frames, or filling the black borders of the frames to form a predetermined shape.
[0059] In box 230, in response to detecting a sending instruction for the selected first interactive object, the terminal device 110 sends the multimedia content associated with the selected first interactive object, prompt word information, and model parameters to the target machine learning model.
[0060] Before sending the multimedia content, prompts, and model parameters associated with the selected first interactive object to the target machine learning model, the target machine learning model needs to be determined.
[0061] In some embodiments, the interactive interface presented by the terminal device 110 further includes a second interactive object, which indicates at least one machine learning model. In response to detecting a selection of a machine learning model among the at least one machine learning model, the terminal device 110 determines the selected machine learning model as the target machine learning model. In some embodiments, the second interactive object may be the name of the machine learning model.
[0062] In some embodiments, the terminal device 110 may send the identifier or name of the determined target machine learning model to the server 120, so that the server 120 can determine the target machine learning model from a plurality of machine learning models.
[0063] Referring again to Figure 3B, interface 300B displays a model selection area 360, in which the name 370 of the selected machine learning model is displayed. As an example, the name of the machine learning model 370 can be presented as a drop-down menu. That is, when end user 150 clicks on the model selection area 360, a drop-down menu will pop up, displaying multiple machine learning models. When end user 150 selects any machine learning model, the corresponding machine learning model name 370 will be displayed in the model selection area 360.
[0064] After determining the target machine learning model, the end user 150 can click the send button 325 to send the multimedia content associated with the selected first interactive object, prompt words, and model parameters to the server 120. The server 120 then sends these same multimedia content, prompt words, and model parameters to the target machine learning model to execute the target task. After the target machine learning model completes the target task, the server 120 sends the execution result back to the end device 110.
[0065] In box 240, terminal device 110, in response to receiving the execution result for the target task from the target machine learning model, presents the execution result in the interactive interface.
[0066] In some embodiments, the interactive interface includes an interactive window for interacting with the target machine learning model. The terminal device 110 may present the selected first interactive object, the prompt word information associated with the selected first interactive object, and the execution result as the context of the interactive window in the interactive window.
[0067] As shown in Figure 3C, the interface 300C presents an input area 320 and an interaction window 390. After the terminal user 150 clicks the send button 325, the first interactive object and prompt information can be stopped in the input area 320, and the first interactive object and prompt information can be displayed in the interaction window 390.
[0068] In response to receiving the execution result for the target task from the target machine learning model, the terminal device 110 can present the execution result 380 in the interaction window 390 as the context of the interaction window 390, so as to facilitate the user to observe the execution result output by the target machine learning model for the target task.
[0069] In the interaction window 390, the prompt message information 340 is configured to be editable. After the end user 150 edits the prompt message information 340, they can click the send button 325 again to send the multimedia content associated with the selected first interactive object, the edited prompt message information, and the model parameters to the target machine learning model. Alternatively, the terminal device can send only the edited prompt message information to the target machine learning model so that the machine learning model can perform the target task again based on the edited prompt message information.
[0070] Similarly, after the target machine learning model returns the execution result, the model parameters 350 can also be edited. After the end user 150 edits the model parameters 350, they can click the send button 325 again to send the multimedia content associated with the selected first interactive object, the prompt word information, and the edited model parameters to the target machine learning model, or to send only the edited model parameters to the target machine learning model so that the machine learning model can execute the target task again based on the edited model parameters.
[0071] It should be understood that after the target machine learning model returns the execution result, the prompt word information 340 and the model parameters 350 can be edited separately or simultaneously.
[0072] Referring again to Figure 3C, after the terminal device 110 sends the multimedia content associated with the selected first interactive object, the prompt word information, and the model parameters to the target machine learning model again, the first interactive object sent again and the prompt word information 340-1 sent again can be presented in the interaction window 390. After the target machine learning model returns the execution result, the returned execution result 380-1 is presented in the interaction window as a new context in the interaction window 390.
[0073] In this embodiment of the disclosure, when evaluating the ability of a machine learning model to understand multimedia content, the user only needs to select the corresponding multimedia content and confirm the sending. The multimedia content, the prompt word information associated with the multimedia content, and the model parameters can be sent to the target machine learning model, and the execution result of the target machine learning model can be returned to the user. This allows the user to intuitively understand the capabilities of the machine learning model with minimal operational cost.
[0074] Figure 4 shows a schematic structural block diagram of an interactive device 400 according to certain embodiments of the present disclosure. The device 400 may be implemented as or included in the terminal device 110. The various modules / components in the device 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0075] As shown in the figure, device 400 includes a first presentation module 410 configured to present an interactive interface. The interactive interface includes at least one first interactive object, each of which is associated with multimedia content, prompt information, and model parameters. The prompt information describes the target task that a target machine learning model needs to perform on the multimedia content, and the model parameters indicate the parameters used by the target machine learning model in performing the target task. Device 400 also includes a second presentation module 420 configured to, in response to detecting a selection of a first interactive object among the at least one first interactive object, present the selected first interactive object, the prompt information associated with the selected first interactive object, and the model parameters in the interactive interface. Device 400 also includes an information sending module 430 configured to, in response to detecting a sending instruction for the selected first interactive object, send the multimedia content associated with the selected first interactive object, the prompt information, and the model parameters to the target machine learning model. Device 400 also includes a third presentation module 440 configured to, in response to receiving an execution result for the target task from the target machine learning model, present the execution result in the interactive interface.
[0076] In some embodiments, the interactive interface includes an interactive window for the target machine learning model, and presenting the execution result in the interactive interface includes presenting the selected first interactive object, the prompt word information associated with the selected first interactive object, and the execution result as the context of the interactive window in the interactive window.
[0077] In some embodiments, the device 400 further includes an acquisition module configured to acquire the multimedia content associated with the selected first interactive object, the prompt word information, and the model parameters.
[0078] In some embodiments, the prompt word information presented in the interactive interface can be edited, and the device 400 further includes a fourth presentation module configured to present the edited prompt word information in the interactive interface in response to detecting that the prompt word information has been edited.
[0079] In some embodiments, the model parameters presented in the interactive interface can be edited, and the device 400 further includes a fifth presentation module configured to present the edited model parameters in the interactive interface in response to detecting that the model parameters have been edited.
[0080] In some embodiments, the interactive interface further includes a second interactive object, which indicates at least one machine learning model.
[0081] In some embodiments, the apparatus 400 further includes a model determination module configured to determine the selected machine learning model as the target machine learning model in response to detecting the selection of a machine learning model among the at least one machine learning models.
[0082] In some embodiments, the model parameters include at least one of the following: sampling parameters, length of the execution result text, number of frames in the multimedia content, resolution used for capturing the multimedia content, frame rate used for capturing the multimedia content, method of capturing frames in the multimedia content, or area for capturing the multimedia content.
[0083] Figure 5 shows a block diagram illustrating an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in Figure 5 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 can be used to implement the terminal device 110 of Figure 1.
[0084] As shown in Figure 5, the electronic device 500 is in the form of a general-purpose electronic device. Components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 500.
[0085] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.
[0086] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0087] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0088] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0089] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0090] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0091] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0092] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0094] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for interaction, comprising: An interactive interface is presented, the interactive interface including at least one first interactive object, each of the at least one first interactive object being associated with multimedia content, prompt word information and model parameters, the prompt word information being used to describe the target task that the target machine learning model needs to perform on the multimedia content, and the model parameters indicating the parameters used by the target machine learning model in performing the target task; In response to detecting a selection of a first interactive object among the at least one first interactive objects, the selected first interactive object, the prompt word information associated with the selected first interactive object, and the model parameters are presented in the interactive interface; In response to detecting a sending instruction for the selected first interactive object, the multimedia content associated with the selected first interactive object, the prompt word information, and the model parameters are sent to the target machine learning model; as well as In response to receiving the execution result for the target task from the target machine learning model, the execution result is presented in the interactive interface.
2. The method according to claim 1, wherein the interactive interface includes an interactive window for the target machine learning model, and presenting the execution result in the interactive interface includes: The selected first interactive object, the prompt information associated with the selected first interactive object, and the execution result are presented in the interactive window as the context of the interactive window.
3. The method of claim 1, wherein before presenting the selected first interactive object, the prompt word information associated with the selected first interactive object, and the model parameters, the method further comprises: Obtain the multimedia content associated with the selected first interactive object, the prompt word information, and the model parameters.
4. The method according to claim 1, wherein the prompt information presented in the interactive interface is editable, and the method further includes: In response to detecting that the prompt word information has been edited, the edited prompt word information is displayed in the interactive interface.
5. The method according to claim 1, wherein the model parameters presented in the interactive interface are editable, and the method further comprises: In response to detecting that the model parameters have been edited, the edited model parameters are displayed in the interactive interface.
6. The method of claim 1, wherein the interactive interface further comprises a second interactive object, the second interactive object indicating at least one machine learning model.
7. The method according to claim 6, further comprising: In response to detecting a selection of a machine learning model among the at least one machine learning model, the selected machine learning model is determined as the target machine learning model.
8. The method of claim 1, wherein the model parameters include at least one of the following: Sampling parameters, Length of the execution result text Capture the frame count of images in multimedia content. The resolution used to capture multimedia content. The frame rate used to capture multimedia content. Methods for capturing video frames from multimedia content, or The area for collecting multimedia content.
9. A device for interaction, comprising: The first presentation module is configured to present an interactive interface, which includes at least one first interactive object. Each of the at least one first interactive object is associated with multimedia content, prompt word information, and model parameters. The prompt word information is used to describe the target task that the target machine learning model needs to perform on the multimedia content. The model parameters indicate the parameters used by the target machine learning model in performing the target task. The second presentation module is configured to, in response to detecting a selection of a first interactive object among the at least one first interactive object, present in the interactive interface the selected first interactive object, the prompt information associated with the selected first interactive object, and the selected first interactive object's prompt word information. Describe the model parameters; The information sending module is configured to, in response to detecting a sending instruction for the selected first interactive object, send the multimedia content associated with the selected first interactive object, the prompt word information, and the model parameters to the target machine learning model; as well as The third presentation module is configured to present the execution result in the interactive interface in response to receiving the execution result for the target task from the target machine learning model.
10. An electronic device, comprising: At least one processing unit; as well as At least one memory is coupled to at least one processing unit and stores instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 8.
12. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.