Interface interaction method and device, equipment and storage medium
By synchronously acquiring media content during audio acquisition and merging it into a single message, the problem of cumbersome information sending steps and insufficient content relevance in existing technologies is solved, thereby improving user experience and interaction accuracy.
Patent Information
- Application Number
- CN202511081076.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies involve cumbersome information transmission steps during voice input, resulting in a poor user experience. Furthermore, the lack of correlation between voice content and captured content affects the accuracy of interaction.
During the audio content acquisition process, the image capture unit is activated based on trigger conditions to acquire media content, and the audio content and media content are merged into a single message for transmission, thereby improving the efficiency and richness of information acquisition.
By processing audio and media content synchronously, the types of message content are enriched, the efficiency of message sending and interactivity are improved, and the relevance and accuracy of the content are ensured.
Smart Images

Figure CN120980053A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to interface interaction methods, apparatuses, devices, and computer-readable storage media. Background Technology
[0002] With the development of computer technology, the Internet has become an important way for people to interact and obtain information. People can not only communicate on the Internet platform, but also achieve various virtual interactions, such as conversational interaction with real users or virtual objects. Summary of the Invention
[0003] In a first aspect of this disclosure, a user interface interaction method is provided. The method includes: in response to receiving a first operation in a session interface, activating an audio acquisition unit to acquire audio content; during the acquisition of the audio content, in response to detecting that operation information meets a trigger condition, activating an image capture unit to acquire media content; and in the session interface, sending at least one message, the at least one message including a first part corresponding to the audio content and a second part corresponding to the media content.
[0004] In a second aspect of this disclosure, an apparatus for user interface interaction is provided. The apparatus includes: an audio acquisition module, a media acquisition module, and a message sending module. The audio acquisition module is configured to, in response to receiving a first operation in a session interface, activate an audio acquisition unit to acquire audio content; the media acquisition module is configured to, during the audio content acquisition process, in response to detecting that operation information meets a trigger condition, activate an image capture unit to acquire media content; and the message sending module is configured to, in the session interface, send at least one message, the at least one message including a first part corresponding to the audio content and a second part corresponding to the media content.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] FIG. 1 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;
[0011] FIGS. 2A-2E Example interfaces according to some embodiments of this disclosure are shown;
[0012] FIGS. 3A-3C An example interface according to an embodiment of the present disclosure is shown;
[0013] FIGS. 4A-4C An example interface according to an embodiment of the present disclosure is shown;
[0014] FIGS. 5A-5C An example interface according to an embodiment of the present disclosure is shown;
[0015] FIG. 6 A flowchart illustrating an example process of interface interaction according to some embodiments of the present disclosure is shown;
[0016] FIG. 7 A schematic structural block diagram of an example device for interface interaction according to some embodiments of the present disclosure is shown; and
[0017] FIG. 8 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation
[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0019] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0020] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0021] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0022] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0023] As mentioned above, with the development of computer technology, the Internet has become an important way for people to interact and obtain information. People can not only communicate on the Internet platform, but also achieve various virtual interactions, such as conversational interaction with real users or virtual objects.
[0024] Most existing interactive products support voice input, such as voice input based on long press. However, these technologies often only allow users to cancel input or convert the voice input to text. When it comes to asking questions or recognizing surrounding objects, it requires separately using the camera to take a picture and then sending the content. This multi-step and lengthy information transmission process results in a poor user experience.
[0025] Furthermore, this method of operation results in fragmented information. In scenarios involving interaction with physical users, content sent by other users may be interspersed between the voice and recorded content, affecting the correlation between the two. In scenarios involving interaction with virtual objects, the virtual object may respond to the voice content before the recorded content is successfully sent, or the virtual object may respond to the recorded content after it has been successfully sent but during the voice input process, leading to inaccurate responses.
[0026] The embodiments of this disclosure propose a user interface interaction scheme. The scheme includes: in response to receiving a first operation in a session interface, activating an audio acquisition unit to acquire audio content; during the audio content acquisition process, in response to detecting that operation information meets a trigger condition, activating an image capture unit to acquire media content; and in the session interface, sending at least one message, the at least one message including a first part corresponding to the audio content and a second part corresponding to the media content.
[0027] In this way, the embodiments of this disclosure can, during the acquisition of audio content, activate the image capturing unit to acquire media content based on triggering conditions, and send messages corresponding to the audio and video content. This can improve the efficiency of acquiring media content, effectively enrich the content types of the sent messages, enhance the richness of the message content, improve the efficiency of sending message content, and enhance the interactivity of the conversation interface.
[0028] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.
[0029] Example Environment
[0030] FIG. 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... FIG. 1 As shown, example environment 100 may include electronic device 110.
[0031] In this example environment 100, electronic device 110 can run an application 120 that supports user interface interaction. Application 120 can be any suitable type of application for user interface interaction, examples of which may include, but are not limited to, content sharing applications, live streaming applications, conversational applications, or other suitable applications. User 140 can interact with application 120 via electronic device 110 and / or its attached devices.
[0032] exist FIG. 1 In environment 100, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting interface interaction.
[0033] In some embodiments, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
[0034] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 that support user interface interaction in electronic devices 110.
[0035] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.
[0036] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0037] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0038] Example Interaction
[0039] FIGS. 2A-2E Example interfaces 200A to 200E according to some embodiments of the present disclosure are shown. Interfaces 200A to 200E may, for example, be provided by... FIG. 1 The electronic device 110 shown is provided.
[0040] In some embodiments, such as FIG. 2A As shown, the electronic device 110 can present an interface 200A. As an example, interface 200A can be an interface for virtual interactive applications, such as a conversational interface corresponding to a physical or virtual user. As an example, interface 200A can also be an interface that presents content including visual content, or an interface that enables virtual social interaction, content creation, and other interactive activities.
[0041] As an example, interface 200A can support multimodal interaction between user 140 and a conversation partner, which can correspond to a physical user or a virtual object. Such interaction modalities include, but are not limited to, text interaction, audio interaction, and visual interaction.
[0042] In some embodiments, the virtual object can be any suitable processing entity, such as an agent or a bot. As an example, the virtual object can be implemented based on any suitable machine learning model, such as a generative model like a language model.
[0043] In some embodiments, virtual objects may also correspond to or be associated with physical users. As an example, virtual objects may be created based on the association information of physical users. Taking the physical user as the first user, the virtual object may be created by the first user. For instance, the first user may create at least one virtual object based on the virtual avatar they wish to present, to perform virtual interactions and achieve different virtual interaction effects and experiences. Virtual objects may also be created by other users based on their historical interaction records with the first user; alternatively, with the authorization of the first user and the provision of association information, other users or the platform providing the virtual interaction service may create at least one virtual object based on the association information provided by the first user.
[0044] Through interface 200A, electronic device 110 can receive interactive operations from user 140. As an example, the interactive operations of user 140 may include at least one of the following: input operation, copy operation, share operation, media playback operation, etc.
[0045] refer to FIG. 2A As shown, in interface 200A, electronic device 110 can provide input component 210. Through input component 210, electronic device 110 can receive at least one type of input content, such as text content, audio content, video content, and image content, input by user 140.
[0046] In some embodiments, the input component 210 includes at least a voice input control 211. Via the voice input control 211, the electronic device 110 can acquire audio content input by the user 140. As an example, the audio content may be audio content pre-stored in the electronic device 110, or audio content pre-stored in other terminals or cloud devices associated with the user 140, or audio content captured by the user 140 in real time.
[0047] In some embodiments, in response to a trigger operation on the voice input control 211, the electronic device 110 may activate the audio acquisition unit to acquire the audio content of the user 140. As an example, the trigger operation may include any feasible operation such as a touch operation (e.g., a single click or double click), a swipe operation, or a long press operation.
[0048] In some embodiments, the audio acquisition unit activated by the electronic device 110 may include any audio acquisition device such as a microphone or a sound acquisition device.
[0049] Taking a long press operation as an example of triggering the voice input control 211, the operation information of the long press operation can include the trigger position of the long press operation. As an example, in response to the duration of the trigger operation on the voice input control 211 reaching a first threshold, the electronic device 110 determines that the received operation is a long press operation on the voice input control 211, and then starts the audio acquisition unit to acquire audio content.
[0050] In some embodiments, in response to the audio acquisition unit being successfully activated, the electronic device 110 can present FIG. 2B The interface 200B is shown. As an example, in response to a long press operation on the voice input control 211, the electronic device 110 can switch the presentation style of the voice input control 211 from a first style to a second style, for example, by... FIG. 2A The style shown has been switched to FIG. 2B The style shown.
[0051] refer to FIG. 2B As shown, during the process of acquiring audio content via the audio acquisition unit, the electronic device 110 can display the audio acquisition duration 221 on the interface 200B so that the user 140 can understand the duration information corresponding to the acquired audio content in real time.
[0052] In some embodiments, after the audio acquisition unit is started to acquire audio content, that is, during the acquisition of audio content, the trigger position of the long press operation can remain in the presentation position area corresponding to the voice input control 211, or it can move to other position areas in the interface 200B.
[0053] like FIG. 2B As shown, in some embodiments, during the acquisition of audio content, the electronic device 110 may also display a first prompt message 231 on the interface 200B. As an example, the first prompt message 231 may indicate at least one change state corresponding to the trigger position of the long-press operation and the operation content it indicates. For example, the first prompt message 231 may be presented as at least one of the following: "Release to send," "Move up to cancel," "Move up to record," "Move up to cancel or record," etc. As an example, "Release to send" may indicate stopping the long-press operation to send the currently acquired audio content; "Move up to cancel" may indicate moving the trigger position of the long-press operation upwards to stop acquiring audio content and delete the acquired audio content; "Move up to record" may indicate moving the trigger position of the long-press operation upwards to obtain image content or video content, etc.
[0054] In some scenarios, in response to detecting that the operation information associated with the voice input control 211 meets the trigger condition, such as the trigger position of a long press operation meeting the trigger condition, the electronic device 110 can activate the image capturing unit to acquire media content. As an example, the media content may include: image content and / or video content, and may also include audio content.
[0055] In some embodiments, in response to the successful activation of the audio acquisition unit, during the acquisition of audio content, the electronic device 110 may also present at least one indicator element, such as... FIG. 2B The first indicator element 212 and the second indicator element 213 are shown.
[0056] As an example, the first indicator element 212 can indicate the acquisition of media content during the acquisition of audio content. As an example, in response to the trigger position of a long press operation moving to the first area where the first indicator element 212 is located, the electronic device 110 can activate the image capturing unit to acquire media content. For example, the electronic device 110 can invoke a built-in camera or other image capturing device to capture images or record videos, etc.
[0057] As an example, the second indicator element 213 can indicate to stop capturing audio content. As an example, in response to the trigger position of the long press operation moving to the second area where the second indicator element 213 is located, the electronic device 110 can trigger the audio capture unit to shut down, thereby stopping the capture of audio content.
[0058] In some embodiments, in response to the trigger position of a long press operation moving to the first area where the first indicator element 212 is located, the electronic device 110 may display... FIG. 2C The interface shown is 200C.
[0059] refer to FIG. 2C As shown, in response to the trigger position of the long press operation moving to the first area where the first indicator element 212 is located, the electronic device 110 can display an image window 240. And the acquired media content 241 is displayed through the image window 240.
[0060] In some embodiments, the electronic device 110 may also display duration information 242 corresponding to the media content in the image window 240. The duration information 242 can characterize the duration for which the trigger position of the long press operation is located in the first area. By displaying the duration information 242, the electronic device 110 can facilitate the user 140 to view the continuous acquisition duration of the media content 241 in real time.
[0061] In some embodiments, the electronic device 110 may determine the type of the acquired media content based on the continuous acquisition duration of the media content. As an example, in response to the continuous acquisition duration of the media content being less than a first duration threshold, the electronic device 110 may determine that the acquired media content is an image; in response to the continuous acquisition duration of the media content being greater than or equal to the first duration threshold, the electronic device 110 may determine that the acquired media content is a video.
[0062] In some embodiments, the type of media content pair can also be determined based on an indicator element. As an example, the first indicator element 212 may include a first sub-element corresponding to an image type and a second sub-element corresponding to a video type. For example, in response to a trigger position moving to the position corresponding to the first sub-element, the electronic device 110 can activate an image capturing unit to acquire image-type media content. Similarly, in response to a trigger position moving to the position corresponding to the second sub-element, the electronic device 110 can acquire video-type media content via the image capturing unit.
[0063] In some embodiments, the first indication element 212 may also be configured to be associated with a second duration threshold. For example, the second duration threshold may be configured as the longest duration for which media content is acquired at one time. For example, in response to the duration information 242 reaching the second duration threshold, the electronic device 110 may shut down the image capturing unit and stop acquiring media content. For example, the electronic device 110 may present the acquired media content in an image window 240 or in other forms.
[0064] In some embodiments, the electronic device 110 may also convert the first indicator element 212 by FIG. 2B The third style shown is switched to FIG. 2C The fourth style is shown. As an example, the fourth style corresponding to the first indicator element 212 can also present media content acquisition progress information, such as a circular progress bar.
[0065] In some embodiments, in response to the trigger position of the long press operation moving to the first area where the first indicator element 212 is located, the electronic device 110 can switch the first prompt message 231 to the second prompt message 232. As an example, the second prompt message 232 may be presented as: "Release to send" and / or "Slide down to cancel".
[0066] In some embodiments, in response to the operation information of a trigger operation (e.g., a long press operation) on the voice input control 211 being updated to indicate that the trigger condition is not met, the electronic device 110 may stop acquiring media content. As an example, in response to the trigger position of the long press operation deviating from the first area where the first indicator element 212 is located, for example, moving outside the first area, the electronic device 110 may turn off the image capturing unit to stop acquiring media content.
[0067] As an example, the trigger position of the long press operation deviating from the first area where the first indicator element 212 is located can include: the trigger position of the long press operation shifting in any direction outside the first area where the first indicator element 212 is located, and not being located in the second area where the second indicator element 213 is located. For example, the trigger position moves to the left or right of the first indicator element 212, or moves below the second indicator element 213, etc.
[0068] In some embodiments, in response to the trigger position offset of the long press operation from the first area where the first indicator element 212 is located, the electronic device 110 may also display a first preview content corresponding to the acquired media content, such as... FIG. 2D The first preview content shown is 243.
[0069] refer to FIG. 2D As shown, in response to the trigger position of the long press operation being updated to no longer meet the trigger condition, the electronic device 110 can switch the second prompt message 232 to the third prompt message 233. As an example, the third prompt message 233 can be presented as "Release to send" and / or "Move up to cancel", etc.
[0070] As an example, in interface 200D, electronic device 110 can also continue to display the audio acquisition duration 221 corresponding to the currently acquired audio content, so that users can understand the current audio duration information in real time.
[0071] According to the scheme disclosed herein, during the acquisition of an audio content segment, the electronic device 110 can support acquiring only one segment of media content. For example, it can support acquiring a relatively complete media content at once, or it can support acquiring different media segments multiple times to form a single media content. As an example, during the acquisition of an audio content segment, the electronic device 110 can also support acquiring multiple media content segments separately.
[0072] In some embodiments, taking the media content acquired in the above process as the first media content as an example, after the electronic device 110 presents the first preview content 243 corresponding to the first media content, in response to the operation information (e.g., the trigger position of the long press operation) being updated again to meet the trigger condition, the electronic device 110 can restart the image capturing unit to acquire the second media content.
[0073] As an example, after presenting the first preview content 243, the electronic device 110 may also present the first indicator element 212, and in response to the trigger position of the long press operation, move back to the position of the first indicator element 212. The electronic device 110 determines that the operation information has been updated to meet the trigger condition, so that the image capturing unit can be restarted to obtain the second media content.
[0074] As an example, after presenting the first preview content 243, the electronic device 110 can also move to the location of the first preview content 243 based on the trigger position of the long press operation, determine that the operation information has been updated to meet the trigger condition, so that the electronic device 110 can restart the image capturing unit to obtain the second media content.
[0075] In some embodiments, the first media content and the second media content may correspond to the same media type. For example, the first media content and the second media content may both be video content or both be image content.
[0076] In some embodiments, the first media content and the second media content may correspond to different media types. For example, the first media content may be an image and the second media content may be a video; or, the first media content may be a video and the second media content may be an image, etc.
[0077] As an example, the electronic device 110 can be pre-configured to have a threshold of media content that can be acquired during the acquisition of an audio content, such as no more than 3 or 5.
[0078] In some embodiments, after presenting the first preview content 243 corresponding to the media content, the electronic device 110 may also activate the image capturing unit to acquire image content as an additional segment of the media content in response to the operation information being updated to meet the triggering conditions. As an example, the image content may be of image type or video type.
[0079] As an example, after presenting the first preview content 243, the electronic device 110 can also present the first indicator element, and in response to the trigger position of the long press operation, move back to the position of the first indicator element 212. The electronic device 110 can determine that the operation information has been updated to meet the trigger condition, so that the image capturing unit can be restarted to obtain additional segments of the image type or video type. Then, the obtained additional segments are combined with the media content corresponding to the first preview content 243. For example, the additional segments and the original media content are spliced together according to the acquisition time order to generate new media content.
[0080] As an example, after presenting the first preview content 243, the electronic device 110 can also move to the location of the first preview content 243 based on the trigger position of the long press operation, determine that the operation information has been updated to meet the trigger condition, so that the electronic device 110 can restart the image capturing unit to obtain additional segments of image type or video type, and combine the additional segments with the media content corresponding to the first preview content 243 to generate new media content.
[0081] In some embodiments, after activating the image capturing unit to acquire media content, in response to the operation information indicating that the trigger position of the long press operation moves to a preset second area, such as the second area where the second indicator element 213 is located, the electronic device 110 can turn off the image capturing unit to stop acquiring media content and can delete the acquired media content. As an example, the electronic device 110 can also delete the acquired audio content.
[0082] In some embodiments, in response to confirmation of the acquired audio content and / or media content, the electronic device 110 may present at least one message corresponding to the acquired audio content and / or media content in a session interface, such as... FIG. 2E As shown. As an example, the confirmation action for the captured audio content and / or media content can be to stop the long press action, such as "release".
[0083] Comprehensive reference FIG. 2D and FIG. 2E As shown, in response to the operation information received in interface 200D satisfying the sending conditions, electronic device 110 can present session interface 200E and send at least one message in session interface 200E. This at least one message includes a first part corresponding to the audio content and a second part corresponding to the media content. As an example, the sending conditions may include any of the following: stopping the long-press operation, the trigger position of the long-press operation moving to a preset area (e.g., a third area indicating message sending, or exceeding the trigger area associated with the voice input control 211, etc.). As an example, if the acquired media content includes second media content, the at least one message may also include a third part corresponding to the second media content, and so on, which will not be elaborated further here.
[0084] In some embodiments, the first part and the second part correspond to different messages. For example... FIG. 2E As shown, the second part corresponding to the media content can correspond to the first message 251, and the first part corresponding to the audio content can correspond to the second message 252.
[0085] In some embodiments, the first part corresponding to the audio content and the second part corresponding to the media content may correspond to different parts of the same message. That is, the electronic device 110 can simultaneously present the message content corresponding to the audio content and the message content corresponding to the media content in the same message.
[0086] In some embodiments, the first part includes text content corresponding to the audio content. As an example, the electronic device 110 may invoke a pre-trained model or speech recognition algorithm to recognize at least a portion of the acquired audio content, convert it into text content, and then present the converted text content as at least a part of a message in the conversation interface 200E.
[0087] In some embodiments, text content corresponding to the audio content can also be generated based on the acquired media content. During the process of converting the acquired audio content into text, the electronic device 110 can also perform speech recognition on the audio content in conjunction with the acquired media content. By using the main objects and background images in the media content as aids, recognition errors caused by polyphonic characters or tonal variations can be effectively avoided, ambiguity caused by different polyphonic characters can be effectively eliminated, thereby effectively improving the accuracy of audio content recognition and ensuring the accuracy of the text content.
[0088] In some embodiments, the session interface 200E may be associated with an agent. As an example, the electronic device 110 may provide response content 260 generated by the agent based on at least one message (e.g., a first message 251 and a second message 252) within the session interface 200E.
[0089] In some embodiments, the response content 260 may correspond to different response states. For example, after at least one message is presented in the interface 200E, the electronic device 110 may sequentially present a first response state, a second response state, etc., corresponding to the response content 260. For example, the first response state may indicate the recognition and analysis status of at least one message, such as "content understanding in progress" or "content analysis in progress"; the second response state may indicate the message text corresponding to the generated response content 260. For example, the electronic device 110 may present the message text corresponding to the response content 260 in the form of streaming output.
[0090] In some embodiments, the response content 260 is also generated based on the temporal relationship between the media content and the audio content. As an example, this temporal relationship indicates the association between a first time period corresponding to the media content and a second time period corresponding to the audio content.
[0091] As an example, for media content acquired in one go, the first time period of 0-13 seconds corresponding to the media content corresponds to 15-28 seconds of the audio content. That is, the acquisition of media content begins at the 15th second of the audio content acquisition process, and the acquisition of 13 seconds of media content is completed at the 28th second of the audio content acquisition process.
[0092] As an example, when an audio content corresponds to multiple media content segments, each segment can correspond to a different fragment of the audio content. For instance, during the acquisition of a 45-second audio content, if two media segments (e.g., 0-5 seconds and 6-15 seconds) are acquired, the 0-5 second segment of the media content can correspond to the 8-13 second segment of the audio content, and the 6-15 second segment can correspond to the 32-41 second segment. For example, electronic device 110 acquires the first media content associated with a first type of flower during the 8-13 second segment of the audio content acquisition. This audio segment can describe inquiries about the first type of flower, such as "What kind of flower is this?", "How can this flower be cultivated to produce more buds?", or "What are the precautions for cultivating this type of plant?". Then, during the 32-41 second segment of the audio content acquisition, it acquires the second media content associated with a second type of flower, and this audio segment can describe inquiries about the second type of flower, such as "Does this flower belong to the same family and genus as the first flower?", or "What are the differences in cultivation techniques between this flower and the first flower?".
[0093] In this way, the embodiments of this disclosure can, during the acquisition of audio content, activate the image capturing unit to acquire media content based on triggering conditions, and send messages corresponding to the audio and video content. This can effectively enrich the content types of the sent messages, improve the richness of the message content, improve the efficiency of message content acquisition, ensure the correlation between the acquired media content and audio content, and enhance the interactivity of the conversation interface.
[0094] FIGS. 3A-3C Example interfaces 300A to 300C according to some embodiments of the present disclosure are shown. Example interfaces 300A to 300C may, for example, be provided by... FIG. 1 The electronic device 110 shown is provided.
[0095] refer to FIG. 3A As shown, in response to a triggering action on the voice input control 311, such as a long press, the electronic device 110 activates the audio acquisition unit to acquire audio content, and can also switch the initial presentation style corresponding to the voice input control 311 to... FIG. 3A The third style shown.
[0096] During the process of audio acquisition by the audio acquisition unit, the electronic device 110 can display the audio acquisition duration 312 on the interface 300A so that the user 140 can understand the duration information corresponding to the acquired audio content in real time. As an example, the audio acquisition duration 312 can be displayed at a position associated with the voice input control 311. For example, the electronic device 110 can display the audio acquisition duration 312 within the display position range of the voice input control 311.
[0097] In some embodiments, during the acquisition of audio content, the electronic device 110 may also present a fourth prompt message 331 on the interface 300A. As an example, the fourth prompt message 331 may indicate at least one change state corresponding to the trigger position of the long-press operation and the operation content it indicates. For example, the fourth prompt message 331 may be presented as at least one of the following: "Release to send," "Move up to cancel," "Move left to record," etc. As an example, "Release to send" may indicate stopping the long-press operation to send the currently acquired audio content; "Move up to cancel" may indicate moving the trigger position of the long-press operation upwards to stop acquiring audio content and delete the acquired audio content; "Move left to record" may indicate moving the trigger position of the long-press operation to the left to obtain image content or video content, etc.
[0098] like FIG. 3A As shown, in some embodiments, in response to the successful activation of the audio acquisition unit, the electronic device 110 may also display an indicator element 321 during the acquisition of audio content. As an example, the indicator element 321 may indicate that media content is being acquired during the acquisition of audio content.
[0099] As an example, in response to the trigger position of a long press operation moving to the location area where the indicator element 321 is located, the electronic device 110 can activate the image capturing unit to acquire media content. For example, the electronic device 110 can invoke a built-in camera or other image capturing device to capture images or record videos, etc.
[0100] In some embodiments, in response to the trigger position of a long press operation moving to the location area where the indicator element 321 is located, the electronic device 110 may display... FIG. 3B The interface shown is 300B. (As shown in the image) FIG. 3B As shown, the electronic device 110 can display an image window 340 in the interface 300B, and can also display the acquired media content 341 through the image window 340.
[0101] In some embodiments, the electronic device 110 may also display duration information 342 corresponding to the media content in the image window 340. Duration information 342 can characterize the duration for which the trigger position of the long-press operation is located in the first area. By displaying duration information 342, the electronic device 110 allows the user 140 to easily view the continuous acquisition duration of the media content 341 in real time. The relationship between duration information 342 and the media type of media content 341 can be referred to... FIG. 2C The relevant descriptions will not be repeated here.
[0102] refer to FIG. 3BAs shown, in some embodiments, in response to moving the trigger position of the long press operation to the location area where the indicator element 321 is located, the electronic device 110 can also move the indicator element 321 from... FIG. 3A The presentation style shown has been switched to FIG. 3B The presentation style shown.
[0103] In some embodiments, in response to moving the trigger position of the long press operation to the location area where the indicator element 321 is located, the electronic device 110 may also switch the fourth prompt message 331 to the fifth prompt message 332. As an example, the fifth prompt message 332 may be presented as "Release to send" and / or "Move up to cancel", etc.
[0104] In some embodiments, in response to a trigger operation (e.g., a long press) on the voice input control 311, the operation information is updated to indicate that the trigger condition is not met, for example, the location area of the indicator element 321 is deviated from. The electronic device 110 may turn off the image capturing unit to stop acquiring media content. As an example, in response to the trigger position of the long press operation moving back to the location area of the indicator element 321, supplementary media content or new media content can be acquired, which will not be elaborated here.
[0105] In some embodiments, in response to the location area of the trigger position offset indicator element 321 of the long press operation, the electronic device 110 may also display preview content corresponding to the acquired media content, such as... FIG. 3C The preview content 343 is shown. As an example, the preview content 343 can be presented at the location of the indicator element 321, or it can be presented in other free areas of the interface 300C, which is not limited here.
[0106] FIGS. 4A-4C Example interfaces 400A to 400C according to some embodiments of the present disclosure are shown. Example interfaces 400A to 400C may, for example, be provided by... FIG. 1 The electronic device 110 shown is provided.
[0107] refer to FIG. 4A As shown, in response to a triggering action on the voice input control 411, such as a long press, the electronic device 110 activates the audio acquisition unit to acquire audio content, and can also switch the initial presentation style corresponding to the voice input control 411 to... FIG. 4A The fourth style shown.
[0108] During the process of the audio acquisition unit acquiring audio content, the electronic device 110 can display the audio acquisition duration 421 on the interface 400A so that the user 140 can understand the duration information corresponding to the acquired audio content in real time.
[0109] In some embodiments, during the acquisition of audio content, the electronic device 110 may also display a sixth prompt message 431 on the interface 400A. As an example, the sixth prompt message 431 may indicate at least one change state corresponding to the trigger position of the long press operation and the operation content indicated therein. For example, the sixth prompt message 431 may be presented as at least one of the following: "Release to send", "Move to the upper right to cancel", "Move to the upper left to record", etc.
[0110] like FIG. 4A As shown, in some embodiments, in response to the successful activation of the audio acquisition unit, the electronic device 110 may also display an indicator element 412 during the acquisition of audio content. As an example, the indicator element 412 may indicate that media content is being acquired during the acquisition of audio content.
[0111] As an example, in response to the trigger position of a long press operation moving to the location area where the indicator element 412 is located, the electronic device 110 can activate the image capturing unit to acquire media content. For example, the electronic device 110 can invoke a built-in camera or other image capturing device to capture images or record videos, etc.
[0112] In some embodiments, during the acquisition of audio content, the electronic device 110 may also display an indicator element 413. For example, the indicator element 413 may indicate that audio content acquisition should be stopped. For example, in response to the trigger position of a long press operation moving to the area where the indicator element 413 is located, the electronic device 110 may trigger the audio acquisition unit to shut down, thereby stopping the acquisition of audio content.
[0113] In some embodiments, in response to the trigger position of a long press operation moving to the location area where the indicator element 412 is located, the electronic device 110 may display... FIG. 4B The interface shown is 400B. (As shown...) FIG. 4B As shown, the electronic device 110 can display an image window 440 in the interface 400B, and can also display the acquired media content 441 through the image window 440.
[0114] In some embodiments, the electronic device 110 may also display duration information 442 corresponding to the media content in the image window 440. Duration information 442 can characterize the duration for which the trigger position of the long-press operation is located in the first area. By displaying duration information 442, the electronic device 110 allows the user 140 to easily view the continuous acquisition duration of the media content 441 in real time. The relationship between duration information 442 and the media type of media content 441 can be referred to... FIG. 2C The relevant descriptions will not be repeated here.
[0115] refer to FIG. 4BAs shown, in some embodiments, in response to moving the trigger position of the long press operation to the location area where the indicator element 412 is located, the electronic device 110 can also move the indicator element 412 from... FIG. 4A The presentation style shown has been switched to FIG. 4B The presentation style shown.
[0116] In some embodiments, in response to a trigger operation (e.g., a long press) on the voice input control 411, the operation information is updated to indicate that the trigger condition is not met, for example, the location area of the indicator element 412 is deviated from. The electronic device 110 may turn off the image capturing unit to stop acquiring media content. As an example, in response to the trigger position of the long press operation moving back to the location area of the indicator element 412, supplementary media content or new media content can be acquired, which will not be elaborated here.
[0117] In some embodiments, in response to the location area of the trigger position offset indicator element 412 of the long press operation, the electronic device 110 may also display preview content corresponding to the acquired media content, such as... FIG. 4C The preview content 443 is shown. As an example, the preview content 443 can be presented at the location of the indicator element 412, or it can be presented in other free areas of the interface 400C, which is not limited here.
[0118] FIGS. 5A-5C Example interfaces 500A to 500C according to some embodiments of the present disclosure are shown. Example interfaces 500A to 500C may, for example, be provided by... FIG. 1 The electronic device 110 shown is provided.
[0119] refer to FIG. 5A As shown, in response to a triggering action on the voice input control 511, such as a long press, the electronic device 110 activates the audio acquisition unit to acquire audio content, and can also switch the initial presentation style corresponding to the voice input control 511 to... FIG. 5A The fifth style shown.
[0120] During the process of audio acquisition by the audio acquisition unit, the electronic device 110 can display the audio acquisition duration 521 on the interface 500A so that the user 140 can understand the duration information corresponding to the acquired audio content in real time.
[0121] In some embodiments, during the acquisition of audio content, the electronic device 110 may also display a seventh prompt message 530 on the interface 500A. As an example, the seventh prompt message 530 may indicate at least one change state corresponding to the trigger position of the long-press operation and the operation content it indicates. For example, the seventh prompt message 530 may be presented as at least one of the following: "Release to send," "Move left to cancel," "Move right to record," etc.
[0122] like FIG. 5A As shown, in some embodiments, in response to the successful activation of the audio acquisition unit, the electronic device 110 may also display an indicator element 512 during the acquisition of audio content. As an example, the indicator element 512 may indicate that media content is being acquired during the acquisition of audio content.
[0123] As an example, in response to the trigger position of a long press operation moving to the location area where the indicator element 512 is located, the electronic device 110 can activate the image capturing unit to acquire media content. For example, the electronic device 110 can invoke a built-in camera or other image capturing device to capture images or record videos, etc.
[0124] In some embodiments, during the acquisition of audio content, the electronic device 110 may also display an indicator element 513. For example, the indicator element 513 may indicate that audio content acquisition should be stopped. For example, in response to the trigger position of a long press operation moving to the area where the indicator element 513 is located, the electronic device 110 may trigger the shutdown of the audio acquisition unit to stop acquiring audio content.
[0125] In some embodiments, in response to the trigger position of a long press operation moving to the location area where the indicator element 512 is located, the electronic device 110 may display... FIG. 5B The interface shown is 500B. (As shown in the image...) FIG. 5B As shown, the electronic device 110 can display an image window 540 in the interface 500B, and can also display the acquired media content 541 through the image window 540.
[0126] In some embodiments, the electronic device 110 may also display duration information 542 corresponding to the media content in the image window 540. Duration information 542 can characterize the duration for which the trigger position of the long-press operation is located in the first area. By displaying duration information 542, the electronic device 110 allows the user 140 to easily view the continuous acquisition duration of the media content 541 in real time. The relationship between duration information 542 and the media type of media content 541 can be referred to... FIG. 2C The relevant descriptions will not be repeated here.
[0127] refer to FIG. 5BAs shown, in some embodiments, in response to moving the trigger position of the long press operation to the location area where the indicator element 512 is located, the electronic device 110 can also move the indicator element 512 from... FIG. 5A The presentation style shown has been switched to FIG. 5B The presentation style shown.
[0128] In some embodiments, in response to a trigger operation (e.g., a long press) on the voice input control 511, the operation information is updated to indicate that the trigger condition is not met, for example, the location area of the indicator element 512 is deviated from. The electronic device 110 may turn off the image capturing unit to stop acquiring media content. As an example, in response to the trigger position of the long press operation moving back to the location area of the indicator element 512, supplementary media content or new media content can be acquired, which will not be elaborated here.
[0129] In some embodiments, in response to the location area of the trigger position offset indicator element 512 of the long press operation, the electronic device 110 may also display preview content corresponding to the acquired media content, such as... FIG. 5C The preview content 543 is shown. As an example, the preview content 543 can be presented at the location of the indicator element 512, or it can be presented in other free areas of the interface 500C, which is not limited here.
[0130] Example Process
[0131] FIG. 6 A flowchart illustrating an example process 600 of interface interaction according to some embodiments of the present disclosure is shown. Process 600 can be implemented at electronic device 110. Reference is made below. FIG. 1 To describe process 600.
[0132] like FIG. 6 As shown in box 610, electronic device 110, in response to receiving a first operation in the session interface, starts the audio acquisition unit to acquire audio content.
[0133] In frame 620, during the acquisition of audio content, electronic device 110, in response to detecting that the operation information meets the triggering condition, starts the image capturing unit to acquire media content.
[0134] In box 630, electronic device 110 sends at least one message in the session interface, the at least one message including a first part corresponding to audio content and a second part corresponding to media content.
[0135] In this way, the embodiments of this disclosure can, during the acquisition of audio content, activate the image capturing unit to acquire media content based on triggering conditions, and send messages corresponding to the audio and video content. This can improve the efficiency of acquiring media content, effectively enrich the content types of the sent messages, increase the richness of the message content, improve the efficiency of sending message content, and enhance the interactivity of the conversation interface.
[0136] In some embodiments, the first operation includes a long press operation on the voice input control in the conversation interface, and the operation information indicates the trigger position of the long press operation.
[0137] In this way, the embodiments of this disclosure can determine the triggering condition for starting the image capturing unit based on the triggering position corresponding to the long press operation of the voice input control, thereby ensuring that media content is acquired during the process of acquiring audio content. This not only ensures the efficiency of media content acquisition, but also effectively ensures the correlation between the acquired media content and the acquired audio content.
[0138] In some embodiments, in response to detecting that the operation information meets the triggering condition, the image capturing unit is activated to acquire media content, including: presenting an indicator element during the acquisition of audio content; and in response to the operation information indicating that the triggering position of the long press operation moves to a first area associated with the indicator element, the image capturing unit is activated to acquire media content.
[0139] In this way, the embodiments of this disclosure can determine whether the operation information meets the triggering conditions based on the relationship between the triggering position of the long press operation and the presentation area associated with the indicator element, thereby activating the image capturing unit to acquire media content, thereby ensuring the timing and efficiency of media content acquisition, and ensuring the correlation between the acquired media content and the collected audio content.
[0140] In some embodiments, process 600 further includes: the electronic device 110 stops acquiring media content in response to the operation information indicating that the trigger position of the long press operation moves outside the first area.
[0141] In this way, the present embodiment can stop acquiring media content based on the first area corresponding to the indicator element when the trigger position of the long press operation deviates. This allows audio content to continue to be acquired even when media content acquisition is stopped, ensuring the continuity of the acquired audio content and avoiding resource occupation and waste caused by acquiring too much unnecessary media content.
[0142] In some embodiments, process 600 further includes: the electronic device 110 moving to a preset second area in response to the operation information indicating the trigger position of the long press operation, and deleting the acquired media content.
[0143] In this way, the embodiments of this disclosure can determine whether to delete the acquired media content based on the trigger position of the long press operation, thereby avoiding the acquisition and sending of invalid media content, ensuring the validity of the acquired media content, and improving the quality of interaction.
[0144] In some embodiments, the type of media content pair is determined based on indicator elements, which include a first indicator element corresponding to the image type and / or a second indicator element corresponding to the video type.
[0145] In this way, the embodiments of this disclosure can select and obtain corresponding types of media content based on different indicator elements, thereby not only enriching the types of media content obtained and improving the interactivity of the interface, but also effectively ensuring the accuracy of the types of media content obtained and ensuring the efficiency of media content acquisition.
[0146] In some embodiments, process 600 further includes: electronic device 110 determining the duration for which operation information satisfies the triggering condition; and determining the type of media content based on a comparison of the duration with a threshold, the type including image type or video type.
[0147] In this way, the embodiments of this disclosure can determine the type of media content based on the duration for which the operation information meets the triggering conditions. This not only enriches the types of acquired media content, but also allows for the conversion of media content types based on duration, thereby enhancing the interactivity of the interface.
[0148] In some embodiments, process 600 further includes: electronic device 110 presenting a first preview content corresponding to media content in response to operation information being updated to not meet triggering conditions.
[0149] In this way, the present embodiment can present a preview of the acquired media content when the operation information is updated to no longer meet the triggering conditions, so as to allow the user to preview it, thereby improving the diversity of the content presented in the interface, enhancing the interactive capabilities based on the acquired media content, and ensuring the effectiveness of the acquired media content.
[0150] In some embodiments, the media content is first media content, and after presenting the first preview content corresponding to the media content, the process 600 further includes: the electronic device 110 responding to the operation information being updated to meet the triggering condition, starting the image capturing unit to acquire the second media content, and sending at least one message further including a third part corresponding to the second media content.
[0151] In this way, the present embodiment can continue to acquire second media content after presenting the preview content corresponding to the first media content, thereby acquiring multiple media content segments during the acquisition process corresponding to an audio content segment, further improving the richness of the acquired content, ensuring the acquisition efficiency of multiple media content segments, and also presenting a third part corresponding to the second media content in the message sent in the conversation interface, thereby further improving the content richness of the message, improving the acquisition efficiency of the message content, ensuring the correlation between the acquired multiple media content segments and the audio content, and improving the interactivity of the conversation interface.
[0152] In some embodiments, the first media content and the second media content correspond to different media types.
[0153] In this way, the embodiments of this disclosure can acquire multiple media content of different types during the acquisition of a piece of audio content, thereby effectively enriching the types of acquired content, improving the content richness of the message, and enhancing the interactivity of the interface.
[0154] In some embodiments, after presenting a first preview of the media content, process 600 further includes: the electronic device 110, in response to the operation information being updated to meet the triggering condition, activating the image capturing unit to acquire image content as an additional segment of the media content.
[0155] In this manner, the embodiments of this disclosure can, after acquiring media content, restart the image capturing unit to acquire additional segments of the media content once the operation information is updated to meet the triggering conditions, so that the additional segments can be presented as part of the media content. Therefore, this solution can acquire different segments in multiple segments and combine them into media content, thereby effectively enriching the methods of acquiring media content, improving the richness of media content, and ensuring the efficiency of media content acquisition.
[0156] In some embodiments, the first part and the second part correspond to different messages; or, the first part and the second part correspond to different parts of the same message.
[0157] In this way, embodiments of this disclosure can present media content and audio content through different messages or the same message, thereby enriching the presentation of media content and audio content in messages, enhancing message diversity, and improving the interactivity of the conversation interface.
[0158] In some embodiments, the first part includes text content corresponding to the audio content.
[0159] In this way, the embodiments of this disclosure can present the text content corresponding to the audio content in the message, so as to present the audio content in the form of text and improve the intuitiveness of the message.
[0160] In some embodiments, the text content is also generated based on media content.
[0161] In this way, the embodiments of this disclosure can generate text content corresponding to audio content by combining media content, effectively avoiding recognition errors caused by polyphonic characters or tone changes, effectively eliminating ambiguity caused by different polyphonic characters, thereby effectively improving the recognition accuracy of audio content and ensuring the accuracy of text content.
[0162] In some embodiments, the session interface is associated with an agent, and process 600 further includes: the electronic device 110 providing response content generated by the agent based on at least one message in the session interface.
[0163] In this way, embodiments of the present disclosure can provide response content generated by the agent in the interactive interface, ensuring the efficiency and accuracy of the response to at least one message, and improving the interactive capabilities of the conversation interface.
[0164] In some embodiments, the response content is also generated based on the temporal relationship between the media content and the audio content, whereby the temporal relationship indicates the association between a first time period corresponding to the media content and a second time period corresponding to the audio content.
[0165] In this way, the embodiments of this disclosure can generate corresponding response content based on the time correspondence between media content and audio content, thereby effectively ensuring the relevance and relevance between the response content and the media content and audio content, ensuring the accuracy and effectiveness of the response content, and improving the interactivity of the interface.
[0166] Example Devices and Equipment
[0167] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. FIG. 7 A schematic structural block diagram of an example device 700 for interface interaction according to certain embodiments of the present disclosure is shown. Device 700 may be implemented as or included in electronic device 110. Various modules / components in device 700 may be implemented by hardware, software, firmware, or any combination thereof.
[0168] like FIG. 7As shown, the device 700 includes: an audio acquisition module 710, a media acquisition module 720, and a message sending module 730. The audio acquisition module 710 is configured to, in response to receiving a first operation in the session interface, activate the audio acquisition unit to acquire audio content; the media acquisition module 720 is configured to, during the audio content acquisition process, in response to detecting that the operation information meets the trigger condition, activate the image capture unit to acquire media content; and the message sending module 730 is configured to, in the session interface, send at least one message, the at least one message including a first part corresponding to the audio content and a second part corresponding to the media content.
[0169] In some embodiments, the first operation includes a long press operation on the voice input control in the conversation interface, and the operation information indicates the trigger position of the long press operation.
[0170] In some embodiments, the media acquisition module 720 is configured to: present an indicator element during the acquisition of audio content; and, in response to the operation information indicating that the trigger position of a long press operation moves to a first area associated with the indicator element, activate the image capture unit to acquire media content.
[0171] In some embodiments, the device 700 further includes a stop acquisition module configured to stop acquiring media content in response to an operation information indicating that the trigger position of a long press operation moves outside the first area.
[0172] In some embodiments, the device 700 further includes a media deletion module configured to delete the acquired media content in response to an operation information indicating that the trigger position of a long press operation moves to a preset second area.
[0173] In some embodiments, the type of media content pair is determined based on indicator elements, which include a first indicator element corresponding to the image type and / or a second indicator element corresponding to the video type.
[0174] In some embodiments, the apparatus 700 further includes a type determination module configured to: determine the duration for which the operation information satisfies the triggering condition; and determine the type of media content based on a comparison of the duration with a threshold, the type including image type or video type.
[0175] In some embodiments, the device 700 further includes a media preview module configured to present a first preview of the media content in response to an update of the operation information indicating that the triggering condition is not met.
[0176] In some embodiments, the media content is first media content, and the device 700 further includes a multiple acquisition module, which is configured to: after presenting the first preview content corresponding to the media content, in response to the operation information being updated to meet the triggering condition, start the image capturing unit to acquire the second media content, and send at least one message that also includes a third part corresponding to the second media content.
[0177] In some embodiments, the first media content and the second media content correspond to different media types.
[0178] In some embodiments, the device 700 further includes a continuous acquisition module configured to: after presenting a first preview of the media content, in response to an operation information update that a trigger condition is met, activate an image capture unit to acquire image content as an additional segment of the media content.
[0179] In some embodiments, the first part and the second part correspond to different messages; or, the first part and the second part correspond to different parts of the same message.
[0180] In some embodiments, the first part includes text content corresponding to the audio content.
[0181] In some embodiments, the text content is also generated based on media content.
[0182] In some embodiments, the conversational interface is associated with an agent, and the device 700 further includes a content response module configured to provide response content generated by the agent based on at least one message in the conversational interface.
[0183] In some embodiments, the response content is also generated based on the temporal relationship between the media content and the audio content, whereby the temporal relationship indicates the association between a first time period corresponding to the media content and a second time period corresponding to the audio content.
[0184] like FIG. 8 As shown, electronic device 800 is in the form of a general-purpose electronic device. Components of electronic device 800 may include, but are not limited to, at least one processor 810 or processing unit, memory 820, storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processor 810 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 800.
[0185] Electronic device 800 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 830 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 800.
[0186] Electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... FIG. 8 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 820 may include computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0187] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 800 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0188] Input device 850 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 860 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 800 can also communicate with one or more external devices (not shown) via communication unit 840 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 800, or with any device that enables electronic device 800 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0189] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0190] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0191] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0192] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0193] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0194] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A user interface interaction method, comprising: In response to receiving the first operation in the session interface, the audio acquisition unit is activated to acquire audio content; During the acquisition of the audio content, in response to the detection that the operation information meets the triggering condition, the image capturing unit is activated to acquire the media content; as well as In the session interface, at least one message is sent, the at least one message including a first part corresponding to the audio content and a second part corresponding to the media content.
2. The method according to claim 1, wherein, The first operation includes a long press operation on the voice input control in the conversation interface, and the operation information indicates the trigger position of the long press operation.
3. The method according to claim 2, wherein, The response to detecting that the operation information meets the triggering condition, and activating the image capturing unit to acquire media content, includes: During the acquisition of the audio content, indicator elements are presented; and In response to the operation information indicating that the trigger position of the long press operation moves to the first area associated with the indicator element, the image capturing unit is activated to acquire the media content.
4. The method according to claim 3, further comprising: In response to the operation information indicating that the trigger position of the long press operation moves outside the first area, the acquisition of the media content is stopped.
5. The method according to claim 4, further comprising: In response to the operation information instructing the trigger position of the long press operation to move to a preset second area, the acquired media content is deleted.
6. The method according to claim 3, wherein, The type of media content is determined based on the indicator elements, which include a first indicator element corresponding to the image type and / or a second indicator element corresponding to the video type.
7. The method according to claim 1, further comprising: Determine the duration for which the operation information satisfies the triggering condition; as well as Based on the comparison between the duration and the threshold, the type of the media content is determined, and the type includes image type or video type.
8. The method according to claim 1, further comprising: In response to the operation information being updated to indicate that the triggering condition is not met, a first preview content corresponding to the media content is presented.
9. The method according to claim 8, wherein, The media content is first media content, and after presenting the first preview content corresponding to the media content, the method further includes: In response to the operation information being updated to meet the triggering condition, the image capturing unit is activated to acquire the second media content, and the at least one message sent further includes a third part corresponding to the second media content.
10. The method according to claim 9, wherein, The first media content and the second media content correspond to different media types.
11. The method according to claim 8, wherein, After presenting the first preview content corresponding to the media content, the method further includes: In response to the operation information being updated to meet the triggering condition, the image capturing unit is activated to acquire image content as an additional segment of the media content.
12. The method according to claim 1, wherein: The first part and the second part correspond to different messages; or The first part and the second part correspond to different parts of the same message.
13. The method according to claim 1, wherein, The first part includes text content corresponding to the audio content.
14. The method according to claim 13, wherein, The text content is also generated based on the media content.
15. The method according to claim 1, wherein, The session interface is associated with the intelligent agent, and the method further includes: The conversation interface provides response content generated by the agent based on the at least one message.
16. The method according to claim 15, wherein, The response content is also generated based on the temporal relationship between the media content and the audio content, wherein the temporal relationship indicates the association between a first time period corresponding to the media content and a second time period corresponding to the audio content.
17. A device for interface interaction, comprising: The audio acquisition module is configured to start the audio acquisition unit to acquire audio content in response to receiving a first operation in the session interface. The media acquisition module is configured to, during the acquisition of the audio content, activate the image capture unit to acquire the media content in response to detecting that the operation information meets the triggering condition; as well as The message sending module is configured to send at least one message in the session interface, the at least one message including a first part corresponding to the audio content and a second part corresponding to the media content.
18. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 16 when executed by the at least one processor.
19. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 16.
20. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 16.