Voice interaction method and device
By integrating voice interaction functions on the media playback pages of short video applications, the problem of users jumping between different pages is solved, efficient and natural voice interaction and personalized content recommendations are achieved, and user experience and interaction efficiency are improved.
Patent Information
- Application Number
- CN202510734366.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-08-26
AI Technical Summary
The voice interaction method in existing short video applications requires users to jump to a dedicated voice assistant page on the video playback page, resulting in cumbersome operations and destroying the user's viewing experience and interaction fluency.
The voice interaction function is integrated on the media playback page. By triggering the start and end events of voice acquisition, voice is collected and user demand types are determined, and feedback results are displayed directly on the media playback page, including the session panel and operation command execution.
It simplifies the voice interaction operation process, improves interaction efficiency and user experience, allows natural language interaction and personalized content recommendations, and enhances the user's sense of participation and convenience of interaction.
Smart Images

Figure CN120544567A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet application technology, and in particular to a voice interaction method and device. Background Art
[0002] With the development of multimedia technology, short video applications have gradually become an indispensable part of people's daily lives. However, current short video application interaction methods are still mainly touch-based, and voice interaction is significantly insufficient. Existing voice interaction solutions typically require users to jump from the video playback page to a dedicated voice assistant page through a separate entrance. This design is not only cumbersome to operate, but also seriously disrupts the user's viewing experience. Users are forced to interrupt their viewing of video content and cross-page operations to achieve basic voice command input, which seriously affects the user's interaction smoothness and interactive experience. Summary of the Invention
[0003] In view of this, the present application provides a voice interaction method and device to simplify the operational process of user voice interaction, thereby improving interaction efficiency and user experience.
[0004] This application provides the following solutions:
[0005] In a first aspect, a voice interaction method is provided, the method comprising:
[0006] Displaying a media playback page, wherein the media playback page plays the first media content;
[0007] In response to an event triggering the start of voice collection, adjusting the playback state of the first media content and then collecting voice;
[0008] In response to an event triggering the end of voice collection, determining a user demand type based on the collected voice;
[0009] Based on the user demand type, the feedback result is displayed on the media playback page.
[0010] Optionally, displaying the feedback result on the media playback page based on the user demand type includes:
[0011] In response to the user demand type being a general demand or a precise demand of a first preset type, a conversation panel is displayed on the media playback page, and feedback results corresponding to the user demand type are displayed based on the conversation panel.
[0012] Optionally, displaying the feedback result on the media playback page based on the user demand type includes:
[0013] In response to the user demand type being a precise demand of the second preset type, a target operation instruction corresponding to the precise demand of the second preset type is determined, the target operation instruction is executed, and the execution result of the target operation instruction is displayed on the media playback page.
[0014] Optionally, if the user demand type is a general demand, displaying the feedback result corresponding to the user demand type based on the conversation panel includes:
[0015] Displaying at least one precise requirement option corresponding to the general requirement on the conversation panel, and in response to an event in which one of the precise requirement options is selected, displaying a feedback result corresponding to the selected precise requirement option on the conversation panel; or
[0016] The conversation panel displays a query text, which is generated based on the general demand. In response to the user further inputting voice or text based on the query text, the conversation panel displays a feedback result for the further input voice or text.
[0017] Optionally, if the user demand type is a precise demand of a first preset type, displaying a feedback result corresponding to the demand type based on the conversation panel includes:
[0018] In response to the precise demand of the first preset type being a dialogue demand, a response text is displayed on the conversation panel, where the response text is generated based on the text corresponding to the voice; or
[0019] In response to the precise requirement of the first preset type being an operation instruction requirement, the operation instruction corresponding to the precise requirement of the first preset type is executed, and the execution result of the operation instruction is displayed on the session panel.
[0020] Optionally, in response to the user demand type being a precise demand of a second preset type, determining a target operation instruction corresponding to the precise demand of the second preset type, executing the target operation instruction, and displaying the execution result of the target operation instruction on the media playback page includes:
[0021] In response to the second preset type of precise requirement being a media content recommendation requirement, determining second media content corresponding to the second preset type of precise requirement, and switching to playing the second media content on the media playback page;
[0022] In response to the second preset type of precise requirement being an operation instruction requirement, the operation instruction corresponding to the second preset type of precise requirement is executed, and the execution result of the operation instruction is displayed on the media playback page.
[0023] Optionally, the determining the second media content corresponding to the precise requirement of the second preset type includes:
[0024] In response to the second preset type of precise requirement being a media recommendation requirement based on a target object or a target attribute, determining second media content to be recommended to the user based on the target object or the target attribute; or
[0025] In response to the second preset type of precise requirement being to play designated media content, the second media content corresponding to the second preset type of precise requirement is determined to be the designated media content.
[0026] Optionally, the adjusting the playback status of the first media content includes at least one of the following:
[0027] Pause the playback of the first media content;
[0028] Lowering the playback volume of the first media content to below a preset threshold;
[0029] The audio output of the first media content is turned off.
[0030] Optionally, after adjusting the playback status of the first media content, the method further includes:
[0031] Prompt information is displayed on the media playback page, where the prompt information is used to prompt the target object to perform voice input.
[0032] Optionally, the method further includes at least one of the following:
[0033] In response to a voice collection end event in the media playback page, restoring the playback state of the first media content to a state before the voice interaction start event;
[0034] During the voice collection process, the voice recognition result corresponding to the collected voice is displayed in a preset area on the media playback page, and the size of the preset area is adjusted as the voice recognition result increases until a preset size limit is reached.
[0035] Optionally, the event that triggers the start of voice collection includes: detecting a long press gesture on the media playback page; the event that triggers the end of voice collection includes: detecting the end of the long press gesture on the media playback page.
[0036] In a second aspect, a voice interaction device is provided, the device comprising:
[0037] a page display unit, configured to display a media playback page, wherein the media playback page plays the first media content;
[0038] a voice collecting unit configured to collect voice after adjusting the playback state of the first media content in response to an event triggering the start of voice collection;
[0039] a type determination unit configured to determine a user demand type based on the collected speech in response to an event triggering the end of speech collection;
[0040] The feedback unit is configured to display feedback results on the media playback page based on the user demand type.
[0041] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the first aspect are implemented.
[0042] In a fourth aspect, an electronic device is provided, including:
[0043] one or more processors; and
[0044] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the first aspects above.
[0045] In a fifth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of any one of the methods described in the first aspect.
[0046] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0047] 1) This embodiment integrates voice interaction functionality into the media playback page. When a voice collection start event is triggered, the playback status of the first media content is adjusted and voice is collected. When a voice collection end event is triggered, the user's request type is determined based on the collected voice. Finally, based on the user's request type, feedback results are displayed on the media playback page. This solution avoids the tedious operation of users jumping between different pages, simplifies the voice interaction process, and improves interaction efficiency and user experience.
[0048] 2) By displaying a conversation panel on the media playback page, the present embodiment effectively processes and provides feedback on both general user requests and specific requests of a first preset type that require detailed interaction. This design not only allows users to interact with the content in natural language without leaving the current media content, but also provides rich feedback information through the conversation panel, such as dialogue responses, multiple options, or further inquiries, greatly improving the smoothness of the user experience and the convenience of interaction.
[0049] 3) The embodiment of the present application can effectively process and feedback the user's precise needs of the second preset type by directly executing the target operation instruction on the media playback page and displaying the execution result, thereby achieving efficient and direct user operation response.
[0050] 4) By displaying multiple precise demand options corresponding to general demands in the conversation panel, the present embodiment allows users to intuitively see different specific options, making it easier to select the demand that meets their needs. This interactive method not only improves the efficiency of users in obtaining accurate information, but also enhances the user's sense of participation and the smoothness of the experience during use. In addition, by displaying the query text in the conversation panel and providing feedback based on the user's further input, the system can gain a deeper understanding of the user's intentions and provide more accurate answers or suggestions, further enhancing the user's interactive experience.
[0051] 5) This embodiment of the present application addresses the precise needs of the first preset type by displaying response text or the results of an operation command in the conversation panel. This allows users to quickly obtain answers to questions or confirm the completion of an operation without leaving the current media content. This approach makes the interaction process more natural and smooth, eliminating the need for users to jump between multiple pages, thereby improving the convenience and efficiency of interaction.
[0052] 6) By directly executing the target operation command on the media playback page and displaying the execution results, the system can quickly respond to the user's precise needs of the second preset type. For media content recommendation needs, the system can quickly filter and play the media content that the user is interested in, saving the user time in searching and selecting. For operation instruction needs, the user can immediately see the operation results, which enhances the intuitiveness and efficiency of the interaction, simplifies the user operation process, and further improves user satisfaction with the media playback page operation.
[0053] 7) When determining the second media content corresponding to the precise need of the second preset type, the embodiments of the present application recommend media content based on analysis of the target object or target attributes, providing users with a more personalized and accurate content recommendation experience. Users do not need to enter specific search terms; the system can recommend relevant media content based on the key information in their instructions. For requests to play specific media content, the system can directly switch to playing the content requested by the user, ensuring that users can quickly obtain the required media and further improving user experience satisfaction.
[0054] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0056] Figure 1 A diagram of the system architecture applicable to the embodiments of the present application;
[0057] Figure 2 Flowchart of the voice interaction method provided in the embodiment of the present application;
[0058] Figure 3 A schematic diagram of the voice interaction function in the media playback page provided in an embodiment of the present application;
[0059] Figure 4 An interactive diagram of a user demand type provided in an embodiment of the present application as general demand;
[0060] Figure 5 Another interaction diagram provided in an embodiment of the present application where the user demand type is general demand;
[0061] Figure 6 An interactive diagram of a user requirement type provided in an embodiment of the present application being a precise requirement of a first preset type;
[0062] Figure 7 Another interactive schematic diagram of an embodiment of the present application in which the user requirement type is a precise requirement of a first preset type;
[0063] Figure 8 An interactive diagram of a user requirement type provided in an embodiment of the present application being a precise requirement of a second preset type;
[0064] Figure 9 Another interactive diagram of an embodiment of the present application in which the user requirement type is a precise requirement of a second preset type;
[0065] Figure 10 A schematic block diagram of a voice interaction device provided in an embodiment of the present application;
[0066] Figure 11 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0067] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0068] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0069] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0070] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0071] Existing voice interaction solutions typically require users to navigate from the video playback page to a dedicated voice assistant page through a separate portal. This design is not only cumbersome but also severely disrupts the user's viewing experience. Users are forced to interrupt their viewing experience and navigate across multiple pages to input basic voice commands, severely impacting the smoothness and user experience of the interaction.
[0072] In view of this, the present application provides a new approach. To facilitate understanding of the present application, the system architecture on which the present application is based is first described. Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown. Figure 1 As shown in , the system architecture may include: a user end, a terminal device and a server end.
[0073] Among them, the user end is set on the terminal device. The user end involved in the embodiment of the present application can be a client running on the terminal device, a small program or a Web application running through a browser, etc.
[0074] Terminal devices may include, but are not limited to, smart mobile terminals, wearable devices, PCs (Personal Computers), smart home devices, etc. Smart mobile devices may include mobile phones, tablets, laptops, PDAs (Personal Digital Assistants), and internet car terminals. Wearable devices may include smart watches, smart glasses, smart bracelets, VR (Virtual Reality) devices, AR (Augmented Reality) devices, mixed reality devices (i.e., devices that support both virtual reality and augmented reality), etc. Smart home devices may include smart TVs and smart refrigerators with displays.
[0075] The client can interact with the server through the network and obtain media content from the server.
[0076] The server side can be the backend server for the client, providing backend services to the client, such as delivering media content. The server side can be a single server, a server cluster consisting of multiple servers, or even a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a hosting product within the cloud computing service ecosystem. It addresses the management difficulties and limited scalability of traditional physical hosting and virtual private server (VPS) services.
[0077] As one of the feasible ways, the user terminal watches the media content published by others through the server terminal, and the user terminal uses the method provided in the embodiment of the present application to perform voice interaction on the media playback page.
[0078] It should be understood that Figure 1 The number of user terminals, terminal devices and server terminals in the embodiment is merely illustrative. Any number of user terminals, terminal devices and server terminals may be provided according to implementation requirements.
[0079] Figure 2 This is a flow chart of the voice interaction method provided in the embodiment of the present application. The method can be performed by Figure 1 The user side of the system shown is executed. Figure 2 As shown in , the method may include the following steps:
[0080] Step 201: Display a media playback page, which plays first media content.
[0081] Step 202: In response to the event that triggers the start of voice collection, the playback state of the first media content is adjusted, and then voice is collected.
[0082] Step 203 : In response to the event triggering the end of voice collection, the user demand type is determined based on the collected voice.
[0083] Step 204: Based on the user's requirement type, the feedback result is displayed on the media playback page.
[0084] As can be seen, the embodiment of the present application integrates a voice interaction function into the media playback page. When the voice collection start event is triggered, the playback status of the first media content is adjusted and the voice is collected. When the voice collection end event is triggered, the user's demand type is determined based on the collected voice. Finally, based on the user's demand type, the feedback result is displayed on the media playback page. This solution avoids the tedious operation of users jumping between different pages, simplifies the operation process of voice interaction, and improves interaction efficiency and user experience.
[0085] The following describes in detail each step of the above process and the effects that can be further produced, in conjunction with the embodiments. It should be noted that the terms "first" and "second" in this disclosure do not have limitations on size, order, or quantity, and are merely used to distinguish between them in name. For example, "first preset type" and "second preset type" are used to distinguish between two preset types in name. For another example, "first media content" and "second media content" are used to distinguish between two media contents in name. And so on.
[0086] First, the above step 201 , namely “displaying a media playback page, and playing the first media content on the media playback page”, is described in detail with reference to an embodiment.
[0087] In the embodiment of the present application, the first media content includes but is not limited to videos, pictures in an album, etc.
[0088] Videos can be either long or short. The distinction between long and short videos is mainly reflected in the length and content of the video. Long videos are usually more than ten minutes long; the content is relatively complete and in-depth, and the content quality is relatively high. For example, movies, TV series, documentaries, etc. are all long videos. Short videos are usually less than ten minutes or a few minutes long; the content is usually brief and concise, focusing on highlighting a certain highlight or quickly conveying information. Short videos focus more on rapid production and rapid dissemination, emphasizing timeliness; for example, some funny videos, life recording videos, advertising videos, short dramas, etc. are all short videos.
[0089] In addition, the above videos can be either horizontal or vertical. Horizontal videos are videos where the width of the video is greater than its height, which better fits the natural field of view of the human eye and is more suitable for viewing on wide-screen devices (such as TVs and computers). Vertical videos are videos where the height of the video is greater than its width, which is more suitable for viewing on narrow-screen devices (such as smartphones and tablets) and is more popular in areas such as social media and short videos.
[0090] Images in an album can be either regular or long. Regular images are typically single, static images, allowing users to view their entire content without any additional interaction. Long images, on the other hand, are a special type of image, consisting of a series of images. Users must scroll or swipe to view the entire image. Long images typically combine multiple image elements to form a coherent content.
[0091] It should be noted here that the “playback” involved in the embodiments of the present application refers to converting the content of the work into images, sounds, or a combination of the two that can be directly perceived by humans.
[0092] The above step 202 , ie, “collecting voice after adjusting the playback state of the first media content in response to the event that triggers the start of voice collection”, is described in detail with reference to the embodiment.
[0093] In an embodiment of the present application, the event for starting voice collection can be detecting a long press gesture on the media playback page. Specifically, a voice interaction control can be added to the media playback page. When the user's finger long presses on the voice interaction control for a preset duration (such as 1 or 2 seconds), the system will determine that the event for starting voice collection has been triggered; or when the user's finger long presses on a specific area on the screen (such as the media playback area) for a preset duration, the system will determine that the event for starting voice collection has been triggered.
[0094] In addition to the long press gesture, the event of triggering the start of voice collection can also be when other preset gestures such as sliding and staying, three-finger clicking, etc. are detected on the media playback page. This embodiment of the present application is not limited to this.
[0095] After triggering this event, in order to ensure the quality of voice collection and the consistency of user experience, the system will adjust the playback status of the first media content currently being played. For example, the volume of the first media content is automatically lowered to below the preset threshold so that the user's voice input can be collected more clearly, while avoiding excessive media volume interfering with the user's voice. For another example, pausing the playback of the first media content can completely eliminate the impact of media playback sound on voice collection, making voice collection more accurate; for another example, turning off the audio output of the first media content can completely avoid the interference of audio on voice collection without affecting video playback, ensuring the accuracy of voice recognition and the user's interactive experience.
[0096] After completing the above playback status adjustment, activate the voice collection function and start collecting the user's voice input to ensure the efficiency and accuracy of voice interaction.
[0097] As an example, Figure 3 As shown. Among them, Figure 3 (a) shows a media playback page for playing the first media content. A voice interaction control is added at the bottom of the media playback page. The voice interaction control is displayed in the form of a button to facilitate user operation. When the user long presses the voice interaction control, the event of starting voice collection is triggered, and the playback status of the first media content is adjusted to optimize the voice input environment, and then voice collection begins. Figure 3 As shown in (b), during voice capture, a progress bar displays the capture status on the media playback page. As capture progresses, the bar grows in length. This visual feedback helps users intuitively understand the current capture progress. Once the bar reaches the preset upper limit, capture stops. This design not only enhances the user experience during voice interaction but also ensures the accuracy and efficiency of voice capture.
[0098] Optionally, after completing the above-mentioned playback state adjustment, a prompt message can be displayed on the media playback page, which is used to remind the target object that voice input is available. This not only enhances the user's operation guidance, but also improves the friendliness of interaction and user experience.
[0099] The prompt message may be in text or voice form, and may be in dynamic or static form. The embodiment of the present application does not impose any specific limitation on its display form.
[0100] Still refer to Figure 3 ,in, Figure 3(b) shows a text prompt message "I'm listening, please speak" and a microphone icon displayed on the media playback page to intuitively inform the user that voice input is currently available. At this time, to avoid cluttering the media playback page, other page elements except the first media content can be hidden, and only page elements related to the voice interaction function can be displayed.
[0101] In addition, during the process of collecting voice, the voice recognition results corresponding to the collected voice can be further displayed in a preset area on the media playback page. The preset area can be resized as the voice recognition results grow until it reaches a preset size limit.
[0102] The preset area can be a text box with a small initial size. As voice recognition results are added, the text box gradually expands until it reaches the maximum size limit. This design not only provides real-time feedback on the user's voice input content, but also ensures the aesthetics and rationality of the page layout.
[0103] Still refer to Figure 3 ,in, Figure 3 (c) and Figure 3 (d) shows the process of displaying the speech recognition results in the preset area of the media playback page. Figure 3 As shown in (c), the preset area is the initial size, and Figure 3 As shown in (d), the preset area is the maximum limit size.
[0104] The above step 203 , namely “determining the user demand type based on the collected voice in response to the event triggering the end of voice collection”, is described in detail with reference to the embodiment.
[0105] In an embodiment of the present application, the event that triggers the end of voice collection can be the detection of the end of a long-press gesture on the media playback page, that is, the user releases the long-pressed finger, which is the event that triggers the end of voice collection. For example, the user releases the finger that was long-pressed on the voice interaction control. In addition, the event that triggers the end of voice collection can also be the detection of a pause in user voice input exceeding a preset duration (such as 3 seconds or 5 seconds). The embodiment of the present application does not limit the specific triggering method.
[0106] Still refer to Figure 3 , when the user cancels Figure 3 Long press the voice interaction control shown in (a) to (d) to stop voice collection.
[0107] Furthermore, when the voice collection completion event is triggered, since there is no interference at this point, the playback state of the first media content can be restored to the state before the voice interaction started. This ensures that after the voice interaction is completed, the user can seamlessly continue watching or listening to the previous media content, without missing any content due to adjustments during the voice interaction. This process improves the consistency of the user experience and makes the user feel a more natural and smooth transition when using the voice interaction function.
[0108] After the event triggering the end of voice collection, the collected voice is processed and analyzed to determine the type of user needs. Specifically, the collected voice information can be converted into text information, and then the text can be semantically understood and the intent can be identified through natural language processing technology. For example, a pre-trained natural language processing model can be used to achieve semantic understanding and intent recognition of text. This means that the system uses deep learning technology and is trained based on a large amount of text data, so that it can understand the semantic information in the text and the user's potential intentions.
[0109] The above step 204 , namely “displaying the feedback result on the media playback page based on the user demand type”, is described in detail with reference to the embodiment.
[0110] In an embodiment of the present application, user demand types may include general demands and precise demands of the user. Among them, general demands refer to vague or broad intentions expressed by the user, which usually do not have a clear specific direction or contain multiple possible interpretations, such as "Hello" or "Adjust the volume". Such demands require further clarification or provide multiple options for users to choose from. In this case, a session panel can be displayed on the media playback page, and feedback results corresponding to the user demand type can be displayed based on the session panel.
[0111] Precise demands, on the other hand, are clear and specific intentions expressed by users, which can usually be executed or responded to directly. Taking into account the different implementation methods of direct execution or response, the precise demands in the embodiments of the present application can be divided into precise demands of a first preset type and precise demands of a second preset type. Among them, the precise demands of the first preset type correspond to scenarios where feedback results need to be displayed through a session panel. In this scenario, the session panel is displayed on the media playback page, and the corresponding feedback results are displayed based on the session panel. For example, when a user asks "Who is the leading actor in this movie", the answer can be displayed in the session panel. The precise demands of the second preset type correspond to scenarios where feedback results do not need to be displayed through a session panel. In this scenario, the target operation instruction corresponding to the precise demand of the second preset type is determined, the target operation instruction is executed, and the execution result of the target operation instruction is displayed on the media playback page. For example, if a user issues a "raise the volume" instruction, the volume-raising operation is directly executed and the operation result is displayed on the media playback page.
[0112] By displaying a conversation panel on the media playback page, the above approach effectively addresses and responds to both general user requests and specific requests of the first preset type that require detailed interaction. This design not only allows users to interact with the content in natural language without leaving the current media content, but also provides rich feedback information through the conversation panel, such as dialogue responses, multiple options, or further inquiries, greatly improving the smoothness of the user experience and the convenience of interaction.
[0113] By directly executing the target operation instructions on the media playback page and displaying the execution results, the user's precise needs of the second preset type can be effectively processed and fed back, achieving efficient and direct user operation response. This processing method is particularly suitable for operations that do not require complex interactions to complete, such as volume adjustment, media playback control, etc., allowing users to quickly see the operation results, enhancing the intuitiveness and efficiency of the interaction, while also simplifying the user operation process, further improving user satisfaction with media playback page operations.
[0114] Furthermore, if the user demand type is identified as a general demand, the feedback result corresponding to the general demand is displayed on the media playback page.
[0115] As an implementable method, if the user demand type is identified as a general demand, at least one precise demand option corresponding to the general demand is displayed in the session panel. In response to the event that one of the precise demand options is selected, the feedback result corresponding to the selected precise demand option is displayed in the session panel.
[0116] That is, multiple precise demand options corresponding to the general demand are displayed in the conversation panel, and the user can select the one that best meets his or her intention from these options, and then the specific feedback result of the option is displayed in the conversation panel.
[0117] As an example, Figure 4 As shown. Among them, Figure 4 (a) shows that the voice recognition result of "I want to watch cat videos" is displayed in the preset area of the media playback page. When the voice recognition result is identified as a general demand, Figure 4 (b) shows a conversation panel, which displays multiple precise options, including "Funny Cat Videos", "Cute Cat Daily Life", and "Cat Science Videos". When the user selects "Funny Cat Videos", Figure 4 As shown in (c), the corresponding funny cat recommendation list is displayed in the conversation panel. When the user selects "Funny Cat Video 1" from the recommendation list, Figure 4 As shown in (d), the media playback page switches to playing the video selected by the user.
[0118] As another feasible method, if the user demand type is identified as a general demand, a query text is displayed in the conversation panel. The query text is generated based on the general demand. In response to the user's further input of voice or text based on the query text, the feedback result for the further input voice or text is displayed in the conversation panel.
[0119] Specifically, a query text can be displayed in the conversation panel based on the user's general needs. This query text is generated based on the user's general needs and is intended to further clarify the user's intention. The user can then provide a more specific response based on this query text via voice or text. The system will then display the corresponding feedback results in real time in the conversation panel based on the user's further input.
[0120] Still refer to Figure 4 ,when Figure 4 When the speech recognition result in (a) is identified as a general demand, Figure 4 In (e), a conversation panel is displayed, and the query text "Do you want to watch funny cat videos, popular science cat videos, or other types of cat videos?" is displayed in the conversation panel. The user answers "funny cat videos". When the user selects "funny cat videos", Figure 4 As shown in (c), the corresponding funny cat recommendation list is displayed in the conversation panel. When the user selects one from the recommendation list, Figure 4 As shown in (d), the media playback page switches to playing the video selected by the user.
[0121] For example, Figure 5 As shown, Figure 5 (a) shows that the voice recognition result of "Hello" is displayed in the preset area of the media playback page. When the voice recognition result is identified as a general demand, Figure 5 (b) displays a conversation panel, in which the inquiry text "Hello, how can I help you today" is displayed to further clarify the user's intention.
[0122] The above-mentioned interaction method can not only improve the efficiency of users in obtaining accurate information, but also enhance the user's sense of participation and the smoothness of the experience during use.
[0123] Furthermore, if the user demand type is a precise demand of the first preset type, the feedback result corresponding to the demand type is displayed based on the conversation panel.
[0124] As an implementable manner, in response to the precise demand of the first preset type being a dialogue demand, a response text is displayed on the conversation panel, where the response text is generated based on the text corresponding to the voice.
[0125] That is, when a user asks a question that needs an answer or engages in a conversation, the corresponding text is generated based on the user's voice input and the answer is displayed in the conversation panel.
[0126] As an example, Figure 6 As shown, Figure 6 (a) shows the user's voice input of "Who is the main actor of this movie" in the media playback page, which is recognized as the precise demand of the first preset type. Figure 6 In (b), a conversation panel is displayed, and the answer text "The starring actors of this movie are actor A and actor B" is displayed in it, which satisfies the user's query requirements in sequence.
[0127] As another achievable manner, in response to the precise requirement of the first preset type being an operation instruction requirement, the operation instruction corresponding to the precise requirement of the first preset type is executed, and the execution result of the operation instruction is displayed on the session panel.
[0128] That is, when the user issues a specific operation instruction, the instruction is executed and the execution result is fed back in the session panel.
[0129] As an example, Figure 7 As shown, Figure 7 (a) shows that the user issues a voice command of "lower volume" in the media playback page. When it is recognized as a precise demand of the first preset type, the volume lowering operation is then performed, and Figure 7 In (b), the conversation panel is displayed, which shows the feedback information of "volume has been lowered" to ensure that the user understands the current operation status.
[0130] Optionally, in response to the user demand type being a precise demand of the second preset type, determining a target operation instruction corresponding to the precise demand of the second preset type, executing the target operation instruction, and displaying the execution result of the target operation instruction on the media playback page can be implemented in the following manner:
[0131] In response to the precise demand of the second preset type being a media content recommendation demand, determining second media content corresponding to the precise demand of the second preset type, and switching to playing the second media content on the media playback page;
[0132] In response to the precise requirement of the second preset type being an operation instruction requirement, the operation instruction corresponding to the precise requirement of the second preset type is executed, and the execution result of the operation instruction is displayed on the media playback page.
[0133] That is, when the user's request type is identified as a precise request of the second preset type, different response methods can be adopted according to the type of request. If the request is a media content recommendation request, the corresponding second media content is determined and switched to play on the media playback page. If it is an operation instruction request, the instruction is executed and the execution result is displayed on the media playback page.
[0134] As an example, Figure 8 As shown, Figure 8 (a) shows that the user issues a voice command "I want to watch cat videos" on the media playback page. When it is recognized as the second preset type of precise demand - media content recommendation demand, the system can filter out one or more cat videos according to the preset recommendation algorithm, and then Figure 8 As shown in (b), the screened cat videos are played sequentially on the media playback page.
[0135] As another example, Figure 9 As shown, Figure 9 (a) shows that the user issues a voice command of "lower volume" on the media playback page. When it is recognized as a precise demand of the second preset type - an operation command demand, the operation of lowering the volume is directly executed, and a prompt message of "Volume Lowered" is displayed on the media playback page.
[0136] Furthermore, when determining the second media content corresponding to the precise requirement of the second preset type, in response to the precise requirement of the second preset type being a media recommendation requirement based on a target object or a target attribute, the second media content recommended to the user is determined based on the target object or the target attribute; or, in response to the precise requirement of the second preset type being to play specified media content, the second media content corresponding to the precise requirement of the second preset type is determined to be the specified media content.
[0137] That is to say, when the user issues a precise demand of the second preset type, different processing methods can be adopted according to the specific content of the demand. If the demand is a media recommendation request based on a target object or target attribute, the recommended media content is determined by analyzing these target features. For example, if the user says "I want to watch a science fiction movie", relevant movies will be screened out and recommended based on the attribute "science fiction". On the contrary, if it is a demand to play specified media content, such as the user explicitly says "play "Wolf Warrior 2", "Wolf Warrior 2" will be directly played as the target media content. This processing logic ensures that the system can accurately respond to different types of user requests, provide personalized media recommendations or directly play specified content.
[0138] It should be noted that in the above description, when the user issues a voice command "I want to watch cat videos" on the media playback page, it can be identified as a general demand (refer to Figure 4 ), can also be identified as a precise requirement of the second preset type (refer to Figure 8 ), because the semantic clarity of the instruction depends on the specific context and the training method of the AI model, and the training method of the AI model depends on the specific needs of the actual application.
[0139] If, in actual applications, the desire is greater to display feedback results in the conversation panel, the AI model will be trained with this goal in mind, identifying the instruction as a general demand so that the user's specific needs can be further clarified in the conversation panel. Conversely, if the actual application focuses more on quickly responding to user media content requests and displaying the results directly on the media playback page, the AI model will be trained to focus on identifying such instructions as a second preset type of precise demand, allowing for direct media content recommendation and playback.
[0140] If, in actual applications, there is a greater desire to display the results of operations directly on the media playback page to improve the intuitiveness and efficiency of the interaction, the AI model training will focus on identifying such instructions as precise requirements of the second preset type, so that the operation can be directly executed and the results can be displayed. On the contrary, if the actual application focuses more on providing detailed operation feedback and interaction records through the conversation panel, the AI model training will focus on identifying such instructions as precise requirements of the first preset type. In this way, users can see feedback information such as "volume has been lowered" in the conversation panel, and can also conduct further voice or text interaction, ensuring the transparency of the operation and the user's right to know.
[0141] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0142] According to another embodiment, a voice interaction device is provided. Figure 10 FIG. 1 shows a schematic block diagram of the voice interaction device according to an embodiment. Figure 10 As shown, the device 1000 includes: a display unit 701 and a first interaction unit 702. The main functions of each component unit are as follows:
[0143] The page display unit 1001 is configured to display a media playback page, wherein the media playback page plays the first media content;
[0144] The voice collecting unit 1002 is configured to collect voice after adjusting the playback state of the first media content in response to an event triggering the start of voice collection;
[0145] The type determination unit 1003 is configured to determine the type of user demand based on the collected voice in response to the event triggering the end of voice collection;
[0146] The feedback unit 1004 is configured to display the feedback result on the media playback page based on the user demand type.
[0147] Optionally, the feedback unit 1004 is configured to:
[0148] In response to the user demand type being a general demand or a precise demand of the first preset type, a conversation panel is displayed on the media playback page, and feedback results corresponding to the user demand type are displayed based on the conversation panel.
[0149] Optionally, the feedback unit 1004 is configured to:
[0150] In response to the user demand type being a precise demand of the second preset type, a target operation instruction corresponding to the precise demand of the second preset type is determined, the target operation instruction is executed, and the execution result of the target operation instruction is displayed on the media playback page.
[0151] Optionally, if the user demand type is a general demand, the feedback unit 1004 displays the feedback result corresponding to the user demand type based on the conversation panel, and is specifically configured to:
[0152] Displaying at least one precise requirement option corresponding to the general requirement on the conversation panel, and in response to an event in which one of the precise requirement options is selected, displaying a feedback result corresponding to the selected precise requirement option on the conversation panel; or
[0153] The conversation panel displays a query text, which is generated based on the general demand. In response to the user further inputting voice or text based on the query text, the conversation panel displays a feedback result for the further input voice or text.
[0154] Optionally, if the user demand type is a precise demand of a first preset type, the feedback unit 1004 displays the feedback result corresponding to the demand type based on the conversation panel, and is configured to:
[0155] In response to the precise demand of the first preset type being a dialogue demand, a response text is displayed on the conversation panel, where the response text is generated based on the text corresponding to the voice; or
[0156] In response to the precise requirement of the first preset type being an operation instruction requirement, the operation instruction corresponding to the precise requirement of the first preset type is executed, and the execution result of the operation instruction is displayed on the session panel.
[0157] Optionally, in response to the user demand type being a precise demand of a second preset type, the feedback unit 1004 determines a target operation instruction corresponding to the precise demand of the second preset type, executes the target operation instruction, and displays the execution result of the target operation instruction on the media playback page, and is configured to:
[0158] In response to the second preset type of precise requirement being a media content recommendation requirement, determining second media content corresponding to the second preset type of precise requirement, and switching to playing the second media content on the media playback page;
[0159] In response to the second preset type of precise requirement being an operation instruction requirement, the operation instruction corresponding to the second preset type of precise requirement is executed, and the execution result of the operation instruction is displayed on the media playback page.
[0160] Optionally, the feedback unit 1004 determines the second media content corresponding to the precise requirement of the second preset type, and is configured to:
[0161] In response to the second preset type of precise requirement being a media recommendation requirement based on a target object or a target attribute, determining second media content to be recommended to the user based on the target object or the target attribute; or
[0162] In response to the second preset type of precise requirement being to play designated media content, the second media content corresponding to the second preset type of precise requirement is determined to be the designated media content.
[0163] Optionally, the voice collection unit 1002 adjusts the playback status of the first media content by at least one of the following:
[0164] Pause the playback of the first media content;
[0165] Lowering the playback volume of the first media content to below a preset threshold;
[0166] The audio output of the first media content is turned off.
[0167] Optionally, after adjusting the playback status of the first media content, the voice collection unit 1002 is further configured to:
[0168] Prompt information is displayed on the media playback page, where the prompt information is used to prompt the target object to perform voice input.
[0169] Optionally, the voice collection unit 1002 further includes at least one of the following configurations:
[0170] In response to a voice collection end event in the media playback page, restoring the playback state of the first media content to a state before the voice interaction start event;
[0171] During the voice collection process, the voice recognition result corresponding to the collected voice is displayed in a preset area on the media playback page, and the size of the preset area is adjusted as the voice recognition result increases until a preset size limit is reached.
[0172] Optionally, the event triggering the start of voice collection includes: detecting a long press gesture on the media playback page;
[0173] The event triggering the end of voice collection includes: detecting the end of the long press gesture on the media playback page.
[0174] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or device embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0175] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0176] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.
[0177] And an electronic device comprising:
[0178] one or more processors; and
[0179] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.
[0180] The present application also provides a computer program product, comprising a computer program, which implements the steps of any one of the methods described in the aforementioned method embodiments when executed by a processor.
[0181] in, Figure 11 The electronic device architecture is shown as an example, and may include a processor 1110, a video display adapter 1111, a disk drive 1112, an input / output interface 1113, a network interface 1114, and a memory 1120. The processor 1110, the video display adapter 1111, the disk drive 1112, the input / output interface 1113, the network interface 1114, and the memory 1120 may be communicatively connected via a communication bus 1130.
[0182] Among them, the processor 1110 can be implemented by a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solutions provided in this application.
[0183] The memory 1120 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1120 can store an operating system 1121 for controlling the operation of the electronic device, and a basic input and output system (BIOS) 1122 for controlling the low-level operation of the electronic device. In addition, a web browser 1123, a data storage management system 1124, and a voice interaction device 1000, etc. can also be stored. The above-mentioned voice interaction device 1000 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided in this application is implemented by software or firmware, the relevant program code is stored in the memory 1120 and is called and executed by the processor 1110.
[0184] The input / output interface 1113 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0185] The network interface 1114 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WIFI, Bluetooth, etc.).
[0186] The bus 1130 comprises a pathway for transmitting information between the various components of the device (eg, the processor 1110 , the video display adapter 1111 , the disk drive 1112 , the input / output interface 1113 , the network interface 1114 , and the memory 1120 ).
[0187] It should be noted that although the above device only shows the processor 1110, video display adapter 1111, disk drive 1112, input / output interface 1113, network interface 1114, memory 1120, bus 1130, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may also include only the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.
[0188] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer program product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0189] The above is a detailed introduction to the technical solutions provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting this application.
Claims
1. A voice interaction method, characterized in that: The method comprises: Displaying a media playback page, wherein the media playback page plays the first media content; In response to an event triggering the start of voice collection, adjusting the playback state of the first media content and then collecting voice; In response to an event triggering the end of voice collection, determining a user demand type based on the collected voice; Based on the user demand type, the feedback result is displayed on the media playback page.
2. The method according to claim 1, characterized in that The displaying of the feedback result on the media playback page based on the user demand type includes: In response to the user demand type being a general demand or a precise demand of a first preset type, a conversation panel is displayed on the media playback page, and feedback results corresponding to the user demand type are displayed based on the conversation panel.
3. The method according to claim 1, characterized in that The displaying of the feedback result on the media playback page based on the user demand type includes: In response to the user demand type being a precise demand of the second preset type, a target operation instruction corresponding to the precise demand of the second preset type is determined, the target operation instruction is executed, and the execution result of the target operation instruction is displayed on the media playback page.
4. The method according to claim 2, characterized in that If the user demand type is a general demand, displaying the feedback result corresponding to the user demand type based on the conversation panel includes: Displaying at least one precise requirement option corresponding to the general requirement on the conversation panel, and in response to an event in which one of the precise requirement options is selected, displaying a feedback result corresponding to the selected precise requirement option on the conversation panel; or The conversation panel displays a query text, which is generated based on the general demand. In response to the user further inputting voice or text based on the query text, the conversation panel displays a feedback result for the further input voice or text.
5. The method according to claim 2, characterized in that If the user demand type is a precise demand of the first preset type, displaying feedback results corresponding to the demand type based on the conversation panel includes: In response to the precise demand of the first preset type being a dialogue demand, a response text is displayed on the conversation panel, where the response text is generated based on the text corresponding to the voice; or In response to the precise requirement of the first preset type being an operation instruction requirement, the operation instruction corresponding to the precise requirement of the first preset type is executed, and the execution result of the operation instruction is displayed on the session panel.
6. The method according to claim 3, characterized in that In response to the user demand type being a precise demand of a second preset type, determining a target operation instruction corresponding to the precise demand of the second preset type, executing the target operation instruction, and displaying the execution result of the target operation instruction on the media playback page, includes: In response to the second preset type of precise requirement being a media content recommendation requirement, determining second media content corresponding to the second preset type of precise requirement, and switching to playing the second media content on the media playback page; In response to the second preset type of precise requirement being an operation instruction requirement, the operation instruction corresponding to the second preset type of precise requirement is executed, and the execution result of the operation instruction is displayed on the media playback page.
7. The method according to claim 6, characterized in that The determining of the second media content corresponding to the precise requirement of the second preset type includes: In response to the second preset type of precise requirement being a media recommendation requirement based on a target object or a target attribute, determining second media content to be recommended to the user based on the target object or the target attribute; or In response to the second preset type of precise requirement being to play designated media content, the second media content corresponding to the second preset type of precise requirement is determined to be the designated media content.
8. The method according to any one of claims 1 to 7, characterized in that The adjusting the playback status of the first media content includes at least one of the following: Pause the playback of the first media content; Lowering the playback volume of the first media content to below a preset threshold; The audio output of the first media content is turned off.
9. The method according to any one of claims 1 to 7, characterized in that After adjusting the playback status of the first media content, the method further includes: Prompt information is displayed on the media playback page, where the prompt information is used to prompt the target object to perform voice input.
10. The method according to any one of claims 1 to 7, characterized in that The method further comprises at least one of the following: In response to a voice collection end event in the media playback page, restoring the playback state of the first media content to a state before the voice interaction start event; During the voice collection process, the voice recognition result corresponding to the collected voice is displayed in a preset area on the media playback page, and the size of the preset area is adjusted as the voice recognition result increases until a preset size limit is reached.
11. The method according to any one of claims 1 to 7, characterized in that The event that triggers the start of voice collection includes: detecting a long press gesture on the media playback page; The event triggering the end of voice collection includes: detecting the end of the long press gesture on the media playback page.
12. An interactive device, characterized in that: The device comprises: a page display unit, configured to display a media playback page, wherein the media playback page plays the first media content; a voice collecting unit configured to collect voice after adjusting the playback state of the first media content in response to an event triggering the start of voice collection; a type determination unit configured to determine a user demand type based on the collected speech in response to an event triggering the end of speech collection; The feedback unit is configured to display feedback results on the media playback page based on the user demand type.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 10.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Voice interaction method and device based on artificial intelligence
CN106941000A
Intelligent automated assistant in a media environment
CN107577385A
Method and device for providing voice service
CN107833574A
Voice control method and electronic equipment
CN113488042A
Speech recognition method and electronic equipment
CN116110400A