Interaction method and apparatus, device and medium
Through semantic analysis of speech information and machine learning model to determine the search results, directly display video information with high matching degree, solving the problem of low interaction efficiency in the existing technology, and achieving faster response and more efficient human-computer interaction.
Patent Information
- Application Number
- PCT/CN2024/144556
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-03
- Filing Date
- 2024-12-31
- Publication Date
- 2025-07-10
AI Technical Summary
The existing voice interaction methods are less efficient in interaction between users and search engines, resulting in longer waiting times and lower interaction efficiency, especially when users are eager to get answers to questions.
By obtaining the user's voice information, semantic analysis and/or text recognition technology is used to determine the content type of the voice information as the search type, and using machine learning models to determine the video information with high matching degree in the video repository, directly displaying the search results in the dialogue interaction page, reducing the dependence on predetermined keywords.
It achieves faster response time and improves the efficiency of human-computer interaction. Users can directly express their search requirements and obtain relevant video information in real time, reducing operational complexity.
Smart Images

Figure CN2024144556_10072025_PF_FP_ABST
Abstract
Description
Method, apparatus, device and medium for interaction
[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, equipment and media for interaction” and application number 2024100160979, filed on January 3, 2024, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] Exemplary implementations of the present disclosure generally relate to interaction processing, and more particularly, to methods, apparatuses, devices, and computer-readable storage media for voice interaction. Background Art
[0003] Currently, various user interaction technologies have been proposed. For example, users can perform human-computer interaction through voice, text, and gestures. In the field of voice interaction, it is already possible to execute corresponding actions and provide corresponding results based on the user's voice input. However, the performance of existing interaction methods is not satisfactory, and therefore it is desired to provide more convenient and effective interaction technology solutions. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for interaction is provided. In this method, a voice message is obtained. If the content type of the voice message is a search type, the search information is determined based on the voice message. Based on the search information, a search result set corresponding to the search information is determined, wherein the search result set includes at least one video. A conversation interaction page is displayed, wherein the conversation interaction page displays at least one video from the search result set.
[0005] In a second aspect of the present disclosure, a device for interaction is provided. The device includes: an acquisition module configured to acquire voice information; an information determination module configured to determine search information based on the voice information if the content type of the voice information is a search type; a result determination module configured to determine a search result set corresponding to the search information based on the search information, wherein the search result set includes at least one video; and a display module configured to display a dialogue interaction page, wherein the dialogue interaction page displays at least one video from the search result set.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor implements the method according to the first aspect of the present disclosure.
[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the implementation of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of various implementations of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] FIG1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;
[0011] FIG2 illustrates a block diagram of a process for interaction according to some implementations of the present disclosure;
[0012] 3 illustrates a block diagram of a process for determining the content type of voice information according to some implementations of the present disclosure;
[0013] FIG4 illustrates a block diagram for determining a search result set according to some implementations of the present disclosure;
[0014] FIG5 shows a block diagram of a page for displaying search results according to some implementations of the present disclosure;
[0015] 6A and 6B respectively illustrate block diagrams for displaying video information according to some implementations of the present disclosure;
[0016] FIG7 shows a block diagram for displaying more video information according to some implementations of the present disclosure;
[0017] 8A and 8B respectively illustrate block diagrams for loading video information according to some implementations of the present disclosure;
[0018] FIG9 shows a flowchart of a method for interaction according to some implementations of the present disclosure;
[0019] FIG10 shows a block diagram of an apparatus for interaction according to some implementations of the present disclosure; and
[0020] FIG11 shows a block diagram of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION
[0021] The following describes implementations of the present disclosure in more detail with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations described herein. Rather, these implementations are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0022] In the description of the implementation of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "an implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". The following may also include other explicit and implicit definitions. As used herein, the term "model" can represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on a variety of technical solutions currently known and / or to be developed in the future.
[0023] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0024] It is understandable that before using the technical solutions disclosed in each implementation of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0025] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0026] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0028] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, a subsequent action may be executed immediately upon the occurrence of the event or the satisfaction of the condition; in other cases, the subsequent action may be executed some time after the occurrence of the event or the satisfaction of the condition.
[0029] Sample Environment
[0030] Currently, a variety of user interaction technology solutions have been proposed. For example, users can perform human-computer interaction via voice, text, and actions. Below, more details about the interaction will be described using the interaction between a user and a search engine as an example. FIG1 shows a block diagram 100 of an application environment according to an exemplary implementation of the present disclosure. As shown in FIG1 , a user 110 can interact with a search engine 120. For example, the user 110 can enter a search request expressed in text or voice. The search engine 120 can perform a search in a repository 130 to find search results 140 that match the search request.
[0031] Taking voice interaction as an example, it's already possible to provide results based on user voice input. In one example, a user can speak predetermined keywords, such as "voice assistant, voice assistant," to activate the voice processing function of search engine 120. Search engine 120 can then respond with "How can I help you?" After the voice assistant responds, the user can then speak the desired search terms: "What should I do if my windshield washer fluid is frozen?" or similar.
[0032] However, the performance of existing interaction methods is unsatisfactory, and multiple rounds of interaction may occur between the user and the search engine 120. This results in long wait times and low interaction efficiency. For example, when a user is eager to obtain an answer to a question, existing voice interaction methods are inefficient. Therefore, it is desirable to provide a more convenient and effective interaction technology solution.
[0033] Summary of the interaction process
[0034] In order to at least partially address the deficiencies in the prior art, a method for interaction is proposed according to an exemplary implementation of the present disclosure. For ease of description, more details will be described below using a video search scenario as an example. See FIG2 for an overview of an exemplary implementation according to the present disclosure, which shows a block diagram 200 of a process for interaction according to some implementations of the present disclosure. As shown in FIG2 , voice information 210 from the user 110 may be obtained. The voice information 210 may include words spoken by the user to describe their needs, for example, an audio file indicating “what to do if the car windshield washer is frozen”.
[0035] The voice information 210 can be analyzed to determine the search type 220 of the voice information 210. Specifically, if the content type of the voice information 210 is determined to be a search type, the corresponding search information 220 can be determined based on the voice information 210. For example, the user's utterance can indicate that the user desires to search for videos on how to handle frozen glass water. Subsequently, a search can be performed in a repository containing a large number of videos based on the search information 220, and a search result set 230 corresponding to the search information 220 can be determined. Here, the search result set 230 can include at least one video information 240. In other words, the search result set 230 can include one or more videos found.
[0036] Furthermore, a conversation interaction page 250 may be displayed, so that at least one video information in the search result set 230 is displayed in the conversation interaction page 250. It should be understood that the number of video information displayed in the conversation interaction page 250 is not limited. For example, only part of the video information in the search result set may be displayed, or all of the video information may be displayed.
[0037] Using the exemplary implementation of the present disclosure, user 110 does not need to use predefined keywords to activate the search engine's voice processing function. Instead, they can directly speak their search requirements. In this case, the search engine can detect the user's voice information in real time. If the voice information is detected to be related to the search type, it can directly identify matching video information and then provide the found video information to the user. This method can reduce response time and improve the efficiency of human-computer interaction.
[0038] Detailed description of the interaction process
[0039] Having described an overview of an example implementation of the present disclosure, more information about the interaction process will be provided below. According to an example implementation of the present disclosure, during the process of determining the content type of the voice message, the voice message may be converted into text data. Semantic analysis and / or text recognition may then be used to determine the content type of the voice message. For further details, see FIG3 , which illustrates a block diagram 300 of the process for determining the content type of voice message according to some implementations of the present disclosure.
[0040] As shown in FIG3 , the voice information 210 can be converted into text data 310 based on a variety of voice recognition technologies currently known and / or to be developed in the future. For example, the text data "What should I do if my car's windshield washer is frozen" can be extracted from the received audio file. Further, a semantic analysis 320 can be performed on the text data to determine the semantic content expressed by the text data. Here, the semantic content can be determined based on a variety of semantic analysis 320 technical solutions currently known and / or to be developed in the future. For example, various semantic components in a sentence can be extracted from the text data 310 to determine the semantic content, and so on. Further, the content type 340 can be determined based on the semantic content. Through semantic analysis, it can be known that the user wants to search for videos on how to deal with frozen windshield washer. Therefore, the content type 340 of the voice information 210 can be determined as a search type, and the corresponding search information, such as "frozen windshield washer", can be determined.
[0041] Alternatively and / or additionally, keywords associated with the search type can be identified in the text data 310 based on the text recognition 330 technical solution to determine the content type 340. For example, multiple sentence patterns representing search types can be predefined, such as, "Please help me find a short video that solves the xxx problem," "I want to find a short video that solves the xxx problem," "What should I do if xxx has a problem," "I have a problem with xxx, is there any solution," "Play a short video to see how to solve the xxx problem," "I want to see how to solve the xxx problem through a video," "How other users solved the xxx problem," and so on. In this case, keywords associated with the search type may include, but are not limited to, "find," "what to do," "how to solve," "how to solve," "why," and so on.
[0042] By utilizing the exemplary implementation of the present disclosure, the content type of the voice information can be determined in a more accurate and effective manner through semantic analysis and / or text recognition, thereby determining the corresponding search information.
[0043] According to an example implementation of the present disclosure, when determining a search result set based on search information, a machine learning model can be used to determine the degree of match between target video information from multiple video information in a video repository and the search information. FIG4 shows a block diagram 400 for determining a search result set according to some implementations of the present disclosure. As shown in FIG4 , a machine learning model 410 can be used to determine the degree of match between the video information and the search information 220. Here, the machine learning model 410 can be a trained machine learning model, and the model can describe the degree of match between the video information and the search information.
[0044] According to an example implementation of the present disclosure, a machine learning model 410 can be used to determine the degree of match between each piece of video information 430 in the video repository and the search information 220. Further, if it is determined that the degree of match between a certain piece of video information and the search information meets a predetermined condition, the video information can be added to the search result set 230. Here, the predetermined condition may include, for example: the degree of match is higher than a predetermined threshold (e.g., 70% or other predetermined value), the ranking of the degree of match is higher than a predetermined threshold (e.g., ranked in the top 10, etc.). As shown in FIG4 , the degrees of match can be sorted in descending order: match 420> match 422>...> match 424. Further, the video information 240, 412,..., and 414 corresponding to the degrees of match 420, 422,..., 424, respectively, can be added to the search result set 230.
[0045] According to an exemplary implementation of the present disclosure, if no video information matching the predetermined criteria is found (i.e., the search result set is empty), a voice prompt may be generated through text-to-speech conversion. For example, the user may be informed that "no matching search results were found." If video information matching the predetermined criteria is found (i.e., the search result set is not empty), the found video information may be displayed on the dialogue interaction page.
[0046] According to an example implementation of the present disclosure, a conversation interaction page may include at least one video information in a search result set, and the at least one video information is determined based on a matching degree. Specifically, the at least one video information may be sorted according to a matching degree between the at least one video information in the search result set and the search information. Furthermore, the video information with the highest matching degree among the at least one sorted video information may be used as the video information displayed on the conversation interaction page. In other words, the displayed video information may be the video information 240 with the highest matching degree. Alternatively and / or additionally, multiple video information may be displayed (e.g., video information ranked in the top 2, etc.).
[0047] For more details, see FIG5 , which shows a block diagram 500 of a page for displaying search results according to some implementations of the present disclosure. As shown in FIG5 , the user's input and the response to the input can be displayed in the search page 510. For example, the text data 512 can represent text data extracted from the user's language information, and the conversation interaction page 250 can provide a response to the user's input. Hereinafter, the video information in the search result set that is displayed in the conversation page 250 can be referred to as the target video information (e.g., the video information 240), and the video information in the search result set that is not displayed in the conversation page 250 (i.e., the video information other than the target video information) can be referred to as the candidate video information.
[0048] In Figure 5, the conversation interaction page 250 may include summary information 522 of a target video from at least one video message, and a play control 520 for playing the target video from the at least one video message. Summary information 522 may include, but is not limited to, at least one of the following: the target video's cover, title, author, duration, and release date. Alternatively and / or in addition, summary information 522 may further include the degree of match between the target video and the search information. Using this exemplary implementation, users can be presented with a variety of information regarding search results, making it easier for them to determine whether the search results meet their needs.
[0049] According to an example implementation of the present disclosure, video information 240 can be presented in a variety of ways. See Figures 6A and 6B for further details, which respectively illustrate block diagrams 600A and 600B for presenting video information according to some implementations of the present disclosure. As shown in Figure 6A, video information 240 can be played directly. In other words, the search results can be automatically played without user initiation. In this way, search results can be provided to users in a simpler and more efficient manner.
[0050] According to an example implementation of the present disclosure, as shown in FIG6B , a play control 520 may be provided, which a user may press to start video playback. Specifically, upon receiving an interactive operation on the play control 520, the target video information may be played. Alternatively and / or additionally, a pause control for stopping playback may be displayed during playback. In this manner, it is convenient for the user to start or pause video playback when needed.
[0051] According to an example implementation of the present disclosure, candidate video information in the search result set 230 can be played. Playing of the candidate video information can be triggered in a variety of ways. In one example, a play control for playing the candidate video information can be provided. For example, a jump control for jumping to the next candidate video information after the target video information can be provided. The candidate video information can be played when an interactive operation (i.e., a play request for playing the candidate video information) is received for the jump control.
[0052] Alternatively and / or additionally, the candidate video information can be played upon detecting a playback end event associated with the target video information. That is, after the target video information has finished playing, the playback of the candidate video information can be automatically initiated. Alternatively and / or additionally, the candidate video information can be played upon detecting a playback pause event associated with the target video information. In this manner, users can adjust the playback content to their needs, thereby obtaining desired search results in a more convenient and efficient manner.
[0053] According to an example implementation of the present disclosure, a list of each video information item in a search result set can be displayed. For example, when a trigger event for displaying candidate video information in a search result set is detected, the candidate video information item can be displayed. Specifically, the candidate information item can be displayed in a variety of ways. When a search result set includes multiple candidate video information items, the candidate video information item can be displayed based on the degree of match between at least one video information item in the search result set and the search information. For example, the candidate video information items can be displayed in descending order of their matching degree.
[0054] According to an exemplary implementation of the present disclosure, the triggering events for displaying candidate video information in the search result set can involve various types. For example, a user can click on the display control 612 shown in Figures 6A and / or 6B to view the candidate information. If a user interaction with the display control 612 is detected, the candidate video information can be displayed based on the degree of match between at least one video information in the search result set and the search information.
[0055] Refer to FIG7 for more details on displaying candidate video information, which shows a block diagram 700 for displaying more video information according to some implementations of the present disclosure. As shown in FIG7 , a list of video information can be displayed, which can include video information 240, video information 412, ..., and video information 414. Indicator 720 can indicate video information that has been played, or can represent video information currently being played, etc. It should be understood that the individual video information 240, 412, ..., and 414 can be arranged in order from high to low according to the degree of matching. In this way, users can be supported to preferentially view video information with a higher degree of matching.
[0056] According to an example implementation of the present disclosure, when a playback end event associated with target video information is detected, that is, after the target video information has finished playing, a list including more candidate video information can be automatically popped up. In other words, after the playback end event, the candidate video information in the search combination can be displayed based on the degree of match between at least one video information in the search result set and the search information. In this case, after the video information with the highest matching degree has finished playing, the list shown in Figure 7 can be automatically displayed. In this way, the complexity of user operations can be reduced, and potential search results can be automatically displayed in order from high to low matching degree.
[0057] According to an example implementation of the present disclosure, a user can interact with the candidate video information. For example, the user can play the video information 412 by single-clicking, double-clicking, and / or other means. When an interactive operation with the video information 412 is detected, the video information 412 can be loaded from the repository. FIG8A shows a block diagram 800A for loading video information according to some implementations of the present disclosure. As shown in FIG8A , in the play page 810, a progress indicator for loading the video information 412 can be displayed, and summary information about the video information 412 (for example, the avatar and name of the user who posted the video information, and the title and duration of the video information, etc.) and display controls for displaying the candidate video information can be displayed.
[0058] In the event of a successful load, the video information 412 can be played. Alternatively and / or additionally, FIG8B shows a block diagram 800B for loading video information according to some implementations of the present disclosure, i.e., FIG8B shows a case where loading fails. In the event of a load failure, the play page 820 can indicate the load failure and display a load control 822. The user can press the load control 822 to reload the video information 412 from the repository. Alternatively and / or additionally, although not shown, the play page 820 can include a return control to return to the list shown in FIG7.
[0059] According to an exemplary implementation of the present disclosure, the technical solution described above can be implemented in a variety of application environments. For example, a user can perform a voice search process on a mobile computing device or a fixed computing device, and the found media data can be displayed on a display device of the mobile computing device.
[0060] Alternatively and / or additionally, the technical solution described above can be implemented on an in-vehicle computing device. In this case, the in-vehicle computing device is deployed in the vehicle and can include one or more display devices. For example, a first display device can be deployed in the front driver's seat of the vehicle, and a second display device can be deployed in the rear passenger seat of the vehicle. Furthermore, one or more acquisition devices can be deployed inside the vehicle to collect voice information from the driver and / or passengers.
[0061] According to an example implementation of the present disclosure, a location associated with voice information can be determined, where the location represents the location where the voice information appears in the vehicle. Furthermore, in the process of displaying a dialogue interaction page, the dialogue interaction page can be displayed at a display device corresponding to the location in the vehicle. Specifically, assuming that a voice message from the driver is detected, it can be determined that the voice message comes from the front row of the vehicle, that is, from the driver. At this time, the video information that the driver expects to obtain (for example, what to do if the windshield washer fluid is frozen) can be displayed on a first display device near the driver. In this case, the video information will not be displayed on the second display device at the rear passenger seat, but the user can continue to watch the video of his interest, and so on.
[0062] Alternatively and / or additionally, assuming that voice information from a passenger is detected, it can be determined that the voice information comes from the rear seat of the vehicle. A video information desired by the passenger (e.g., a movie, etc.) can be displayed on a second display device near the passenger. In this case, the video information is not displayed on the first display device at the front driver's seat, thereby avoiding interference with the driver's normal driving operations.
[0063] Using the exemplary implementation of the present disclosure, users do not need to use predefined keywords to activate the search engine's voice processing function, but can instead directly express their desired search requirements. In this case, the search engine can detect voice information from the user in real time, and if it detects that the voice information relates to the search information, it can directly determine matching video information and provide the corresponding video information to the user. In this way, response time can be reduced and the efficiency of human-computer interaction can be improved. Furthermore, in a vehicle environment, users can be prompted to pay attention to driving safety using text and / or voice.
[0064] According to an exemplary implementation of the present disclosure, a voice interaction function can be provided on the login page of an in-vehicle device, allowing users to ask questions and watch videos in real time while in the vehicle. Furthermore, the vehicle manufacturer can provide video guidance on vehicle knowledge, which can be played on display devices at various locations in the vehicle, thereby helping users improve their problem-solving skills more quickly.
[0065] Example Process
[0066] FIG9 illustrates a flow chart of a method 700 for interaction according to some implementations of the present disclosure. At block 910, a voice message is obtained. At block 920, if the content type of the voice message is a search type, search information is determined based on the voice message. At block 930, a search result set corresponding to the search information is determined based on the search information, where the search result set includes at least one video. At block 940, a conversation interaction page is displayed, where at least one video from the search result set is displayed on the conversation interaction page.
[0067] According to an example implementation of the present disclosure, the content type of voice information is determined based on the following steps: converting the voice information into text data; performing semantic analysis on the text data to determine the semantic content expressed by the text data, and determining the content type based on the semantic content; or identifying keywords associated with the search type in the text data to determine the content type.
[0068] According to an example implementation of the present disclosure, determining a search result set corresponding to the search information based on the search information includes: utilizing a machine learning model to determine a degree of matching between target video information among multiple video information in a video repository and the search information; and adding the target video information to the search result set if it is determined that the degree of matching between the target video information and the search information satisfies a predetermined condition.
[0069] According to an example implementation of the present disclosure, the conversation interaction page includes: summary information of target video information in at least one video information, and playback controls for playing the target video information in at least one video information, the summary information including at least any one of the following: cover, title, author, duration, and release time of the target video information.
[0070] According to an example implementation of the present disclosure, the method further includes: if an interactive operation on a playback control is received, playing the target video information.
[0071] According to an example implementation of the present disclosure, the method further includes: playing candidate video information other than the target video information in the search result set if at least any one of the following is detected: a play request for playing the candidate video information; a play end event associated with the target video information; and a play pause event associated with the target video information.
[0072] According to an example implementation of the present disclosure, the method further includes: if a triggering event for displaying candidate video information in the search result set is detected, displaying the candidate video information according to a matching degree between at least one video information in the search result set and the search information.
[0073] According to an example implementation of the present disclosure, displaying candidate video information includes: if a playback end event associated with the target video information is detected, displaying the candidate video information in the search combination according to the matching degree between at least one video information in the search result set and the search information; or the dialogue interaction page further includes a display control for displaying the candidate video information in the search result set, and if an interactive operation for the display control is received, displaying the candidate video information according to the matching degree between at least one video information in the search result set and the search information.
[0074] According to an example implementation of the present disclosure, a video information in a search result set is displayed on a conversation interaction page, and the video information is determined based on the following steps: sorting at least one video information according to the degree of match between at least one video information in the search result set and the search information; and selecting the video information with the highest degree of match from the at least one sorted video information as the video information displayed on the conversation interaction page.
[0075] According to an example implementation of the present disclosure, the method is executed at an on-board computing device in a vehicle, and the method further includes: determining a location associated with voice information, the location representing an occurrence location of the voice information in the vehicle; and wherein displaying the conversation interaction page further includes: displaying the conversation interaction page at a display device in the vehicle corresponding to the location.
[0076] Example devices and equipment
[0077] Figure 10 shows a block diagram of an apparatus 1000 for interaction according to some implementations of the present disclosure. The apparatus includes: an acquisition module 1010 configured to acquire voice information; an information determination module 1020 configured to determine search information based on the voice information if the content type of the voice information is a search type; a result determination module 1030 configured to determine a search result set corresponding to the search information based on the search information, wherein the search result set includes at least one video; and a presentation module 1040 configured to present a dialogue interaction page, wherein the dialogue interaction page presents at least one video from the search result set.
[0078] According to an example implementation of the present disclosure, the content type of voice information is determined based on the following modules: a conversion module, configured to convert the voice information into text data; a first type determination module, configured to perform semantic analysis on the text data to determine the semantic content expressed by the text data, and determine the content type based on the semantic content; or a second type determination module, configured to identify keywords associated with the search type in the text data to determine the content type.
[0079] According to an example implementation of the present disclosure, the result determination module includes: a matching degree determination module, configured to use a machine learning model to determine the matching degree between target video information and search information among multiple video information in a video repository; and an adding module, configured to add the target video information to the search result set if it is determined that the matching degree between the target video information and the search information meets a predetermined condition.
[0080] According to an example implementation of the present disclosure, the conversation interaction page includes: summary information of target video information in at least one video information, and playback controls for playing the target video information in at least one video information, the summary information including at least any one of the following: cover, title, author, duration, and release time of the target video information.
[0081] According to an exemplary implementation of the present disclosure, the apparatus further includes: a first playback module configured to play target video information if an interactive operation directed to a playback control is received.
[0082] According to an example implementation of the present disclosure, the device further includes: a second playback module, configured to play candidate video information other than the target video information in the search result set if at least any one of the following is detected: a playback request for playing the candidate video information; a playback end event associated with the target video information; and a playback pause event associated with the target video information.
[0083] According to an example implementation of the present disclosure, the device further includes: a candidate video information display module, which is configured to display the candidate video information according to the matching degree between at least one video information in the search result set and the search information if a trigger event for displaying the candidate video information in the search result set is detected.
[0084] According to an example implementation of the present disclosure, the candidate video information display module further includes: a first display module, configured to display the candidate video information in the search combination according to the matching degree between at least one video information in the search result set and the search information if a playback end event associated with the target video information is detected; and a second display module, configured to display the candidate video information according to the matching degree between at least one video information in the search result set and the search information if an interactive operation for the display control is received, and the dialogue interaction page further includes a display control for displaying the candidate video information in the search result set.
[0085] According to an example implementation of the present disclosure, a video information in a search result set is displayed on a conversation interaction page, and the video information is determined based on the following modules: a sorting module, configured to sort at least one video information according to a degree of match between at least one video information in the search result set and the search information; and a selection module, configured to select the video information with the highest degree of match from the at least one sorted video information as the video information displayed on the conversation interaction page.
[0086] According to an example implementation of the present disclosure, the apparatus is implemented at an on-board computing device in a vehicle, and the apparatus further includes: a location determination module configured to determine a location associated with voice information, the location representing an appearance location of the voice information in the vehicle; and a location-based display module configured to display a dialogue interaction page at a display device corresponding to the location in the vehicle.
[0087] FIG11 shows a block diagram of a device 1100 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1100 shown in FIG11 is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. The computing device 1100 shown in FIG11 can be used to implement the methods described above.
[0088] As shown in FIG11 , computing device 1100 is in the form of a general-purpose computing device. Components of computing device 1100 may include, but are not limited to, one or more processors or processing units 1110, memory 1120, storage device 1130, one or more communication units 1140, one or more input devices 1150, and one or more output devices 1160. Processing unit 1110 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 1120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 1100.
[0089] The computing device 1100 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device 1100, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1120 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1130 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device 1100.
[0090] The computing device 1100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 11 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 1120 may include a computer program product 1125 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0091] The communication unit 1140 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1100 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 1100 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.
[0092] Input device 1150 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 1160 may be one or more output devices, such as a display, a speaker, or a printer. Computing device 1100 may also communicate with one or more external devices (not shown) via communication unit 1140, as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with computing device 1100, or with any device that allows computing device 1100 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0093] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0094] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0095] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0096] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0097] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0098] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. An interaction method, comprising: Obtaining a voice message; If the content type of the voice message is a search type, determining search information according to the voice message; Determining a search result set corresponding to the search information according to the search information, wherein the search result set includes at least one video message; and Displaying a dialogue interaction page, wherein at least one video message in the search result set is displayed in the dialogue interaction page.
2. The method according to claim 1, wherein the content type of the voice message is determined based on the following steps: Converting the voice message into text data; Performing semantic analysis on the text data to determine the semantic content expressed by the text data, and determining the content type based on the semantic content; or Identifying keywords associated with the search type in the text data to determine the content type.
3. The method according to claim 1, wherein determining the search result set corresponding to the search information according to the search information includes: Using a machine learning model to determine the matching degree between a target video message in a plurality of video messages in a video repository and the search information; And If it is determined that the matching degree between the target video message and the search information meets a predetermined condition, adding the target video message to the search result set.
4. The method according to claim 1, wherein the dialogue interaction page comprises: The summary information of the target video message in the at least one video message, and a playback control for playing the target video message in the at least one video message, wherein the summary information includes at least any one of the following: the cover, title, author, time length, and release time of the target video message.
5. The method according to claim 4, further comprising: If an interaction operation on the playback control is received, playing the target video message.
6. The method according to claim 5, further comprising: If at least any one of the following is detected, playing a candidate video message other than the target video message in the search result set: A playback request for playing the candidate video message; A playback end event associated with the target video message; and A playback pause event associated with the target video message.
7. The method according to claim 5, further comprising: If a trigger event for displaying a candidate video message in the search result set is detected, displaying the candidate video message according to the matching degree between at least one video message in the search result set and the search information.
8. The method according to claim 7, wherein displaying the candidate video message includes: If a playback end event associated with the target video message is detected, displaying the candidate video message in the search combination according to the matching degree between at least one video message in the search result set and the search information; Or The dialogue interaction page further includes a display control for displaying the candidate video message in the search result set, and if an interaction operation on the display control is received, displaying the candidate video message according to the matching degree between at least one video message in the search result set and the search information.
9. The method according to claim 3, wherein one video information in the search result set is displayed in the dialogue interaction page, and the video information is determined based on the following steps: Sort the at least one video information according to the matching degree between the at least one video information in the search result set and the search information; and From the sorted at least one video information, use the video information with the highest matching degree as the video information displayed in the dialogue interaction page.
10. The method according to claim 1, wherein the method is performed at an in-vehicle computing device in a vehicle, and the method further comprises: Determine the position associated with the voice information, where the position represents the occurrence position of the voice information in the vehicle; And where displaying the dialogue interaction page further includes: displaying the dialogue interaction page at a display device corresponding to the position in the vehicle.
11. An interaction device, comprising: An acquisition module configured to acquire a voice information; An information determination module configured to, if the content type of the voice information is a search type, determine search information according to the voice information; A result determination module configured to determine a search result set corresponding to the search information according to the search information, wherein the search result set includes at least one video information; and A display module configured to display a dialogue interaction page, wherein at least one video information in the search result set is displayed in the dialogue interaction page.
12. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to execute the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, on which a computer program is stored, the computer program, when executed by a processor, causing the processor to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Video playing method based on voice semantics
CN111797274A
Search result display method, device and equipment and non-instantaneous computer storage medium
CN115037959A
Media data display method and device, computer equipment and storage medium
CN116975322A