Method and device for interaction, equipment and medium
The search type is determined through semantic analysis of speech information and machine learning model, and the matching video information is directly displayed in the dialogue interaction page, solving the problem of low voice interaction efficiency in the prior art and achieving more efficient video search and display.
Patent Information
- Application Number
- CN202410016097.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
The existing voice interaction methods are inefficient in the interaction between users and search engines, resulting in longer waiting times and lower interaction efficiency.
By obtaining the user's voice information, semantic analysis and/or text recognition technology is used to determine the content type of the voice information as the search type, and using machine learning models to determine the matching video information in the video repository, and directly display the search results in the dialogue interaction page.
Reduces response time and improves the efficiency of human-computer interaction. Users do not need to use predetermined keywords to activate the search engine and directly state the search needs to obtain relevant video information.
Smart Images

Figure CN120256673A_ABST
Abstract
Description
Technical Field
[0001] Exemplary implementations of the present disclosure generally relate to interaction processing, and particularly to methods, devices, equipment, and computer-readable storage media for voice interaction. Background Art
[0002] Currently, a variety of user interaction technology solutions have been proposed. For example, users can perform human-computer interaction via voice, text, and actions, etc. In the field of voice interaction, corresponding actions can already be performed based on voice input from users to provide corresponding results. However, the performance of existing interaction methods is not satisfactory, so it is desired to provide more convenient and effective interaction technology solutions. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for interaction is provided. In this method, a voice message is obtained. If the content type of the voice message is a search type, search information is determined according to the voice message. According to the search information, a search result set corresponding to the search information is determined, where the search result set includes at least one video message. A dialogue interaction page is displayed, where at least one video message in the search result set is displayed in the dialogue interaction page.
[0004] In a second aspect of the present disclosure, a device for interaction is provided. The device includes: an obtaining module configured to obtain a voice message; an information determination module configured to determine search information according to the voice message if the content type of the voice message is a search type; a result determination module configured to determine a search result set corresponding to the search information according to the search information, where the search result set includes at least one video message; and a display module configured to display a dialogue interaction page, where at least one video message in the search result set is displayed in the dialogue interaction page.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, where at least one memory is coupled to at least one processing unit and stores instructions for execution by at least one processing unit, and when the instructions are executed by at least one processing unit, the electronic device is caused to execute the method according to the first aspect of the present disclosure.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the processor is caused to implement the method according to the first aspect of the present disclosure.
[0007] It should be understood that the content described in this content part is not intended to limit the key features or important features of the implementation of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In the following, with reference to the accompanying drawings and in conjunction with the following detailed description, the above and other features, advantages, and aspects of various implementations of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0009] Figure 1 A block diagram of an application environment according to an exemplary implementation of the present disclosure is shown;
[0010] Figure 2 A block diagram of a process for interaction according to some implementations of the present disclosure is shown;
[0011] Figure 3 A block diagram of a process for determining the content type of voice information according to some implementations of the present disclosure is shown;
[0012] Figure 4 A block diagram of a process for determining a search result set according to some implementations of the present disclosure is shown;
[0013] Figure 5 A block diagram of a page for displaying search results according to some implementations of the present disclosure is shown;
[0014] Figure 6A and Figure 6B Block diagrams of a process for displaying video information according to some implementations of the present disclosure are shown respectively;
[0015] Figure 7 A block diagram of a process for displaying more video information according to some implementations of the present disclosure is shown;
[0016] Figure 8A and Figure 8B Block diagrams of a process for loading video information according to some implementations of the present disclosure are shown respectively;
[0017] Figure 9 A flowchart of a method for interaction according to some implementations of the present disclosure is shown;
[0018] Figure 10 A block diagram of a device for interaction according to some implementations of the present disclosure is shown; and
[0019] Figure 11 A block diagram of a device capable of implementing multiple implementations of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The implementation manners of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some implementation manners of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementation manners set forth herein. On the contrary, these implementation manners are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and implementation manners of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0021] In the description of the implementation manners of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an implementation manner" or "the implementation manner" should be understood as "at least one implementation manner". The term "some implementation manners" should be understood as "at least some implementation manners". There may also be other explicit and implicit definitions hereinafter. As used herein, the term "model" may represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on a variety of technical solutions known currently and / or to be developed in the future.
[0022] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0023] It can be understood that before using the technical solutions disclosed in the implementation manners of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to the relevant laws and regulations.
[0024] For example, when receiving the user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require the acquisition and use of the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0025] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0026] It should be understood that the above notification and the process of obtaining user authorization are merely illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0027] As used herein, the term "responsive to" indicates a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the execution timing of subsequent actions performed in response to the event or condition and the time when the event occurs or the condition is established may not necessarily be strongly correlated. For example, in some cases, the subsequent action can be immediately executed when the event occurs or the condition is established; while in other cases, the subsequent action can be executed after a period of time after the event occurs or the condition is established.
[0028] Example environment
[0029] Currently, a variety of user interaction technical solutions have been proposed. For example, a user can perform human-computer interaction via voice, text, actions, etc. In the following, more details about the interaction will be described by taking the interaction between a user and a search engine as an example. Figure 1 FIG. 100 is a block diagram showing an application environment according to an exemplary implementation manner of the present disclosure. As Figure 1 shown, a user 110 can interact with a search engine 120. For example, the user 110 can input a search request expressed in text or voice, etc. The search engine 120 can perform a search in a repository 130 to find search results 140 that match the search request.
[0030] Taking voice interaction as an example, corresponding results can already be provided based on voice input from a user. In one example, a user can say a predetermined keyword, for example, use keywords such as "voice assistant, voice assistant" to wake up the voice processing function of the search engine 120. The search engine 120 can answer "What can I do for you?" After the voice assistant responds, the user can say the content to be searched: "What should I do if the car windshield washer fluid freezes?", etc.
[0031] However, the performance of existing interaction manners is not satisfactory, and there may be multiple rounds of interaction between the user and the search engine 120. This results in a long waiting time and low interaction efficiency. For example, in the case where a user is eager to obtain an answer to a question, the processing efficiency of existing voice interaction manners is low. Therefore, it is desirable to provide a more convenient and effective interaction technical solution.
[0032] Summary of the interaction process
[0033] To at least partially address the deficiencies in the prior art, according to an exemplary implementation of the present disclosure, a method for interaction is proposed. For ease of description, in the following, more details will be described by taking the video search scenario as an example. Refer to Figure 2 Describe the outline of an exemplary implementation of the present disclosure, Figure 2 FIG. 200 is a block diagram showing a process for interaction according to some implementations of the present disclosure. As Figure 2 shown, voice information 210 from user 110 can be obtained. The voice information 210 may include words spoken by the user to describe their own needs. For example, an audio file indicating "What should I do if the windshield washer fluid freezes?"
[0034] The voice information 210 can be analyzed to determine the search type 220 of the voice information 210. Specifically, if it is determined that the content type of the voice information 210 is a search type, the corresponding search information 220 can be determined based on the voice information 210. For example, it can be determined from the words spoken by the user that the user expects to search for videos related to dealing with frozen windshield washer fluid. Subsequently, a search can be performed in a repository including a large number of videos according to the search information 220, and a set of search results 230 corresponding to the search information 220 can be determined. Here, the set of search results 230 may include at least one video information 240. That is, the set of search results 230 may include one or more pieces of found video information.
[0035] Furthermore, a dialogue interaction page 250 can be displayed to display at least one video information in the set of search results 230 in the dialogue interaction page 250. It should be understood that the number of video information displayed in the dialogue interaction page 250 is not limited here. For example, only some of the video information in the set of search results may be displayed, or all of the video information may be displayed.
[0036] Using the exemplary implementation of the present disclosure, user 110 does not have to use a predetermined keyword to activate the voice processing function of the search engine, but can directly speak out the search requirement. At this time, the search engine can detect the voice information from the user in real time, and when it detects that the voice information involves a search type, directly determine the matching video information and then provide the found video information to the user. In this way, the response time can be reduced and the efficiency of human-computer interaction can be improved.
[0037] Detailed description of the interaction process
[0038] A summary of an example implementation according to the present disclosure has been described. Hereinafter, more information about the interaction process will be provided. According to an example implementation of the present disclosure, in the process of determining the content type of voice information, the voice information can be converted into text data. Furthermore, semantic analysis and / or text recognition are utilized to determine the content type of the voice information. Refer to Figure 3 Describe more details, the Figure 3 FIG. 300 is a block diagram showing a process for determining the content type of voice information according to some implementations of the present disclosure.
[0039] As Figure 3 shown, the voice information 210 can be converted into text data 310 based on various speech recognition technologies that are currently known and / or will be developed in the future. For example, the text data "What should I do if the windshield washer fluid freezes" can be extracted from the received audio file. Further, semantic analysis 320 can be performed on the text data to determine the semantic content expressed by the text data. Here, the semantic content can be determined based on various semantic analysis 320 technical solutions that are currently known and / or will be developed in the future. For example, various semantic components in the sentence can be extracted from the text data 310, and then the semantic content can be determined, and so on. Further, the content type 340 can be determined based on the semantic content. Through semantic analysis, it can be known that the user expects to search for videos on how to deal with the frozen windshield washer fluid. Thus, the content type 340 of the voice information 210 can be determined as the search type, and the corresponding search information, for example, "windshield washer fluid freezes", can be determined.
[0040] Alternatively and / or additionally, keywords associated with the search type can be identified in the text data 310 based on the text recognition 330 technical solution to determine the content type 340. For example, various sentence patterns indicating the search type can be predefined, such as "Please help me find short videos on solving xxx problems", "I want to find short videos on solving xxx problems", "What should I do if xxx has a problem", "xxx has a problem, is there any solution", "Play short videos to see how to solve xxx problems", "I want to see through videos how to solve xxx problems", "How do other users solve xxx problems", and so on. At this time, the keywords associated with the search type can include, for example, but are not limited to: "find", "what should I do", "how to solve", "how to resolve", "why", and so on.
[0041] Using the example implementation of the present disclosure, through semantic analysis and / or text recognition, the content type of the voice information can be determined in a more accurate and effective manner, and then the corresponding search information can be determined.
[0042] According to an example implementation of the present disclosure, in the process of determining a search result set based on search information, a machine learning model can be used to determine the matching degree between target video information among multiple video information in a video repository and the search information. Figure 4 FIG. 400 is a block diagram showing a method for determining a search result set according to some implementations of the present disclosure. As Figure 4 described, a machine learning model 410 can be used to determine the matching degree between video information and search information 220. Here, the machine learning model 410 can be a trained machine learning model, and this model can describe the matching degree between video information and search information.
[0043] According to an example implementation of the present disclosure, the machine learning model 410 can be used to separately determine the matching degree between each piece of video information 430 in the video repository and the search information 220. Further, when it is determined that the matching degree between a certain piece of video information and the search information meets a predetermined condition, this piece of video information can be added to the search result set 230. Here, the predetermined condition can include, for example: the matching degree is higher than a predetermined threshold (for example, 70% or other predetermined values), and the ranking of the matching degree is higher than a predetermined threshold (for example, the ranking is in the top 10, etc.). As Figure 4 shown, the matching degrees can be sorted in descending order: matching degree 420 > matching degree 422 >... > matching degree 424. Further, the video information 240, 412,..., and 414 corresponding to the matching degrees 420, 422,..., and 424 can be added to the search result set 230 respectively.
[0044] According to an example implementation of the present disclosure, if no video information whose matching degree meets the predetermined condition is found (that is, the search result set is empty), a voice prompt can be generated through text-to-speech conversion. For example, the user can be informed that "no matching search results were found", etc. If video information whose matching degree meets the predetermined condition is found (that is, the search result set is non-empty), the found video information can be displayed on the dialogue interaction page.
[0045] According to an example implementation of the present disclosure, the dialogue interaction page can include at least one piece of video information in the search result set, and this at least one piece of video information is determined based on the matching degree. Specifically, at least one piece of video information in the search result set can be sorted according to the matching degree between the at least one piece of video information and the search information. Further, from the sorted at least one piece of video information, the piece of video information with the highest matching degree can be used as the video information to be displayed on the dialogue interaction page. In other words, the displayed video information can be the video information 240 with the highest matching degree. Alternatively and / or additionally, multiple pieces of video information can be displayed (for example, the video information ranked in the top 2, etc.).
[0046] See Figure 5 for more details, the Figure 5 shows a block diagram 500 of a page for presenting search results according to some implementations of the present disclosure. As Figure 5 shown, the user input and responses to the input can be presented in the search page 510. For example, the text data 512 can represent text data extracted from the user's language information, and the dialogue interaction page 250 can provide responses to the user input. Hereinafter, the video information presented in the dialogue page 250 in the search result set can be referred to as target video information (e.g., video information 240), and the video information not presented in the dialogue page 250 in the search result set (i.e., video information other than the target video information) can be referred to as candidate video information.
[0047] In Figure 5 , the dialogue interaction page 250 can include: summary information 522 of the target video information in at least one video information, and playback controls 520 for playing the target video information in at least one video information. Here, the summary information 522 can include, but is not limited to, at least any one of the following: the cover, title, author, time length, and release time of the target video information, etc. Alternatively and / or additionally, the summary information 522 can further include the degree of match between the target video information and the search information. Using the exemplary implementations of the present disclosure, multi-faceted information about the search results can be presented to the user, thereby facilitating the user to determine whether the search results meet their own needs.
[0048] According to an exemplary implementation of the present disclosure, the video information 240 can be presented in multiple ways. See Figure 6A and Figure 6B for more details, the Figure 6A and Figure 6B respectively show block diagrams 600A and 600B for presenting video information according to some implementations of the present disclosure. As Figure 6A shown, the video information 240 can be directly played, in other words, the found search results can be automatically played without the user starting. In this way, the search results can be provided to the user in a simpler and more effective manner.
[0049] According to an exemplary implementation of the present disclosure, as Figure 6BAs shown, a playback control 520 can be provided, and the user can press the playback control 520 to start video playback. Specifically, in response to an interaction operation on the playback control 520, the target video information can be played. Alternatively and / or additionally, a pause control for stopping playback can be shown during playback. In this way, it is convenient for the user to start or pause video playback when needed.
[0050] According to an example implementation of the present disclosure, candidate video information in a search result set 230 can be played. There are various ways to trigger the playback of candidate video information. In one example, a playback control for playing candidate video information can be provided. For example, a jump control for jumping to the next candidate video information after the target video information can be provided. In response to an interaction operation on the jump control (i.e., a playback request for playing candidate video information), the candidate video information can be played.
[0051] Alternatively and / or additionally, candidate video information can be played when a playback end event associated with the target video information is detected. That is, after the playback of the target video information has been completed, the playback of candidate video information can be automatically started. Alternatively and / or additionally, candidate video information can be played when a playback pause event associated with the target video information is detected. In this way, it is convenient for the user to adjust the playback content according to their own needs, so as to obtain the desired search results in a more convenient and effective manner.
[0052] According to an example implementation of the present disclosure, a list of each video information in the search result set can be shown. For example, in response to a trigger event for showing candidate video information in the search result set, the candidate video information can be shown. Specifically, the candidate information can be shown in various ways. When the search result set includes multiple candidate video information, the candidate video information can be shown according to the matching degree between at least one video information in the search result set and the search information. For example, each candidate video information can be shown in descending order of the matching degree.
[0053] According to an example implementation of the present disclosure, the trigger event for showing candidate video information in the search result set can involve various types. For example, the user can click on the display control 612 as shown in Figure 6A and / or Figure 6B to view the candidate information. In response to the detected interaction operation of the user with the display control 612, the candidate video information can be shown according to the matching degree between at least one video information in the search result set and the search information.
[0054] See Figure 7 for more details about showing candidate video information. TheFigure 7 FIG. 700 is a block diagram showing more video information according to some implementations of the present disclosure. As Figure 7 shown, a list of video information can be presented, which may include video information 240, video information 412, …, and video information 414. The indicator 720 can indicate the video information that has been played, or can represent the video information currently being played, and so on. It should be understood that each of the video information 240, 412, …, and 414 can be arranged in descending order of matching degree. In this way, it is possible to support the user to preferentially view video information with a higher matching degree.
[0055] According to an example implementation of the present disclosure, in the case of detecting a playback end event associated with the target video information, that is, after the target video information has been played, a list including more candidate video information can be automatically popped up. In other words, after the playback end event, the candidate video information in the search set can be presented according to the matching degree between at least one video information in the search result set and the search information. At this time, after the video information with the highest matching degree has been played, the list as Figure 7 shown can be automatically presented. In this way, the complexity of user operations can be reduced, and potential search results can be automatically presented in descending order of matching degree.
[0056] According to an example implementation of the present disclosure, the user can interact with the candidate video information. For example, the user can play the video information 412 by clicking, double-clicking, and / or other means. In the case of detecting an interaction operation for the video information 412, the video information 412 can be loaded from the repository. Figure 8A FIG. 800A is a block diagram showing the loading of video information according to some implementations of the present disclosure. As Figure 8A shown, in the playback page 810, a progress indicator for loading the video information 412 can be presented, and summary information about the video information 412 (such as the avatar, name of the user who released the video information, and the title, duration of the video information, etc.) and a display control for presenting candidate video information can be presented.
[0057] In the case of successful loading, the video information 412 can be played. Alternatively and / or additionally, Figure 8B FIG. 800B is a block diagram showing the loading of video information according to some implementations of the present disclosure, that is, this Figure 8BShows a situation where the loading fails. In the case of loading failure, the playback page 820 may indicate the loading failure and display a loading control 822. The user may press the loading control 822 to reload the video information 412 from the repository. Alternatively and / or additionally, although not shown, the playback page 820 may include a return control to return to the list as Figure 7 shown.
[0058] According to an example implementation of the present disclosure, the above-described technical solution may be executed in a variety of application environments. For example, a user may execute a voice search process at a mobile computing device or a fixed computing device, and at this time, the found media data may be displayed on the display device of the mobile computing device.
[0059] Alternatively and / or additionally, the above-described technical solution may be executed at an in-vehicle computing device. At this time, the in-vehicle computing device is deployed in a vehicle, and the in-vehicle computing device may include one or more display devices. For example, a first display device may be deployed at the front driver's position of the vehicle, and a second display device may be deployed at the rear passenger's position of the vehicle. Further, one or more collection devices may be deployed inside the vehicle to collect voice information from the driver and / or passengers.
[0060] According to an example implementation of the present disclosure, a position associated with the voice information may be determined, where this position represents the occurrence position of the voice information in the vehicle. Further, during the process of displaying the dialogue interaction page, the dialogue interaction page may be displayed on the display device corresponding to this position in the vehicle. Specifically, assuming that voice information from the driver is detected, it may be determined that the voice information comes from the front row position of the vehicle, that is, from the driver. At this time, the video information that the driver expects to obtain (for example, what to do if the windshield washer fluid freezes) may be displayed on the first display device near the driver. In this case, the video information will not be displayed on the second display device at the rear passenger's position, but the user can continue to watch the video that they are interested in, and so on.
[0061] Alternatively and / or additionally, assuming that voice information from a passenger is detected, it may be determined that the voice information comes from the rear row position of the vehicle. The video information that the passenger expects to obtain (for example, a certain movie, etc.) may be displayed on the second display device near the passenger. At this time, the video information will not be displayed on the first display device at the front driver's position. In this way, it is possible to avoid interfering with the normal driving operation of the driver.
[0062] Using the exemplary implementation of the present disclosure, the user does not have to use a predetermined keyword to activate the voice processing function of the search engine, but can directly state their search needs. At this time, the search engine can detect the voice information from the user in real time, and when it detects that the voice information relates to search information, directly determine the matching video information and then provide the corresponding video information to the user. In this way, the response time can be reduced and the efficiency of human-computer interaction can be improved. Further, in a vehicle environment, the user can be prompted to pay attention to driving safety in text and / or voice form.
[0063] According to an exemplary implementation of the present disclosure, a voice interaction function can be provided on the login page of the in-vehicle device to meet the user's needs for real-time Q&A and video viewing in the vehicle. In addition, the vehicle manufacturer can provide video guidance on vehicle knowledge, and thus play videos on the display devices at various positions in the vehicle in the form of videos, thereby helping the user to more quickly improve the ability to solve problems.
[0064] Example process
[0065] Figure 9 The flowchart of a method 700 for interaction according to some implementations of the present disclosure is shown. At block 910, a voice message is obtained. At block 920, if the content type of the voice message is a search type, search information is determined based on the voice message. At block 930, according to the search information, a search result set corresponding to the search information is determined, where the search result set includes at least one video information. At block 940, a dialogue interaction page is displayed, where at least one video information in the search result set is displayed in the dialogue interaction page.
[0066] According to an exemplary implementation of the present disclosure, the content type of the voice message is determined based on the following steps: converting the voice message into text data; performing semantic analysis on the text data to determine the semantic content expressed by the text data, and determining the content type based on the semantic content; or identifying keywords associated with the search type in the text data to determine the content type.
[0067] According to an exemplary implementation of the present disclosure, determining the search result set corresponding to the search information according to the search information includes: using a machine learning model to determine the matching degree between the target video information in the multiple video information in the video repository and the search information; and if it is determined that the matching degree between the target video information and the search information meets a predetermined condition, adding the target video information to the search result set.
[0068] According to an exemplary implementation of the present disclosure, the conversation interaction page includes: summary information of the target video information among at least one video information, and a playback control for playing the target video information among at least one video information, and the summary information includes at least any one of the following: the cover, title, author, duration, and release time of the target video information.
[0069] According to an exemplary implementation of the present disclosure, the method further includes: if an interaction operation for the playback control is received, playing the target video information.
[0070] According to an exemplary implementation of the present disclosure, the method further includes: if at least any one of the following is detected, playing candidate video information other than the target video information in the search result set: a playback request for playing the candidate video information; a playback end event associated with the target video information; and a playback pause event associated with the target video information.
[0071] According to an exemplary implementation of the present disclosure, the method further includes: if a trigger event for displaying candidate video information in the search result set is detected, displaying the candidate video information according to the matching degree between at least one video information in the search result set and the search information.
[0072] According to an exemplary implementation of the present disclosure, displaying the candidate video information includes: if a playback end event associated with the target video information is detected, displaying the candidate video information in the search combination according to the matching degree between at least one video information in the search result set and the search information; or the conversation interaction page further includes a display control for displaying the candidate video information in the search result set, and if an interaction operation for the display control is received, displaying the candidate video information according to the matching degree between at least one video information in the search result set and the search information.
[0073] According to an exemplary implementation of the present disclosure, one video information in the search result set is displayed in the conversation interaction page, and the video information is determined based on the following steps: sorting at least one video information according to the matching degree between at least one video information in the search result set and the search information; and using the video information with the highest matching degree as the video information to be displayed in the conversation interaction page from the sorted at least one video information.
[0074] According to an exemplary implementation of the present disclosure, the method is executed at an in-vehicle computing device in a vehicle, and the method further includes: determining a position associated with the voice information, where the position represents the occurrence position of the voice information in the vehicle; and where displaying the conversation interaction page further includes: displaying the conversation interaction page at a display device corresponding to the position in the vehicle.
[0075] Example devices and equipment
[0076] Figure 10 The block diagram of the device 1000 for interaction according to some implementations of the present disclosure is shown. The device includes: an acquisition module 1010 configured to acquire a voice message; an information determination module 1020 configured to determine search information according to the voice message if the content type of the voice message is a search type; a result determination module 1030 configured to determine a search result set corresponding to the search information according to the search information, wherein the search result set includes at least one video message; and a display module 1040 configured to display a dialogue interaction page, wherein at least one video message in the search result set is displayed in the dialogue interaction page.
[0077] According to an example implementation of the present disclosure, the content type of the voice message is determined based on the following modules: a conversion module configured to convert the voice message into text data; a first type determination module configured to perform semantic analysis on the text data to determine the semantic content expressed by the text data and determine the content type based on the semantic content; or a second type determination module configured to identify keywords associated with the search type in the text data to determine the content type.
[0078] According to an example implementation of the present disclosure, the result determination module includes: a matching degree determination module configured to use a machine learning model to determine the matching degree between a target video message in multiple video messages in a video repository and the search information; and an addition module configured to add the target video message to the search result set if it is determined that the matching degree between the target video message and the search information meets a predetermined condition.
[0079] According to an example implementation of the present disclosure, the dialogue interaction page includes: summary information of the target video message in at least one video message, and a playback control for playing the target video message in at least one video message, and the summary information includes at least any one of the following: the cover, title, author, time length, and release time of the target video message.
[0080] According to an example implementation of the present disclosure, the device further includes: a first playback module configured to play the target video message if an interaction operation on the playback control is received.
[0081] According to an example implementation of the present disclosure, the apparatus further includes: a second playback module configured to play candidate video information other than the target video information in the search result set if at least any one of the following is detected: a playback request for playing the candidate video information; a playback end event associated with the target video information; and a playback pause event associated with the target video information.
[0082] According to an example implementation of the present disclosure, the apparatus further includes: a candidate video information display module configured to display candidate video information in the search result set according to the matching degree between at least one video information in the search result set and the search information if a trigger event for displaying the candidate video information in the search result set is detected.
[0083] According to an example implementation of the present disclosure, the candidate video information display module further includes: a first display module configured to display candidate video information in the search combination according to the matching degree between at least one video information in the search result set and the search information if a playback end event associated with the target video information is detected; and a second display module configured to display candidate video information according to the matching degree between at least one video information in the search result set and the search information if an interaction operation on the display control is received, and the dialogue interaction page further includes a display control for displaying candidate video information in the search result set.
[0084] According to an example implementation of the present disclosure, one video information in the search result set is displayed in the dialogue interaction page, and the video information is determined based on the following modules: a sorting module configured to sort at least one video information according to the matching degree between at least one video information in the search result set and the search information; and a selection module configured to use the video information with the highest matching degree as the video information displayed in the dialogue interaction page from the sorted at least one video information.
[0085] According to an example implementation of the present disclosure, the apparatus is implemented at an in-vehicle computing device in a vehicle, and the apparatus further includes: a location determination module configured to determine a location associated with the voice information, where the location represents the occurrence location of the voice information in the vehicle; and a location-based display module configured to display the dialogue interaction page at a display device corresponding to the location in the vehicle.
[0086] Figure 11 The block diagram of a device 1100 capable of implementing multiple implementations of the present disclosure is shown. It should be understood that Figure 11 The shown computing device 1100 is merely exemplary and should not constitute any limitation to the functions and scope of the implementations described herein. Figure 11The computing device 1100 shown can be used to implement the methods described above.
[0087] As Figure 11 shown, the computing device 1100 is in the form of a general-purpose computing device. The components of the computing device 1100 can include, but are not limited to, one or more processors or processing units 1110, a memory 1120, a storage device 1130, one or more communication units 1140, one or more input devices 1150, and one or more output devices 1160. The processing unit 1110 can be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 1120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing ability of the computing device 1100.
[0088] The computing device 1100 generally includes multiple computer storage media. Such media can be any available media accessible to the computing device 1100, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1120 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1130 can be removable or non-removable media and can include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device 1100.
[0089] The computing device 1100 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 11 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 1120 can include a computer program product 1125 having one or more program modules configured to perform the various methods or actions of the various implementations of the present disclosure.
[0090] The communication unit 1140 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1100 can be implemented in a single computing cluster or multiple computer machines that are capable of communicating via a communication link. Thus, the computing device 1100 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0091] The input device 1150 can be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 1160 can be one or more output devices such as a display, speakers, printer, etc. The computing device 1100 can also communicate with one or more external devices (not shown) as needed via the communication unit 1140, external devices such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the computing device 1100, or communicate with any device that enables the computing device 1100 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0092] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the methods described above. According to an exemplary implementation of the present disclosure, there is provided a computer program product having a computer program stored thereon, and the program implements the methods described above when executed by a processor.
[0093] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0094] These computer-readable program instructions may be provided to a processing unit of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0095] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0096] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction may include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system for performing the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0097] The implementations of the present disclosure have been described above. The description is illustrative, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skilled persons in the art to understand the implementations disclosed herein.
Claims
1. An interaction method, comprising: Obtaining a voice message; If the content type of the voice message is a search type, determining search information according to the voice message; Determining a search result set corresponding to the search information according to the search information, wherein the search result set includes at least one video message; and Displaying a dialogue interaction page, wherein at least one video message in the search result set is displayed in the dialogue interaction page.
2. The method according to claim 1, wherein the content type of the voice message is determined based on the following steps: Converting the voice message into text data; Performing semantic analysis on the text data to determine the semantic content expressed by the text data, and determining the content type based on the semantic content; or Identifying keywords associated with the search type in the text data to determine the content type.
3. The method according to claim 1, wherein determining the search result set corresponding to the search information according to the search information includes: Using a machine learning model to determine the matching degree between a target video message in a plurality of video messages in a video repository and the search information; And If it is determined that the matching degree between the target video message and the search information meets a predetermined condition, adding the target video message to the search result set.
4. The method according to claim 1, wherein the dialogue interaction page comprises: Summary information of the target video message among the at least one video message, and a playback control for playing the target video message among the at least one video message, where the summary information includes at least any one of the following: the cover, title, author, time length, and release time of the target video message.
5. The method according to claim 4, further comprising: If an interaction operation on the playback control is received, play the target video message.
6. The method according to claim 5, further comprising: If at least any one of the following is detected, play a candidate video message other than the target video message in the search result set: A playback request for playing the candidate video message; A playback end event associated with the target video message; and A playback pause event associated with the target video message.
7. The method according to claim 5, further comprising: If a trigger event for displaying a candidate video message in the search result set is detected, display the candidate video message according to the matching degree between at least one video message in the search result set and the search information.
8. The method according to claim 7, wherein displaying the candidate video message includes: If a playback end event associated with the target video message is detected, display the candidate video message in the search combination according to the matching degree between at least one video message in the search result set and the search information; Or The dialogue interaction page further includes a display control for displaying the candidate video message in the search result set, and if an interaction operation on the display control is received, display the candidate video message according to the matching degree between at least one video message in the search result set and the search information.
9. The method according to claim 3, wherein one video information in the search result set is displayed in the dialogue interaction page, and the video information is determined based on the following steps: Sort the at least one video information according to the degree of match between the at least one video information in the search result set and the search information; and From the sorted at least one video information, use the video information with the highest degree of match as the video information to be displayed in the dialogue interaction page.
10. The method according to claim 1, wherein the method is performed at an in-vehicle computing device in a vehicle, and the method further comprises: Determine the position associated with the voice information, where the position indicates the occurrence position of the voice information in the vehicle; And wherein displaying the dialogue interaction page further includes: displaying the dialogue interaction page at a display device corresponding to the position in the vehicle.
11. An interaction device, comprising: An acquisition module configured to acquire a voice information; An information determination module configured to determine search information according to the voice information if the content type of the voice information is a search type; A result determination module configured to determine a search result set corresponding to the search information according to the search information, wherein the search result set includes at least one video information; and A display module configured to display a dialogue interaction page, wherein at least one video information in the search result set is displayed in the dialogue interaction page.
12. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the electronic device to execute the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, having stored thereon a computer program, which when executed by a processor causes the processor to implement the method according to any one of claims 1 to 10.