Information interaction method and device, electronic equipment and storage medium
By introducing an intelligent Q&A portal into short video/image collection platforms and utilizing AI-generated guidance data, the system addresses users' information-related questions when pausing or swiping, enabling personalized question guidance and efficient information acquisition, thereby improving user experience and information delivery efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
- Filing Date
- 2025-12-25
- Publication Date
- 2026-05-26
Smart Images

Figure CN122087181A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to an information interaction method, apparatus, electronic device and storage medium. Background Technology
[0002] In the realm of short video / photo galleries, users pausing to watch or swiping to stay on screen are common scenarios in the content consumption process.
[0003] In the relevant technologies for this scenario, the platform often displays visual elements such as image search icons, floating product recognition entrances, or related recommended content cards on the interface when the user pauses viewing or swipes. Its core is to use a function entry + result list as the main display format. For example, clicking the image search icon will jump to the image search results page, and clicking the product recognition entrance will pop up a list of product links.
[0004] However, in reality, when users pause or swipe, they often have questions about the content's background, key points, and details. The aforementioned technical solutions only focus on traffic redirection and conversion, providing only visual entry points for functions like product search and image search, without designing corresponding interactive display formats for addressing these informational questions. This not only results in a limited functionality for the scenario but also restricts the targeting of information services, making it difficult to meet users' needs for immediate and accurate information access in the current context. Furthermore, it reduces the platform's user experience and information delivery efficiency in this scenario. Summary of the Invention
[0005] This disclosure provides an information interaction method, apparatus, and electronic device to at least solve the problems of low user experience and low information transmission efficiency in related technologies. The technical solution of this disclosure is as follows:
[0006] According to a first aspect of the present disclosure, an information interaction method is provided, comprising:
[0007] Play the target video data in the display interface;
[0008] At least one intelligent Q&A entry is displayed on the video playback page of the target video data, and the intelligent Q&A entry displays an intelligent Q&A icon and question guidance data;
[0009] In response to a trigger operation for any of the intelligent question-and-answer entry points, intelligent question-and-answer dialogue information is displayed through an information display unit. The intelligent question-and-answer dialogue information includes the question guidance data corresponding to the intelligent question-and-answer entry point and the response data to the question guidance data. The question guidance data is guidance data generated using artificial intelligence analysis.
[0010] In one embodiment, the at least one intelligent question-and-answer entry point is displayed in the form of question-and-answer tags, and displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes:
[0011] The question and answer tags corresponding to the at least one smart question and answer entry are displayed on the video playback page in a preset sorting manner. The preset sorting manner includes sorting in descending order according to the matching degree between the question guidance data corresponding to the smart question and answer entry and the user account, or sorting in descending order according to the trigger popularity of the smart question and answer entry.
[0012] In one embodiment, the video playback page includes a text display area for displaying the publishing account information and title information of the target video data. Displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes:
[0013] If the user account's account tag meets the display conditions, hide the publishing account information and title information of the target video data displayed in the text display area.
[0014] The intelligent question-and-answer icon and the trigger control corresponding to each intelligent question-and-answer entry are displayed in the text display area. The trigger control displays the question guidance data corresponding to the intelligent question-and-answer entry.
[0015] In one embodiment, the video playback page includes a comment information display area, and displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes:
[0016] The comment information of the target video data is displayed in the comment information display area. The comment information includes at least one target comment information, which is the comment information that triggers a question and answer event. The target comment information includes text comment data and a smart question and answer identifier.
[0017] In one embodiment, when the text comment data is of the query type, the text comment data serves as the question guidance data for the intelligent question-and-answer entry point; or, when the text comment data is not of the query type, the intelligent question-and-answer identifier also displays question guidance data, which is guidance data generated based on the text comment data.
[0018] In one embodiment, displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes:
[0019] When the target video data is paused, or when the comment data of the target video data displayed in the comment information display area includes target comment data that triggers a question-and-answer event, at least one smart question-and-answer entry point is displayed on the video playback page of the target video data.
[0020] In one embodiment, the intelligent question-and-answer dialogue information further includes at least one piece of follow-up question guidance data, which includes at least one of the following:
[0021] Question guidance data corresponding to untriggered intelligent question-answering entry points;
[0022] The question guidance data is obtained by further processing the question guidance data and / or response data corresponding to the triggered intelligent question answering entry.
[0023] In one embodiment, the response data includes text response data and reference video data, wherein the reference video data is video data associated with the question guidance data and / or response data.
[0024] According to a second aspect of the present disclosure, an information interaction device is provided, comprising:
[0025] The playback unit is configured to play the target video data in the display interface.
[0026] The first display unit is configured to display at least one intelligent question-and-answer entry on the video playback page of the target video data, wherein the intelligent question-and-answer entry displays an intelligent question-and-answer identifier and question guidance data;
[0027] The second display unit is configured to perform a trigger operation in response to any of the intelligent question-and-answer entry points, and to display intelligent question-and-answer dialogue information through the information display unit. The intelligent question-and-answer dialogue information includes the question guidance data corresponding to the intelligent question-and-answer entry point and the response data to the question guidance data. The question guidance data is guidance data generated based on artificial intelligence analysis.
[0028] In one embodiment, the at least one intelligent question-and-answer entry point is displayed in the form of question-and-answer tags, and displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes:
[0029] The question and answer tags corresponding to the at least one smart question and answer entry are displayed on the video playback page in a preset sorting manner. The preset sorting manner includes sorting in descending order according to the matching degree between the question guidance data corresponding to the smart question and answer entry and the user account, or sorting in descending order according to the trigger popularity of the smart question and answer entry.
[0030] In one embodiment, the video playback page includes a text display area for displaying the publishing account information and title information of the target video data. Displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes:
[0031] If the user account's account tag meets the display conditions, hide the publishing account information and title information of the target video data displayed in the text display area.
[0032] The intelligent question-and-answer icon and the trigger control corresponding to each intelligent question-and-answer entry are displayed in the text display area. The trigger control displays the question guidance data corresponding to the intelligent question-and-answer entry.
[0033] In one embodiment, the video playback page includes a comment information display area, and displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes:
[0034] The comment information of the target video data is displayed in the comment information display area. The comment information includes at least one target comment information, which is the comment information that triggers a question and answer event. The target comment information includes text comment data and a smart question and answer identifier.
[0035] In one embodiment, when the text comment data is of the query type, the text comment data serves as the question guidance data for the intelligent question-and-answer entry point; or, when the text comment data is not of the query type, the intelligent question-and-answer identifier also displays question guidance data, which is guidance data generated based on the text comment data.
[0036] In one embodiment, displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes:
[0037] When the target video data is paused, or when the comment data of the target video data displayed in the comment information display area includes target comment data that triggers a question-and-answer event, at least one smart question-and-answer entry point is displayed on the video playback page of the target video data.
[0038] In one embodiment, the intelligent question-and-answer dialogue information further includes at least one piece of follow-up question guidance data, which includes at least one of the following:
[0039] Question guidance data corresponding to untriggered intelligent question-answering entry points;
[0040] The question guidance data is obtained by further processing the question guidance data and / or response data corresponding to the triggered intelligent question answering entry.
[0041] In one embodiment, the response data includes text response data and reference video data, wherein the reference video data is video data associated with the question guidance data and / or response data.
[0042] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement any of the information interaction methods provided in the first aspect.
[0043] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the information interaction methods provided in the first aspect.
[0044] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including instructions that, when executed by a processor of an electronic device, enable the electronic device to perform any of the information interaction methods provided in the first aspect.
[0045] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:
[0046] The information interaction method, apparatus, electronic device, and storage medium provided in this disclosure play target video data in a display interface and display at least one intelligent question-and-answer entry point on the video playback page of the target video data. The intelligent question-and-answer entry point displays an intelligent question-and-answer identifier and question guidance data. In response to a trigger operation for any intelligent question-and-answer entry point, intelligent question-and-answer dialogue information is displayed through an information display unit. The intelligent question-and-answer dialogue information includes the question guidance data corresponding to the intelligent question-and-answer entry point and response data to the question guidance data. The question guidance data is guidance data generated based on artificial intelligence analysis. By employing the information interaction method, apparatus, electronic device, and storage medium provided in this disclosure, and leveraging the multimodal deep understanding capabilities of a large model, personalized question guidance data displayed through intelligent question-and-answer entry points guides users to initiate intelligent dialogues to obtain accurate information in user video consumption scenarios. Furthermore, through a seamless process of scene triggering, question-and-answer interaction, and deep feedback, combined with the deep analysis of multimodal video information by artificial intelligence, accurate matching of question guidance content with user needs and high-quality responses based on video context are achieved. Ultimately, this significantly reduces the operational threshold and time cost for users to obtain information, and greatly improves the accuracy and overall efficiency of information transmission.
[0047] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0048] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0049] Figure 1 This is a flowchart illustrating an information interaction method according to an exemplary embodiment.
[0050] Figure 2 This is a schematic diagram of an intelligent question-and-answer dialogue page according to an exemplary embodiment.
[0051] Figure 3 This is a schematic diagram of a video playback pause interface according to an exemplary embodiment.
[0052] Figure 4 This is a schematic diagram of a video playback pause interface according to another exemplary embodiment.
[0053] Figure 5 This is a detailed flowchart illustrating step 104 according to an exemplary embodiment.
[0054] Figure 6 This is a schematic diagram of a video playback interface that hides the publishing account and title, according to an exemplary embodiment.
[0055] Figure 7 This is a schematic diagram of a video playback interface containing comment information, according to an exemplary embodiment.
[0056] Figure 8 This is a block diagram illustrating an information interaction device according to an exemplary embodiment.
[0057] Figure 9 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation
[0058] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0059] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0060] Figure 1 This is a flowchart illustrating an information interaction method according to an exemplary embodiment. This embodiment uses the application of the method to a terminal as an example for illustration. It is understood that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes steps 102 to 106, wherein:
[0061] Step 102: Play the target video data in the display interface.
[0062] In this embodiment, the target video data refers to video content provided by the platform that is available for user browsing, including short videos, image gallery videos, etc. It can consist of image data such as video keyframes, single images in the image gallery, and entity elements in the scene; audio data such as background noise, narration, and sound effects; and text data such as subtitles, on-screen text, video titles, and descriptions. This multimodal information data can serve as a data source for subsequent multimodal semantic understanding. The display interface is the video playback carrier on the terminal, which can be the main interface of a short video app (Application), an image gallery playback page, a mini-program playback window, etc. It has event monitoring and interface call capabilities, enabling it to capture user interaction behavior in real time and trigger subsequent processes.
[0063] Step 104: Display at least one intelligent Q&A entry point on the video playback page of the target video data. The intelligent Q&A entry point displays an intelligent Q&A icon and question guidance data.
[0064] For example, during video browsing, users can interact with the target video data, such as pausing playback. Upon detecting this interaction, the terminal will display one or more intelligent Q&A entry points containing question-guided data on the video playback page. These interaction events are user behaviors generated during video browsing and can be identified as triggers indicating a user's desire to delve deeper into the content. The detection of these events can be achieved by collecting user behavior data through terminal-level tracking and combining it with behavior recognition algorithms for real-time capture. The intelligent Q&A entry point serves as the interactive carrier for the question-guided data, and its display format can be a floating label, a semi-transparent pop-up label, etc., located within the video playback page.
[0065] Step 106: In response to a trigger operation for any intelligent question-and-answer entry point, intelligent question-and-answer dialogue information is displayed through the information display unit. The intelligent question-and-answer dialogue information includes the question guidance data corresponding to the intelligent question-and-answer entry point and the response data to the question guidance data. The question guidance data is guidance data generated by artificial intelligence analysis.
[0066] In this embodiment, the user can trigger the intelligent Q&A entry point through interactive methods such as clicking or long-pressing. After the terminal detects the triggering operation, it will display intelligent Q&A dialogue information in the information display unit. The information display unit can be an intelligent dialogue information page, or a half-screen or floating card used to display intelligent Q&A dialogue information. This embodiment does not specifically limit the information display unit; any carrier capable of displaying intelligent Q&A dialogue information is applicable. For clarity, the following embodiments use an intelligent Q&A dialogue page as an example for explanation. For instance, after the terminal detects a triggering operation on the intelligent Q&A entry point, it can jump to the intelligent Q&A dialogue page, where intelligent Q&A dialogue information is displayed. Wherein, as... Figure 2 As shown, the intelligent question-and-answer dialogue information includes the question guidance data corresponding to the intelligent question-and-answer entry and the response data from AI (Artificial Intelligence).
[0067] Among them, the question-guided data is based on user needs and video content. It is generated by deep analysis of the video stream in real time or preprocessing through a multimodal large language model. The analysis objects include multimodal information such as image entities, speech-to-text, and subtitle keywords of the target video data, as well as related data of information interaction events, such as users' historical question and answer preferences, pause position screen features, popular questions of similar users, etc., so as to fully understand the core theme, key entities, style, emotion and potential questions of the video.
[0068] The response data is generated by transforming multimodal information from the video into a unified semantic vector using a multimodal large model. This vector is then combined with question-guided data to construct inference prompts. These prompts are input into the large language model to generate initial responses. After content review and format standardization, the final response data is presented to the user. This approach encourages users to actively engage in question-and-answer interactions, increasing user dwell time and multi-turn interactions, thereby improving user engagement and stickiness.
[0069] In one example, the intelligent question-and-answer dialogue page can also be equipped with a back button, which allows users to jump back to the original video playback page with one click. The original video playback page can retain the progress and scene state at the time of pause. This disclosure does not limit the specific types of multimodal large models and large language models; any model with multimodal semantic understanding and natural language generation capabilities is applicable to this disclosure.
[0070] For example, Figure 3 This is a video showcasing a cream-colored interior design style. When the user pauses the video, this pause can be seen as a strong signal that the user has questions and wants to learn more. At this point, relevant content from the current frame can be recommended, displaying the corresponding intelligent Q&A entry point. This entry point can be displayed as a tag, with guided questions such as "Ask AI what interior design style this is," "Ask AI the difference between cream-colored and vintage styles," and "Ask AI for a 90㎡ cream-colored interior price estimate." "Ask AI" is the intelligent Q&A identifier, indicating that the current tag is an entry point for intelligent Q&A, allowing users to engage in a Q&A dialogue. After clicking the above tags, the page will redirect to an AI multi-turn dialogue page (i.e., the intelligent Q&A dialogue page). Figure 2 As shown, on the intelligent Q&A dialogue page, questions are automatically sent based on question-guided data, and responses are automatically generated by AI based on the question-guided data. Users can subsequently engage in multiple rounds of dialogue to obtain information on this intelligent Q&A dialogue page, or return to the video playback page to continue watching the video.
[0071] The information interaction method provided in this disclosure plays target video data in a display interface and displays at least one intelligent question-and-answer entry point on the video playback page of the target video data. The intelligent question-and-answer entry point displays an intelligent question-and-answer identifier and question guidance data. In response to a trigger operation for any intelligent question-and-answer entry point, intelligent question-and-answer dialogue information is displayed through an information display unit. The intelligent question-and-answer dialogue information includes the question guidance data corresponding to the intelligent question-and-answer entry point and response data to the question guidance data. The question guidance data is generated using artificial intelligence analysis. By employing the information interaction method provided in this disclosure, leveraging the multimodal deep understanding capabilities of a large model, in a user's video consumption scenario, personalized question guidance data displayed through the intelligent question-and-answer entry point guides the user to initiate an intelligent dialogue to obtain accurate information. Furthermore, through a seamless process of scene triggering, question-and-answer interaction, and deep feedback, combined with the deep analysis of multimodal information in the video by the large model, accurate matching of question guidance content with user needs and high-quality responses based on the video context are achieved. Ultimately, this significantly reduces the operational threshold and time cost for users to obtain information, and greatly improves the accuracy and overall efficiency of information transmission.
[0072] In one exemplary embodiment, at least one intelligent question-and-answer entry point is displayed in the form of question-and-answer tags. In step 104, at least one intelligent question-and-answer entry point is displayed on the video playback page of the target video data, including:
[0073] Display the question and answer tags corresponding to at least one smart question and answer entry point on the video playback page in a preset sorting manner. The preset sorting manner includes sorting in descending order according to the matching degree between the question guidance data corresponding to the smart question and answer entry point and the user account, or sorting in descending order according to the trigger popularity of the smart question and answer entry point.
[0074] In this embodiment, question-and-answer tags can serve as display carriers for intelligent question-and-answer entry points, and are displayed on the video playback page according to a preset sorting method. The question-and-answer tags can employ a lightweight, highly recognizable visual design, such as using high-contrast fonts for the text, font sizes adapted to the terminal screen size, and each tag can be accompanied by an AI identifier icon (i.e., an intelligent question-and-answer identifier) to clearly identify its functional attributes.
[0075] The preset sorting method can be to sort the question guidance data corresponding to the intelligent Q&A entry point in descending order based on the matching degree between the data and the user account. The matching degree between the question guidance data and the user account can be determined by calculating the similarity between the related information of the question guidance data and the user account. The related information of the user account can include user historical behavior data (historical Q&A records, types of Q&A tags clicked, video theme preferences, dwell time distribution, etc.) and account attribute data (interest areas filled in during registration, content preference tags set by the platform, etc.).
[0076] For example, the core keywords of each question-guided data can be extracted through a multimodal large model. For example, the keywords for "cream-style decoration budget" are "cream-style", "decoration", and "budget". A semantic vector is generated, and an interest vector is generated based on the user account's associated information. Finally, the cosine similarity between the semantic vector and the interest vector is calculated. The similarity score is the matching degree, and the question and answer tags are arranged in descending order according to the score on the video playback page.
[0077] The preset sorting method can also be a descending order based on the trigger popularity of the smart Q&A entry point. Trigger popularity refers to the frequency with which the Q&A tags corresponding to the smart Q&A entry point are triggered by users in similar scenarios within a preset statistical period. This can include the cumulative number of tag clicks, the click-through rate after tag exposure, and the effective trigger rate of completing the Q&A process after clicking. Furthermore, such as... Figure 4 As shown, the question and answer tags corresponding to the intelligent question and answer entry can coexist with the regular image search and be displayed together in the video playback area.
[0078] The information interaction method provided in this disclosure, during video playback or image gallery consumption, sorts question and answer tags based on user profiles and trigger popularity through multimodal parsing and a large language model, generating personalized intelligent question and answer entry points. This makes the presented question guidance data more closely aligned with individual needs. Through semantic question tags, users can find questions of interest more quickly and get direct answers, improving the efficiency of user information acquisition and reducing the cost of user information acquisition.
[0079] In one exemplary embodiment, the video playback page includes a text display area for displaying the publishing account information and title information of the target video data. (See also...) Figure 5 As shown, in step 104, displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data may include the following steps 502 to 504, wherein:
[0080] Step 502: If the user account's account tag meets the display conditions, hide the publishing account information and title information of the target video data displayed in the text display area.
[0081] Step 504: Display the intelligent question and answer icon and the trigger control corresponding to each intelligent question and answer entry in the text display area. The trigger control displays the question guidance data corresponding to the intelligent question and answer entry.
[0082] In this embodiment, the video playback page includes a text display area for displaying the publishing account information and title information of the target video data. When the user's account tags meet the display conditions, trigger controls corresponding to each intelligent Q&A entry can be displayed through this text display area. These trigger controls can display question guidance data and an intelligent Q&A identifier corresponding to the intelligent Q&A entry. Users can enter the intelligent Q&A dialogue page by clicking or long-pressing the trigger control and engage in intelligent dialogue based on the question guidance data. When displaying the trigger controls corresponding to the intelligent Q&A entry, the publishing account information and title information of the target video data are hidden.
[0083] The user account's account tags can be structured data tags generated by the platform based on user behavior data and attribute information. These can include interest tags (such as home renovation enthusiasts, food explorers, and digital product reviewers), behavioral tags (such as high-frequency Q&A users, in-depth content consumers, and active users who pause and interact), and attribute tags (such as membership level and content preference categories). The display conditions are the basis for determining whether to enable the display of trigger controls corresponding to each intelligent Q&A entry point in the text area. These conditions can be based on user tags, scenarios, and content. For example, if a user's account tag is an interest tag, specifically "home renovation enthusiast," and its semantic similarity to the current target video data's theme tag ("cream-style home renovation") is ≥0.7, and the user has triggered a pause operation, then the user's account tag meets the display conditions.
[0084] For example, such as Figure 6 As shown, the user account's interest tag is interior design, and the target video's theme tag is "creamy style interior design." Assuming the semantic similarity between the user account's tag and the target video's theme tag is 0.85, if the user actively clicks the pause button, then all display conditions are met. At this point, the original posting account information and video title "90㎡ Creamy Style Interior Design Completed!" in the text display area are hidden, and the corresponding trigger controls for the two questions "What kind of interior design is this?" and "Difference between creamy style and vintage style" will be displayed in this text display area.
[0085] The information acquisition method provided in this disclosure reuses the original text display area of the video playback page without adding a new independent layout space. This avoids the problem of the intelligent Q&A entrance obscuring the video screen, reduces operation costs through conditional triggering, and ensures that the display of the intelligent Q&A entrance is targeted through precise matching of account tags and scenarios, which can further improve users' willingness to interact and the efficiency of information acquisition.
[0086] In one exemplary embodiment, the video playback page includes a comment information display area. In step 104, at least one intelligent question-and-answer entry point is displayed on the video playback page of the target video data, including:
[0087] The comment information display area displays comment information for the target video data. The comment information includes at least one target comment, which is the comment that triggered the question-and-answer event. The target comment includes text comment data and a smart question-and-answer identifier.
[0088] In this embodiment of the disclosure, the video playback page may include a comment information display area. The comment information display area is used to display native comment information and target comment information. The comment information display area can retain the original functions of the area, and at the same time, it can also filter out target comment information from a large number of comments for differentiated display. For example, a smart question and answer label can be added to the text comment data corresponding to the target comment information, and the user can be guided to trigger question and answer interaction based on the target comment information through the smart question and answer label.
[0089] The comment information display area prioritizes displaying all valid comments on the target video data according to a preset sorting strategy (such as descending time or descending popularity). Target comment information can be comments that trigger question-and-answer events; that is, the content expressed by users in the comments is strongly related to the target video content, representative, and can serve as a trigger for interactive question-and-answer sessions. Target comment information can be obtained from native comment information through modal big data modeling and semantic analysis techniques, or it can be question-guided data strongly related to the target video data obtained based on multimodal big data modeling analysis.
[0090] For example, target comment information can be displayed in the following forms:
[0091] In the comment information display area, in addition to valid original comment information, only target comment information filtered from the original comment information is displayed; or, in addition to valid original comment information, the comment information display area can display both target comment information filtered from the original comment information and target comment information obtained based on question-guided data; or, in the comment information display area, in addition to valid original comment information, only target comment information filtered from the original comment information is displayed, and when the user pauses the video, the target comment information obtained based on question-guided data is displayed on the video playback page.
[0092] For example, such as Figure 7 As shown, the target video data is a real-life display of a 90㎡ cream-themed interior design. One comment in the comment information display area asks, "What kind of interior design style is this?". Assuming that multimodal large model analysis shows the semantic similarity between this comment and the core entity "cream-themed interior design" in the target video data meets the filtering criteria (e.g., greater than a preset similarity threshold), this comment is selected as the target comment. When displayed in the comment information display area, the comment text data will be fully preserved, and a "Ask AI" icon will be added after the text data. When the user clicks the icon, an intelligent question-and-answer dialogue page will automatically pop up, with the guiding question being "What kind of interior design style is this?". AI will then respond to this guiding question by combining visual features, subtitles, and other information.
[0093] The information acquisition method provided in this disclosure transforms scattered information needs in the comment section into standardized and reusable intelligent Q&A entry points by reusing the native comment information in the comment information display area. This reduces the cost of the platform actively generating question guidance data, improves the accuracy of Q&A content through real user feedback, and further strengthens the interaction between users and between users and the platform by leveraging the interactive attributes of the comment section, thereby improving the efficiency of information transmission.
[0094] In one exemplary embodiment, when the text comment data is of the query type, the text comment data serves as the question guidance data for the intelligent question and answer entry point; or, when the text comment data is not of the query type, the intelligent question and answer entry point identifier also displays question guidance data, which is guidance data generated based on the text comment data.
[0095] In this embodiment, in the comment information display area, when the text comment data, i.e., the native comment information, is determined to be of the inquiry type, it already contains a clear information-related question and request, and has been verified by the preliminary screening process to be strongly related to the video content. Therefore, this text comment data can be directly used as the question guidance data for the intelligent question-and-answer entry point without the need for additional generation. When the text comment data, i.e., the native comment information, is not of the inquiry type, it is necessary to mine the user's potential information needs behind the comment, such as the desire to understand relevant details or compare similar content implied in the expression of opinions. Targeted question guidance data is then generated through a multimodal big data model and embedded in the intelligent question-and-answer entry point identifier.
[0096] For example, such as Figure 7 As shown, if the text comment data is "What kind of decorating style is this?", which is a query type, then this text comment data can be directly used as the question guidance data for the intelligent question-and-answer entry point. Simply add the intelligent question-and-answer tag "Ask AI" after the comment. If the text comment data is "A cream-style living room looks particularly spacious", which is a non-query type, then based on the non-query type text comment data, combined with the multimodal information of the target video, such as keyframe images, speech-to-text, subtitle keywords, core entities, etc., the necessary context is constructed to generate the corresponding question guidance data. This ensures that it is strongly related to the comment content and does not deviate from the video theme. The intelligent question-and-answer tag is displayed in the format of "Ask AI: Question Guidance Data", such as "Ask AI: How to design a cream-style living room to make it look more spacious".
[0097] The information acquisition method provided in this disclosure, through a differentiated question-guided data strategy, not only fully utilizes the explicit demand value of inquiry-type comments, but also explores the potential demand of non-inquiry-type comments, significantly expanding the coverage scenarios of the intelligent question-and-answer entry point. At the same time, the guiding data is generated based on real user comments, which is more in line with the needs of ordinary users, improving the accuracy of question-and-answer interaction and user participation.
[0098] In one exemplary embodiment, at least one intelligent question-and-answer entry point is displayed on the video playback page of the target video data, including:
[0099] When the target video data is paused, or when the target video data's comment data displayed in the comment information display area includes the target comment data that triggered the question-and-answer event, at least one smart question-and-answer entry point is displayed on the target video data's video playback page.
[0100] In this embodiment of the disclosure, when the target video data is paused, that is, when the user actively switches the video playback status from playing to paused, this pause operation is a direct and strong signal that the user has an information need. Users usually pause the video actively because they need to focus on the details of a certain frame or think about the content of that frame. At this time, at least one intelligent question and answer entry can be displayed on the video playback page of the target video data so that users can efficiently and directly obtain the information they need through the intelligent question and answer entry.
[0101] If the comment data of the target video data displayed in the comment information display area includes target comment data that triggers a question-and-answer event, that is, if the platform detects target comment data that triggers a question-and-answer event from the currently displayed comment data, since the target comment data is a filtered comment that meets the conditions and contains informational questions or potential needs, then the corresponding intelligent question-and-answer entry can also be displayed on the video playback page based on the target comment data, thereby providing users with a convenient and efficient way to obtain information.
[0102] The information interaction method provided in this disclosure has independent and complementary triggering logics for actively pausing and displaying the intelligent Q&A entry based on target comment data. Through the above-mentioned information interaction events, users' willingness to interact and the efficiency of information acquisition can be further improved.
[0103] In one exemplary embodiment, the intelligent question-answering dialogue information further includes at least one piece of follow-up question guidance data, which includes at least one of the following:
[0104] Question guidance data corresponding to untriggered intelligent question-answering entry points;
[0105] The question guidance data is obtained by further processing the question guidance data and / or response data corresponding to the triggered intelligent question answering entry.
[0106] In this embodiment of the disclosure, the intelligent question-and-answer dialogue information may include follow-up question guidance data in addition to the question guidance data and AI response data corresponding to the intelligent question-and-answer entry. The follow-up question guidance data may be the question guidance data corresponding to the untriggered intelligent question-and-answer entry, or it may be an extended and related question generated by performing semantic analysis and in-depth mining on the triggered question guidance data and / or AI response data through a large model, thereby predicting the user's potential subsequent needs and upgrading from passive response to active guidance.
[0107] The question guidance data for untriggered intelligent Q&A entry points comes from the question guidance data for intelligent Q&A entry points displayed on the video playback page that the user has not clicked to trigger. That is, after a user selects a Q&A tag on the video page to enter the dialogue page, the system filters content strongly related to the current dialogue topic from the remaining untriggered tags as follow-up question options, ensuring that the follow-up question is consistent with the core theme of the video and avoids deviating from the user's focus. Follow-up question guidance data can be displayed below the AI response data, titled with relevant follow-up questions, using a horizontally scrolling tag group layout. When a user clicks on any tag, the corresponding question guidance data is automatically sent to the AI as a new question. The AI generates a targeted response in real time, while simultaneously updating the follow-up question guidance data list, removing triggered tags and adding new relevant options.
[0108] Based on the question guidance data and / or response data corresponding to the triggered intelligent question-answering entry, extended processing is performed to obtain the question guidance data. This is generated by a large model using the semantic vector of the triggered question guidance data, keywords and core entities of the AI response data as input, combined with multimodal information from the target video data, and based on a preset extended prompt template, to create an initial list of follow-up questions. When the user clicks, a question is automatically sent. When the AI responds, it will refer to the previous dialogue context. Each time an extended follow-up question is triggered, the system will generate a new extended follow-up question based on the new dialogue content, continuously expanding the depth of question-answering.
[0109] The information interaction method provided in this embodiment, through the design of two types of follow-up questions to guide data, not only makes full use of the question resources generated in the early stage, but also extends the potential needs through the large model, effectively extending the time users stay in the question-and-answer scenario, improving the completeness of information acquisition and the coherence of the interactive experience, and further strengthening the platform's intelligent service capabilities.
[0110] In one exemplary embodiment, the response data includes text response data and reference video data, which is video data associated with the question guidance data and / or response data.
[0111] In this embodiment, the response data output by the intelligent question-and-answer dialogue page can adopt a dual-modal presentation format of precise text answers and concrete video supplements. Structured text directly responds to user questions, while reference video data concretizes abstract information and contextualizes complex information, achieving a deep integration of the precision of text answers and the intuitiveness of video content, significantly reducing the cost of information comprehension for users. The two types of data are displayed collaboratively and complement each other, jointly forming a complete and hierarchical question-and-answer response system. The reference video data consists of video content selected by the platform from its video library that is related to the core requirements of the question guidance data and / or the key information of the text response data. Through dynamic video visuals, the abstract information in the text response is presented concretely, compensating for the limitations of purely text-based answers.
[0112] It should be understood that, although Figures 1-7 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figures 1-7 At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.
[0113] It is understood that the same / similar parts between the various embodiments of the methods described above in this specification can be referred to each other. Each embodiment focuses on the differences from other embodiments, and relevant parts can be referred to the description of other method embodiments.
[0114] Figure 8 This is a block diagram 800 illustrating an information interaction device according to an exemplary embodiment. (Refer to...) Figure 8 The device includes a playback unit 802, a first display unit 804, and a second display unit 806, wherein:
[0115] Playback unit 802 is configured to play target video data in the display interface;
[0116] The first display unit 804 is configured to display at least one intelligent question and answer entry in the video playback page of the target video data, wherein the intelligent question and answer entry displays an intelligent question and answer icon and question guidance data.
[0117] The second display unit 806 is configured to perform a trigger operation in response to any intelligent question-and-answer entry point, and to display intelligent question-and-answer dialogue information through the information display unit. The intelligent question-and-answer dialogue information includes the question guidance data corresponding to the intelligent question-and-answer entry point and the response data to the question guidance data. The question guidance data is guidance data generated by artificial intelligence analysis.
[0118] The information interaction device provided in this embodiment plays target video data in a display interface and displays at least one intelligent question-and-answer entry point on the video playback page of the target video data. The intelligent question-and-answer entry point displays an intelligent question-and-answer identifier and question guidance data. In response to a trigger operation for any intelligent question-and-answer entry point, an intelligent question-and-answer dialogue information is displayed through an information display unit. The intelligent question-and-answer dialogue information includes the question guidance data corresponding to the intelligent question-and-answer entry point and response data to the question guidance data. The question guidance data is generated using artificial intelligence analysis. By employing the information interaction device provided in this embodiment, leveraging the multimodal deep understanding capabilities of a large model, in a user's video consumption scenario, personalized question guidance data displayed through the intelligent question-and-answer entry point guides the user to initiate an intelligent dialogue to obtain accurate information. Furthermore, through a seamless process of scene triggering, question-and-answer interaction, and deep feedback, combined with the deep analysis of multimodal information in the video by the large model, accurate matching of question guidance content with user needs and high-quality responses based on the video context are achieved. Ultimately, this significantly reduces the operational threshold and time cost for users to obtain information, and greatly improves the accuracy and overall efficiency of information transmission.
[0119] In one exemplary embodiment, the at least one intelligent question-and-answer entry point is displayed in the form of question-and-answer tags, and displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes:
[0120] The question and answer tags corresponding to the at least one smart question and answer entry are displayed on the video playback page in a preset sorting manner. The preset sorting manner includes sorting in descending order according to the matching degree between the question guidance data corresponding to the smart question and answer entry and the user account, or sorting in descending order according to the trigger popularity of the smart question and answer entry.
[0121] In one exemplary embodiment, the video playback page includes a text display area for displaying the publishing account information and title information of the target video data. Displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes:
[0122] If the user account's account tag meets the display conditions, hide the publishing account information and title information of the target video data displayed in the text display area.
[0123] The intelligent question-and-answer icon and the trigger control corresponding to each intelligent question-and-answer entry are displayed in the text display area. The trigger control displays the question guidance data corresponding to the intelligent question-and-answer entry.
[0124] In one exemplary embodiment, the video playback page includes a comment information display area, and displaying at least one intelligent Q&A entry point on the video playback page of the target video data includes:
[0125] The comment information of the target video data is displayed in the comment information display area. The comment information includes at least one target comment information, which is the comment information that triggers a question and answer event. The target comment information includes text comment data and a smart question and answer identifier.
[0126] In one exemplary embodiment, when the text comment data is of the query type, the text comment data serves as the question guidance data for the intelligent question-and-answer entry point; or, when the text comment data is not of the query type, the intelligent question-and-answer identifier also displays question guidance data, which is guidance data generated based on the text comment data.
[0127] In one exemplary embodiment, displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes:
[0128] When the target video data is paused, or when the comment data of the target video data displayed in the comment information display area includes target comment data that triggers a question-and-answer event, at least one smart question-and-answer entry point is displayed on the video playback page of the target video data.
[0129] In one exemplary embodiment, the intelligent question-and-answer dialogue information further includes at least one piece of follow-up question guidance data, which includes at least one of the following:
[0130] Question guidance data corresponding to untriggered intelligent question-answering entry points;
[0131] The question guidance data is obtained by further processing the question guidance data and / or response data corresponding to the triggered intelligent question answering entry.
[0132] In one exemplary embodiment, the response data includes text response data and reference video data, wherein the reference video data is video data associated with the question guidance data and / or response data.
[0133] Figure 9This is a block diagram illustrating an electronic device 900 for an information interaction method according to an exemplary embodiment. For example, the electronic device 900 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0134] Reference Figure 9 The electronic device 900 may include one or more of the following components: processing component 902, memory 904, power supply component 906, multimedia component 908, audio component 910, input / output (I / O) interface 912, sensor component 914, and communication component 916.
[0135] Processing component 902 typically controls the overall operation of electronic device 900, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 902 may include one or more processors 920 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 902 may include one or more modules to facilitate interaction between processing component 902 and other components. For example, processing component 902 may include a multimedia module to facilitate interaction between multimedia component 908 and processing component 902.
[0136] Memory 904 is configured to store various types of data to support the operation of electronic device 900. Examples of such data include instructions for any application or method operating on electronic device 900, contact data, phonebook data, messages, pictures, videos, etc. Memory 904 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, optical disk, or graphene storage.
[0137] Power supply component 906 provides power to various components of electronic device 900. Power supply component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 900.
[0138] Multimedia component 908 includes a screen that provides an output interface between the electronic device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 908 includes a front-facing camera and / or a rear-facing camera. When the electronic device 900 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0139] Audio component 910 is configured to output and / or input audio signals. For example, audio component 910 includes a microphone (MIC) configured to receive external audio signals when electronic device 900 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 904 or transmitted via communication component 916. In some embodiments, audio component 910 also includes a speaker for outputting audio signals.
[0140] I / O interface 912 provides an interface between processing component 902 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0141] Sensor assembly 914 includes one or more sensors for providing state assessments of various aspects of electronic device 900. For example, sensor assembly 914 can detect the on / off state of electronic device 900, the relative positioning of components such as the display and keypad of electronic device 900, changes in position of electronic device 900 or its components, the presence or absence of user contact with electronic device 900, orientation or acceleration / deceleration of device 900, and temperature changes of electronic device 900. Sensor assembly 914 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 914 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 914 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0142] Communication component 916 is configured to facilitate wired or wireless communication between electronic device 900 and other devices. Electronic device 900 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 6G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 916 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 916 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0143] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0144] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions, which can be executed by a processor 920 of an electronic device 900 to perform the above-described method. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0145] In an exemplary embodiment, a computer program product is also provided, the computer program product including instructions that can be executed by a processor 920 of an electronic device 900 to perform the above-described method.
[0146] It should be noted that the above-mentioned apparatus, electronic equipment, computer-readable storage medium, computer program product, etc., may also include other implementation methods according to the description of the method embodiments. For specific implementation methods, please refer to the description of the relevant method embodiments, which will not be elaborated here.
[0147] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0148] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An information exchange method, characterized in that, include: Play the target video data in the display interface; At least one intelligent Q&A entry is displayed on the video playback page of the target video data. The intelligent Q&A entry displays an intelligent Q&A icon and question guidance data. In response to a trigger operation for any of the intelligent question-and-answer entry points, intelligent question-and-answer dialogue information is displayed through an information display unit. The intelligent question-and-answer dialogue information includes the question guidance data corresponding to the intelligent question-and-answer entry point and the response data to the question guidance data. The question guidance data is guidance data generated using artificial intelligence analysis.
2. The method according to claim 1, characterized in that, The at least one intelligent question-and-answer entry point is displayed in the form of question-and-answer tags. Displaying at least one intelligent question-and-answer entry point on the video playback page of the target video data includes: The question and answer tags corresponding to the at least one smart question and answer entry are displayed on the video playback page in a preset sorting manner. The preset sorting manner includes sorting in descending order according to the matching degree between the question guidance data corresponding to the smart question and answer entry and the user account, or sorting in descending order according to the trigger popularity of the smart question and answer entry.
3. The method according to claim 1, characterized in that, The video playback page includes a text display area for displaying the account information of the target video data and the title information of the target video data. The video playback page of the target video data displays at least one intelligent question-and-answer entry point, including: If the user account's account tag meets the display conditions, hide the publishing account information and title information of the target video data displayed in the text display area. The intelligent question-and-answer icon and the trigger control corresponding to each intelligent question-and-answer entry are displayed in the text display area. The trigger control displays the question guidance data corresponding to the intelligent question-and-answer entry.
4. The method according to any one of claims 1 to 3, characterized in that, The video playback page includes a comment information display area, and the display of at least one intelligent Q&A entry point on the video playback page of the target video data includes: The comment information of the target video data is displayed in the comment information display area. The comment information includes at least one target comment information, which is the comment information that triggers a question and answer event. The target comment information includes text comment data and a smart question and answer identifier.
5. The method according to claim 4, characterized in that, When the text comment data is of the query type, the text comment data serves as the question guidance data for the intelligent question-and-answer entry point; or, when the text comment data is not of the query type, the intelligent question-and-answer identifier also displays question guidance data, which is guidance data generated based on the text comment data.
6. The method according to claim 1, characterized in that, The display of at least one intelligent question-and-answer entry point on the video playback page of the target video data includes: When the target video data is paused, or when the comment data of the target video data displayed in the comment information display area includes target comment data that triggers a question-and-answer event, at least one smart question-and-answer entry point is displayed on the video playback page of the target video data.
7. The method according to claim 1, characterized in that, The intelligent question-and-answer dialogue information also includes at least one piece of follow-up question guidance data, which includes at least one of the following: Question guidance data corresponding to untriggered intelligent question-answering entry points; The question guidance data is obtained by further processing the question guidance data and / or response data corresponding to the triggered intelligent question answering entry.
8. The method according to claim 1, characterized in that, The response data includes text response data and reference video data, wherein the reference video data is video data associated with the question guidance data and / or response data.
9. An information interaction device, characterized in that, include: The playback unit is configured to play the target video data in the display interface. The first display unit is configured to display at least one intelligent question-and-answer entry on the video playback page of the target video data, wherein the intelligent question-and-answer entry displays an intelligent question-and-answer identifier and question guidance data; The second display unit is configured to perform a trigger operation in response to any of the intelligent question-and-answer entry points, and to display intelligent question-and-answer dialogue information through the information display unit. The intelligent question-and-answer dialogue information includes the question guidance data corresponding to the intelligent question-and-answer entry point and the response data to the question guidance data. The question guidance data is guidance data generated based on artificial intelligence analysis.
10. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the information interaction method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the information interaction method as described in any one of claims 1 to 8.