Information search method and device, electronic equipment and storage medium

By using intelligent agents to generate suggestive keywords based on video content, search text, and account information, the problem of time-consuming judgment in video searches is solved, achieving the effect of quickly obtaining information and reducing decision-making costs.

CN122064840APending Publication Date: 2026-05-19BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511905789.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Users need to spend a considerable amount of time determining whether the video content search contains the information they need, resulting in low information acquisition efficiency and high decision-making costs.

Method used

The suggested keywords generated by the intelligent agent are displayed in the search interface based on the video content, search text, and the currently logged-in account. The suggested keywords are used to describe the video content and follow preset constraints, including quantity, word count, and arrangement, providing video content features, matching points, and account association points.

Benefits of technology

It improves information acquisition efficiency, reduces user decision-making costs, enhances search efficiency and user experience, and enables users to quickly determine the match between video content and their own needs through intuitive keyword prompts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122064840A_ABST
    Figure CN122064840A_ABST
Patent Text Reader

Abstract

The invention provides an information search method and device, electronic equipment and a storage medium, and belongs to the technical field of computers. The method comprises the steps that in response to a search operation in a search interface, a target video is displayed based on a search text corresponding to the search operation, and the target video is matched with the search text; and based on the target video, the search text and the current login account, displaying a prompt keyword extracted by the intelligent agent, the prompt keyword being used for describing the target video according to a preset constraint condition, and the preset constraint condition being used for constraining at least one of a display format of the prompt keyword or an information extraction direction of the intelligent agent. The prompt keyword can intuitively represent the personalized bright spot of the target video and better fit the preference and demand of the current user, and the prompt keyword is preposed and displayed, so that the user can quickly judge the matching degree between the video and the content required by the user without consuming time to watch the video, the information acquisition efficiency is improved, the information screening cost is reduced, and the user experience is improved. And the overall search efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an information search method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the rapid development of artificial intelligence technology, app search has become an important way for users to obtain information. The search function within an app can provide search results, such as videos, related to the user's input text. However, due to the nature of video content presentation, users often need to spend a considerable amount of time watching, or even multiple times, to determine if the video contains the information they need. This makes information retrieval from video content inefficient, increases the user's decision-making cost, and leads to low actual search efficiency. Summary of the Invention

[0003] This disclosure provides an information search method, apparatus, electronic device, and storage medium, which improves the efficiency of information retrieval from video content, reduces user decision-making costs, and improves actual search efficiency. The technical solution of this disclosure is as follows.

[0004] According to one aspect of the embodiments of this disclosure, an information search method is provided, the method comprising: In response to a search operation in the search interface, a target video is displayed based on the search text corresponding to the search operation, wherein the target video matches the search text; Based on the target video, the search text, and the currently logged-in account, prompt keywords extracted by the agent are displayed. The prompt keywords are used to describe the target video according to preset constraints. The preset constraints are used to constrain at least one of the display format of the prompt keywords or the information extraction direction of the agent.

[0005] According to another aspect of the present disclosure, an information search device is provided, the device comprising: The first display unit is configured to perform a search operation in response to the search interface, and display a target video based on the search text corresponding to the search operation, wherein the target video matches the search text; The second display unit is configured to display prompt keywords extracted by the agent based on the target video, the search text, and the currently logged-in account. The prompt keywords are used to describe the target video according to preset constraints, which are used to constrain at least one of the display format of the prompt keywords or the information extraction direction of the agent.

[0006] In some embodiments, the preset constraints correspond to a search scenario, and the search scenario is determined by the search text; The second display unit is configured to display the prompt keywords based on the search scenario to which the search text belongs, according to the corresponding preset constraints.

[0007] In some embodiments, the display format includes at least one of the following: a threshold for the number of prompt keywords, a threshold for the number of characters, and an arrangement method.

[0008] In some embodiments, the prompt keywords are used to characterize at least one of the following: the content features of the target video; the matching points between the target video and the search text; and the association points between the target video and the currently logged-in account.

[0009] In some embodiments, the apparatus further includes: The generation unit is configured to perform the following operations: based on the search text, the target video, and the currently logged-in account, obtain multi-dimensional initial information, which describes at least two of the following: search requirements, video content, and account preferences; input the multi-dimensional initial information into a large language model to obtain the suggested keywords.

[0010] In some embodiments, the multi-dimensional initial information includes at least two of the following: search keywords and search intent information obtained from the search text; video keyframes, video text, and comment extraction information obtained from the target video; and historical behavior records and interest tags obtained from the currently logged-in account.

[0011] In some embodiments, the preset constraints correspond to a search scenario, and the search scenario is determined by the search text; The generation unit is further configured to perform operations based on the search scenario to which the search text belongs, inputting the corresponding preset constraints and the multi-dimensional initial information into the large language model to obtain the suggested keywords.

[0012] In some embodiments, the second display unit is configured to switch between displaying prompt keywords for the target video and the video title of the target video.

[0013] In some embodiments, the second display unit is configured to perform a scrolling display of the prompt keywords and the video title; or, in response to a swipe operation, to switch the display of the prompt keywords and the video title.

[0014] In some embodiments, the second display unit is configured to display the prompt keywords in the video frame of the target video.

[0015] In some embodiments, the second display unit is configured to display a label for each prompt keyword in the video frame; or, to display a carousel card in the video frame, the carousel card being used to cyclically display the prompt keywords.

[0016] In some embodiments, the apparatus further includes: The third display unit is configured to perform a trigger operation in response to the prompt keyword, play the target video in the playback interface of the target video, and display the prompt keyword and summary information, wherein the summary information is used to describe the video content of the corresponding node in the target video.

[0017] In some embodiments, the summary information includes multiple summary texts; The third display unit is also configured to perform a trigger operation on any summary text, play the target video from the node corresponding to the summary text, and highlight prompt keywords that are semantically related to the summary text.

[0018] According to another aspect of the embodiments of this disclosure, an electronic device is provided, the electronic device comprising: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the aforementioned information search method.

[0019] According to another aspect of the present disclosure, a computer-readable storage medium is provided that, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the above-described information search method.

[0020] According to another aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described information search method.

[0021] This disclosure provides an information search method that, in response to a search operation, displays a target video matching the search text and provides suggested keywords for the target video. Since the suggested keywords are structured information generated by an agent based on the video content, search text, and the currently logged-in account, they can intuitively represent the personalized highlights of the target video, better aligning with the user's preferences and needs. Displaying the suggested keywords prominently in the search interface allows users to quickly determine the match between the video and their desired content without spending time watching the video, improving information acquisition efficiency and reducing information filtering costs. Simultaneously, preset constraints ensure the standardization of keyword display and the effectiveness of extraction, improving overall search efficiency and user experience.

[0022] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0024] Figure 1 This is a schematic diagram illustrating the implementation environment of an information search method according to an exemplary embodiment; Figure 2 This is a flowchart illustrating an information search method according to an exemplary embodiment; Figure 3 This is a flowchart illustrating another information search method according to an exemplary embodiment; Figure 4 This is a schematic diagram illustrating a prompt keyword display according to an exemplary embodiment; Figure 5 This is a schematic diagram illustrating another display of prompt keywords according to an exemplary embodiment; Figure 6 This is a schematic diagram illustrating an overlay display according to an exemplary embodiment; Figure 7 This is a schematic diagram illustrating a playback interface according to an exemplary embodiment; Figure 8 This is a block diagram illustrating an information search device according to an exemplary embodiment; Figure 9 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0025] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0026] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0027] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the search text, account information, and historical behavior records involved in this disclosure were obtained with full authorization.

[0028] Figure 1 This is a schematic diagram illustrating an implementation environment for an information search method according to an exemplary embodiment. See also... Figure 1 The implementation environment specifically includes: terminal 101 and server 102. Terminal 101 can be connected to server 102 via wireless network or wired network.

[0029] Terminal 101 can be at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), and laptop computer.

[0030] Applications can be installed and run on terminal 101. These applications can respond to search operations and display search results. Such applications include, but are not limited to, short video applications, audio / video applications, social media applications, shopping applications, live streaming applications, etc. Users can log in to the application through terminal 101 to access its services. The application is associated with server 102, which provides background services to terminal 101. For example, during application operation, terminal 101 displays corresponding multimedia resources based on the data stream pushed by the application's background server 102 (such as the search result data stream).

[0031] Terminal 101 can refer to one of a plurality of terminals, and this embodiment uses terminal 101 as an example. Those skilled in the art will know that the number of terminals can be more or less. For example, there can be several terminals, or dozens or hundreds of terminals, or more. This embodiment does not limit the number of terminals or the type of devices.

[0032] Server 102 can be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Optionally, the number of servers may be more or fewer, and this disclosure does not limit this. Of course, server 102 may also include other functional servers to provide more comprehensive and diversified services. In some embodiments, server 102 undertakes the main computing work, and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work, and terminal 101 undertakes the main computing work; or, server 102 and terminal 101 collaborate on computing using a distributed computing architecture. Server 102 can be connected to terminal 101 and other terminals via a wireless network or a wired network. Optionally, the number of servers may be more or fewer, and this disclosure does not limit this.

[0033] It should be noted that the above implementation environment is only an example, and the method provided in this embodiment can also be executed in other implementation environments. This embodiment does not limit this.

[0034] Figure 2 This is a flowchart illustrating an information search method according to an exemplary embodiment, such as... Figure 2 As shown, this method is executed by the terminal and includes the following steps.

[0035] In step 201, the terminal responds to the search operation in the search interface, and displays the target video based on the search text corresponding to the search operation, and the target video matches the search text.

[0036] In this embodiment, the search interface is a visual interface for users to interact with the system through search. It receives the user's search text and can be found in various applications such as e-commerce platforms, social media platforms, online music platforms, and short video platforms. This embodiment does not limit the scope of the search interface. Users can submit search text through text input, voice-to-text input, etc. A search operation refers to the user's action in the search interface to obtain the required information, such as entering search text and clicking the search button, or entering voice input and then confirming the search. Search text is the text content entered by the user through the search interface to express their needs. Search text can be keywords, phrases, complete sentences, etc., for example, search text such as "cold symptoms," "push-up tutorial," "how to cook beef deliciously," etc. The target video refers to a video that the system selects from the video resource library based on the content characteristics of the search text and that is associated with the search text. The matching logic can be based on video title keywords, preset tags, screen content themes, or subtitle information. If the target video matches the search text in terms of theme, content, or keywords, this embodiment does not limit how video search is implemented. This example uses the target video being any video in the search results. The search results may also include other multimedia resources, which will not be discussed here.

[0037] In step 202, the terminal displays prompt keywords extracted by the agent based on the target video, search text, and the currently logged-in account. The prompt keywords are used to describe the target video according to preset constraints. The preset constraints are used to constrain at least one of the following: the display format of the prompt keywords or the information extraction direction of the agent.

[0038] In this embodiment, the currently logged-in account refers to the identity identifier used by the user to currently log in to the application. The currently logged-in account is associated with the user's personalized data, such as search history, preference settings, and permission information. The intelligent agent is used to automatically extract key information from the target video, search text, and associated data of the currently logged-in account to generate prompt keywords. Exemplarily, the intelligent agent is implemented using a large language model. Besides the large language model, the intelligent agent may also include other sub-modules, such as computer vision models, etc., without limitation. The prompt keywords can summarize the characteristics of the target video. Since the intelligent agent not only relies on the target video to generate prompt keywords but also combines the associated information of the search text and the currently logged-in account, these features include not only the video content features of the target video but also features used to characterize the correlation between the target video, the search text, and the currently logged-in account information.

[0039] Preset constraints refer to the set of rules used to limit the generation and display of prompt keywords, ensuring their effectiveness and standardization. Among these, the display format limits how prompt keywords are presented on the interface, including at least one of the following: a threshold for the number of prompt keywords, a threshold for the number of characters, and an arrangement method. For example, it might limit the display to a maximum of four prompt keywords, a maximum of five characters for each keyword, or an arrangement in the form of horizontal labels. The information extraction direction refers to the focus dimension of the agent when extracting keywords. For example, for fitness tutorial videos, the information extraction direction indicates the target audience and the advantages and disadvantages of the movements; for food review videos, the information extraction direction indicates the location of the restaurant and the taste of the food.

[0040] This disclosure provides an information search method that, in response to a search operation, displays a target video matching the search text and provides suggested keywords for the target video. Since the suggested keywords are structured information generated by an agent based on the video content, search text, and the currently logged-in account, they can intuitively represent the personalized highlights of the target video, better aligning with the user's preferences and needs. Displaying the suggested keywords prominently in the search interface allows users to quickly determine the match between the video and their desired content without spending time watching the video, improving information acquisition efficiency and reducing information filtering costs. Simultaneously, preset constraints ensure the standardization of keyword display and the effectiveness of extraction, improving overall search efficiency and user experience.

[0041] In some embodiments, preset constraints correspond to search scenarios, and the search scenarios are determined by the search text. Display the prompt keywords extracted by the agent, including: Based on the search context to which the search text belongs, suggestive keywords are displayed according to the corresponding preset constraints.

[0042] In this embodiment of the disclosure, since the user's focus differs significantly in different search scenarios, the scenario is determined by the search text and matched with the corresponding preset constraints, so that the keyword extraction direction and display format are highly consistent with the current search scenario. This allows users in different scenarios to obtain targeted key information, further reducing the cost of information acquisition and decision-making for users and improving information acquisition efficiency.

[0043] In some embodiments, the display format includes at least one of the following: a threshold for the number of prompt keywords, a threshold for the number of characters, and an arrangement method.

[0044] In this embodiment of the disclosure, by limiting the number, word count, and arrangement, keyword redundancy or information incompleteness can be avoided. For example, controlling the number prevents information overload, limiting the word count ensures that core information is highlighted, and a reasonable arrangement facilitates quick browsing. These constraints on the display format make the display of prompt keywords more regular and readable, making prompt keywords easier to read and filter. Users can focus on key information in a short time, improving information acquisition efficiency and user experience.

[0045] In some embodiments, prompt keywords are used to characterize at least one of the following: Content characteristics of the target video; Matching points between the target video and the search text; The connection between the target video and the currently logged-in account.

[0046] In this disclosed embodiment, the core content of the video is conveyed to the user, the convergence point between the video and the search needs is clearly defined, and personalized highlights are emphasized based on account preferences. This not only satisfies the user's cognitive needs regarding the video itself but also takes into account search targeting and personalized adaptation. These keyword information dimensions are comprehensive, providing users with decision-making basis from multiple perspectives, helping them to more comprehensively and accurately determine whether the video meets their expectations, and improving search decision-making efficiency.

[0047] In some embodiments, the method for generating prompt keywords includes: Based on the search text, target video, and currently logged-in account, obtain multi-dimensional initial information. The multi-dimensional initial information is used to describe at least two of the search requirements, video content, and account preferences. Inputting multi-dimensional initial information into a large language model yields suggested keywords; In this embodiment of the disclosure, core data from the search end, video end, and user end are integrated through multi-dimensional initial information, providing more comprehensive data support for keyword generation. The keywords are generated by inputting into a large language model, avoiding the one-sidedness of keywords caused by a single information source. This makes the generated keywords more in line with the user's real needs and the essence of the video, providing users with more valuable information guidance and reliable support for users to make quick judgments.

[0048] In some embodiments, the multidimensional initial information includes at least two of the following: Search keywords and search intent information obtained from the search text; Video keyframes, video text, and comment extraction information obtained from the target video; Historical behavior records and interest tags obtained from the currently logged-in account.

[0049] In this embodiment of the disclosure, the multi-dimensional initial information covers the elements of the entire search chain, providing high-quality input for the large language model, ensuring the reliability and relevance of subsequent AI keyword extraction, improving the effectiveness of suggested keywords from the source, and thus further improving the efficiency of user decision-making.

[0050] In some embodiments, preset constraints correspond to search scenarios, and the search scenarios are determined by the search text. Inputting multi-dimensional initial information into a large language model yields suggested keywords, including: Based on the search context to which the search text belongs, the corresponding preset constraints and multi-dimensional initial information are input into the large language model to obtain suggested keywords.

[0051] In this embodiment of the disclosure, the preset constraints corresponding to the scenario and multi-dimensional initial information are synchronously input into the large language model. This allows the model to extract keywords while adhering to the format and direction requirements of scenario adaptation and based on comprehensive information. The generated keywords are more in line with the usage needs of the current search scenario, improving the scenario relevance and information effectiveness of the suggested keywords, thereby further enhancing the auxiliary value for user decision-making.

[0052] In some embodiments, displaying prompt keywords extracted by the agent includes: Switch between displaying prompt keywords for the target video and the video title of the target video.

[0053] In this embodiment, the display of keywords and video titles can be switched. This not only retains the convenience of traditional title display, but also allows users to quickly obtain core information through keywords. Users can choose to view the title or keywords according to their own needs and obtain information flexibly. This satisfies the information acquisition preferences of different users, avoids the limitations of a single display method, improves the flexibility of information display, and enhances the user experience.

[0054] In some embodiments, switching the display of prompt keywords for the target video and the video title of the target video includes: A scrolling prompt display of keywords and video titles; or, In response to a swipe gesture, the tooltip keywords and video title are displayed in a toggle.

[0055] In this embodiment, two switching modes are provided: scrolling display and swipe trigger. Scrolling display can automatically present two types of information without requiring active user operation, while swipe trigger meets the user's self-control needs and allows switching the content to be viewed at any time as needed. This not only adapts to the operating habits and usage scenarios of different users, but also enables users to more conveniently obtain information from video titles or prompt keywords, improving operational convenience and thus improving human-computer interaction efficiency.

[0056] In some embodiments, displaying prompt keywords extracted by the agent includes: The target video will display keyword prompts.

[0057] In this embodiment of the disclosure, the prompt keywords are superimposed on the video screen, making the keywords more intuitive and enabling users to grasp the keyword information while watching the video. This further reduces the visual cost of information acquisition and improves the efficiency of information transmission.

[0058] In some embodiments, displaying prompt keywords in the video frame of the target video includes: Display a tag for each prompt keyword in the video frame; or, A carousel card is displayed in the video frame, which is used to display prompt keywords in a loop.

[0059] In this embodiment of the disclosure, the label display makes the keywords more intuitive, while the carousel cards display multiple keywords without taking up too much screen space. Both methods take into account the needs of information transmission and screen preview. That is, they not only ensure the effective transmission of keyword information, but also avoid the screen preview being affected by keyword obstruction. This improves the visual richness of the interface and enhances the user experience.

[0060] In some embodiments, the method further includes: In response to the triggering operation of the prompt keyword, the target video is played in the target video playback interface, and the prompt keyword and summary information are displayed. The summary information is used to describe the video content of the corresponding node in the target video.

[0061] In this embodiment of the disclosure, triggering keywords indicates that the user is interested in information in that direction. At this time, playing the video and simultaneously displaying the keywords and summary can help the user quickly browse the video, improve the relevance and efficiency of video viewing, and further improve the efficiency of information transmission.

[0062] In some embodiments, the summary information includes multiple summary texts; the method further includes: In response to a trigger action on any summary text, the target video is played from the corresponding node of the summary text, and prompt keywords that are semantically related to the summary text are highlighted.

[0063] In this embodiment of the disclosure, the summary text helps users quickly locate and play the target segment, while the highlighting of related keywords helps users focus on the core information of the segment. This improves the accuracy of video viewing and the efficiency of information acquisition, and also improves the efficiency of finding specific segments of the video. It helps users accurately obtain the specific content information they need, and is especially suitable for filtering the core content of long videos, thus improving the user experience.

[0064] The above Figure 2The diagram shown is a flowchart of an information search method according to this disclosure. The information search scheme provided by this disclosure will be further elaborated below. Figure 3 This is a flowchart illustrating another information search method according to an exemplary embodiment, see [link to flowchart]. Figure 3 This method is executed by the terminal and includes the following steps.

[0065] In step 301, the terminal responds to the search operation in the search interface and displays the target video based on the search text corresponding to the search operation, and the target video matches the search text.

[0066] In this embodiment of the disclosure, the principle of the terminal displaying the target video in step 301 is the same as that in step 201, and will not be repeated here.

[0067] In step 302, the terminal obtains the prompt keywords extracted by the agent based on the target video, search text, and the currently logged-in account. The prompt keywords are used to describe the target video according to preset constraints. The preset constraints are used to constrain at least one of the following: the display format of the prompt keywords or the information extraction direction of the agent.

[0068] In this embodiment, the prompt keywords can be generated by the terminal or a backend server associated with the terminal, and there is no limitation on this. The number of prompt keywords can be one or more, organized into structured key points. The information extraction direction limited by preset constraints can be various, such as extracting prominent features, key steps, precautions, advantages and disadvantages, etc. For example, if the target video is a travelogue, the prompt keywords may include "city tour, relaxed pace, real-life experience, emotional healing"; if the target video is a game commentary, the prompt keywords may include "highlight operations, new version explanation, player insights, fast-paced"; if the target video is workplace skills, the prompt keywords may include "must-see for office workers, practical case studies, improving communication skills, condensed practical information". It should be noted that the prompt keywords here are only illustrative examples, and the prompt keywords can also be replaced with short sentences with subject-verb-object structures, as long as the word count meets the display format constraints.

[0069] Using an intelligent agent, personalized core information is automatically extracted from the target video based on the target video, search text, and associated data of the currently logged-in account, and structured prompt keywords are generated. This is essentially structured highlight generation and extraction based on the intelligent agent. The intelligent agent is constrained by preset constraints during the generation process. For example, the preset constraints and the aforementioned data are input into the intelligent agent to generate prompt keywords that meet the constraints; alternatively, the intelligent agent is trained using a training set that satisfies the preset constraints, enabling the model to generate prompt keywords that meet the constraints during the model inference phase. It should be noted that the above processing methods for the intelligent agent are illustrative examples. Post-processing filtering, embedding constraint layers within the model, and other methods can also be used to generate prompt keywords that meet the constraints, which will not be limited or elaborated upon here. The following section further describes the process of generating prompt keywords.

[0070] In some embodiments, the data is first pre-processed and then synthesized to generate suggested keywords. Accordingly, multi-dimensional initial information is obtained based on the search text, target video, and currently logged-in account; this multi-dimensional initial information is input into a large language model to obtain suggested keywords. The multi-dimensional initial information describes at least two of the following: search intent, video content, and account preferences. By integrating core data from the search engine, video platform, and user end using multi-dimensional initial information, more comprehensive data support is provided for keyword generation. Inputting this information into the large language model avoids the bias caused by a single information source, making the generated keywords more closely aligned with the user's actual needs and the essence of the video, providing users with more valuable information guidance and reliable support for quick decision-making.

[0071] For example, the multi-dimensional initial information includes at least two of the following: search keywords and search intent information obtained from the search text; video keyframes, video text, and comment extraction information obtained from the target video; and historical behavior records and interest tags obtained from the currently logged-in account.

[0072] Regarding the information obtained from the search text: search keywords refer to the core words extracted from the search text. For example, if the search text is "beginner baking cake tutorial", the keywords could be "beginner", "baking", "cake" and "tutorial". Search intent information refers to the judgment result of the search purpose, which can be obtained by the search intent judgment model. For example, it can be divided into information query, skill learning, and product purchase.

[0073] The information obtained from the target video is divided into information from the video content itself and information extracted from video comments. Information from the video content itself includes video text and video keyframes. Video text includes at least one of audio subtitles and visual text. Audio subtitles, such as narration and dialogue, are obtained from the target video through speech recognition; visual text, such as recipe steps, is identified from the target video through keyframe analysis. Video keyframes include important frames, which are identified from the target video through keyframe analysis algorithms. These keyframe analysis algorithms can be based on inter-frame differences, etc., and are not limited thereto. The above-mentioned initial information from the video content itself is only illustrative and may also include audio extracted from the target video. Information extracted from comments is core content extracted from the user comment section of the target video, including frequently discussed topics and sentiments, such as frequently discussed topics like "egg whites to pass the time" and sentiments like "the steps are clear and easy to understand." This information can indicate "points of contention" or "highlights" that users are generally concerned about.

[0074] Regarding the information obtained from the currently logged-in account: Historical behavior records include historical search records, browsing records, favorite records, and purchase records, all of which have been authorized by the user. Interest tags are tags automatically generated based on account behavior or actively set by the user, which can identify the user's interest areas, such as "experienced baking enthusiast" or "fitness beginner." Through historical behavior records, the focus of keyword generation can be adjusted personalized during the intelligent agent's processing. For example, if the currently logged-in account has the interest tag "fitness beginner," the generated suggestion keywords will include "beginner-friendly." This information comes from the associated information of the currently logged-in account, which specifically refers to the raw data related to user-authorized behaviors and preferences bound to the account.

[0075] It should be noted that the above multi-dimensional information is for illustrative purposes only and does not constitute a limitation. The multi-dimensional initial information covers all elements of the search process, providing high-quality input for the large language model, ensuring the reliability and relevance of subsequent AI keyword extraction, and improving the effectiveness of suggested keywords from the source, thereby further enhancing user decision-making efficiency.

[0076] In some embodiments, the suggested keywords are used to characterize at least one of the following: the content characteristics of the target video, intuitively presenting the core information of the video, such as the keywords for a game commentary video including "tight-paced"; the matching point between the target video and the search text, clarifying the fit between the video and the user's search needs, such as the keywords for a game commentary video including "new version explanation" if the user searches for a new version gameplay introduction; and the association point between the target video and the currently logged-in account, characterizing the adaptability of the video to the user's personalized preferences, such as the keywords for a game commentary video including "must-see for beginners" if the user is a novice player.

[0077] This connects to the aforementioned multi-dimensional initial information. Because this information encompasses various aspects of search, video, and user-related data, the suggested keywords can not only represent the video content but also further characterize the relationship between the video and the user's search needs and preferences. This not only conveys the core content of the video to the user but also clarifies the point of convergence between the video and the search needs, highlighting personalized features based on account preferences. It not only satisfies the user's cognitive needs regarding the video itself but also ensures targeted and personalized search results. Compared to single-dimensional information prompts, these multi-dimensional keywords provide users with decision-making support from multiple perspectives, helping them to more comprehensively and accurately determine whether a video meets their expectations, thus improving search decision-making efficiency.

[0078] In some embodiments, different preset constraints are applied for different search scenarios. Accordingly, the preset constraints correspond to the search scenario, which is determined by the search text. For example, the search scenario to which the search text belongs is determined by the search intent indicated by the search text. Based on the search scenario to which the search text belongs, suggested keywords are displayed according to the corresponding preset constraints.

[0079] For example, if the search scenario is a knowledge learning scenario, such as searching for courses or tutorials (e.g., "Python beginner" or "Photoshop tips"), then preset constraints instruct the agent (e.g., a large language model) to extract a knowledge point outline, and the suggested keywords should at least include the knowledge point outline. If the search scenario is an e-commerce live streaming and product search scenario, such as searching for product reviews or usage tutorials, then preset constraints instruct the agent to extract structured information indicating the advantages and disadvantages of the product and its target audience, and the suggested keywords should at least include structured information such as "advantages, disadvantages, and target audience" to assist in purchasing decisions. If the search scenario is a news and information search scenario, then preset constraints instruct the agent to extract the core elements of the event, extract key points from news videos, and the suggested keywords should at least include "time, location, people, and results" to enable users to quickly understand the core elements of the event. The above examples illustrate different information extraction directions for the agent for different search scenarios. In addition, different suggested keyword display formats can also be set, which will not be elaborated here.

[0080] Since users' focus varies significantly across different search scenarios, the scenario is determined by the search text and matched with corresponding preset constraints. This ensures that the keyword extraction direction and display format are highly consistent with the current search scenario, allowing users in different scenarios to obtain targeted key information. This further reduces the cost of information acquisition and decision-making for users and improves information acquisition efficiency.

[0081] In some embodiments, keyword generation is guided by preset constraints and multi-dimensional initial information. Accordingly, based on the search scenario to which the search text belongs, the corresponding preset constraints and multi-dimensional initial information are input into the large language model to obtain suggested keywords. This allows the model to extract keywords while adhering to the format and direction requirements of scenario adaptation, and based on comprehensive information, the generated keywords are more in line with the usage needs of the current search scenario, improving the scenario relevance and information effectiveness of the suggested keywords, thereby further enhancing their auxiliary value for user decision-making.

[0082] It should be noted that the above content uses multi-dimensional initial information and preset constraints as an example to illustrate keyword generation under different search scenarios, but it does not constitute a limitation. In some other embodiments, a method of binding search scenarios with corresponding constraints can be adopted: a corresponding model is configured for different search scenarios, and each model is specifically trained using a training set that meets the preset constraints of that scenario (the data in the training set contains scene-adapted display formats, extraction directions, and other constraint attributes); when the search scenario to which the search text belongs is detected, the model bound to that scenario is called to generate prompt keywords that meet the corresponding constraints. In yet another embodiment, the agent (such as a large language model) determines the search scenario and applies the corresponding constraints: after the agent receives multiple data (including search text, target video, and associated information of the currently logged-in account), or after the large language model receives multi-dimensional initial information, it autonomously identifies the search scenario corresponding to the search text through its own semantic analysis capabilities, and then automatically generates prompt keywords that fit the scenario based on the scene and constraint mapping rules built into the model. The scene and constraint mapping rules indicate the correspondence between different search scenarios and preset constraints. This explanation uses the example of an agent receiving multiple data points as input. In other words, a large language model is part of the agent's implementation, and the agent can also include a model for acquiring multi-dimensional initial information.

[0083] In some embodiments, the display format includes at least one of a threshold for the number of suggested keywords, a threshold for the number of characters, and an arrangement method. This avoids keyword redundancy or incomplete information. For example, controlling the number of keywords prevents information overload, limiting the number of characters ensures that core information is highlighted, and a reasonable arrangement facilitates quick browsing. These constraints on the display format make the display of suggested keywords more regular and readable, making them easier to read and filter. Users can focus on key information in a short time, improving information acquisition efficiency and user experience. The above display format is only an illustrative example and does not constitute a limitation.

[0084] In step 303, the terminal displays a prompt keyword.

[0085] In this embodiment, after the terminal obtains the suggested keywords, the suggested keywords are displayed visually in the search results page (or still the search interface described above), allowing users to quickly grasp the core value of the video before clicking to play it. It should be noted that the search results can be displayed in a single column or two columns; correspondingly, the suggested keywords can be displayed not only in single-column videos but also in two-column videos. Additionally, an AI identifier can be displayed along with the suggested keywords to indicate to the user that the keyword was generated by AI.

[0086] For example, see the single-column video. Figure 4 As shown, Figure 4 This is a schematic diagram illustrating the display of prompt keywords according to an exemplary embodiment. See also: Figure 4 As shown in Figure (a), the user-input search text is "deadlift tutorial," and the generated and displayed keyword suggestions are "beginner-friendly, power point explanation, stance explanation." These keywords are organized into multiple word forms. See also... Figure 4 As shown in Figure (b), the user's search text is "how to make braised noodles with green beans," and the generated and displayed suggestion keywords are "this way of making braised noodles with green beans, non-stick, and with chewy and flavorful noodles." Here, the suggestion keywords are organized into short sentences. In this method, the suggestion keywords and the original video title are displayed together. For an example of a two-column video, see [link to example]. Figure 5 As shown, Figure 5 This is an illustration of another suggestion keyword display according to an exemplary embodiment. The user inputs the search text "how to cook beef," and the generated and displayed suggestion keywords are "beginner-friendly, ready in half an hour, spicy, perfect with rice." In this method, the original video title is replaced by the suggestion keywords. It should be noted that the above illustration uses the suggestion keywords located in a preset area as an example; this display position is merely an example and can be displayed in other locations without limitation.

[0087] Regarding the display of prompt keywords, two more exemplary methods are given below, see (1) and (2), but these are not intended to be limiting.

[0088] (1) The terminal switches to display the target video's prompt keywords and the target video's title.

[0089] The video title is the original title associated with the target video, which differs from the newly generated prompt keywords. It supports switching between keyword and video title display, retaining the convenience of traditional title display while allowing users to quickly obtain core information through keywords. Users can choose to view either the title or keywords according to their needs, flexibly acquiring information. This satisfies different users' information acquisition preferences, avoids the limitations of a single display method, improves the flexibility of information display, and enhances the user experience.

[0090] In some embodiments, the switching method includes: scrolling to display prompt keywords and video titles, such as scrolling vertically or horizontally; or, responding to a swipe operation, switching between displaying prompt keywords and video titles, such as displaying both pieces of information in a swipeable component, supporting switching between left / right swipes or up / down swipes. Scrolling displays automatically present both types of information without requiring active user intervention, while swipe-triggered displays satisfy the user's need for self-control, allowing them to switch between viewed content at any time as needed. This not only adapts to different users' operating habits and usage scenarios but also enables users to more conveniently obtain information from video titles or prompt keywords, improving operational convenience and thus enhancing human-computer interaction efficiency.

[0091] (2) Display prompt keywords in the video frame of the target video on the terminal.

[0092] The video can be in autoplay mode or displaying a cover image; this embodiment does not impose any limitations on this. Overlaying the prompt keywords onto the video screen makes the keywords more intuitive, allowing users to grasp the keyword information while watching the video, further reducing the visual cost of information acquisition and improving information transmission efficiency.

[0093] In some embodiments, the overlay method includes: displaying a label for each prompt keyword in the video frame; or, displaying a carousel card in the video frame, where the carousel card is used to cycle through the prompt keywords, and a carousel switching card is used to switch between displaying different prompt keywords vertically or horizontally. Label display makes the keywords more intuitive, while the carousel card displays multiple keywords without occupying too much screen space. Both methods balance the needs of information delivery and screen preview, ensuring not only the effective delivery of keyword information but also avoiding the obstruction of the screen preview by keywords. This improves the visual richness of the interface and enhances the user experience.

[0094] For ease of description, see Figure 6 As shown, Figure 6 This is a schematic diagram illustrating an overlay display according to an exemplary embodiment. See also... Figure 6 As shown in Figure (a), multiple keyword suggestion labels are displayed in area 601. See Figure (a). Figure 6 As shown in Figure (b), a carousel card is displayed in area 602. It should be noted that both single-column and double-column videos can be displayed using tags and carousel cards; the illustration here is merely illustrative.

[0095] The two display methods described above are merely illustrative examples. The prompt keyword can be displayed in any preset position on the interface, and there are no restrictions on this.

[0096] In step 304, in response to the triggering operation of the prompt keyword, the terminal plays the target video in the target video playback interface and displays the prompt keyword.

[0097] In this embodiment, summary information can also be displayed on the playback interface of the target video (i.e., the video details page). This summary information describes the video content of the corresponding node in the target video. For example, prompt keywords and summary information are displayed in a pop-up video description panel, which showcases various information extracted by AI. Since triggering keywords indicates user interest in information in that direction, playing the video and simultaneously displaying the keywords and summary helps users quickly browse the video, improving the relevance and efficiency of video viewing, and further enhancing information delivery efficiency.

[0098] For a clearer description of the playback interface, please refer to [link / reference]. Figure 7 As shown, Figure 7 This is a schematic diagram illustrating a playback interface according to an exemplary embodiment. Region 701 displays prompt keywords, and region 702 displays summary information.

[0099] In some embodiments, the summary information includes multiple summary texts. Accordingly, in response to a trigger operation on any summary text, the terminal plays the target video from the corresponding node of the summary text and highlights semantically related keywords. The highlighting can employ various methods such as highlighting, flashing, or changing the background color, without limitation. Semantically related keywords can be words appearing in the summary text or words semantically equivalent and related to the summary text. For example, if the summary text is "Beginners should pay attention to their posture," then semantically related keywords could include "beginner"; if the summary text is "Put chili peppers in the pot," then semantically related keywords could include "spicy." The summary text helps users quickly locate and play the target segment, while the highlighting of related keywords helps users focus on the core information of the segment. This improves the accuracy of video viewing and the efficiency of information retrieval, and also improves the efficiency of finding specific video segments, helping users accurately obtain the specific content information they need. This is especially suitable for filtering the core content of long videos, thus improving the user experience.

[0100] This disclosure provides an information search method that, in response to a search operation, displays a target video matching the search text and provides suggested keywords for the target video. Since the suggested keywords are structured information generated by an agent based on the video content, search text, and the currently logged-in account, they can intuitively represent the personalized highlights of the target video, better aligning with the user's preferences and needs. Displaying the suggested keywords prominently in the search interface allows users to quickly determine the match between the video and their desired content without spending time watching the video, improving information acquisition efficiency and reducing information filtering costs. Simultaneously, preset constraints ensure the standardization of keyword display and the effectiveness of extraction, improving overall search efficiency and user experience.

[0101] More specifically, this approach allows users to assess whether content meets their needs without watching the entire video, directly jumping to segments of interest and saving time. Structured highlights (i.e., suggestive keywords) present core information in a concise and clear format, helping users quickly filter the most relevant videos, leading to more accurate decision-making. In skill-based and recipe-based scenarios, users can quickly obtain specific steps or precautions, reducing trial-and-error costs and improving learning and practice efficiency. This method, by combining user-personalized information, search term characteristics, and multiple result characteristics, uses AI capabilities to extract highlight information and present it in a structured manner, significantly reducing users' decision-making costs and browsing time, improving information acquisition efficiency, and enhancing users' decision-making capabilities.

[0102] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.

[0103] Figure 8 This is a block diagram illustrating an information search device according to an exemplary embodiment. Figure 8 As shown, the device includes: a first display unit 801 and a second display unit 802.

[0104] The first display unit 801 is configured to perform a search operation in response to the search interface, display a target video based on the search text corresponding to the search operation, and the target video matches the search text. The second display unit 802 is configured to display prompt keywords extracted by the agent based on the target video, search text, and the currently logged-in account. The prompt keywords are used to describe the target video according to preset constraints, which are used to constrain at least one of the display format of the prompt keywords or the information extraction direction of the agent.

[0105] In some embodiments, preset constraints correspond to search scenarios, and the search scenarios are determined by the search text. The second display unit 802 is configured to display suggested keywords based on the search scenario to which the search text belongs and according to the corresponding preset constraints.

[0106] In some embodiments, the display format includes at least one of the following: a threshold for the number of prompt keywords, a threshold for the number of characters, and an arrangement method.

[0107] In some embodiments, the prompt keywords are used to characterize at least one of the following: the content features of the target video; the matching points between the target video and the search text; and the association points between the target video and the currently logged-in account.

[0108] In some embodiments, the apparatus further includes: The generation unit is configured to perform actions based on the search text, target video, and currently logged-in account to obtain multi-dimensional initial information. The multi-dimensional initial information describes at least two of the search requirements, video content, and account preferences. The multi-dimensional initial information is then input into a large language model to obtain suggested keywords. In some embodiments, the multi-dimensional initial information includes at least two of the following: search keywords and search intent information obtained from the search text; video keyframes, video text, and comment extraction information obtained from the target video; and historical behavior records and interest tags obtained from the currently logged-in account.

[0109] In some embodiments, preset constraints correspond to search scenarios, and the search scenarios are determined by the search text. The generation unit is also configured to perform tasks based on the search scenario to which the search text belongs, inputting the corresponding preset constraints and multi-dimensional initial information into the large language model to obtain suggested keywords.

[0110] In some embodiments, the second display unit 802 is configured to switch between displaying prompt keywords for the target video and the video title of the target video.

[0111] In some embodiments, the second display unit 802 is configured to perform a scrolling display of prompt keywords and video titles; or, in response to a swipe operation, to switch between displaying prompt keywords and video titles.

[0112] In some embodiments, the second display unit 802 is configured to display prompt keywords in the video frame of the target video.

[0113] In some embodiments, the second display unit 802 is configured to perform the action of displaying a label for each prompt keyword in the video frame; or, to display a carousel card in the video frame, the carousel card being used to cyclically display the prompt keywords.

[0114] In some embodiments, the apparatus further includes: The third display unit is configured to perform a response to the triggering operation of the prompt keywords, play the target video in the target video playback interface, and display the prompt keywords and summary information. The summary information is used to describe the video content of the corresponding node in the target video.

[0115] In some embodiments, the summary information includes multiple summary texts; The third display unit is also configured to perform a trigger operation on any summary text, play the target video from the corresponding node of the summary text, and highlight prompt keywords that are semantically related to the summary text.

[0116] This disclosure provides an information search device that, in response to a search operation, displays a target video matching the search text and provides suggested keywords for the target video. Since the suggested keywords are structured information generated by an agent based on the video content, search text, and the currently logged-in account, they can intuitively represent the personalized highlights of the target video, better aligning with the user's preferences and needs. Displaying the suggested keywords prominently in the search interface allows users to quickly determine the match between the video and their desired content without spending time watching the video, improving information acquisition efficiency and reducing information filtering costs. Simultaneously, preset constraints ensure the standardization of keyword display and the effectiveness of extraction, improving overall search efficiency and user experience.

[0117] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure.

[0118] It should be noted that the information search device provided in the above embodiments is only an example of the division of the above functional units. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the information search device and the information search method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0119] Regarding the information search device in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0120] Figure 9 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Typically, the electronic device 900 includes a processor 901 and a memory 902.

[0121] Processor 901 may include one or more processing cores, such as a quad-core processor, a nine-core processor, etc. Processor 901 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 901 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 901 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 901 may also include an AI processor for handling computational operations related to machine learning.

[0122] The memory 902 may include one or more computer-readable storage media, which may be non-transitory. The memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 902 are used to store at least one program code, which is executed by the processor 901 to implement the information search method provided in the method embodiments of this disclosure.

[0123] In some embodiments, the electronic device 900 may optionally include a peripheral device interface 903 and at least one peripheral device. The processor 901, memory 902, and peripheral device interface 903 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 903 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907, and a power supply 908.

[0124] Peripheral device interface 903 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 901 and memory 902. In some embodiments, processor 901, memory 902 and peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 901, memory 902 and peripheral device interface 903 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0125] The radio frequency (RF) circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 904 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 904 can communicate with other electronic devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 904 may also include circuitry related to NFC (Near Field Communication), which is not limited in this disclosure.

[0126] Display screen 905 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 905 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 901 for processing. In this case, display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 905, which serves as the front panel of electronic device 900; in other embodiments, there may be at least two display screens 905, respectively disposed on different surfaces of electronic device 900 or in a folded design; in still other embodiments, display screen 905 may be a flexible display screen, disposed on a curved or folded surface of electronic device 900. Furthermore, display screen 905 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 905 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0127] The camera assembly 906 is used to acquire images or videos. Optionally, the camera assembly 906 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the electronic device, and the rear-facing camera is located on the back of the electronic device. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 906 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.

[0128] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 901 for processing, or input to the radio frequency circuit 904 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located in a different part of the electronic device 900. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert the electrical signals from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 907 may also include a headphone jack.

[0129] Power supply 908 is used to supply power to the various components in electronic device 900. Power supply 908 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 908 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0130] Those skilled in the art will understand that Figure 9 The structure shown does not constitute a limitation on the electronic device 900, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0131] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 902 including instructions, which can be executed by a processor 901 of an electronic device 900 to complete the aforementioned information search method. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.

[0132] A computer program product includes a computer program that, when executed by a processor, implements the aforementioned information search method.

[0133] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0134] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. An information search method, characterized in that, The method includes: In response to a search operation in the search interface, a target video is displayed based on the search text corresponding to the search operation, wherein the target video matches the search text; Based on the target video, the search text, and the currently logged-in account, prompt keywords extracted by the agent are displayed. The prompt keywords are used to describe the target video according to preset constraints. The preset constraints are used to constrain at least one of the display format of the prompt keywords or the information extraction direction of the agent.

2. The information search method according to claim 1, characterized in that, The preset constraints correspond to the search scenario, and the search scenario is determined by the search text; The displayed prompt keywords extracted by the intelligent agent include: Based on the search scenario to which the search text belongs, the suggested keywords are displayed according to the corresponding preset constraints.

3. The information search method according to claim 1, characterized in that, The display format includes at least one of the following: the number threshold of the prompt keywords, the number of characters, and the arrangement method.

4. The information search method according to claim 1, characterized in that, The prompt keywords are used to represent at least one of the following: The content characteristics of the target video; The matching points between the target video and the search text; The connection point between the target video and the currently logged-in account.

5. The information search method according to claim 1, characterized in that, The method for generating the prompt keywords includes: Based on the search text, the target video, and the currently logged-in account, multi-dimensional initial information is obtained, which is used to describe at least two of the search requirements, video content, and account preferences. The multi-dimensional initial information is input into the large language model to obtain the prompt keywords.

6. The information search method according to claim 5, characterized in that, The multi-dimensional initial information includes at least two of the following: The search keywords and search intent information obtained from the search text; Video keyframes, video text, and comment extraction information obtained from the target video; Historical behavior records and interest tags obtained from the currently logged-in account.

7. The information search method according to claim 5, characterized in that, The preset constraints correspond to the search scenario, and the search scenario is determined by the search text; The step of inputting the multi-dimensional initial information into the large language model to obtain the prompt keywords includes: Based on the search scenario to which the search text belongs, the corresponding preset constraints and the multi-dimensional initial information are input into the large language model to obtain the suggested keywords.

8. The information search method according to any one of claims 1 to 7, characterized in that, The displayed prompt keywords extracted by the intelligent agent include: Switch between displaying the prompt keywords for the target video and the video title of the target video.

9. The information search method according to claim 8, characterized in that, The switching display of the prompt keywords and the video title of the target video includes: The prompt keywords and the video title are displayed in a scrolling manner; or, In response to the swipe gesture, the prompt keywords and the video title are switched on and off.

10. The information search method according to any one of claims 1 to 7, characterized in that, The displayed prompt keywords extracted by the intelligent agent include: The prompt keywords are displayed in the video frame of the target video.

11. The information search method according to claim 10, characterized in that, The display of the prompt keywords in the video frame of the target video includes: The video frame displays a tag for each prompt keyword; or, A carousel card is displayed in the video frame, which is used to display the prompt keywords in a loop.

12. The information search method according to claim 1, characterized in that, The method further includes: In response to the triggering operation of the prompt keyword, the target video is played in the playback interface of the target video, and the prompt keyword and summary information are displayed. The summary information is used to describe the video content of the corresponding node in the target video.

13. The information search method according to claim 12, characterized in that, The summary information includes multiple summary texts; the method further includes: In response to a trigger operation on any summary text, the target video is played from the node corresponding to the summary text, and prompt keywords that are semantically related to the summary text are highlighted.

14. An information search device, characterized in that, The device includes: The first display unit is configured to perform a search operation in response to the search interface, and display a target video based on the search text corresponding to the search operation, wherein the target video matches the search text; The second display unit is configured to display prompt keywords extracted by the agent based on the target video, the search text, and the currently logged-in account. The prompt keywords are used to describe the target video according to preset constraints, which are used to constrain at least one of the display format of the prompt keywords or the information extraction direction of the agent.

15. An electronic device, characterized in that, The electronic device includes: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the information search method as described in any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the information search method as described in any one of claims 1 to 13.

17. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the information search method as described in any one of claims 1 to 13.