Generating a prompt for a user to link annotations
By generating link annotations and prompts through generative models to process user and content data, the problem of users having difficulty obtaining detailed information is solved, the information quality of the search results interface and the reliability of user insights are improved, and the time spent on ineffective searches is reduced.
Patent Information
- Application Number
- CN202411558442.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-11-04
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2044-11-04
AI Technical Summary
In existing technologies, users find it difficult to obtain detailed information related to their interests from search results pages, and it is also difficult to obtain insights from other users, resulting in time-consuming and ineffective searches and insufficient information.
By processing user and content data through generative models, predictive suggestions are generated to provide link annotations in the search results interface. Insightful annotation suggestions are generated by leveraging user search history, browsing history, and previous annotations to guide users in providing more detailed information.
It improved the quality of information in the search results interface, reduced wasted browsing time, provided more detailed and reliable user insights, and reduced the number of subsequent queries and page redirects.
Smart Images

Figure CN119669554B_ABST
Abstract
Description
[0001] Related applications
[0002] This application is based on and claims priority to U.S. non-provisional application 18 / 392,648, filed December 21, 2023, which claims the benefit of U.S. provisional application serial number 63 / 596,484, filed November 6, 2023. The applicant claims priority and benefit for each of these applications, and the entire contents of all such applications are incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to generating prompts for obtaining link annotations. More specifically, this disclosure relates to determining when and how to prompt users for annotations about links associated with web resources, which can then be provided to other users. Background Technology
[0004] Understanding search results from a search results page can be difficult because titles and text snippets may provide limited information that may not be relevant to the user's interests, potentially leading to time-consuming web resource searches that may not yield the desired information. Obtaining additional information about a web resource can also be challenging, potentially involving additional searches that may or may not identify relevant information.
[0005] Additionally, gaining user insights can be difficult. Specifically, users may struggle to determine which words to use. Furthermore, these words may not be relevant to other users' interests and / or may not be rich enough to generate the desired results. Summary of the Invention
[0006] Various aspects and advantages of embodiments of this disclosure will be set forth in part in the following description, or may be learned from the description or by practicing the embodiments.
[0007] One example aspect of this disclosure relates to a computing system for comment suggestion generation and input retrieval, the system including: one or more processors; and one or more non-transitory computer-readable media that jointly store instructions that, when executed by the one or more processors, cause the computing system to perform operations. Operations may include: obtaining content data. The content data may be associated with a web resource. Operations may include: processing the content data with a generative model to generate a predicted suggestion. The suggestion may include a predicted text string associated with commenting on the web resource. Operations may include: providing the predicted suggestion for display using an input suggestion interface. The input suggestion interface may be configured to receive input. Operations may include: obtaining comment input data from a user computing system via the input suggestion interface. In some implementations, the comment input data may include user-generated comments on the web resource. Operations may include: storing data associated with the comment input data together with data associated with the web resource. The data associated with the comment input data may be stored in a searchable database to be provided for display in response to the web resource being offered as a search result.
[0008] In some implementations, operations may include: obtaining user data. User data may be associated with a specific user. A user computing system may be associated with a specific user. Processing content data with a generative model to generate predictive suggestions may include: processing content data and user data with a generative model. User data may include user search history data. The generative model may generate predictive suggestions based on information related to the topic of a web resource previously searched by a specific user. In some implementations, user data may include user browser history data. The generative model may generate predictive suggestions based on other web resources previously viewed by a specific user, including information related to the topic of the web resource. Operations may include: generating a graphics card based on user data, content data, and comment input data. The graphics card may include a user profile identifier for a specific user and data associated with the comment input data. Operations may include: storing the graphics card. The graphics card may include a graphical background generated by an image generation model based on the comment input data.
[0009] In some implementations, the operations may include: obtaining a search query, determining that a web resource is associated with the search query, and providing specific search results for display. Specific search results may include links to the web resource, the web resource's title, and data associated with the comment input data. Storing the data associated with the comment input data together with the data associated with the web resource may include: generating web resource annotations, and storing the web resource annotations together with multiple other web resource annotations associated with the web resource. In some implementations, the operations may include: providing web resource annotations and multiple other web resource annotations in an annotation interface that provides web resource annotations and multiple other web resource annotations across multiple graphics cards. The generative model may include an autoregressive language model. The generative model may be prompted to generate questions describing the request for information about the web resource.
[0010] Another exemplary aspect of this disclosure relates to a computer-implemented method for providing link annotation suggestions. The method may include: obtaining context data by a computing system including one or more processors. The context data may be associated with a specific content display instance. The specific content display instance may include a specific user viewing a specific content item. The method may include: determining an input request action by the computing system based on the context data. The input request action may include providing an input input interface to the user to obtain user input. The method may include: processing the context data with a generative language by the computing system to generate predictive suggestions. In some implementations, the predictive suggestions may include a natural language request for information generated based on the context data. The method may include: providing predictive suggestions by the computing system in the input input interface, and obtaining user-generated content by the computing system via the input input interface. The method may include: generating link annotations by the computing system based on the user-generated content. Link annotations may be generated to be provided for display in a search results interface in response to a specific content item being identified as a search result.
[0011] In some implementations, context data may be associated with the type of content being provided for display. Context data may be associated with a specific user associated with a particular content display instance. Context data may include search history data. The content being provided for display may be associated with a specific web resource. Context data may be associated with interaction data of links to a specific web resource across multiple social networking platforms. Input request actions may be determined based on interaction data. In some implementations, context data may include user data and content data. Input request actions may be determined based on a topic associated with the content being provided for display, which is one of several topics known to a specific user, determined based on user data. Context data may include previous annotations generated by a specific user. Predictive prompts may include a structure based on previous annotations and previous structures.
[0012] Another example aspect of this disclosure relates to one or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. Operations may include: obtaining a first search query at a first time and determining that a web resource responds to the first search query. Operations may include: obtaining content data. The content data may be associated with the web resource. Operations may include: processing the content data with a generative model to generate a predictive suggestion. The suggestion may include a predicted text string associated with commenting on the web resource. Operations may include: providing the predictive suggestion for display within an input suggestion interface. The input suggestion interface may include an input field. Operations may include: obtaining comment input data from a user's computing system via the input suggestion interface. The comment input data may include user-generated content. Operations may include: storing the user-generated content. Operations may include: obtaining a second search query at a second time. The second time may differ from the first time. Operations may include: determining that the web resource responds to the second search query and providing the user-generated content in a search results interface having data describing the web resource.
[0013] Another example aspect of this disclosure relates to a computing system for graphics card generation. The system may include: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. Operations may include: obtaining card data. The card data may describe content in the graphics card. The content may be associated with one or more topics. Operations may include: processing the card data to determine one or more entity tags associated with the content. The one or more entity tags may be associated with one or more topics. Operations may include: accessing a media content item database to obtain one or more media content items. The one or more media content items may be obtained based on determining that the one or more media content items are associated with one or more entity tags associated with the content. Operations may include: providing one or more media content items for display. The one or more media content items may be provided for display in an interactive user interface. The one or more media content items may be selectable to be inserted into the graphics card.
[0014] In some implementations, the graphics card may be associated with a link annotation. The link annotation may include user-generated content that links to a specific web resource. Operations may include: obtaining input selection associated with one or more media content items, generating an enhanced graphics card, and providing the enhanced graphics card for display. The enhanced graphics card may include at least a portion of the graphics card's content and at least a portion of one or more media content items. In some implementations, operations may include: obtaining adjustment input. The adjustment input may be associated with a request to enhance the enhanced graphics card. Operations may include: generating an updated graphics card based on the adjustment input. The updated graphics card may include an enhanced graphics card with one or more adjustments. Operations may include: providing the updated graphics card for display. One or more adjustments may include at least one of the following: a layout change of the enhanced graphics card, a cropping change of one or more media content items, a size change of one or more content items, a color change, or a template change.
[0015] In some implementations, the media content item database may include a user-specific database. This user-specific database may be associated with a specific user. This specific user may have already generated at least a portion of the content. In some implementations, the user-specific database may include an image library associated with a specific user. This image library may be stored on a server computing system associated with a specific content item storage platform. In some implementations, the user-specific database may include a local storage database on the user's computing device. The media content item database may include multiple media content items. In some implementations, these multiple content items may have been preprocessed to generate multiple corresponding metadata datasets.
[0016] Another example aspect of this disclosure relates to a computing system. The system may include: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. Operations may include: providing an input drafting interface for display. The input drafting interface may include a graphical user interface including multiple attribute options and text input boxes. The multiple attribute options may be associated with multiple candidate attributes for content item generation. Operations may include: obtaining a selection of a specific attribute option among the multiple attribute options via the input drafting interface. The specific attribute option may be associated with a specific candidate attribute. Operations may include: obtaining text input via the text input boxes of the input drafting interface. The text input may be associated with a schematic diagram for content item generation. Operations may include: processing the specific attribute options and text input with a generative model to generate a model-generated content item. The model-generated content item may include specific candidate attributes. In some implementations, the model-generated content item may be associated with a schematic diagram. Operations may include: providing the model-generated content item for display via the input drafting interface.
[0017] In some implementations, the operation may further include: obtaining input selections via an input drafting interface, and generating enhanced graphics cards based on the input selections. The enhanced graphics cards may include graphics cards enhanced to include content items generated by the model. The operation may include: providing the enhanced graphics cards for display. Multiple candidate attributes may include multiple different styles. Multiple different styles may be associated with at least one of multiple different art styles or multiple different writing styles.
[0018] In some implementations, multiple candidate attributes can include multiple different tones. These different tones can be associated with at least one of multiple different emotions or multiple different rhythm types. In some implementations, the generative model can be obtained from a generative model database based on the selection of specific attribute options. Specific attribute soft cues can be obtained based on the selection of specific attribute options. Specific attribute soft cues can include a learned set of parameters. The learned parameter set can be processed by the generative model to generate model-generated content items.
[0019] In some implementations, the first search query and the second search query can be different. Comment input data can include multimodal data. Multimodal data can include text data and image data.
[0020] Other aspects of this disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.
[0021] These and other features, aspects, and advantages of the various embodiments of this disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the disclosure and, together with the description, serve to explain the relevant principles. Attached Figure Description
[0022] Referring to the accompanying drawings, a detailed discussion of embodiments is set forth in this specification for those skilled in the art, in which:
[0023] Figure 1 A block diagram of an example link annotation generation system according to an example embodiment of the present disclosure is depicted.
[0024] Figure 2 A block diagram of an example user prompting system according to an example embodiment of the present disclosure is depicted.
[0025] Figure 3 A flowchart is depicted for an example method for performing link annotation hints according to an example embodiment of the present disclosure.
[0026] Figure 4 An illustration depicts an example prompt based on an example embodiment of the present disclosure.
[0027] Figure 5A An illustration depicts an example annotation interface with topic hints according to an example embodiment of the present disclosure.
[0028] Figure 5B An illustration depicts an example annotation interface with similar article prompts according to an example embodiment of this disclosure.
[0029] Figure 5C An illustration depicts an example prediction prompt according to an example embodiment of the present disclosure.
[0030] Figures 6A to 6C An illustration depicts an example link annotation entry point according to an example embodiment of this disclosure.
[0031] Figure 7 A flowchart is depicted for an example method for performing link annotation generation according to an example embodiment of the present disclosure.
[0032] Figure 8 A flowchart illustrating an example method for performing the example method shown in the link annotation, according to an example embodiment of the present disclosure, is depicted.
[0033] Figures 9A to 9G An illustration depicts an example graphics card interface according to an example embodiment of the present disclosure.
[0034] Figure 10An illustration depicts an example card generation interface according to an example embodiment of the present disclosure.
[0035] Figure 11 An illustration depicts an example content item generation interface according to an example embodiment of the present disclosure.
[0036] Figure 12 An illustration depicts an example image suggesting an interface according to an example embodiment of the present disclosure.
[0037] Figure 13 A flowchart is depicted for an example method for performing image suggestion according to an example embodiment of the present disclosure.
[0038] Figure 14 A flowchart is depicted for an example method for performing content item generation according to an example embodiment of the present disclosure.
[0039] Figure 15A A block diagram of an example computing system that performs the execution link annotation hints according to an example embodiment of the present disclosure is depicted.
[0040] Figure 15B A block diagram of an example computing system that performs the execution link annotation hints according to an example embodiment of the present disclosure is depicted.
[0041] The repeated reference numerals across multiple figures are intended to identify the same features in various implementations. Detailed Implementation
[0042] Generally, this disclosure relates to generating prompts for user data entry. Specifically, the systems and methods disclosed herein can utilize context determination (e.g., determining the context in which a user is likely to provide a comment and / or determining the comment gap and / or content gap for a particular link) to determine the input entry interface to provide (e.g., a link comment input entry interface), and can utilize a generative model (e.g., a large language model) to generate prompts based on user data (e.g., user search history and / or user browsing history) and / or content data (e.g., the topic and / or type of content). For example, a user can be prompted to provide a comment on a specific web resource (and / or other content item) on a search results page, during web resource browsing, and / or at the next search instance. Prompts can be generated based on previous user comments, previously viewed content, the topic and / or type of content to provide the user with formatted request information that leads to the generation of insightful comments.
[0043] Link comments can provide additional information about web resources without requiring the user to consult the web resource itself, and these comments can be provided by other users. Systems and methods can determine when to provide link comment prompts to users based on contexts identified as associated with the inclusion of valuable comments. For example, a particular user might provide more reliable and / or more detailed information about a specific topic based on prior knowledge and / or previously generated comments. Additionally and / or alternatively, specific content types can be identified as associated with user comments and / or user confusion.
[0044] The prompts provided to users can "inspire" them to provide more detailed information and / or guide them to leave comments about specific topics and / or features of the web resource. The generative model can process user data and / or content data to generate predictive prompts. Specifically, the generative model can utilize the user's search history, browsing history, previous comments, and / or other user data to generate suggested comments, prompting questions, and / or comment templates. Alternatively and / or additionally, the generative model can utilize semantic understanding of the web resource, topic classification, content type classification, other comments and / or other content data associated with the web resource to generate suggested comments, prompting questions, and / or comment templates.
[0045] The input interface can provide predictive prompts to the user. The input interface can then obtain input from the user (e.g., comment input data) to generate user-generated content describing the link annotations. In some implementations, graphic cards can be generated based on the link annotations. Graphic cards can include the user-generated content of the link annotations, user profile identifiers (e.g., name and / or image), link information, and / or graphic backgrounds. Link annotations and / or graphic cards can be stored in association with web resources. The stored link annotations and / or graphic cards can then be retrieved in response to one or more users searching for web resources and / or one or more users interacting with the annotation interface.
[0046] Understanding search results from search results pages can be difficult because titles and text snippets may provide limited information that may not be relevant to the user's interests, potentially leading to time-consuming web resource searches that may not yield the desired information. Additionally and / or alternatively, obtaining additional information about web resources can be difficult, which may include additional searches that may or may not identify relevant information. Social media posts, blog posts, and / or commentary on web resources and / or entities associated with them may lack detail, be misleading, and / or lack context and / or perspective.
[0047] Link annotations (e.g., link annotations obtained from users and / or generated by generative models) can provide additional information about web resources, informing other users of their relevance to their requests. Link annotations may be provided on search results pages and / or displayed in annotation interfaces accessible from search results pages and / or from web resources. Link annotations may be provided in graphical cards, text panels inline with text fragments, and / or in other formats.
[0048] Determining when and how to generate suggestions for annotations can be based on factors such as: identifying annotation differences (e.g., whether there are many annotations related to an article on blogging and / or social media platforms, but relatively few in the annotation interface of a search platform); identifying user-specific interests (e.g., whether the resource is similar to other articles the user has viewed in the past); resource trends (e.g., whether the resource and / or similar resources have been commented on previously); identifying resources worth annotating (e.g., does the annotation provide usefulness?); and / or other determinations. Suggestions can be generated based on generative language model processing, which may include processing previous annotations (e.g., other annotations made by the user and / or other users), processing search queries, processing web resources, and / or processing other data.
[0049] Link annotation prompts can be used to initiate and / or facilitate the collection of information about web resources. This information can then be provided to other users, who can identify user comments, user summaries, and / or other user-identified details. The obtained annotations can then be presented in the search results interface and / or the discovery feed. Link annotation prompts can be determined based on the identification of annotation differences, user-specific interests, resource trends, the identification of resources worthy of annotation, and / or other determinations.
[0050] Obtaining additional information from other users can be useful for users to determine the topic and quality of search results that may not be discernible from traditional search result displays; however, obtaining helpful and detailed information can be difficult. An interface for obtaining detailed information from relevant users can be provided by determining when to generate prompts for users regarding annotations and by generating context-aware prompts.
[0051] Incentivizing users to contribute information (e.g., comments, reviews, insights, etc.) when they see search results can be challenging. A suggestion generation system can be leveraged to deliver the right suggestions to the right users at the right time and place, based on insights. Specifically, suggestions can help create posts with desired characteristics (e.g., the required level of detail, the desired topic, and / or other features). The suggestion generation system disclosed herein can facilitate the generation (or creation) of comments by specific users, and / or can facilitate the generation (or creation) of comments for specific web resources (and / or content items) via the suggestion generation system.
[0052] In some implementations, the systems and methods disclosed herein can be used to prompt users to generate self-generated content. Self-generated content can include user recipes, user tutorials, user graphics, lifestyle updates, shared links, and / or other user-generated content. Self-generated content can be generated in a free-form format and / or based on model-generated prompts. In some implementations, one or more machine learning models can be used to generate content templates, and / or one or more machine learning models can be used to enhance user-provided content (e.g., refactoring and / or redesigning text, images, audio, interface elements, and / or video).
[0053] In some implementations, link annotations and / or interactions with them can be used to adjust web resource ranking, web resource tagging, web resource embedding, and / or web resource indexing. For example, in some implementations, link annotations can be processed to determine the quality of a web resource. Quality determination can be based on processing link annotations using one or more machine learning models (e.g., sentiment analysis models, language models, classification models, etc.). One or more machine learning models can be used to process link annotations to determine topics associated with the web resource, determine web resource bias, web resource usefulness, and / or web resource orientation. Link annotations can be used to suggest additional content, can be embedded for embedding-based search, and / or can be used for query suggestions.
[0054] Link annotations in the annotation interface can be ranked and / or displayed based on interaction, quality determined by a machine learning model, responsiveness to queries, level of detail, and / or other attributes. In some implementations, user-generated link annotations can be provided to all other users, to users within that user's social network, and / or only to users identified as associated with that user based on interests, location, and / or activity.
[0055] Link annotations can be used for multiple different content items and are not limited to web resources. For example, the systems and methods disclosed herein can be used to generate prompts and / or interfaces for obtaining, inspiring, and / or generating link annotations for sources of local files (e.g., documents, images, videos, etc. on a device), intranet files, and / or other content items, which may include folders on external drives, documents in the cloud, etc.
[0056] In some implementations, the input interface may include an open input interface that provides one or more options for providing user input. Alternatively and / or additionally, the input interface may include multiple features and / or options for generating user-generated content, which may be used for linking annotations and / or stand-alone content. The input interface may include a stand-alone content item user interface that allows users to add images, links, and / or content of different template types and may be interactive. Interactive user interfaces may include image suggestions, template suggestions, text suggestions, layout suggestions, link suggestions, widget suggestions, and / or other options (e.g., other types of suggestions).
[0057] Image suggestions may include processing user-input text, data associated with web resources, generated prompts, stock photo libraries, and / or user-associated image databases (e.g., online image galleries associated with the user and / or local images on the user's computing device) to determine images relevant to a specific context (e.g., related to user input, web resources, and / or generated prompts). Image suggestions may include identifying one or more entities, topics, and / or features associated with web resources and / or user input, and then processing stock photo libraries and / or image databases associated with the user to determine one or more specific images associated with one or more entities, topics, and / or features associated with web resources and / or user input. For example, if a web resource discusses pasta recipes, stock image libraries and / or user image libraries may be searched for images depicting pasta, cooking, pasta ingredients, and / or kitchens. Another example may include determining that the text of a generated prompt can be associated with a trip to Mexico, and identifying and suggesting one or more images from the user's image library based on location metadata, feature detection, optical character recognition, and / or other identification techniques that can be used to identify the one or more images as associated with a trip to Mexico. In some implementations, image suggestions can be based on generating cue embeddings and / or web resource embeddings, and then performing an embedding search based on multiple image embeddings associated with one or more image databases. Suggestion determination and display can be performed against images, videos, document files, audio, text data, templates, and / or other data.
[0058] Additionally and / or alternatively, the interactive input interface may include a "Help Me Write" feature. The "Help Me Write" feature may be a selectable user interface feature that provides an interface to a generative language model for generating text for user-generated content. The "Help Me Write" feature may include dropdown menus for selecting specific tone, style, format, length, and / or other attributes for the model-generated text. The "Help Me Write" feature may process user input to adjust and / or change the style, tone, format, language, vocabulary, length, and / or simplicity of the input text. For example, a user may select a tone from multiple tone options, enter a text string, and the input interface may provide the text string and the selected tone hints to a generative language model (e.g., a large language model) to generate a model-generated text response, which can then be used for user-generated content (e.g., link comments and / or stand-alone content). Alternatively and / or additionally, the input interface may interact with different generative language models associated with different attributes in response to the selection of different attribute options. Different generative models can be trained and / or tuned for specific attributes.
[0059] The systems and methods disclosed herein provide several technical effects and benefits. As an example, the systems and methods can provide an interactive user interface that can be used to generate prompts and obtain user input data. Specifically, the systems and methods disclosed herein can utilize one or more machine learning models to determine when to request link annotations and generate prompts for requesting information. For example, the generative model can process user data, content data, and / or other contextual data to determine the information request action to be performed. Additionally and / or alternatively, the generative model can generate prompts for requesting information based on user data, content data, and / or other contextual data. Prompts can be provided to the user, user input can be received, and link annotations can be generated and stored.
[0060] Another technical benefit of the systems and methods disclosed herein is the ability to utilize user and content data to determine which users can provide credible information about a particular web resource and / or when to prompt users to provide information. For example, it can be determined that a user is knowledgeable on a particular topic and / or a frequent commenter on a given type of content. Based on this determination, the user can be prompted to provide a link comment for a given web resource. Alternatively and / or additionally, the topic of the content, the type of the content, and / or other interactions with the content can be used to determine whether a web resource is “appropriate” for commenting. Prompts can be generated using generative models to provide both user-aware and content-aware prompts.
[0061] The systems and methods disclosed in this paper address the problems arising from the acquisition, processing, and transmission of data from multiple databases across multiple sources by computing systems. The vast amounts of data available to users may provide the possibility of false positives, misleading information, and / or a lack of verification. Text snippets, titles, and / or sample images in search results interfaces can provide some detail about the content of web resources; however, information from other users can provide further insights into the topic, trustworthiness, and / or expected content, which can be leveraged to reduce instances of users navigating and consulting irrelevant web resources.
[0062] Another example of the technical effects and benefits involves improved computational efficiency and functional enhancements of computing systems. For instance, the systems and methods disclosed herein can utilize annotation generation to provide an interface that offers information about links, which can alleviate the tedium of searching search results by providing user-based verification. The reduced number of subsequent queries and page redirects can reduce latency at the user's device and lower the computational cost of the search engine. Exemplary embodiments of this disclosure will now be discussed in more detail with reference to the accompanying drawings.
[0063] Figure 1 A block diagram of an example link annotation generation system 10 according to an exemplary embodiment of the present disclosure is depicted. In some implementations, the link annotation generation system 10 is configured to receive and / or obtain context data 12 describing user data, content data, and / or other context data associated with a web page and / or viewing stance, and as a result of receiving the context data 12, generate, determine, and / or provide a predictive prompt 18 describing a generated natural language request for information from a user. Therefore, in some implementations, the link annotation generation system 10 may include a generation model 16 operable to process the context data 12 to generate the predictive prompt 18, which includes a text string describing a question, command, template, and / or suggestion.
[0064] Specifically, the link annotation generation system 10 can obtain context data 12. Context data 12 may include user data (e.g., data associated with users viewing search results pages, entering search queries, and / or viewing discovery feeds), content data (e.g., data associated with content in web resource 14), and / or other context data (e.g., time, query trends, comment trends, news, etc.). Context data 12 may include user search history data, user browsing history data, user purchase history data, user profile data, user annotation history data, web resource topic tag data, web resource content type tags, other annotations about the web resource, and / or other data. Context data 12 can be generated using personalized machine learning models and / or one or more other machine learning models.
[0065] Contextual data 12 can be obtained based on web resources 14 being provided for display, previously viewed, and / or associated with search results on search results pages. Contextual data 12 can be obtained and / or generated based on one or more user interactions, one or more global trends, and / or based on web resources 14 being associated with specific types of content (e.g., editorials, tutorials, blogs, news articles, sports score trackers, etc.).
[0066] Generative model 16 (e.g., an autoregressive language model, a diffusion model, and / or one or more other generative models) can process contextual data 12 to generate predictive prompts 18. Generative model 16 may include language models (e.g., large language models, visual language models, and / or other language models), text-to-image generation models, and / or other generative models. Predictive prompts 18 may include text data, image data, audio data, latently encoded data, and / or multimodal data. Predictive prompts 18 may include questions to which a user can respond, templates for drafting annotations, and / or one or more optional annotation options. For example, predictive prompts 18 may include questions generated based on semantic analysis and / or topic determination for the web resource, templates generated based on annotations previously generated by a specific user and / or other users, and / or optional annotation options based on previous comments provided for similar web resources. In some implementations, predictive prompts 18 may describe a request for a specific type of information about web resource 14. Alternatively and / or additionally, predictive prompts 18 may describe general information about web resource 14. Prediction hints 18 may include novel text strings that have not been previously provided by the user and / or that are related to web resource 14. Prediction hints 18 may include multiple prediction characters, words, pixels, signals, and / or structures.
[0067] Predictive prompts 18 can be provided for display in the input interface. Users can interact with the input interface to generate user-generated content 20, which can be sent to a server computing system (e.g., a search engine computing system). User-generated content 20 may include text data, image data, audio data, video data, latently encoded data, and / or multimodal data. User-generated content 20 may describe annotations related to web resource 14. Annotations may describe comments, opinions, reviews, verifications, and / or indications of quality and / or topic. User-generated content 20 may include annotations displayed in a graphics card with one or more graphics, one or more widgets, one or more links, one or more media content items, and / or a graphic background.
[0068] The link annotation generation system 10 can index link annotations 22 along with web resource 14. This index can be used to provide user-generated content 20, including link annotations, for display when providing search results for web resource 14. Alternatively and / or additionally, the user-generated content 20 can be stored in an annotation database for display in the annotation interface when selected by one or more users.
[0069] Figure 2 A block diagram of an example user prompting system 200 according to an example embodiment of the present disclosure is depicted. The user prompting system 200 is similar to... Figure 1 The link annotation generation system 10 differs from the user prompt system 200 in that it further includes an action determination block 230.
[0070] User suggestion system 200 can obtain content data 224 and user data 226. This data can be obtained in response to search queries, back events returning to the search results page (e.g., returning to the search results page after viewing a web resource), the next search instance, the web resource as the next instance of the search results, and / or based on other triggering events. Content data 224 can describe the content within the web resource, which may include text, images, videos, layouts, audio files, transitions, latent encoded data, related links / web resources, interaction history with the web resource, and / or other data associated with the web resource and / or other similar web resources. User data 226 may include data describing the following: user search history (e.g., logs of previous queries obtained from the user), user browsing history (e.g., logs of previously visited web pages and / or platforms), user application history (e.g., logs of previous interactions with applications), user purchase history (e.g., logs of previously obtained products and / or services), user profile (e.g., user identifier, user preferences, user name, user account, and / or user contacts), annotation history (e.g., logs of previously provided / generated annotations), and / or social media networks and / or activities.
[0071] The context determination block 228 can be used to process content data 224 and / or user data 226 to determine the context. The context determination block 228 may include one or more machine learning models and / or one or more deterministic functions. The context determination block 228 can generate context data.
[0072] Action determination block 230 can be used to process context data to determine the input request action to be performed. Action determination block 230 may include one or more machine learning models and / or one or more deterministic functions. Context determination and / or action determination may be performed based on heuristic methods.
[0073] Input request actions may include generating prompts and providing users with an input interface that includes those prompts to obtain a link annotation for a given web resource. Input request actions may be determined based on: the likelihood of a user responding, the user's trustworthiness, the user's experience, the user's knowledge, differences in association between the user and previous annotation providers, content gaps describing differences in annotations for a specific web resource compared to similar web resources, comment gaps describing differences in the amount of annotations compared to interactions with the link on one or more blogs or social media platforms, the web resource's subject matter, the search topic, the content type, the content's intent, and / or other data.
[0074] Specific content types (e.g., news articles, short stories, movies, short dramas, blog posts, and / or social media posts) may be identified as more likely to be interacted with for link annotation generation, and / or may be identified as potentially providing greater benefits from link annotation. Additionally and / or alternatively, interactions with the web resource on other platforms can be identified. If the quantity and / or quality of interactions on other platforms are determined to meet a threshold difference compared to the current platform, input request actions can be identified more frequently. For example, thresholds for input request actions can be adjusted based on interactions on other platforms. Alternatively and / or additionally, thresholds can be adjusted based on search and / or viewing trends associated with the web resource.
[0075] The input request action can be immediately followed by a determination to execute and / or can be provided as a "nudge" at a later time, which can be a determined time when the probability of a response is higher (e.g., when the user is in a specific location (e.g., at home), when the user's calendar is empty, a specific time of day when phone activity increases, and / or when the next user search instance). The "nudge" can be provided via device notification, email, and / or app-based notification.
[0076] Then, the user suggestion system 200 can determine, based on the input request action, to generate a suggestion for requesting comments from the user using contextual data. Contextual data may include user data 226 (e.g., data associated with a user viewing a search results page, entering a search query, and / or viewing a discovery feed), content data 224 (e.g., data associated with content in a web resource), and / or other contextual data (e.g., time, query trends, comment trends, news, etc.). Contextual data may include user search history data (e.g., a list of previously searched search queries, which may include queries associated with the same topic as the web resource), user browsing history data (e.g., a list of previously viewed web pages, which may include web pages associated with the same topic as the web resource), user purchase history data, user profile data (e.g., user's name, occupation, education, preferences, etc.), user comment history data, web resource topic tag data, web resource content type tags, other comments about the web resource, and / or other data. Contextual data may be generated using personalized machine learning models and / or one or more other machine learning models.
[0077] Contextual data can be obtained based on web resources provided for display, previously viewed, and / or associated with search results on search results pages. Contextual data can be obtained and / or generated based on one or more user interactions, one or more global trends, and / or based on web resources associated with specific types of content (e.g., editorials, tutorials, blogs, news articles, sports score trackers, etc.).
[0078] Generative model 216 (e.g., text generation model, image generation model, audio generation model, video generation model, and / or multimodal media content item generation model) can process contextual data to generate predictive prompts 218. Generative model 216 may include language models (e.g., large language models, visual language models, and / or other language models), text-to-image generation models, and / or other generative models. Predictive prompts 218 may include text data, image data, audio data, latently encoded data, and / or multimodal data. Predictive prompts 218 may include questions to which a user can respond, templates for drafting annotations, and / or one or more selectable annotation options. For example, predictive prompt 218 may include questions generated based on semantic analysis and / or topic determination of the web resource (e.g., for an article about a cold case, prompts may include “what are your thoughts on the analysis of the cold case?”, “who do you think committed the crime?”, “was the article understandable and comprehensive with regards to the forensic evidence?”, etc.). In some implementations, predictive prompt 218 may include templates generated based on comments previously generated by a particular user and / or other users (e.g., if a user typically begins their comments with a preposition, the prompt may include a preposition-starting template that mimics the style and tone of previous comments). Additionally and / or alternatively, predictive prompt 218 may include optional annotation options based on previous comments provided for similar web resources (e.g., “the analysis of the case was thorough, understandable, and had a plausible conclusion”, “forensic analysis lacked basis in verified science”, “the article is more of a fan fiction that an actual article”, etc.).In some implementations, prediction prompt 218 may describe a request for a specific type of information about the web resource (e.g., for a politician's biography, "what were your thoughts on their upbringing?", "please provide your insight on the epilogue", "in my experience as a congressional historian, this biography is accurate / untrustworthy / well-written / poorly-structured", etc.). Alternatively and / or additionally, prediction prompt 218 may describe general information about the web resource. Prediction prompt 218 may include novel text strings not previously provided by the user and / or related to the web resource. Prediction prompt 218 may include multiple prediction characters, words, pixels, signals, and / or structures.
[0079] In some implementations, specific items, details, and / or structure of the search query can be used to determine the user's level of experience and / or knowledge on a particular topic. The search query can be included in contextual data, and the generative model 216 can generate predictive tips 218 that reflect the determined level of experience and / or knowledge. Additionally and / or alternatively, previous search queries can be used to determine search query chains to determine search intent. The search intent can then be used to generate predictive tips 218 associated with the search intent.
[0080] In some implementations, the predicted prompt 218 may vary based on the user's tendency to provide link annotations, based on the user's previous annotations, based on the user's trust level, and / or other user data. If the user has never provided link annotations before and / or has only provided a small number of link annotations previously, the predicted prompt 218 can be configured as a general prompt, a multiple-choice prompt, and / or a dialog format. Predictive prompts 218 can be generated for experienced users to provide them with direct prompts, annotation templates, and / or options based on previous interactions.
[0081] Predictive prompts 218 can be provided for display in the input interface. Users can interact with the input interface to generate user-generated content 220, which can be sent to a server computing system (e.g., a search engine computing system). User-generated content 220 may include text data, image data, audio data, video data, latently encoded data, and / or multimodal data. User-generated content 220 may describe annotations related to web resources. Annotations can describe comments, opinions (e.g., “I think the trade discussed in the article was fair based on the long-term outcomes for both teams.”), reviews (e.g., “the short story lacked proper pacing and direction with the main character having zero character growth”), verifications (e.g., “the facts in this article match those of other reputable sources”), and / or indications of quality and / or subject matter (e.g., “such a well-written play on the perils of love in war-stricken towns”). User-generated content 220 may include annotations displayed in a graphics card with one or more graphics, one or more widgets, one or more links, one or more media content items, and / or a graphic background. For example, the text for link annotations can be provided as stylized text with color, thus juxtaposing it with an image generated by the model and used as a background.
[0082] In some implementations, the generative model 216 can process user-generated content 220 and generate follow-up prompts. These follow-up prompts can request additional information and / or provide options for further customization.
[0083] User suggestion system 200 can store comments 222 along with web resources. An index can be used to provide user-generated content 220, including link comments, for display when providing search results for web resources. Alternatively and / or additionally, user-generated content 220 can be stored in a comment database for display in the comment interface when selected by one or more users.
[0084] For example, a specific user and / or other users can enter a search query. The search engine system can determine that the web resource responds to the search query. Search results associated with the web resource can be provided in a search results interface (e.g., a graphics card) that contains data such as the web resource's title, media clips, and descriptive link annotations.
[0085] Figure 3 A flowchart is depicted for an example method performed according to an example embodiment of this disclosure. Although Figure 3 The steps performed in a particular order are depicted for illustrative and discussion purposes, but the method of this disclosure is not limited to the particular illustrated order or arrangement. The steps of method 300 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of this disclosure.
[0086] At position 302, the computing system can obtain content data. Content data can be associated with a web resource. Content data can describe the content of the web resource, which may include text data, image data, video data, audio data, latently encoded data, and / or multimodal data. Content data may include data describing the following: the topic of the web resource, the type of content, additional comments received from other users, web resource metadata, the author of the web resource, and / or entities associated with the web resource. Content data may include content tags, the overall content of the web resource, a summary of the content, media clips, and / or content embeddings.
[0087] At 304, the computational system can use a generative model to process content data to generate a predicted prompt. The prompt may include a predicted text string associated with commenting on a web resource. The generative model may include an autoregressive language model. In some implementations, the generative model may be prompted to generate a question describing a request for information about the web resource. The generative model may include a transformer model. The generative model may have been trained, configured, and / or prompted to perform semantic understanding on the web resource, and then generate a prompt (e.g., a question) based on the semantic understanding. The predicted prompt may include multiple predicted characters, which may be specifically determined based on the content data. The predicted prompt may query for the quality of the web resource. The predicted prompt may query for opinions and / or comments on the web resource.
[0088] In some implementations, the computing system can obtain user data. User data can be associated with a specific user. Processing content data with a generative model to generate predictive suggestions can include processing both content data and user data using a generative model. User data can include user search history, user browsing history, user social networks, user preferences, user profile information, user location, user purchase history, and / or user connections. The generative model can generate predictive suggestions based on information about a specific user's previous searches related to the topic of a web resource. Alternatively and / or additionally, the generative model can generate predictive suggestions based on other web resources that a specific user has previously viewed, including information related to the topic of the web resource. The generative model can determine that a specific user is associated with a specific topic, a specific type of content, a specific opinion, and / or a specific commentary context, and can generate predictive suggestions based on that determination.
[0089] At point 306, the computational system may use an input suggestion interface to provide prediction suggestions for display. The input suggestion interface can be configured to receive input. The input suggestion interface may include multiple selectable user interface elements. In some implementations, the input suggestion interface may include an input field for receiving input from a user. The input suggestion interface may include an upload element for uploading media content items (e.g., documents, images, text, video, audio files, etc.). In some implementations, the input suggestion interface may include an interface for users to provide input to a generative model to generate model-generated annotations based on user input. Alternatively and / or additionally, the multiple selectable user interface elements may include one or more selectable templates for generating user-generated content (e.g., user-generated annotations). One or more templates may be generated based on content items (e.g., previously generated annotations) previously generated by a specific user.
[0090] At point 308, the computing system can obtain comment input data from the user's computing system via an input prompt interface. Comment input data may include text data, image data, audio data, latently encoded data, and / or multimodal data. Comment input data may include one or more selections, one or more text strings, and / or one or more uploaded files. Comment input data may include user-generated comments on web resources. User-generated comments may include comments on the quality of the web resource, the topic of the web resource, and / or other aspects of the web resource.
[0091] At point 310, the computing system may store data associated with the comment input data together with data associated with the web resource. Storing data associated with the comment input data together with data associated with the web resource may include: generating web resource annotations, and storing the web resource annotations together with multiple other web resource annotations associated with the web resource. The comment input data may be indexed in association with the web resource, and may be stored in a database to be provided along with search results for the web resource. Data associated with the comment input data may be stored in a searchable database to be provided for display in response to the web resource being identified and / or provided as a search result.
[0092] In some implementations, the computing system can provide web resource annotations and multiple other web resource annotations in an annotation interface that provides web resource annotations and multiple other web resource annotations across multiple graphics cards.
[0093] Additionally and / or alternatively, the computing system may generate a graphics card based on user data, content data, and comment input data. The graphics card may include a user profile identifier for a specific user and data associated with the comment input data. The computing system may then store the graphics card. The graphics card may include a graphical background generated using an image generation model based on the comment input data.
[0094] In some implementations, the computing system can obtain a search query, determine that a web resource is associated with the search query, and provide specific search results for display. Specific search results may include links to the web resource, the web resource's title, and data associated with comment input data.
[0095] Figure 4 An illustration depicts an example prompt according to an exemplary embodiment of the present disclosure. Specifically, Figure 4 The example input interface 402 is depicted. The example input interface 402 may be provided in response to determining that an input request action should be performed and generating predictive prompts. The input interface 402 may include a link to and / or a reference to a web resource 404, a writing panel 406 for displaying input provided during user-generated content generation, one or more selectable predictive prompts 408, and / or one or more user interface elements for providing audio input and / or multimedia input (e.g., images).
[0096] One or more optional predictive hints 408 may include predictive hints generated by processing the content of web resource 404 using a generative language model. Alternatively and / or additionally, one or more optional predictive hints 408 may include predictive hints generated by processing annotations associated with articles similar to those provided by web resource 404 using a generative language model.
[0097] Suggested predictive prompts can be provided in multiple formats. Additionally and / or alternatively, the number and / or length of predictive prompts can vary based on content, user, and / or other contextual data. For example, three options 412 or ten options 414 could be provided. Annotation options may include multiple selectable prompts that can be chosen by the user as link annotations and / or as part of a user's link annotations.
[0098] Input interface 402 can be used to receive text input (e.g., via a graphical keyboard interface), audio input (e.g., via one or more microphones), selections (e.g., selections of user interface elements associated with predictive hint annotation options), and / or media content item input (e.g., image uploads). The received input can be provided for display in a preview window of writing panel 406, and then the received input can be used to generate user-generated content items, which may include link annotations.
[0099] Figure 5A An illustration depicts an example annotation interface with topic prompts according to an example embodiment of the present disclosure. The example annotation interface can be used to learn more about web resource 502, read other users' thoughts on web resource 502, and / or generate and provide new annotations about web resource 502. Specifically, Figure 5A The document depicts a link to web resource 502, a previously provided comment 504, one or more prompts 506 for requesting information from the user, and an approval user interface element 508. The link to web resource 502 may include a thumbnail, URL, and title. The previously provided comment 504 may include an image and / or text provided by another user. The previously provided comment 504 may provide interactive data, including approvals, comments, and / or likes. The one or more prompts 506 for requesting information from the user may include discussion topics in which the user responds to a user's comment. The approval user interface element 508 can be used to interact with web resource 502 and / or the previously provided comment 504.
[0100] Figure 5BAn illustration depicts an example annotation interface with similar article prompts according to an example embodiment of this disclosure. The example annotation interface can be used to learn more about web resource 502, read other users' thoughts on web resource 502 (e.g., opinions on the topic and / or quality of the web resource), and / or generate and provide new annotations about web resource 502 (e.g., users can provide their own details and / or thoughts about the web resource). Specifically, Figure 5B The document depicts a link to web resource 502, a previously provided comment 504, a user interface element 510 for commenting on articles similar to web resource 502, and suggested articles 512. The link to web resource 502 may include a thumbnail, URL, and title. The previously provided comment 504 may include an image and / or text provided by another user. The previously provided comment 504 may provide interactive data, including approvals, comments, and / or likes. The user interface element 510 for commenting on articles similar to web resource 502 can be selected to open an input field for providing link comments on one or more similar web resources (which may include suggested articles 512).
[0101] Figure 5C An illustration depicts an example predictive prompt according to an exemplary embodiment of this disclosure. Specifically, a generative model can be used to process content data associated with a web resource and / or user data associated with a specific user being prompted, to generate one or more predictive prompts. Figure 5C Example question prompt 514 and starting point prompt 516 are depicted. Question prompt 514 may include questions to solicit information from the user regarding a specific topic and / or subtopic. Starting point prompt 516 may include introductory sentences for the user to build their link notes based on. When provided in the search results interface and / or annotation interface, user-generated content including link notes may include starting prompts.
[0102] Figures 6A to 6C An illustration depicts an example link annotation entry point according to an exemplary embodiment of this disclosure. Entry points to the link annotation generation interface may be provided in various different interfaces and / or media. Specifically, Figure 6A A sample search results interface with sample search results 602 is depicted. Contextual data associated with the search instance (e.g., user data, search queries, and / or search results) can be processed to determine how to provide link annotation suggestions. Figure 6A The first example prompt 604 is depicted in response to the determination that the user has previously searched for this topic. Additionally, Figure 6AA second example tip 606 is described in response to determining interaction trends and / or determining that a particular web resource is frequently interacted with for the type of content generated for annotation.
[0103] Figure 6B Three distinct entry points for linking and generating comment suggestions are described. The first entry point 608 may include a dropdown menu in a browser application. The user can select an option dropdown menu and choose a comment generation option from it. The second entry point 610 may be included within a search results interface. The user can select an option within the search results interface to comment on previously viewed content. The third entry point 612 may include an overlay user interface element that can be provided in a viewing window of the web resource's content. This overlay user interface element may be provided by the web resource, the browser, and / or the operating system of the user's computing device.
[0104] Figure 6C The search entry point is described. Specifically, widgets and / or graphical panes can be provided to users to offer a review of previous searches. This can include selecting dropdown list element 614 to choose a web resource from a list of viewed content 616, indicating which web resource was helpful for the user's specific search. This selection can be used to generate link annotations, re-rank web resources, and / or navigate to a link annotation generation interface.
[0105] Figure 7 A flowchart is depicted for an example method performed according to an example embodiment of this disclosure. Although Figure 7 The steps performed in a particular order are depicted for illustrative and discussion purposes, but the method of this disclosure is not limited to the particular illustrated order or arrangement. The steps of method 700 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of this disclosure.
[0106] At point 702, the computing system can obtain context data. Context data can be associated with a specific content display instance. A specific content display instance can include a specific user viewing a specific content item. In some implementations, context data can be associated with the type of content being provided for display (e.g., articles, academic papers, blog posts, encyclopedia entries, videos, media content libraries, and / or other types of content). Additionally and / or alternatively, context data can be associated with a specific user associated with a specific content display instance. Context data can include search history data, browsing history data, user profile data, purchase history data, social network data, and / or other user data.
[0107] At point 704, the computing system can determine the input request action based on context data. The input request action may include providing an input interface to the user to obtain user input. The content being provided for display may be associated with a specific web resource. Context data may be associated with interaction data (e.g., linked posts, comments, reposts, likes, and / or mentions) on multiple social networking platforms targeting a specific web resource. In some implementations, the input request action may be determined based on interaction data. Context data may include user data and content data. Additionally and / or alternatively, the input request action may be determined based on a topic associated with the content being provided for display, chosen from multiple topics known to the specific user based on user data.
[0108] At point 706, the computational system can use a generative language model to process contextual data to generate predictive prompts. Predictive prompts may include natural language requests for information generated based on the contextual data. The contextual data may include previous annotations generated by a specific user. Predictive prompts may include a structure based on previous annotations and a previous structure. In some implementations, the generative model may process content data to generate predictive prompts based on the content of a specific web resource.
[0109] At point 708, the computing system may provide predictive suggestions within the input field. Predictive suggestions may be provided adjacent to the input fields used to receive and display input text and / or images. The input field may include a panel adjacent to search results for a specific web resource. Alternatively and / or additionally, the input field may be provided as a pop-up interface, and / or may be redirected based on one or more inputs.
[0110] At point 710, the computing system can obtain user-generated content via an input field. User-generated content may include text data, image data, video data, audio data, latently encoded data, statistical data, and / or multimodal data. User-generated content can be obtained via an upload field and / or an input field.
[0111] At point 712, the computing system can generate link annotations based on user-generated content. In some implementations, the computing system can generate a graphic card including the link annotations. The graphic card may include a graphic background that can be selected by the user and / or automatically generated. The graphic background may be generated based on the content of the web resource, the content of the link annotations, and / or the type of annotation. The graphic card and / or link annotations may be stored to be provided along with data associated with the web resource. The link annotations may be generated to be provided for display in the search results interface in response to a specific content item being identified as a search result.
[0112] Figure 8A flowchart is depicted for an example method performed according to an example embodiment of this disclosure. Although Figure 8 The steps performed in a particular order are depicted for illustrative and discussion purposes, but the method of this disclosure is not limited to the particular illustrated order or arrangement. The steps of method 800 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of this disclosure.
[0113] At point 802, the computing system can obtain the first search query immediately and determine that the web resource responds to the first search query. The first search query may include text queries, image queries, audio queries, embedded queries, and / or multimodal queries. The web resource may be identified via a search engine that can perform keyword searches, embedded-based searches, and / or other search techniques. The topic, question, and / or intent of the web resource responding to the first search query can be determined.
[0114] At point 804, the computational system can obtain content data and process it using a generative model to generate predictive tips. The content data can be associated with a web resource. Tips can include predicted text strings associated with commenting on the web resource. The content data can describe the web resource content as a whole, media segments, content summaries, content tags, metadata, and / or previously provided link annotations associated with the web resource. The generative model can process the content data to determine topics, viewpoints, intent, titles, structure, target audience, content type, and / or other content details. Predictive tips can then be generated based on this determination.
[0115] At point 806, the computing system may provide predictive prompts for display within an input prompt interface and obtain comment input data from the user computing system via the input prompt interface. The input prompt interface may include input fields. In some implementations, the input prompt interface may include multiple user interface elements for drafting content (e.g., annotations). The comment input data may include user-generated content. In some implementations, the comment input data may include multimodal data. Multimodal data may include text data and image data.
[0116] At point 808, the computing system can store user-generated content. User-generated content can be indexed along with links to web resources. Alternatively and / or additionally, web resources can be indexed along with user-generated content. User-generated content can be stored along with other user comments associated with a specific web resource and / or a specific user.
[0117] At point 810, the computing system can obtain a second search query at a second time and determine that the web resource responds to the second search query. The second time may differ from the first time. In some implementations, the first and second search queries may differ. The second search query may include text queries, image queries, audio queries, embedded queries, and / or multimodal queries. The web resource may be identified via a search engine that can perform keyword searches, embedded-based searches, and / or other search techniques. The topic, question, and / or intent of the web resource responding to the second search query can be determined.
[0118] At point 812, the computing system can provide user-generated content in a search results interface that includes data describing web resources. User-generated content can be provided along with links, titles, and text snippets from the web resources.
[0119] Figures 9A to 9G Illustrations depict an example graphical interface according to an exemplary embodiment of this disclosure. The systems and methods disclosed herein can be used to generate graphical interfaces for user-generated content, which may include link annotations and / or self-contained content.
[0120] Figure 9A Two example graphic cards are depicted. The depicted graphic cards may include a user profile identifier 902 (e.g., user profile image and name), a body of card 904 (e.g., a stylized textual link annotation overlaid on the graphic background), widget interface elements 906 (e.g., selectable user interface elements for redirecting to web resources and / or additional content), and / or interactive information 908 (e.g., likes, comments, and / or saves for the graphic card). The body of card 904 may be configured by a specific user and / or may be automatically generated based on link annotations, web resources, and / or user preferences. Widget interface elements 906 may include links to web pages, links to image galleries, links to videos, selectable elements for opening pop-ups, additional annotations, and / or other data.
[0121] Figure 9B This describes example multi-page user-generated content and example video user-generated content. Multi-page user-generated content can include multiple graphics cards that can be looped through to display the user-generated content. Video user-generated content can include a video overlaid with graphics and / or text.
[0122] The widget interface element 906 may include links to web resources, links to one or more other web resources, video elements that can be selected to provide video for display, media content display elements for providing media content (e.g., video, images, audio files, and / or other media) for display, comments on web resources, links to other annotations, structured content items (e.g., structured recipes and / or structured calculators), lists (e.g., ingredient lists), map location cards (e.g., maps associated with web resources and / or links to web applications), knowledge panels, and / or links to shopping interfaces.
[0123] Figure 9C Example interactions with the graphics card are depicted. For example, a user profile identifier 902 can be selected to peek at the user's profile 910. Additionally and / or alternatively, a video widget element 912 can be selected to expand a video for playback and / or to navigate to a video player interface. The graphics card can be selected to minimize the plugin 914. The knowledge panel widget element 916 can be selected to expand the knowledge panel to provide additional information for display. Interactive elements 918 can be selected to like, comment, save, and / or share user-generated content.
[0124] Figure 9D The prominence level of the widget interface element (e.g., a plug-in element) is depicted. Specifically, the ingredients and description of the depicted smoothie can be provided in a medium-prominence interface element 922 (e.g., in a detailed view state), a low-prominence interface element 924 (e.g., in a collapsed state), and / or a high-prominence graphics panel 926 (e.g., in an expanded state). In some implementations, the user can interact with the widget interface element to change between levels of detail and / or sizes.
[0125] Figure 9E The search results interface 930 depicts the annotated search results. Annotated search results can be provided in a separate tab, adjacent to other search results, and / or in a category-specific panel. Annotated search results 932 can be selected to navigate to an immersive viewer 936 that displays an expanded view of user-generated content.
[0126] Figure 9F Different graphics card displays and / or annotation interface displays are described. Graphics cards can be displayed in a vertically scrolling single-width format 940, an offset vertically scrolling double-width format 942, a horizontally scrolling rotating interface within other search result formats 944, and / or an aligned vertically scrolling double-width format 946. The format can be based on the theme, interface type, user preferences, and / or context.
[0127] Figure 9GDifferent customization options can be described. For example, a graphics card customization interface can be provided to generate graphics cards, which may include editing text 952, editing layout 954, editing images 956, and / or other customization options. Specifically, the interactive interface may include multiple options (and / or features) for content generation. Multiple interface features may include text, images, audio, video, templates, and / or other input options. Multiple interface features may include content suggestions, template suggestions, and / or one or more generative model interfaces for generative model-assisted generation (e.g., a large language model for rewriting text and / or actively generating text; an image generation model for generating novel images based on web resources, user input, and / or generated prompts; an audio generation model for generating narration, songs, and / or other audio; and / or a graphics card generation model for processing web resources, generated prompts, and / or user input to generate graphics cards that can be suggested to the user for link annotations and / or self-contained content). Multiple interface features may include customization options for customizing layouts, fonts, interface element sizes, images, text, transitions, hues, shadows, and / or other user-generated content features. Multiple interface features may include options for adding action user interface elements to the graphics card of user-generated content. Action user interface elements may include selectable options for performing one or more actions (e.g., API calls, navigation to different applications, search, generating content items using a generative model, etc.).
[0128] The systems and methods disclosed herein may include image suggestions and / or image generation for generating graphics cards. For example, the systems and methods may determine that images from a database (e.g., a server database, a local database, and / or a user image library) are associated with web resources, prompts, and / or link annotations. These images can then be offered as suggestions for use in the graphics card. Alternatively and / or additionally, the systems and methods may provide an image generation model (e.g., a text-to-image generation model) interface to generate images to be included in the graphics card. For example, an image generation model interface may be provided to a user who can provide prompts to the image generation model, and the image generation model can generate model-generated images that can then be used in the graphics card.
[0129] In some implementations, one or more machine learning models can be used to perform fact-checking on web resources and / or link annotations. These machine learning models may include one or more generative models that can use application programming interfaces (APIs) to make API calls to obtain information and / or interact with other applications.
[0130] In some implementations, generative models can be used to generate one or more model-generated link annotations, which can be indexed along with web resources to provide link annotation examples and / or semantically understood annotations. Alternatively and / or additionally, generative models can be used to rewrite and / or suggest link annotations and / or standalone content. An interactive user interface may include an interface for interacting with the generative model to generate content (e.g., text, images, and / or other data). The interactive user interface may include options for selecting tone, style, format, vocabulary, genre, and / or other attributes to tune the generative model to generate content with specific attributes. For example, the interactive user interface may be configured to generate suggestions for the generative model based on user input, link annotation suggestions, and / or web resources.
[0131] The search results interface and / or discovery interface can provide statistics on the volume of a particular search, the volume of web resource selections, and / or trends in link and / or search query interactions.
[0132] In some implementations, the system and method may include training and / or utilizing one or more contribution propensity models. The contribution propensity model may learn and / or determine the user credibility (e.g., a user's relevant experience, expertise, and / or trustworthiness) for a particular user and / or a particular set of users. Additionally and / or alternatively, the contribution propensity model may learn and / or determine the tendency to provide link annotations.
[0133] Contribution propensity models can be trained to detect the likelihood and credibility of contributions, the usefulness of annotations, and / or other attributes associated with users, web resources, and / or context. Contribution propensity models can be trained on labeled datasets, unlabeled datasets, and / or mixed datasets. In some implementations, contribution propensity models can be trained on interaction data for a contribution prediction task, on the output of a validation model for a credibility determination task, and / or on click-through rates for a usefulness determination task.
[0134] Figure 10 An illustration depicts an example card generation interface 1000 according to an exemplary embodiment of the present disclosure. Specifically, a user can select an option to generate a graphical card for a link annotation. Based on this selection, a template 1002 can be selected and provided for display. A specific template can be selected based on a web resource associated with the link annotation (e.g., based on the content of the web resource), based on user interaction history, based on query history, based on user profile data, and / or based on other data.
[0135] Card generation interface 1000 may include a drop-up menu 1004 associated with multiple suggestions for linking annotation generation, which may include thematic ideas. The user can pull up this menu to access an expanded view 1006 of the suggestions. Expanded view 1006 may include multiple selectable suggestions that can be chosen to generate text, images, and / or layouts to be inserted into the graphics card. For example, the suggestion "How to Water Your Monstera Like a Pro" can be selected. Content items associated with the selected suggestion can be inserted into the graphics card, which can then transition to editing interface 1008. Editing interface 1008 may include options for editing text, style, layout, font, color, and / or other edits.
[0136] Figure 11 An illustration depicts an example content item generation interface 1100 according to an exemplary embodiment of the present disclosure. Specifically, when modifying a graphics card template, one or more content items can be generated using the input drafting interface.
[0137] For example, a user can select an option to open the content item generation interface 1100. At 1102, the user can select one or more attributes from a drop-down menu. These attributes can be associated with the attributes of the requested content item being generated. These attributes can be associated with the tone and / or style of the content. At 1104, the user can generate and / or provide text input. The text input can be associated with a topic, intent, information, and / or other cueing details. A generative model can process one or more attributes and text input to generate model-generated content items. The model-generated content items can have one or more attributes and can be tailored to the topic, intent, information, and / or other cueing details of the text input.
[0138] At point 1106, the model-generated content item can be displayed below the text input, and can be provided along with multiple options. These options may include editing one or more attributes, editing the text input, reprocessing data, saving the model-generated content item, exiting the interface, inserting the model-generated content item into the graphics card, and / or other options. At point 1108, a modified graphics card can be provided for display, where the model-generated content item is inserted into the graphics card based on user selection. The user can then edit the layout, size, color, font, and / or orientation of the model-generated content item and / or other content on the graphics card.
[0139] Figure 12An illustration depicts an example image suggestion interface 1200 according to an exemplary embodiment of the present disclosure. Specifically, the image suggestion interface 1200 can acquire card data, scene data, and / or input data. The card data, scene data, and / or input data can then be processed to determine one or more images (and / or other media content items) to be offered as suggestions for insertion into a graphics card.
[0140] For example, at 1202, a graphics card is provided for display, which has options for inserting additional text, stickers, and / or images. The user can then select the option to add an image. At 1204, an image selection interface can be provided for display, which may include a default image, camera album images, and / or image suggestions based on the graphics card's text, the content of web resources associated with linked annotations, user history, and / or other data. For example, the association of images with the graphics card's text may be determined based on identifying multiple images from the user's image library that are associated with locations referenced in the graphics card's text (e.g., Mexico). At 1206, identified images can be provided for display for selection. The user can select a specific image from the identified images, which can be processed and inserted into the graphics card. At 1208, the selected image can be cropped and inserted into the graphics card for display.
[0141] Figure 13 A flowchart is depicted for an example method performed according to an example embodiment of this disclosure. Although Figure 13 For illustrative and discussion purposes, the steps performed in a particular order are described, but the method of this disclosure is not limited to the specifically described order or arrangement. The various steps of method 1300 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of this disclosure.
[0142] At 1302, the computing system can obtain card data. Card data can describe the content in the graphics card. Content can be associated with one or more topics. Graphics cards can be associated with link annotations. Link annotations can include user-generated content tagged to a specific web resource. Graphics cards can include a background, one or more images, one or more text strings, and / or one or more user interface elements. The background can include a single color, multiple colors, images, and / or other data. One or more user interface elements can include selectable widgets for providing additional information for display and / or for performing one or more actions. Content can include text data, image data, video data, latently encoded data, multimodal data, and / or other data.
[0143] At point 1304, the computing system can process the card data to determine one or more entity tags associated with the content. One or more entity tags can be associated with one or more topics. The card data can be processed using one or more machine learning models (e.g., generative models, classification models, and / or other models) to generate entity tags. Entity tags can be associated with one or more objects, one or more companies, one or more locations, one or more individuals, one or more structures, and / or other entities.
[0144] At 1306, the computing system can access a media content item database to obtain one or more media content items. The one or more media content items can be obtained based on determining that one or more media content items are associated with one or more entity tags associated with the content. The media content item database may include a user-specific database. In some implementations, the user-specific database may be associated with a specific user. The specific user may have generated at least a portion of the content. The user-specific database may include an image library associated with the specific user. The image library may be stored on a server computing system associated with a specific content item storage platform. Alternatively and / or additionally, the user-specific database may include a local storage database on the user's computing device. The media content item database may include multiple media content items. The multiple content items may have been preprocessed to generate multiple corresponding metadata datasets. Determining that one or more media content items are associated with one or more entity tags associated with the content may include: determining whether one or more media content items include features associated with entity tags. Features may be determined based on metadata, image processing, and / or other techniques. One or more media content items may include one or more images, one or more videos, one or more animations, one or more audio files, and / or one or more other content items.
[0145] At 1308, the computing system may provide one or more media content items for display. One or more media content items may be provided in an interactive user interface. One or more media content items may be selected to be inserted into a graphics card. The interactive user interface may provide multiple media content items for display, which may include media content items associated with the user, web media content items, and / or other media content items.
[0146] In some implementations, the computing system can obtain input selections associated with one or more media content items and generate an enhanced graphics card. The enhanced graphics card may include at least a portion of the content of the graphics card and at least a portion of one or more media content items. The computing system can then provide the enhanced graphics card for display.
[0147] Alternatively and / or additionally, the computing system may obtain enhanced input. The enhanced input may be associated with a request to enhance the enhanced graphics card. The computing system may generate an updated graphics card based on the enhanced input. The updated graphics card may include an enhanced graphics card with one or more enhancements. The computing system can then provide the updated graphics card for display. One or more enhancements may include at least one of the following: layout changes to the enhanced graphics card, cropping changes to one or more media content items, resizing changes to one or more content items, color changes, or stencil changes.
[0148] Figure 14 A flowchart is depicted for an example method performed according to an example embodiment of this disclosure. Although Figure 14 The steps performed in a particular order are depicted for illustrative and discussion purposes, but the method of this disclosure is not limited to the order or arrangement specifically shown. The various steps of method 1400 may be omitted, rearranged, combined and / or adapted in various ways without departing from the scope of this disclosure.
[0149] At point 1402, the computing system may provide an input drafting interface for display. The input drafting interface may include a graphical user interface (GUI) comprising multiple attribute options and text input boxes. The multiple attribute options may be associated with multiple candidate attributes for generating content items. The multiple candidate attributes may include tone, style, length, content type, and / or other details. The input drafting interface may include a preview window for viewing the current state of a graphics card. Graphic cards may be associated with link annotations. In some implementations, graphics cards may include card templates that may have been modified based on one or more user inputs. For example, the user may have added images, text, audio, video, widgets, and / or other data.
[0150] At point 1404, the computing system can obtain a selection of a specific attribute option from multiple attribute options via an input drafting interface. The specific attribute option can be associated with a specific candidate attribute. In some implementations, the multiple candidate attributes can include multiple different styles. Multiple different styles can be associated with at least one of multiple different artistic styles or multiple different writing styles. Alternatively and / or additionally, the multiple candidate attributes can include multiple different tones. Multiple different tones can be associated with at least one of multiple different emotions and / or multiple different rhythm types. The specific candidate attribute can include the tone and / or style requested for generating content items. This selection can be based on the selection of a specific attribute option from a drop-down menu that provides multiple attribute options for display.
[0151] At point 1406, the computing system can obtain text input via a text input box on the input drafting interface. The text input can be associated with a prompt diagram generated from the content item. In some implementations, the text input can be automatically filled based on the graphics card content, the user context, and / or prompts / suggestions.
[0152] At point 1408, the computational system can use a generative model to process specific attribute options and text input to generate content items generated by the model. The generated content items may include specific candidate attributes. The generated content items may be associated with a given schematic diagram. The generated content items may include text data, image data, audio data, multimodal data, and / or other data. In some implementations, the generative model may be obtained from a generative model database based on the selection of specific attribute options. For example, the generative model database may store multiple different generative models associated with multiple candidate attributes. Each of the multiple different generative models can be configured, trained, and / or tuned to generate content items associated with the corresponding candidate attribute. Alternatively and / or additionally, the generative model may be a general generative model trained for multiple content generation tasks. Additionally and / or alternatively, specific attribute soft cues may be obtained based on the selection of specific attribute options. Specific attribute soft cues may include a learned set of parameters. The learned parameter set can be processed by the generative model to generate model-generated content items.
[0153] At point 1410, the computing system can provide model-generated content items for display via an input drafting interface. Providing model-generated content items for display may include providing options for inserting the model-generated content items into the graphics card. The input drafting interface may include multiple post-processing editing options, which may include options for changing size, color, font, cropping, resolution, saturation, shading, and / or other details.
[0154] In some implementations, the computing system can obtain input selections via an input drafting interface and generate an enhanced graphics card based on those selections. The enhanced graphics card can include graphics cards enhanced to include content items generated from the model. The computing system can then provide the enhanced graphics card for display.
[0155] Figure 15A A block diagram of an example computing system 100, illustrating an exemplary embodiment of the present disclosure, is depicted, with execution link notes indicating its purpose. System 100 includes a user computing system 102, a server computing system 130, and / or a third computing system 150 communicatively coupled via a network 180.
[0156] User computing system 102 may include any type of computing device, such as, for example, a personal computing device (e.g., a laptop computer or desktop computer), a mobile computing device (e.g., a smartphone or tablet computer), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0157] User computing system 102 includes one or more processors 112 and memory 114. The one or more processors 112 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. Memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 114 may store data 116 and instructions 118 executed by processor 112 to cause user computing system 102 to perform operations.
[0158] In some implementations, the user computing system 102 may store or include one or more machine learning models 120. For example, machine learning model 120 may or may otherwise include various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including nonlinear and / or linear models. Neural networks may include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks.
[0159] In some implementations, one or more machine learning models 120 may be received from server computing system 130 via network 180, stored in user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, user computing system 102 may implement multiple parallel instances of a single machine learning model 120 (e.g., performing parallel machine learning model processing across multiple instances of input data and / or detected features).
[0160] More specifically, one or more machine learning models 120 may include one or more detection models, one or more classification models, one or more segmentation models, one or more augmentation models, one or more generative models, one or more natural language processing models, one or more optical character recognition models, and / or one or more other machine learning models. One or more machine learning models 120 may include one or more Transformer models. One or more machine learning models 120 may include one or more neural radiation field models, one or more diffusion models, and / or one or more autoregressive language models.
[0161] One or more machine learning models 120 can be used to detect one or more object features. The detected object features can be classified and / or embedded. The classification and / or embedding can then be used to perform a search to determine one or more search results. Alternatively and / or additionally, one or more detected features can be used to determine which indicators (e.g., user interface elements indicating detected features) should be provided to indicate that features have been detected. The user can then select the indicator to perform feature classification, embedding, and / or search. In some implementations, classification, embedding, and / or search can be performed before the indicator is selected.
[0162] In some implementations, one or more machine learning models 120 may process image data, text data, audio data, and / or latently encoded data to generate output data, which may include image data, text data, audio data, and / or latently encoded data. One or more machine learning models 120 may perform optical character recognition, natural language processing, image classification, object classification, text classification, audio classification, context determination, action prediction, image correction, image enhancement, text enhancement, sentiment analysis, object detection, error detection, inpainting, video stabilization, audio correction, audio enhancement, and / or data segmentation (e.g., mask-based segmentation).
[0163] Additionally or alternatively, one or more machine learning models 140 may be included in or otherwise stored and implemented by server computing system 130, which communicates with user computing system 102 according to a client-server relationship. For example, machine learning model 140 may be implemented by server computing system 130 as part of a web service (e.g., viewfinder service, visual search service, image processing service, ambient computing service, and / or overlay application service). Thus, one or more models 120 may be stored and implemented at user computing system 102, and / or one or more models 140 may be stored and implemented at server computing system 130.
[0164] User computing system 102 may also include one or more user input components 122 for receiving user input. For example, user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). Touch-sensitive components can be used to implement a virtual keyboard. Other example user input components include microphones, conventional keyboards, or other devices that a user can use to provide user input.
[0165] In some implementations, the user computing system may store and / or provide one or more user interfaces 124, which may be associated with one or more applications. The one or more user interfaces 124 may be configured to receive input and / or provide data for display (e.g., image data, text data, audio data, one or more user interface elements, augmented reality experiences, virtual reality experiences, and / or other data for display). The user interface 124 may be associated with one or more other computing systems (e.g., server computing system 130 and / or third-party computing system 150). The user interface 124 may include a viewfinder interface, a search interface, a generative model interface, a social media interface, and / or a media content gallery interface.
[0166] User computing system 102 may include one or more sensors 126 and / or receive data from said one or more sensors. The one or more sensors 126 may be housed in a housing assembly that houses one or more processors 112, memory 114, and / or one or more hardware components that may store one or more software packages and / or cause execution of said one or more software packages. The one or more sensors 126 may include one or more image sensors (e.g., cameras), one or more lidar sensors, one or more audio sensors (e.g., microphones), one or more inertial sensors (e.g., inertial measurement units), one or more biosensors (e.g., heart rate sensors, pulse sensors, retinal sensors, and / or fingerprint sensors), one or more infrared sensors, one or more position sensors (e.g., GPS), one or more touch sensors (e.g., conductive touch sensors and / or mechanical touch sensors), and / or one or more other sensors. The one or more sensors may be used to obtain data associated with the user environment (e.g., images of the user environment, records of the environment, and / or the user's location).
[0167] User computing system 102 may include user computing device 104 and / or portions thereof. User computing device 104 may include mobile computing devices (e.g., smartphones or tablets), desktop computers, laptop computers, smart wearable devices, and / or smart home appliances. Additionally and / or alternatively, the user computing system may acquire data from one or more user computing devices 104 and / or generate data using those devices. For example, a smartphone camera may be used to capture image data describing the environment, and / or an overlay application on user computing device 104 may be used to track and / or process data provided to the user. Similarly, one or more sensors associated with a smart wearable device may be used to acquire data about the user and / or about the user's environment (e.g., a camera housed in the user's smart glasses may be used to acquire image data). Additionally and / or alternatively, data may be acquired and uploaded from other user devices that may be specifically used for data acquisition or generation.
[0168] Server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 134 can store data 136 and instructions 138 executed by processor 132 to cause server computing system 130 to perform operations.
[0169] In some implementations, the server computing system 130 includes one or more server computing devices or is otherwise implemented by said one or more server computing devices. In instances where the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0170] As described above, the server computing system 130 may store or otherwise include one or more machine learning models 140. For example, model 140 may be, or may otherwise include, various machine learning models. Example machine learning models include neural networks or other multi-layered nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. (Reference) Figure 15B Discuss example model 140.
[0171] Additionally and / or alternatively, server computing system 130 may include search engine 142 and / or be communicatively connected to search engine, which may be used to crawl one or more databases (and / or resources). Search engine 142 may process data from user computing system 102, server computing system 130, and / or third-party computing system 150 to determine one or more search results associated with input data. Search engine 142 may perform item-based search, tag-based search, Boolean-based search, image search, embedding-based search (e.g., nearest neighbor search), multimodal search, and / or one or more other search techniques.
[0172] Server computing system 130 may store and / or provide one or more user interfaces 144 for obtaining input data and / or providing output data to one or more users. The one or more user interfaces 144 may include one or more user interface elements, which may include input fields, navigation tools, content tiles, optional tiles, widgets, data display carousels, dynamic animations, information pop-ups, image enhancement, text-to-speech, speech-to-text, augmented reality, virtual reality, feedback loops, and / or other interface elements.
[0173] User computing system 102 and / or server computing system 130 may train models 120 and / or 140 via interaction with a third-party computing system 150 communicatively coupled via network 180. The third-party computing system 150 may be separate from or part of the server computing system 130. Alternatively and / or additionally, the third-party computing system 150 may be associated with one or more web resources, one or more web platforms, one or more other users, and / or one or more scenarios.
[0174] The third-party computing system 150 may include one or more processors 152 and memory 154. The one or more processors 152 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158 executed by the processor 152 to cause the third-party computing system 150 to perform operations. In some implementations, the third-party computing system 150 includes one or more server computing devices or is otherwise implemented by such one or more server computing devices.
[0175] Network 180 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over Network 180 can be conducted using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL) over any type of wired and / or wireless connection.
[0176] The machine learning models described in this specification can be used for a variety of tasks, applications, and / or use cases.
[0177] In some implementations, the input to the machine learning model of this disclosure may be image data. The machine learning model can process the image data to generate output. As an example, the machine learning model can process image data to generate image recognition output (e.g., image data identification, latent embedding of image data, encoded representation of image data, hashing of image data, etc.). As another example, the machine learning model can process image data to generate image segmentation output. As another example, the machine learning model can process image data to generate image classification output. As another example, the machine learning model can process image data to generate image data modification output (e.g., image data alteration, etc.). As another example, the machine learning model can process image data to generate encoded image data output (e.g., encoded and / or compressed representation of image data, etc.). As another example, the machine learning model can process image data to generate magnified image data output. As another example, the machine learning model can process image data to generate prediction output.
[0178] In some implementations, the input to the machine learning model of this disclosure may be text or natural language data. The machine learning model may process the text or natural language data to generate output. For example, the machine learning model may process natural language data to generate a language-encoded output. For another example, the machine learning model may process text or natural language data to generate a latent text embedding output. For another example, the machine learning model may process text or natural language data to generate a translation output. For another example, the machine learning model may process text or natural language data to generate a classification output. For another example, the machine learning model may process text or natural language data to generate a text segmentation output. For another example, the machine learning model may process text or natural language data to generate a semantic intent output. For another example, the machine learning model may process text or natural language data to generate an amplified text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language). For yet another example, the machine learning model may process text or natural language data to generate a predictive output.
[0179] In some implementations, the input to the machine learning model of this disclosure may be speech data. The machine learning model may process the speech data to generate output. As an example, the machine learning model may process speech data to generate speech recognition output. As another example, the machine learning model may process speech data to generate speech translation output. As another example, the machine learning model may process speech data to generate latent embedding output. As another example, the machine learning model may process speech data to generate encoded speech output (e.g., encoded and / or compressed representations of speech data, etc.). As another example, the machine learning model may process speech data to generate amplified speech output (e.g., speech data of higher quality than the input speech data, etc.). As another example, the machine learning model may process speech data to generate text representation output (e.g., text representations of the input speech data, etc.). As another example, the machine learning model may process speech data to generate predictive output.
[0180] In some implementations, the input to the machine learning model of this disclosure may be sensor data. The machine learning model may process the sensor data to generate output. As an example, the machine learning model may process sensor data to generate identification output. As another example, the machine learning model may process sensor data to generate prediction output. As another example, the machine learning model may process sensor data to generate classification output. As another example, the machine learning model may process sensor data to generate segmentation output. As another example, the machine learning model may process sensor data to generate segmentation output. As another example, the machine learning model may process sensor data to generate visualization output. As another example, the machine learning model may process sensor data to generate diagnostic output. As another example, the machine learning model may process sensor data to generate detection output.
[0181] In some cases, the input includes visual data, and the task is a computer vision task. In other cases, the input includes pixel data for one or more images, and the task is an image processing task. For example, an image processing task could be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the probability that one or more images depict an object belonging to that object class. An image processing task could be object detection, where the image processing output identifies one or more regions in one or more images, and for each region, identifies the probability that the region depicts an object of interest. As another example, an image processing task could be image segmentation, where the image processing output defines a corresponding probability for each class in a predetermined set of categories for each pixel in one or more images. For example, the category set could be foreground and background. As another example, the category set could be object classes. As another example, an image processing task could be depth estimation, where the image processing output defines a corresponding depth value for each pixel in one or more images. As another example, an image processing task could be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene depicted at that pixel between the images in the network input for each pixel in one of the input images.
[0182] A user computing system may include multiple applications (e.g., applications 1 to N). Each application may include its own corresponding machine learning library and machine learning model. For example, each application may include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.
[0183] Each application can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a field manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is application-specific.
[0184] User computing system 102 may include multiple applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. In some implementations, each application may use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).
[0185] The central intelligence layer may include multiple machine learning models. For example, a corresponding machine learning model (e.g., a model) may be provided for each application, and this machine learning model is managed by the central intelligence layer. In other implementations, two or more applications may share a single machine learning model. For example, in some implementations, the central intelligence layer may provide a single model (e.g., a single model) for all applications. In some implementations, the central intelligence layer is included within the operating system of the computing system 100 or otherwise implemented by the operating system.
[0186] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data repository for the computing system 100. The central device data layer can communicate with many other components of the computing device, such as one or more sensors, a field manager, a device status component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0187] Figure 15BA block diagram of an example computing system 50, illustrating an exemplary embodiment of the present disclosure, is depicted. Specifically, the example computing system 50 may include one or more computing devices 52, which may be used to acquire and / or generate one or more datasets. These datasets may be processed by a sensor processing system 60 and / or an output determination system 80 to provide feedback to a user, who can provide information about features in one or more acquired datasets. The one or more datasets may include image data, text data, audio data, multimodal data, latently encoded data, etc. The one or more datasets may be acquired via one or more sensors associated with the one or more computing devices 52 (e.g., one or more sensors within the computing devices 52). Additionally and / or alternatively, the one or more datasets may be stored data and / or retrieved data (e.g., data retrieved from web resources). For example, a user may interact with images, text, and / or other content items. Interaction with the content items may then be used to generate one or more determinations.
[0188] One or more computing devices 52 may acquire and / or generate one or more datasets based on image capture, sensor tracking, data storage retrieval, content download (e.g., downloading images or other content items from web resources via the Internet), and / or via one or more other technologies. Sensor processing system 60 may be used to process one or more datasets. Sensor processing system 60 may use one or more machine learning models, one or more search engines, and / or one or more other processing technologies to perform one or more processing techniques. One or more processing techniques may be performed in any combination and / or individually. One or more processing techniques may be performed serially and / or in parallel. Specifically, context determination block 62 may be used to process one or more datasets, which can determine the context associated with one or more content items. Context determination block 62 may identify and / or process metadata, user profile data (e.g., preferences, user search history, user browsing history, user purchase history, and / or user input data), previous interaction data, global trend data, location data, time data, and / or other data to determine the specific context associated with a user. A context can be associated with an event, an identified trend, a specific action, a specific type of data, a specific environment, and / or another context associated with a user and / or data retrieved or obtained.
[0189] The sensor processing system 60 may include an image preprocessing block 64. The image preprocessing block 64 may be used to adjust one or more values of the acquired and / or received image to prepare the image for processing by one or more machine learning models and / or one or more search engines 74. The image preprocessing block 64 may resize the image, adjust saturation values, adjust resolution, strip and / or add metadata, and / or perform one or more other operations.
[0190] In some implementations, the sensor processing system 60 may include one or more machine learning models, which may include a detection model 66, a segmentation model 68, a classification model 70, an embedding model 72, and / or one or more other machine learning models. For example, the sensor processing system 60 may include one or more detection models 66 that can be used to detect specific features in a processed dataset. Specifically, one or more detection models 66 may be used to process one or more images to generate one or more bounding boxes associated with the detected features in the one or more images.
[0191] Additionally and / or alternatively, one or more segmentation models 68 may be used to segment one or more portions of a dataset from one or more datasets. For example, one or more segmentation models 68 may utilize one or more segmentation masks (e.g., manually generated and / or generated based on one or more bounding boxes) to segment a portion of an image, a portion of an audio file, and / or a portion of text. Segmentation may include isolating one or more detected objects and / or removing one or more detected objects from an image.
[0192] One or more classification models 70 can be used to process image data, text data, audio data, latently encoded data, multimodal data, and / or other data to generate one or more classifications. The one or more classification models 70 may include one or more image classification models, one or more object classification models, one or more text classification models, one or more audio classification models, and / or one or more other classification models. The one or more classification models 70 can process data to determine one or more classifications.
[0193] In some implementations, one or more embedding models 72 can be used to process data to generate one or more embeddings. For example, one or more embedding models 72 can be used to process one or more images to generate one or more image embeddings in an embedding space. One or more image embeddings can be associated with one or more image features of one or more images. In some implementations, one or more embedding models 72 can be configured to process multimodal data to generate multimodal embeddings. One or more embeddings can be used for classification, search, and / or learning the embedding space distribution.
[0194] The sensor processing system 60 may include one or more search engines 74, which can be used to perform one or more searches. The one or more search engines 74 may crawl one or more databases (e.g., one or more local databases, one or more global databases, one or more private databases, one or more public databases, one or more specialized databases, and / or one or more general-purpose databases) to determine one or more search results. The one or more search engines 74 may perform feature matching, text-based search, embedding-based search (e.g., k-nearest neighbor search), metadata-based search, multimodal search, web resource search, image search, text search, and / or application search.
[0195] Additionally and / or alternatively, the sensor processing system 60 may include one or more multimodal processing blocks 76, which may be used to assist in processing multimodal data. The one or more multimodal processing blocks 76 may include generating multimodal queries and / or multimodal embeddings for processing by one or more machine learning models and / or one or more search engines 74.
[0196] The output determination system 80 can then be used to process the output of the sensor processing system 60 to determine one or more outputs to be provided to the user. The output determination system 80 may include heuristic-based determination, machine learning model-based determination, user-selection-based determination, and / or context-based determination.
[0197] Output determination system 80 can determine how and / or where to provide one or more search results in search results interface 82. Additionally and / or alternatively, output determination system 80 can determine how and / or where to provide one or more machine learning model outputs in machine learning model output interface 84. In some implementations, one or more search results and / or one or more machine learning model outputs can be provided for display via one or more user interface elements. One or more user interface elements can be overlaid on the displayed data. For example, one or more detection indicators can be overlaid on detected objects in the viewfinder. One or more user interface elements can be selected to perform one or more additional search and / or one or more additional machine learning model processes. In some implementations, user interface elements can be provided as application-specific user interface elements and / or can be uniformly provided across different applications. One or more user interface elements can include pop-up displays, interface overlays, interface tiles and / or slices, carousels, audio feedback, animations, interactive widgets, and / or other user interface elements.
[0198] Additionally and / or alternatively, data associated with the output of the sensor processing system 60 can be used to generate and / or provide augmented reality and / or virtual reality experiences 86. For example, one or more acquired datasets can be processed to generate one or more augmented reality rendering assets and / or one or more virtual reality rendering assets, which can then be used to provide augmented reality and / or virtual reality experiences 86 to a user. Augmented reality experiences can render information associated with the environment into the appropriate environment. Alternatively and / or additionally, objects associated with the processed dataset can be rendered into the user environment and / or virtual environment. Rendering dataset generation can include training one or more neural radiation field models to learn three-dimensional representations of one or more objects.
[0199] In some implementations, one or more action prompts 88 can be determined based on the output of the sensor processing system 60. For example, a search prompt, purchase prompt, generate prompt, reservation prompt, call prompt, redirection prompt, and / or one or more other prompts can be determined to be associated with the output of the sensor processing system 60. The one or more action prompts 88 can then be provided to the user via one or more optional user interface elements. In response to the selection of one or more optional user interface elements, a corresponding action of the corresponding action prompt can be performed (e.g., a search can be performed, a purchase application programming interface can be used, and / or another application can be opened).
[0200] In some implementations, one or more generative models 90 may be used to process one or more datasets and / or the output of the sensor processing system 60 to generate model-generated content items, which may then be provided to a user. This generation may be based on user selection and / or may be performed automatically (e.g., automatically based on one or more conditions, which may be associated with a threshold amount of unrepresented search results).
[0201] One or more generative models 90 may include language models (e.g., large language models and / or visual language models), image generation models (e.g., text-to-image generation models and / or image augmentation models), audio generation models, video generation models, graphics generation models, and / or other data generation models (e.g., other content generation models). One or more generative models 90 may include one or more Transformer models, one or more convolutional neural networks, one or more recurrent neural networks, one or more feedforward neural networks, one or more generative adversarial networks, one or more self-attention models, one or more embedding models, one or more encoders, one or more decoders, and / or one or more other models. In some implementations, one or more generative models 90 may include one or more autoregressive models (e.g., machine learning models trained to generate predicted values based on previously generated behavioral data) and / or one or more diffusion models (e.g., machine learning models trained to generate predicted data based on generating and processing distributional data associated with the input data).
[0202] One or more generative models 90 can be trained to process input data and generate model-generated content items, which may include multiple predicted words, pixels, signals, and / or other data. The model-generated content items may include novel content items that differ from any existing work. One or more generative models 90 may utilize learned representations, sequences, and / or probability distributions to generate content items, which may include phrases, storylines, settings, objects, characters, beats, lyrics, and / or other aspects not included in existing content items.
[0203] One or more generative models 90 may include visual language models.
[0204] Visual language models can be trained, tuned, and / or configured to process image and / or text data to generate natural language output. Visual language models can leverage pre-trained large language models (e.g., large autoregressive language models) and one or more encoders (e.g., one or more image encoders and / or one or more text encoders) to provide detailed natural language output that mimics human-written natural language.
[0205] Visual language models can be used for zero-shot image classification, few-shot image classification, image captioning, multimodal query extraction, multimodal question answering, and / or can be tuned and / or trained for multiple different tasks. Visual language models can perform visual question answering, image captioning generation, feature detection (e.g., content monitoring (e.g., for inappropriate content)), object detection, scene recognition, and / or other tasks.
[0206] Visual language models can leverage pre-trained language models and then tune them for multimodality. Training and / or tuning of visual language models can include image-text matching, masked language modeling, multimodal fusion with cross-attention, contrastive learning, prefix language model training, and / or other training techniques. For example, a visual language model can be trained to process images to generate predicted text similar to ground-value text data (e.g., ground-value captions for an image). In some implementations, a visual language model can be trained to replace masked lexical units of a natural language template with text lexical units describing features depicted in an input image. Alternatively and / or additionally, training, tuning, and / or model inference can include multi-layered connections of visual and text embedding features. In some implementations, a visual language model can be trained and / or tuned by jointly learning image embedding and text embedding generation, which can include training and / or tuning a system to map embeddings to a joint feature embedding space that maps text features and image features to a shared embedding space. Joint training can include image-text pair parallel embeddings and / or can include triple training. In some implementations, it can be used as a prefix for language models and / or for processing images.
[0207] The output determination system 80 can use the data augmentation block 92 to process one or more datasets and / or the output of the sensor processing system 60 to generate augmented data. For example, the data augmentation block 92 can be used to process one or more images to generate one or more augmented images. Data augmentation can include data correction, data cropping, removal of one or more features, addition of one or more features, resolution adjustment, lighting adjustment, saturation adjustment, and / or other enhancements.
[0208] In some implementations, one or more datasets and / or the output of the sensor processing system 60 can be stored based on the determination of the data storage block 94.
[0209] The output of the output determination system 80 can then be provided to the user via one or more output components of the user computing device 52. For example, one or more user interface elements associated with one or more outputs can be provided for display via the visual display of the user computing device 52.
[0210] This process can be performed iteratively and / or continuously. One or more user inputs to the provided user interface elements can modulate and / or influence the continuous processing loop.
[0211] This paper discusses technologies related to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from these systems. The inherent flexibility of computer-based systems allows for a wide range of possible configurations, combinations, and divisions of tasks and functions among and within components. For example, the processes discussed in this paper can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0212] While the subject matter has been described in detail with respect to various specific example embodiments, each example is provided by way of explanation and not as a limitation of this disclosure. Modifications, alterations, and equivalents of these embodiments will be readily apparent to those skilled in the art upon understanding the foregoing. Therefore, this disclosure does not exclude such modifications, alterations, and / or additions to the subject matter that will be readily understood by those of ordinary skill in the art. For example, features shown or described as part of one embodiment may be used with another embodiment to produce yet another embodiment. Therefore, this disclosure is intended to cover such modifications, alterations, and equivalents.
Claims
1. A computing system for review prompt generation and input retrieval, the system comprising: one or more processors; and one or more non-transitory computer-readable media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations comprising: obtaining content data, wherein the content data is associated with a web resource, wherein the web resource comprises one or more content items; processing the content data with a generation model to generate a predicted prompt, wherein the prompt comprises a predicted string of text associated with commenting on the web resource; providing the predicted prompt for display with an input prompt interface, wherein the input prompt interface is configured to receive input; obtaining comment input data from a user computing system via the input prompt interface, wherein the comment input data comprises a user-generated comment on the web resource; obtaining user data, wherein the user data is associated with a particular user, wherein the user computing system is associated with the particular user; generating a graph card based on the user data, the content data, and the comment input data, wherein the graph card comprises a user profile identifier of the particular user and data associated with the comment input data, wherein the graph card comprises a graphical background generated based on the comment input data with a graphical image generation model; and storing data associated with the comment input data with data associated with the web resource and the graph card, wherein the data associated with the comment input data is stored in a searchable database to be provided for display in response to the web resource being provided as a search result.
2. The system of claim 1, wherein the user data comprises user search history data, and wherein the generation model generates the predicted prompt based on the particular user previously searching for information associated with a topic of the web resource.
3. The system of claim 1, wherein the user data comprises user browser history data, and wherein the generation model generates the predicted prompt based on the particular user previously viewing other web resources comprising information associated with a topic of the web resource.
4. The system of claim 1, wherein the operations further comprise: obtaining a search query; determining that the web resource is associated with the search query; providing a particular search result for display, wherein the particular search result comprises a link to the web resource, a title of the web resource, and data associated with the comment input data.
5. The system of claim 1, wherein storing the data associated with the comment input data with the data associated with the web resource and the graph card comprises: generating a web resource annotation; and storing the web resource annotation with a plurality of other web resource annotations associated with the web resource. 6. The system of claim 5, wherein the operations further comprise: providing the web resource annotation and the plurality of other web resource annotations in an annotation interface, the annotation interface providing the web resource annotation and the plurality of other web resource annotations in a plurality of graphical cards.
7. The system of claim 1, wherein the operations further comprise: obtaining a selection of a property of a request to augment the user-generated comment; processing the user-generated comment and the property of the request with the generation model to generate a model-generated content item; and augmenting the graphical card to include the model-generated content item.
8. The system of claim 1, wherein the operations further comprise: processing the graphical card to determine one or more entity tags associated with a topic of the graphical card; accessing a media content item database to obtain one or more media content items based on the one or more entity tags; and providing the one or more media content items for display.
9. The system of claim 1, wherein the operations further comprise: providing a graphical card customization interface for display, wherein the graphical card customization interface includes a plurality of options for editing the graphical card.
10. The system of claim 1, wherein the generation model comprises a self- recurrent language model, and wherein the generation model is prompted to generate a question that describes a request for information about the web resource.
11. A computer-implemented method for linking annotation prompts, the method comprising: obtaining, by a computing system comprising one or more processors, context data, wherein the context data is associated with a particular content display instance, wherein the particular content display instance comprises a particular user viewing a particular content item of a particular web resource; determining, by the computing system, an input request action based on the context data, wherein the input request action comprises providing an input entry interface to a user to obtain user input; processing, by the computing system, the context data with a generation language model to generate a predicted prompt, wherein the predicted prompt comprises a natural language request for information generated based on the context data; providing, by the computing system, the predicted prompt in the input entry interface; obtaining, by the computing system, user-generated content via the input entry interface; obtaining, by the computing system, user data, wherein the user data is associated with a particular user, wherein the user-generated content is associated with the particular user; generating, by the computing system, a graphical card based on the user data and the user-generated content, wherein the graphical card comprises a user profile identifier of the particular user and data associated with the user-generated content, wherein the graphical card comprises a graphical background generated based on the user-generated content with an image generation model; and generating, by the computing system, a linking annotation included in the graphical card based on the user-generated content, wherein the linking annotation is generated to be provided for display in a search results interface in response to the particular content item being determined as a search result.
12. The method of claim 11, wherein the contextual data is associated with a type of content being provided for display.
13. The method of claim 11, wherein the contextual data is associated with the particular user associated with the particular content display instance, wherein the contextual data comprises search history data.
14. The method of claim 11, wherein content being provided for display is associated with a particular web resource, wherein the contextual data is associated with interaction data for links to the particular web resource on a plurality of social networking platforms, and wherein the input request action is determined based on the interaction data.
15. The method of claim 11, wherein the contextual data comprises user data and content data, and wherein the input request action is determined based on a topic associated with content being provided for display is one of a plurality of topics that the particular user is determined to know based on the user data.
16. The method of claim 11, wherein the contextual data comprises a prior annotation generated by the particular user, and wherein the predictive prompt comprises a structure based on a prior structure of the prior annotation.
17. One or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations comprising: obtaining a first search query at a first time; determining a web resource responsive to the first search query, wherein the web resource comprises one or more content items; obtaining content data, wherein the content data is associated with the web resource; processing the content data with a generative model to generate a predictive prompt, wherein the prompt comprises a predicted string of text associated with commenting on the web resource; providing the predictive prompt for display within an input prompt interface, wherein the input prompt interface comprises an input entry box; obtaining comment input data from a user computing system via the input prompt interface, wherein the comment input data comprises user-generated content; obtaining user data, wherein the user data is associated with a particular user, wherein the user computing system is associated with the particular user; generating a graphic card based on the user data, the content data, and the comment input data, wherein the graphic card comprises a user profile identifier of the particular user and data associated with the comment input data, wherein the graphic card comprises a graphic background generated based on the comment input data with an image generative model; storing the graphic card; obtaining a second search query at a second time, wherein the second time is different from the first time; determining the web resource responsive to the second search query; and providing the graphic card in a search results interface having data describing the web resource.
18. The one or more non-transitory computer-readable media of claim 17, wherein the first search query and the second search query are different.
19. The one or more non-transitory computer-readable media of claim 17, wherein the review input data comprises multi-modal data.
20. The one or more non-transitory computer-readable media of claim 19, wherein the multi-modal data comprises textual data and image data.
Citation Information
Patent Citations
System and method for presenting comments with media
CN103136326A
Comment determination method and device
CN107066536A