Generating prompt for user link note

The computing system addresses the challenge of obtaining relevant information about web resources by generating prediction prompts for user-generated comments, which are then stored with web resource data, enhancing user experience and search result accuracy.

JP2025080748AActive Publication Date: 2025-05-26GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024178950
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-10-11
Publication Date
2025-05-26
Estimated Expiration
2044-10-11

AI Technical Summary

Technical Problem

Users face difficulties in obtaining relevant information about web resources due to limited and irrelevant title and text fragments in search results, leading to time-consuming reviews without finding desired information. Additionally, users struggle to generate insightful notes that are relevant to other users.

Method used

A computing system that generates comment prompts and retrieves inputs by processing content data with a generative model to create prediction prompts. These prompts are provided to users through an input prompt interface, allowing them to generate user-generated comments which are then stored alongside the web resource data in a searchable database.

Benefits of technology

The system enhances user experience by providing additional relevant information about web resources through user-generated comments, reducing the time spent on reviewing irrelevant information and improving the accuracy of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025080748000001_ABST
    Figure 2025080748000001_ABST
Patent Text Reader

Abstract

To enhance understanding of search results from search result pages.SOLUTION: A system and method for generating prompts for user data input entry can include obtaining context data. The context data can be processed to determine whether an input entry interface is to be provided. In response to determination that the input entry interface is to be provided, the context data or other data associated with a content display instance can be processed with a generative model to generate a prompt that can be provided to a user. User input data can then be obtained and stored to be provided to other users.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims priority based on U.S. Non-Provisional Application No. 18 / 392,648, filed on December 21, 2023, and claims the benefit of U.S. Provisional Application No. 63 / 596,484, filed on November 6, 2023. The applicant claims the priority and the benefit of each of such applications and incorporates them herein by reference in their entirety.

[0002] The present disclosure generally relates to generating prompts for obtaining link notes. More specifically, the present disclosure relates to determining when and how to prompt a user to provide a note regarding a link associated with a web resource, which can then be provided to other users.

Background Art

[0003] Understanding search results from a search results page can be difficult because the title and fragments of text may provide limited information that may not be relevant to the user's interests, leading to a review of web resources that may take time and not result in the desired information. It can be difficult to obtain additional information about a web resource and may involve additional searches that do not identify relevant information or do not identify any information at all.

[0004] Furthermore, it can be difficult to obtain user insights. In particular, a user may struggle to determine which words to use. Additionally, the words may not be directed towards points of interest to other users and / or may not be rich enough to generate the desired results.

Summary of the Invention

[0005] Aspects and advantages of embodiments of the present disclosure are partially shown in the following description, or can be learned from the description, or can be learned through the practice of the embodiments.

[0006] One exemplary aspect of the present disclosure is directed to a computing system for generating comment prompts and retrieving inputs. The system can include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations can include obtaining content data. The content data can be associated with a web resource. The operations can include processing the content data with a generative model to generate a prediction prompt. The prompt can include a predicted text string related to a comment to the web resource. The operations can include providing the prediction prompt to an input prompt interface for display. The input prompt interface can be configured to receive an input. The operations can include obtaining comment input data from a user computing system via the input prompt interface. In some embodiments, the comment input data can include a user-generated comment to the web resource. The operations can include storing data associated with the comment input data together with data associated with the web resource. The data associated with the comment input data can be stored in a searchable database provided for display in response to the web resource provided as a search result.

[0007] In some embodiments, the operation can include obtaining user data. The user data can be associated with a specific user. The user computing system may be associated with a specific user. To generate a prediction prompt, processing the content data by a generation model can include processing the content data and the user data by the generation model. The user data can include user search history data. The generation model can generate a prediction prompt based on the fact that a specific user has previously searched for information associated with a topic of a web resource. In some embodiments, the user data can include user browser history data. The generation model can generate a prediction prompt based on the fact that a specific user has previously viewed other web resources that include information associated with a topic of a web resource. The operation can include generating a graphic card based on the user data, the content data, and the comment input data. The graphic card can include a user profile identifier of a specific user and data associated with the comment input data. The operation can include storing the graphic card. The graphic card can include a graphic background generated by an image generation model based on the comment input data.

[0008] In some embodiments, the operation can include obtaining a search query, determining that a web resource is associated with the search query, and providing specific search results for display. The specific search results can include a link to the web resource, a title of the web resource, and data associated with comment input data. Storing data associated with the comment input data together with data associated with the web resource can include generating a web resource note and storing the web resource note together with a plurality of other web resource notes associated with the web resource. In some embodiments, the operation can include providing the web resource note and the plurality of other web resource notes to a note interface that provides the web resource note and the plurality of other web resource notes to a plurality of graphics cards. The generation model can include an autoregressive language model. The generation model can be prompted to generate questions that describe requests for information about the web resource.

[0009] Other exemplary aspects of the present disclosure are directed to a computer-implemented method for link note prompts. The method can include obtaining context data by a computing system that includes one or more processors. The context data can be associated with a particular content display instance. The particular content display instance can include a particular user who is viewing a particular content item. The method can include determining an input request action by the computing system based on the context data. The input request action can include providing an input entry interface to the user to obtain user input. The method can include processing the context data with a generative language model by the computing system to generate a prediction prompt. In some embodiments, the prediction prompt can include a natural language request for information generated based on the context data. The method can include providing the prediction prompt at the input entry interface by the computing system, and obtaining user-generated content through the input entry interface by the computing system. The method can include generating a link note by the computing system based on the user-generated content. A link note can be generated for display at a search result interface in response to a particular content item being determined as a search result.

[0010] In some embodiments, context data can be associated with the type of content being provided for display. The context data can be associated with a particular user associated with a particular content display instance. The context data can include search history data. The content being provided for display can be associated with a particular web resource. The context data can be associated with interaction data of links to a particular web resource across multiple social network platforms. An input request action can be determined based on the interaction data. In some embodiments, the context data can include user data and content data. An input request action can be determined based on the topic associated with the content being provided for display being one of a plurality of topics determined that a particular user has knowledge of based on the user data. The context data can include previous notes generated by a particular user. A prediction prompt can include a structure based on a previous structure for the previous notes.

[0011] Other exemplary aspects of the present disclosure are directed to one or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations can include obtaining a first search query at a first time and determining that a web resource is responsive to the first search query. The operations can include obtaining content data. The content data can be associated with the web resource. The operations can include processing the content data with a generation model to generate a prediction prompt. The prompt can include a predicted text string associated with a comment to the web resource. The operations can include providing the prediction prompt for display within an input prompt interface. The input prompt interface can include an input entry box. The operations can include obtaining comment input data from a user computing system via the input prompt interface. The comment input data can include user-generated content. The operations can include storing the user-generated content. The operations can include obtaining a second search query at a second time. The second time can be different from the first time. The operations can include determining that the web resource is responsive to the second search query and providing the user-generated content to a search result interface with data describing the web resource.

[0012] Other exemplary aspects of the present disclosure are directed to a computing system for generating a graphics card. The system can include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations can include obtaining card data. The card data can describe the content within the graphics card. The content can be associated with one or more topics. The operations can include processing the card data to determine one or more entity tags associated with the content. The one or more entity tags can be associated with one or more topics. The operations can include accessing a media content item database to obtain one or more media content items. The one or more media content items can be obtained based on a determination that the one or more media content items are associated with one or more entity tags related to the content. The operations can include providing the one or more media content items for display. The one or more media content items can be provided for display in an interactive user interface. The one or more media content items can be selectable to be inserted into the graphics card.

[0013] In some embodiments, the graphics card can be associated with a link note. The link note can include user-generated content tagged to a particular web resource. The operations can include obtaining an input selection related to one or more media content items, generating an extended graphics card, and providing the extended graphics card for display. The extended graphics card can include at least a portion of the content of the graphics card and at least a portion of one or more media content items. In some embodiments, the operations can include obtaining an adjustment input. The adjustment input can be associated with a request to extend the extended graphics card. The operations can include generating an updated graphics card based on the adjustment input. The updated graphics card can include the extended graphics card with one or more adjustments made. The operations can include providing the updated graphics card for display. The one or more adjustments can include at least one of a layout change of the extended graphics card, a cropping change of one or more media content items, a size change of one or more content items, a color change, or a template change.

[0014] In some embodiments, the media content item database can include a user-specific database. The user-specific database can be associated with a particular user. A particular user may have generated at least a portion of the content. In some embodiments, the user-specific database can include an image gallery associated with a particular user. The image gallery can be stored in a server computing system associated with a particular content item storage platform. In some embodiments, the user-specific database can include a local storage database of a user computing device. The media content item database can include a plurality of media content items. In some embodiments, the plurality of content items may be preprocessed to generate a plurality of respective metadata sets.

[0015] Other exemplary aspects of the present disclosure are directed to a computing system. The system can include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations can include providing an input draft interface for display. The input draft interface can include a graphical user interface that includes a plurality of attribute options and a text input box. The plurality of attribute options can be associated with a plurality of candidate attributes for content item generation. The operations can include obtaining a selection of a particular attribute option among the plurality of attribute options via the input draft interface. The particular attribute option can be associated with a particular candidate attribute. The operations can include obtaining text input via the text input box of the input draft interface. The text input can be associated with a prompt intent for content item generation. The operations can include processing the particular attribute option and the text input using a generation model to generate a model-generated content item. The mode-generated content item can include the particular candidate attribute. In some embodiments, the model-generated content item can be associated with the prompt intent. The operations can include providing the model-generated content item for display via the input draft interface.

[0016] In some embodiments, the operation can further include obtaining an input selection via an input draft interface and generating an extended graphics card based on the input selection. The extended graphics card can include a graphics card extended to include model generation content items. The operation can include providing the extended graphics card for display. The plurality of candidate attributes can include a plurality of different styles. The plurality of different styles can be associated with at least one of a plurality of different artistic styles or a plurality of different writing styles.

[0017] In some embodiments, the plurality of candidate attributes can include a plurality of different tones. The plurality of different tones can be associated with at least one of a plurality of different emotions or a plurality of different pace types. In some embodiments, the generation model can be obtained from a generation model database based on the selection of a specific attribute option. A specific attribute soft prompt can be obtained based on the selection of a specific attribute option. The specific attribute soft prompt can include a set of learned parameters. The set of learned parameters can be processed by the generation model to generate model generation content items.

[0018] In some embodiments, the first search query and the second search query may be different. The comment input data can include multimodal data. The multimodal data can include text data and image data.

[0019] Other aspects of the present disclosure are directed to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.

[0020] These and other features, aspects, and advantages of the various embodiments of the present disclosure will become better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description of the embodiments, serve to explain the relevant principles.

[0021] A detailed description of embodiments directed to those skilled in the art is set forth in this specification with reference to the accompanying drawings.

Brief Description of the Drawings

[0022]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 5C

Figure 6A

Figure 6B

Figure 6C

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 9D

Figure 9E

Figure 9F

Figure 9G

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15A

Figure 15B

[0023] Reference numbers repeated throughout the several drawings are intended to identify the same features in various embodiments.

[0024] Generally, the present disclosure is directed to generating prompts for user data entry. In particular, the systems and methods disclosed herein can utilize context determination (e.g., determining the context in which a user is likely to provide a note and / or determining comment gaps and / or content gaps for specific links) to determine an input entry interface (e.g., a link note input entry interface), and provided to utilize a generation model (e.g., a large language model) to generate prompts based on user data (e.g., user search history and / or user browsing history) and / or content data (e.g., content topic and / or content type). For example, a user may be prompted to provide notes about a particular web resource (and / or other content item) on a search results page, during a review of a web resource, and / or in a next search instance. The prompts can be generated based on previous user notes, previously viewed content, content topic, and / or content type, and can provide the user with prompts that request information in a format that generates insightful notes.

[0025] A link note can provide additional information about a web resource without reviewing the web resource, and the link note can be provided by other users. The system and method can determine when to provide a link note prompt to a user based on a context determined to be associated with the capture of valuable notes. For example, a particular user may be able to provide more reliable and / or more detailed information about a particular topic based on previously acquired knowledge and / or based on previously generated notes. Additionally and / or alternatively, a particular content type can be determined to be associated with a user's comment and / or a user's confusion.

[0026] The prompt provided to the user can "inspire" the user to provide more detailed information and / or can instruct the user to leave a note about a particular topic and / or feature of the web resource. The generation model can process user data and / or content data to generate a prediction prompt. Specifically, the generation model can utilize the user's search history, the user's browsing history, the user's previous notes, and / or other user data to generate a proposed note, a question to prompt a response, and / or a template for a note. Alternatively and / or additionally, the generation model can utilize a semantic understanding of the web resource, topic classification, content type classification, other notes associated with the web resource, and / or other content data to generate a proposed note, a question to prompt a response, and / or a template for a note.

[0027] The input entry interface can provide a prediction prompt to the user. Next, the input entry interface can obtain input from the user (e.g., comment input data) and generate user-generated content that describes the link note. In some embodiments, a graphics card can be generated based on the link note. The graphics card can include user-generated content of the link note, a user profile identifier (e.g., name and / or image), link information, and / or a graphic background. The link note and / or the graphics card can be stored in association with a web resource. The stored link note and / or graphics card can then be obtained in response to one or more users searching for the web resource and / or one or more users interacting with the note interface.

[0028] Understanding search results from a search results page can be difficult because titles and text fragments may provide limited information that may not be relevant to the user's interests, leading to a review of web resources that may take time without obtaining the desired information. Additionally and / or alternatively, it may be difficult to obtain additional information about a web resource, which may involve additional searches where relevant information may or may not be identified. Social media posts, blog posts, and / or reviews of web resources and / or entities associated with web resources may lack details, be misleading, and / or lack context and / or perspective.

[0029] Link notes (e.g., link notes obtained from users and / or generated by a generation model) can provide additional information about web resources, which can notify other users of their relevance to their requests. Link notes can be provided on search result pages and / or displayed in a note interface accessible from search result pages and / or web resources. Link notes can be provided in a graphic card, in a text fragment and an inline text panel, and / or in other formats.

[0030] Determining when and how to prompt a user to generate a note can be based on determining note discrepancies (e.g., for an article, there are many notes on a blog platform and / or social media platform, but relatively few for the note interface of a search platform), determining user-specific interests (e.g., whether the resource is similar to other articles the user has viewed in the past), resource trends (e.g., whether this resource and / or similar resources have been commented on before), determining the resources for which notes should be made (e.g., whether the note will provide usefulness), and / or other determinations. The prompt can be generated based on generative language model processing, which can include processing previous notes (e.g., other notes by the user and / or other users), processing search queries, processing web resources, and / or processing other data.

[0031] Using the prompt of the link note, it is possible to start and / or prompt the collection of information about web resources that can then be provided to other users, which can identify the opinions of the user, the summary of the user, and / or the identified details of other users. The obtained notes can then be provided in a search result interface and / or a discovery feed. The prompt of the link note can be determined based on determining note discrepancies, determining user-specific interests, resource trends, determining the resources for which notes should be made, and / or other determinations.

[0032] Obtaining additional information from other users can be useful to the user for determining the topics and quality of search results that may be difficult to distinguish from conventional search result displays. However, it may be difficult to obtain useful and detailed information. It is possible to provide an interface for obtaining detailed information from relevant users by determining when to prompt the user to generate a note and utilizing the generation of context-aware prompts.

[0033] It may be difficult to prompt the user to provide information (such as comments, reviews, insights, etc.) when the user views search results. Using a prompt generation system, it is possible to target the appropriate user with the appropriate prompt at the appropriate time / location based on insights. In particular, the prompt can help create posts having desired characteristics (such as a desired level of detail regarding a desired topic, and / or other characteristics). The prompt generation system disclosed herein can facilitate generating (or creating) a note for a particular user and / or can facilitate generating (or creating) a note for a particular web resource (and / or content item) via the prompt generation system.

[0034] In some embodiments, the systems and methods disclosed herein can be utilized to prompt the user to generate stand-alone content. Stand-alone content can include user recipes, user tutorials, user graphics, life updates, link sharing, and / or other user-generated content. Stand-alone content can be generated in free form and / or based on a prompt for model generation. In some embodiments, one or more machine-learned models can be utilized to generate a content template and / or can be utilized to expand user-provided content (e.g., reconstruct and / or restyle text, images, audio, interface elements, and / or video).

[0035] In some embodiments, link notes and / or interactions with link notes can be used to adjust the ranking, tagging, embedding, and / or indexing of web resources. For example, in some embodiments, link notes can be processed to determine the quality of a web resource. The determination of quality can be based on processing the link note using one or more machine-learned models (e.g., sentiment analysis models, language models, classification models, etc.). Processing the link note with one or more machine-learned models can identify topics related to the web resource, and identify the bias, usefulness, and / or direction of the web resource. Link notes can be used to propose additional content, can be embedded for embedding-based search, and / or can be used in query proposal.

[0036] Link notes in the note interface can be ranked and / or displayed based on interactions, quality determined by machine-learned models, responsiveness to queries, level of detail, and / or other attributes. In some embodiments, link notes generated by a user can be provided only to that user, to all other users, only to users within the user's social network, and / or only to users determined to be associated with the user based on interests, location, and / or activity.

[0037] Link notes can be used for multiple different content items and may not be limited to web resources. For example, using the systems and methods disclosed herein, prompts and / or interfaces can be generated to obtain, inspire, and / or generate link notes for other content item sources that can include local files (e.g., documents, images, videos on a device), intranet files, and / or folders on an external drive, cloud documents, etc.

[0038] In some embodiments, the input interface can include an open input interface that provides one or more options for user input. Alternatively and / or additionally, the input interface can include a plurality of features and / or options for generating user-generated content that can be utilized in linked notes and / or stand-alone content. The input interface can enable a user to add images, links, and / or content of different template types, and can include a user interface for independent content items that can be interactive. The interactive user interface can include suggestions for images, templates, text, layouts, links, widgets, templates, and / or other options (e.g., suggestions of other types).

[0039] Image proposals can include processing user input text, data related to web resources, generated prompts, stock photo libraries, and / or an image database associated with the user (e.g., an online image gallery associated with the user and / or local images on the user's computing device) to determine images relevant to a particular context (e.g., related to user input, web resources, and / or generated prompts). Image proposals can include determining one or more entities, topics, and / or features associated with a web resource and / or user input, and then processing a stock photo library and / or image database(s) associated with the user to determine one or more specific images associated with the one or more entities, topics, and / or features associated with the web resource and / or user input. For example, if a web resource describes a pasta recipe, images depicting pasta, cooking, pasta ingredients, and / or a kitchen can be searched for in a stock image gallery and / or user image gallery. Another example can include determining text of a generated prompt that can be associated with a trip to Mexico, and one or more images from the user's image gallery can be identified and proposed based on location metadata, feature detection, optical character recognition, and / or other determination techniques that can be utilized to identify one or more images associated with a trip to Mexico. In some embodiments, image proposals can be based on generating prompt embeddings and / or web resource embeddings and then performing an embedding search based on the plurality of image embeddings associated with one or more image databases. The determination and display of proposals can be made for images, videos, document files, audio, text data, templates, and / or other data.

[0040] Additionally and / or alternatively, the interactive input interface can include a "Help Me Write" function. The "Help Me Write" function can be a selectable user interface feature that provides a generative language model interface for generating text for user-generated content. The "Help Me Write" function can include a drop-down menu for selecting specific tones, styles, formats, lengths, and / or other attributes for the model-generated text. The "Help Me Write" function can process user input to adjust and / or change the style, tone, format, language, vocabulary, length, and / or level of conciseness of the input text. For example, the user can select a tone from multiple tone options and enter a text string, and the input interface can provide the text string and the selected tone prompt to a generative language model (e.g., a large language model), and then generate a model-generated text response that can be used for user-generated content (e.g., link notes and / or sand arrow content). Additionally and / or alternatively, the input interface can interface with different generative language models associated with different attributes in response to the selection of different attribute options. Different generative models can be trained and / or adjusted for specific attributes.

[0041] The systems and methods of the present disclosure provide several technical effects and advantages. As an example, the systems and methods can provide an interactive user interface that can be utilized to generate prompts and obtain user input data. Specifically, the systems and methods disclosed herein can utilize one or more machine-learned models to determine when to request a link note and generate a prompt for requesting information. For example, a generation model can process user data, content data, and / or other context data to determine that a request for an information action should be made. Additionally and / or alternatively, the generation model can generate a prompt for requesting information based on user data, content data, and / or other context data. The prompt can be provided to the user, user input can be received, and a link note can be generated and stored.

[0042] Other technical advantages of the systems and methods of the present disclosure include the ability to utilize user data and content data to determine which users can provide reliable information regarding a particular web resource and / or when to prompt a user to provide information. For example, a user can be determined to be knowledgeable about a particular topic and / or a poster of common notes for a given type of content. Based on the determination, the user can be prompted to provide a link note to a given web resource. Alternatively and / or additionally, the topic of the content, the type of the content, and / or other interactions with the content can be utilized to determine that a web resource is "expert" at commenting. The prompt can be generated by a generation model to provide a prompt that is both user-aware and content-aware.

[0043] The systems and methods disclosed herein address problems arising from computer systems that acquire, process, and transmit data from multiple databases from multiple sources. The vast amount of data available to users has the potential to introduce misinformation, false leads, and / or lack of verification. Fragments of text, titles, and / or exemplary images of search result interfaces may provide some details about the content of web resources. However, information from other users can provide further insights regarding topics, reliability, and / or expectations, and this can be utilized to reduce instances of irrelevant web resources that users navigate and review.

[0044] Other examples of technical effects and benefits relate to improvements in computational efficiency and the functionality of computing systems. For example, the systems and methods disclosed herein can provide an interface that provides information about links, which may reduce the review of redundant search results by leveraging the generation of notes to provide user-based validation. As the amount of follow-up queries decreases and the amount of page redirects decreases, latency on user devices can be reduced and the computational cost of search engines can be reduced. Here, with reference to the figures, exemplary embodiments of the present disclosure will be described in more detail.

[0045] FIG. 1 shows a block diagram of an exemplary link note generation system 10 according to an exemplary embodiment of the present disclosure. In some embodiments, the link note generation system 10 is configured to receive and / or obtain context data 12 that describes user data, content data, and / or other context data associated with a web page and / or a viewing stance. As a result of receiving the context data 12, a prediction prompt 18 that describes a generated natural language request for information from a user is generated, determined, and / or provided. Thus, in some embodiments, the link note generation system 10 includes a generation model 16 operable to process the context data 12 to generate a prediction prompt 18 that includes text strings that describe questions, commands, templates, and / or proposed comments.

[0046] In particular, the link note generation system 10 can obtain the context data 12. The context data 12 can include user data (e.g., data associated with a user who is viewing a search results page, entering a search query, and / or viewing a discovery feed), content data (e.g., data associated with the content of the web resource 14), and / or other context data (e.g., time, query trends, comment trends, news, etc.). The context data 12 can include user search history data, user browsing history data, user purchase history data, user profile data, user note history data, topic label data of web resources, content type labels of web resources, other notes about web resources, and / or other data. The context data 12 can be generated using a personalized machine-learned model and / or one or more other machine-learned models.

[0047] Context data 12 can be obtained based on web resources 14 that are provided for display, previously reviewed, and / or associated with search results on a search results page. Context data 12 can be obtained and / or generated based on one or more user interactions, one or more global trends, and / or web resources 14 associated with a particular type of content (e.g., treatises, tutorials, blogs, news articles, sports score trackers, etc.).

[0048] The generation model 16 (e.g., an autoregressive language model, a diffusion model, and / or one or more other generation models) can process the context data 12 to generate a prediction prompt 18. The generation model 16 can include a language model (e.g., a large language model, a vision-language model, and / or other language models), a text-to-image generation model, and / or other generation models. The prediction prompt 18 can include text data, image data, audio data, latent encoding data, and / or multimodal data. The prediction prompt 18 can include a question that a user can respond to, a template for drafting a note, and / or one or more selectable note options. For example, the prediction prompt 18 can include a question generated based on a semantic analysis and / or topic determination of a web resource, a template generated based on previously generated notes by a particular user and / or other users, and / or selectable note options based on previous comments provided for a similar web resource. In some embodiments, the prediction prompt 18 can be something that describes a request for a particular type of information regarding the web resource 14. Alternatively and / or additionally, the prediction prompt 18 can describe general information regarding the web resource 14. The prediction prompt 18 can include a novel text string that has not been previously provided by the user and / or is not provided in relation to the web resource 14. The prediction prompt 18 can include a plurality of predicted characters, words, pixels, signals, and / or structures.

[0049] The prediction prompt 18 can be provided for display in the input entry interface. The user can interact with the input entry interface to generate user-generated content 20 that can be transmitted to a server computing system (e.g., a search engine computing system). The user-generated content 20 can include text data, image data, audio data, video data, potential encoding data, and / or multimodal data. The user-generated content 20 can describe a note regarding the web resource 14. The note can describe an interpretation, opinion, review, verification, and / or the display of a quality and / or topic. The user-generated content 20 can include one or more graphics, one or more widgets, one or more links, one or more media content items, and / or a note displayed on a graphic card having a graphic background.

[0050] The link note generation system 10 can index the web resource 14 into the link note 22. By utilizing the indexing, when providing search results of the web resource 14, user-generated content 20 including a link note for display can be provided. Alternatively and / or additionally, the user-generated content 20 can be stored in a note database and can be displayed on a note interface when selected by one or more users.

[0051] FIG. 2 shows a block diagram of an exemplary user prompt system 200 according to an exemplary embodiment of the present disclosure. The user prompt system 200 is similar to the link note generation system 10 of FIG. 1, except that the user prompt system 200 further includes an action determination block 230.

[0052] The user prompt system 200 can obtain content data 224 and user data 226. The data can be obtained in response to a search query, an on-back event to a search result page (e.g., returning to the search result page after displaying a web resource), the next search instance, the next instance of a web resource that is a search result, and / or other trigger events. The content data 224 can describe the content of a web resource and can include text, images, videos, layouts, audio files, transitions, potential encoding data, related links / web resources, an interaction history of the web resource, and / or other data related to the web resource and / or other similar web resources. The user data 226 can include a user search history (e.g., a log of previous queries obtained from the user), a user browsing history (e.g., a log of previously visited web pages and / or platforms), a user application history (e.g., a log of previous interactions with an application), a user purchase history (e.g., a log of previously acquired products and / or services), a user profile (e.g., a user identifier, user preferences, user title, user account, and / or user contacts), a note history (e.g., a log of previously provided / generated notes), and / or data describing a social media network and / or activity.

[0053] The content data 224 and / or the user data 226 can be processed by a context determination block 228 to determine a context. The context determination block 228 can include one or more machine-learned models and / or one or more deterministic functions. The context determination block 228 can generate context data.

[0054] The context data can be processed by the action decision block 230 to determine that an input request action is to be performed. The action decision block 230 can include one or more machine learned models and / or one or more deterministic functions. The context decision and / or action decision can be made based on heuristics.

[0055] The input request action can include generating a prompt and providing an input entry interface to the user using the prompt to obtain a link note of a given web resource. The input request action can be determined based on the likelihood that the user will respond, the reliability of the user, the experience of the user, the knowledge of the user, different users associated with the previous note provider, the content gap describing the difference in notes for a particular web resource versus a similar web resource, the comment gap describing the difference in the interaction between one or more blogs or social media platforms and the amount of links to notes, the topic of the web resource, the topic of the search, the content type, the intent of the content, and / or other data.

[0056] Certain content types (e.g., news articles, short stories, movies, skits, blog posts, and / or social media posts) can be determined to be more likely to be interacted with for link note generation and / or to receive more benefit from link notes. Additionally and / or alternatively, interactions with web resources on other platforms can be determined. If the amount and / or quality of the interaction on other platforms is determined to meet a threshold difference compared to the current platform, the input request action can be determined more frequently. For example, the threshold for the input request action can be adjusted based on the interaction on other platforms. Alternatively and / or additionally, the threshold can be adjusted based on search and / or viewing trends related to the web resource.

[0057] The input request action may be performed immediately after the determination and / or may be provided later as a "nudge", which may be the time when it is determined that the likelihood of a response is higher (e.g., when the user is at a specific location (e.g., home), when the user's calendar is empty, a specific time of day when phone activity has increased, and / or at the next user search instance). The "nudge" may be provided via device notifications, emails, and / or application-based notifications.

[0058] Next, the user prompt system 200 can utilize the context data to generate a prompt for requesting a note from the user based on the determination of the input request action. The context data can include user data 226 (e.g., data associated with a user who is viewing a search result page, entering a search query, and / or viewing a discovery feed), content data 224 (e.g., data related to the content of a web resource), and / or other context data (e.g., time, query trends, comment trends, news, etc.). The context data can include user search history data (e.g., a list of previously searched search queries, which may include queries related to the same topic as the web resource), the user's browsing history data (e.g., a list of previously viewed web pages, which may include web pages associated with the same topic as the web resource), the user's purchase history data, the user profile data (e.g., the user's name, occupation, education, preferences, etc.), the user note history data, the topic label data of the web resource, the content type label of the web resource, other notes related to the web resource, and / or other data. The context data can be generated using a personalized machine-learned model and / or one or more other machine-learned models.

[0059] Context data can be obtained based on web resources that are provided for display, previously reviewed, and / or associated with the search results of a search results page. Context data can be obtained and / or generated based on one or more user interactions, one or more global trends, and / or web resources associated with a particular type of content (e.g., treatises, tutorials, blogs, news articles, sports score trackers, etc.).

[0060] The generation model 216 (e.g., a text generation model, an image generation model, an audio generation model, a video generation model, and / or a multimodal media content item generation model) can process context data to generate a prediction prompt 218. The generation model 216 can include a language model (e.g., a large language model, a vision language model, and / or other language models), a text-to-image generation model, and / or other generation models. The prediction prompt 218 can include text data, image data, audio data, latent encoding data, and / or multimodal data. The prediction prompt 218 can include questions that a user can respond to, templates for drafting notes, and / or one or more selectable note options. For example, the prediction prompt 218 can include questions generated based on semantic analysis and / or topic determination of web resources (e.g., in the case of an article about an unsolved case, the prompt may include "What do you think about the analysis of the unsolved case?", "Who do you think committed the crime?", "Was the article easy to understand and comprehensive regarding forensic evidence?", etc.). In some embodiments, the prediction prompt 218 can include templates generated based on notes previously generated by a particular user and / or other users (e.g., if a user typically prefixes their notes, the prompt may include a template that starts with a preposition emulating the style and tone of previous notes). Additionally and / or alternatively, the prediction prompt 218 can include selectable note options based on previous comments provided for similar web resources (e.g., "The analysis of the case was comprehensive and understandable, and a reasonable conclusion was obtained", "The forensic analysis lacked a basis for scientific verification", "This article is more of a fan creation than an actual article", etc.).In some embodiments, the prediction prompt 218 may describe a request for a specific type of information about a web resource (e.g., for a biography of a politician, "What did you think about their background?", "Please state your opinion on the epilogue.", "In my experience as a parliamentary historian, this biography is accurate / unreliable / well-written / poorly structured", etc.). Alternatively and / or additionally, the prediction prompt 218 may describe general information about the web resource. The prediction prompt 218 can include new text strings that have not been previously provided by the user and / or are not provided in relation to the web resource. The prediction prompt 218 can include a plurality of predicted characters, words, pixels, signals, and / or structures.

[0061] In some embodiments, the specific terms, details, and / or structure of the search query can be utilized to determine the level of the user's experience and / or knowledge regarding a particular topic. The search query may be included in the context data, and the generation model 216 can generate a prediction prompt 218 that reflects the determined level of experience and / or knowledge. Additionally and / or alternatively, previous search queries can be utilized to determine a chain of search queries and to determine the search intent. Next, the search intent can be utilized to generate a prediction prompt 218 associated with the search intent.

[0062] In some embodiments, the prediction prompt 218 may vary based on the user's tendency to provide link notes, based on the user's previous notes, based on the user's credibility, and / or based on other user data. If the user has not previously provided link notes and / or has only provided a few link notes previously, the prediction prompt 218 can be composed of general prompts, multiple-choice prompts, and / or formatted conversations. The prediction prompt 218 for an experienced user can be generated to provide direct prompts, note templates, and / or options to the user based on previous interactions.

[0063] The prediction prompt 218 can be provided for display in the input entry interface. The user can interact with the input entry interface to generate user-generated content 220 that can be sent to a server computing system (e.g., a search engine computing system). The user-generated content 220 can include text data, image data, audio data, video data, potential encoding data, and / or multimodal data. The user-generated content 220 can describe a note regarding a web resource. The note can describe an interpretation, an opinion (e.g., "I think the trade discussed in this article was fair based on the long-term results of both teams"), a review (e.g., "This short story has an inappropriate pace and direction, and there is no growth of the protagonist at all"), a verification (e.g., "The facts in this article are consistent with other reliable sources of information"), and / or an indication of quality and / or topic (e.g., "A very well-written play depicting the perils of love in a town devastated by war"). The user-generated content 220 can include notes displayed on one or more graphics, one or more widgets, one or more links, one or more media content items, and / or a graphics card with a graphic background. For example, the text of a link note can be provided in a stereotyped text having a color alongside a model-generated image used as a background.

[0064] In some embodiments, the generation model 216 can process the user-generated content 220 and generate a follow-up prompt. The follow-up prompt may request additional information and / or may provide options for further customization.

[0065] The user prompt system 200 can store the note 222 together with the web resource. When providing search results of the web resource by utilizing indexing, user-generated content 220 including link notes for display can be provided. Alternatively and / or additionally, the user-generated content 220 can be stored in a note database and displayed on a note interface when selected by one or more users.

[0066] For example, a specific user and / or other users can enter a search query. The search engine system may determine that the web resource responds to the search query. The search results associated with the web resource can be provided in a search result interface having data describing the title of the web resource, media fragments, and link notes (e.g., a graphics card).

[0067] FIG. 3 shows a flowchart diagram of an exemplary method for functioning according to an exemplary embodiment of the present disclosure. Although FIG. 3 shows steps performed in a specific order for purposes of illustration and description, the method of the present disclosure is not limited to the specifically shown order or arrangement. The various steps of method 300 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0068] At 302, the computing system can obtain content data. The content data can be associated with a web resource. The content data can describe the content of the web resource and can include text data, image data, video data, audio data, potential encoding data, and / or multimodal data. The content data can include the topic of the web resource, the type of content, other notes received from other users, the metadata of the web resource, the producer of the web resource, and / or descriptive data of entities related to the web resource. The content data can include content labels, the entire web resource content, a summary of the content, media fragments, and / or content embeds.

[0069] At 304, the computing system can process the content data with a generative model to generate a prediction prompt. The prompt can include a predicted text string related to a comment on the web resource. The generative model can include an autoregressive language model. In some embodiments, the generative model can be prompted to generate a question that describes a request for information about the web resource. The generative model may include a transformer model. The generative model may be trained, configured, and / or prompted to perform semantic understanding on the web resource and then generate a prompt (e.g., a question) based on the semantic understanding. The prediction prompt can include a plurality of predicted characters that can be specifically determined based on the content data. The prediction prompt may ask about the quality of the web resource. The prediction prompt may ask about the opinion and / or review of the web resource.

[0070] In some embodiments, a computing system can obtain user data. The user data can be associated with a particular user. To generate a prediction prompt, processing content data by a generation model can include processing the content data and the user data by the generation model. The user data can include user search history data, user browsing history data, the user's social network, user preferences, user profile information, the user's location, user purchase history, and / or user connections. The generation model can generate a prediction prompt based on the particular user having previously searched for information associated with the topic of a web resource. Alternatively and / or additionally, the generation model can generate a prediction prompt based on the particular user having previously viewed other web resources that include information related to the topic of a web resource. The generation model can determine that a particular user is associated with a particular topic, a particular type of content, a particular opinion, and / or a particular context for commenting, and generate a prediction prompt based on the determination.

[0071] At 306, the computing system can provide a prediction prompt for display using an input prompt interface. The input prompt interface can be configured to receive input. The input prompt interface can include a plurality of selectable user interface elements. In some embodiments, the input prompt interface can include an input entry box for receiving input from a user. The input prompt interface can include an upload element for uploading media content items (e.g., documents, images, text, videos, audio files, etc.). In some embodiments, the input prompt interface can include an interface for a user to provide input to a generative model to generate model-generated notes based on the user input. Alternatively and / or additionally, the plurality of selectable user interface elements may include one or more selectable templates for generating user-generated content (e.g., user-generated notes). The one or more templates can be generated based on content items (e.g., previously generated notes) previously generated by a particular user.

[0072] At 308, the computing system can obtain comment input data from a user computing system via the input prompt interface. The comment input data can include text data, image data, audio data, latent encoding data, and / or multimodal data. The comment input data can include one or more selections, one or more text strings, and / or one or more uploaded files. The comment input data can include user-generated comments regarding web resources. The user-generated comments can include comments regarding the quality of the web resource, the topic of the web resource, and / or other aspects of the web resource.

[0073] At 310, the computing system can store data associated with the comment input data together with data associated with the web resource. Storing the data associated with the comment input data together with the data associated with the web resource can include generating a web resource note and storing the web resource note together with a plurality of other web resource notes associated with the web resource. The comment input data can be indexed in relation to the web resource, stored in a database, and search results of the web resource can be provided. The data associated with the comment input data can be stored in a searchable database and provided for display in response to a web resource being determined and / or provided as a search result.

[0074] In some embodiments, the computing system can provide the web resource note and the plurality of other web resource notes to a note interface that provides the web resource note and the plurality of other web resource notes to a plurality of graphics cards.

[0075] Additionally and / or alternatively, the computing system can generate a graphics card based on user data, content data, and comment input data. The graphics card can include a user profile identifier of a particular user and data associated with the comment input data. The computing system can then store the graphics card. The graphics card can include a graphic background generated by an image generation model based on the comment input data.

[0076] In some embodiments, the computing system can obtain a search query, determine that a web resource is associated with the search query, and provide specific search results for display. The specific search results can include a link to the web resource, a title of the web resource, and data associated with the comment input data.

[0077] Figure 4 shows a diagram of an exemplary prompt according to an exemplary embodiment of the present disclosure. Specifically, an exemplary input entry interface 402 is shown in Figure 4. The exemplary input entry interface 402 can be provided in response to determining that an input request action should be performed and a prediction prompt should be generated. The input entry interface 402 can include a link and / or reference to a web resource 404, a configuration panel 406 for displaying the input provided during the generation of user-generated content, one or more selectable prediction prompts 408, and / or one or more user interface elements for providing one or more audio inputs and / or multimedia inputs (e.g., images).

[0078] One or more selectable prediction prompts 408 can include prediction prompts generated by processing the content of the web resource 404 using a generative language model. Alternatively and / or additionally, one or more selectable prediction prompts 408 can include prediction prompts generated by processing notes associated with articles similar to the articles provided by the web resource 404 using a generative language model.

[0079] The proposed prediction prompts can be provided in multiple formats. Additionally or alternatively, the number and / or length of the prediction prompts can vary based on the content, the user, and / or other contextual data. For example, three options 412 can be provided, or ten options 414 can be provided. The note options can include multiple selectable prompts that the user can select as link notes and / or as part of link notes.

[0080] The input entry interface 402 can be utilized to receive text input (e.g., via a graphical keyboard interface), audio input (e.g., via one or more microphones), selections (e.g., selection of user interface elements associated with predictive prompt note options), and / or input of media content items (e.g., uploading of an image). The received input can be provided for display in the preview window of the configuration panel 406 and can then be utilized to generate a user-generated content item that can include a link note.

[0081] FIG. 5A shows a diagram of an exemplary note interface for prompting a topic according to an exemplary embodiment of the present disclosure. The exemplary note interface can be utilized to learn more about the web resource 502, read the thoughts of other users regarding the web resource 502, and / or generate and provide new notes regarding the web resource 502. Specifically, FIG. 5A shows a link to the web resource 502, previously provided notes 504, one or more prompts 506 for requesting information from the user, and recommended user interface elements 508. The link to the web resource 502 can include a thumbnail, a URL, and a title. The previously provided notes 504 can include images and / or text provided by other users. The previously provided notes 504 can be provided with interaction data including recommendations, comments, etc. One or more prompts 506 for requesting information from the user can include topics for discussion for the user to respond to the note. The recommended user interface elements 508 can be utilized to interact with the web resource 502 and / or the previously provided notes 504.

[0082] FIG. 5B shows a diagram of an exemplary note interface that prompts for similar articles according to an exemplary embodiment of the present disclosure. Using the exemplary note interface, one can further learn about web resource 502, read the thoughts of other users regarding web resource 502 (e.g., opinions regarding the topic of the web resource and / or the quality of the web resource), and / or generate and provide new notes regarding web resource 502 (e.g., a user can provide their details and / or thoughts regarding the web resource). Specifically, FIG. 5B shows a link to web resource 502, previously provided note 504, user interface element 510 for commenting on articles similar to web resource 502, and proposed article 512. The link to web resource 502 can include a thumbnail, URL, and title. The previously provided note 504 can include images and / or text provided by other users. The previously provided note 504 can be provided along with interaction data including recommendations, comments, etc. The user interface element 510 for commenting on articles similar to web resource 502 can be selectable to open an input entry interface for providing link notes regarding one or more similar web resources, which may include proposed article 512.

[0083] Figure 5C shows a diagram of an exemplary prediction prompt according to an exemplary embodiment of the present disclosure. Specifically, content data related to a web resource and / or user data related to a specific user being prompted can be processed by a generative model to generate one or more prediction prompts. Figure 5C shows an exemplary question prompt 514 and a starting point prompt 516. The question prompt 514 can include a question for asking the user to provide information about a specific topic and / or subtopic. The starting point prompt 516 can include an introductory sentence for the user to build from their link note. User-generated content including link notes can include a starting prompt when provided in a search result interface and / or a note interface.

[0084] Figures 6A - 6C show diagrams of exemplary link note entry points according to exemplary embodiments of the present disclosure. The entry points to the link note generation interface can be provided in a plurality of different interfaces and / or media. Specifically, Figure 6A shows an exemplary search result interface of exemplary search results 602. Context data related to a search instance (e.g., user data, search query, and / or search results) can be processed to determine that a link note prompt is to be provided. Figure 6A shows a first exemplary prompt 604 provided in response to determining that the user has previously searched for this topic. Further, Figure 6A shows a second exemplary prompt 606 provided in response to determining an interaction tendency and / or determining that a specific web resource is a content type that is frequently interacted with for note generation.

[0085] Figure 6B shows three different entry points for the prompt and generation of link notes. The first entry point 608 can include a drop-down menu of the browser application. The user can select the optional drop-down and then select a note generation option from that drop-down. The second entry point 610 may be included within the search results interface. The user can select an option within the search results interface to comment on previously viewed content. The third entry point 612 can include an overlay user interface element that can be provided in the viewing window of the content of the web resource. The overlay user interface element can be provided by the web resource, the browser, and / or the operating system of the user computing device.

[0086] Figure 6C represents a search review entry point. Specifically, a widget and / or graphical pane can be provided to the user to provide a review of previous search experiences, which can include selecting the drop-down list element 614, selecting a web resource from the viewed content list 616, and indicating which web resources were useful for a particular search. The selection can be utilized to generate a link note, re-rank the web resources, and / or navigate to the link note generation interface.

[0087] Figure 7 shows a flowchart diagram of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. Figure 7 shows steps that are executed in a particular order for purposes of illustration and explanation, but the method of the present disclosure is not limited to the specifically shown order or arrangement. The various steps of method 700 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0088] At 702, the computing system can obtain context data. The context data can be associated with a specific content display instance. A specific content display instance can include a specific user who is viewing a specific content item. In some embodiments, the context data can be associated with the type of content provided for display (e.g., article, academic paper, blog post, encyclopedia entry, video, media content library, and / or other types of content). Additionally or alternatively, the context data can be associated with a specific user associated with a specific content display instance. The context data can include search history data, browsing history data, user profile data, purchase history data, social network data, and / or other user data.

[0089] At 704, the computing system can determine an input request action based on the context data. The input request action can include providing an input entry interface to the user to obtain user input. The content provided for display can be associated with a specific web resource. The context data can be associated with interaction data of links to specific web resources of multiple social network platforms (e.g., link posts, comments, reposts, likes, and / or mentions). In some embodiments, the input request action can be determined based on the interaction data. The context data can include user data and content data. Additionally or alternatively, the input request action can be determined based on that the topic associated with the content provided for display is one of multiple topics determined that a specific user has knowledge of based on the user data.

[0090] At 706, the computing system can process the context data with a generative language model to generate a prediction prompt. The prediction prompt can include a natural language request for information generated based on the context data. The context data can include previous notes generated by a particular user. The prediction prompt can include a structure based on a previous structure for the previous notes. In some embodiments, the generative model can process the content data to generate a prediction prompt based on the content of a particular web resource.

[0091] At 708, the computing system can provide the prediction prompt to an input entry interface. The prediction prompt can be provided adjacent to an input entry box for receiving and displaying input text and / or an image. The input entry interface can include a panel adjacent to search results for a particular web resource. Alternatively and / or additionally, the input entry interface can be provided in a pop-up interface and / or can be redirected based on one or more inputs.

[0092] At 710, the computing system can obtain user-generated content via the input entry interface. The user-generated content can include text data, image data, video data, audio data, latent encoding data, statistical data, and / or multimodal data. The user-generated content can be obtained via an upload interface and / or via the input entry box.

[0093] At 712, the computing system can generate link notes based on user-generated content. In some embodiments, the computing system can generate a graphics card that includes the link notes. The graphics card can include a graphic background that can be selected by the user and / or automatically generated. The graphic background can be generated based on the content of the web resource, the content of the link note, and / or the type of note. The graphics card and / or the link note can be stored to provide data related to the web resource. In response to a particular content item being determined as a search result, a link note can be generated for display on a search result interface.

[0094] FIG. 8 shows a flowchart diagram of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. FIG. 8 shows steps performed in a particular order for purposes of illustration and explanation, but the methods of the present disclosure are not limited to the particularly shown order or arrangement. The various steps of method 800 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0095] At 802, the computing system can obtain a first search query at a first time and determine that a web resource responds to the first search query. The first search query can include a text query, an image query, an audio query, an embedded query, and / or a multimodal query. The web resource can be identified by a search engine, and the search engine can perform keyword search, embedding-based search, and / or other search techniques. The web resource can be determined to respond to the topic, question, and / or intent of the first search query.

[0096] At 804, the computing system can obtain content data and process the content data with a generative model to generate a prediction prompt. The content data can be associated with a web resource. The prompt can include a predicted text string associated with a comment to the web resource. The content data can describe the entire content of the web resource, a media fragment, a summary of the content, content labels, metadata, and / or the content of previously provided link notes related to the web resource. The generative model can process the content data to determine a topic, perspective, intent, theme, structure, intended audience, type of content, and / or other content details. Next, a prediction prompt can be generated based on the determination.

[0097] At 806, the computing system can provide the prediction prompt for display within an input prompt interface and obtain comment input data from the user computing system via the input prompt interface. The input prompt interface can include an input entry box. In some embodiments, the input prompt interface can include a plurality of user interface elements for drafting content (e.g., notes). The comment input data can include user-generated content. In some embodiments, the comment input data can include multimodal data. The multimodal data can include text data and image data.

[0098] At 808, the computing system can store user-generated content. The user-generated content can be indexed with a link to the web resource. Alternatively and / or additionally, the web resource can be indexed with the user-generated content. The user-generated content can be stored with other user notes related to a particular web resource and / or a particular user.

[0099] At 810, the computing system can obtain a second search query at a second time and determine that the web resource is responding to the second search query. The second time can be different from the first time. In some embodiments, the first search query and the second search query can be different. The second search query can include a text query, an image query, an audio query, an embedded query, and / or a multimodal query. The web resource can be identified by a search engine, which may perform keyword search, embedding-based search, and / or other search techniques. The web resource can be determined to respond to the topic, question, and / or intent of the second search query.

[0100] At 812, the computing system can provide user-generated content to a search result interface with data describing the web resource. The user-generated content can include a link to the web resource, a title, and a text fragment.

[0101] Figures 9A - 9G show diagrams of an exemplary graphic card interface according to an exemplary embodiment of the present disclosure. Using the systems and methods disclosed herein, a graphic card for user-generated content that can include link notes and / or stand-alone content can be generated.

[0102] FIG. 9A shows two exemplary graphic cards. The graphic cards shown may include a user profile identifier 902 (e.g., a user profile image and name), a body of the card 904 (e.g., a templated text link note overlaid on a graphic background), widget interface elements 906 (e.g., selectable user interface elements for redirecting to web resources and / or additional content), and / or interaction information 908 (e.g., likes, comments, and / or saves to the graphic card). The body of the card 904 may be configured by a particular user and / or may be automatically generated based on link notes, web resources, and / or user preferences. The widget interface elements 906 can include links to web pages, links to image galleries, links to videos, selectable elements for opening a popup interface, additional notes, and / or other data.

[0103] FIG. 9B shows exemplary multi-page user-generated content and exemplary video user-generated content. The multi-page user-generated content can include a plurality of graphic cards that can cycle to display the user-generated content. The video user-generated content can include a video with text overlaid on the graphics and / or video.

[0104] The widget interface element 906 can include a link to a web resource, a link to one or more other web resources, a video element selectable to provide a video for display, a media content display element (e.g., video, image, audio file, and / or other media) for providing media content for display, a review of a web resource, a link to other notes, a structured content item (e.g., a structured recipe and / or a structured computer), a list (e.g., a materials list), a map place card (e.g., a map associated with a web resource and / or a link to a web application), a knowledge panel, and / or a link to a shopping interface.

[0105] FIG. 9C shows an exemplary interaction with a graphics card. For example, the user profile identifier 902 can be selected to peak the user's profile 910. Additionally and / or alternatively, the video widget element 912 can be selected to deploy a video for playback and / or to navigate to a video player interface. The graphics card can be selected to minimize the add-on 914. The knowledge panel widget element 916 can be selected to expand the knowledge panel and provide additional information for display. The interaction element 918 can be selected to like, comment, save, and / or share user-generated content.

[0106] FIG. 9D shows the prominence level of a widget interface element (e.g., an add-on element). Specifically, the illustrated smoothie materials and instructions can be provided in a medium prominence interface element 922 (e.g., in a detailed view state), a low prominence interface element 924 (e.g., in a collapsed state), and / or a high prominence graphic panel 926 (e.g., in an expanded state). In some embodiments, the user can interact with the widget interface element to transition between levels of detail and / or size.

[0107] Figure 9E shows the search results of notes in the search result interface 930. The search results of notes can be provided in a separate tab adjacent to other search results and / or within a category-specific panel. The note search results 932 may be selectable to navigate to an immersive viewer 936 that displays an enlarged view of user-generated content.

[0108] Figure 9F shows different graphic card displays and / or note interface displays. The graphic card can be displayed in a single-width format 940 that is scrollable vertically, a two-width format 942 that is offset and scrollable vertically, a carousel interface 944 that is scrollable horizontally within another search result format, and / or a two-width format 946 that is aligned and scrollable vertically. The format may be based on the topic, type of interface, user preferences, and / or context.

[0109] FIG. 9G can show different customization options. For example, a graphics card customization interface can be provided to generate a graphics card, which can include text editing 952, layout editing 954, image editing 956, and / or other customization options. Specifically, the interactive interface can include multiple options (and / or features) for content generation. The multiple interface features can include text, images, audio, video, templates, and / or other input options. The multiple interface features can include one or more generation model interfaces for content proposal, template proposal, and / or generation model assisted generation (e.g., a large language model for rewriting text and / or generating text proactively, an image generation model for generating new images based on web resources, user input, and / or generated prompts, an audio generation model for generating narration, songs, and / or other audio, and / or a graphics card generation model for processing web resources, generated prompts, and / or user input to generate a graphics card that can be proposed to the user for link notes and / or stand-alone content). The multiple interface features can include customization options for customizing features of user-generated content such as layout, fonts (multiple), interface element sizes (multiple), images (multiple), text, transitions (multiple), tone, shadow, and / or others. The multiple interface features can include an option to add action user interface elements to the graphics card of user-generated content. The action user interface elements can include selectable options for performing one or more actions (e.g., API calls, navigation to different applications, search, content item generation using a generation model, etc.).

[0110] The systems and methods disclosed herein can include the proposal and / or generation of images for generating a graphics card. For example, the systems and methods can determine that an image from a database (e.g., a server database, a local database, and / or a user image gallery) is associated with a web resource, a prompt, and / or a link note. The image can then be provided as a proposal for use with the graphics card. Alternatively and / or additionally, the systems and methods can provide an image generation model (e.g., a text-to-image generation model) interface for generating an image to include in the graphics card. For example, the image generation model interface can be provided to a user, and the user may provide a prompt to the image generation model, and the image generation model can then generate a model-generated image that can be used with the graphics card.

[0111] In some embodiments, one or more trained machine learning models can be utilized to fact-check web resources and / or link notes. The one or more trained machine learning models can include one or more generation models that can utilize an application programming interface for API calls to obtain information and / or interact with other applications.

[0112] In some embodiments, one or more model-generated link notes can be generated that utilize a generative model to index web resources, provide examples of link notes, and / or provide semantic understanding notes. Alternatively and / or additionally, a generative model can be utilized to rewrite and / or propose link notes and / or sand-along content. The interactive user interface can include an interface for interacting with the generative model to generate content (e.g., text, image(s), and / or other data). The interactive user interface can include options for selecting other attributes for adjusting the generative model to generate content having a tone, style, format, lexicon, genre, and / or specific attributes. For example, the interactive user interface can be configured to generate a prompt for the generative model based on user input, link note prompts, and / or web resources.

[0113] The search results interface and / or discovery interface can provide statistics regarding a particular amount of searches, an amount of web resource selections, and / or a tendency of interactions of links and / or search queries.

[0114] In some embodiments, the system and method can include training and / or utilizing one or more contribution tendency models. The contribution tendency model can learn and / or determine user credibility (e.g., the relevant experience, expertise, and / or reliability of the user) for a particular user and / or a particular set of users. Additionally and / or alternatively, the contribution tendency model can learn and / or determine a tendency to provide link notes.

[0115] A contribution tendency model can be trained to detect the likelihood of contribution, credibility, note usefulness, and / or other attributes associated with a user, web resource, and / or context. The contribution tendency model can be trained on a labeled dataset, an unlabeled dataset, and / or a hybrid dataset. In some embodiments, the contribution tendency model can be trained with interaction data for learning a contribution prediction task, trained with the output from a verification model for a credibility determination task, and / or trained with a click-through rate for a usefulness determination task.

[0116] FIG. 10 shows a diagram of an exemplary card generation interface 1000 according to an exemplary embodiment of the present disclosure. Specifically, a user may select an option to generate a graphic card for a link note. Based on that selection, a template 1002 may be selected and provided for display. A particular template may be selected based on a web resource associated with the link note (e.g., based on the content of the web resource), based on user interaction history, based on query history, based on user profile data, and / or based on other data.

[0117] The card generation interface 1000 may include a pull-up menu 1004 associated with a plurality of proposed prompts for link note generation that may include topic ideas. The user may pull up the menu to provide an enlarged view 1006 of the proposed prompts. The enlarged view 1006 can include a plurality of selectable proposed prompts that can be selected to generate text, images, and / or layouts to be inserted into the graphic card. For example, a proposal such as "How to water Monstera like a pro" can be selected. Content items associated with the selected prompt proposal can be inserted into the graphic card, and the graphic card can transition to an editing interface 1008. The editing interface 1008 can include options for editing text, style, layout, font, color, and / or other edits.

[0118] FIG. 11 shows a diagram of an exemplary content item generation interface 1100 according to an exemplary embodiment of the present disclosure. Specifically, when modifying a graphic card template, one or more content items can be generated using an input draft interface.

[0119] For example, the user may select an option to open the content item generation interface 1100. At 1102, the user may select one or more attributes from a drop-down menu. The one or more attributes may be associated with the required attributes of the content item to be generated. The one or more attributes may be associated with the tone and / or style of the content. At 1104, the user may generate and / or provide text input. The text input can be associated with details of a topic, intention, information, and / or other prompts. The one or more attributes and the text input can be processed by a generation model to generate a model-generated content item. The model-generated content item can have one or more attributes and can be targeted at details of the topic, intention, information, and / or other prompts of the text input.

[0120] At 1106, the model-generated content item can be provided for display below the text input and can be provided with a plurality of options. The plurality of options can include editing of the one or more attributes, editing of the text input, reprocessing of data, saving of the mode-generated content item, exiting the interface, inserting the model-generated content item into a graphics card, and / or other options. At 1108, a modified graphics card can be provided for display together with the model-generated content item inserted into the graphics card based on user selection. Next, the user can edit the layout, size, color, font, and / or orientation of the model-generated content item and / or other content of the graphics card.

[0121] FIG. 12 shows a diagram of an exemplary image proposal interface 1200 according to an exemplary embodiment of the present disclosure. Specifically, the image proposal interface 1200 can acquire card data, context data, and / or input data. Next, the card data, context data, and / or input data can be processed to determine one or more images (and / or other media content items) to be provided as proposals for insertion into the graphics card.

[0122] For example, at 1202, a graphics card is provided for display with the option to insert additional text, stickers, and / or images. Next, the user may select an image addition option. At 1204, an image selection interface can be provided for display, and the image selection interface can include default images, camera roll images, and / or image proposals based on the text of the graphics card, the content of web resources associated with the link note, the user history, and / or other data. For example, a plurality of images from the user's image gallery can be determined to be related to the text of the graphics card based on the determination that an image is associated with a location (e.g., Mexico) referenced by the text of the graphics card. At 1206, the identified images can be provided for display for selection. The user may select a particular image from the identified images, and the image can be processed and inserted into the graphics card. At 1208, the selected image can be trimmed and inserted into the graphics card for display.

[0123] FIG. 13 shows a flowchart diagram of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. FIG. 13 shows steps performed in a particular order for purposes of illustration and description, but the methods of the present disclosure are not limited to the particular order or arrangement shown. The various steps of method 1300 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0124] In 1302, the computing system can obtain card data. The card data can describe the content within the graphics card. The content can be associated with one or more topics. The graphics card can be associated with a link note. The link note can include user-generated content tagged with specific web resources. The graphics card can include a background, one or more images, one or more text strings, and / or one or more user interface elements. The background can include a single color, multiple colors, an image, and / or other data. One or more user interface elements can include selectable widgets for providing additional information for display and / or for performing one or more actions. The content can include text data, image data, video data, potential encoding data, multimodal data, and / or other data.

[0125] In 1304, the computing system can process the card data to determine one or more entity tags associated with the content. The one or more entity tags can be associated with one or more topics. The card data can be processed by one or more machine-learned models (e.g., a generation model, a classification model, and / or other models) to generate entity tags. The entity tags can be associated with one or more objects, one or more companies, one or more locations, one or more individuals, one or more structures, and / or other entities.

[0126] At 1306, the computing system can access a media content item database to obtain one or more media content items. One or more media content items can be obtained based on determining that the one or more media content items are associated with one or more entity tags related to the content. The media content item database can include a user-specific database. In some embodiments, the user-specific database can be associated with a particular user. The particular user may have generated at least a portion of the content. The user-specific database can include an image gallery associated with the particular user. The image gallery can be stored in a server computing system associated with a particular content item storage platform. Alternatively and / or additionally, the user-specific database can include a local storage database of the user computing device. The media content item database can include a plurality of media content items. The plurality of content items may have been preprocessed to generate a plurality of respective metadata sets. Determining that one or more media content items are associated with one or more entity tags related to the content can include determining whether the one or more media content items include features associated with the entity tags. The features can be determined based on metadata, image processing, and / or other technologies. The one or more media content items can include one or more images, one or more videos, one or more animations, one or more audio files, and / or one or more other content items.

[0127] At 1308, the computing system can provide one or more media content items for display. In an interactive user interface, one or more media content items can be provided. One or more media content items can be selectable to be inserted into a graphics card. The interactive user interface can provide a plurality of media content items for display, which can include media content items related to the user, web media content items, and / or other media content items.

[0128] In some embodiments, the computing system can obtain an input selection associated with one or more media content items and generate an extended graphics card. The extended graphics card can include at least a portion of the content of the graphics card and at least a portion of one or more media content items. Next, the computing system can provide the extended graphics card for display.

[0129] Additionally or alternatively, the computing system can obtain an adjustment input. The adjustment input can be associated with a request to extend the extended graphics card. The computing system can generate an updated graphics card based on the adjustment input. The updated graphics card can include the extended graphics card with one or more adjustments. Next, the computing system can provide the updated graphics card for display. The one or more adjustments can include at least one of a layout change of the extended graphics card, a cropping change of one or more media content items, a size change of one or more content items, a color change, or a template change.

[0130] FIG. 14 shows a flowchart diagram of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. FIG. 14 shows steps performed in a particular order for purposes of illustration and explanation, but the method of the present disclosure is not limited to the particular order or arrangement shown. The various steps of method 1400 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0131] At 1402, the computing system can provide an input draft interface for display. The input draft interface can include a graphical user interface that includes a plurality of attribute options and a text input box. The plurality of attribute options can be associated with a plurality of candidate attributes for content item generation. The plurality of candidate attributes can include tone, style, length, content type, and / or other details. The input draft interface can include a preview window for viewing the current state of the graphics card. The graphics card can be associated with a link note. In some embodiments, the graphics card can include a card template that may have been changed based on one or more user inputs. For example, the user may have added an image, text, audio, video, widget, and / or other data.

[0132] At 1404, the computing system can obtain a selection of a particular attribute option among a plurality of attribute options via an input draft interface. The particular attribute option can be associated with a particular candidate attribute. In some embodiments, the plurality of candidate attributes can include a plurality of different styles. The plurality of different styles can be associated with at least one of a plurality of different artistic styles or a plurality of different writing styles. Alternatively and / or additionally, the plurality of candidate attributes can include a plurality of different tones. The plurality of different tones can be associated with at least one of a plurality of different emotions and / or a plurality of different pace types. The particular candidate attribute can include the tone and / or style required to generate the content item. The selection can be obtained based on the selection of a particular attribute option from a drop-down menu that provides the plurality of attribute options for display.

[0133] At 1406, the computing system can obtain text input via a text input box of the input draft interface. The text input can be associated with the prompt intent for content item generation. In some embodiments, the text input can be automatically input based on the content of the graphics card, based on the user context, and / or based on the prompt proposal.

[0134] At 1408, the computing system can process specific attribute options and text inputs using a generative model to generate model-generated content items. The mode-generated content items can include specific candidate attributes. The model-generated content items can be associated with a prompt intent. The model-generated content items can include text data, image data, audio data, multimodal data, and / or other data. In some embodiments, the generative model can be obtained from a generative model database based on the selection of specific attribute options. For example, the generative model database can store a plurality of different generative models associated with a plurality of candidate attributes. Each of the plurality of different generative models can be configured, trained, and / or adjusted to generate content items associated with their respective candidate attributes. Alternatively and / or additionally, the generative model can be a general generative model trained for a plurality of content generation tasks. Additionally or alternatively, a specific attribute soft prompt can be obtained based on the selection of specific attribute options. The specific attribute soft prompt can include a set of learned parameters. The set of learned parameters can be processed by the generative model to generate model-generated content items.

[0135] At 1410, the computing system can provide model-generated content items for display via an input draft interface. Providing the model-generated content items for display can include providing an option to insert the model-generated content items into a graphics card. The input draft interface can include a plurality of post-processing editing options, and the plurality of post-editing options can include options to change size, color, font, cropping, resolution, saturation, coloring, and / or other details.

[0136] In some embodiments, a computing system can obtain an input selection via an input draft interface and generate an extended graphics card based on the input selection. The extended graphics card can include a graphics card extended to include model generation content items. Next, the computing system can provide the extended graphics card for display.

[0137] FIG. 15A shows a block diagram of an exemplary computing system 100 for performing a link note prompt, according to an exemplary embodiment of the present disclosure. System 100 includes a user computing system 102, a server computing system 130, and / or a third computing system 150 communicatively coupled via a network 180.

[0138] The user computing system 102 can include any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0139] The user computing system 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 114 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing system 102 to perform operations.

[0140] In some embodiments, the user computing system 102 can store or include one or more machine-learned models 120. For example, the machine-learned model 120 can be various machine-learned models such as a neural network (e.g., a deep neural network), or other types of machine-learned models including non-linear models and / or linear models, or can include them otherwise. The neural network can include a feed-forward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks.

[0141] In some embodiments, one or more machine-learned models 120 can be received from server computing system 130 via network 180, stored in user computing device memory 114, and then used by one or more processors 112, or otherwise implemented. In some embodiments, user computing system 102 can implement multiple parallel instances of a single machine-learned model 120 (e.g., to perform parallel machine-learned model processing across multiple instances of input data and / or detected features).

[0142] More specifically, one or more machine-learned models 120 can include one or more detection models, one or more classification models, one or more segmentation models, one or more augmentation models, one or more generation models, one or more natural language processing models, one or more optical property recognition models, and / or one or more other machine-learned models. One or more machine-learned models 120 can include one or more transformer models. One or more machine-learned models 120 may include one or more neural radiance field models, one or more diffusion models, and / or one or more autoregressive language models.

[0143] One or more pre-trained machine learning models 120 can be utilized to detect features of one or more objects. The detected object features may be classified and / or embedded. Next, classification and / or embedding can be utilized to perform a search and determine one or more search results. Alternatively and / or additionally, one or more detected features can be utilized to determine whether an indicator (e.g., a user interface element indicating the detected feature) should be provided to indicate that the feature has been detected. Next, the user may select the indicator to cause classification, embedding, and / or search of the feature to be performed. In some embodiments, classification, embedding, and / or search may be performed before the indicator is selected.

[0144] In some embodiments, one or more pre-trained machine learning models 120 can process image data, text data, audio data, and / or latent encoding data to generate output data that can include image data, text data, audio data, and / or latent encoding data. One or more pre-trained machine learning models 120 may perform optical character recognition, natural language processing, image classification, object classification, text classification, audio classification, context determination, action prediction, image correction, image enhancement, text enhancement, sentiment analysis, object detection, error detection, inpainting, video stabilization, audio correction, audio enhancement, and / or data segmentation (e.g., mask-based segmentation).

[0145] Additionally or alternatively, one or more machine-learned models 140 can be included in a server computing system 130 that communicates with the user computing system 102 according to a client-server relationship, or can be stored and implemented in other ways. For example, the machine-learned model 140 can be implemented by the server computing system 130 as part of a web service (such as a viewfinder service, a visual search service, an image processing service, an ambient computing service, and / or an overlay application service). Thus, one or more models 120 can be stored and implemented in the user computing system 102, and / or one or more models 140 can be stored and implemented in the server computing system 130.

[0146] The user computing system 102 can also include one or more user input components 122 that receive user input. For example, the user input component 122 can be a touch sensor-based component (such as a touch sensor-based display screen or a touch pad) that is sensitive to a user input object (such as a finger or a stylus). The touch sensor-based component can function to implement a virtual keyboard. Other exemplary user input components include a microphone, a conventional keyboard, or other means by which a user can provide user input.

[0147] In some embodiments, the user computing system can store and / or provide one or more user interfaces 124 that may be associated with one or more applications. The one or more user interfaces 124 can be configured to receive input and / or provide data for display (e.g., image data, text data, audio data, one or more user interface elements, augmented reality experiences, virtual reality experiences, and / or other data for display). The user interface 124 may be associated with one or more other computing systems (e.g., server computing system 130 and / or third-party computing system 150). The user interface 124 can include a viewfinder interface, a search interface, a generative model interface, a social media interface, and / or a media content gallery interface.

[0148] The user computing system 102 may include one or more sensors 126 and / or may receive data from the sensors 126. The one or more sensors 126 may be housed in a housing component that houses one or more processors 112, a memory 114, and / or one or more hardware components, and the hardware components may store one or more software and / or enable the performance thereof. The one or more sensors 126 can include one or more image sensors (e.g., cameras), one or more lidar sensors, one or more audio sensors (e.g., microphones), one or more inertial sensors (e.g., inertial measurement units), one or more biological sensors (e.g., heartbeat sensors, pulse sensors, retinal sensors, and / or fingerprint sensors), one or more infrared sensors, one or more position sensors (e.g., GPS), one or more touch sensors (e.g., conductive touch sensors and / or mechanical touch sensors), and / or one or more other sensors. The one or more sensors can be utilized to obtain data related to the user's environment (e.g., an image of the user's environment, a record of the environment, and / or the user's location).

[0149] The user computing system 102 may include, and / or be part of, a user computing device 104. The user computing device 104 may include a mobile computing device (e.g., a smartphone or a tablet), a desktop computer, a laptop computer, smart wearables, and / or a smart appliance. Additionally and / or alternatively, the user computing system may obtain data from, and / or generate data using, one or more user computing devices 104. For example, the camera of a smartphone may be utilized to capture image data that describes the environment, and / or the overlay application of the user computing device 104 may be utilized to track and / or process data provided to the user. Similarly, one or more sensors associated with smart wearables may be utilized to obtain data regarding the user and / or the user's environment (e.g., image data may be obtained by a camera housed in the user's smart glasses). Additionally and / or alternatively, the data may be obtained and uploaded from other user devices that may be specialized for data acquisition or generation.

[0150] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor, or multiple processors operably connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0151] In some embodiments, the server computing system 130 includes, or is otherwise implemented by, one or more server computing devices. If the server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0152] As noted above, the server computing system 130 can store, or otherwise include, one or more machine-learned models 140. For example, the model 140 can be, or otherwise include, various machine-learned models. Exemplary machine-learned models include neural networks or other multi-layer non-linear models. Exemplary neural networks include feed-forward neural networks, deep neural networks, regression neural networks, and convolutional neural networks. An exemplary model 140 is described with reference to FIG. 15B.

[0153] Additionally and / or alternatively, the server computing system 130 can include, and / or be communicatively coupled to, a search engine 142 that can be utilized to crawl one or more databases (and / or resources). The search engine 142 can process data from the user computing system 102, the server computing system 130, and / or the third-party computing system 150 to determine one or more search results associated with the input data. The search engine 142 can perform term-based search, label-based search, boolean-based search, image search, embedding-based search (e.g., nearest neighbor search), multimodal search, and / or one or more other search techniques.

[0154] The server computing system 130 may store and / or provide one or more user interfaces 144 for obtaining input data and / or providing output data to one or more users. The one or more user interfaces 144 can include one or more user interface elements, which can include input fields, navigation tools, content tips, selectable tiles, widgets, data display carousels, dynamic animations, information pop-ups, image magnification, text-to-speech, speech-to-text, augmented reality, virtual reality, feedback loops, and / or other interface elements.

[0155] The user computing system 102 and / or the server computing system 130 can train the models 120 and / or 140 through interaction with a third-party computing system 150 communicatively coupled via the network 180. The third-party computing system 150 may be separate from the server computing system 130 or may be part of the server computing system 130. Alternatively and / or additionally, the third-party computing system 150 may be associated with one or more web resources, one or more web platforms, one or more other users, and / or one or more contexts.

[0156] The third-party computing system 150 may include one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 154 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the third-party computing system 150 to perform operations. In some embodiments, the third-party computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0157] The network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication via the network 180 can be performed via any type of wired and / or wireless connection using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, secure HTTP, SSL).

[0158] The machine-learned models described herein can be used in a variety of tasks, applications, and / or use cases.

[0159] In some embodiments, the input to the machine-learned model(s) of the present disclosure can be image data. The machine-learned model(s) can process the image data to generate an output. By way of example, the machine-learned model(s) can process the image data to generate an image recognition output (e.g., recognition of image data, potential embedding of image data, encoded representation of image data, hash of image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an image segmentation output. As another example, the machine-learned model(s) can process the image data to generate an image classification output. As another example, the machine-learned model(s) can process the image data to generate an image data modification output (e.g., modification of image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an encoded image data output (e.g., encoded and / or compressed representation of image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an upscaled image data output. As another example, the machine-learned model(s) can process the image data to generate a prediction output.

[0160] In some embodiments, the input to the machine-learned model(s) of the present disclosure can be text or natural language data. The machine-learned model(s) can process the text or natural language data to generate an output. By way of example, the machine-learned model(s) can process natural language data to generate a language encoding output. As another example, the machine-learned model(s) can process text or natural language data to generate a latent text embedding output. As another example, the machine-learned model(s) can process text or natural language data to generate a transformation output. As another example, the machine-learned model(s) can process text or natural language data to generate a classification output. As another example, the machine-learned model(s) can process text or natural language data to generate a text segmentation output. As another example, the machine-learned model(s) can process text or natural language data to generate a semantic output. As another example, the machine-learned model(s) can process text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data that is of higher quality than the input text or natural language). As another example, the machine-learned model(s) can process text or natural language data to generate a prediction output.

[0161] In some embodiments, the input to the machine-learned model(s) of the present disclosure can be audio data. The machine-learned model(s) can process the audio data to generate an output. As an example, the machine-learned model(s) can process the audio data to generate an audio recognition output. As another example, the machine-learned model(s) can process the audio data to generate an audio conversion output. As another example, the machine-learned model(s) can process the audio data to generate a potential embedding output. As another example, the machine-learned model(s) can process the audio data to generate an encoded audio output (e.g., an encoded and / or compressed representation of the audio data, etc.). As another example, the machine-learned model(s) can process the audio data to generate an upscaled audio output (e.g., audio data of higher quality than the input audio data, etc.). As another example, the machine-learned model(s) can process the audio data to generate a text representation output (e.g., a text representation of the input audio data, etc.). As another example, the machine-learned model(s) can process the audio data to generate a prediction output.

[0162] In some embodiments, the input to the machine-learned model(s) of the present disclosure can be sensor data. The machine-learned model(s) can process the sensor data to generate an output. As an example, the machine-learned model(s) can process the sensor data to generate a recognition output. As another example, the machine-learned model(s) can process the sensor data to generate a prediction output. As another example, the machine-learned model(s) can process the sensor data to generate a classification output. As another example, the machine-learned model(s) can process the sensor data to generate a segmentation output. As another example, the machine-learned model(s) can process the sensor data to generate a visualization output. As another example, the machine-learned model(s) can process the sensor data to generate a diagnostic output. As another example, the machine-learned model(s) can process the sensor data to generate a detection output.

[0163] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data of one or more images and the task is an image processing task. For example, the image processing task can be image classification, and the output is a set of scores, where each score corresponds to a different object class and represents the likelihood that one or more images depict an object belonging to that object class. The image processing task can be object detection, and the image processing output identifies one or more regions of one or more images and, for each region, the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, and the image processing output defines, for each pixel of one or more images, the respective likelihoods of each category within a given set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, and the image processing output defines, for each pixel of one or more images, a respective depth value. As another example, the image processing task can be motion estimation, the network input includes a plurality of images, and the image processing output defines, for each pixel of one of the input images, the motion of the scene represented by the pixels between the images within the network input.

[0164] The user computing system can include several applications (e.g., applications 1 - N). Each application can include its own respective machine learning library and one or more machine-learned models. For example, each application can include a machine-learned model. Exemplary applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, and the like.

[0165] Each application can communicate with some other components of a computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some embodiments, each application can communicate with each device component using an API (e.g., a public API). In some embodiments, the API used by each application is specific to that application.

[0166] The user computing system 102 can include several applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some embodiments, each application can communicate with the central intelligence layer (and the model(s) stored therein) using an API (e.g., a common API across all applications).

[0167] The central intelligence layer can include several machine-learned models. For example, each machine-learned model (e.g., a model) can be provided for each application and managed by the central intelligence layer. In other embodiments, two or more applications can share a single machine-learned model. For example, in some embodiments, the central intelligence layer can provide a single model (e.g., a single model) for all of the applications. In some embodiments, the central intelligence layer is included within the operating system of the computing system 100 or, otherwise, is implemented by the operating system of the computing system 100.

[0168] The Central Intelligence Layer can communicate with the Central Device Data Layer. The Central Device Data Layer may be a centralized repository of data for the computing system 100. The Central Device Data Layer may communicate with some other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some embodiments, the Central Device Data Layer can communicate with each device component using an API (e.g., a private API).

[0169] FIG. 15B shows a block diagram of an exemplary computing system 50 that executes a link note prompt, according to an exemplary embodiment of the present disclosure. Specifically, the exemplary computing system 50 can obtain and / or generate one or more data sets that can be processed by a sensor processing system 60 and / or an output determination system 80, and can include one or more computing devices 52 that can be utilized to provide feedback to a user regarding information about the characteristics of the one or more obtained data sets. The one or more data sets can include image data, text data, audio data, multimodal data, latent encoding data, and the like. The one or more data sets can be obtained via one or more sensors associated with one or more computing devices 52 (e.g., one or more sensors of the computing device 52). Additionally and / or alternatively, the one or more data sets can be stored data and / or retrieved data (e.g., data retrieved from a web resource). For example, an image, text, and / or other content item can be interacted with by a user. Next, one or more decisions can be generated using the item interacted with the content.

[0170] One or more computing devices 52 can obtain and / or generate one or more data sets based on image capture, sensor tracking, data storage retrieval, content downloading (e.g., downloading an image or other content item from a web resource over the Internet), and / or via one or more other techniques. The one or more data sets can be processed by a sensor processing system 60. The sensor processing system 60 may perform one or more processing techniques using one or more machine-learned models, one or more search engines, and / or one or more other processing techniques. The one or more processing techniques may be performed in any combination and / or individually. The one or more processing techniques may be performed sequentially and / or in parallel. In particular, the one or more data sets can be processed by a context determination block 62, which may determine the context associated with one or more content items. The context determination block 62 can identify and / or process metadata, user profile data (e.g., preferences, user search history, user browsing history, user purchase history, and / or user input data), previous interaction data, global trend data, location data, time data, and / or other data to determine a particular context related to the user. The context can be associated with an event, a determined trend, a particular action, a particular type of data, a particular environment, and / or other context associated with the user and / or the retrieved or obtained data.

[0171] The sensor processing system 60 may include an image preprocessing block 64. The image preprocessing block 64 can be used to adjust one or more values of the acquired image and / or the received image to prepare an image to be processed by one or more trained machine learning models and / or one or more search engines 74. The image preprocessing block 64 may resize the image, adjust saturation values, adjust resolution, remove and / or add metadata, and / or perform one or more other operations.

[0172] In some embodiments, the sensor processing system 60 can include one or more trained machine learning models, which may include a detection model 66, a segmentation model 68, a classification model 70, an embedding model 72, and / or one or more other trained machine learning models. For example, the sensor processing system 60 may include one or more detection models 66 that can be used to detect specific features of a processed data set. Specifically, one or more images can be processed by one or more detection models 66 to generate one or more bounding boxes associated with the features detected in the one or more images.

[0173] Additionally and / or alternatively, one or more segmentation models 68 can be used to segment one or more parts of a data set from one or more data sets. For example, one or more segmentation models 68 can use one or more segmentation masks (e.g., one or more segmentation masks generated manually and / or based on one or more bounding boxes) to segment a part of an image, a part of an audio file, and / or a part of text. Segmentation may include separating one or more detected objects and / or removing one or more detected objects from an image.

[0174] One or more classification models 70 can be utilized to process image data, text data, audio data, latent encoding data, multimodal data, and / or other data to generate one or more classifications. The one or more classification models 70 can include one or more image classification models, one or more object classification models, one or more text classification models, one or more audio classification models, and / or one or more other classification models. The one or more classification models 70 can process data to determine one or more classifications.

[0175] In some embodiments, the data may be processed by one or more embedding models 72 to generate one or more embeddings. For example, one or more images can be processed by one or more embedding models 72 to generate one or more image embeddings in an embedding space. The one or more image embeddings can be associated with one or more features of the one or more images. In some embodiments, the one or more embedding models 72 can be configured to process multimodal data to generate multimodal embeddings. The one or more embeddings can be utilized for classification, retrieval, and / or learning of embedding space distribution.

[0176] The sensor processing system 60 may include one or more search engines 74 that can be used to perform one or more searches. The one or more search engines 74 can crawl one or more databases (e.g., one or more local databases, one or more global databases, one or more private databases, one or more public databases, one or more dedicated databases, and / or one or more general databases) to determine one or more search results. The one or more search engines 74 can perform feature matching, text-based search, embedding-based search (e.g., k-nearest neighbor search), metadata database search, multimodal search, web resource search, image search, text search, and / or application search.

[0177] Additionally and / or alternatively, the sensor processing system 60 may include one or more multimodal processing blocks 76 that can be used to assist in processing multimodal data. The one or more multimodal processing blocks 76 may include generating multimodal queries and / or multimodal embeddings that are processed by one or more machine-learned models and / or one or more search engines 74.

[0178] The output(s) of the sensor processing system 60 can then be processed by an output determination system 80 to determine one or more outputs to provide to the user. The output determination system 80 may include heuristic-based determination, machine-learned model-based determination, user-selection-based determination, and / or context-based determination.

[0179] The output determination system 80 may determine a method and / or location for providing one or more search results in the search result interface 82. Additionally and / or alternatively, the output determination system 80 may determine a method and / or location for providing the output of one or more machine-learned models in the machine-learned model output interface 84. In some embodiments, one or more search results and / or the output of one or more machine-learned models may be provided for display via one or more user interface elements. The one or more user interface elements may be overlaid on the displayed data. For example, one or more detection indicators may be overlaid on the detected objects in the viewfinder. The one or more user interface elements may be selectable to perform one or more additional searches and / or one or more additional machine-learned model processes. In some embodiments, the user interface elements may be provided as user interface elements dedicated to a particular application and / or may be provided uniformly across different applications. The one or more user interface elements can include a pop-up display, an interface overlay, an interface style and / or chip, a carousel interface, an audio feedback, an animation, an interactive widget, and / or other user interface elements.

[0180] Additionally and / or alternatively, an augmented reality experience and / or a virtual reality experience 86 may be generated and / or provided using data associated with the output(s) of the sensor processing system 60. For example, one or more acquired data sets may be processed to generate one or more augmented reality rendering assets and / or one or more virtual reality rendering assets, which may then be used to provide an augmented reality experience and / or a virtual reality experience 86 to a user. The augmented reality experience may render information related to the environment into each environment. Additionally and / or alternatively, objects associated with the processed data set(s) may be rendered within the user environment and / or virtual environment. Generation of the rendering data set may include training one or more neural radiance field models to learn a three-dimensional representation of one or more objects.

[0181] In some embodiments, one or more action prompts 88 may be determined based on the output(s) of the sensor processing system 60. For example, a search prompt, a purchase prompt, a generation prompt, a reservation prompt, a call prompt, a redirect prompt, and / or one or more other prompts may be determined to be associated with the output(s) of the sensor processing system 60. Next, the one or more action prompts 88 may be provided to the user via one or more selectable user interface elements. In response to selection of the one or more selectable user interface elements, the respective action of each action prompt may be performed (e.g., a search may be performed, a purchase application programming interface may be utilized, and / or other applications may be launched).

[0182] In some embodiments, one or more datasets and / or outputs of the sensor processing system 60 may be processed by one or more generative models 90 to generate model-generated content items, which may then be provided to the user. Generation may be prompted based on user selection and / or may be performed automatically (e.g., automatically performed based on one or more conditions that may be associated with an identified threshold amount of search results).

[0183] One or more generative models 90 can include language models (e.g., large language models and / or vision-language models), image generation models (e.g., text-to-image generation models and / or image enhancement models), audio generation models, video generation models, graph generation models, and / or other data generation models (e.g., other content generation models). One or more generative models 90 can include one or more transformer models, one or more convolutional neural networks, one or more recurrent neural networks, one or more feedforward neural networks, one or more generative adversarial networks, one or more self-attention models, one or more embedding models, one or more encoders, one or more decoders, and / or one or more other models. In some embodiments, one or more generative models 90 can include one or more autoregressive models (e.g., machine-learned models trained to generate predicted values based on previous behavior data) and / or one or more diffusion models (e.g., machine-learned models trained to generate predicted data based on generating and processing distribution data associated with input data).

[0184] One or more generative models 90 can be trained to process input data and generate model-generated content items that can include multiple predicted words, pixels, signals, and / or other data. The model-generated content items can include new content items that are not the same as any existing works. One or more generative models 90 can utilize learned representations, sequences, and / or probability distributions to generate content items, which can include phrases, storylines, settings, objects, characters, beats, lyrics, and / or other aspects not included in existing content items.

[0185] One or more generative models 90 may include a vision-language model.

[0186] The vision-language model can be trained, tuned, and / or configured to process image data and / or text data to generate natural language output. The vision-language model can utilize a pre-trained large language model (e.g., a large autoregressive language model) with one or more encoders (e.g., one or more image encoders and / or one or more text encoders) to provide fine-grained natural language output that emulates human-generated natural language.

[0187] The vision-language model can be utilized for zero-shot image classification, few-shot image classification, image captioning, multimodal query distillation, multimodal question answering, and / or can be tuned and / or trained for multiple different tasks. The vision-language model can perform visual question answering, image caption generation, feature detection (e.g., content monitoring (such as inappropriate content)), object detection, scene recognition, and / or other tasks.

[0188] A vision-language model can utilize a pre-trained language model and then adjust the language model for multimodality. Training and / or adjustment of the vision-language model can include image-text matching, masked language modeling, multimodal fusion with cross-attention, contrastive learning, prefix language model training, and / or other training techniques. For example, the vision-language model can be trained to process an image to generate predictive text similar to ground truth text data (e.g., a ground truth caption of the image). In some embodiments, the vision-language model can be trained to replace masked tokens of a natural language template with text tokens that describe features depicted in the input image. Alternatively and / or additionally, training, adjustment, and / or model estimation may include multi-layer concatenation of visual and text embedding features. In some embodiments, the vision-language model can be trained and / or adjusted by jointly learning the generation of image embeddings and text embeddings, which can include training and / or adjusting a system that maps embeddings to a shared feature embedding space that maps text features and image features to the shared embedding space. Joint training can include parallel embeddings of image-text pairs and / or can include triplet training. In some embodiments, an image can be utilized and / or processed as a prefix of a language model.

[0189] The output determination system 80 can process the output(s) of one or more datasets and / or the sensor processing system 60 using the data augmentation block 92 to generate augmented data. For example, one or more images can be processed by the data augmentation block 92 to generate one or more augmented images. Data augmentation can include data correction, data cropping, removal of one or more features, addition of one or more features, resolution adjustment, lighting adjustment, saturation adjustment, and / or other augmentations.

[0190] In some embodiments, one or more datasets and / or outputs(s) of the sensor processing system 60 may be stored based on a determination of the data storage block 94.

[0191] Next, the output(s) of the output determination system 80 may be provided to the user via one or more output components of the user computing device 52. For example, one or more user interface elements associated with the one or more outputs may be provided for display via the visual display of the user computing device 52.

[0192] The process may be performed iteratively and / or continuously. One or more user inputs to the provided user interface elements may condition and / or affect the continuous processing loop.

[0193] The techniques described herein refer to servers, databases, software applications, and other computer-based systems, as well as the actions performed and the information sent to and sent from such systems. The flexibility inherent in computer-based systems allows for a wide variety of configurations, combinations, and divisions of tasks and functions among components. For example, the processes discussed herein can be implemented using a single device or component, or multiple devices or components operating in combination. Databases and applications can be implemented in a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0194] Although the subject matter of the present disclosure has been described in detail with respect to its various specific and exemplary embodiments, each example is provided for illustrative purposes and is not intended to limit the present disclosure. Those skilled in the art, upon understanding the foregoing, can readily make changes and modifications to such embodiments and easily create equivalents. Accordingly, the present disclosure does not exclude including such modifications, changes, and / or additions to the subject matter, as would be readily apparent to those skilled in the art. For example, features illustrated or described as part of one embodiment can be used with other embodiments to create still further embodiments. Accordingly, the present disclosure is intended to cover such changes, modifications, and equivalents.

Claims

1. 1. A computing system for comment prompt generation and input retrieval, the system comprising: one or more processors; and one or more non-transitory computer-readable storage media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations including: Obtaining content data, the content data being associated with a web resource; processing the content data with a generative model to generate predictive prompts, the prompts including predictive text strings associated with commenting on the web resource; providing the predictive prompts for display on an input prompt interface, the input prompt interface configured to receive input; obtaining comment input data from a user computing system via the input prompt interface, the comment input data including user-generated comments on the web resource; storing data associated with the comment entry data together with data associated with the web resource, the data associated with the comment entry data being stored in a searchable database that is provided for display in response to the web resource being provided as a search result; A computing system comprising:

2. The operation includes: obtaining user data, the user data being associated with a particular user and the user computing system being associated with the particular user; The system of claim 1 , wherein processing the content data with the generative model to generate the predictive prompt comprises processing the content data and the user data with the generative model.

3. 3. The system of claim 2, wherein the user data includes user search history data, and the generative model generates the predictive prompts based on the particular user having previously searched for information associated with a topic of the web resource.

4. 3. The system of claim 2, wherein the user data includes user browser history data, and the generative model generates the predictive prompts based on the particular user's previous viewing of other web resources that contain information associated with a topic of the web resource.

5. The operation includes: generating a graphics card based on the user data, the content data, and the comment input data, the graphics card including a user profile identifier for the particular user and data associated with the comment input data; storing said graphics card; The system of claim 2 further comprising:

6. The system of claim 5 , wherein the graphics card includes a graphical background generated with an image generation model based on the comment input data.

7. The operation includes: Obtaining a search query; and determining that the web resource is associated with the search query; providing specific search results for display, the specific search results including a link to the web resource, a title of the web resource, and data associated with the comment input data; The system of claim 1 further comprising:

8. Storing the data associated with the comment entry data with the data associated with the web resource includes: Generating a web resource note; and and storing the web resource note with a plurality of other web resource notes associated with the web resource.

9. The operation, providing the web resource note and the plurality of other web resource notes to a note interface that provides the web resource note and the plurality of other web resource notes to a plurality of graphic cards; The system of claim 8 further comprising:

10. The system of claim 1 , wherein the generative model comprises an autoregressive language model, and the generative model is prompted to generate a question that describes a request for information about the web resource.

11. 1. A computer-implemented method for linked note prompts, the method comprising: obtaining, by a computing system including one or more processors, context data associated with a particular content presentation instance, the particular content presentation instance including a particular user viewing a particular content item; determining, by the computing system, a prompting action based on the context data, the prompting action including providing an input entry interface to a user for obtaining user input; processing, by the computing system, the contextual data with a generative language model to generate a predictive prompt, the predictive prompt including a natural language request for information generated based on the contextual data; providing, by the computing system, the predictive prompt in the input entry interface; obtaining, by the computing system, user-generated content via the input entry interface; generating, by the computing system, link notes based on the user-generated content, the link notes being generated such that they are provided for display in a search result interface in response to the particular content item being determined as a search result. method.

12. The method of claim 11 , wherein the contextual data is associated with a type of content being provided for display.

13. The method of claim 11 , wherein the contextual data is associated with the particular user associated with the particular content display instance, and the contextual data includes search history data.

14. 12. The method of claim 11, wherein the content being offered for display is associated with a particular web resource, the contextual data is associated with interaction data of links of the particular web resource across multiple social network platforms, and the prompt action is determined based on the interaction data.

15. 12. The method of claim 11, wherein the contextual data includes user data and content data, and the prompt action is determined based on a topic associated with content being offered for display being one of a plurality of topics about which the particular user is determined to have knowledge based on the user data.

16. The method of claim 11 , wherein the contextual data includes previous notes generated by the particular user, and the predictive prompts include a structure based on previous structures for the previous notes.

17. One or more non-transitory computer-readable storage media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations including: Obtaining a first search query at a first time; determining that a web resource is responsive to the first search query; obtaining content data, the content data being associated with the web resource; processing the content data with a generative model to generate predictive prompts, the prompts including predictive text strings associated with commenting on the web resource; providing the predictive prompt for display within an input prompt interface, the input prompt interface including an input entry box; obtaining comment input data from a user computing system via the input prompt interface, the comment input data including user generated content; storing the user generated content; obtaining a second search query at a second time, the second time being different from the first time; determining that the web resource is responsive to the second search query; providing data describing the web resource to the user-generated content of a search result interface; One or more non-transitory computer-readable storage media.

18. 20. The one or more non-transitory computer-readable storage media of claim 17, wherein the first search query and the second search query are different.

19. 20. The one or more non-transitory computer-readable storage media of claim 17, wherein the comment input data comprises multi-modal data.

20. 20. The one or more non-transitory computer-readable storage media of claim 19, wherein the multi-modal data comprises textual data and image data.

Citation Information

Patent Citations

  • Generating action elements suggesting content for ongoing tasks

    JP2023008982A

  • Engaging users by personalized composing-content recommendation

    US11010436B1

  • System and method for managing search results

    US8214380B1