Generating segment groups based on the selection of parts of a web page
By generating and storing graphic cards with segment groups, including address and location data, it solves the problem that users find it difficult to quickly locate and save web page content, realizes efficient access and low-cost sharing, and improves the computing efficiency and user experience of the computing system.
Patent Information
- Application Number
- CN202380041117.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-04-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-04-25
AI Technical Summary
In the prior art, when saving web page content, it is difficult for users to quickly locate and access the original location of the saved data, and they need to find the data source through search queries or browsing history, resulting in inconvenience in use.
By generating segment groupings, including address and location data of content items, stored in the form of a graphical card, users can easily access and navigate to specific parts of the web page, using a graphical user interface and machine learning model to process gesture data to determine selected content items.
It realizes rapid positioning and accessing saved web page content, reduces transmission costs, and improves the computing efficiency and user experience of the computing system.
Smart Images

Figure CN119234215B_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims the benefit and priority of U.S. Non - Provisional Patent Application No. 18 / 081,814, filed on Dec. 15, 2022, which claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 344,783, filed on May 23, 2022. U.S. Non - Provisional Patent Application No. 18 / 081,814 and U.S. Provisional Patent Application No. 63 / 344,783 are hereby incorporated by reference in their entirety. Technical Field
[0003] The present disclosure generally relates to generating interactive segment groupings in response to user input. More specifically, the present disclosure relates to obtaining user input to select content items to be saved with a segment grouping, which can later be selected to provide a portion of a web page that includes the content items. Background Art
[0004] Saving text, images, and / or audio from a web page allows a user to experience the text, images, and / or audio locally again without being connected to the Internet. However, the saving process may provide limited context to the data source, and in instances where the user wishes to view the context in which the saved data was obtained, the user must use the data as a search query, navigate their browsing history, or try to remember how they initially accessed the web page. Additionally, when the web page source is found, the user may still need to examine much of the web page content to find exactly where in the web page the saved data originally originated. Summary of the Invention
[0005] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.
[0006] One example aspect of the present disclosure relates to a computing system. The system may include: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations may include providing data that describes a graphical user interface. The graphical user interface may include a graphical window for displaying a web page. In some implementations, the web page may include a plurality of content items. The operations may include obtaining input data. The input data may include a request to save one or more of the plurality of content items. The operations may include generating a fragment grouping. In some implementations, the fragment grouping may include one or more content items. The fragment grouping may include address data. The address data may describe the web address of the web page. The fragment grouping may include location data. The location data may describe the location of one or more content items within the web page. The operations may include storing the fragment grouping. The fragment grouping may be associated with a particular user.
[0007] In some implementations, the operations may include receiving a fragment request that provides a fragment interface. The fragment interface may include interactive elements associated with the fragment grouping. The operations may include providing the fragment interface for display, receiving an interface selection that selects an interactive element, and providing a portion of the web page for display. The portion of the web page may include the location of one or more content items within the web page. In some implementations, the operations may include receiving an insertion input. The insertion input may include a user input that requests inserting one or more content items into a different interface. The operations may include providing the fragment grouping to a third-party server computing system.
[0008] In some implementations, generating the fragment grouping may include obtaining one or more content items and generating a graphic card. The graphic card may describe the one or more content items. Generating the fragment grouping may include storing the graphic card as a graphical representation of the fragment grouping. In some implementations, the location data may include at least one of a scroll position, a start node, or an end node. The scroll position may describe the location of one or more content items relative to other portions of the web page. The start node may describe the location where one or more content items begin. The end node may describe the location where one or more content items end.
[0009] In some implementations, the one or more content items may include at least one of an image, a video, a graphical description of a product, or an audio. The operations may include processing the one or more content items to determine an entity associated with the one or more content items and generating an entity label based on the entity. The fragment grouping may include the entity label. In some implementations, storing the fragment grouping may include locally storing the fragment grouping on a mobile computing device.
[0010] Another example aspect of the present disclosure relates to a computer-implemented method. The method may include obtaining, by a computing system including one or more processors, input data. The input data may describe a selection of a content item associated with a segment grouping. The method may include obtaining, by the computing system, address data and location data associated with the segment grouping. The address data may be associated with a web page. The content item may be associated with the web page. In some implementations, the location data may describe the location of the content item within the web page. The method may include obtaining, by the computing system, web page data. The web page data may be obtained at least in part based on the address data. The method may include determining, by the computing system, the location within the web page associated with the content item and providing, by the computing system, a portion of the web page. The portion of the web page may include the location of the content item.
[0011] In some implementations, the address data may include a uniform resource locator. The location data may include a text snippet. Determining, by the computing system, the location within the web page associated with the content item may include adding, by the computing system, the text snippet to the uniform resource locator to generate a shortcut link and entering, by the computing system, the shortcut link into a browser.
[0012] In some implementations, the segment grouping may be generated by: providing, by the computing system, a graphical user interface. The graphical user interface may include a graphical window for displaying a web page. The web page may include a plurality of content items. The segment grouping may be generated by: obtaining, by the computing system, selection data; generating, by the computing system, the segment grouping; and storing, by the computing system, the segment grouping in a user database. The selection data may include a request to save a content item among the plurality of content items. In some implementations, providing, by the computing system, a portion of the web page may include providing one or more indicators to the portion of the web page. The one or more indicators may indicate the content item associated with the segment grouping. The one or more indicators may include highlighted text associated with the content item. In some implementations, the segment grouping may be associated with a user account of a particular user. The user account may be associated with one or more platforms.
[0013] Another example aspect of the present disclosure relates to one or more non-transitory computer-readable media jointly storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations may include providing a graphical user interface. The graphical user interface may include a graphical window for displaying a web page. In some implementations, the web page may include a plurality of content items. The operations may include receiving gesture data. In some implementations, the gesture data may describe a gesture associated with a portion of the web page. The operations may include processing the gesture data to determine the selected content item. The selected content item may be associated with the portion of the web page. The operations may include generating a segment grouping based on the gesture data. In some implementations, the segment grouping may include the selected content item.
[0014] In some implementations, the fragment grouping may include address data and location data. The address data may be associated with a web page. The location data may describe the location of the selected content item within the web page. The gesture may include a circular gesture that encloses a portion of the web page. In some implementations, processing the gesture data to determine the selected content item may include: determining the portion of the web page enclosed by the gesture; determining the focus of that portion; and determining that the selected content item is associated with the focus of that portion. In some implementations, processing the gesture data to determine the selected content item may include processing the gesture data with a machine learning model to determine the selected content item. The gesture data may describe a touch input to a touch screen display of a mobile computing device.
[0015] Other aspects of the present disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.
[0016] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] With reference to the drawings, a detailed discussion of the embodiments directed to those of ordinary skill in the art is set forth in this specification, in which:
[0018] Figure 1A A block diagram of an example computing system that performs fragment grouping generation in accordance with an example embodiment of the present disclosure is depicted.
[0019] Figure 1B A block diagram of an example computing device that performs fragment grouping generation in accordance with an example embodiment of the present disclosure is depicted.
[0020] Figure 1C A block diagram of an example computing device that performs fragment grouping generation in accordance with an example embodiment of the present disclosure is depicted.
[0021] Figure 2 An illustration of an example fragment grouping generation interface in accordance with an example embodiment of the present disclosure is depicted.
[0022] Figure 3 An illustration of an example gesture interaction in accordance with an example embodiment of the present disclosure is depicted.
[0023] Figure 4 An illustration of an example fragment grouping generation interface in accordance with an example embodiment of the present disclosure is depicted.
[0024] Figure 5Depicts an illustration of an example segment grouping generation and favorite addition interface according to an example embodiment of the present disclosure.
[0025] Figure 6 Depicts a flowchart of an example method for performing segment grouping generation according to an example embodiment of the present disclosure.
[0026] Figure 7 Depicts a flowchart of an example method for performing segment grouping interaction according to an example embodiment of the present disclosure.
[0027] Figure 8 Depicts a flowchart of an example method for performing segment grouping generation based on gesture input according to an example embodiment of the present disclosure.
[0028] Figure 9 Depicts an illustration of an example favorite addition interface according to an example embodiment of the present disclosure.
[0029] Figure 10A Depicts an illustration of an example segment grouping interaction according to an example embodiment of the present disclosure.
[0030] Figure 10B Depicts an illustration of an example segment grouping search according to an example embodiment of the present disclosure.
[0031] Figure 10C Depicts an illustration of an example highlighting interface according to an example embodiment of the present disclosure.
[0032] Figure 11 Depicts an illustration of an example graphics card according to an example embodiment of the present disclosure.
[0033] Figure 12 Depicts an illustration of an example summary segment grouping generation interface according to an example embodiment of the present disclosure.
[0034] Figure 13 Depicts an illustration of an example segment grouping suggestion interface according to an example embodiment of the present disclosure.
[0035] Figure 14 Depicts an illustration of an example segment grouping generation and sharing interaction according to an example embodiment of the present disclosure.
[0036] Figure 15 Depicts an illustration of an example graphics card customization interface according to an example embodiment of the present disclosure.
[0037] Figure 16 Depicts a block diagram of an example segment grouping generation system according to an example embodiment of the present disclosure.
[0038] Reference numerals that are repeated across multiple figures are intended to identify the same features in various implementations. Detailed implementation manners
[0039] Overview
[0040] Generally speaking, the present disclosure relates to generating interactive segment groups in response to user input. More specifically, the present disclosure relates to obtaining user input to select content items to be saved together with a segment group, and the segment group can be later selected to provide a portion of a web page including the content items. For example, a user can select a portion of a web page and / or a data file. Then, the data describing the selected portion can be stored together with the information associated with the web page / data file and the position of the portion relative to the web page / data file. The stored data set can be a segment group including a graphical representation of the selected portion, and the graphical representation can include a graphic card having text and / or images from the selected portion. The segment group can be stored for later reference and / or can be shared with other users. The segment group can enable a user to view the selected portion and then navigate to a specific position of the selected portion in the original web page / data file after selection. These systems and methods can include providing a graphical user interface. The graphical user interface can include a graphical window for displaying a web page. In some implementations, the web page can include multiple content items. These systems and methods can include obtaining input data. The input data can include a request to save one or more of the multiple content items. A segment group can be generated. The segment group can include one or more content items, address data, and location data. In some implementations, the address data can describe the web address of the web page. The location data can describe the position of one or more content items within the web page. The segment group can be stored in a user database.
[0041] For example, the systems and methods disclosed herein can include providing a graphical user interface. The graphical user interface can include a graphical window for displaying a web page. In some implementations, the web page can include multiple content items. The displayed web page can include text, one or more images, one or more interactive user interface elements, one or more videos, and / or one or more audio clips. The graphical user interface can be part of a browser application.
[0042] These systems and methods can obtain input data. The input data can include a request to save one or more of a plurality of content items. In some implementations, one or more content items can include at least one of an image, a video, a graphical description of a product, or audio. The graphical user interface can be updated to include one or more user interface elements for interacting with one or more content items, and the one or more user interface elements can include saving one or more content items and can include different options for the saved formatting. Alternatively and / or additionally, a pop-up interface can be provided in response to the request. The input data can describe a gesture associated with one or more content items (e.g., a circle around one or more content items).
[0043] A segment grouping can be generated. The segment grouping can be generated based on the input data. The segment grouping can include one or more content items, address data, and location data. Generating the segment grouping can include processing one or more content items with one or more machine learning models to generate a semantic understanding output that can be used for summarizing, annotating, and / or classifying.
[0044] One or more content items can include text data (e.g., text data describing words, sentences, quotes, and / or paragraphs), image data (e.g., an image of an object (e.g., a product)), video data (e.g., a video and / or one or more frames of a video), audio data (e.g., waveform data), and / or latent encoded data. One or more content items can include multimodal data. The address data can include a resource locator associated with a web page (e.g., a URL) and / or a file address. The location data can describe the location of one or more content items within a web page and / or a file. For example, the location data can indicate the start and end of one or more content items within a web page and / or a file.
[0045] In some implementations, generating the segment grouping can include obtaining one or more content items and generating a graphic card. The graphic card can describe one or more content items. The graphic card can include text data overlaid on a color and / or an image. The color can be determined based on the primary color of the web page. In some implementations, the color can be predefined and / or can be determined based on surrounding content items. The image can be an image determined based on the identified theme of the content item. Alternatively and / or additionally, the image can be an image close to the content item. In some implementations, the graphic card can include a font determined based on the font used in the web page. The graphic card can include a background, text describing the content item (e.g., the content item and / or a summary of the content item), and / or text and / or a logo describing the source and / or the identified entity. In some implementations, the text size of the text in the graphic card can be based on the amount of text in the content item.
[0046] Address data may describe the web address of a web page. The address data may include a Uniform Resource Identifier and / or a Uniform Resource Locator. In some implementations, the address data may include data that describes the source of a content item.
[0047] Location data may describe the location of one or more content items within a web page. In some implementations, the location data may include at least one of a scroll position, a start node, or an end node. The scroll position may describe the location of one or more content items relative to other parts of the web page. In some implementations, the start node may describe the location where one or more content items begin. The end node may describe the location where one or more content items end. The location data may include a text fragment (Tomayac et al., "Scroll to TextFragment", published on May 20, 2022, 9:40 PM, GITHUB, https: / / github.com / WICG / scroll-to-text-fragment), which may be used to indicate the location of a content item. In some implementations, the text fragment may include one or more text instructions associated with the location. The text instructions may include a string of start data (e.g., the first text associated with the content item and / or the first pixel associated with the content item) and / or end data (e.g., the last text associated with the content item and / or the last pixel associated with the content item). The text instructions may be used to search for a set of data in the web page that matches the start and / or end of the content item.
[0048] In some implementations, these systems and methods may include processing one or more content items to determine an entity associated with the one or more content items. An entity label may be generated based on the entity. The fragment grouping may include the entity label.
[0049] The fragment grouping may be stored in a user database. In some implementations, storing the fragment grouping in a user database may include locally storing the fragment grouping on a mobile computing device. Additionally and / or alternatively, a graphics card may be stored as a graphical representation of the fragment grouping. The graphics card may be automatically generated and may be user-customizable. The graphics card may start from a template that may be customized based on other content in the web page, based on user input, and / or based on other contexts. In some implementations, the graphics card may include multimodal data. Alternatively and / or additionally, the fragment grouping may be stored on a server computing system. In some implementations, if a content item references another content item, these systems and methods may obtain the additional content item and save the additional content item in the fragment grouping.
[0050] In some implementations, these systems and methods may include receiving a fragment request that provides a fragment interface. The fragment interface may include interactive elements associated with a fragment group. The fragment interface may be provided for display. An interface selection may be received, and the interface selection may describe a selection of an interactive element. Then, a portion of a web page may be provided for display. In some implementations, the portion of the web page may include the location of one or more content items within the web page.
[0051] In some implementations, these systems and methods may include receiving an insertion input. The insertion input may include user input requesting that a content item be inserted into a different interface. The fragment group may be provided to a third-party server computing system.
[0052] Additionally and / or alternatively, these systems and methods may include adding the fragment group to a favorites list. The fragment group may be added to the favorites list based on the received user input. Alternatively and / or additionally, the fragment group may be automatically added to the favorites list based on the determined entity associated with the content item. In some implementations, the fragment group may be added to the favorites list based on the source of the content item (e.g., the type of web page, the type of media provider, and / or the type of content item).
[0053] In some implementations, a fragment group may be generated based on content items obtained from sources other than a web page (e.g., a mobile application, a large data file (e.g., a downloaded video or book), and / or another data source).
[0054] These systems and methods may include providing a specific portion of a web page for display in response to an interaction with the fragment group. For example, these systems and methods may include obtaining input data. The input data may describe a selection of a content item associated with the fragment group. Address data and location data associated with the fragment group may be obtained. The address data may be associated with the web page. The content item may be associated with the web page. Additionally and / or alternatively, the location data may describe the location of the content item within the web page. These systems and methods may obtain web page data. The web page data may be obtained at least in part based on the address data. The location within the web page associated with the content item may be determined. Then, a portion of the web page may be provided. The portion of the web page may include the location of the content item.
[0055] These systems and methods can obtain input data. In some implementations, the input data can describe a selection of content items associated with a segment grouping. The segment grouping can be associated with a user account of a particular user. Additionally and / or alternatively, the user account can be associated with one or more platforms. The segment grouping can include content items and deep links. Alternatively and / or additionally, the segment grouping can include segments (e.g., content items and / or media data generated based on content items (e.g., graphics cards and / or summaries of content items)), address data (e.g., uniform resource locators and / or uniform resource identifiers), and / or metadata. The segment grouping can include location data (e.g., metadata indicating the location of a content item within a web page, text segments for identifying start data and end data, one or more pointers, and / or one or more scroll position data).
[0056] In some implementations, a segment grouping can be generated by providing a graphical user interface for display. The graphical user interface can include a graphical window for displaying a web page. In some implementations, the web page can include multiple content items. Segment grouping generation can include obtaining selection data. The selection data can include a request to save a content item among multiple content items. In some implementations, segment grouping generation can include generating a segment grouping and storing the segment grouping in a user database.
[0057] These systems and methods can obtain address data and location data associated with a segment grouping based on the input data. The address data can be associated with a web page. The content item can be associated with the web page. Additionally and / or alternatively, the location data can describe the location of the content item within the web page.
[0058] Then web page data can be obtained. The web page data can be obtained at least in part based on the address data (e.g., by using a uniform resource locator to navigate to the web page). Alternatively and / or additionally, a file can be obtained based on the address data (e.g., the address data can describe a file location and can be used to obtain the specific file).
[0059] The location within the web page (and / or file) associated with the content item can be determined based on the obtained location data. The location can be determined based on text segments, one or more pointers, and / or via web page processing.
[0060] In some implementations, the address data can include a uniform resource locator. Additionally and / or alternatively, the location data can include one or more text segments. Then, determining the location within the web page associated with the content item can include adding the text segments to the uniform resource locator to generate a shortcut link and entering the shortcut link into a browser.
[0061] Then, a portion of the web page can be provided for display. The portion of the web page can include the locations of content items. In some implementations, providing a portion of the web page can include providing one or more indicators to the portion of the web page. The one or more indicators can indicate the content items associated with the segment grouping. In some implementations, the one or more indicators can include highlighted text associated with the content item.
[0062] These systems and methods can include processing gesture data. For example, these systems and methods can include providing a graphical user interface. The graphical user interface can include a graphical window for displaying a web page. In some implementations, the web page can include multiple content items. These systems and methods can receive gesture data. The gesture data can describe a gesture associated with a portion of the web page. The gesture data can be processed to determine the selected content item. The selected content item can be associated with the portion of the web page. A segment grouping can be generated based on the gesture data. In some implementations, the segment grouping can include the selected content item.
[0063] Specifically, a graphical user interface can be provided for display. The graphical user interface can include a graphical window for displaying a web page. In some implementations, the web page can include multiple content items. The web page can be associated with a uniform resource locator and / or source code. The multiple content items can include structured text data (e.g., body paragraphs and / or one or more headings), whitespace, one or more images, and / or audio content items.
[0064] Then, gesture data can be received. The gesture data can describe a gesture associated with a portion of the web page. In some implementations, the gesture can include a circular gesture surrounding the portion of the web page. The gesture data can describe a touch input to a touch screen display of a mobile computing device.
[0065] The gesture data can be processed to determine the selected content item. The selected content item can be associated with the portion of the web page. For example, the gesture can surround an image and one or more text lines in a web page that includes multiple lines and / or multiple images.
[0066] In some implementations, processing the gesture data to determine the selected content item can include: determining the portion of the web page surrounded by the gesture; determining the focus of the portion; and determining that the selected content item is associated with the focus of the portion.
[0067] Alternatively and / or additionally, processing the gesture data to determine the selected content item can include processing the gesture data with a machine learning model to determine the selected content item. The machine learning model can be trained to determine the start and end of the selected content item based on proximity to gesture boundaries, syntax, structural data, whitespace, and / or semantic cohesion.
[0068] In some implementations, the selected content item may be determined based on a determined matching region associated with a rectangle determined based on a circle. The rectangle may be a rectangle of data based on the syntactic composition of a web page. In some implementations, the selected content item may be determined based on the determined word boundaries and / or based on the determined media content boundaries. The determination may be based on hypertext markup language code boundaries. In some implementations, the source code of the web page may be parsed, and the parsed data may be processed.
[0069] Alternatively and / or additionally, the determination may include calculating the area of a gesture rectangle associated with a gesture. The areas of one or more content item elements may be determined. The intersection areas between the area of the gesture rectangle and each of the areas of different content item elements may be determined. The element with the largest intersection may have the highest selection probability and may thus be determined as the selected content item.
[0070] A segment grouping may then be generated based on the gesture data. The segment grouping may include the selected content item. In some implementations, the segment grouping may include address data and location data. The address data may be associated with the web page, and the location data may describe the location of the content item within the web page.
[0071] A save interface may then be provided for display. The save interface may provide an interactive interface element that may be selected to add the segment grouping to a favorite. A favorites interface may then be provided. A drag input may then be received that drags a graphical representation of the segment grouping to a graphical tile describing a particular favorite. The segment grouping may then be stored in the favorite (e.g., the segment grouping may be stored along with a relationship label linking the segment grouping to the particular favorite). In some implementations, when the graphical representation of the segment grouping is dragged to a particular favorite, the graphical representation may change in size and / or scale. The change in size and scale may provide an intuitive indication of the favorite addition while providing an aesthetic display.
[0072] In some implementations, the content item may be processed to determine an entity associated with the content item. The entity may be determined by processing the content item (e.g., text data, image data, audio data, video data, latent encoding data, and / or link data) with a machine learning model (e.g., an image classification model, an object classification model, a text classification model (e.g., a natural language processing model), a segmentation model, a semantic model, and / or a detection model) to generate entity data (e.g., a classification). Relationship data may then be generated based on the entity data and added to the segment grouping. The relationship data may include one or more entity labels and / or references to other related segment groupings and / or related web pages or content items.
[0073] The segment groups can be searchable within one or more applications and / or databases. Additionally and / or alternatively, the segment groups can be shared via a messaging application, a social media application, and / or via another application.
[0074] In some implementations, one or more machine learning models can be used to process content items to generate tags that can be stored in segment groups. The tags can then be used as searchable tags to display the segment groups in response to a search query. The tags can be determined based on the content of the content items (e.g., the identified words, the identified objects in an image, the characteristics of a video frame or an audio stream, etc.).
[0075] In some implementations, segment group generation occurs after selecting a segment group generation interface element, which can open a segment group generation interface. Alternatively and / or additionally, the user can select a content item and one or more (e.g., two) pop - up windows can be set with various options for interacting with the content item, and one of the options can include segment group generation. Alternatively and / or additionally, search results associated with the content item can be provided.
[0076] In some implementations, a screenshot request can be received. A screenshot can be generated and uploaded to a new client, and lemmas can be generated. The lemmas can be used to receive data associated with the screenshot. The lemmas, the screenshot, and the screenshot details can be used to generate segment groups.
[0077] These systems and methods can be implemented to generate segment groups based on data sources other than just web pages. For example, these systems and methods can be used to generate segment groups based on content items saved in data files stored locally and / or on a server computing system. The generated segment groups can include content items (and / or graphics cards), address data, and location data. The address data can describe the location where the data file is saved (e.g., the name of the drive and the name of any folders (e.g., G:\ResearchPapers\Quantum\Spin)). The location data can describe the location of the content item within the data file. The location data can include start data and end data, which can be used to find matching data in the data file and then navigate to and highlight the matching data. Alternatively and / or additionally, the location data can include one or more pointers.
[0078] These systems and methods can enable a user to select a subset, an excerpt, and / or a portion of an object (e.g., text, an image, and / or a video that can be part of a larger web page). In some implementations, these systems and methods can split a portion of text from a larger body of text, can split a portion of an image, can isolate frames in a video, and / or can split a portion of an audio file.
[0079] In some implementations, these systems and methods can be used as extensions in browser applications and / or as features located on top of another application, such that segment grouping generation can be used to display content in a variety of different applications. For example, these systems and methods can be built into the operating system of a computing device to allow segment grouping generation to occur in response to selections made in a variety of different applications (e.g., map applications, browser applications, social media applications, etc.).
[0080] Additionally and / or alternatively, these systems and methods can be utilized by a variety of different types of computing devices. For example, these systems and methods can be utilized by mobile computing devices, desktop computing devices, smart wearable devices (e.g., smart glasses), and / or other computing devices. These systems and methods can be used in virtual reality interfaces and augmented reality interfaces.
[0081] In some implementations, a segment grouping can include user context data that describes the context of the user when the segment grouping was generated. For example, a computing device can include multiple sensors that can collect data about the context of the user. In some implementations, physical location data of the user's computing device can be obtained and stored in the segment grouping to provide further context to the segments. The physical location data can be provided in a graphics card and / or can be provided as an optional data set that can be viewed during segment grouping interactions.
[0082] These systems and methods can locally store segment groupings on the user device and / or can store segment groupings on a server computing system. Local storage of segment groupings can be used to ensure that segment groupings remain private to the user and can provide offline access. Additionally, metadata related to the collection and generation of segment groupings can remain private and secure.
[0083] These systems and methods can be provided via browser extensions, via overlay applications, and / or via built-in application features. These systems and methods can be used on mobile devices, desktop devices, smart wearable devices, and / or other computing devices.
[0084] A segment grouping can include other metadata associated with selected portions, web pages, and / or one or more contexts of the user (e.g., the application being used, the time of day, the geographical location of the user, and / or user profile data).
[0085] These systems and methods can be executed on a server computing system. Alternatively and / or additionally, these systems and methods can be executed locally on the user computing device. In some implementations, the user computing device can be communicatively connected via a network and can transmit data to perform cloud-based computing. Segment groupings can be stored locally and / or can be stored on a server.
[0086] The systems and methods of the present disclosure provide a variety of technical effects and benefits. As an example, these systems and methods can generate and store segment groups. Specifically, the systems and methods disclosed herein can obtain input data, determine content items associated with the input data (e.g., text, images, videos, and / or audio), generate segment groups, and store the segment groups. The segment groups can include graphical representations of the content items that, when selected, can direct a user to the portion of the web page from which the content item originated. The generation and saving of segment groups can enable easy access to the saved content while maintaining links to more contexts regarding the content items.
[0087] Another technical benefit of the systems and methods of the present disclosure is the ability to use segment groups to share hierarchical information at a relatively low transmission cost. For example, these systems and methods can generate segment groups. The segment groups can be shared with a second user who can initially view the content items. The second user can then select a content item to navigate to the web page and be routed to the specific portion of the web page from which the content item originated, which can allow the second user to obtain more context regarding the content item. Since the content items, URLs, and text segments can be transmitted, the provision of hierarchical information can be accomplished at a relatively low transmission cost. The second user can interact with the segment group, view the content items individually, and then can select the segment group to use the URL in combination with the text segment to navigate to the portion of the web page with the content item highlighted or otherwise indicated. Sending the entire web page file with highlights can involve more upload and download during transmission.
[0088] Another example of technical effects and benefits relates to an improvement in computational efficiency and the functionality of computing systems. For example, the systems and methods disclosed herein can use segment groups to mitigate the amount of data stored for saving content items and related web page contexts. Specifically, the segment groups can include compressed versions of the content items, URLs, and text segments, instead of saving compressed versions of entire web pages that can include a large number of content items and embedded data. Additionally, searching across an entire collection of segment groups is computationally cheaper than searching across multiple compressed web pages.
[0089] Referring now to the drawings, example embodiments of the present disclosure will be discussed in more detail.
[0090] Example Apparatus and System
[0091] Figure 1A A block diagram of an example computing system 100 that performs segment group generation in accordance with an example embodiment of the present disclosure is depicted. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.
[0092] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop computer or a desktop computer), a mobile computing device (e.g., a smart phone or a tablet computer), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0093] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple processors operatively connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0094] In some implementations, the user computing device 102 can store or include one or more grouping generation models 120. For example, the grouping generation model 120 can be or otherwise include various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including non-linear models and / or linear models. Neural networks can include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Refer to Figures 2 to 4 for a discussion of example grouping generation models 120.
[0095] In some implementations, the one or more grouping generation models 120 can be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 can implement multiple parallel instances of a single grouping generation model 120 (e.g., to perform parallel fragment grouping generation across multiple instances selected for a segment).
[0096] More specifically, the grouping generation model can receive input data, determine one or more selected content items, generate fragment groupings, and / or determine one or more labels. In some implementations, the grouping generation model can process the selected content items and generate a summary to be added to the fragment grouping.
[0097] Additionally or alternatively, one or more grouping generation models 140 may be included in or otherwise stored and implemented by the server computing system 130, which communicates with the user computing device 102 according to a client-server relationship. For example, the grouping generation model 140 may be implemented by the server computing system 140 as part of a web service (e.g., a snippet grouping generation service). Thus, one or more models 120 may be stored and implemented at the user computing device 102, and / or one or more models 140 may be stored and implemented at the server computing system 130.
[0098] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to a user input object (e.g., a finger or a stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other devices by which a user may provide user input.
[0099] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or multiple processors operatively connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc. and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0100] In some implementations, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. In cases where the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0101] As described above, the server computing system 130 may store or otherwise include one or more machine-learned grouping generation models 140. For example, the model 140 may be or may otherwise include various machine learning models. Example machine learning models include neural networks or other multi-layer non-linear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Refer to Figures 2 to 4 for a discussion of example models 140.
[0102] The user computing device 102 and / or the server computing system 130 may train the models 120 and / or 140 via interaction with a training computing system 150 communicatively coupled via a network 180. The training computing system 150 may be separate from the server computing system 130 or may be part of the server computing system 130.
[0103] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple processors operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc. and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes one or more server computing devices or is otherwise implemented by one or more server computing devices.
[0104] The training computing system 150 may include a model trainer 160 that trains the machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques (e.g., such as error backpropagation). For example, a loss function may be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update the parameters over multiple training iterations.
[0105] In some implementations, performing error backpropagation may include performing truncated backpropagation through time. The model trainer 160 may perform various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0106] Specifically, the model trainer 160 may train the grouping generation model 120 and / or 140 based on a collection of training data 162. The training data 162 may include, for example, training inputs (e.g., training gestures), training web pages, training text data, training image data, ground truth graphics cards, ground truth segment groupings (e.g., ground truth segments, ground truth address data, and / or ground truth location data), and / or ground truth entity labels.
[0107] In some implementations, if the user has provided consent, training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some instances, this process may be referred to as personalizing the model.
[0108] The model trainer 160 includes computer logic for providing the desired functionality. The model trainer 160 may be implemented using hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium (such as RAM, a hard disk, or optical or magnetic media).
[0109] The network 180 may be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. Generally, communication over the network 180 may be performed using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, secure HTTP, SSL) over any type of wired and / or wireless connection.
[0110] The machine learning models described in this specification may be used for a variety of tasks, applications, and / or use cases.
[0111] In some implementations, the input to the machine learning models of the present disclosure may be image data. The machine learning models may process the image data to generate an output. As an example, the machine learning models may process the image data to generate an image recognition output (e.g., recognition of the image data, latent embedding of the image data, encoded representation of the image data, hash of the image data, etc.). As another example, the machine learning models may process the image data to generate an image segmentation output. As another example, the machine learning models may process the image data to generate an image classification output. As another example, the machine learning models may process the image data to generate an image data modification output (e.g., change of the image data, etc.). As another example, the machine learning models may process the image data to generate an encoded image data output (e.g., encoded representation and / or compressed representation of the image data, etc.). As another example, the machine learning models may process the image data to generate an enlarged image data output. As another example, the machine learning models may process the image data to generate a prediction output.
[0112] In some implementations, the input to the machine learning model of the present disclosure can be text or natural language data. The machine learning model can process the text or natural language data to generate an output. As an example, the machine learning model can process natural language data to generate a language encoding output. As another example, the machine learning model can process text or natural language data to generate a latent text embedding output. As another example, the machine learning model can process text or natural language data to generate a translation output. As another example, the machine learning model can process text or natural language data to generate a classification output. As another example, the machine learning model can process text or natural language data to generate a text segmentation output. As another example, the machine learning model can process text or natural language data to generate a semantic intent output. As another example, the machine learning model can process text or natural language data to generate amplified text or natural language output (e.g., text or natural language data with higher quality than the input text or natural language, etc.). As another example, the machine learning model can process text or natural language data to generate a prediction output.
[0113] In some implementations, the input to the machine learning model of the present disclosure can be speech data. The machine learning model can process the speech data to generate an output. As an example, the machine learning model can process speech data to generate a speech recognition output. As another example, the machine learning model can process speech data to generate a speech translation output. As another example, the machine learning model can process speech data to generate a latent embedding output. As another example, the machine learning model can process speech data to generate an encoded speech output (e.g., an encoded representation and / or a compressed representation of the speech data, etc.). As another example, the machine learning model can process speech data to generate amplified speech output (e.g., speech data with higher quality than the input speech data, etc.). As another example, the machine learning model can process speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine learning model can process speech data to generate a prediction output.
[0114] In some implementations, the input to the machine learning model of the present disclosure can be latent encoded data (e.g., an input latent space representation, etc.). The machine learning model can process the latent encoded data to generate an output. As an example, the machine learning model can process the latent encoded data to generate a recognition output. As another example, the machine learning model can process the latent encoded data to generate a reconstruction output. As another example, the machine learning model can process the latent encoded data to generate a search output. As another example, the machine learning model can process the latent encoded data to generate a reclustering output. As another example, the machine learning model can process the latent encoded data to generate a prediction output.
[0115] In some implementations, the input to the machine learning model of the present disclosure can be statistical data. The machine learning model can process the statistical data to generate an output. As an example, the machine learning model can process the statistical data to generate an identification output. As another example, the machine learning model can process the statistical data to generate a prediction output. As another example, the machine learning model can process the statistical data to generate a classification output. As another example, the machine learning model can process the statistical data to generate a segmentation output. As another example, the machine learning model can process the statistical data to generate a visualization output. As another example, the machine learning model can process the statistical data to generate a diagnostic output.
[0116] In some cases, the machine learning model can be configured to perform a task that includes encoding the input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task can be an audio compression task. The input can include audio data, and the output can include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task can include generating an embedding for the input data (e.g., input audio or visual data).
[0117] In some cases, the input includes visual data, and the task is a computer vision task. In some cases, the input includes pixel data of one or more images, and the task is an image processing task. For example, the image processing task can be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images depict an object belonging to the object class. The image processing task can be object detection, where the image processing output identifies one or more regions in one or more images, and for each region, identifies the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, where the image processing output defines, for each pixel in one or more images, a corresponding likelihood for each class in a predefined set of classes. For example, the set of classes can be foreground and background. As another example, the set of classes can be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines a corresponding depth value for each pixel in one or more images. As another example, the image processing task can be motion estimation, where the network input includes multiple images, and the image processing output defines, for each pixel in one of the input images, the motion of the scene depicted at the pixel between the images in the network input.
[0118] In some cases, the tasks include encrypting or decrypting input data. In some cases, the tasks include a microprocessor performing tasks such as branch prediction or memory address translation.
[0119] Figure 1A An example computing system that can be used to implement the present disclosure is shown. Other computing systems can also be used. For example, in some implementations, the user computing device 102 can include a model trainer 160 and a training data set 162. In such implementations, the model 120 can be both locally trained and used at the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the model 120 based on user-specific data.
[0120] Figure 1B A block diagram of an example computing device 10 executing in accordance with an example embodiment of the present disclosure is depicted. The computing device 10 can be a user computing device or a server computing device.
[0121] The computing device 10 includes multiple applications (e.g., Application 1 to Application N). Each application contains its own machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.
[0122] As Figure 1B shown, each application can communicate with multiple other components of the computing device (e.g., such as one or more sensors, a context manager, a device status component, and / or additional components). In some implementations, each application can communicate with each device component using an API (e.g., a common API). In some implementations, the API used by each application is specific to that application.
[0123] Figure 1C A block diagram of an example computing device 50 executing in accordance with an example embodiment of the present disclosure is depicted. The computing device 50 can be a user computing device or a server computing device.
[0124] The computing device 50 includes multiple applications (e.g., Application 1 to Application N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like. In some implementations, each application can use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).
[0125] The central intelligence layer includes multiple machine learning models. For example, as Figure 1CAs shown, a corresponding machine learning model (e.g., a model) can be provided for each application, and the machine learning model is managed by the central intelligent layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligent layer can provide a single model (e.g., a single model) for all applications. In some implementations, the central intelligent layer is included within or otherwise implemented by the operating system of the computing device 50.
[0126] The central intelligent layer can communicate with the central device data layer. The central device data layer can be a centralized data repository of the computing device 50. As Figure 1C shown, the central device data layer can communicate with multiple other components of the computing device (e.g., such as one or more sensors, a context manager, a device status component, and / or additional components). In some implementations, the central device data layer can use an API (e.g., a private API) to communicate with each device component.
[0127] Example Model Arrangement
[0128] Figure 2 Depicts an illustration of an example fragment grouping generation interface 200 according to an example embodiment of the present disclosure. The fragment grouping generation interface can be configured to receive text selections, image selections, screenshots, video selections, and / or audio selections. Then an option menu can pop up, and the option menu has multiple options including a save option. The save option can be selected, which can prompt the generation of a graphics card and provide the graphics card for display. A fragment grouping can be generated. Then the user can select in which favorite to save the fragment grouping.
[0129] For example, at 202, a web page display window is provided for display, a portion of the web page including a text content item is selected, and a user interface element is selected to be saved. At 204, a fragment grouping is generated, and the generated graphical representation (e.g., a graphics card with text indicating the title and location of the portion) is provided for display. Additionally, at 204, the user interface includes a pop-up window for providing options to add to favorites, which can include adding to a pre-existing favorite and / or generating a new favorite. These favorites can be automatically generated and / or manually generated.
[0130] At 206, a portion of a search results web page is selected (e.g., a portion of a knowledge graph can be selected), and a save option can be selected. At 208, a graphics card is provided for display, where the graphics card is generated based on the selected portion of the search results page. For example, the graphics card can include a picture associated with the selected portion as a background, and the selected text is displayed in the foreground. Additionally and / or alternatively, the graphics card can include a search query and the associated search engine. A favorites addition option can be provided.
[0131] At 210, a screenshot input is received and save options are provided for display. At 212, a save input is received and the generated graphical representation and the option to add to favorites are provided for display. The graphical representation may include at least a portion of the screenshot and a banner including source information (e.g., the title of the web page, the entity associated with the screenshot information, and / or the source of the screenshot information).
[0132] The generated segment groupings may be stored together with the generated graphical representation. Then, the user may select a segment grouping to view an enlarged graphical representation, view the saved content item, and / or navigate to a specific point in the web page where the content item in the resource originated.
[0133] Figure 3 An illustration of an example gesture interaction 300 according to an example embodiment of the present disclosure is depicted. For example, a user input requesting a clipping interface (e.g., selection of a "clip" user interface element 302) may be received. A segment grouping generation interface may be provided, and a graphical animation 304 may be provided to demonstrate how to select a content item. The graphical animation may demonstrate that a content item can be selected using a sliding gesture with a circular motion. An instruction user interface element 306 may be provided as a pop-up window to provide instructions on how to clip and save.
[0134] Figure 4 An illustration of an example segment grouping generation interface 400 according to an example embodiment of the present disclosure is depicted. In some implementations, a request for a segment grouping generation interface may be based on the received user input, and a segment grouping generation interface may be provided. The segment grouping generation interface may include clipping options or summarization options that may allow for the generation of segment groupings based on the selected content items. The segment grouping generation interface 400 may include a web page display window 402 and an interaction pane 404, which may be provided in response to a web page interaction and may provide a plurality of web page interaction options for display. The interaction options may be predefined and uniform across web pages. Alternatively and / or additionally, the interaction options may vary based on the content of the web page, based on the selected content item, and / or based on the entity associated with the web page. The interaction options may include adding to favorites, summarizing the selected portion and / or the entire web page, clipping a portion of the web page, saving an image, finding more web pages similar to the current web page, comparing a product in the web page with other similar products, and / or tracking price options.
[0135] Figure 5Depicts an illustration of an example snippet grouping generation and favorite addition interface 500 according to an example embodiment of the present disclosure. The systems and methods disclosed herein can be used to generate and save snippet groupings. The snippet groupings can be added to favorites and can be searchable for later use. The snippet groupings can include generated graphic cards. In some implementations, the selected content items can be processed by one or more machine learning models to generate a summary of the content items, and the summary can be used to generate the graphic cards. The graphic cards can be customizable. The graphic cards can include a background color and / or a background image determined based on the source web page. Alternatively and / or additionally, the background color and / or the background image can be user-selected. Similarly, the font can be automatically determined, predefined, and / or user-selected.
[0136] At 502, a snippet grouping (including a graphic representation) is generated based on a selected portion of a web page. The generated snippet grouping is then added to the "Inspo" favorites of media content items and snippet groupings.
[0137] At 504, a summary is generated for the portion of the web page and a graphic card is generated based on semantic understanding and / or the determined entities. The generated graphic card can be saved as part of the generated snippet grouping and can be added to favorites. The favorites can be associated with a specific entity and / or a specific type of entity.
[0138] At 506, different options for sharing and / or customizing the generated snippet grouping are provided for display. For example, the template, text, and / or background of the graphic representation can be customized. The sharing options can include adding to favorites, adding to notes, sending via text message, sending via email, copying, and / or air dropping.
[0139] At 508, the snippet grouping can be posted to social media and / or posted to a local search interface. For example, a user can utilize a search application that can display multiple web search results in response to a query and / or can display one or more of the generated snippet groupings (e.g., the generated snippet groupings of a specific user and / or the generated snippet groupings of associated users (e.g., friends or users near a specific user)) in response to a query.
[0140] Then, the generated snippet grouping can be shared via a messaging application, a social media application, and / or various other methods. In some implementations, the snippet grouping can be posted to the web and can be used as a new format for web search results.
[0141] Figure 9Depicts an illustration of an example favorites addition interface 900 in accordance with an example embodiment of the present disclosure. The favorites addition interface may provide a graphical representation (e.g., a graphic card) of content items and / or snippet groupings for display. In response to selecting a save element, a plurality of favorites may be provided for display. The user may then add snippet groupings to a particular favorite. The snippet groupings may then be stored in the particular favorite.
[0142] Then, a particular favorite may be opened, and a graphical representation of the generated snippet groupings may be provided for display together with other graphical representations associated with other snippet groupings. The favorites addition interface 900 may include (at a first size) displaying a graphic card for display during snippet grouping generation 902. Then, a pop-up window 904 for favorites addition may be provided for display when a user interface element is selected. When a snippet grouping is added to a particular favorite, the favorite 906 may then be provided for display (at a second size) together with a plurality of thumbnails depicting different snippet groupings in the favorite, the plurality of thumbnails including a thumbnail depicting the generated graphic card.
[0143] Figure 10A Depicts an illustration of an example snippet grouping interaction 1020 in accordance with an example embodiment of the present disclosure. Once a snippet grouping is generated, it may be interacted with to navigate to a web page associated with the content item. The location data of the snippet grouping may be used to navigate to a particular portion of the web page that is the source of the content item. The content item may be highlighted when displayed. For example, the graphic card 1022 may be selected. Then, the address data and location data of the snippet grouping associated with the graphic card may be obtained. The address data and location data may then be utilized to open the web page 1024 to reach the exact location of one or more content items of the selected snippet grouping, where the one or more content items are highlighted.
[0144] Figure 10B Depicts an illustration of an example snippet grouping search 1040 in accordance with an example embodiment of the present disclosure. The generated snippet groupings may be provided as search results during a local search and / or when searching the web. For example, a user may enter a search query, which may be processed to determine a plurality of suggested queries and a plurality of suggested snippet groupings, which may be provided for display (e.g., as shown in 1042) when further input is received.
[0145] Figure 10C Depicts an illustration of an example highlighting interface 1080 in accordance with an example embodiment of the present disclosure. In some implementations, in response to receiving user input requesting removal of the highlight, the highlight (e.g., as shown in 1082) may be removed (e.g., as shown in 1084).
[0146] Figure 11 Depicts an illustration of an example graphics card 1100 according to an example embodiment of the present disclosure. The graphics card may be automatically generated and / or may be generated based on one or more user inputs. The graphics card may include a portion assigned to a content item and / or a portion assigned to an attribute of a source of the content item. The graphics card may include text, images, video, and / or audio. The graphics card may vary based on the selected content item. The text may be the selected text and / or may be a summary of data obtained from a web page. In some implementations, the background and / or image may be automatically selected / generated and / or may be user selected. Specifically, Figure 11 Depicts text-based graphics cards 1102 in two formats, which have resource attributes and associated titles / search queries; image-based graphics cards 1104, which have an image, an entity thumbnail, a title or caption, and resource attributes; and screenshot-based graphics cards 1106, which include a screenshot, an entity logo, a title of a corresponding article, and resource attributes.
[0147] Figure 12 Depicts an illustration of an example snippet grouping generation interface 1200 according to an example embodiment of the present disclosure. Snippet grouping generation may include generating a graphics card. The graphics card may include text that describes a summary of the selected data in a web page. The summary may be generated by processing the selected data with a machine learning model. For example, at 1202, a portion of a web page is selected. A user may select a summary user interface element to process the selected portion to generate a summary. The summary may include the language from the passage and / or may include different languages that summarize the passage in simple terms or more user-specific terms. At 1204, the summarized text is used to generate a graphics card to be stored with the generated snippet grouping, the graphics card including information associated with the web page and a specific location of the selected portion within the web page.
[0148] Figure 13 Depicts an illustration of an example snippet grouping suggestion interface 1300 according to an example embodiment of the present disclosure. In some implementations, the snippet grouping suggestion interface may include providing suggested graphics cards associated with the suggested snippet groupings (e.g., as shown at 1304, these graphics cards may be displayed in response to a selection of a bookmark option as shown at 1302). The suggested snippet groupings may be based on a particular user's past user interactions and / or based on other users' past interactions (e.g., popular snippets). The suggested snippet groupings may be selected by the user computing system and stored locally.
[0149] Figure 14Depicts an illustration of an example snippet grouping generation and sharing interface 1400 in accordance with an example embodiment of the present disclosure. The snippet grouping may be generated in response to a user's selection of a content item (e.g., the selection in 1402). The generated snippet grouping may then be shared (e.g., shared via the text message option of a sharing interface 1404 that provides multiple sharing options). The snippet grouping may be shared via a messaging application, a social media application, and / or may be inserted into another data file (e.g., a note, a text document, and / or a slide set). The shared snippet grouping may be sent along with a graphics card and a download option for local saving (e.g., as shown in 1406). Alternatively and / or additionally, the graphics card may be selectable to navigate to one or more content items of the snippet grouping in a native web page.
[0150] Figure 15 Depicts an illustration of an example graphics card customization interface 1500 in accordance with an example embodiment of the present disclosure. The graphics card customization interface may involve multiple templates that may be selected to customize the graphics card. The templates may include different images, different colors, different fonts, and / or different layouts. For example, an initial graphics card 1502 may include text in a first font, text in a first size, and a first background. A change template request may be received, and an enhanced graphics card 1504 with a different text font, a different text size, and / or a different background may be generated.
[0151] Figure 16 Depicts a block diagram of an example snippet grouping generation system 1600 in accordance with an example embodiment of the present disclosure. The snippet grouping generation system 1600 may include obtaining input data 1602 (e.g., input data describing a selection of one or more content items). The input data 1602 may be processed to determine the selected content item 1604. Based on the determined selection, the content item may be obtained, and a graphics card 1606 may be generated. Address data 1608 may be generated and / or obtained. The address data 1608 may include uniform resource locator data. Location data 1610 may be generated and / or obtained. The location data 1610 may include text snippet data that may include a scroll position, a start of a content item, and an end of a content item, and the text snippet data may be used to locate and highlight the content item within a source page. The graphics card 1606, the content item, the address data 1608, and the location data 1610 may be used to generate a snippet grouping 1612. The snippet grouping 1612 may be processed to determine one or more entity tags 1614 of the snippet grouping 1612 based on the content item and / or based on the source of the content item. The entity tags 1614 may include relationship data that links the snippet grouping 1612 to other snippet groupings associated with the same entity. The snippet grouping 1612 with the entity tags 1614 may then be stored 1616 (e.g., locally stored and / or stored on a server computing system).
[0152] One or more of the determinations and / or one or more of the generations may be performed at least in part based on one or more machine learning models. For example, determining the selected content item 1604, obtaining the content item and / or generating the graphics card 1606, generating the location data 1610, and / or determining the entity label 1614 may be performed by one or more machine learning models.
[0153] Example Method
[0154] Figure 6 A flowchart depicting an example method for execution in accordance with an example embodiment of the present disclosure is shown. Although Figure 6 the steps are depicted in a particular order for purposes of illustration and discussion, the methods of the present disclosure are not limited to the particular order or arrangement shown. The various steps of method 600 may be omitted, rearranged in various ways, combined, and / or adjusted without departing from the scope of the present disclosure.
[0155] At 602, the computing system may provide data describing a graphical user interface. The graphical user interface may include a graphical window for displaying a web page. In some implementations, the web page may include a plurality of content items.
[0156] At 604, the computing system may obtain input data. The input data may include a request to save one or more of the plurality of content items. In some implementations, one or more of the content items may include at least one of an image, a video, a graphical description of a product, or audio.
[0157] At 606, the computing system may generate a segment grouping. The segment grouping may be generated based on the input data. The segment grouping may include one or more content items, address data, and location data.
[0158] One or more of the content items may include text data (e.g., text data describing words, sentences, quotations, and / or paragraphs), image data (e.g., an image of an object (e.g., a product)), video data (e.g., a video and / or one or more frames of a video), audio data (e.g., waveform data), and / or latent encoded data. One or more of the content items may include multimodal data.
[0159] In some implementations, generating the snippet grouping may include obtaining one or more content items and generating a graphic card. The graphic card may describe the one or more content items. The graphic card may include text data overlaid on a color and / or an image. The color may be determined based on the primary color of the web page. In some implementations, the color may be predefined and / or may be determined based on surrounding content items. The image may be an image determined based on a determined theme of the content item. Alternatively and / or additionally, the image may be an image close to the content item. In some implementations, the graphic card may include a font determined based on the font used in the web page. The graphic card may include a background, text describing the content item (e.g., the content item and / or a summary of the content item), and / or text and / or a logo describing the source and / or the determined entity. In some implementations, the text size of the text in the graphic card may be based on the amount of text in the content item.
[0160] The address data may describe the web address of the web page. The address data may include a Uniform Resource Identifier and / or a Uniform Resource Locator.
[0161] The location data may describe the location of the one or more content items within the web page. In some implementations, the location data may include at least one of a scroll position, a start node, or an end node. The scroll position may describe the location of the one or more content items relative to other parts of the web page. In some implementations, the start node may describe the location where the one or more content items start. The end node may describe the location where the one or more content items end. The location data may include a text fragment (published by Tomayac et al. (May 20, 2022, 9:40 PM), "Scroll to TextFragment", GITHUB, https: / / github.com / WICG / scroll-to-text-fragment), which may be used to indicate the location of the content item. In some implementations, the text fragment may include one or more text instructions associated with the location. The text instructions may include a string of code start data (e.g., the first text associated with the content item and / or the first pixel associated with the content item) and / or end data (e.g., the last text associated with the content item and / or the last pixel associated with the content item). The text instructions may be used to search for a data set in the web page that matches the start and / or end of the content item.
[0162] In some implementations, the computing system may include processing the one or more content items to determine an entity associated with the one or more content items. An entity label may be generated based on the entity. The snippet grouping may include the entity label.
[0163] At 608, the computing system may store a segment grouping. The segment grouping may be associated with a particular user. The association with the particular user may include storing the segment grouping together with metadata indicating the user and / or associating the segment grouping with a particular user profile. The particular user may be the user who provided the input data that was processed to determine the segment grouping generation request. The segment grouping may be stored in a user database. In some implementations, storing the segment grouping in the user database may include locally storing the segment grouping on a mobile computing device. Additionally and / or alternatively, a graphics card may be stored as a graphical representation of the segment grouping. The graphics card may be automatically generated and may be user-customizable. The graphics card may start from a template that may be customized based on other content in a web page, based on user input, and / or based on other contexts. In some implementations, the graphics card may include multimodal data. Alternatively and / or additionally, the segment grouping may be stored on a server computing system. In some implementations, if a content item references another content item, these systems and methods may obtain the additional content item and save the additional content item in the segment grouping.
[0164] In some implementations, these systems and methods may include receiving a segment request that provides a segment interface. The segment interface may include interactive elements associated with the segment grouping. The segment interface may be provided for display. An interface selection may be received, and the interface selection may describe a selection of an interactive element. Then a portion of a web page may be provided for display. In some implementations, the portion of the web page may include the location of one or more content items within the web page.
[0165] In some implementations, these systems and methods may include receiving an insertion input. The insertion input may include user input requesting to insert a content item into a different interface. The segment grouping may be provided to a third-party server computing system.
[0166] Additionally and / or alternatively, these systems and methods may include adding the segment grouping to a favorites list. The segment grouping may be added to the favorites list based on the received user input. Alternatively and / or additionally, the segment grouping may be automatically added to the favorites list based on a determined entity associated with the content item. In some implementations, the segment grouping may be added to the favorites list based on the source of the content item (e.g., the type of web page, the type of media provider, and / or based on the type of content item).
[0167] In some implementations, a segment grouping may be generated based on content items obtained from sources other than a web page (e.g., a mobile application, a large data file (e.g., a downloaded video or book), and / or another data source).
[0168] Figure 7 A flowchart of an example method for performing according to an example embodiment of the present disclosure is depicted. AlthoughFigure 7 For purposes of illustration and discussion, steps are depicted as being performed in a particular order, but the methods of the present disclosure are not limited to the particular order or arrangement shown. The various steps of method 700 may be omitted, rearranged in various ways, combined, and / or adjusted without departing from the scope of the present disclosure.
[0169] At 702, the computing system may obtain input data. In some implementations, the input data may describe a selection of content items associated with a segment grouping. The segment grouping may be associated with a user account of a particular user. Additionally and / or alternatively, the user account may be associated with one or more platforms. The segment grouping may include content items and deep links. Alternatively and / or additionally, the segment grouping may include segments (e.g., content items and / or media data generated based on content items (e.g., graphics cards and / or summaries of content items)), address data (e.g., uniform resource locators and / or uniform resource identifiers), and / or metadata. The segment grouping may include location data (e.g., metadata indicating the location of a content item within a web page, text segments for identifying start data and end data, one or more pointers, and / or one or more scroll position data).
[0170] In some implementations, a segment grouping may be generated by providing a graphical user interface for display. The graphical user interface may include a graphical window for displaying a web page. In some implementations, the web page may include multiple content items. Segment grouping generation may include obtaining selection data. The selection data may include a request to save a content item among the multiple content items. In some implementations, segment grouping generation may include generating a segment grouping and storing the segment grouping in a user database.
[0171] At 704, the computing system may obtain address data and location data associated with the segment grouping. The address data may be associated with a web page. The content item may be associated with the web page. Additionally and / or alternatively, the location data may describe the location of the content item within the web page.
[0172] At 706, the computing system may obtain web page data. The web page data may be obtained at least in part based on the address data (e.g., by using a uniform resource locator to navigate to the web page).
[0173] At 708, the computing system may determine the location within the web page associated with the content item. The location may be determined based on text segments, one or more pointers, and / or via web page processing.
[0174] In some implementations, the address data may include a Uniform Resource Locator. Additionally and / or alternatively, the location data may include one or more text fragments. Determining a location within a web page associated with a content item may then include adding the text fragment to the Uniform Resource Locator to generate a shortcut link and entering the shortcut link into a browser.
[0175] At 710, the computing system may provide a portion of a web page. The portion of the web page may include the location of a content item. In some implementations, providing the portion of the web page may include providing one or more indicators to the portion of the web page. The one or more indicators may indicate the content item associated with the fragment grouping. In some implementations, the one or more indicators may include highlighted text associated with the content item.
[0176] Figure 8 A flowchart of an example method for performing in accordance with an example embodiment of the present disclosure is depicted. Although Figure 8 For purposes of illustration and discussion, steps are depicted as being performed in a particular order, but the methods of the present disclosure are not limited to the particular order or arrangement shown. The various steps of method 800 may be omitted, rearranged in various ways, combined, and / or adjusted without departing from the scope of the present disclosure.
[0177] At 802, the computing system may provide a graphical user interface. The graphical user interface may include a graphical window for displaying a web page. In some implementations, the web page may include a plurality of content items.
[0178] At 804, the computing system may receive gesture data. The gesture data may describe a gesture associated with a portion of the web page. In some implementations, the gesture may include a circular gesture surrounding the portion of the web page. The gesture data may describe a touch input to a touch screen display of a mobile computing device.
[0179] At 806, the computing system may process the gesture data to determine the selected content item. The selected content item may be associated with the portion of the web page.
[0180] In some implementations, processing the gesture data to determine the selected content item may include: determining the portion of the web page surrounded by the gesture; determining the focus of the portion; and determining that the selected content item is associated with the focus of the portion.
[0181] Alternatively and / or additionally, processing the gesture data to determine the selected content item may include processing the gesture data with a machine learning model to determine the selected content item.
[0182] In some implementations, the selected content item can be determined based on the determined matching region associated with a rectangle determined based on a circle. The rectangle can be a rectangle of data based on the syntactic composition of a web page. In some implementations, the selected content item can be determined based on the determined word boundaries and / or based on the determined media content boundaries. This determination can be based on the HyperText Markup Language code boundaries. In some implementations, the source code of the web page can be parsed, and the parsed data can be processed.
[0183] Alternatively and / or additionally, the determination can include calculating the area of a gesture rectangle associated with a gesture. The area of one or more content item elements can be determined. The intersection area between the area of the gesture rectangle and each of the areas of different content item elements can be determined. The element with the largest intersection can have the highest selection probability and can thus be determined as the selected content item.
[0184] At 808, the computing system can generate a fragment grouping based on the gesture data. The fragment grouping can include the selected content item. In some implementations, the fragment grouping can include address data and location data. The address data can be associated with the web page, and the location data can describe the location of the content item within the web page.
[0185] Then a save interface can be provided for display. The save interface can provide an interactive interface element that can be selected to add the fragment grouping to a favorite. Then a favorites interface can be provided. Then a drag input can be received that drags a graphical representation of the fragment grouping to a graphical tile that describes a specific favorite. Then the fragment grouping can be stored in the favorite (e.g., the fragment grouping can be stored along with a relationship label that links the fragment grouping to the specific favorite). In some implementations, when the graphical representation of the fragment grouping is dragged to a specific favorite, the graphical representation can change in size and / or scale. The change in size and scale can provide an intuitive indication of the favorite addition while providing an aesthetic display.
[0186] In some implementations, the content item can be processed to determine an entity associated with the content item. The entity can be determined by processing the content item (e.g., text data, image data, audio data, video data, latent encoding data, and / or link data) with a machine learning model (e.g., an image classification model, an object classification model, a text classification model (e.g., a natural language processing model), a segmentation model, a semantic model, and / or a detection model) to generate entity data (e.g., a classification). Then relationship data can be generated based on the entity data and added to the fragment grouping. The relationship data can include one or more entity labels and / or references to other related fragment groupings and / or related web pages or content items.
[0187] The segment grouping can be searchable within one or more applications and / or databases. Additionally and / or alternatively, the segment grouping can be shared via a messaging application, a social media application, and / or via another application.
[0188] In some implementations, one or more machine learning models can be used to process content items to generate tags that can be stored in the segment grouping. The tags can then be used as searchable tags to display the segment grouping in response to a search query. The tags can be determined based on the content of the content item (e.g., identified words, identified objects in an image, characteristics of a video frame or audio stream, etc.).
[0189] In some implementations, segment grouping generation occurs after selecting a segment grouping generation interface element, which can open a segment grouping generation interface. Alternatively and / or additionally, a user can select a content item and one or more (e.g., two) pop - up windows can be set with various options for interacting with the content item, and one of the options can include segment grouping generation. Alternatively and / or additionally, search results associated with the content item can be provided.
[0190] Additional Disclosure
[0191] The techniques discussed herein relate to servers, databases, software applications, and other computer - based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer - based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functions among and within components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0192] Although the subject matter has been described in detail with respect to various specific example embodiments of the subject matter, each example is provided by way of explanation and not as a limitation of the disclosure. Those skilled in the art can readily generate changes, alterations, and equivalents of these embodiments upon understanding the foregoing. Accordingly, the disclosure does not exclude including such modifications, alterations, and / or additions to the subject matter that would be readily understood by those of ordinary skill in the art. For example, features shown or described as part of one embodiment can be used with another embodiment to yield yet another embodiment. Accordingly, the disclosure is intended to cover these changes, alterations, and equivalents.
Claims
1. A computing system, the system comprising: One or more processors; And One or more non-transitory computer-readable media, the one or more non-transitory computer-readable media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations including: Providing data describing a graphical user interface, wherein the graphical user interface includes a graphical window for displaying a web page, and wherein the web page includes a plurality of content items; Obtaining input data, wherein the input data includes a request to save one or more of the plurality of content items; Generating a fragment group, wherein the fragment group includes: The one or more content items; Address data, wherein the address data describes the web address of the web page; and Location data, wherein the location data describes the location of the one or more content items within the web page; Wherein generating the fragment group includes: Determining an image associated with the one or more content items; Determining text describing the one or more content items; and Generating a graphics card, wherein the graphics card includes text overlaid on the image; Storing the fragment group, wherein the fragment group is associated with a specific user; Obtaining a selection of a specific sharing option from one or more sharing options; and Based on the selection of the specific sharing option, sending the fragment group to one or more other computing systems, wherein the fragment group includes the graphics card, and wherein the graphics card is selectable to automatically navigate to the location of the selected content item within the web page based on the address data and the location data.
2. The system of claim 1, wherein the operations further include: Receiving a fragment request providing a fragment interface, wherein the fragment interface includes interactive elements associated with the fragment group; Providing the fragment interface for display; Receiving an interface selection selecting the interactive elements; and Providing a portion of the web page for display, wherein the portion of the web page includes the location of the one or more content items within the web page.
3. The system of claim 1, wherein the operations further include: Receiving an insertion input, wherein the insertion input includes user input requesting to insert the one or more content items into a different interface; And Providing the fragment group to a third-party server computing system.
4. The system of claim 1, wherein the location data includes at least one of a scroll position, a start node, or an end node, wherein the scroll position describes the location of the one or more content items relative to other portions of the web page, wherein the start node describes the location where the one or more content items begin, and wherein the end node describes the location where the one or more content items end.
5. The system of claim 1, wherein the one or more content items include at least one of an image, a video, a graphical description of a product, or an audio.
6. The system of claim 1, wherein the operations further include: Process the one or more content items to determine an entity associated with the one or more content items; Generate an entity label based on the entity; And where the fragment grouping includes the entity label.
7. The system according to claim 1, wherein storing the fragment groups comprises: Locally store the fragment grouping on the mobile computing device.
8. A computer-implemented method, the method comprising: Obtain, by a computing system including one or more processors, input data, where the input data describes a selection of a graphics card associated with a fragment grouping, where the graphics card includes text data and an image, where the text data describes a summary of a content item, where the image is associated with the content item, and where the graphics card includes a summary overlaid on the image; Obtain, by the computing system, address data and location data associated with the fragment grouping, where the address data is associated with a web page, where the content item is associated with the web page, and where the location data describes the location of the content item within the web page; Obtain, by the computing system, web page data, where the web page data is obtained at least in part based on the address data; Determine, by the computing system, the location within the web page associated with the content item at least in part based on the location data; and Provide, by the computing system, a portion of the web page, where the portion of the web page includes the location of the content item, where providing the portion of the web page includes: Providing a web page for display in a graphics window; and Automatically navigating to the location within the web page associated with the content item based on the location data; and Providing one or more indicators of the location of the content item within the web page provided for display.
9. The method of claim 8, where the address data includes a uniform resource locator, where the location data includes a text fragment, and where determining, by the computing system, the location within the web page associated with the content item includes: Adding, by the computing system, the text fragment to the uniform resource locator to generate a shortcut link; And Entering, by the computing system, the shortcut link into a browser.
10. The method of claim 8, where the fragment grouping is generated by: Providing, by the computing system, a graphical user interface, where the graphical user interface includes a graphics window for displaying the web page, where the web page includes a plurality of content items; Obtaining, by the computing system, selection data, where the selection data includes a request to save the content item among the plurality of content items; Generating, by the computing system, the fragment grouping; And Storing, by the computing system, the fragment grouping in a user database.
11. The method of claim 8, where the one or more indicators include highlighted text associated with the content item.
12. The method of claim 8, where the fragment grouping is associated with a user account of a particular user, where the user account is associated with one or more platforms.
13. One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations including: Providing a graphical user interface, wherein the graphical user interface includes a graphical window for displaying a web page, wherein the web page includes a plurality of content items; Receiving gesture data, wherein the gesture data describes a gesture associated with a portion of the web page; Processing the gesture data to determine a selected content item, wherein the selected content item is associated with the portion of the web page; And Generating a snippet group based on the gesture data, wherein the snippet group includes the selected content item, wherein the snippet group includes a graphic card associated with the selected content item, wherein the graphic card depicts data describing the selected content item associated with the portion of the web page, and wherein the graphic card is selectable to automatically navigate to the location of the selected content item within the web page based on address data and location data, wherein the graphic card is generated by: Processing the selected content item to determine a theme associated with the selected content item; Determining an image based on the theme associated with the selected content item; and Generating a graphic card, wherein the graphic card includes the image as a background, and wherein the graphic card includes text associated with at least a subset of the selected content items overlaid on the image.
14. The one or more non-transitory computer-readable media of claim 13, wherein the snippet group includes address data and location data, wherein the address data is associated with the web page, and wherein the location data describes the location of the selected content item within the web page.
15. The one or more non-transitory computer-readable media of claim 13, wherein the gesture includes a circular gesture surrounding the portion of the web page.
16. The one or more non-transitory computer-readable media of claim 13, wherein processing the gesture data to determine the selected content item includes: Determining the portion of the web page surrounded by the gesture; Determining the focus of the portion; And Determining that the selected content item is associated with the focus of the portion.
17. The one or more non-transitory computer-readable media of claim 13, wherein processing the gesture data to determine the selected content item includes: Processing the gesture data with a machine learning model to determine the selected content item.
18. The one or more non-transitory computer-readable media of claim 13, wherein the gesture data describes a touch input to a touch screen display of a mobile computing device.
Citation Information
Patent Citations
Presenting search result information
CN101490677A
Systems and methods for improved searching and categorizing of media content items based on a destination for the media content item
CN113557504A