Search query generation using machine learning models

Large-Scale Language Models assist in generating suggested search queries by analyzing media attributes, addressing the inefficiency of basic queries in large media collections, thereby improving search precision and reducing resource usage.

JP2026525154APending Publication Date: 2026-07-29GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
GOOGLE LLC
Filing Date
2025-04-30
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Users face challenges in efficiently searching through large collections of media items on mobile devices or cloud storage due to the vast number of results returned by basic queries, often requiring multiple searches to find specific images or videos.

Method used

A method utilizing Large-Scale Language Models (LLMs) to generate suggested search queries by identifying attributes from media items, scoring templates based on attribute categories and media item counts, and providing descriptive text for improved search queries.

Benefits of technology

This approach reduces the need for multiple searches by generating complex queries that yield specific results, enhancing user skill in creating effective search queries and optimizing computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026525154000001_ABST
    Figure 2026525154000001_ABST
Patent Text Reader

Abstract

A method performed by a computer includes identifying attributes from a group of media items associated with a user account. The method further includes generating one or more templates with the attributes. The method further includes generating templates with combinations of attributes. The method further includes scoring each template based on the number of attribute categories in the corresponding template and the number of media items in the group of media items containing the corresponding attributes. The method further includes selecting one or more templates based on the corresponding scores. The method further includes providing one or more templates as input to a large language model. The method further includes using the large language model to output descriptive text based on one or more templates. The method further includes providing the descriptive text to the user as a suggested search query in the user interface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 640,690, titled "Search Query Generation Using Machine Learning Models," filed on April 30, 2024, the entire contents of which are incorporated herein by reference.

Background Art

[0002] With the popularity of smartphones, consumers often store thousands of media items on their mobile devices or back them up to cloud storage. Users search for photos and videos using basic queries such as "March 2024" and "John," but due to the vast number of media items stored for the user, the search may return too many results to find what the user is looking for.

[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. In the scope described in this background art section, the achievements of the inventors named in this specification, as well as aspects of this specification that may not meet the requirements of the prior art at the time of filing, are not to be recognized as the prior art to the present disclosure, either explicitly or implicitly.

Summary of the Invention

[0004] A method performed by a computer includes identifying attributes from a group of media items associated with a user account. The method further includes generating one or more templates with the attributes. The method further includes generating templates with combinations of attributes. The method further includes scoring each template based on the number of attribute categories in the corresponding template and the number of media items in the group of media items containing the corresponding attributes. The method further includes selecting one or more templates based on the corresponding scores. The method further includes providing one or more templates as input to a large language model. The method further includes using the large language model to output descriptive text based on one or more templates. The method further includes providing the descriptive text to the user as a suggested search query in a user interface for media items associated with a user account.

[0005] In some embodiments, the method further includes receiving a selection of search queries proposed by the user via a user interface and providing the user with search results from a database of media items associated with the user account that match the descriptive text. In some embodiments, the method further includes determining how many times each attribute appears in media items of a group of media items, and scoring templates is based on how many times each corresponding attribute appears in media items. In some embodiments, the method further includes determining the number of media items that contain the corresponding attribute and, for each template, calculating a minimum threshold by dividing the number of categories by the number of media items in the group of media items, and scoring each template is based on multiplying each template by the minimum threshold corresponding to the number of media items, and selecting one or more templates based on the corresponding score is based on selecting one or more templates in response to the number of media items containing the corresponding attribute exceeding the corresponding score.

[0006] In some embodiments, one or more templates are provided to a large language model along with prompts, the prompts being based on features selected from a group consisting of capitalization emphasis, repetition emphasis, repetitive prompting, negative instructions, structure, constraints, delimiters, and combinations thereof. In some embodiments, one or more templates are provided to the large language model along with a level of randomness to be associated with the descriptive text. In some embodiments, one or more templates are provided to the large language model along with a level of complexity to be associated with the descriptive text. In some embodiments, the method further includes pre-calculating attributes that are part of a group of media items associated with a user account. In some embodiments, the method further includes receiving a request from the user for media items containing one or more specific attributes, and providing the user with a group of media items based on a group of media items containing one or more specific attributes, the proposed search query being provided along with the group of media items. In some embodiments, the media items include one or more screenshots captured by a user device associated with the user account.

[0007] The system includes one or more processors and memory coupled to one or more processors, the memory storing instructions that, when executed by the processors, cause one or more processors to perform an action. The action includes identifying attributes from a group of media items associated with a user account, inputting them into a template with combinations of attributes, scoring each template based on the number of attribute categories in the corresponding template and the number of media items in the group of media items containing the corresponding attributes, selecting one or more templates based on the corresponding scores, providing one or more templates as input to a large language model, using the large language model to output descriptive text based on one or more templates, and providing the descriptive text to the user as a suggested search query in a user interface for media items associated with the user account.

[0008] In some embodiments, the operation further includes receiving a selection of search queries proposed by the user via a user interface and providing the user with search results from a database of media items associated with the user account that match the descriptive text. In some embodiments, the operation further includes determining how many times each attribute appears in media items of a group of media items, and scoring templates is based on how many times each corresponding attribute appears in media items. In some embodiments, the operation further includes determining how many media items contain the corresponding attribute and, for each template, calculating a minimum threshold based on the number of attribute categories in the corresponding template, and scoring each template is based on multiplying each template by the minimum threshold corresponding to the number of media items, and selecting one or more templates based on the corresponding score is based on selecting one or more templates in response to the number of media items containing the corresponding attribute exceeding the corresponding score. In some embodiments, one or more templates are provided to a large language model along with prompts, which are based on features selected from a group consisting of capitalization emphasis, repetition emphasis, repetitive prompting, negative indication, structure, constraint, delimiter, and combinations thereof.

[0009] Non-temporary computer-readable media includes instructions stored internally, which, when executed by one or more computers, cause one or more computers to perform an action. The action includes identifying attributes from a group of media items associated with a user account, inputting them into a template with combinations of attributes, scoring each template based on the number of attribute categories in the corresponding template and the number of media items in the group of media items containing the corresponding attributes, selecting one or more templates based on the corresponding scores, providing one or more templates as input to a large language model, using the large language model to output descriptive text based on one or more templates, and providing the descriptive text to the user as a suggested search query in a user interface for media items associated with a user account.

[0010] In some embodiments, the operation further includes receiving a selection of search queries proposed by the user via a user interface and providing the user with search results from a database of media items associated with the user account that match the descriptive text. In some embodiments, the operation further includes determining how many times each attribute appears in media items of a group of media items, and scoring templates is based on how many times each corresponding attribute appears in media items. In some embodiments, the operation further includes determining how many media items contain the corresponding attribute and, for each template, calculating a minimum threshold based on the number of attribute categories in the corresponding template, and scoring each template is based on multiplying each template by the minimum threshold corresponding to the number of media items, and selecting one or more templates based on the corresponding score is based on selecting one or more templates in response to the number of media items containing the corresponding attribute exceeding the corresponding score. In some embodiments, one or more templates are provided to a large language model along with prompts, which are based on features selected from a group consisting of capitalization emphasis, repetition emphasis, repetitive prompting, negative indication, structure, constraint, delimiter, and combinations thereof. [Brief explanation of the drawing]

[0011] [Figure 1] This is a block diagram of an exemplary network environment according to some embodiments described herein. [Figure 2] This is a block diagram of an exemplary computing device according to some embodiments described herein. [Figure 3] This specification shows an exemplary user interface, including a group of media items, according to several embodiments described herein. [Figure 4]This specification shows exemplary prompts sent to a machine learning model along with one or more templates, according to some embodiments described herein. [Figure 5] Two exemplary user interfaces, including proposed search queries, are shown according to several embodiments described herein. [Figure 6] This specification shows an exemplary user interface, including proposed search queries based on an initial search request from a user, according to several embodiments described herein. [Figure 7] A flowchart illustrating an exemplary method for generating a proposed search query, according to several embodiments described herein, is shown. [Figure 8] A flowchart of another exemplary method for generating the proposed search query, according to some embodiments described herein, is shown. [Modes for carrying out the invention]

[0012] overview Mobile devices use increasingly sophisticated technology to capture photos and videos. Users often store thousands of media items on their mobile devices or back them up to cloud storage. Media items can include any type of media item, such as images, documents, video files / data, audio files / data, and / or video and audio files / data. Users search for media items such as photos and videos using basic queries like "March 2024" and "John," but because of the enormous number of media items stored for users, search results may return too many results to find what the user is looking for. Searches can be improved by adding details to images (e.g., adding descriptions, tags, etc.), but users may not remember enough about the images to improve their search queries.

[0013] The techniques described below advantageously use Large-Scale Language Models (LLMs) to guide users in implementing more complex search queries, outputting suggested search queries for finding media items. Complex search queries can yield specific results that match the user's information needs. Furthermore, as a result of teaching users to implement more complex search queries, users become more skilled at creating them. Instead of performing multiple searches to find a specific image, the suggested search queries use fewer computing resources to obtain the desired search results.

[0014] Network environment Figure 1 shows a block diagram of an exemplary network environment 100. In some embodiments, the network environment 100 includes a media server 101 and user devices 115 coupled to the network 105. Users 125a, 125n may be associated with their respective user devices 115a, 115n. In some embodiments, the network environment 100 may include other servers or devices not shown in Figure 1. In Figure 1 and the remaining figures, letters following a reference number, such as "115a", represent a reference to the element having that particular reference number. Reference numbers in the text without following letters, such as "115", represent a general reference to embodiments of the element prefixed with that reference number.

[0015] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicably coupled to the network 105 via signal lines 102. The signal lines 102 may be a wired connection such as Ethernet®, coaxial cable, or fiber optic cable, or a wireless connection such as Wi-Fi®, Bluetooth®, or other wireless technology. In some embodiments, the media server 101 sends and receives data to and from one or more user devices 115a, 115n via the network 105.

[0016] The media server 101 may include a media application 103a, a machine learning model 120, and a database 199. Although the machine learning model 120 is shown as being stored in the same media server 101 as the media application 103a, in some embodiments the machine learning model 120 is stored in a separate server.

[0017] The machine learning model 120 is trained to provide text in response to queries. For example, the machine learning model 120 may be a large-scale language model (LLM) designed for natural language processing tasks such as language generation. The LLM is trained to receive one or more templates and prompts and to output descriptive text based on one or more templates. The machine learning model 120 receives one or more templates with attributes from the media application 103 (for example, from media application 103a on a media server, or from media applications 103b, 103c stored on user devices 115a, 115b). The machine learning model 120 outputs descriptive text.

[0018] The trained machine learning model 120 may include one or more model forms or structures. For example, the model form or structure may include any type of neural network, such as a linear network, a deep learning neural network implementing multiple layers (e.g., a “hidden layer” between the input and output layers, with each layer being a linear network), a convolutional neural network (e.g., a network that divides or partitions input data into multiple parts or tiles, processes each tile individually using one or more neural network layers, and aggregates the results of processing each tile), or a sequence-to-sequence neural network (e.g., a network that takes sequential data such as words in a text or frames in a video as input and produces a resulting sequence as output).

[0019] The model format or structure may specify the connectivity between various nodes and the organization of the nodes into layers. For example, the nodes of a first layer (e.g., the input layer) may receive data as input data or application data. Such data may include, for example, one or more pixels per node when the trained model is used, for example, to analyze an initial image. Subsequent intermediate layers may receive the outputs of the nodes of the previous layer as input, according to the connectivity specified in the model format or structure. These layers may also be referred to as hidden layers. For example, the first layer may output segmentation between the foreground and background. The final layer (e.g., the output layer) produces the output of the machine learning model. In some embodiments, the model format or structure also specifies the number and / or type of nodes in each layer.

[0020] In another embodiment, the trained model can include one or more models. One or more of the models can include a plurality of nodes arranged in layers according to a model structure or format. In some embodiments, the nodes can be computational nodes without memory, configured, for example, to process one unit of input to produce one unit of output. The computation performed by the nodes can include, for example, multiplying weights to each of a plurality of node inputs, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce a node output. In some embodiments, the computation performed by the nodes can also include applying a class / activation function to the adjusted weighted sum. In some embodiments, the class / activation function can be a non-linear function. In various embodiments, such computations can include operations such as matrix multiplication. In some embodiments, the computation by a plurality of nodes can be performed in parallel, for example, using a plurality of processor cores of a multi-core processor, individual processing units of a graphics processing unit (GPU), or a dedicated neural circuit. In some embodiments, the nodes can include memory and, for example, can be capable of storing and using one or more previous inputs when processing subsequent inputs. For example, a node having memory can include a long short-term memory (LSTM) node. The LSTM node can use the memory to maintain a "state" that enables the node to function like a finite state machine (FSM).

[0021] In some embodiments, the trained model may include embeddings or weights of individual nodes. For example, the model may start as a plurality of nodes organized into layers as specified by the model format or structure. In initialization, each weight may be applied to the connection between nodes that are connected according to the model format, e.g., each pair of nodes in successive layers of a neural network. For example, each weight may be randomly assigned or initialized to a default value. Next, the model may be trained, e.g., using training data, to produce a result.

[0022] Training may include applying supervised learning techniques. In supervised learning, the training data can include a plurality of inputs (e.g., templates), and a ground truth output corresponding to each input (e.g., a corresponding descriptive text). Based on the comparison between the output of the model and the ground truth output, the values of the weights are automatically adjusted, e.g., to increase the probability that the model generates the ground truth output of an image.

[0023] In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In some embodiments, the trained model may include a set of fixed weights downloaded from, e.g., a server that provides the weights. In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In embodiments where data is omitted, the trained model may be based on, e.g., prior training by a developer, a third party, etc. In some embodiments, the trained model may include a set of fixed weights downloaded from, e.g., a server that provides the weights.

[0024] Database 199 may store machine learning models, training datasets, media items, etc. Database 199 may also store social network data associated with user 125, user preferences of user 125, etc.

[0025] The user device 115 may be a computing device that includes memory coupled to a hardware processor. For example, the user device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or other electronic device that can access the network 105.

[0026] In the illustrated embodiment, user device 115a is connected to network 105 via signal line 108, and user device 115n is connected to network 105 via signal line 110. Media application 103 may be stored as media application 103b on user device 115a and / or as media application 103c on user device 115n. Signal lines 108 and 110 may be wired connections such as Ethernet, coaxial cable, or fiber optic cable, or wireless connections such as Wi-Fi®, Bluetooth®, or other wireless technologies. User devices 115a and 115n are accessed by users 125a and 125n, respectively. User devices 115a and 115n in Figure 1 are used as examples. Although Figure 1 shows two user devices 115a and 115n, this disclosure applies to system architectures having one or more user devices 115.

[0027] The media application 103 may be stored in the media server 101 or the user device 115. In some embodiments, the operations described herein are performed in the media server 101 or the user device 115. In some embodiments, some operations may be performed on the media server 101 and some on the user device 115.

[0028] The operation is performed according to user settings. For example, user 125a may specify that the user's images and / or other data should be stored locally only on the user device 115a and not on the media server 101.

[0029] The transmission of user data (e.g., templates, attributes, etc.) to the media server 101, any temporary or permanent storage of such data by the media server 101, and the performance of actions on such data by the media server 101 will only occur if the user consents to the transmission, storage, and performance of actions by the media server 101. User data from the media will not be used in advertising, responses will not be reviewed by humans unless the user provides feedback or addresses abuse or harm, and user data will not be used to train machine learning models other than those used to provide suggested search queries. Users are given the option to change their settings at any time, for example, they can enable or disable the use of the media server 101.

[0030] In some embodiments, a machine learning model (e.g., a neural network or other type of model) is stored locally on the user device 115 and used with the permission of a specific user when used for one or more operations. Server-side models are used only when permitted by the user. Furthermore, trained models may be provided for use on the user device 115. During use, on-device training of the model may be performed if permitted by user 125. If permitted by user 115, updated model parameters may be sent to the media server 101, for example, to enable federative learning. Model parameters do not include any user data.

[0031] The media application 103 identifies attributes from a group of media items associated with a user account. Attributes of media items (e.g., images, documents, video files / data, audio files / data, and / or video and audio files / data) may include one or more of the following: objects, people, landmarks, actions, etc. Attributes can be identified by performing independent component analysis.

[0032] The media application 103 generates a template with a combination of attributes. The template may also be a set of attributes. For example, the first template may include {Sara, John, climbing}, and the second template may include {John, eating}.

[0033] The media application 103 scores each template based on the number of attribute categories in the corresponding template and the number of media items in the group of media items containing the corresponding attributes. For example, the first template has three attribute categories (i.e., Sara, John, and climbing). In one embodiment, generating a score may include calculating a minimum threshold based on the number of attribute categories in the corresponding template and multiplying that threshold by the number of media items. For example, for the first template, the minimum threshold may be 0.66 based on the number of attribute categories being three, and the score is the product of the minimum threshold and the number of media items, i.e., 0.66 × 4 = 2.64. For the second template, the second template has two attribute categories (i.e., John and eating), and the minimum threshold is 0.5, resulting in a score of the product of the minimum threshold and the number of media items, i.e., 0.5 × 1 = 0.5.

[0034] The media application 103 selects one or more templates based on the corresponding scores. In the first example, a score of 2.64 results in the need for at least three media items containing Sara, John, and climbing. In the second example, a score of 0.5 results in the need for at least one media item containing John and eating. Since only two images contain Sara, John, and climbing, and one of the images contains John and eating, the media application 103 selects the second template.

[0035] The media application 103 provides one or more templates as input to the machine learning model 120. The machine learning model 120 outputs descriptive text based on one or more templates. For example, the machine learning model 120 may output "John is eating." The media application 103 provides the descriptive text to the user as a suggested search query.

[0036] The example above shows a simple query, but queries can become complex in various cases, such as "John eating a hamburger on a sidewalk in Manhattan," "John eating shrimp noodles at a Japanese restaurant," or "John eating a taco and holding a soda in his other hand, with a taco stand in the background."

[0037] In some embodiments, the media application 103 may be implemented using hardware including a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, the media application 103a may be implemented using a combination of hardware and software.

[0038] Computing devices Figure 2 is a block diagram of an exemplary computing device 200 that may be used to perform one or more of the features described herein. The computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, the computing device 200 is a media server 101 used to run a media application 103a. In another example, the computing device 200 is a user device 115.

[0039] In some embodiments, the computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, display 241, camera 243, and storage device 245, all coupled via a bus 218. The processor 235 may be coupled to the bus 218 via signal line 222, the memory 237 via signal line 224, the I / O interface 239 via signal line 226, the display 241 via signal line 228, the camera 243 via signal line 230, and the storage device 245 via signal line 232.

[0040] The processor 235 may be one or more processors and / or processing circuits for executing program code and controlling the basic operation of the computing device 200. “Processor” includes any suitable hardware system, mechanism, or component for processing data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., single-core, dual-core, or multi-core configurations), multiple processing units (e.g., in a multi-processor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a composite programmable logic device (CPLD), dedicated circuitry for implementing a function, a dedicated processor for performing processing based on a neural network model, a neural circuit, a processor optimized for matrix calculations (e.g., matrix multiplication), or a system having other such systems. In some embodiments, the processor 235 may include one or more coprocessors for performing neural network processing. In some embodiments, the processor 235 may be a processor that processes data to produce a probabilistic output, for example, the output produced by the processor 235 may be inaccurate or accurate within a range from an expected output. The processing does not need to be limited to a specific geographical location or have temporal constraints. For example, the processor may perform its functions in real time, offline, or batch mode. Multiple parts of the processing may be performed at different times and in different locations by different (or the same) processing systems. The computer can be any processor that communicates with memory.

[0041] Memory 237 may be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erase-read-only memory (EEPROM), or flash memory, which is typically located separately from and / or integrated with the processor 235, and is provided in the computing device 200 for access by the processor 235 and is suitable for storing instructions for execution by the processor or a set of processors. Memory 237 can store software running on the computing device 200 by the processor 235, including media applications 103.

[0042] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, and the like. One or more methods disclosed herein can operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application ("App") that runs on a mobile computing device, and so on.

[0043] Application data 266 may be data generated by other applications 264 or the hardware of computing device 200. For example, application data 266 may include images used by an image library application and user actions identified by other applications 264 (e.g., a social networking application).

[0044] The I / O interface 239 can provide the functionality to allow the computing device 200 to interface with other systems and devices. Interfaced devices may be included as part of the computing device 200, or they may be separate and communicate with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 245), and input / output devices can communicate via the I / O interface 239. In some embodiments, the I / O interface 239 can be connected to interface devices such as input devices (keyboards, pointing devices, touchscreens, microphones, scanners, sensors, etc.) and / or output devices (display devices, speaker devices, printers, monitors, etc.).

[0045] Some examples of interface devices that can be connected to the I / O interface 239 include a display 241 that can be used to display content, such as images, videos, and / or user interfaces of output applications described herein, and to receive touch (or gesture) input from the user. For example, the display 241 may be used to display a user interface, including graphical guides, on a viewfinder. The display 241 may include any suitable display device, such as a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, cathode ray tube (CRT), television, monitor, touchscreen, three-dimensional display screen, or other visual display device. For example, the display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in the shape factor of eyeglasses or a headset device, or a monitor screen of a computer device.

[0046] Camera 243 may be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video that the I / O interface 239 sends to the media application 103.

[0047] The storage device 245 is a database that stores data related to the media application 103. For example, the storage device 245 may store a training dataset containing labeled images, a machine learning model, the output from the machine learning model, media items, and so on.

[0048] Media applications Figure 2 shows an exemplary media application 103 stored in memory 237. In some embodiments, the media application 103 includes a user interface module 202, an event module 204, an attribute module 206, and a template module 208.

[0049] The user interface module 202 generates a user interface that displays information about media items. Media items may include images, videos, screenshots, etc. Media items may be captured by the camera 243 or obtained from other sources. Media items may be displayed chronologically, according to groups created by the event module 204, etc.

[0050] In some embodiments, the user interface module 202 generates a user interface that includes questions regarding the identification of people and / or pets appearing in media items. For example, the user interface may ask the user to identify the names of people in the images and their relationship to the user. The user interface module 202 may ask for identification of people and / or pets based on the presence of people and / or pets in a predetermined number of images (e.g., 10%). The user interface module 202 will ask for permission to use user data before asking for the identity of people and / or pets in media items, unless such permission has been previously provided by the user.

[0051] The user interface includes options for performing a search for media items. In some embodiments, the user interface generates a suggested search query for the user. The user interface module 202 requests permission from the user before generating the suggested search query. If the user does not provide permission for the use of their user data, the media application 103 does not generate a suggested search query.

[0052] Event module 204 generates groups of media items. These groups may be based on various factors, such as periodic grouping of media items (e.g., monthly, weekly, yearly) based on the date associated with the media item (e.g., capture data, last modified date). Event module 204 may also generate groups of media items based on different events. For example, event module 204 may generate groups of media items based on travel (e.g., a trip to London), events (e.g., a wedding), people or pets (e.g., the user's children over the years), or themes (e.g., a beach adventure, a running trail throughout the season, the progress of a home renovation plan, etc.).

[0053] In some embodiments, the event module 204 outputs a group of media items based on a user request. For example, a user may request media items associated with a specific person. The event module 204 generates a group containing all media items that include the specific person (e.g., images or videos depicting the person).

[0054] In some embodiments, the event module 204 includes a machine learning model trained to receive media items as input and output groups of media items based on different events. The machine learning model may include different types of machine learning models, such as the examples described above with reference to the machine learning model in Figure 1. In some embodiments, the event machine learning model is a classifier that receives media items along with information about the media items and uses the information to output an event signal indicating that an event may have occurred. The information may include the results of performing optical character recognition on an image to identify text in the image that indicates a particular event. For example, an image of a menu captured at dinner might include the term "wedding" on it. In some embodiments, the information may include the results of performing object recognition to identify objects associated with the event. For example, a media item captured at a baby shower might have a gift associated with the baby.

[0055] In some embodiments, metadata associated with a media item may be provided as additional input to a machine learning model if the user permits such use of the metadata. The metadata may include the location and / or time of the video capture, whether the video was shared via a social network, image sharing application, messaging application, etc., depth information associated with one or more video frames, sensor values ​​of one or more sensors of the camera that captured the video, e.g., accelerometer, gyroscope, light sensor, or other sensors, and user permission factors such as the user's identity (if the user consents). For example, if the video was captured at night in an outdoor location with the camera pointing upwards, such metadata may be associated with an astronomical event, as it may indicate that the camera was pointing to the sky at the time the video was captured.

[0056] In some embodiments, the event module 204 trains a machine learning model to identify groups of images based on event predictions. In some embodiments, the event machine learning model may use a combination of metadata, optical character recognition, and other signals as input to the event machine learning model. For example, the metadata might indicate that a person is using firecrackers and the date is July 4th, and as a result, the event machine learning model outputs an event signal corresponding to Independence Day. In some embodiments, the event machine learning model also outputs the event type for one or more events. Continuing the above example, the event machine learning model outputs an event signal and the probability that the event is Independence Day or a public holiday.

[0057] The attribute module 206 identifies attributes from a group of media items associated with a user account. Attributes may include people and pets in the media items (e.g., depicted in pixels of an image media item, described in the attributes of a text or text media item), the location where the media item was captured, the time the media item was taken, objects within the media item, text within the media item, events or activities associated with the media item, and / or actions performed in the media item (e.g., eating, swimming, running). In some embodiments, one or more attributes are determined from labels provided by the user, such as when the user identifies different people in an image. The attribute module 206 may use attributes within media items to infer the identification of other attributes. For example, an image of a user in front of a restaurant with the restaurant's name may be used by the attribute module 206 to identify an event or action involving eating, and the restaurant's setting.

[0058] In some embodiments, the attribute module 206 performs independent component analysis of the media item to identify objects. Independent component analysis is a technique used to separate a mixed signal into its independent components. In some embodiments, the attribute module 206 performs object recognition to identify objects within the media item.

[0059] In response to obtaining user consent, attribute module 206 may pre-calculate the attributes of a group of media items. For example, attribute module 206 may pre-calculate the attributes of media items monthly, etc., in response to a user creating a photo album, or in response to attribute module 206 generating a specific group of media items (e.g., a group of media items on a theme such as hiking, vacation, or events). Pre-calculating attributes favorably reduces the time between a user requesting a suggested search query and receiving a suggested search query generated by media application 103.

[0060] In some embodiments, the attribute module 206 generates a histogram that includes the identification of attributes within each media item and the number of times each attribute appears within a group of media items. Table 1 includes an example of attributes identified from six media items, and Table 2 includes a count of the number of times each attribute appears within all six media items.

[0061] [Table 1]

[0062] [Table 2]

[0063] Template module 208 generates templates with combinations of attributes. In some embodiments, template module 208 generates templates using a fixed set of attribute categories. The attribute categories may include different combinations of attribute categories such as PERSON, PLACE, DATE, EVENT, ACTIVITY, and SCENE. Template module 208 may combine attribute categories in specific combinations derived from common user search queries. For example, templates may include "PERSON,PERSON,ACTIVITY", "PERSON,PERSON,EVENT", "EVENT,DATE", "PERSON,ACTIVITY,DATE", etc. Continuing with the example from the table above, template module 208 generates a template with attributes by inputting {"person1", "person2", "skiing"} into the {PERSON,PERSON,EVENT} template.

[0064] In some embodiments, the template module 208 calculates a minimum threshold based on the number of attribute categories in the corresponding template. For example, the minimum threshold for two attribute categories is 0.5, for three attribute categories it is 0.66, and for four attribute categories it is 0.75.

[0065] In some embodiments, the minimum threshold is calculated using the following formula:

[0066]

number

[0067] In the formula, N is the number of categories, and N ≥ 2. The minimum threshold formula uses the pigeonhole principle to ensure that there is at least one media item that matches the generated combination. In some embodiments, 2 ≤ N ≤ 4.

[0068] The template module 208 scores each template based on the number of attribute categories in the corresponding template and the number of media items in the group of media items that contain the corresponding attribute. If the number of media items containing the corresponding attribute exceeds the corresponding score, the template module 208 selects the template. In some embodiments, if multiple people are included in a template, the multiple people do not need to be included in the same media item in order to be considered in the number of media items that have the attribute.

[0069] Figure 3 shows an exemplary user interface 300 of a group of media items according to some embodiments described herein. The group of media items includes a first image 305 of person 1 skiing, a second image 307 of person 1 and person 2 skiing, a third image 309 of person 2 climbing, a fourth image 311 of person 2 skiing while person 1 is present with two other spectators, a fifth image 313 of person 1 kayaking, and a sixth image 315 of person 1 and person 2 skiing.

[0070] Continuing with the above example, for the template "PERSON,PERSON,ACTIVITIES", there are three attribute categories, which correspond to a minimum threshold of 0.66. For a group of six media items, the attribute must appear in at least four media items, since 6 * 0.66 = 4 media items. Since person1, person2, and skiing appear in four media items (305, 307, 311, 315), they are selected as the template {"person1", "person2", "skiing"}. Since climbing and kayaking appear in only one media item, {"person1", "climbing"} and {"person2", "kayaking"} are not selected as templates.

[0071] The template module 208 provides one or more templates as input to the machine learning model 120. Continuing the example above, the machine learning model 120 receives {"person1", "person2", "skiing"} as input and outputs "person1 and person2 skiing together". In other examples, {"person2", "climbing", "December", "2023"} is provided as input and the machine learning model 120 outputs "person2 climbing in December 2023", {"person1", "person2", "wedding"} is provided as input and the machine learning model 120 outputs "person1 and person2 during their wedding", and {"person1", "person2", "Christmas"} is provided as input and the machine learning model 120 outputs "person1 celebrating Christmas with person2".

[0072] In some embodiments, the template module 208 provides one or more templates along with prompts, which are based on at least one feature selected from the group consisting of capitalization emphasis, repetition emphasis, repetitive prompting, negative instructions, structure, constraints, and delimiters. Figure 4 shows an exemplary prompt 400 sent to the machine learning model 120 along with one or more templates according to some embodiments described herein.

[0073] In some embodiments, the template module 208 provides the machine learning model 120 with one or more templates along with a level of randomness. The level of randomness may be a numerical scale (e.g., 0 for no randomness, 5 for moderate randomness, and 10 for most random), a word scale (e.g., no randomness, a little randomness, moderate randomness, very random, etc.), or another paradigm. The level of randomness is used by the machine learning model 120 to modify the descriptive text. For example, a low level of randomness may result in concatenating the search results in a natural language style, such as the example "person1 celebrating Christmas with person2" above.

number

[0074] In some embodiments, the template module 208 provides the machine learning model 120 with one or more templates along with a specified level of complexity. The level of complexity may be a numerical scale (e.g., 0 for simplest, 5 for moderate complexity, and 10 for most complex), a word scale (e.g., simplest, somewhat complex, moderate complexity, very complex, etc.), or another paradigm. The level of complexity is used by the machine learning model 120 to modify the descriptive text. For example, a lower level of complexity may include descriptive text with fewer words added to the template than a higher level of complexity. The level of complexity and the level of randomness may be specified based on user preferences (e.g., input via a user interface generated by the user interface module 202), feedback from search results, etc. In some embodiments, if the user rejects proposed search results with high complexity but accepts proposed search results with lower complexity, the template module 208 may set the level of complexity to the level at which the user is most likely to accept the proposed search results.

[0075] The machine learning model 120 outputs descriptive text. The user interface module 202 may provide the user with the descriptive text as a suggested search query in the user interface for media items associated with the user account, or use the descriptive text to evaluate search quality, etc. In some embodiments, the user interface module 202 receives a selection of suggested search queries from the user via the user interface and provides the user with search results that match the descriptive text from a database associated with the media items associated with the user account, such as a database that is part of the storage device 245.

[0076] In some embodiments, the user may provide feedback used to refine the user's preferences for suggested search queries. For example, if the user modifies a suggested search query, the modification may be used by the machine learning model 120 and / or template module 208 to improve suggested search queries for the user in the future.

[0077] Exemplary User Interface Figure 5 shows two exemplary user interfaces 500, 525, including suggested search queries, according to some embodiments described herein. In the first user interface 500, the user interface module 202 generates a list of suggested search queries 502 based on a group of media items 510. The user may enter their own search query in the text field 505 based on having been inspired by the list of suggested search queries 502, or they may select one of the suggested search queries 502.

[0078] In the second user interface 525, the user interface module 202 generates a list of suggestions 531 based on two different groups of media items 527, 529. Each suggestion is associated with a corresponding search button 533, 535, 537 for selecting a specific search suggestion.

[0079] Figure 6 shows an exemplary user interface 600, including suggested search queries based on an initial search request from a user, according to several embodiments described herein. In this example, the user provides a request in the text field 602 for a media item containing one or more specific attributes. In this case, the specific attributes are for a media item containing the user's daughter, Ava.

[0080] The user interface module 202 provides suggested search queries 605, corresponding search buttons 612 and 614 for further refinement of search results, and search results 610. The list of suggestions 605 is based on attributes identified in media items corresponding to the initial search for media items containing the user's daughter.

[0081] Exemplary Method Figure 7 shows a flowchart of an exemplary method 700 for generating a proposed search query. Method 700 can be performed by the computing device 200 in Figure 2. In various embodiments, Method 700 is performed by the user device 115, the media server 101, or partially on the user device 115 and partially on the media server 101.

[0082] In Figure 7, a group of media items 705 is used by a media application 103 to extract and aggregate attributes 710. An attribute histogram 715 is used 725 to create a template from the attributes and generate a template. The generated template 725 (based on a group of various media items) is used by a template generator 730 to provide the template 725 to a large language model 735. The large language model 735 returns a natural language style search query 740. The natural language style search queries 740 may be used 745 to evaluate search quality. For example, they may be used as model examples of search quality used to teach the user how to produce search results. The natural language style search queries 740 may also be provided to the user as search suggestions 750.

[0083] Figure 8 shows a flowchart of another exemplary method 800 for generating the proposed search query. Method 800 may be performed by the computing device 200 in Figure 2. In various embodiments, flowchart 800 is performed by the user device 115, the media server 101, or partly on the user device 115 and partly on the media server 101.

[0084] Method 800 in Figure 8 may begin with block 802. In block 802, a request is received for a proposed search query for a group of media items associated with a user account. In some embodiments, the media items include one or more screenshots captured by a user device associated with the user account. Block 804 may follow block 802.

[0085] In block 804, the authorization interface element is made visible. For example, the media application 103 may display the authorization interface element before fulfilling the request for the proposed search query. In some embodiments where the proposed search query is provided without the user requesting the proposed search query, the authorization interface element is displayed before the media application displays the proposed search query. Block 804 may be followed by block 806.

[0086] In block 806, it is determined whether permission for user data has been granted by the user. If permission has not been granted by the user, block 808 follows block 806, and method 800 terminates there. If permission has been granted, block 810 may follow block 806.

[0087] In block 810, attributes from a group of media items associated with the user are identified. In some embodiments, attributes that are part of a group of media items associated with the user account are pre-calculated. Block 810 may be followed by block 812.

[0088] In block 812, a template with a combination of attributes is generated. Block 814 may follow block 812.

[0089] In block 814, each template is scored based on the number of attribute categories in the corresponding template and the number of media items in the group of media items that contain the corresponding attribute. In some embodiments, method 800 further includes determining how many times each attribute appears in the media items of the group of media items, and scoring the templates is based on how many times each corresponding attribute appears in the media items. In some embodiments, method 800 further includes determining how many media items contain the corresponding attribute and, for each template, calculating a minimum threshold based on the number of attribute categories in the corresponding template, and scoring each template includes multiplying each template by the minimum threshold corresponding to the number of media items, and selecting one or more templates based on the corresponding score includes selecting one or more templates in response to the number of media items containing the corresponding attribute exceeding the corresponding score. Block 814 may be followed by block 816.

[0090] In block 816, one or more templates are selected based on their corresponding scores. Block 818 may follow block 816.

[0091] In block 818, one or more of the selected templates are provided to the large language model as input. In some embodiments, one or more templates are provided to the large language model along with prompts, which are based on features selected from a group consisting of capitalization emphasis, repetition emphasis, repetitive prompting, negative instructions, structure, constraints, delimiters, and combinations thereof. In some embodiments, one or more templates are provided to the large language model along with a level of randomness to be associated with the descriptive text. In some embodiments, one or more templates are provided to the large language model along with a level of complexity to be associated with the descriptive text. Block 818 may be followed by block 820.

[0092] In block 820, the large-scale language model outputs descriptive text based on one or more templates. Block 822 may follow block 820.

[0093] In block 822, the descriptive text is provided to the user as a suggested search query in a user interface for media items associated with a user account. In some embodiments, method 800 further includes receiving a selection of suggested search queries from the user via the user interface and providing the user with search results from a database of media items associated with the user account that match the descriptive text. In some embodiments, method 800 further includes receiving a request from the user for media items containing one or more specific attributes and providing the user with a group of media items based on a group of media items containing one or more specific attributes, the suggested search query is provided along with the group of media items.

[0094] In addition to the above description, users may be provided with controls that allow them to make choices regarding whether and when the systems, programs, or functions described herein may enable the collection of user information (e.g., information about the user's social networks, social behavior, or activities, occupation, user preferences, or the user's current location), and whether content or communications are sent to the user from the server. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, a user's identity may be processed so that personally identifiable information cannot be determined, or if location information is obtained (e.g., at the city, zip code, or state level), the user's geographic location may be generalized so that the user's specific location cannot be determined. Thus, users may control what information is collected about them, how that information is used, and what information is provided to them.

[0095] In the above description, many specific details have been included for illustrative purposes to provide a complete understanding of this specification. However, it will be apparent to those skilled in the art that this disclosure can be implemented without these specific details. In some cases, structures and devices are shown in block diagram form to avoid obscuring this specification. For example, embodiments may be described above with reference primarily to user interfaces and specific hardware. However, embodiments can be applied to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.

[0096] Any reference in this specification to “some embodiments” or “some examples” means that certain features, structures, or characteristics described in relation to the embodiments or examples may be included in at least one embodiment of this specification. The phrase “in some embodiments” appearing in various places in this specification does not necessarily refer to the same embodiments.

[0097] Some parts of the detailed explanation above are presented in terms of algorithms and symbolic representations of operations on data bits in computer memory. These descriptions and representations of algorithms are means used by those skilled in the art to most effectively communicate the content of their work to others skilled in the art. Here, and also generally, an algorithm is considered to be a self-consistent set of steps that lead to a desired result. These steps are steps that require the physical manipulation of physical quantities. Usually, though not essential, these quantities take the form of electrical or magnetic data that can be stored, transferred, combined, compared, or otherwise manipulated. For reasons of general use, it is sometimes convenient to refer to these data as bits, values, elements, symbols, characters, terms, numbers, etc.

[0098] However, it should be recognized that all these terms and similar terms should correspond to appropriate physical quantities and are merely convenient labels applied to those quantities. Unless otherwise specified, as will become clear from the following discussions, throughout this specification, discussions using terms such as “process,” “calculate,” “compute,” “determine,” or “display” refer to the actions and processes of a computer system or similar electronic computing device that manipulate data represented as physical (electronic) quantities in the registers and memory of the computer system and convert it into other data similarly represented as physical quantities in the memory or registers of the computer system, or other such information storage devices, transmission devices, or display devices.

[0099] Embodiments of this specification may also relate to a processor for performing one or more steps of the methods described above. The processor may be a dedicated processor that is selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-temporary computer-readable storage medium, including but not limited to any type of disk including an optical disk, ROM, CD-ROM, magnetic disk, RAM, EPROM, EEPROM, magnetic card or optical card, flash memory including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0100] The specification may take the form of several entirely hardware embodiments, several entirely software embodiments, or several embodiments that include both hardware and software elements. In some embodiments, the specification is implemented with software including, but not limited to, firmware, resident software, and microcode.

[0101] Furthermore, the specification may take the form of a computer program product accessible from a computer-enabled medium or computer-readable medium, which provides program code for use by or in connection with a computer or any instruction execution system. For the purposes of this specification, the computer-enabled medium or computer-readable medium may be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, instruction execution unit, or instruction execution device.

[0102] A data processing system suitable for storing or executing program code would include at least one processor directly or indirectly connected to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, bulk storage, and cache memory providing temporary storage for at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution.

Claims

1. A method performed by a computer, the method is Identifying attributes from a group of media items associated with a user account, Inputting into a template having the aforementioned combination of attributes, Each template is scored based on the number of attribute categories within the corresponding template and the number of media items within the group of media items containing the corresponding attributes. Based on the corresponding score, select one or more of the aforementioned templates, Providing one or more templates as input to a large-scale language model, Using the aforementioned large-scale language model, output descriptive text based on one or more templates, A method comprising providing the descriptive text to the user as a suggested search query in a user interface for the media item associated with the user account.

2. Receiving the selection of the proposed search query from the user via the user interface, The method according to claim 1, further comprising providing the user with search results from a database of media items associated with the user account that match the descriptive text.

3. The method further includes determining the number of times each attribute appears in the media items of the group of media items, The method according to claim 1, wherein scoring the template is based on the number of times each corresponding attribute appears in the media item.

4. Determining the number of media items that include the corresponding attribute, For each template, the method further includes calculating a minimum threshold based on the number of attribute categories in the corresponding template, Scoring each template involves multiplying each template by the minimum threshold corresponding to the number of media items, The method according to claim 1, wherein selecting one or more of the templates based on the corresponding score includes selecting one or more of the templates in response to the number of media items containing the corresponding attribute exceeding the corresponding score.

5. The method according to claim 1, wherein the one or more templates are provided to the large language model along with a prompt, the prompt being based on features selected from a group consisting of capitalization emphasis, repetition emphasis, repetitive prompting, negative instruction, structure, constraint, delimiter, and combinations thereof.

6. The method according to claim 1, wherein the one or more templates are provided to the large language model along with a level of randomness to be associated with the descriptive text.

7. The method according to claim 1, wherein one or more templates are provided to the large language model along with a level of complexity to be associated with the descriptive text.

8. The method according to claim 1, further comprising pre-calculating the attributes which are part of a group of media items associated with the user account.

9. Receiving user requests for media items that include one or more specific attributes, The method according to claim 1, further comprising providing the group of media items to the user based on the group of media items having one or more specific attributes, wherein the proposed search query is provided together with the group of media items.

10. The method according to claim 1, wherein the media item includes one or more screenshots captured by a user device associated with the user account.

11. It is a system, One or more processors, The system comprises one or more processors and a memory coupled thereto, wherein the memory stores instructions that, when executed by the processors, cause one or more processors to perform an operation, and the operation is, Identifying attributes from a group of media items associated with a user account, Inputting into a template having the aforementioned combination of attributes, Each template is scored based on the number of attribute categories within the corresponding template and the number of media items within the group of media items containing the corresponding attributes. Based on the corresponding score, select one or more of the aforementioned templates, Providing one or more templates as input to a large-scale language model, Using the aforementioned large-scale language model, output descriptive text based on one or more templates, A system comprising providing the descriptive text to the user as a suggested search query in a user interface for the media item associated with the user account.

12. The aforementioned operation is, Receiving the selection of the proposed search query from the user via the user interface, The system according to claim 11, further comprising providing the user with search results from a database of media items associated with the user account that match the descriptive text.

13. The aforementioned operation is, The method further includes determining the number of times each attribute appears in the media items of the group of media items, The system according to claim 11, wherein scoring the template is based on the number of times each corresponding attribute appears in the media item.

14. The aforementioned operation is, Determining the number of media items that include the corresponding attribute, For each template, the method further includes calculating a minimum threshold based on the number of attribute categories in the corresponding template, Scoring each template involves multiplying each template by the minimum threshold corresponding to the number of media items, The system according to claim 11, wherein selecting one or more of the templates based on the corresponding score includes selecting one or more of the templates in response to the number of media items containing the corresponding attribute exceeding the corresponding score.

15. The system according to claim 11, wherein one or more templates are provided to the large language model along with a prompt, the prompt being based on features selected from a group consisting of capitalization emphasis, repetition emphasis, repetitive prompting, negative instructions, structure, constraints, delimiters, and combinations thereof.

16. A non-temporary computer-readable medium that, when executed by one or more computers, stores instructions that cause the one or more computers to perform an action, wherein the action is: Identifying attributes from a group of media items associated with a user account, Inputting into a template having the aforementioned combination of attributes, Each template is scored based on the number of attribute categories within the corresponding template and the number of media items within the group of media items containing the corresponding attributes. Based on the corresponding score, select one or more of the aforementioned templates, Providing one or more templates as input to a large-scale language model, Using the aforementioned large-scale language model, output descriptive text based on one or more templates, Non-temporary computer-readable media, including providing the descriptive text to the user as a suggested search query in a user interface for the media item associated with the user account.

17. The aforementioned operation is, Receiving the selection of the proposed search query from the user via the user interface, The computer-readable media according to claim 16, further comprising providing the user with search results from a database of media items associated with the user account that match the descriptive text.

18. The aforementioned operation is, The method further includes determining the number of times each attribute appears in the media items of the group of media items, The computer-readable medium according to claim 16, wherein scoring the template is based on the number of times each corresponding attribute appears in the media item.

19. The aforementioned operation is, Determining the number of media items that include the corresponding attribute, For each template, the method further includes calculating a minimum threshold based on the number of attribute categories in the corresponding template, Scoring each template involves multiplying each template by the minimum threshold corresponding to the number of media items, The computer-readable media according to claim 16, wherein selecting one or more of the templates based on the corresponding score includes selecting one or more of the templates in response to the number of media items containing the corresponding attribute exceeding the corresponding score.

20. The computer-readable medium according to claim 16, wherein the one or more templates are provided to the large language model along with a prompt, the prompt being based on features selected from a group consisting of capitalization emphasis, repetition emphasis, repetitive prompting, negative instructions, structure, constraints, delimiters, and combinations thereof.