Text generation from images using machine learning

JP2026527578APending Publication Date: 2026-08-14GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-12
Publication Date
2026-08-14

Smart Images

  • Figure 2026527578000001_ABST
    Figure 2026527578000001_ABST
Patent Text Reader

Abstract

Embodiments described herein relate to methods, devices, and computer-readable media for generating text. In some embodiments, the method includes providing a prompt to a Large Language Model (LLM). The prompt may include a task description and image metadata associated with one or more images. The method further includes providing user-specific data to the LLM. The method further includes having the LLM generate text in response to the prompt based on the task description, the metadata of one or more images, and the user-specific data. The text is personalized based on user account data and responds to the task description. The method further includes making the text available for display in a user interface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Application No. 63 / 519,735, filed on August 15, 2023, entitled "GENERATING TEXT FROM IMAGES USING MACHINE LEARNING", the content of which is hereby incorporated by reference in its entirety.

Background Art

[0002] Users can take pictures of images and videos using smartphones and cameras. The images and videos taken by the user can be stored in the user's image library. The user may organize the images into albums or collections, may edit the images, and may share the images with other users. The user may input text describing the image and / or the event associated with the image. Manually entering text can be cumbersome and inefficient.

[0003] The description of the background art provided herein is for the purpose of schematically presenting the context of the present disclosure. Within the scope described in this background art section, the achievements of the inventors named herein, as well as aspects of this document that may not meet the requirements of prior art at the time of filing, are not expressly or implicitly admitted as prior art to the present disclosure.

Summary of the Invention

[0004] Various embodiments relate to computer-based methods that include providing prompts to a Large-Scale Language Model (LLM), the prompts including a task description and metadata associated with one or more images, the one or more images being images from an image library associated with a user account. The method further includes providing data from the user account to the LLM, the data including user-specific data including one or more of a person's name, a place name, an activity type, an object name, and a date. The method further includes having the LLM generate text in response to the prompt based on the task description, the metadata of one or more images, and the user account data. The text is personalized based on the user account data and responds to the task description. The method further includes making the text available for display in a user interface.

[0005] In some embodiments, the method further includes generating captions for one or more images using an image caption model and providing the captions to an LLM. In these embodiments, the text generated by the LLM is further based on the respective captions.

[0006] In some embodiments, one or more images are located in an image album. In these embodiments, the method further includes receiving a user request for an album title for the image album and, in response to the user request, selecting a task description that specifies that text is to be used as the album title. The text includes one or more proposed album titles.

[0007] In some embodiments, the method further includes receiving a search query from a user and identifying one or more images from an image library associated with the user account that match the search query. In these embodiments, the user interface includes at least one of the one or more images.

[0008] In some embodiments, the method further includes receiving user input text and providing the user input text to an LLM, where the LLM generates text based at least in part on the user input text.

[0009] In some embodiments, the method further includes receiving user input text after making text and at least one of one or more images available for display in the user interface, and providing the user input text to the LLM. The method further includes having the LLM generate corrected text based at least partially on the user input text, and making the corrected text available for display in the user interface. In some embodiments, the method further includes automatically updating user account data based on the user input text.

[0010] In some embodiments, the method further includes accessing a knowledge repository based on metadata associated with one or more images to obtain additional information, and generating text is further based on that additional information.

[0011] In some embodiments, the user interface includes at least one of one or more images. In some embodiments, at least a portion of the text is superimposed on at least one of one or more images.

[0012] Some embodiments include a computing device comprising a processor and memory coupled to the processor for storing instructions. When an instruction is executed by the processor, it causes the processor to perform an action which includes providing a prompt to a Large Language Model (LLM), the prompt including a task description and metadata associated with one or more images, the one or more images being images from an image library associated with a user account. The action further includes providing data from the user account to the LLM, the data including user-specific data which includes one or more of the following: person names, place names, activity types, object names, and dates. The action further includes having the LLM generate text in response to the prompt based on the task description, the metadata of one or more images, and the user account data, the text which is personalized based on the user account data and responds to the task description. The action further includes making the text available for display in a user interface.

[0013] In some embodiments, the operation further includes generating captions for one or more images using an image caption model, providing the captions to an LLM, and the LLM generating text based on these captions.

[0014] In some embodiments, one or more images reside in an image album, and the operation further includes receiving a user request for an album title of the image album, and in response to the user request, selecting a task description that specifies that the text is to be used as the album title, and the text includes the proposed one or more album titles. In some of these embodiments, the operation further includes receiving a search query from the user, and identifying one or more images from the image library associated with the user account that match the search query, and the user interface includes at least one of the one or more images.

[0015] In some embodiments, the operation further includes receiving user input text and providing the user input text to the LLM, and the generation of text by the LLM is at least partially based on the user input text.

[0016] In some embodiments, the operation further includes receiving user input text after making text and at least one of one or more images available for display in the user interface; providing the user input text to the LLM; having the LLM generate corrected text based at least partially on the user input text; and making the corrected text available for display in the user interface.

[0017] In some embodiments, the operation further includes accessing a knowledge repository based on metadata associated with one or more images to retrieve additional information, and generating text is further based on that additional information.

[0018] Some embodiments include a non-temporary computer-readable medium storing instructions, which, when executed by a processor, cause the processor to perform an action including providing a prompt to a Large Language Model (LLM), the prompt including a task description and metadata associated with one or more images, the one or more images being images from an image library associated with a user account. The action further includes providing data from the user account to the LLM, the data including user-specific data including one or more of a person's name, a place name, an activity type, an object name, and a date. The action further includes having the LLM generate text in response to the prompt based on the task description, the metadata of one or more images, and the user account data, the text being personalized based on the user account data and responding to the task description. The action further includes making the text available for display in a user interface.

[0019] In some embodiments, the instruction causes the processor to perform further actions, which include generating captions for one or more images using an image caption model and providing the captions to the LLM, which then generates text based on the respective captions.

[0020] In some embodiments, an instruction causes a processor to perform further actions, the further actions including receiving a search query from a user and identifying one or more images from an image library associated with the user account that match the search query, and the user interface includes at least one of the one or more images. [Brief explanation of the drawing]

[0021] [Figure 1] This is a block diagram of an exemplary network environment that may be used in one or more embodiments described herein. [Figure 2]A block diagram showing an exemplary method according to some embodiments. [Figure 3] An exemplary system architecture according to some embodiments. [Figure 4] An exemplary image is shown. [Figure 5A] An exemplary user interface according to some embodiments is shown. [Figure 5B] An exemplary user interface according to some embodiments is shown. [Figure 6] A block diagram of an exemplary device that may be used in one or more embodiments described herein.

Mode for Carrying Out the Invention

[0022] The embodiments described herein relate to the use of machine learning techniques for generating text based on images and / or image metadata. In some embodiments, a large language model (LLM) may generate text in response to a prompt that includes a task description (e.g., "generate an album title for these images," "generate a summary description for these images," etc.) and metadata for a set of images. The LLM may also be provided with user-specific data, user input text, and / or image captions as inputs in addition to the prompt.

[0023] The described embodiments utilize knowledge about an image and a user associated with the image to automatically generate text for various tasks. For example, the task may be to generate a poem that describes a set of images (e.g., "My Summer (Poem)"), a description of a birthday party (e.g., "Jake turned 30. His close friends Ryan, Anil, and Gazala attended. The black forest cake was wonderful. Ryan played the piano. Jake received a guitar as a surprise present!") that can be used as a visual story of the birthday party along with multiple images, create actionable elements from an image (e.g., "This is a receipt from Store X for Item Y. The total expenditure was $10.24 including $1.24 in taxes. Click here to add to your expense report"), etc.

[0024] The described embodiments can automatically generate text for a given set of images based on information about the image and a user associated with the image. The described embodiments reduce the computational resources required to render a user interface that a user can use to input text.

[0025] FIG. 1 shows a block diagram of an exemplary network environment 100 that may be used in some embodiments described herein. In some embodiments, network environment 100 includes one or more server systems, such as server system 102 in the example of FIG. 1. Server system 102 can communicate, for example, with network 130. Server system 102 may include server device 104 and database 106 or other storage devices. In some embodiments, server device 104 may provide an image management application 156b and one or more machine learning models 158b. In FIG. 1 and the remaining figures, a letter following a reference number, such as "156a", represents a reference to an element having that particular reference number. A reference number in the text without a subsequent letter, such as "156", represents a general reference to an embodiment of an element having that reference number.

[0026] The network environment 100 may also include one or more client devices, such as client devices 120, 122, 124, and 126, that can communicate with each other and / or with the server system 102 via the network 130. The network 130 may be any type of communication network, including one or more of the Internet, a local area network (LAN), a wireless network, a switch or hub connection, etc. In some embodiments, the network 130 may include peer-to-peer communication between devices using, for example, a peer-to-peer wireless protocol (e.g., Bluetooth®, Wi-Fi Direct, etc.). An example of peer-to-peer communication between two client devices 120 and 122 is shown by arrow 132.

[0027] For ease of explanation, Figure 1 shows one block of server system 102, server device 104, and database 106, as well as four blocks of client devices 120, 122, 124, and 126. Blocks 102, 104, and 106 may represent multiple systems, server devices, and network databases, and the blocks may be provided in configurations different from those shown. For example, server system 102 may represent multiple server systems that can communicate with other server systems via network 130. In some embodiments, server system 102 may include, for example, a cloud host server. In some examples, database 106 and / or other storage devices may be provided in a server system block(s), which is separate from server device 104 and can communicate with server device 104 and other server systems via network 130.

[0028] Furthermore, there may be any number of client devices. Each client device may be any type of electronic device, such as a desktop computer, laptop computer, portable or mobile device, mobile phone, smartphone, tablet computer, television, TV set-top box or entertainment device, wearable device (e.g., display glasses or goggles, wristwatch, headset, armband, jewelry, etc.), personal digital assistant (PDA), media player, game device, etc. Some client devices may also have a local database similar to database 106 or other storage. In some embodiments, the network environment 100 may not have all of the illustrated components and / or may have other components, including other types of elements, instead of or in addition to the components described herein.

[0029] In various embodiments, end users U1, U2, U3, and U4 may communicate with and / or with the server system 102 using their respective client devices 120, 122, 124, and 126. In some examples, users U1, U2, U3, and U4 may interact through applications running on their respective client devices and / or the server system 102, and / or through network services implemented on the server system 102, such as social networking services or other types of network services. For example, each client device 120, 122, 124, and 126 may communicate data with one or more server systems, such as the server system 102.

[0030] In some embodiments, the server system 102 may provide appropriate data to client devices so that each client device can receive communication content or shared content uploaded to the server system 102 and / or network services. In some examples, users U1 to U4 may interact via audio or video conferencing, voice, video, or text chat, or other communication modes or communication applications.

[0031] The network services implemented by the server system 102 may include a system that enables users to perform various communications, form links and associations, upload and post shared content such as images (e.g., individual images, image collections such as image albums, stories when a series of images are shown sequentially, image works such as animations), text, video, audio, and other types of content, and / or perform other functions. For example, a client device can view received data such as content posts that are transmitted to or streamed to the client device and originate from different client devices (or directly from different client devices) via the server and / or network services, or originate from the server system and / or network services. In some embodiments, client devices can communicate directly with each other, for example, using peer-to-peer communication between client devices as described above. In some embodiments, “user” may include one or more programs or virtual entities and persons that interface with the system or network.

[0032] In some embodiments, any of the client devices 120, 122, 124, and / or 126 may provide one or more applications. For example, as shown in Figure 1, client device 120 may provide an image management application 156a. Client device 120 may also include one or more machine learning models 158a. Client devices 122-126 may also provide similar applications.

[0033] The image management application 156a may be implemented using the hardware and / or software of the client device 120. In different embodiments, the image management application 156a may be, for example, a standalone client application running on any of the client devices 120 to 126, or it may operate in conjunction with the image management application 156b provided on the server system 102. The image management application 156a may provide a variety of functions related to images and / or videos. For example, such functions may include taking images using a camera, programmatically analyzing images and assigning labels indicating the subject of the images, modifying images, storing images in an image library or database and providing a user interface for viewing images, and generating, viewing, and sharing image-based works (e.g., stories, animations, social media posts, etc., in which a series of images are shown sequentially) or image collections such as image albums.

[0034] In some embodiments, the image management application 156 may allow a user to manage a library or database for storing images. For example, a user may use the backup function of the image application 156a on a client device (e.g., any of client devices 120-126) to back up local images or videos on a client device to a server device, for example, server device 104. For example, a user may manually select one or more images or videos to back up, or specify backup settings that identify the images or videos to be backed up. Backing up images or videos to a server device may include, for example, coordinating with the image application 156b on server device 104 to send the images or videos to the server for storage on the server.

[0035] In some embodiments, client device 120 and / or client devices 122-126 may include one or more machine learning models 158. Client machine learning model 158a may be implemented using the hardware and / or software of client device 120. In various embodiments, machine learning model 158a may be directly available on any of the client devices 120-126, or it may operate in conjunction with machine learning model 158b provided on server system 102.

[0036] The machine learning model 158 may provide various functions related to digital maps. For example, such functions may include automatically generating image labels (e.g., based on recognizing one or more objects or entities in an image) and generating captions for images (e.g., an image caption model may be trained to generate short text suitable for use as a caption based on images, image metadata, and / or labels generated by other models or other techniques). In some embodiments, a language model (e.g., a large-scale language model) may be included in the machine learning model 158.

[0037] In some embodiments, the language model may be a machine learning model capable of generating responses to text prompts provided as input to the model. In some embodiments, the text prompts may include a task description and image metadata for a task, and the language model may generate output text. In some embodiments, the language model may provide an application programming interface (API) for other applications, allowing other applications, such as an image management application 156, to utilize the language model to generate text by providing prompts. The language model 158 may utilize data, such as images, image metadata including image labels, and data from user accounts (e.g., of a user associated with a client device 120). The data may be stored locally on the client device 120 and / or retrieved from a server device 104. In some embodiments, the language model may be a multimodal model capable of receiving non-text data as input, such as images, videos, binary files, or other types of data.

[0038] In different embodiments, the client device 120 and / or the server device 104 may provide other applications (not shown). These other applications may provide various types of functionality, such as calendars, address books, email, web browsers, shopping, transportation (e.g., taxi, train, and flight reservations), entertainment (e.g., music players, video players, and game applications), and social networking (e.g., messaging or chat, voice / video calls, and image / video sharing). In some embodiments, one or more of the applications may be standalone applications running on the client device 120. In some embodiments, one or more of the other applications may access a server system, such as the server system 102, which provides data and / or functionality for the other applications. In various embodiments, data associated with one or more of the other applications may be stored in a user account. With the user's permission, data in the user account may be provided to the machine learning model 158.

[0039] The user interfaces on client devices 120, 122, 124, and / or 126 may enable the display of user content and other content, including not only images, image albums, videos, data, and other content, but also communications, privacy settings, notifications, and other data. Such user interfaces can be displayed using software on the client devices, software on the server devices, and / or a combination of client and server software running on server device 104, for example, application software or client software that communicates with server system 102. The user interfaces can be displayed by the display devices of the client devices or server devices, such as touchscreens or other display screens, projectors, etc. In some embodiments, an application program running on the server system can communicate with the client devices to receive user input on the client devices and output data such as video data and audio data on the client devices.

[0040] Other embodiments of the features described herein can use any type of system and / or service. For example, other networked (e.g., internet-connected) services can be used instead of, or in addition to, social networking services. The features described herein can be utilized by any type of electronic device. Some embodiments can provide one or more of the features described herein on one or more client devices or server devices that are disconnected from or intermittently connected to a computer network. In some examples, a client device including or connected to a display device can display posts of content stored in a local storage device to the client device, for example, posts of content previously received over a communication network.

[0041] Images referred to herein may include digital images having pixels with one or more pixel values ​​(e.g., color values, lightness values, etc.). Images may be still images (e.g., still photographs, images having a single frame, etc.), dynamic images (e.g., animations, motion images, animated GIFs, cinemagraphs where part of the image is moving and other parts are static, etc.), or videos (e.g., images or sequences of image frames which may include sound). In the remainder of this specification, images will be referred to as static images, but it should be understood that the techniques described herein are applicable to dynamic images, videos, etc. For example, embodiments described herein can be used with still images (e.g., photographs, or other images), videos, or dynamic images.

[0042] Figure 2 is a flowchart illustrating an exemplary method 200 for generating text according to several embodiments. In some embodiments, method 200 can be performed, for example, on the server system 102 shown in Figure 1. In some embodiments, some or all of method 200 can be performed on one or more client devices 120, 122, 124, or 126 shown in Figure 1, one or more server devices, and / or on both server devices and client devices. In the examples described, the execution system includes one or more digital processors or processing circuits ("processors"), and one or more storage devices (e.g., database 106 or other storage). In some embodiments, one or more different components of the server and / or client can perform different blocks or other parts of method 200. In some examples, the first device is described as performing a block of method 200. In some embodiments, one or more blocks of method 200 can be performed by one or more other devices (e.g., other client devices or server devices) that can transmit results or data to the first device.

[0043] In some embodiments, Method 200, or part thereof, may be automatically initiated by the system. In some embodiments, the execution system is a first device. For example, Method (or part thereof) may be executed periodically, or on the basis of one or more specific events or conditions, such as a user request, the creation of an image album, the addition of new images and / or videos to the image library associated with the user account, the client device entering an idle state, a predetermined period of time elapsed since the last execution of Method 200, and / or the occurrence of one or other conditions that may be specified in the settings read by Method.

[0044] Method 200 may begin in block 202. In block 202, it is checked whether user consent (e.g., user permission) has been obtained for the use of user data in embodiments of Method 200. For example, user data may include images or videos stored on a client device (e.g., any of client devices 120-126), such as videos stored or accessed by the user using the client device, image metadata, an image library associated with a user account, user data related to the use of an image management application, user preferences, etc. One or more blocks of the method described herein may use such user data in some embodiments.

[0045] If user consent is obtained from the relevant user who may use user data in Method 200, in block 204 it is determined that the blocks of the Method herein may be executed with the use of user data available, as described for those blocks, and the Method proceeds to block 212. If user consent is not obtained, in block 206 it is determined that the blocks should be executed without using user data, and the Method proceeds to block 212. In some embodiments, if user consent is not obtained, the blocks are executed without using user data, using synthetic data and / or generic or publicly accessible and publicly available data. In some embodiments, if user consent is not obtained, the remaining blocks of Method 200 are not executed.

[0046] In block 212, a prompt is provided to the Large Language Model (LLM). The prompt may include a task description and the associated image metadata for one or more images.

[0047] In some embodiments, prompts may be selected from a prompt library. Each prompt in the prompt library may contain a task description. Some examples of such prompts include: "Generate an album title for an image with the following metadata," "Generate a short phrase summarizing the main attributes of an image with the following metadata," "Write a title based on the following data (including emojis)," "A catchy phrase based on the following data," "Suggest a phrase of 10 words or less for the following image," "Image data of images captured on the following trip. Generate a 10-line description of the trip," "These photos were taken at a wedding. Write a short story describing the bride and groom, the location, the decorations, and the food," and so on.

[0048] The prompt library may contain a set of prompts tested for use with the LLM. For example, such tests may be performed using a set of known images, using known attributes, and optionally using human-written text based on a set of known images. The human-written text may be written by a human as a response to a specific task description. Various test prompts may be provided to the LLM, and response text may be obtained from the LLM. The response text may be rated (e.g., assigned a score) based on known attributes and / or human-written text. The set of prompts may include test prompts that yield high-rated responses (e.g., satisfy a score threshold) for a set of known images, while excluding other test prompts.

[0049] In some embodiments, an automated evaluation may be performed based on the characteristics of the generated title and a set of known images, and a machine learning model may operate on that title to select prompts to include in the prompt library. In some embodiments, A / B testing may be used to determine which prompts to include in the library, for example, prompts that cause the LLM to output response text that the user should select (e.g., as a caption, as an album title, etc.). Excessive response text generated by the LLM in response to other prompts may also be included in the prompt library.

[0050] In some embodiments, the task description in the prompt may specify the text to be generated by the LLM and the context in which the generated text will be used. For example, the context of use may be an image caption (for a single image), individual image captions for two or more images, an album title for an album containing one or more images, a snippet describing a set of images, or a search summary of images that match a search query. In some embodiments, the prompt may specify attributes of the text, such as the length of the text, the language of the text, the style (e.g., funny, rhyming, formal), or the type (e.g., haiku, short poem).

[0051] In some embodiments, prompts provided to the LLM for a task may be customized (modified) based on image metadata (for example, with respect to images provided to the LLM along with the prompt). If an image contains certain metadata, the task prompt may be modified differently than if the image does not contain such metadata. The presence or absence of certain metadata, and the value of the metadata, may be used when modifying the prompt.

[0052] The prompt may also include image metadata for each of the one or more images associated with it. These one or more images may be from the image library associated with the user account. For example, the image library may include images captured by the user and automatically added to the library, images shared with the user by other users, and images manually added to the library by the user. The user's image library may contain various types of images (e.g., still images, motion images, animations, videos, screenshots, document images, etc.). Image metadata may include multiple attributes and associated values. These attributes and associated values ​​may be stored as the image metadata for the image.

[0053] For example, when an image is taken, the camera type and shooting settings (e.g., aperture, focal length), the date of shooting, and the location of shooting may be stored as image metadata. If the image is modified, the metadata may include attributes such as the time / date of modification and the modification tool (e.g., image editing application).

[0054] Furthermore, in some embodiments, images can be programmatically analyzed to determine various attributes of the image. For example, with user permission, face detection and / or face recognition techniques can be performed to identify faces in the image. If the user allows it, a name associated with the face (e.g., specified by the user for the image or from a previous image in which a face was depicted) can be stored as metadata.

[0055] In some embodiments, metadata may also include label attributes having multiple values. The values ​​of label attributes may indicate various aspects of the image, such as the type of environment (indoor / outdoor, sunny, cloudy, rainy, etc.), objects in the image (e.g., trees, sky, dog, etc.), image type (e.g., photograph, animation, video, screenshot, receipt, document, etc.), or any other value. Labels may be determined based on various image processing techniques such as object recognition, image segmentation, face detection, and face recognition. In some embodiments, attributes may have associated confidence scores, with higher confidence scores associated with a higher probability of the attribute value being accurate. An example of image metadata is as follows: {face: Jake, Kali; location: 44.38867, -121.2155; date: July 4, 2021; labels: dog, animal, pet, child, nature, cloudy, summer, tree, sky, eye, plant, sand; camera: f / 2.01 / 3602.22mm ISO35}.

[0056] Block 214 may follow Block 212.

[0057] In block 214, user account data is provided to the LLM. In some embodiments, the data may include user-specific data such as one or more of the following: person names, place names, activity types, object names, and dates. For example, user-specific data may be stored as part of the user account. Person names may include the user's name and the names of other users known to the user (e.g., spouse, children, parents, relatives, friends, colleagues, etc.). In some embodiments, person names may include the names of pets or other animals. In some embodiments, person names may be provided by the user or obtained (with permission) from the user's contact list, etc.

[0058] In some embodiments, location names in user account data may include places associated with the user. For example, location names may include information describing the user's home (e.g., home address, city, state, country, etc.), workplace, the homes of the user's acquaintances, and places the user has visited (e.g., restaurants, transit points, stadiums, auditoriums, museums, tourist attractions, outdoor locations, etc.). In some embodiments, activity types may include activities the user engages in or is interested in, such as running, skating, cycling, skiing, hiking, dancing, singing, etc. In some embodiments, object names may include objects the user owns or likes, such as models of cars, bicycles or other vehicles, musical instruments, books, etc.

[0059] In some embodiments, the dates in the user account data may include dates that are important to the user, such as birthdays (for the user, spouse, children, parents, relatives, friends, colleagues, etc.), anniversaries (e.g., wedding anniversaries, work anniversaries, etc.), and event dates (e.g., concerts, trips, sporting events, etc. that the user attended).

[0060] In various embodiments, user account data may include data provided by the user and / or data automatically determined from user activity, if permitted by the user. For example, a user may provide their home address to a digital map application. Automatic determination of data from user activity may be performed with the permission of a particular user, using any appropriate analytical technique such as natural language processing, text analysis, semantic analysis, text summarization, or media summarization. For example, if a user's query to a digital map includes a place name, the place name may be included in the user account data. In another example, if a user saves reminders about a flight, the flight date and origin and / or destination may be included in the user account data.

[0061] User information is accessed and processed according to user-provided settings and regulations, subject to specific user permissions. Users are given options to specify what data may be stored, for how long, and for what purpose such data will be used. User account data is provided to the LLM for the specific purpose of generating text. Block 214 may be followed by Block 216.

[0062] In block 216, captions may be generated on one or more images and provided to the LLM. For example, one or more images (and associated metadata) may be provided to an image caption model, e.g., a machine-trained model capable of generating captions (e.g., short strings) based on image input. Each caption may correspond to a specific image among the one or more images and may indicate the subject of the image. In some embodiments, block 216 may not be performed. Block 218 may follow block 216.

[0063] In block 218, user input text may be received and provided to the LLM. For example, user input text may be received via keyboard, as voice input, or through the selection of user interface elements. For example, the user may provide text that provides additional information about one or more images (which may not be present in, for example, image metadata and / or user account data). Such additional information may include, for example, identifying people present in the images (e.g., "The bearded man in the green jacket is Ryan," "The horse I'm riding is Daisy"), objects in the images (e.g., "The boat is Ryan's boat, named Voyager"), or other information (e.g., "These photos are of Ryan and Gazala's wedding," "These are receipts for my trip to London," etc.). In some embodiments, with the user's permission, user account data may be updated based on user input text. For example, the name "Ryan" may be associated with "The bearded man in the green jacket" in the user's images. Block 220 may follow block 214.

[0064] In block 220, the text is generated by the LLM. The text responds to prompts and is based on the task description, metadata for one or more images, and user account data. The text is personalized based on the user account data and responds to the task description. If a caption is provided to the LLM (by the execution of block 216), or if user input text is provided to the LLM (by the execution of block 218), the caption and / or user input text are also provided to the LLM.

[0065] For example, in the case of an image containing a boy named Jake and a dog named Kali, data from the user account may indicate that the user has several photos of Jake and Kali (e.g., multiple images in the user's image library labeled "Jake" and "Kari"). Furthermore, image metadata (e.g., location coordinates) may indicate that the image was taken in Telbon, Oregon. The user's image library may not contain previous images of Jake and Kali in Oregon, which suggests that this is likely their first time in Oregon. The age of the boy named Jake may be specified in the user account data. The date of the image (July 4, 2021) suggests that this is likely a holiday (and the user account data may indicate that the user's home is in another state). The presence of sand indicates that the boy and dog are playing in the sand. The image metadata and data from the user account provide contextual information to the LLM, and the task description provides an indication of the desired output type.

[0066] For example, in response to a task description in a prompt indicating that the text should be a caption of fewer than 10 words for a single image, the LLM may generate personalized text such as "Jake and Kali are building a sandcastle" or "Jake and Kali are playing in the sandbox." In some embodiments, the generated text may be, for example,

number

number

number

number

number

[0067] In embodiments where block 216 is executed and image captions are provided to the LLM, the generated text may be further based on the image captions. For example, if one or more images include images with captions such as "Rodeo Time in Terrebon", "Jake's First Rodeo", "Cozy Desert Cabin", "Ran and Family Horse Riding", "Sunset at Smith Rock State Park", "Hiking in the Mountains", and "Ran and BBQ", and the task description is to generate a title for an album containing one or more images, the album title may be "Outdoor Fun in Oregon - BBQ and Horse Riding" or "Ran and Family Adventure in Terrebon".

[0068] In embodiments where block 218 is executed and user input text is provided to the LLM, the generated text may be further based on the user input text. For example, the generated text may utilize the user input text, such as "Surfing with Ryan on the Voyager" or "Ryan and Gazala are so lovely."

[0069] In block 222, the text (generated by LLM in block 220) is made visible in the user interface. For example, the text may be displayed as a summary of one or more images, as the title of an album containing one or more images, as the title of a specific image, or superimposed on a specific image. Block 222 may follow block 214.

[0070] The various blocks of Method 200 may be combined, divided into multiple blocks, or executed in parallel. For example, blocks 212 and 214 may be executed in parallel. In another example, blocks 216 and 218 may be executed in parallel. In some embodiments, blocks 216 and / or 218 may not be executed. Method 200, or any part thereof, may be repeated any number of times with additional inputs. For example, Method 200 may be repeated for different prompts.

[0071] In some embodiments, one or more images for which image metadata is provided to the LLM may be in an image album. In these embodiments, a user request for an album title for the image album may be received. In response to the user request, a task description may be selected, which is included in the prompt, specifying that the text (generated by the LLM) is intended to be used as the album title. One or more other parameters, such as minimum or maximum length (in words or characters), tone or style (e.g., quirky, formal, funny), language, etc., may be specified for the text to be generated. In these embodiments, the generated text may include several options (suggested album titles generated by the LLM) that are displayed to the user in the user interface. The user may select a particular title from the options as the album, or modify the title and assign it to the album. The user may also choose to regenerate the title. In response to a user request, method 200 may be repeated to generate new text.

[0072] In some embodiments, the method may further include receiving search queries from the user, for example, via text input, voice input, image input, etc. For example, the user may specify the search "Oregon" (images associated with Oregon), "baseball" (images associated with baseball), "my birthday" (images associated with the user's birthday), "mountain," etc. In the case of image input, the user may provide a query image as the search.

[0073] In these embodiments, one or more images (for which metadata is provided to the LLM) may be identified as images from an image library associated with a user account that matches the search query. In these examples, the text generated by the LLM may describe the search results. For example, in response to the search query "Oregon," one or more identified images may include three different sets of images from three different trips the user took to Oregon. The generated text may include text describing each trip, e.g., "Portland Food Truck Trail," "Waterfalls," and "Outdoors in Terrebon." The user interface may include, for example, the text and the relevant images from one or more images within a section associated with each trip. Thus, the generated text may be snippets that provide a summary of the matching results, organized into sections.

[0074] In another example, in response to the text query "mountain," the generated text may include an essay about various mountains characterized in the user's image library, along with images from the user's image library related to mountains. In yet another example, the generated text in response to the text query "pumpkin carving" may include multiple sentences, each with a related image from one or more images, and the user interface may display the images sequentially (based on image timestamps) along with the individual sentences, thereby providing a sequential display that shows the step-by-step video of the pumpkin carving activity. In various embodiments, the task description in the prompt may include the context in which the generated text is used (e.g., an essay, section snippet, step-by-step video), and the generated text may be based on the context.

[0075] In some embodiments, after generating text and making it available for display in the user interface, Method 200 may further include receiving user input text. For example, the user may provide additional information that may not be included in the image metadata and / or user account data, or that may not be reflected in the generated text if present. In these embodiments, the LLM may generate corrected text based at least in part on the user input text. For example, if the initially generated text is "Dog and Jake", the user may provide the additional information "The dog in the picture is Kali", and the corrected text may be "Kali and Jake". In other examples, the user input text may include a command, e.g., "Text with emojis is needed". The corrected text may include emojis. In these embodiments, the corrected text may be displayed in the user interface. The user may provide input text and request corrections any number of times.

[0076] In some embodiments, Method 200 may further include, with the user's permission, updating user account data based on user input text. For example, the update may include "The user owns a dog named Kali" or "The user lives in California." Updating the data in this way may allow the text generated by the LLM in future executions of Method 200 to automatically incorporate the updated information and avoid errors.

[0077] In some embodiments, before generating text, the method may further include accessing a knowledge repository based on metadata associated with one or more images to retrieve additional information. The knowledge repository may include, for example, search indexes (of an internet search engine), digital maps (including, for example, locations and related data such as location, type of location), entity databases (including, for example, entity identifiers and related data such as entity characteristics such as dog breed names and corresponding traits), product catalogs (containing, for example, data on various products), and so on. For example, if the metadata includes a location associated with an image (e.g., latitude / longitude), additional information indicating entities associated with that location (e.g., restaurant X, library Y, bookstore Z) may be fetched from the digital map. In these embodiments, the additional information may be provided to the LLM (e.g., the LLM may implement an API or code for accessing the knowledge repository, or the additional information may be provided by prompt). In these embodiments, generating text is further based on the additional information. For example, if the location coordinates are associated with an Italian restaurant named "Roberto's Pizzeria," the generated text may include the restaurant name. For example, the text generated for all food images associated with that location might be "You're in Roberto's Food Paradise."

[0078] Figure 3 shows an exemplary system architecture 300 in several embodiments. The system architecture 300 includes an image library 302, a large language model 310, an image caption model 308, user account data 312, and a user interface module 330.

[0079] The image library 302 may store, for example, multiple images 304 associated with the user account of any of the users of client devices 120-126 that use the image management application 156. In various embodiments, the image library 302 may be stored in the server system 102 (e.g., database 106), on the client devices 120-126, or in a combination of the server system 102 and client device 120, according to user settings and permissions.

[0080] In some embodiments, the image library 302 may store image metadata associated with each image. The image metadata 306 may include metadata acquired with the image, such as location, timestamp, camera settings, camera type, and image format (e.g., raw, JPEG, etc.). In some embodiments, the image metadata may include labels generated based on programmatically analyzing the image using object recognition, text recognition, entity recognition, image quality assessment, and / or other techniques. In some embodiments, the image metadata may include user-assigned labels, such as "Jake" or "Kali".

[0081] During operation, metadata 320 (of one or more images) from an image library may be provided to the LLM 310. Furthermore, in some embodiments, the image 316 (optionally together with the metadata 320) may be provided to an image caption model 308 (e.g., a trained machine learning model that generates captions for input images) to generate an image caption 318. The image caption 318 may be provided to the LLM 310.

[0082] In some embodiments, user-specific data 316 from user account 312 may be provided to LLM 310. For example, user-specific data 316 may include a person's name, a place name, an activity type, an object name, a date, or other information from user account 312. User-specific data 316 may be stored locally on the client device (120-126) or in a server system (e.g., database 106).

[0083] A prompt 314 is provided to the LLM 310. The prompt 314 may include a task description. In some embodiments, the user interface module 330 may provide a user interface that includes options for the user to select one or more actions, and the prompt 314 may be selected based on the user selection. For example, the user may, from the options, browse to several images (e.g., displayed in the user interface) and select "Suggest album titles". In another example, the user may select "Search" and enter a search term, and use this search term to select a prompt 314. In some embodiments, a prompt library (not shown) may be provided, and prompts may be selected from the prompt library. In some embodiments, image metadata 320 may be provided together with or as part of the prompt 314.

[0084] In some embodiments, the user interface module 330 may allow the user to provide user input text 324. In these embodiments, the user input text 324 may be provided to the LLM 310.

[0085] LLM310 may generate text 322 based on the provided input (e.g., prompt 314, image metadata 320, user-specific data 316, image caption 318, user input text 324, etc.). The generated text 322 may be provided to a user interface module 330 which can display it in the user interface. For example, if the user selects "Suggest album title," the generated text may be displayed in the user interface as a suggested album title. In another example, if the user selects "Search," the generated text may be displayed in the search interface as a summary or snippet describing the search results (which may include images from the image library 302 that match the search query).

[0086] In some embodiments, the user may provide user-input text 324 after viewing the generated text 322. In these embodiments, the LLM 310 may generate additional text 322 based on the user-input text (while other inputs remain the same). The generation and correction of text by the LLM 310 based on the user-input text may be performed any number of times.

[0087] In various embodiments, the LLM310 (or a part thereof) and / or the image caption model (308) may be implemented on client devices (120-126), on server device 104, or in a combination of client and server devices.

[0088] Figure 4 shows exemplary images from an image library. As can be seen from the figure, images 402-404 depict specific individuals and other entities, respectively. For example, image 402 represents a child and a dog, image 404 represents a child and a man, image 406 represents a group of people riding horses, and image 408 represents a group of people eating outdoors. In the example in Figure 4, image 402 may be associated with the following image metadata: {Face: Jake, Kali; Location: 44.38867,-121.2155; Date: July 4, 2021; Labels: dog, animal, pet, child, nature, cloudy, summer, tree, sky, eye, plant, sand; Camera: f / 2.01 / 3602.22mm ISO35}. The image metadata identifies the child as "Jake" and the dog as "Kari". The image metadata further indicates the location (latitude and longitude) of the image, the date the image was taken, etc. Image metadata also includes details about the camera and camera settings used to take the image. Image metadata also includes labels associated with the image, such as "child," "dog," or "sand."

[0089] Figure 5A shows an exemplary user interface 500. User interface 500 includes image 402. User interface 500 further includes text 502 generated by LLM 310 using, for example, method 200. In the example shown in Figure 5A, the generated text is "Kari helped Jake build a castle" (here, emojis associated with a dog and a castle are included in the text). As can be seen from the figure, when image metadata associated with image 402 ("dog", "child"), user-specific data (e.g., "My dog ​​is Kali", "My son's name is Jake"), and prompts (e.g., "Generate a short phrase with emojis") are provided as input to LLM, the generated text describes the image content, including attributes that are not present in the image but are included in the generated text (e.g., a person's name). Furthermore, the depicted activity ("building a castle") is included in the generated text even if the input data does not include the name of the activity.

[0090] Figure 5B shows another exemplary user interface 510. User interface 510 includes multiple images 512 arranged in a picture pile. User interface 510 further includes three text boxes 514, 516, and 518, each containing text generated by LLM 310 (based on images 512, associated metadata, user-specific data, prompts, etc.) using, for example, method 200. In the example shown in Figure 5B, three options (514-518) are presented, and the user can select one of the options. The text in each box is a suggested album title, and the picture pile represents the contents of the album (may contain any number of images other than the four images shown in Figure 5B). Alternatively, the user may enter their own album title in text box 520. The selected title is associated with the album and stored.

[0091] Figure 6 is a block diagram of an exemplary device 600 that may be used to implement one or more features described herein. In one example, device 600 may be used to implement a client device, for example, one of the client devices 115 shown in Figure 1. Alternatively, device 600 may implement a server device, for example, server 101. In some embodiments, device 600 may be used to implement a client device, a server device, or both a client device and a server device. Device 600 may be any suitable computer system, server, or other electronic or hardware device described above.

[0092] One or more methods described herein can be performed as standalone programs that can run on any type of computing device, programs that run on a web browser, or mobile applications ("Apps") that run on mobile computing devices (e.g., mobile phones, smartphones, tablet computers, wearable devices (watches, armbands, jewelry, headwear, virtual reality goggles or glasses, augmented reality goggles or glasses, head-mounted displays, etc.), laptop computers, etc.). In one example, a client / server architecture can be used, for example, where the mobile computing device (as a client device) sends user input data to a server device and receives final output data from the server for output (e.g., display). In another example, all calculations can be performed within a mobile app (and / or other app) on the mobile computing device. In yet another example, calculations can be divided between the mobile computing device and one or more server devices.

[0093] In some embodiments, device 600 includes a processor 602, memory 604, and an input / output (I / O) interface 606. The processor 602 may be one or more processors and / or processing circuits that execute program code to control the basic operation of device 600. "Processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., single-core, dual-core, or multi-core configuration), multiple processing units (e.g., a multiprocessor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a composite programmable logic device (CPLD), dedicated circuitry for implementing a function, a dedicated processor for performing processing based on a neural network model, a neural circuit, a processor optimized for matrix calculations (e.g., matrix multiplication), or a system having other such components. In some embodiments, the processor 602 may include one or more coprocessors that perform neural network processing. In some embodiments, processor 602 may be a processor that processes data to produce a probabilistic output, for example, the output produced by processor 602 may be inaccurate or accurate within a range from the expected output. Processing does not need to be limited to a specific geographical location or have temporal constraints. For example, the processor may perform its functions in "real-time," "offline," or "batch mode." Parts of the processing may be performed by different (or the same) processing systems at different times and locations. The computer may be any processor that communicates with memory.

[0094] Memory 604 is typically provided within device 600 for access by processor 602 and may be any suitable processor-readable storage medium suitable for storing instructions for execution by the processor, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), or flash memory, located separately from and / or integrally with processor 602. Memory 604 can store software that runs on server device 600 by processor 602, including an operating system 608, a machine learning application 630, other applications 612, and application data 614. Other applications 612 may include applications such as a data display engine, a web host engine, an image display engine, a notification engine, and a social network engine. In some embodiments, the machine learning application 630 and other applications 612 may each include instructions that enable processor 602 to execute some or all of the functions described herein, for example, the method shown in Figure 2.

[0095] Other applications 612 may include, for example, image editing applications, media display applications, communication applications, web hosting engines or applications, mapping applications, media sharing applications, and the like. One or more methods disclosed herein can operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, and as a mobile application ("App") that runs on a mobile computing device.

[0096] In various embodiments, the machine learning application 634 may utilize a Bayesian classifier, a support vector machine, a neural network, or other learning techniques. In some embodiments, the machine learning application 630 may include a trained model 634, an inference engine 636, and data 632. In some embodiments, data 632 may include training data, for example, data used to generate the trained model 634. For example, the training data may include any type of data, such as text, images, audio, or video.

[0097] In some embodiments, the trained model 634 may be a large-scale language model. A large-scale language model may have a large number of parameters (e.g., thousands, millions, or billions of parameters). The large-scale language model may be trained to respond to prompts in natural language text. The large-scale language model may be trained on a large corpus of text. In some embodiments, the corpus of text used to train the large-scale language model may include image descriptions, stories with images, image titles, emojis associated with images, and social media posts containing text and images.

[0098] Training data may be obtained from any source, such as a data repository specifically marked for training, or data for which permission has been granted to be used as training data for machine learning. In embodiments where one or more users are permitted to use their respective user data to train a machine learning model, e.g., a trained model 634, the training data may include such user data. In embodiments where users are permitted to use their respective user data, the data 632 may include permitted data such as images (e.g., photographs or other user-generated images).

[0099] In some embodiments, the training data may include synthetic data generated for training purposes, such as data not based on user input or user activity in the context being trained, for example, data generated from simulated photographs or other computer-generated images. In some embodiments, the machine learning application 630 excludes the data 632. For example, in these embodiments, the trained model 634 may be generated, for example, on a different device and provided as part of the machine learning application 630. In various embodiments, the trained model 634 may be provided as a data file containing the model structure or format and associated weights. The inference engine 636 can read the data file of the trained model 634 and implement a neural network using node connectivity, layers, and weights based on the model structure or format specified for the trained model 634.

[0100] In some embodiments, the trained model 634 may include one or more model forms or structures. For example, the model form or structure may include any type of neural network, such as a linear network, a deep neural network implementing multiple layers (e.g., a "hidden layer" between the input and output layers, where each layer is a linear network), a convolutional neural network (e.g., a network that divides or partitions input data into multiple parts or tiles, processes each tile separately using one or more neural network layers, and aggregates the results from the processing of each tile), or a sequence-to-sequence neural network (e.g., a network that takes sequential data as input, such as words in a sentence or frames in a video, and produces a resulting sequence as output). The model form or structure may specify the connectivity between various nodes and the organization of the nodes into layers.

[0101] For example, a node in the first layer (e.g., the input layer) may receive data as input data 632 or application data 614. For example, if the trained model 634 is a large language model, the input data may include prompts (e.g., text prompts). Subsequent intermediate layers may receive the output of a node in the previous layer as input, according to the connectivity specified in the model form or structure. These layers are sometimes called hidden or latent layers.

[0102] The final layer (e.g., the output layer) generates the output of the machine learning application. For example, the output may be generated text. In some embodiments, the model format or structure also specifies the number and / or type of nodes in each layer.

[0103] In different embodiments, the trained model 634 may include multiple nodes arranged in layers according to the model structure or form. In some embodiments, a node may be a memoryless computation node configured to process one unit input and produce one unit output. The computation performed by the node may include, for example, multiplying each of the multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce a node output. In some embodiments, the computation performed by the node may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computation may include operations such as matrix multiplication. In some embodiments, computations by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multicore processor, individual processing units of a GPU, or dedicated neural circuits. In some embodiments, a node may include memory, for example, to store one or more previous inputs and use them when processing subsequent inputs. For example, a node with memory may include a Long Short-Term Memory (LSTM) node. LSTM nodes can use memory to maintain "states" that allow the nodes to function like finite state machines (FSMs). Models with such nodes can be useful when processing sequential data, such as words in a sentence or paragraph, or frames in video, speech, or other audio.

[0104] In some embodiments, the trained model 634 may include embeddings or weights for individual nodes. For example, the model may be initialized as a group of nodes organized into layers, as specified by the model format or structure. In initialization, each weight may be applied to the connections between nodes connected according to the model format, for example, the connections between each pair of nodes in a series of layers of a neural network. For example, each weight may be assigned randomly or initialized to a default value. The model can then be trained using data 632, for example, to produce results.

[0105] For example, training may involve applying supervised learning techniques. In supervised learning, training data may include multiple inputs (e.g., a set of grayscale images) and expected outputs corresponding to each input (e.g., a ground truth image or other set of color images corresponding to the grayscale images). Based on a comparison of the model's output with the expected output, the weight values ​​are automatically adjusted, for example, to increase the probability that the model will produce the expected output when similar inputs are provided.

[0106] In some embodiments, training may involve applying unsupervised learning techniques. In unsupervised learning, only input data may be provided, and a model may be trained to distinguish the data, for example, to cluster the input data into multiple groups, each group containing input data that is similar in some way.

[0107] In some embodiments, unsupervised learning can be used to generate knowledge representations that can be used by, for example, a machine learning application 630. For example, unsupervised learning can be used to generate embeddings that can be utilized by the machine learning application 630. In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In embodiments where data 632 is omitted, the machine learning application 630 may include a trained model 634 based on prior training, for example, by the developer of the machine learning application 630, or by a third party. In some embodiments, the trained model 634 may include a fixed set of weights (for example, downloaded from a server that provides the weights).

[0108] The machine learning application 630 also includes an inference engine 636. The inference engine 636 is configured to provide inference by applying a trained model 634 to data such as application data 614. In some embodiments, the inference engine 636 may include software code executed by a processor 602. In some embodiments, the inference engine 636 may specify a circuit configuration (e.g., for a programmable processor, for a field-programmable gate array (FPGA), etc.) that allows the processor 602 to apply the trained model. In some embodiments, the inference engine 636 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, the inference engine 636 may provide an application programming interface (API) which is used by the operating system 608 and / or other applications 612 to invoke the inference engine 636, for example, by applying the trained model 634 to application data 614 to generate inference. For example, the inference of an LLM model may be the generated text.

[0109] The machine learning application 630 may offer several technical advantages. For example, if the trained model 634 is generated based on unsupervised learning, the trained model 634 can be applied by the inference engine 636 to produce knowledge representations (e.g., numerical representations) from input data, e.g., application data 614. For example, a model trained for image analysis may generate representations of images with a smaller data size (e.g., 1KB) than the input images (e.g., 10MB). In some embodiments, such representations may help reduce the processing cost (e.g., computational cost, memory usage, etc.) for generating outputs (e.g., text generated in response to prompts including labels, classifications, image metadata, and task descriptions).

[0110] In some embodiments, such representations may be provided from the output of the inference engine 636 as input to different machine learning applications that generate outputs. In some embodiments, the knowledge representations generated by the machine learning application 630 may be provided to different devices that perform further processing, for example, over a network. In such embodiments, providing knowledge representations rather than images can offer technical benefits, such as reducing costs and enabling faster data transmission. In other examples, a model trained to cluster documents may generate document clusters from input documents. Document clusters may be suitable for further processing (e.g., determining whether a document is relevant to a topic, determining a classification category of a document, etc.) without needing to access the original documents, thus saving computational costs.

[0111] In some embodiments, the machine learning application 630 may be implemented offline. In these embodiments, a trained model 634 may be generated in a first stage and provided as part of the machine learning application 630. In some embodiments, the machine learning application 630 may be implemented online. For example, in such embodiments, an application calling the machine learning application 630 (e.g., one or more of the operating system 608, other applications 612) may utilize the inferences generated by the machine learning application 630, for example, by providing the inferences to a user and generating a system log (e.g., actions taken by the user based on the inferences, if permitted by the user, or the results of further processing, if used as input for further processing). The system log may be generated periodically, for example, hourly, monthly, quarterly, etc., and may be used, with the user's permission, to update the trained model 634, for example, by updating the embedding of the trained model 634.

[0112] In some embodiments, the machine learning application 630 may be implemented in a way that can adapt to a specific configuration of the device 600 on which the machine learning application 630 is executed. For example, the machine learning application 630 may determine the computation graph that utilizes available computing resources, e.g., the processor 602. For example, if the machine learning application 630 is implemented as a distributed application across multiple devices, the machine learning application 630 may determine which computations should be performed on each device in a computationally optimized manner. In another example, the machine learning application 630 may determine that the processor 602 includes a GPU with a specific number (e.g., 1000) GPU cores and implement an inference engine accordingly (e.g., as 1000 individual processes or threads).

[0113] In some embodiments, the machine learning application 630 may implement an ensemble of trained models. For example, the trained models 634 may include multiple trained models, each applicable to the same input data. In these embodiments, the machine learning application 630 may select a particular trained model based on, for example, available computational resources, success rate of prior inference, etc. In some embodiments, the machine learning application 630 may run an inference engine 636 so that multiple trained models are applied. In these embodiments, the machine learning application 630 may combine the outputs from the application of individual models, for example, by using a voting technique that scores the individual outputs from each application of a trained model, or by selecting one or more specific outputs. Furthermore, in these embodiments, the machine learning application may apply a time threshold (e.g., 0.5 ms) to apply individual trained models and may utilize only those individual outputs available within the time threshold. Outputs not received within the time threshold may not be utilized, for example, may be discarded. For example, such an approach may be suitable when there is a specified time limit while the machine learning application is being invoked, for example, by the operating system 608 or one or more applications 612.

[0114] In different embodiments, the machine learning application 630 can generate different types of output. For example, the machine learning application 630 can provide representations or clusters (e.g., numerical representations of input data), labels (e.g., for input data including images, documents, etc.), phrases or sentences (e.g., descriptions of images or videos, suitable for use as responses to input sentences, etc.), images (e.g., in response to an input image, e.g., a grayscale image, a color or other stylized image generated by the machine learning application), audio or video (e.g., in response to an input video, the machine learning application 630 may generate an output video with a specific effect applied, e.g., an output video rendered in the style of a comic book or a specific artist, if the trained model 634 is trained using training data from a comic book or a specific artist). In some embodiments, the machine learning application 630 can generate output based on a format specified by a calling application, e.g., the operating system 608 or one or more applications 612. In some embodiments, the calling application may be another machine learning application. For example, such a configuration could be used in a generative adversarial network where the calling machine learning application is trained using the output from the machine learning application 630, and vice versa.

[0115] Alternatively, any of the software in memory 604 may be stored in any other suitable storage location or computer-readable medium. Furthermore, memory 604 (and / or other connected storage devices) may store one or more messages, one or more classifications, electronic encyclopedias, dictionaries, thesauruses, knowledge bases, message data, grammars, user preferences, and / or other instructions and data used in the features described herein. Memory 604 and any other type of storage (magnetic disks, optical disks, magnetic tapes, or other tangible media) may be considered “storage” or “storage devices.”

[0116] The I / O interface 606 can provide functionality that enables the server device 600 to interface with other systems and devices. Interfaced devices may be included as part of device 600 or may be separate and communicate with device 600. For example, network communication devices, storage devices (e.g., memory and / or database 106), and input / output devices can communicate via the I / O interface 606. In some embodiments, the I / O interface can be connected to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone, camera, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, motor, etc.).

[0117] Some examples of interfaced devices that can be connected to the I / O interface 606 may include one or more display devices 620, which can be used to display content, for example, images, videos, and / or the user interface of an output application described herein. The display device 620 may be any suitable display device, and may be connected to device 600 via a local connection (e.g., a display bus) and / or a network connection. The display device 620 may include any suitable display device, such as an LCD, LED, or plasma display screen, a CRT, a television, a monitor, a touchscreen, a 3D display screen, or other visual display device. For example, the display device 620 may be a flat display screen provided in a mobile device, a multi-display screen provided in a goggles or headset device, or a monitor screen in a computer device.

[0118] The I / O interface 606 can interface with other input and output devices. Some examples include one or more cameras capable of capturing images. Some embodiments may provide a microphone for capturing sound (e.g., as part of captured images, voice commands, etc.), an audio speaker device for outputting sound, or other input and output devices.

[0119] For simplicity of explanation, Figure 6 shows one block for each of the processor 602, memory 604, I / O interface 606, and software blocks 608, 612, and 630. These blocks may represent one or more processors or processing circuits, operating systems, memory, I / O interfaces, applications, and / or software modules. In other embodiments, device 600 may not have all of the illustrated components and / or may have other components, including other types of elements, instead of or in addition to the components shown herein. While some components are described as performing the blocks and operations described in some embodiments herein, any suitable component or combination of components of environment 100, device 600, a similar system, or any suitable processor(s) associated with such a system may perform the described blocks and operations.

[0120] The methods described herein can be executed by computer program instructions or code that can be executed on a computer. For example, the code can be executed by one or more digital processors (e.g., microprocessors or other processing circuits) and can be stored in computer program products including non-temporary computer-readable media (e.g., storage media) such as magnetic, optical, electromagnetic, or semiconductor storage media, including semiconductor or solid-state memory, magnetic tape, removable computer diskettes, random-access memory (RAM), read-only memory (ROM), flash memory, rigid magnetic disks, optical disks, solid-state memory drives, etc. Program instructions can also be contained in electronic signals and provided as electronic signals, for example, in the form of software as a service (SaaS) delivered from a server (e.g., a distributed system and / or a cloud computing system). Alternatively, one or more methods can be executed in hardware (e.g., logic gates), or in combination of hardware and software. Exemplary hardware may include programmable processors (e.g., field-programmable gate arrays (FPGAs), composite programmable logic devices), general-purpose processors, graphics processors, application-specific integrated circuits (ASICs), etc. It can perform one or more actions as part of or a component of an application running on the system, or as an application or software that operates in conjunction with other applications and the operating system.

[0121] The description is given with respect to specific embodiments, but these specific embodiments are merely illustrative and not limiting. Concepts shown in the examples may apply to other examples and embodiments.

[0122] In addition to the above description, the user may be provided with controls that allow the user to make choices regarding both whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., information about the user's image, user's address book, social networks, social behavior or activities, occupation, user preferences, or user location), and whether the user receives content or communications from the server. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, a user's identity may be processed so that personally identifiable information cannot be determined, or, if location information is obtained, the user's geographical location may be generalized (to the city, zip code, or state level, etc.) so that the user's specific location cannot be determined. Thus, the user may have control over what information is collected about them, how that information is used, and what information is provided to them.

[0123] It should be noted that the functional blocks, operations, features, methods, devices, and systems described herein may be integrated or divided into different combinations of systems, devices, and functional blocks, as is known to those skilled in the art. Routines of particular embodiments may be implemented using any preferred programming language and programming technique. Different programming techniques, e.g., procedural or object-oriented, may be employed. Routines may be executed on a single processing device or multiple processors. Steps, operations, or calculations may be presented in a specific order, but the order may be modified in different particular embodiments. In some embodiments, multiple steps or operations shown herein as sequential may be executed simultaneously.

Claims

1. A method performed by a computer, This includes providing a prompt to a Large-Scale Language Model (LLM), wherein the prompt is: Task description and, The method further includes metadata associated with one or more images, wherein the one or more images are images from an image library associated with a user account, and the method further includes The method further includes providing the user account data to the LLM, wherein the data includes user-specific data including one or more of a person's name, a place name, an activity type, an object name, and a date, and the method further includes The method includes generating text in response to the prompt using the LLM based on the task description, the metadata of one or more images, and the data of the user account, wherein the text is personalized based on the data of the user account and responds to the task description, and the method further includes To ensure that the aforementioned text is displayed in the user interface, A method performed by a computer, including [this].

2. Using an image caption model, generate captions for each of the one or more images, To provide the caption to the LLM, It further includes, The generation of the aforementioned text by the LLM is further based on the respective captions. The method performed by a computer as described in claim 1.

3. The aforementioned one or more images are located in an image album, and the method further, Receiving a user request for the album title of the aforementioned image album, This includes selecting the task description that specifies that the text is to be used as the album title in response to the user request, The aforementioned text includes one or more proposed album titles, The method performed by a computer as described in claim 1.

4. Receiving search queries from users, Identifying one or more images from the image library associated with the user account that match the search query, It further includes, The user interface includes at least one of the one or more images, The method performed by a computer as described in claim 1.

5. Receiving user input text and The user input text is provided to the LLM, The further includes, and the generation of the text by the LLM is at least partially based on the user input text. The method performed by a computer as described in claim 1.

6. After the text and at least one of the one or more images are displayed on the user interface, Receiving user input text and The user input text is provided to the LLM, The LLM generates text corrected at least partially based on the user input text, The corrected text is displayed in the user interface, The method performed by a computer according to claim 1, further comprising:

7. The method performed by a computer according to claim 6, further comprising automatically updating the data of the user account based on the user input text.

8. A method performed by a computer according to claim 1, further comprising accessing a knowledge repository based on the metadata associated with one or more images in order to obtain additional information, wherein generating the text is further based on the additional information.

9. The computer-based method according to claim 1, wherein the user interface includes at least one of the one or more images.

10. The computer-based method according to claim 9, wherein at least a portion of the text is superimposed on at least one of the one or more images.

11. A computing device, Processor and The system includes a memory coupled to the processor, the memory, when executed by the processor, stores instructions that cause the processor to perform an action, and the action is, This includes providing a prompt to a Large-Scale Language Model (LLM), wherein the prompt is: Task description and, The operation includes metadata associated with one or more images, wherein the one or more images are images from an image library associated with a user account, and the operation further includes The operation further includes providing the user account data to the LLM, the data including user-specific data including one or more of a person's name, a place name, an activity type, an object name, and a date, and the operation further includes: The operation includes generating text in response to the prompt using the LLM based on the task description, the metadata of one or more images, and the data of the user account, wherein the text is personalized based on the data of the user account and responds to the task description, and the operation further includes: To ensure that the aforementioned text is displayed in the user interface, Computing devices, including [this].

12. The aforementioned operation further, Using an image caption model, generate captions for each of the one or more images, To provide the caption to the LLM, Includes, The generation of the aforementioned text by the LLM is further based on the respective captions above. The computing device according to claim 11.

13. The aforementioned one or more images are in an image album, and the aforementioned operation further, Receiving a user request for the album title of the aforementioned image album, This includes selecting the task description that specifies that the text is to be used as the album title in response to the user request, The aforementioned text includes one or more proposed album titles, The computing device according to claim 11.

14. The aforementioned operation further, Receiving search queries from users, This includes identifying one or more images that match the search query from the image library associated with the user account, The user interface includes at least one of the one or more images, The computing device according to claim 11.

15. The aforementioned operation further, Receiving user input text and The user input text is provided to the LLM, The generation of the text by the LLM is at least partially based on the user input text. The computing device according to claim 11.

16. After the text and at least one of the one or more images are displayed on the user interface, the operation further: Receiving user input text and The user input text is provided to the LLM, The LLM generates corrected text based at least partially on the user input text, The corrected text is displayed in the user interface, The computing device according to claim 11, including the following:

17. The computing device according to claim 11, wherein the operation further includes accessing a knowledge repository based on the metadata associated with the one or more images in order to obtain additional information, and the generation of the text is further based on the additional information.

18. A non-temporary computer-readable medium in which instructions causing the processor to perform an action are stored when executed by the processor, wherein the action is This includes providing a prompt to a Large-Scale Language Model (LLM), wherein the prompt is: Task description and, The operation includes metadata associated with one or more images, wherein the one or more images are images from an image library associated with a user account, and the operation further includes The operation further includes providing the user account data to the LLM, the data including user-specific data including one or more of a person's name, a place name, an activity type, an object name, and a date, and the operation further includes: The operation includes generating text in response to the prompt using the LLM based on the task description, the metadata of one or more images, and the data of the user account, wherein the text is personalized based on the data of the user account and responds to the task description, and the operation further includes: To ensure that the aforementioned text is displayed in the user interface, Non-temporary computer-readable media, including [specific examples of such media].

19. The instruction causes the processor to perform a further operation, and the further operation is Using an image caption model, generate captions for each of the one or more images, To provide the caption to the LLM, Includes, The generation of the aforementioned text by the LLM is further based on the respective captions. The non-temporary computer-readable medium according to claim 18.

20. The instruction causes the processor to perform a further operation, and the further operation is Receiving search queries from users, Identifying one or more images from the image library associated with the user account that match the search query, Includes, The user interface includes at least one of the one or more images, The non-temporary computer-readable medium according to claim 18.