System and method for analyzing text extracted from image and performing appropriate transformations on extracted text
By analyzing image text content using image processing and machine learning models and generating appropriate responses, this technology solves the problem of low efficiency in existing image text analysis technologies, achieving real-time response and efficient data processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2024-09-06
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies struggle to effectively analyze and respond to user requests for text content within images, especially in extracting text from real-time videos or images and providing appropriate services or actions.
By using image processing techniques and machine learning models, the text content in an image is analyzed to determine its characteristics, and appropriate response types, such as summaries, explanations, or query answers, are generated and displayed to the user based on these characteristics.
It enables real-time response to real-time video or images, improves computing efficiency, reduces the need for users to select response types, saves processor time and battery power, and efficiently processes data in network-limited situations.
Smart Images

Figure CN121941985A_ABST
Abstract
Description
[0001] Priority Statement
[0002] This application is based on and claims priority to U.S. non-provisional application 18 / 463,951, filed on September 8, 2023, which is incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to performing appropriate transformations on text. More specifically, this disclosure relates to recognizing text in an image, extracting the text, and performing one of a variety of transformations on the text based on its characteristics. Background Technology
[0004] As computing devices advance, they can be used to provide an increasing number of services to users. In some examples, computing devices can be used to capture and display images. These images may include components of interest to the user. The computing system can be enabled to perform various services or transformations associated with the components of interest. It would be useful if the computing system (or the applications on it) could perform appropriate services based on one or more characteristics of the components of interest. Summary of the Invention
[0005] Aspects and advantages of embodiments of this disclosure will be set forth in part in the description which follows, or may be learned from the description or by practice of the embodiments.
[0006] One example aspect of this disclosure relates to a computing system. The system may include: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. These operations may include acquiring an image depicting a first set of text content. These operations further include determining one or more characteristics of the first set of text content. These operations further include determining a response type from a plurality of response types based on the one or more characteristics. These operations further include generating a model input, wherein the model input includes data describing the first set of text content and a prompt associated with the response type. These operations further include providing the model input as input to a machine learning language model. These operations further include receiving a second set of text as output of a machine learning language model, the output of which is a result of the machine learning language model processing the model input. These operations further include providing the second set of text to a user for display, wherein the second set of text content is associated with the response type.
[0007] Another example aspect of this disclosure relates to a computer-implemented method. The method includes acquiring an image by a computing system having one or more processors, wherein the image depicts a first set of text content. The method further includes determining one or more characteristics of the first set of text content by the computing system. The method further includes determining a response type from a plurality of response types by the computing system based on the one or more characteristics. The method further includes generating a model input by the computing system, wherein the model input includes data describing the first set of text content and a cue associated with the response type. The method further includes providing the model input by the computing system as input to a machine learning language model. The method further includes receiving a second set of text as output by the machine learning language model, the output of which is a result of the machine learning language model processing the model input. The method further includes providing the second set of text to a user for display by the computing system, wherein the second set of text content is associated with the response type.
[0008] Another example aspect of this disclosure relates to one or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to operate. These operations may include acquiring an image depicting a first set of text content. These operations further include determining one or more characteristics of the first set of text content. These operations further include determining a response type from a plurality of response types based on the one or more characteristics. These operations further include generating model input, wherein the model input includes data describing the first set of text content and a prompt associated with the response type. These operations further include providing the model input as input to a machine learning language model. These operations further include receiving a second set of text as output of a machine learning language model, the output of which is a result of the machine learning language model processing the model input. These operations further include providing the second set of text to a user for display, wherein the second set of text content is associated with the response type.
[0009] These and other features, aspects, and advantages of the various embodiments of this disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the disclosure and, together with the description, serve to explain the relevant principles. Attached Figure Description
[0010] Referring to the accompanying drawings, a detailed discussion of embodiments is set forth in this specification for those skilled in the art, in which:
[0011] Figure 1A block diagram of an example computing system according to an example embodiment of the present disclosure is depicted, which uses a machine learning model to respond to a user request for text extracted from an image;
[0012] Figure 2A An example interface of an application (e.g., an image recognition and analysis application) according to some embodiments of the present disclosure is shown, which has interface elements that indicate the visual search features of the application;
[0013] Figure 2B An example interface of an application (e.g., a virtual assistant application) according to an embodiment of the present disclosure is shown, the application having a user interface for displaying a summary of text content;
[0014] Figure 3A An example user interface according to an example embodiment of this disclosure is depicted;
[0015] Figure 3B An example interface of an application (e.g., a virtual assistant application) according to an embodiment of the present disclosure is shown, which has a user interface for displaying an interpretation of text content.
[0016] Figure 4A An example user interface according to an example embodiment of this disclosure is depicted;
[0017] Figure 4B An example interface of an application (e.g., a virtual assistant application) according to an embodiment of the present disclosure is shown, the application having a user interface for displaying an interpretation of text content;
[0018] Figure 5 This is an example image analysis system according to an example embodiment of the present disclosure;
[0019] Figure 6 A block diagram of an example computing device implemented according to an example embodiment of the present disclosure is depicted;
[0020] Figure 7 A block diagram depicting an example computing device 700 performing according to an example embodiment of the present disclosure;
[0021] Figure 8 An example flowchart depicts a method for providing an appropriate response based on the characteristics of an image and its contained text, according to an example embodiment of the present disclosure.
[0022] The repeated reference numerals across multiple figures are intended to identify the same features in various implementations. Detailed Implementation
[0023] In general, this disclosure relates to systems and methods for analyzing textual content in images and providing appropriate actions or services in response to user requests. Specifically, the systems and methods disclosed herein can utilize image processing techniques (e.g., optical character recognition or similar techniques) and machine learning models to provide analysis and additional content for text included in images. For example, the systems and methods disclosed herein can be used to acquire image data, process the image data to extract textual content (e.g., characters or tokens in the image), determine one or more characteristics of the textual content (or image), and, in response to a user request, provide services to the user based on one or more characteristics of the text (or image) using a large language model. Services may include one or more of the following: summarizing, answering queries associated with the textual content or image, or interpreting the text.
[0024] In some examples, the image processing system can acquire an image. In some examples, the image may be a portion of live video captured by a camera associated with a user's computing device and represent a real-time representation of an area of the user's computing device. In other examples, the image may be a stored image file accessed through the user's computing device. The image may be displayed on the screen of the user's computing device. The image processing system may determine that the image contains text and use one or more text recognition techniques to extract the text.
[0025] Image processing systems can determine one or more characteristics associated with text, images, and / or input received from a user. Based on these characteristics, the image processing system can determine the type of response associated with the image and any extracted text content. For example, the image processing system can determine whether the appropriate response type is a summary response, a query-answering response, or an explanatory response. In some implementations, the image processing system can update the interface displaying the image to include interactive elements that allow the user to request the determined type of response (or request a different response). Therefore, the image processing system can infer the correct response type to be performed on the text. The inferred response type can be displayed by user selection or can be performed automatically.
[0026] Image processing systems can generate model inputs based on text, images, and / or inferred response types (e.g., the type of user request). These model inputs can be fed into a machine learning model (e.g., a large language model). The machine learning model can then output a response to the request. In some examples, the request might be a summary of the text content included in an image. If so, the output of the machine learning model could be a summary of the text. In this case, the output would contain less text than the text extracted from the image. In other cases, if a different response type is requested (such as an explanation or query-answer type), the output by the machine learning model might, in some situations, be larger than the text content input into the machine learning model.
[0027] The model's output can be displayed to the user in the user interface of the user's computing device. In some examples, the output may be displayed near or overlapping the image where the text was initially found.
[0028] More specifically, a user computing device can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (e.g., a smartphone or tablet computer), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device. In some examples, a user computing device may include an image capture sensor, such as a camera.
[0029] In some implementations, a user computing device may use an integrated camera to acquire image data of its environment or specific objects within that environment. In some implementations, the captured or acquired image data may include text within the image. In other examples, the image may be accessed via a communication network. The user may wish to interact with the image or receive services associated with it.
[0030] For example, if an image contains a large amount of dense text, a user may have questions about its content and meaning, or may want to summarize or interpret some or all of the dense text. The user's computing device may include an image analysis system that can extract text from the image to obtain a first set of text content. This first set of text content can be generated using OCR technology. However, other techniques can also be used to extract text content from the image.
[0031] Once the text content has been extracted, the image analysis system can determine one or more characteristics of the image, the initial set of text content, or other user input to determine the appropriate response type. For example, the image analysis system can determine the density associated with the text. Density can be measured by the number of words and the amount of space those words occupy on the screen. Therefore, if the image includes a large amount of text within a small area of the screen, the density of the initial set of text content can be determined to be relatively high. Another characteristic can be based on the content of the initial set of text content.
[0032] For example, if an image analysis system determines that the content of an image is related to a problem or learning (or teaching), the system can determine that interpreting the text may be appropriate. The user interface can be updated to add user interface elements associated with the determined appropriate response. For example, if the system determines that a summary is appropriate, the user interface can be updated to include a "Summary" button.
[0033] In some examples, one of these features could be text entered in a query field provided by the image display application. The query field can allow the user to enter a query related to the image, text content extracted from the image, or both. In some examples, the user can enter the query via voice communication. The image analysis system can determine whether the query is related to the content of the image. If so, the image analysis system can determine the appropriate response is the query response type and can update the user interface to include query interface elements.
[0034] Once the user interface has been updated to include the appropriate response element (e.g., a summary button, an explanation button, or a query response button), the user can select that element. The user can select a response element (e.g., a summary button), and in response, the image analysis system can generate a response request. The response request can be based on the element selected by the user. For example, if the user selects the summary button, the image analysis system can generate a summary request. The summary request may include an initial set of textual information, information about the image or its content, and instructions indicating that the request is a summary request. This information can be included as input to the model. This model input can be sent to a machine learning model as input.
[0035] In some examples, the machine learning model is implemented on a remote server system, and the input is transmitted to the remote server system via a communication network. In other examples, the machine learning model is stored on a computing device, and the model may simply be transmitted to the model within the device. The model input may be a prompt to a large language model. The prompt may include an initial set of text data, an indication of the response type, information about the image content, and any associated contextual information. Contextual information may include information about the user (if the user consents to such information), information describing previous requests and corresponding responses, and so on.
[0036] A machine learning model can receive model input. It can process this input and generate output. The specific output can be based on the initial set of text information and the response type. For example, if the response type is a summary, the output could be a summary of the initial set of text information. In this case, the output could contain less text than the initial set of text data. In another example, the request type could be an explanation request. If so, the output could be text explaining the content of the initial set of text. If the response type is a response type, the output could be a response to a user's query about the image content.
[0037] The output can be displayed to the user via a user interface. In some examples, the output of a machine learning model is displayed in the user interface near or overlaid on the image. For instance, a summary of text in an image can be displayed in the user interface near the summarized text.
[0038] The systems and methods disclosed herein provide several technical effects and benefits. As an example, the systems and methods can provide real-time responses to real-time video or images. Specifically, the systems and methods disclosed herein can acquire image data, process the image data, determine an appropriate response type, and use a machine learning model to generate an appropriate response to display to the user. The technical benefit of the systems and methods disclosed herein lies in their ability to utilize information generated by an image processing system to determine one or more characteristics of one or more images and text included within the images, thereby determining an appropriate response for the user. This improves computational efficiency and enhances the operation of the computing system.
[0039] For example, the systems and methods disclosed herein can automatically select the response type to be provided to the user (via elements inserted into the updated interface). This reduces the need for the user (in many cases) to select a specific response type, resulting in easier application use. Furthermore, accurately estimating the appropriate response type allows for more efficient use of processor time and battery power. Additionally, this determination can be performed locally on the user's computing device. Local processing on the user's computing device limits the amount of data transferred over the network to the server computing system for processing, which may be more efficient or effective for computing systems with limited network access.
[0040] Therefore, the proposed system addresses the technical problem of how to effectively analyze and extract valuable information from textual content in images, and subsequently provide relevant and appropriate services or actions in response to user requests. Specifically, the system uses image processing techniques and machine learning models to analyze the textual content in images. These techniques are inherently technical because they involve specific algorithms, computations, and manipulations of data. The proposed system provides technical effects by extracting textual content from images, determining one or more characteristics of the textual content, and providing services based on these characteristics using a large language model. These operations involve processing and transforming data in a way that achieves concrete and tangible results. For example, the system can provide summaries, answers to queries, or textual explanations—meaningful outputs serving practical purposes. Furthermore, the system's ability to obtain image data from both real-time video feeds and stored image files, and its ability to update the interface to include interactive elements for user requests, further demonstrates its technical nature. These features involve the specific hardware configuration and software instructions necessary to implement the system.
[0041] Exemplary embodiments of this disclosure will now be discussed in further detail with reference to the accompanying drawings.
[0042] Figure 1 A block diagram of an example computing system 100 according to an exemplary embodiment of the present disclosure is depicted, which uses a machine learning model to respond to user requests for text extracted from an image. System 100 includes a user computing device 102 and a server computing system 130 communicatively coupled via a network 180.
[0043] User computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop computer or desktop computer), a mobile computing device (e.g., a smartphone or tablet computer), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0044] User computing device 102 includes one or more processors 112 and memory 114. The one or more processors 112 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 114 can store data 116 and instructions 118, which are executed by processor 112 to cause user computing device 102 to perform operations.
[0045] In some implementations, the user computing device 102 may store or include one or more machine learning models 120 for responding to user requests associated with text content extracted from images. In some implementations, the user computing device 102 may store or include one or more machine learning models 120. For example, the machine learning model 120 may be, or may otherwise include, various machine learning models such as neural networks (e.g., deep neural networks), large language models (LLMs), or other types of machine learning models, including nonlinear and / or linear models. Neural networks may include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include multi-head self-attention models (e.g., transformer models). Reference Figure 6 and Figure 7 Example machine learning model 120 is discussed.
[0046] In some implementations, one or more machine learning models 120 may be received from server computing system 130 via network 180, stored in user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, user computing device 102 may implement multiple parallel instances of a single machine learning model 120 (e.g., multiple instances of model 120 perform parallel optimization on user interactions and task selection of a large language model).
[0047] More specifically, in some implementations, machine learning model 120 may include a large-scale language model. The large-scale language model may be, or otherwise include, a model trained on a large corpus of language training data in a manner that enables the large-scale language model to perform a variety of language tasks. For example, the large-scale language model may be trained to perform tasks such as summarizing, conversational, simplification, opposing viewpoints, explanation, and requiring the model to respond to a query. Specifically, the large-scale language model may be trained to process various outputs to generate language output. For example, the large-scale language model may process model inputs that may include a first set of text content extracted from an image, a query, a summary request, an explanation request, and image data. In some examples, image data may be provided as context for the primary request (e.g., summarizing, explaining, or responding to a query entered by the user).
[0048] More specifically, in some embodiments, the machine learning model 120 may process a first set of text content, a request, and in some cases, image content as input to determine an appropriate response that includes a second set of text content. For example, the machine learning model 120 may be trained to summarize, interpret, or respond to queries associated with the first set of text content.
[0049] Alternatively or concurrently, in some embodiments, the machine learning model 120 may be, or otherwise include, a model trained to analyze a first set of text content provided by the user. For example, the machine learning model 120 may be trained to process the first set of text content to generate a second set of text content in response to a request from the user. As another example, the machine learning model 120 may be trained to process requested data (e.g., a user may request a summary, explanation, or submission of a query) to generate an appropriate second set of text data in response to the request. The machine learning model may also receive image data as context for the request.
[0050] Alternatively, one or more machine learning models 140 may be included in or otherwise stored and implemented by server computing system 130, which communicates with user computing device 102 according to a client-server relationship. For example, machine learning model 140 may be implemented by server computing system 130 as part of a web service (e.g., a service that provides responses to user requests). Thus, one or more machine learning models 120 may be stored and implemented at user computing device 102, and / or one or more machine learning models 140 may be stored and implemented at server computing system 130.
[0051] User computing device 102 may also include one or more user input components 122 for receiving user input. For example, user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). Touch-sensitive components may be used to implement a virtual keyboard. Other example user input components include microphones, conventional keyboards, or other components through which a user can provide user input.
[0052] Server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 134 can store data 136 and instructions 138, which are executed by processor 132 to cause server computing system 130 to perform operations.
[0053] In some implementations, the server computing system 130 includes one or more server computing devices or is otherwise implemented by such one or more server computing devices. Where the server computing system 130 includes multiple server computing devices, these server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0054] As described above, the server computing system 130 may store or otherwise include one or more machine learning models 140. For example, machine learning models 140 may be, or may otherwise include, various machine learning models. Example machine learning models include neural networks or other multi-layer nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include multi-head self-attention models (e.g., Transformer models).
[0055] Network 180 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over Network 180 can be carried via any type of wired and / or wireless connection using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).
[0056] The machine learning models described in this specification can be used for a variety of tasks, applications, and / or use cases.
[0057] In some implementations, the input to the machine learning model of this disclosure can be one or more of the following: text content extracted from an image, a specific request from a user, image data, or other data provided by the user and used as context for the request. As an example, the machine learning model can receive input data comprising a first set of text content extracted from an image and a summary request. The machine learning model can process the input data and, in response to the summary request, output a summary of the first set of text content. The summary can be a second set of text content. The second set of text content can have less text content than the first set.
[0058] In another example, the machine learning model might receive input data consisting of a first set of text content extracted from an image and an interpretation request (i.e., a request to interpret the text content included in the image). In some examples, the input data might be included in prompts for the machine learning model. The machine learning model might process the input data and, in response to the interpretation request, output an interpretation of the first set of text content. The interpretation could be a second set of text content. The second set of text content could contain more text than the first set. In some examples, additional contextual information (such as the age of the user submitting the request) could be used to generate an age-appropriate interpretation for a particular first set of text content. In some examples, the output might also include other media as part of the interpretation. For example, the output of the machine learning model could include text content, images, animations, videos, audio content, and so on.
[0059] In some examples, the machine learning model may receive input data that includes an initial set of text content extracted from an image, image data from the image, and a query received from the user. The query can be text or natural language data as input to the machine learning model. For example, a user may select a query input field included in an interface element (e.g., a button in the interface to initiate a query input interface) and then type (or speak) the question into the interface using an audio-based interface. The machine learning model can process the text or natural language data of the query, the initial set of text content, and any image data provided as context to generate output. The machine learning model can process the input data and output a query response in response to a query request.
[0060] The query response can be a second set of text content. In some examples, the query response may include images, animations, videos, audio content, etc. The query response can have more text content than the first set of text content. In some examples, additional contextual information, such as the age or location of the user who submitted the request (if the user chooses to provide this information), can be used to generate an appropriate response for the model input.
[0061] Figure 1 An example computing system that can be used to implement this disclosure is shown. Other computing systems may also be used. For example, in some implementations, the machine learning model 120 may be both trained and used locally at the user computing device 102.
[0062] Figure 2A An example user interface 200A of an application (e.g., an image recognition and analysis application) according to some embodiments of the present disclosure is shown, the application having interface elements that indicate the visual search characteristics of the application. Specifically, user interface 200A depicts the interface of the application that displays images and enables the user to make requests based on these images. As depicted, user interface 200A includes a variety of interface elements with which the user can interact. For example, the user interface includes a toolbar 212 (e.g., bars linking to various features of user interface 200A), an image display area 202, and one or more interface elements (e.g., query element 206).
[0063] In some examples, toolbar 212 may allow users to select one or more different request type modes (e.g., summary, explanation, search query, etc.). Therefore, if a user wishes to make a specific request, they can select an associated label or icon in toolbar 212. In some examples, the application's interface can be updated based on the specific icon or label selected by the user in the toolbar. For example, if the user selects the "Translate" label, the user interface may include elements that allow the user to select any target language for the translation. Similarly, if the user selects a specific item such as "Search" or "Query," the user interface may include query input elements.
[0064] In some implementations, the user interface may include an image display area 202 for displaying an image. The displayed image may include at least one portion containing text content 204. The text content 204 may be analyzed to extract a first set of text content. The user may select a user interface element 206. For example, if the user selects the "Summary" tab, the displayed user interface element 206 may be a "Summary" button.
[0065] If the user selects the "Summary" button, the user's computing device can generate input (e.g., a prompt) to the machine learning model. In response, the machine learning model can generate output. The output can be a second set of text content.
[0066] Figure 2B An example user interface 200B of an application (e.g., a virtual assistant application) according to an embodiment of the present disclosure is shown. This application has a user interface for displaying summaries of text content. The user interface 200B can be displayed when a user has requested a summary of text included in an image. Specifically, the user interface 200B includes a text summary interface element 210. The text summary interface element 210 may include the output of a machine learning model. The output of the machine learning model may be a summary of text from... Figure 2A The text summary shown is extracted from the image. Generally, the text summary will have less text content than the text extracted from the image.
[0067] In some examples, the user interface may also include interface elements that serve as links 208 to image displays. For example, if a user requests a summary of the text in an image, the user interface can be updated from user interface 200A, where the image is displayed, to user interface 200B, where the text summary interface element 210 is displayed. To facilitate the user switching back to user interface 200A, where the image is displayed, a link 208 to the image display is provided in the user interface 210 where the text summary interface element 210 is displayed.
[0068] Figure 3AAn example user interface 300A according to an exemplary embodiment of the present disclosure is depicted. Specifically, the user interface 300A depicts an interface of an application that displays images and enables a user to make requests based on these images. As depicted, the user interface 300A includes a variety of interface elements with which the user can interact. For example, the user interface includes a toolbar 212 (e.g., bars linking to various features of the user interface 300A), an image display area 302, and one or more interface elements (e.g., an interpretation request element 306).
[0069] In some examples, toolbar 212 allows users to select one or more different request types (e.g., summary, explanation, search query, etc.). Therefore, if a user wishes to make a specific type of request, they can select an associated label or icon in toolbar 212. In some examples, the application's interface can be updated based on the specific icon or label selected by the user in the toolbar. In this example, the user selected "Explanation," and the "Explanation" label is displayed in bold. If the user selects another label, that label will be displayed in bold, and the interface can be updated to reflect the selected label.
[0070] In some implementations, the user interface may include an image display area for displaying an image. The displayed image may include at least one portion containing text 304 within the image. The text 304 within the image may be analyzed to extract a first set of text content. The user may select an interpretation request element 306. As mentioned above, this particular user interface element (interpretation request element 306) may only be displayed when the "interpret" label is highlighted.
[0071] If the user selects interpretation request element 306, the user's computing device can generate input to a machine learning model. In response, the machine learning model can generate output. The output can be a second set of text content and can include interpretations associated with the image, the text in the image, or both.
[0072] Figure 3B An example user interface 300B of an application (e.g., a virtual assistant application) according to an embodiment of the present disclosure is shown. This application has a user interface for displaying interpretations of text content. The user interface 300B can be displayed when a user has requested an interpretation of text included in an image. Specifically, the user interface 300B includes an interpretation element 308. The interpretation element 308 may include the output of a machine learning model. The output of the machine learning model may be an interpretation of... Figure 3AThe text displayed in the image describes an explanation of one or more concepts. This text can be extracted from the image and included in the prompts input into the machine learning model. In addition to text content, explanations (e.g., the output of the machine learning model) can also include images, audio content, animations, video content, interactive content, and more.
[0073] In some examples, user interface 300B may also include interface elements as links 310 to image displays. For example, if a user requests a summary of the text in an image, the user interface can be updated from user interface 300A, where the image is displayed, to user interface 300B, where the explanation is displayed 308. To facilitate the user switching back to the page in user interface 300A where the image is displayed, a link 310 to the image display is provided in the user interface where the explanation is displayed 308.
[0074] Figure 4A An example user interface 400A according to an exemplary embodiment of the present disclosure is depicted. Specifically, the user interface 400A depicts an interface of an application that displays images and enables a user to make requests associated with those images. As depicted, the user interface 400A includes a variety of interface elements with which the user can interact. For example, the user interface includes a toolbar 212 (e.g., bars linking to various features of the user interface 400A), an image display area 402, and one or more interface elements (e.g., a query request element 406).
[0075] In some examples, toolbar 212 allows users to select one or more different request types (e.g., summary, explanation, search query, etc.). Therefore, if a user wishes to make a specific type of request, they can select an associated label or icon in toolbar 212. In some examples, the application's interface can be updated based on the specific icon or label selected by the user in the toolbar. In this example, the user selected "Query," and the query label is displayed in bold. If the user selects a different label, that label will be displayed in bold, and the interface can be updated to reflect the selected label.
[0076] In some implementations, the user interface 400A may include an image display area for displaying an image. The displayed image may include at least one portion containing text content 404. The text content 404 may be analyzed to extract a first set of text content. The user may select a query request element 406. As mentioned above, this particular user interface element may only be displayed when the "Explanation" label is highlighted.
[0077] If the user selects query request element 406, the user's computing device can generate input for a machine learning model. In response, the machine learning model can generate output. The output can be a second set of text content.
[0078] Figure 4B An example user interface 400B of an application (e.g., a virtual assistant application) according to an embodiment of the present disclosure is shown. This application has a user interface for displaying interpretations of text content. The user interface 400B can be displayed when a user wishes to submit a query associated with an image and the text included in the image. Specifically, the user interface 400B includes a query response element 408. The query response element 408 may include the output of a machine learning model. The output of the machine learning model may be an interpretation of the user's input. Figure 4A The image displayed in the image is used as background input to the machine learning model in response to a submitted query. In addition to text content, responses can also include images, audio content, animations, video content, interactive content, web search results, and more.
[0079] In some examples, user interface 400B may also include interface elements that serve as links 410 to image displays. For example, if a user requests a summary of the text in an image, the user interface can be updated from user interface 400A, where the image is displayed, to user interface 400B, where the query response 408 is displayed. To facilitate the user switching back to user interface 400A, where the image is displayed, a link 410 to the image display is provided in the user interface where the query response element 408 is displayed.
[0080] Figure 5 This is an example image analysis system according to an exemplary embodiment of the present disclosure. The image analysis system 500 may include an image display system 502, a text extraction system 504, a feature analysis system 506, a prompt generation system 508, a machine learning model 510, and a response system 512.
[0081] Image display system 502 can display images (or video composed of multiple images) on the interface of a user computing device. In some examples, the displayed images are images previously captured by a camera associated with the user computing device. In some examples, these images are part of live video currently being captured by the user computing device. In some examples, the images were previously captured or captured by another user computing device and obtained by the current user computing device via a computer network. In some examples, the displayed images may include text content. Image display system 502 may include applications for capturing, displaying, and analyzing images.
[0082] Once the image display system 302 has displayed an image containing text, the text extraction system 504 can extract the text from the image. In some examples, the text in the image is automatically extracted using an OCR process. In other examples, text is extracted only when requested by the user. The extracted text may be referred to as the first set of text content.
[0083] In some examples, once the text has been extracted, the feature analysis system 506 can analyze the initial set of text content, images, and any input provided by the user to determine one or more features associated with the text / image. Features may include the language of the text, the density of the text, the context of the image / video (e.g., whether the image is learning material), any queries submitted by the user while the text / image is being displayed, and so on. For a specific example, the feature analysis system 506 may determine the density (e.g., words per pixel or another metric) associated with the text in the displayed image.
[0084] The feature analysis system 506 can determine the appropriate response type based on one or more features. For example, if the text density exceeds a threshold, the feature analysis system 506 can determine that the appropriate response type is a summary response type. Once the feature analysis system 506 has determined a suitable or appropriate response type, the image display system 502 can update the user interface to include elements associated with that response type. For example, if the response type is a summary response type, the image display system 502 can update the user interface to include a "Summary" button.
[0085] Users can enter requests via a user interface. In some examples, the user interface includes user interface elements associated with one or more response types. Each user interface element allows a user to request a specific type of response. In some examples, users can select the type of service to request by choosing a one-to-one label displayed below an image. If a user selects a particular label, the user interface can be updated to include user-selectable user interface elements associated with the specific request type associated with that label.
[0086] For example, if a user selects the "Summarize" tab, the user interface element could be a "Summarize this" button displayed near or above the text in the image. The user can select this button to request the system to summarize the text. In other examples, the user interface element is a general request element where the user can type in a prompt for a specific request. For instance, a user can request the system to interpret difficult text by opening a query input field and using natural language. In another example, when viewing difficult text in an image, the user can select a search icon that opens a window where they can type their interpretation request (or other response). The user's computing device can use natural language processing techniques to understand the request and generate an appropriate response.
[0087] Once a request has been received, the prompt generation system 508 can generate prompts based on that request to serve as input for a machine learning model. For example, the prompts may include an initial set of text content, the determined response type, information about the image from which the text is extracted as background, and any additional prompts received from the user regarding the request.
[0088] The prompt can be used as input to a machine learning model. For example, the machine learning model could be a large language model that takes the prompt as input and outputs a response based on the data included in the prompt. In some examples, the machine learning model could be hosted on a remote computing system, and inputting this data could involve transmitting the prompt to the remote computing system using one or more communication networks.
[0089] In some examples, the machine learning model can process the prompt and output a response. In some examples, the response includes a second set of text content. For example, if the request is a summary of text, the output could be a summary of the text. If the request is an explanation of the first set of text content, the output could be an explanation of the first set of text content. In some examples, the request is a query type where the user asks a question about text and / or images. In this example, the output could be a response to the query.
[0090] The response system 512 can receive the output of the machine learning model 510. The output can then be displayed to the user on the user interface of the user computing device. In some examples, the user computing device can display the output on a separate interface page from the original image. In some examples, the model's output can be displayed in the same user interface as the image.
[0091] Figure 6 A block diagram depicts an example computing device 600 performing according to an exemplary embodiment of the present disclosure. The computing device 600 may be a user computing device or a server computing device.
[0092] The computing device 600 includes multiple applications (e.g., applications 1 to N). Each application contains its own machine learning library and machine learning model. For example, each application may include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, search applications, query response applications, image display applications, etc.
[0093] like Figure 6 As shown, each application can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can use an API (e.g., a public API) to communicate with each device component. In some implementations, the API used by each application is application-specific.
[0094] Figure 7 A block diagram depicts an example computing device 700 performing according to an exemplary embodiment of the present disclosure. The computing device 700 may be a user computing device or a server computing device.
[0095] The computing device 700 includes multiple applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. In some implementations, each application may use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).
[0096] The central intelligence layer comprises multiple machine learning models. For example, a corresponding machine learning model can be provided for each application, and this model is managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model for all applications. In some implementations, the central intelligence layer is included within the operating system of the computing device or otherwise implemented by that operating system.
[0097] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized repository for data from the computing device 700. For example... Figure 7 As shown, the central device data layer can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer may use an API (e.g., a private API) to communicate with each device component.
[0098] Figure 8An example flowchart depicts a method for providing an appropriate response based on characteristics of an image and its contained text, according to an exemplary embodiment of this disclosure. One or more portions of the method may be implemented by one or more computing devices (such as, for example, the computing devices described herein). Furthermore, one or more portions of the method may be implemented as an algorithm on the hardware components of the devices described herein. Figure 8 Elements executed in a specific order are depicted for illustrative and discussion purposes. Those skilled in the art will understand using the disclosure provided herein that elements of any method discussed herein can be adjusted, rearranged, expanded, omitted, combined, and / or modified in various ways without departing from the scope of this disclosure. The method may be performed by one or more computing devices (such as…) Figure 1 and Figure 5 The computing device described may be implemented by one or more computing devices.
[0099] User computing devices (e.g., Figure 1 The user computing device 102 may include one or more processors, memory, and one or more sensors. The user computing device 102 (e.g., Figure 1 The user computing device 102 in the middle may include other components that together enable the user computing device 102 (e.g., Figure 1 The user computing device 102 in the image is able to analyze the image, determine one or more response types, and respond to user requests based on the image, the determined response types, and input from the user.
[0100] In some examples, the user computing device may acquire an image at 802, where the image depicts a first set of text content. In some examples, an optical character recognition process is used to generate text data representing the content of the first set of text content from the image.
[0101] In some examples, the user computing device (e.g., Figure 1 The user computing device 102 can determine one or more characteristics of the first set of text content at 803. Characteristics may include text density, content in an image, input from the user, etc.
[0102] In some examples, the user computing device (e.g., Figure 1The user computing device 102 can determine a response type from multiple response types at 804 based on one or more characteristics. In some examples, the multiple response types include summary responses, explanation responses, and query responses. In some examples, the user computing device can determine the density of a first set of text content within the image. In response to determining that the density of the first set of text content within the image meets a threshold, the user computing device can update the user interface to include a summary user interface element. In some examples, the user request is made through user input, whereby the user selects a "summary" user interface element displayed near the image in the user computing device's user interface.
[0103] In some examples, the user computing device (e.g., Figure 1 The user computing device 102 can generate model input for a machine learning model in response to user input associated with elements of the user interface. In some examples, the user request is made through user input that the user selects a "Summary" user interface element displayed near an image in the user computing device's user interface. In some examples, the second set of text content has less text content than the first set of text content.
[0104] In some examples, the user computing device (e.g., Figure 1 The user computing device 102 in the middle can generate model input at 806, wherein the model input includes data describing the first set of text content and prompts associated with the response type.
[0105] In some examples, the user computing device can provide model input at 808 as input to a machine learning language model. In some examples, the machine learning model is a large language model. The user computing system can generate model input in response to a user request. In some examples, the model input for the machine learning model is multimodal. In some examples, the machine learning language model runs on a remote server system, and the model input is transmitted to the remote server system, while a second set of text content is received from the remote server system.
[0106] The user computing device may receive at 810 a second set of text as the output of a machine learning language model, the output of which serves as the result of the machine learning language model processing the model input. The user computing device may provide the second set of text to the user for display at 812, wherein the content of the second set of text includes a summary of the content of the first set of text.
[0107] In some examples, the user's computing device updates the user interface to display a second set of text content. This second set of text content may contain less text than the first set.
[0108] The technologies discussed in this paper refer to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide range of possible configurations, combinations, and partitions of tasks and functions between and within components. For example, the processes discussed in this paper can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0109] While the subject matter has been described in detail with respect to various specific example embodiments, each example is provided by way of explanation and not limitation. Modifications, alterations, and equivalents of such embodiments will readily arise for those skilled in the art upon acquiring an understanding of the foregoing. Therefore, this disclosure does not exclude such modifications, alterations, and / or additions to the subject matter that will be readily apparent to those skilled in the art. For example, features shown or described as part of an embodiment may be used with another embodiment to produce further embodiments. Therefore, this disclosure is intended to cover such modifications, alterations, and equivalents.
Claims
1. A computing system, the system comprising: One or more processors; as well as One or more non-transitory computer-readable media, the one or more non-transitory computer-readable media collectively storing instructions, the instructions causing the computing system to perform operations when executed by the one or more processors, the operations including: Obtain an image, wherein the image depicts the first set of text content; Determine one or more characteristics of the first group of text content; Based on one or more of the aforementioned characteristics, the response type is determined from multiple response types; Generate model input, wherein the model input includes data describing the first set of text content and a prompt associated with the response type; The model input is provided as input to the machine learning language model; Receive a second set of text as the output of the machine learning language model, the output of the machine learning language model being the result of the machine learning language model processing the model input; and The second set of text is provided to the user for display, wherein the content of the second set of text is associated with the response type.
2. The computing system as described in claim 1, wherein, The multiple response types include summary responses, explanation responses, and query responses.
3. The computing system as described in claim 2, wherein, The determined response type is a summary response, and the second set of text content is a summary of the first set of text content.
4. The computing system as described in claim 3, wherein, The one or more characteristics of the first group of text content include the density of the first group of text content, and the operation of determining the response type from multiple response types based on the one or more characteristics further includes: Determine the density of the first group of text content within the image; In response to determining that the density of the first group of text content within the image satisfies a threshold, the response type is determined to be a summary response type; and Update the user interface to include a summary of user interface elements.
5. The computing system as described in claim 4, wherein, Determining the density of the first group of text content within the image further includes: Determine the area in the image that includes the first set of text content; Determine the total area of the image; and Determine the percentage of the image that includes the first set of text content.
6. The computing system as claimed in claim 4, wherein, Determining the density of the first group of text content within the image further includes: Determine the total number of characters visible in the image.
7. The computing system of claim 6, wherein, Determining the density of the first group of text content within the image further includes: Determine the number of characters per pixel in the image.
8. The computing system as claimed in claim 4, wherein, The computing system generates the model input in response to a user request.
9. The computing system of claim 8, wherein, The user request is made by user input, and the user selects the summary user interface element displayed near the image in the user interface of the user computing device.
10. The computing system of claim 3, wherein, The second group of text content has less text content than the first group of text content.
11. The computing system of claim 1, wherein, The optical character recognition process is used to generate text data representing the content of the first set of text from the image.
12. The computing system of claim 1, wherein, The machine learning language model is a large-scale language model.
13. The computing system of claim 1, wherein the operation further comprises: Update the user interface to display the second set of text content.
14. The computing system of claim 1, wherein, The input to the machine learning model is multimodal.
15. The computing system of claim 1, wherein, The machine learning language model runs on a remote server system, and the model input is transmitted to the remote server system, while the second set of text content is received from the remote server system.
16. A computer-implemented method for responding to a query about an image, the method comprising: An image is obtained by a computing system having one or more processors, wherein the image depicts a first set of text content; The computing system determines one or more characteristics of the first group of text content; The computing system determines the response type from multiple response types based on one or more of the aforementioned characteristics; The computing system generates model input, wherein the model input includes data describing the first set of text content and prompts associated with the response type; The computing system provides the model input as input to the machine learning language model; The computing system receives a second set of text as the output of the machine learning language model, and the output of the machine learning language model is the result of the machine learning language model processing the model input. as well as The computing system provides the second set of text to the user for display, wherein the content of the second set of text is associated with the response type.
17. The computer-implemented method as described in claim 16, wherein, The image includes image content, and the model input includes data describing the image content.
18. The computer-implemented method as described in claim 16, wherein, The determined response type is a query response, and the second set of text content is a request response to a query about the content within the image and the first set of text content.
19. The computer-implemented method as described in claim 18, wherein, Determining a response type from multiple response types based on one or more of the aforementioned characteristics further includes: The computing system accesses the text input by the user into the text field; The computing system analyzes the text to determine whether the text is a question related to the image; In response to determining that the text is associated with the image, the computing system determines that the response type is a query response type; and The computing system updates the user interface to include query interface elements.
20. The computer-implemented method as described in claim 16, wherein, The first set of text content and the image content are used as context by the machine learning language model in response to the query.
21. The computer-implemented method as described in claim 16, wherein, The user query was received via voice communication.
22. The computer-implemented method as described in claim 16, wherein, The user query is received when the image is displayed on the display of the user's computing device.
23. A computing system, the system comprising: One or more processors; as well as One or more non-transitory computer-readable media, the one or more non-transitory computer-readable media collectively storing instructions, the instructions causing the computing system to perform operations when executed by the one or more processors, the operations including: Obtain an image, wherein the image depicts the first set of text content; Determine one or more characteristics of the first group of text content; Based on one or more of the aforementioned characteristics, the response type is determined from multiple response types; Generate model input, wherein the model input includes data describing the first set of text content and a prompt associated with the response type; The model input is provided as input to the machine learning language model; Receive a second set of text as the output of the machine learning language model, the output of the machine learning language model being the result of the machine learning language model processing the model input; and The second set of text is provided to the user for display, wherein the content of the second set of text is associated with the response type.