Social platform image understanding method and system for assisting visually impaired users

By building a multimodal large language model, combining screenshot images and context information of the social platform, generating detailed descriptions and conducting interactive Q&A, the problem of inability to meet the diverse information needs of visually impaired users in the existing technology is solved, and comprehensive information acquisition and social participation of visually impaired users are achieved.

CN118779442BInactive Publication Date: 2025-05-13ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411248975.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing image description systems cannot fully consider the context information of the image and the preferences of different visually impaired individuals, resulting in the inability to meet the diverse information needs of visually impaired users, affecting their understanding of image content.

Method used

By obtaining screenshot images and context information of social platforms, a multimodal large language model is built, a detailed description is generated and interactive Q&A is performed, and the answer text is optimized according to user preferences.

Benefits of technology

It achieves comprehensive and specific style preference information acquisition for visually impaired users, and improves their understanding of image content and sense of social participation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118779442B_ABST
    Figure CN118779442B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for assisting visually impaired users in understanding images on social platforms, including: obtaining screenshot images of social platforms to construct a data set; constructing a corresponding interactive network based on a large language model framework, wherein the interactive recognition network includes a graphic information generation module, an image exploration module, and an interactive question-and-answer module; using the data set to train the interactive network to obtain a multimodal large model for assisting in understanding image information; inputting screenshots of information to be read on the social platform and the user's voice text into the multimodal large model to output the corresponding answer text. The present invention also provides a social platform image understanding system. The method provided by the present invention can help visually impaired users quickly understand and fully understand the information on social platforms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning technology, and in particular, relates to a social platform image understanding method and system for assisting visually impaired users. Background Art

[0002] With hundreds of millions of photos uploaded every day on social platforms such as social networking services, interacting with visual content has become an important part of the current online social experience. Images on social networking services not only show personal life and emotions, but also serve as an important medium for information dissemination and social interaction. However, for visually impaired people with blindness or low vision, they are often excluded from mainstream social circles because they have difficulty accessing the rich information in images.

[0003] Traditional image description methods usually rely on manually generated alternative text for images, resulting in a large demand for human labor and long processing time.

[0004] Patent document CN116030264A discloses a method and device for assisting visually impaired people to understand pictures. First, an image uploaded by a user is obtained and features are extracted from the image. Then, the number of feature points of the image, the distribution rate of feature points of the image, the image height and the image width are marked, and the image determination coefficient is calculated. Then, a standard image determination coefficient and a determination threshold are set, and the determination coefficient is calculated using the image determination coefficient and the standard image determination coefficient. All the determination coefficients are combined into a determination set, and feature matching is performed on the determination threshold and the determination set. If they match, the image data corresponding to the matched determination coefficient is extracted. Finally, the text information of the image data corresponding to the matched determination coefficient is extracted, and the text information is converted into voice information, and the voice is provided to the visually impaired.

[0005] Patent document CN114220034A discloses an image processing method, device, terminal and storage medium, including identifying the screen display content of the terminal through a screen reading service; executing image prompts when the screen reading service identifies that the screen display content contains images; identifying the image content in the screen display content and obtaining image content description text when receiving a voice for instructing image recognition in the screen display content; and playing a voice corresponding to the image content description text through a voice service.

[0006] However, the general image descriptions generated by the image description systems in the above inventions are too one-size-fits-all, ignoring the impact of the image's contextual information (such as text, comments, etc.) and the preferences of different visually impaired individuals on the image description. Therefore, they cannot fully meet the diverse information needs of visually impaired users, which in turn affects their understanding of the image content. In addition, the current invention focuses on providing one-way narrative image descriptions, which prevents visually impaired individuals from deeply exploring and verifying their understanding of the image by interacting with the system, reducing the interpretability of the system and the visually impaired users' trust in the descriptions, thereby hindering their social participation. Summary of the invention

[0007] The purpose of the present invention is to provide a method and system for assisting visually impaired users in understanding images on social platforms, which can help visually impaired users quickly and comprehensively understand the information on social platforms.

[0008] In order to achieve the first objective of the present invention, the following technical solution is provided: a method for assisting visually impaired users in understanding social platform images, comprising the following steps:

[0009] Acquire screenshot images of a social platform to construct a data set, wherein the data set includes image elements in the screenshot images and corresponding context information and element types;

[0010] Building a corresponding interactive network based on the large language model framework, the interactive recognition network includes a graphic information generation module, an image exploration module and an interactive question-answering module;

[0011] The graphic information generation module is used to generate a detailed description corresponding to the context information in the screenshot image;

[0012] The image exploration module generates corresponding salient features based on the selected target image elements and the corresponding detailed descriptions, wherein the salient features include enhanced text features and enhanced image features, and classifies and sorts the input images based on the salient features and element types to output corresponding recommendation information, wherein the recommendation information includes key points of the target image elements and corresponding brief descriptions;

[0013] The interactive question-and-answer module is used to obtain the input voice text and generate an answer text that meets the preset user preferences based on the results output by the graphic information generation module and the image exploration module, as well as the input voice text;

[0014] The interactive network is trained using the dataset to obtain a large multimodal model for assisting in understanding image information.

[0015] The screenshots of the information to be read on the social platform and the user's voice text are input into the multimodal large model to output the corresponding answer text.

[0016] The present invention obtains context and user-preferred image descriptions, groups objects in the image according to semantic associations, and guides users to explore objects in order of importance, thereby helping visually impaired users to quickly grasp the key points of the image and gain a comprehensive understanding.

[0017] Specifically, the image types in the dataset include activities and experiences, expression of emotions and opinions, sharing of items, interpersonal relationships, portraits, and artistic creations.

[0018] Specifically, the graphic information generation module includes a visual encoder, a projection matrix and a language generation module. The visual encoder is used to extract features from the input screenshot image to generate visual features. The projection matrix is ​​used to convert the generated visual features into language embedding tags. The language generation module is used to generate a corresponding detailed description based on the generated language embedding tags and the input voice text.

[0019] Specifically, the visual encoder uses a contrastive language-image pre-trained visual encoder.

[0020] Specifically, the salient features use a self-attention mechanism to extract and enhance the input target image elements and detailed descriptions to obtain corresponding enhanced text features and enhanced image features.

[0021] Specifically, in the attention mechanism, the image features are used as the key and value, and the text features are used as the query to perform cross-attention operations to obtain enhanced text features for fusion;

[0022] With text features as keys and values, image features are used as queries to perform cross-attention operations to obtain enhanced image features for fusion.

[0023] Specifically, the recommendation information calculates the similarity between the enhanced image features and the enhanced text features as a feature index, and uses the feature index and image type to predict image elements and detailed descriptions and extract noun phrases.

[0024] Specifically, the image exploration module generates an embedded representation of each image element in the screenshot image through an image segmentation model, and uses the coordinates of the area clicked by the user as a prompt embedding, generates a corresponding segmentation mask based on the generated embedded representation and the prompt embedding, and outputs its related confidence score, and takes the image element corresponding to the segmentation mask with the highest confidence score as the target image element.

[0025] Specifically, the user preferences include first-person or third-person narrative, aesthetic and emotional analysis, and description style, wherein the description style includes a diverse style, a neutral style, or a conservative style.

[0026] In order to achieve the second objective of the present invention, the following technical solution is provided: a social platform image understanding system, which is implemented by the above-mentioned social platform image understanding method for assisting visually impaired users, comprising:

[0027] The data collection module is used to obtain the screenshot images of the current social platform, as well as the context information and image type in the image;

[0028] Style setting module, used to set the user's interaction style preferences;

[0029] A data analysis module, which analyzes the collected screenshot images, context information and image type and the image elements and element types selected by the user to generate a response text;

[0030] The question-answer interaction module is used to obtain the user's voice text and optimize the generated answer text based on the preset interaction style preferences to output the answer text that meets the user's preferences.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] By obtaining image descriptions through an image description function that considers context and user preferences, visually impaired users can be assisted in obtaining comprehensive information with specific style preferences, thereby further enhancing their sense of social participation. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A flowchart of a method for assisting visually impaired users in understanding social platform images provided in this embodiment;

[0034] Figure 2 A schematic diagram of the workflow of the multimodal large model provided in this embodiment;

[0035] Figure 3 A schematic diagram of the framework of the graphic information generation module provided in this embodiment;

[0036] Figure 4 A schematic diagram of the framework of the image exploration module provided in this embodiment;

[0037] Figure 5 A schematic diagram of the framework of the image segmentation model in the image exploration module provided in this embodiment;

[0038] Figure 6 This is a demonstration process of the social platform image understanding system provided in this embodiment. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0040] like Figure 1 As shown, a social platform image understanding method for assisting visually impaired users provided in this example includes image description that considers context and user preferences, key point-prioritized image exploration, and open-ended visual question answering to help visually impaired users effectively and easily understand the image content on social platforms.

[0041] The specific process is as follows:

[0042] A screenshot image of a social platform is obtained to construct a data set, wherein the data set includes image elements in the screenshot image and corresponding context information and element types.

[0043] Building a corresponding interactive network based on the large language model framework, the interactive recognition network includes a graphic information generation module, an image exploration module and an interactive question-answering module;

[0044] The graphic information generation module is used to generate a detailed description corresponding to the context information in the screenshot image;

[0045] The image exploration module generates corresponding salient features based on the selected target image elements and the corresponding detailed descriptions, wherein the salient features include enhanced text features and enhanced image features, and classifies and sorts the input images based on the salient features and element types to output corresponding recommendation information, wherein the recommendation information includes key points of the target image elements and corresponding brief descriptions;

[0046] The interactive question-and-answer module is used to obtain the input voice text, and generate an answer text that meets the preset user preferences based on the results output by the graphic information generation module and the image exploration module, as well as the input voice text.

[0047] The interactive network is trained using the dataset to obtain a large multimodal model for assisting in understanding image information.

[0048] The screenshots of the information to be read on the social platform and the user's voice text are input into the multimodal large model to output the corresponding answer text.

[0049] Furthermore, the image types in the dataset are evaluated through the Voiceover function in the iOS operating system and existing commercial multimodal large models (such as Wenxin Yiyan or GPT-4V) as well as the cosine similarity between Dunnett's image descriptions and human-generated image descriptions, including activities and experiences (AE), expression of emotions and opinions (EO), item sharing (GS), interpersonal relationships (IR), portraits (PP), and artistic creations (AC).

[0050] like Figure 2 The flowchart of the multimodal big model provided in this embodiment is shown, in which before browsing pictures in social media, users can preset corresponding user preferences including first-person or third-person narrative methods, aesthetic and emotional analysis, and description styles, wherein the description styles include diverse styles, neutral styles or conservative styles.

[0051] Furthermore, the analysis of aesthetics and sentiment was evaluated by five human raters according to the following criteria: 1- incorrect emotion polarity (e.g., positive or negative) and inappropriate or offensive comments; 2- correct emotion polarity, but the comments had low relevance to the image; 3- correct emotion type (e.g., joy or gratitude) with reasonable comments about the content; 4- correct emotion type and intensity (e.g., disgust or disgust) with unique and insightful comments.

[0052] Compared with the commercial model, the performance of the model provided in this embodiment in sentiment analysis (M=2.17, SD=0.87) is significantly better than the existing commercial large model (M=3.47, SD=0.55), and the Wilcoxon signed rank test shows that Z=9.699, and p<0.001 overall. This score shows that by integrating contextual information and focusing on the key aspects of the image, the model provided in this embodiment can accurately determine the emotional type of the post and generate relevant comments.

[0053] The model provided in this embodiment can obtain descriptions of these images through an image description function that takes into account context and user preferences. These descriptions include context information of the images, information requirements for different types of images, and user preferences.

[0054] If the user needs to further explore specific image elements after obtaining the description of the image, for example, if he wants to understand the spatial distribution information of different objects in the image, he can explore the image through the key point priority image exploration function.

[0055] After exploring the image, users may have additional personalized needs for the image. At this time, users can freely ask questions about the social posts in the image through an open interactive question-and-answer module.

[0056] like Figure 3 As shown, the graphic information generation module provided in this embodiment is constructed based on the LLaVA architecture, and includes a visual encoder, a projection matrix and a language generation module. The visual encoder is used to extract features of the input screenshot image to generate visual features, and the projection matrix is ​​used to convert the generated visual features into language embedding tags. The language generation module is used to generate a corresponding detailed description according to the generated language embedding tags and the input voice text.

[0057] The visual encoder is responsible for processing the image through the contrastive language-image pre-trained visual encoder (ViT-L / 14) to generate visual feature representation.

[0058] Then, the projection matrix is ​​used to convert the visual features into language embedding tags. Finally, the language generation module processes the language embedding tags and the language instructions input at the beginning to generate a language response. After receiving the above image and context information, the multimodal large model generates a summary description based on user preferences to summarize the important information in this post. Among them, different output styles are achieved by controlling the temperature of LLaVA. The higher the temperature, the more creative the generated content; the lower the temperature, the more conservative and stable the generated content.

[0059] like Figure 4 As shown, the image exploration module provided by this embodiment is set in this embodiment to perform an image exploration task after continuously clicking the same position of the screen, that is, the objects in the picture are grouped according to semantic associations, and the user is guided to explore the objects by touching in sequence according to the priority of importance, so as to help them quickly grasp the key information of the image and form a comprehensive understanding. Subsequently, referring to the user's information needs for the category of images, the number of key features of each object is counted and the importance of each object is ranked according to the number of key features and the contextual information of the image. Objects with similar semantics and equal importance are merged into one. The object names will be stored in the database in the form of an array.

[0060] After determining the priority, the model sends the name of the first object in the array as a cue word to the object detection model to obtain the coordinates of the object. The front end designates the area covered by the object as the only interactive area based on the coordinates of the object.

[0061] More specifically, the image exploration module receives two inputs, image and text, at the same time. The text input uses Bert to extract text features; the image input uses Swin Transformer to extract image features. Then, the image features and text features enter the Feature Enhancer for processing. Feature Enhancer uses multiple feature enhancement layers for feature fusion. The first layer is the self-attention layer, which is used to aggregate and enhance information within the image features and text features. That is, for image features, the self-attention mechanism is used to calculate the correlation between each position in the image feature and other positions, and the features are re-weighted based on these correlations. For text features, the self-attention mechanism is used: the correlation between each word in the text feature and other words is calculated, and the features are re-weighted based on these correlations.

[0062] The enhanced image features and text features are output and sent to the second layer, the image-to-text cross-attention layer. The image-to-text cross-attention layer is used to fuse the information in the image features into the text features.

[0063] The image-to-text cross-attention layer calculates the correlation between image features and text features. It uses image features as Key and Value, and text features as Query for cross-attention calculation, and finally outputs the updated text features. The third layer is the image-to-text cross-attention layer, which is used to fuse the information in the text features into the image features. It receives two inputs: the enhanced image features and the updated text features. The image cross-attention layer calculates the correlation between text features and image features, uses text features as Key and Value, and image features as Query for cross-attention calculation, and finally outputs the updated image features.

[0064] The fourth layer is a feedforward neural network, which is used to further perform nonlinear transformation and enhancement on features. It receives two inputs: updated image features and updated text features. Each feature passes through a two-layer feedforward neural network, and finally outputs the updated image features and text features.

[0065] Finally, the dot product is used to calculate the similarity score between each image feature and each text feature, and then the highest similarity score of each image feature is extracted from the similarity score matrix, and finally the image feature index that is most relevant to the input text is selected. These selected features will be used as the initial query for the cross-modal decoder, so that the decoder can better combine image and text information for subsequent object detection and noun phrase extraction. Specifically, each query of the cross-modal query first passes through a self-attention layer to focus on different parts of itself to enhance feature representation. Then, the query passes through an image cross-attention layer to interact with the image features to extract image information related to the query. Next, the query passes through a text cross-attention layer to interact with the text features to extract text information related to the query. Finally, the query passes through a feed-forward neural network for further processing and feature extraction. The Cross-Modality Decoder outputs an updated cross-modal query, which contains features that fuse image and text information, and is used to predict object bounding boxes and related noun phrases. Finally, the object detection model uses multiple loss functions to calculate the error between the model prediction and the true annotation to guide model optimization.

[0066] like Figure 5 As shown in the figure, the image segmentation model selected for the image exploration module in this embodiment includes three modules: image encoder, prompt encoder and mask decoder. The input of the entire model is a picture, which first enters the image encoder. The image encoder is used to process high-resolution images and create image embeddings. Specifically, it generates an embedded representation of the image through the Vision Transformer (ViT) pre-trained with MAE. The prompt encoder is used to process various types of prompts (user input) and embed them into a format that interacts with the image embedding. The mask decoder is used to generate segmentation masks from the image embedding and the prompt embedding, and outputs valid masks and their associated confidence scores, indicating the accuracy of each mask. Finally, the mask with the highest score is selected as the final result.

[0067] Once the user has finished exploring the elements inside an object, they can swipe right to exit and the system will guide them to the next object. In addition, the user can explore the previous or next object by swiping left or right.

[0068] The specific operation is as follows: the location information of the object in the image is generated from the object name to guide the user to explore the object in the area where the object is located. In addition, the user can explore the various elements contained in the object by clicking on the object. When the element is selected, the system uses the image segmentation model Segment Anything Model to return a blue mask representing the area covered by the element. The system then transmits the object and the blue mask to the multimodal large model and asks it to describe the elements blocked by the blue mask.

[0069] This embodiment further provides a social platform image understanding system, which is implemented by the social platform image understanding method for assisting visually impaired users provided in the above embodiment, including:

[0070] A data collection module is used to obtain screenshot images of the current social platform, as well as context information and image types in the screenshot images;

[0071] Style setting module, used to set the user's interaction style preferences;

[0072] A data analysis module, which analyzes the collected screenshot images, context information and image type and the image elements and element types selected by the user to generate a response text;

[0073] The question-answer interaction module is used to obtain the user's voice text and optimize the generated answer text based on the preset interaction style preferences to output the answer text that meets the user's preferences.

[0074] like Figure 6 As shown in the figure, it is a demonstration flow chart of the social platform image understanding system, where Figure 6 The a in the text is used to set user preferences in advance, including whether to describe in the first person or the third person, the type of description style, and whether aesthetic analysis and sentiment analysis are required.

[0075] like Figure 6 As shown in b in the figure, the user selects and clicks on the picture in the interface and then performs voice input to generate a summary description, open-ended visual question and answer content, and a single image description.

[0076] like Figure 6 As shown in c in , after the user clicks on the same picture three times, key point-prioritized image search content is generated, which includes a description of each element in the image.

[0077] In addition, the terms "upper", "lower", "inner", "outer", "front", "rear" are used for descriptive purposes only and should not be understood as indicating or implying relative importance. Unless otherwise specifically stated, the relative steps, numerical expressions and values ​​of the components and steps described in these embodiments do not limit the scope of the present invention.

[0078] Of course, the above description is only a specific embodiment of the present invention and is not intended to limit the scope of implementation of the present invention. All equivalent changes or modifications made according to the structure, characteristics and principles described in the patent application scope of the present invention should be included in the patent application scope of the present invention.

[0079] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for assisting visually impaired users in understanding social platform images, characterized in that: The following steps are involved: Acquire screenshot images of a social platform to construct a data set, wherein the data set includes image elements in the screenshot images and corresponding context information and element types; Building a corresponding interactive network based on the large language model framework, the interactive network includes a graphic information generation module, an image exploration module and an interactive question-answering module; The graphic information generation module is used to generate a detailed description corresponding to the context information in the screenshot image; The image exploration module generates corresponding salient features based on the selected target image elements and the corresponding detailed descriptions, wherein the salient features include enhanced text features and enhanced image features, and classifies and sorts the screenshot images based on the salient features and element types to output corresponding recommendation information, wherein the recommendation information includes key points of the target image elements and corresponding brief descriptions; The image exploration module generates an embedded representation of each image element in the screenshot image through an image segmentation model, and embeds a prompt according to the coordinates of the area clicked by the user, generates a corresponding segmentation mask based on the generated embedded representation and the prompt embedding, and outputs its related confidence score, and takes the image element corresponding to the segmentation mask with the highest confidence score as the target image element; The salient features use a self-attention mechanism to extract and enhance the input target image elements and detailed descriptions to obtain corresponding enhanced text features and enhanced image features. In the attention mechanism, the image features are used as keys and values, and the text features are used as queries for cross-attention operations to obtain enhanced text features for fusion; Using text features as keys and values, and image features as queries, we perform cross-attention operations to obtain enhanced image features for fusion. The recommendation information is obtained by calculating the similarity between the enhanced image features and the enhanced text features as feature indexes, and using the feature indexes and element types to predict image elements and detailed descriptions and extract noun phrases; The interactive question-and-answer module is used to obtain the input voice text and generate an answer text that meets the preset user preferences based on the results output by the graphic information generation module and the image exploration module, as well as the input voice text; The interactive network is trained using the dataset to obtain a large multimodal model for assisting in understanding image information. The screenshots of the information to be read on the social platform and the user's voice text are input into the multimodal large model to output the corresponding answer text.

2. The method for assisting visually impaired users in understanding social platform images according to claim 1, characterized in that: The types of elements in the dataset include activities and experiences, expressions of emotions and opinions, sharing of items, interpersonal relationships, portraits of people, and artistic creations.

3. The method for assisting visually impaired users in understanding social platform images according to claim 1, characterized in that: The graphic information generation module includes a visual encoder, a projection matrix and a language generation module. The visual encoder is used to extract features from the input screenshot image to generate visual features. The projection matrix is ​​used to convert the generated visual features into language embedding tags. The language generation module is used to generate a corresponding detailed description based on the generated language embedding tags and the input voice text.

4. The method for assisting visually impaired users in understanding social platform images according to claim 3, characterized in that: The visual encoder uses a contrastive language-image pre-trained visual encoder.

5. The method for assisting visually impaired users in understanding social platform images according to claim 1, characterized in that: The user preferences include first-person or third-person narrative, analysis of aesthetics and emotions, and description style, wherein the description style includes a diverse style, a neutral style, or a conservative style.

6. A social platform image understanding system, characterized in that: The method is implemented by the social platform image understanding method for assisting visually impaired users as claimed in any one of claims 1 to 5, comprising: A data collection module is used to obtain screenshots of the current social platform and context information in the screenshots; Style setting module, used to set the user's interaction style preferences; The data analysis module analyzes the image elements and element types selected by the user based on the collected screenshot images and context information to generate the answer text; The question-answer interaction module is used to obtain the user's voice text and optimize the generated answer text based on the preset interaction style preferences to output the answer text that meets the user's preferences.

Citation Information

Patent Citations

  • Method and device for assisting visually impaired person in understanding pictures

    CN116030264A

  • Image processing method and device, terminal and storage medium

    CN114220034A

  • Segmenting content displayed on a computing device into regions based on pixels of a screenshot image that captures the content

    US20170330336A1