Method, device and medium for visual perception

The model architecture addresses localization limitations in language models by integrating region tokenization, enhancing localized understanding and reducing computational demands, thus improving performance in real-world applications.

WO2025217935A1PCT designated stage Publication Date: 2025-10-23BEIJING YOUZHUJU NETWORK TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/088971
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Current language models lack effective localization capabilities, leading to challenges in real-world applications such as robotics and augmented reality, and require excessive computational resources for high-resolution image processing.

Method used

A model architecture that integrates region tokenization to identify and encode Regions of Interest (ROIs) into region visual tokens, decoupling localization and recognition, and using region visual tokens to ground textual outputs, reducing computational demands by handling high-resolution images through image tokenization.

Benefits of technology

The proposed solution enhances localized understanding and reduces computational load while maintaining localization accuracy, enabling seamless integration of localization and recognition capabilities, outperforming comparable models on referring and grounding benchmarks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024088971_23102025_PF_FP_ABST
    Figure CN2024088971_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the disclosure provide a solution for visual perception. A method includes extracting a feature map from a target image. The method further includes determining, based on the feature map, a plurality of image regions of the target image within which respective objects are located. The method further includes determining, from the feature map, a plurality of region visual tokens for the plurality of image regions, respectively. The method further includes providing a model input to a machine learning model, to obtain a model output of the machine learning model. The model input at least comprises the feature map, the plurality of region visual tokens, and a user query for the target image. The method further includes generating a response to the user query based on the model output.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD, DEVICE AND MEDIUM FOR VISUAL PERCEPTIONFIELD

[0001] The disclosed example embodiments relate generally to machine learning and, more particularly, to a method, apparatus, device and computer readable storage medium for visual perception.BACKGROUND

[0002] Multimodal language models have extended the realm of artificial general intelligence from language to the visual domain, igniting new possibilities. Leveraging the core strengths of language models, they shine in vision-language tasks that demand sophisticated comprehension and reasoning, including image captioning and visual question answering.SUMMARY

[0003] In a first aspect of the present disclosure, there is provided a method for visual perception. The method comprises: extracting a feature map from a target image; determining, based on the feature map, a plurality of image regions of the target image within which respective objects are located; determining, from the feature map, a plurality of region visual tokens for the plurality of image regions, respectively; providing a model input to a machine learning model, to obtain a model output of the machine learning model, the model input at least comprising the feature map, the plurality of region visual tokens, and a user query for the target image; and generating a response to the user query based on the model output.

[0004] In a second aspect of the present disclosure, there is provided an apparatus for visual perception. The apparatus comprises: a feature map extracting module configured to extract a feature map from a target image; an image region determining module configured to determine, based on the feature map, a plurality of image regions of the target image within which respective objects are located; a region visual token determining module configured to determine, from the feature map, a plurality of region visual tokens for the plurality of image regions, respectively; a model input providing module configured to provide a model input to a machine learning model, to obtain a model output of the machine  learning model, the model input at least comprising the feature map, the plurality of region visual tokens, and a user query for the target image; a response generating module configured to generate a response to the user query based on the model output.

[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions executable by the at least one processing unit. The instructions, upon execution by the at least one processing unit, cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program which, when executed by a processor, causes the method of the first aspect to be implemented.

[0007] It would be appreciated that the content described in the Summary section of the present invention is neither intended to identify key or essential features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily envisaged through the following description.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent in combination with the accompanying drawings and with reference to the following detailed description. In the drawings, the same or similar reference symbols refer to the same or similar elements, where:

[0009] FIG. 1 illustrates a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;

[0010] FIG. 2 illustrates a schematic diagram of an example model architecture for visual perception in accordance with some embodiments of the present disclosure;

[0011] FIG. 3 illustrates a schematic diagram of an example of model input and model output in accordance with some embodiments of the present disclosure;

[0012] FIG. 4 illustrates a flow chart of a process for visual perception in accordance with some embodiments of the present disclosure;

[0013] FIG. 5 illustrates a block diagram of an apparatus for visual perception according  to some embodiments of the present disclosure; and

[0014] FIG. 6 illustrates a block diagram of an electronic device in which one or more embodiments of the present disclosure can be implemented.DETAILED DESCRIPTION

[0015] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it would be appreciated that the present disclosure may be implemented in various forms and should not be interpreted as limited to the embodiments described herein. On the contrary, these embodiments are provided for a more thorough and complete understanding of the present disclosure. It would be appreciated that the drawings and embodiments of the present disclosure are only for the purpose of illustration and are not intended to limit the scope of protection of the present disclosure.

[0016] In the description of the embodiments of the present disclosure, the term “including” and similar terms would be appreciated as open inclusion, that is, “including but not limited to” . The term “based on” would be appreciated as “at least partially based on”. The term “one embodiment” or “the embodiment” would be appreciated as “at least one embodiment” . The term “some embodiments” would be appreciated as “at least some embodiments” . Other explicit and implicit definitions may also be included below. As used herein, the term “model” can represent the matching degree between various data. For example, the above matching degree can be obtained based on various technical solutions currently available and / or to be developed in the future.

[0017] It will be appreciated that the data involved in this technical proposal (including but not limited to the data itself, data acquisition or use) shall comply with the requirements of corresponding laws, regulations, and relevant provisions.

[0018] It will be appreciated that before using the technical solution disclosed in each embodiment of the present disclosure, users should be informed of the type, the scope of use, the use scenario, etc. of the personal information involved in the present disclosure in an appropriate manner in accordance with relevant laws and regulations, and the user’s authorization should be obtained.

[0019] For example, in response to receiving an active request from a user, a prompt  message is sent to the user to explicitly prompt the user that the operation requested operation by the user will need to obtain and use the user's personal information. Thus, users may select whether to provide personal information to the software or the hardware such as an electronic device, an application, a server, or a storage medium that perform the operation of the technical solution of the present disclosure according to the prompt information.

[0020] As an optional but non-restrictive implementation, in response to receiving the user's active request, the method of sending prompt information to the user may be, for example, a pop-up window in which prompt information may be presented in text. In addition, pop-up windows may also contain selection controls for users to choose “agree” or “disagree” to provide personal information to electronic devices.

[0021] It will be appreciated that the above notification and acquisition of user authorization process are only schematic and do not limit the implementations of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0022] As used herein, the term “model” can learn a correlation between respective inputs and outputs from training data, so that a corresponding output can be generated for a given input after training is completed. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural networks model is an example of a deep learning-based model. As used herein, “model” may also be referred to as “machine learning model” , “learning model” , “machine learning network” , or “learning network” , and these terms are used interchangeably herein.

[0023] “Neural networks” are a type of machine learning network based on deep learning. Neural networks are capable of processing inputs and providing corresponding outputs, typically comprising input and output layers and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically comprise many hidden layers, thereby increasing the depth of the network. The layers of neural networks are sequentially connected so that the output of the previous layer is provided as input to the latter layer, where the input layer receives the input of the neural network, and the output of the output layer serves as the final output of the neural network. Each layer of a neural network comprises one or more nodes (also known as processing  nodes or neurons) , each of which processes input from the previous layer.

[0024] Usually, machine learning can roughly comprise three stages, namely training stage, test stage, and application stage (also known as inference stage) . During the training stage, a given model can be trained using a large scale of training data, iteratively updating parameter values until the model can obtain consistent inference from the training data that meets the expected objective. Through the training, the model can be considered to learn the correlation between input and output (also known as input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the test stage, test inputs are applied to the trained model to test whether the model can provide correct outputs, thereby determining the performance of the model. In the application stage, the model can be used to process actual inputs and determine corresponding outputs based on the parameter values obtained from training.

[0025] FIG. 1 illustrates a block diagram of an example environment 100 in which various embodiments of the present disclosure may be implemented. In the environment 100 of FIG. 1, a computer system 110 applies a machine learning model 105 to perform visual perception. The machine learning model 105 is configured to receive and process a target image 112 and a user query 114 with respect to the target image 112, and then generate a response 116 to the user query 114. For example, the user query 114 may be a question regarding to an object or a certain region in the target image 112. The response 116 may be a reply to the question based on the visual information in the target image 112. In some embodiments, the machine learning model 105 may further receive a user-defined region input (e.g., bounding box) and generate long-form responses that are grounded to visual context.

[0026] In FIG. 1, the computer system 110 may include any computing system with computing capability, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices may include any type of mobile terminals, fixed terminals, or portable terminals, including mobile phones, desktop computers, laptops, netbooks, tablets, media computers, multimedia tablets, or any combination of the aforementioned, including accessories and peripherals of these devices or any combination thereof. Servers include but are not limited to mainframe, edge computing nodes, computing devices in cloud environment, etc.

[0027] It should be understood that the structure and function of each element in the environment 100 is described for illustrative purposes only and does not imply any limitations on the scope of the present disclosure.

[0028] As mentioned above, language models shine in vision-language tasks that demand sophisticated comprehension and reasoning. However, despite these achievements, current language models typically fall short of localization capabilities, thus cannot ground understanding to the visual context. Such limitations constrain the model from fulfilling its potential in real-world applications like robotics, autonomous driving, and augmented reality.

[0029] Considering the gap, one stream of research attempts to augment the language models to directly output quantized object coordinates for localization. While this method is simple in design, the substantial computational demands of language models make it challenging to process high-resolution image inputs, which are essential for accurate localization. Besides, the nature of sequence outputs in the language models is not well-suited for dense prediction tasks such as segmentation. These concerns elicit another stream of research, which incorporates an external localization module to decode bounding boxes or masks. This approach circumvents aforementioned issues but introduces additional latency in inference as it requires processing the image input twice with the language models and the localization module, respectively.

[0030] It is expected to explore a new paradigm for grounded language models. Drawing inspiration from open-vocabulary object detection, the grounding task may be decomposed into two sub-problems: discovering the object (localization) and relating the object to texts (recognition) . It is noted that localization alone requires little semantic understanding but demands perceptual skills, which is typically out of the scope of a language model's expertise. It is inspired to decouple localization and recognition within language models. But instead of using external modules, it is proposed to exploit the spatial understanding capability in the visual tokenizer of language models for localization. This perceive-then-understand design also resembles human vision process from a high-level perspective.

[0031] To address the above limitations, embodiments of the present disclosure propose an improved solution. In the solution, a model architecture may be introduced with localized and fine-grained visual perception abilities. Specifically, the embodiments of the present  disclosure incorporate region tokenization alongside standard image tokenization to identify and encode potential Regions of Interests (ROIs) into region visual tokens. During this process, location information is extracted from the image and associated with region visual tokens, with each region visual token anchored to the underlying ROI. This allows the model architecture 200 to ground its textual output by simply referring to region visual tokens, alleviating the need for the model architecture 200 to meticulously regress object coordinates. In some embodiments, the tokenizer of models can also encode user-specified region inputs (i.e., bounding boxes) into region visual tokens, which are directly inserted into user instructions to initiate referential dialogue.

[0032] Specifically, the solution comprises extracting a feature map from a target image. The solution further comprises determining, based on the feature map, a plurality of image regions of the target image within which respective objects are located. The solution further comprises determining, from the feature map, a plurality of region visual tokens for the plurality of image regions, respectively. The solution further comprises providing a model input to a machine learning model, to obtain a model output of the machine learning model. The model input at least comprises the feature map, the plurality of region visual tokens, and a user query for the target image. The solution further comprises generating a response to the user query based on the model output.

[0033] Compared to previous methods that augment language models for localization, the embodiments of the present disclosure circumvent the heavy computation of language models when handling high-resolution input by settling localization to the image tokenization process. That is, the embodiments of the present disclosure may use high-resolution images for tokenizer input and downsampled image tokens for language model input, which saves computation without sacrificing localization accuracy. Besides, unlike methods adopting separate designs for modeling grounding outputs and referring inputs, the embodiments of the present disclosure seamlessly unify the two capabilities with the use of region visual tokens.

[0034] From the data perspective, to improve the localized understanding of the embodiments of the present disclosure, an extensive collection of datasets with region-level annotations is adopted for training, which encompasses a range of region semantics from objects and relationships to detailed region descriptions. In addition, to remedy the lack of long-form grounded data, a visually grounded chat dataset is constructed for instruction  finetuning, which is the first grounded chat dataset constructed with both visual and textual prompts, leveraging the powerful content generation model for data generation.

[0035] The comprehensive experiments demonstrate the superiority of the design, with results showing that it outperforms the comparable language models on established referring and grounding benchmarks.

[0036] The example embodiments of the present disclosure will be descried with reference to the accompanying drawings, which are capable of understanding user-defined region inputs and generating visually grounded outputs. Specifically, first an example architecture will be described. Then how to format region input and output will be introduced. Finally, the learning pipelines will be detailed.

[0037] Reference is now made to FIG. 2, which illustrates a schematic diagram of an example model architecture 200 for visual perception in accordance with some embodiments of the present disclosure. The model architecture 200 may be implemented at the computer system 110 of FIG. 1, to implement the visual perception task.

[0038] The model architecture 200 may take the target image 112 and the user query 114 as the model input and generate the response 116 as the model output for the visual perception task. As shown in FIG. 2, the target image 112 may include objects, such as a man, a woman, a cat and flowers. The user query 114 may include a question, for example, “what is in the region? ” . Then, the response 116 may be provided as “A cat is walking towards flowers” .

[0039] The model architecture 200 may comprise an image encoder model 210 for scene-level image tokenization. The image encoder model 210 may take the target image 112 as input and extract a feature map 212 from the target image 112. In some examples, the image encoder model 210 may split the feature map 212 into a plurality of image tokens, which is provided as a part of the model input to the machine learning model 105.

[0040] In some embodiments, the image encoder model 210 may be constructed as any suitable model structure that is suitable for visual feature extraction. As some examples, the image encoder model 210 may be based on convolutional neural network, vision transformer, contrastive language-image pre-training model, or the like. As a specific, the image encoder model 210 may be based on a pretrained DINOv2 model. In some cases,  considering that the use of higher-resolution images may lead to extended sequences of visual input for the language model, to save computations, a plurality of neighbor patch tokens may be merged into a single image token.

[0041] The model architecture 200 may further comprise a region proposer model 220 for discovering ROIs in the target image 112. The region proposer model 220 may take the feature map 212 of the image encoder model 210 as input and determine a plurality of image regions 222 of the target image 112 based on the feature map output by the image encoder model 210. There may be objects located within the target image 112. Each of the image regions 222 proposed by the region proposer model 220 may be a region that is believed to present an object therein. For example, as shown in FIG. 2, an object, e.g., a man, is located in one image region 222. Another object, e.g., a woman, is located in another image region 222.

[0042] In some embodiments, to achieve localized understanding of the target image 112, the region proposer model 220 is integrated into the image tokenization process. In some embodiments, feature maps are extracted from the last layers of the image encoder model 210 and rescaled to build a hierarchical feature pyramid as the input for the region proposer model 220. The region proposer generates 300 region proposals for each image, which are then refined through NMS and objectness scores before being input to the region encoder model 230. In some examples, the region proposer model 220 may be implemented as a class-agnostic detector head using the Deformable DETR (DDETR) transformer. The original classification head of DDETR is replaced by a binary classifier to assign scores to region proposals based on their localization quality. In other embodiments, the region proposer model 220 may be constructed in any other suitable model structures that can be implemented as object detection in images.

[0043] The model architecture 200 may further comprise a region encoder model 230 for region-level image tokenization. The region encoder model 230 may take the feature map 212 and the plurality of image regions 222 as input and determine, from the feature map 212, a plurality of region visual tokens 232 (represented as <Region1>, <Region2>…<Regionn>) for the plurality of image regions 222, respectively.

[0044] The region encoder model 230 may translate region proposals (e.g., bounding boxes) , coming from both the user query 114 and the region proposer model 220, into region  visual tokens 232. In some embodiments, the feature maps may be selected from the last three layers of the image encoder model 210 to create a hierarchical feature pyramid. A multiscale ROIAlign module as implemented is utilized to crop and fuse these hierarchical features into unified region visual tokens. Compared with alternative ways to represent regional inputs, such as numerical representation of positions and discrete location tokens, the region visual tokens 232 offer distinct benefits as they are semantically aligned with the underlying the image regions 222, which render them more intuitive for the machine learning model 105 to comprehend.

[0045] The model architecture 200 may further comprise a language tokenizer 240. The language tokenizer 240 may take the user query 114 as input and convert it into a plurality of text tokens 242, which may be provided to the machine learning model 105. The language tokenizer 240 may be implemented as a model that can tokenize text information.

[0046] The model architecture 200 may further comprise a machine learning model 105 for modeling multimodal input and output. The feature map 212 (corresponding to image tokens) from the image encoder model 210, the plurality of region visual tokens 232 and the user query 114 (i.e., the text tokens 242) may be provided as model input to the machine learning model 105. The machine learning model 105 may output the response 116. In some embodiments, the feature map 212, the region visual tokens 232, and the text tokens 242 may be combined to generate a prompt input to the machine learning model 105. In some embodiments, the prompt input may be generate using certain predefined prompt template, as will be discussed below.

[0047] In some embodiments, the machine learning model 105 may be constructed based on a language model, e.g., a large-scale language model. The machine learning model 105 is considered as a generative model which can process input information, understand the semantics in the input information, and generate corresponding output as a response.

[0048] The machine learning model 105 may be constructed based on any suitable model structure that can perform semantic understanding and content generation. Besides, the image tokens and region visual tokens are projected into the feature space of the machine learning model 105 by using a multilayer perceptron (MLP) layer (s) .

[0049] It would be appreciated that depending on the requirements, the model architecture 200 could contain fewer or additional modules in certain embodiments.

[0050] In summary, the model architecture 200 tokenizes the target image input into both global image tokens and local region visual tokens. The region visual tokens, which are naturally grounded to the target image, serve as the building blocks for localized understanding in the model architecture 200. By integrating region visual tokens into user instructions and model responses, the referential dialogue and grounded chat abilities are unlocked.

[0051] Beyond textual only instructions and responses, the embodiments of the present disclosure may offer the flexibility to accept user-specified regions as input (referring) and generate visually grounded answers (grounding) . Specifically, although different in task formulations, both referring and grounding are unified into one format with the use of region visual tokens. In other words, an image referring task and an image grounding task may be supported.

[0052] In some embodiments, the user query 114 may be converted into a plurality of text tokens by the language tokenizer 240. The user query 114 may be corresponding to an image referring task and that the user query 114 may indicate at least one image region of the target image 112. Then the model input may be modified.

[0053] For example, as shown in FIG. 2, the user query 114 may include “What is in the region? ” , which is converted into text tokens, such as “What” , “is” , “in” , “region” , etc. The user query 114 may indicate an image region, i.e., region 3 of the target image 112. The region 3 may include a cat corresponding to a region visual token, i.e., <Region3>. One of those text tokens may be changed, e.g., the text token “region” is modified to the region visual token <Region3>. Then the modified text tokens may be provided to the machine learning model 105.

[0054] In some embodiments, the user query 114 may be corresponding to an image grounding task, the model input may further comprise a token indicating the image grounding task. For example, a special token [grounding] may be used to inform the machine learning model 105 to generate a grounded response.

[0055] Referring to FIG. 3, which illustrates a schematic diagram of an example 300 of model input and model output in accordance with some embodiments of the present disclosure. As shown in FIG. 3, the user query 114 may include “Can you describe this image in detail? ” . The user query 114 doesn’t indicate any image region of the target image  112 specified by the user. Then the token [grounding] may be provided to the machine learning model 105 along with the user query 114.

[0056] In the tokenization process, each region visual token is inherently anchored to a concrete location in the target image, corresponding to its region proposal. This connection allows the language model to ground its text output to particular regions in the image by simply referring to the associated region visual tokens. However, as region visual tokens are continuous embeddings, they cannot be directly integrated into the codebook of the language model and referenced in the text output.

[0057] In some embodiments, to bridge the gap, the model input may further comprise respective region indices for identifying the plurality of region visual tokens, respectively. A region index associated with a region visual token may be placed adjacent to the region visual token in the model input. For example, a set of region indices “<r1>, <r2>…<rn>” is introduced to register region visual tokens. As illustrated below, any region in the model output may be referred by addressing the region indices.

[0058] In some embodiments, the model output may comprise a text sequence with at least one region index. The machine learning model 105 may generate the response 116 by referring to at least one image region corresponding to the at least one region index in the target image 112.

[0059] For example, as shown in FIG. 2, the user query 114 may include “What is in the region? ” . The underlined word “region” may represent an image region (hereinafter also referred to as region 3) . The response 116 may include “A cat is walking towards flowers” , where the words “a cat” represent an object in the region 3 that is understood by the machine learning model 105.

[0060] Regarding to the model input and model output, the following will provide two specific examples.

[0061] As an example, if the user query 114 indicates an image referring task, for a region pointed out by the user, it may be encoded into a region visual token and assigning a region index to it. User-specified regions may be incorporated into the instructions by inserting corresponding region visual tokens. A simple example of referential dialogue is given below, where <region3>comes from user-specified region input.

[0062] As shown in FIG. 2, the model input may be constructed using a prompt input as “User: Here is an image with region crops from it. Image: <image>. Regions: <r1><Region>, <r2><Region>... <rn><Region>. What is in the <r3><region3>? ” , where <image> and <Region> stand for placeholders of image tokens and region visual tokens, which are replaced by corresponding visual tokens before being fed into the machine learning model 105. The corresponding model output may include “A cat is walking towards flowers” .

[0063] As another example, if the user query 114 indicates an image grounding task, a special token [grounding] may be used to inform the machine learning model 105 to generate a grounded response. As shown in FIG. 3, the model input may include “User: Here is an image with region crops from it. Image: <image>. Regions: <r1><Region>, <r2><Region>, ... <rn><Region>. [grounding] Can you describe this image in detail? ” , where <image> and <Region> stand for placeholders of image tokens and region visual tokens, which are replaced by corresponding visual tokens before being fed into the machine learning model 105. The corresponding model output may include “In this image, we see a man<roi><Region1>< / roi> is talking with a woman<roi><Region2>< / roi> standing beside him. There are flowers<roi><Region4>< / roi> nearby. A cat<roi><Region3>< / roi> is walking towards the flowers” , where  and  mark the start and end of the grounded phrase, <roi> and < / roi> are used to enclose the referenced regions.

[0064] The above has described the model architecture 200 and the model input and model output formatting. The training of model architecture 200 is partitioned into three stages: a first training stage for detection pretraining for localization ability, a second training stage for alignment pretraining for image-level and region-level vision-language alignment, and an instruction finetuning stage for enhanced conversation capability. In some embodiments, the image encoder model 210 and the region proposer model 220 may be trained jointly with a first loss function at least based on a loss between a plurality of ground-truth image regions in a sample image and a plurality of predicted image regions determined for the sample image by the region proposer model 220.

[0065] The first training stage only involves the image encoder model 210 and the region proposer model 220, which collectively constitute a DDETR-like detector. The image encoder model 210 may be kept frozen during this training stage. In some embodiments, to endow the region proposer model 220 with localization capability, an extensive collection  of detection datasets, including COCO, Objects365, OpenImages, and V3Det, is utilized for large-scale pretraining. Notably, category information is omitted from the training process, with a primary focus on box supervision.

[0066] In some embodiments, considering traditional detection data are typically limited to object-level annotations, the training is complemented with a two million subset of SA1B data filtered by GLEE. Original mask annotations of SA1B are transformed into bounding boxes for consistency. The inclusion of this enriched dataset encourages the region proposer model 220 to produce region proposals across a wide spectrum of granularities, encompassing not only object instances but also their constituent parts and various background stuff.

[0067] In some embodiments, the region encoder model 230 and at least a part of the machine learning model 105 may be trained jointly in the second training stage after the first training stage, and the image encoder model 210 and the region proposer model 220 may be fixed during the second training stage.

[0068] To align vision and language feature space of models, the model architecture 200 are pretrained on a wide range of vision-language tasks. For example, for image-level alignment, a dataset with paired images and captioning is leveraged for detailed image captioning. For region-level alignment, a dataset with the region labeled is engaged for referring expression comprehension (REC) , a dataset with paired regions and captioning is used for region captioning, and a dataset with grounded regions is used for grounded caption generation.

[0069] In some embodiments, the machine learning model 105 may be fine-tuned in a second training stage after the second training stage. The image encoder model 210, the region proposer model 220, and the region encoder model 230 may be fixed during the second training stage. For example, to maintain training efficiency, finetuning efforts are focused on the MLP projection layer and the region encoder model 230, while other modules are kept frozen throughout the training.

[0070] Based on alignment pretraining, the training data may be refined to focus exclusively on high-quality datasets and proceed to unfreeze the language model for finetuning purposes. For example, at this stage, LLaVA Instruct and ShareGPT-4V may be incorporated to improve the conversational and instruction-following capabilities.  Furthermore, a grounded chat dataset may be used to facilitate synergy of chatting and grounding abilities.

[0071] In summary, a major difference between the training of the model architecture 200 and current multimodal language models is the integration of dedicated detection pretraining, which endows the model architecture 200 with robust and precise localization ability. Thanks to the decoupled architecture of location and understanding, the need is circumvented to involve the language model during detection pretraining. Such a strategic design allows the model architecture 200 to benefit from pretraining on millions of bounding box annotations -a task that would be computationally prohibitive for classic multimodal language models.

[0072] Last but not least, according to the embodiments of the present disclosure, the model architecture 200 is introduced with grounded and fine-grained visual perception ability. Beyond holistic image understanding, the model architecture 200 is adept at region-level tasks such as region captioning and visual grounding. Such capabilities are built upon a localized visual tokenization mechanism, where an image input is decomposed into regions of interest and subsequently encoded into region visual tokens. By integrating region visual tokens into user instructions and model responses, the model architecture 200 is seamlessly enabled to understand user-specified region inputs and ground its textual output to images. Besides, to enhance the grounded chat ability of the model architecture 200, a visually grounded instruction dataset may be curated by leveraging the powerful content generation model and visual prompting techniques. Compared with multimodal language models that rely on the language model or external module for localization, the model architecture 200 consistently demonstrates superior performances in standard referring and grounding benchmarks, highlighting the advantages.

[0073] FIG. 4 illustrates a flowchart of a process 400 for visual perception in accordance with some embodiments of the present disclosure. The process 400 may be implemented at the computer system 110 of FIG. 1.

[0074] At block 410, the computer system 110 extracts a feature map from a target image.

[0075] At block 420, the computer system 110 determines, based on the feature map, a plurality of image regions of the target image within which respective objects are located.

[0076] At block 430, the computer system 110 determines, from the feature map, a plurality of region visual tokens for the plurality of image regions, respectively.

[0077] At block 440, the computer system 110 provides a model input to a machine learning model, to obtain a model output of the machine learning model. The model input at least comprises the feature map, the plurality of region visual tokens, and a user query for the target image.

[0078] At block 450, the computer system 110 generates a response to the user query based on the model output.

[0079] In some embodiments, the model input further comprises respective region indices for identifying the plurality of region visual tokens, respectively. A region index associated with a region visual token is placed adjacent to the region visual token in the model input.

[0080] In some embodiments, the model output comprises a text sequence with at least one region index, and the computer system 110 generates the response by referring to at least one image region corresponding to the at least one region index in the target image.

[0081] In some embodiments, the computer system 110, in response to determining that the user query is corresponding to an image grounding task, determines the model input to further comprise a token indicating the image grounding task.

[0082] In some embodiments, the user query is converted into a plurality of text tokens to be comprised in the model input. The computer system 110, in response to determining that the user query is corresponding to an image referring task and that the user query indicates at least one image region of the target image specified by a user, modifies the model input by replacing at least one of the plurality of text tokens corresponding to the at least one image region with at least one region visual token corresponding to the at least one image region. The modified model input is provided to the machine learning model.

[0083] In some embodiments, the feature map is extracted using an image encoder model, and the plurality of image regions are determined using a region proposer model. The image encoder model and the region proposer model are trained jointly with a first loss function at least based on a loss between a plurality of ground-truth image regions in a sample image and a plurality of predicted image regions determined for the sample image by the region  proposer model.

[0084] In some embodiments, the image encoder model and the region proposer model are trained in a first training stage, and the plurality of region visual tokens are determined using a region encoder model. The region encoder model and at least a part of the machine learning model are trained jointly in a second training stage after the first training stage, and the image encoder model and the region proposer model are fixed during the second training stage.

[0085] In some embodiments, the machine learning model is fine-tuned in a second training stage after the second training stage. The image encoder model, the region proposer model, and the region encoder model are fixed during the second training stage.

[0086] In some embodiments, the machine learning model is constructed based on a language model.

[0087] FIG. 5 shows a block diagram of an apparatus 500 for visual perception in accordance with some embodiments of the present disclosure. The apparatus 500 may be implemented, for example, or included at the computer system 110 of FIG. 1. Various modules / components in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0088] As shown, the apparatus 500 includes a feature map extracting module 510 configured to extract a feature map from a target image.

[0089] The apparatus 500 further includes an image region determining module 520 configured to determine, based on the feature map, a plurality of image regions of the target image within which respective objects are located.

[0090] The apparatus 500 further includes a region visual token determining module 530 configured to determine, from the feature map, a plurality of region visual tokens for the plurality of image regions, respectively.

[0091] The apparatus 500 further includes a model input providing module 540 configured to provide a model input to a machine learning model, to obtain a model output of the machine learning model. The model input at least comprises the feature map, the plurality of region visual tokens, and a user query for the target image.

[0092] The apparatus 500 further includes a response generating module 550 configured to generate a response to the user query based on the model output.

[0093] The apparatus 500 may further comprises corresponding modules that are configured to perform the operations of the process 400 and other embodiments as described herein.

[0094] FIG. 6 illustrates a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure can be implemented. It would be appreciated that the electronic device 600 shown in FIG. 6 is only an example and should not constitute any restriction on the function and scope of the embodiments described herein. The electronic device 600 may be used, for example, to implement the computer system 110 of FIG. 1. The electronic device 600 may also be used to implement the apparatus 500 of FIG. 5.

[0095] As shown in FIG. 6, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be an actual or virtual processor and can execute various processes according to the programs stored in the memory 620. In a multiprocessor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 600.

[0096] The electronic device 600 typically includes a variety of computer storage medium. Such medium may be any available medium that is accessible to the electronic device 600, including but not limited to volatile and non-volatile medium, removable and non-removable medium. The memory 620 may be volatile memory (for example, a register, cache, a random access memory (RAM) ) , a non-volatile memory (for example, a read-only memory (ROM) , an electrically erasable programmable read-only memory (EEPROM) , a flash memory) or any combination thereof. The storage device 630 may be any removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which can be used to store information and / or data (such as training data for training) and can be accessed within the electronic device 600.

[0097] The electronic device 600 may further include additional removable / non- removable, volatile / non-volatile, transitory / non-transitory storage medium. Although not shown in FIG. 6, a disk driver for reading from or writing to a removable, non-volatile disk (such as a “floppy disk” ) , and an optical disk driver for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each driver may be connected to the bus (not shown) by one or more data medium interfaces. The memory 620 may include a computer program product 625, which has one or more program modules configured to perform various methods or acts of various embodiments of the present disclosure.

[0098] The communication unit 640 communicates with a further computing device through the communication medium. In addition, functions of components in the electronic device 600 may be implemented by a single computing cluster or multiple computing machines, which can communicate through a communication connection. Therefore, the electronic device 600 may be operated in a networking environment using a logical connection with one or more other servers, a network personal computer (PC) , or another network node.

[0099] The input device 650 may be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as required. The external device, such as a storage device, a display device, etc., communicate with one or more devices that enable users to interact with the electronic device 600, or communicate with any device (for example, a network card, a modem, etc. ) that makes the electronic device 600 communicate with one or more other computing devices. Such communication may be executed via an input / output (I / O) interface (not shown) .

[0100] According to example implementation of the present disclosure, a computer-readable storage medium is provided, on which a computer-executable instruction or computer program is stored, where the computer-executable instructions or the computer program is executed by the processor to implement the method described above. According to example implementation of the present disclosure, a computer program product is also provided. The computer program product is physically stored on a non-transient computer-readable medium and includes computer-executable instructions, which are executed by the processor to implement the method described above.

[0101] Various aspects of the present disclosure are described herein with reference to the flow chart and / or the block diagram of the method, the device, the equipment and the computer program product implemented in accordance with the present disclosure. It would be appreciated that each block of the flowchart and / or the block diagram and the combination of each block in the flowchart and / or the block diagram may be implemented by computer-readable program instructions.

[0102] These computer-readable program instructions may be provided to the processing units of general-purpose computers, special computers or other programmable data processing devices to produce a machine that generates a device to implement the functions / acts specified in one or more blocks in the flow chart and / or the block diagram when these instructions are executed through the processing units of the computer or other programmable data processing devices. These computer-readable program instructions may also be stored in a computer-readable storage medium. These instructions enable a computer, a programmable data processing device and / or other devices to work in a specific way. Therefore, the computer-readable medium containing the instructions includes a product, which includes instructions to implement various aspects of the functions / acts specified in one or more blocks in the flowchart and / or the block diagram.

[0103] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other devices, so that a series of operational steps can be performed on a computer, other programmable data processing apparatus, or other devices, to generate a computer-implemented process, such that the instructions which execute on a computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks in the flowchart and / or the block diagram.

[0104] The flowchart and the block diagram in the drawings show the possible architecture, functions and operations of the system, the method and the computer program product implemented in accordance with the present disclosure. In this regard, each block in the flowchart or the block diagram may represent a part of a module, a program segment or instructions, which contains one or more executable instructions for implementing the specified logic function. In some alternative implementations, the functions marked in the block may also occur in a different order from those marked in the drawings. For example, two consecutive blocks may actually be executed in parallel, and sometimes can also be  executed in a reverse order, depending on the function involved. It should also be noted that each block in the block diagram and / or the flowchart, and combinations of blocks in the block diagram and / or the flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by the combination of dedicated hardware and computer instructions.

[0105] Each implementation of the present disclosure has been described above. The above description is example, not exhaustive, and is not limited to the disclosed implementations. Without departing from the scope and spirit of the described implementations, many modifications and changes are obvious to ordinary skill in the art. The selection of terms used in this article aims to best explain the principles, practical application or improvement of technology in the market of each implementation, or to enable other ordinary skill in the art to understand the various embodiments disclosed herein.

Claims

1.A method for visual perception, comprising:extracting a feature map from a target image;determining, based on the feature map, a plurality of image regions of the target image within which respective objects are located;determining, from the feature map, a plurality of region visual tokens for the plurality of image regions, respectively;providing a model input to a machine learning model, to obtain a model output of the machine learning model, the model input at least comprising the feature map, the plurality of region visual tokens, and a user query for the target image; andgenerating a response to the user query based on the model output.2.The method of claim 1, wherein the model input further comprises respective region indices for identifying the plurality of region visual tokens, respectively, a region index associated with a region visual token being placed adjacent to the region visual token in the model input.3.The method of claim 2, wherein the model output comprises a text sequence with at least one region index, and wherein generating a response to the user query comprises:generating the response by referring to at least one image region corresponding to the at least one region index in the target image.4.The method of claim 1, further comprising:in response to determining that the user query is corresponding to an image grounding task, determining the model input to further comprise a token indicating the image grounding task.5.The method of claim 1, wherein the user query is converted into a plurality of text tokens to be comprised in the model input, the method further comprising:in response to determining that the user query is corresponding to an image referring task and that the user query indicates at least one image region of the target image specified by a user, modifying the model input by replacing at least one of the plurality of text tokens corresponding to the at least one image region with at least one region visual token corresponding to the at least one image region; andwherein the modified model input is provided to the machine learning model.6.The method of claim 1, wherein the feature map is extracted using an image encoder model, and the plurality of image regions are determined using a region proposer model, andwherein the image encoder model and the region proposer model are trained jointly with a first loss function at least based on a loss between a plurality of ground-truth image regions in a sample image and a plurality of predicted image regions determined for the sample image by the region proposer model.7.The method of claim 6, wherein the image encoder model and the region proposer model are trained in a first training stage, and the plurality of region visual tokens are determined using a region encoder model, andwherein the region encoder model and at least a part of the machine learning model are trained jointly in a second training stage after the first training stage, and the image encoder model and the region proposer model are fixed during the second training stage.8.The method of claim 7, wherein the machine learning model is fine-tuned in a second training stage after the second training stage, and wherein the image encoder model, the region proposer model, and the region encoder model are fixed during the second training stage.9.The method of claim 1, wherein the machine learning model is constructed based on a language model.10.An electronic device, comprising:at least one processing unit; andat least one memory coupled to the at least one processing unit and storing instructions executable by the at least one processing unit, the instructions, upon execution by the at least one processing unit, causing the electronic device to perform actions comprising:extracting a feature map from a target image;determining, based on the feature map, a plurality of image regions of the target image within which respective objects are located;determining, from the feature map, a plurality of region visual tokens for the plurality of image regions, respectively;providing a model input to a machine learning model, to obtain a model output of the machine learning model, the model input at least comprising the feature map, the plurality of region visual tokens, and a user query for the target image; andgenerating a response to the user query based on the model output.11.A computer-readable storage medium, having a computer program stored thereon which, upon execution by an electronic device, causes the device to perform a method according to any of claims 1 to 9.12.A computer program product being tangibly stored on a computer-readable medium and comprising computer-executable instructions which, when executed by a device, cause the device to implement a method according to any of claims 1 to 9.

Citation Information

Patent Citations

  • Image processing method and device and computer storage medium

    CN112528905A

  • Image description generation method and device, equipment, medium and product

    CN114627353A

  • Image classification method and system, medium and electronic equipment

    CN116843961A

  • Multi-modal knowledge graph construction method based on double-flow coding and comparative learning

    CN117787400A

  • Mark recognition system and method for identification of one or more marks on an object

    US20020048403A1