Hybrid model for vision system

By introducing mobile convolution blocks and transformer blocks of hybrid models into computer vision systems, the problem of poor performance on small data sets is solved, and more efficient feature extraction and generation is achieved, computing resource consumption is reduced, and feature resolution and data scalability are improved.

CN120259680APending Publication Date: 2025-07-04FACE CUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411821993.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-02
Filing Date
2024-12-11
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing computer vision systems have poor performance on small data sets, high computing resource consumption, and difficulty in effectively extracting and generating feature maps.

Method used

Using a hybrid model, including mobile convolutional blocks (MBConv) and transformer blocks (TFB), the image data is downsampled and feature extraction is performed through a three-level network structure to reduce the computing resource requirements, and at the same time, the transformer blocks are used to generate feature maps.

Benefits of technology

Improves performance on small datasets, reduces computing resource requirements, and improves feature resolution and data scalability, enhancing zero-sample evaluation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259680A_ABST
    Figure CN120259680A_ABST
Patent Text Reader

Abstract

Methods and systems for generating feature maps from images are disclosed. The vision system includes a vision model for processing the image according to a neural network to generate the feature map. The visual model comprises: a first convolution block, which is used for down-sampling an image data set to obtain first-level convolution data; a second convolution block for downsampling the first level of convolution data to obtain second level of convolution data, where one or both of the first convolution block and the second convolution block is a moving convolution block (MBConv) comprising: a first Gaussian error linear unit (GELU) layer, a depth-by-depth convolution (DWConv) layer, and a resized convolution layer; and a transformer block (TFB) that generates the feature map from the second level convolution data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments described herein generally relate to methods and systems for training and implementing a vision model having a hybrid model for identifying features in an image. More specifically, the embodiments described herein relate to a hybrid model having (multiple) mobile convolutional blocks and (multiple) transformer blocks for extracting features from an input image and / or generating (multiple) feature maps based on the input image. Background Art

[0002] Computer vision systems have been developed using machine learning models implemented by (multiple) neural networks trained to identify (multiple) objects in an image. The macroscopic architecture of a neural network typically includes multiple computational blocks that sequentially process an image input into (multiple) feature maps that can be identified by subsequent operations to generate human-readable material (e.g., text, image, audio) based on the feature maps. Other computer systems or models can utilize (multiple) feature maps as input to extract or generate useful information about the image input for further processing.

[0003] A neural network can be developed to have an architecture that provides a framework for processing input image data. At a macroscopic level, the network architecture can include a set of blocks for processing the input image data and selectively weighting and extracting features from the input data of a previous layer and outputting the extracted data to the next layer. A labeled dataset (e.g., a database having image-text pairs where the image and its corresponding label or description of the image content are specified) can be used to train the selection process of the neural network. In the microscopic design of the network, each block can include more layers of nodes that can include more blocks for processing data. To characterize the performance of one or more architecture designs, the system can be benchmarked in terms of accuracy (e.g., error rate) against a labeled dataset (or a subset thereof), speed, and / or computational resources required for training and / or completing the encoding process for generating an output. Summary of the Invention

[0004] The embodiments described herein generally relate to methods and systems for training and implementing a vision model having a hybrid model for identifying features in an image. More specifically, the embodiments described herein relate to a hybrid model having (multiple) mobile convolutional blocks and (multiple) transformer blocks for extracting features from an input image and / or generating (multiple) feature maps based on the input image.

[0005] It should be understood that by adopting the macro-architecture and micro-architecture of the present disclosure, the vision model of the system can exhibit better data and model scalability as well as feature resolution than alternative models (e.g., vanilla vision transformers (ViT), convolutional and self-attention (CoAtNet), etc.). For example, scalability may be related to the performance of the vision model on smaller datasets rather than larger datasets. In an embodiment, the vision model includes a convolutional stem and a three-stage network for extracting features from the (multiple) images and / or generating feature maps based on the provided (multiple) images. It should be understood that in some embodiments, by adopting the three-stage network, the vision model can provide a feature map with an output stride of 16 (i.e., downsampling by a factor of 16 from the image data before being processed by the vision model).

[0006] It should be further understood that by successively downsampling the image data (or the stemmed image data from the convolutional stem block) with two-stage mobile convolutional blocks (MBConv), the architecture of the present disclosure has a smaller number of MBConv blocks than alternative designs, such that the computational resources required by the vision model can be lower than that of a vision model with a larger number of MBConv blocks. In an embodiment, the transformer block (TFB) is arranged as the last block to obtain the feature map to achieve superior data and model scalability.

[0007] It should be understood that the vision model of the present disclosure can be used as an image encoder of a vision-language model (VLM) to extract context information from image data and output machine-readable data (e.g., a feature map representing the context information) and / or human-readable text based on the input image data. The vision model part of the VLM can be obtained from a trained neural network having the architecture described in the present disclosure. For example, the neural network can be trained with a labeled dataset, a pre-trained model (e.g., contrastive language-image pre-training (CLIP)), or other models with text / image embeddings.

[0008] It should be further understood that by adopting this architecture, the neural network can have fewer parameters (e.g., compared to the number of parameters in a neural network using an alternative architecture), thereby allowing additional transformer blocks to be stacked into a deeper architecture and / or reducing the computational power required to train the network.

[0009] Obtaining the vision model of the vision-language model (VLM) may include training a neural network to minimize the following loss function:

[0010]

[0011] In this loss function, a batch of N image-text pairs can be {(I1,T1),...,(IN , T N )}, where I i and T i represent the image and text of the i-th pair. The goal is to embed the image in each pair into x i and the text into y i for alignment, where and As discussed herein, a vision model and / or a text model can be trained to minimize the above loss function.

[0012] In an embodiment, by adopting a simple convolutional backbone (e.g., having two identical 3-by-3 convolutional layers) and a three-stage network hybrid architecture, the vision model can be trained using fewer computational resources and / or trained faster compared to alternative architectures, while reducing the number of parameters, e.g., having half, one-fourth, or one-tenth of the parameters. Compared to a model trained using an alternative architecture, the vision model can further have a higher zero-shot evaluation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings illustrate various embodiments of the systems and methods of the present disclosure, as well as embodiments of various other aspects. Any ordinary person skilled in the art will understand that the element boundaries shown in the figures (e.g., boxes, multiple sets of boxes, or other shapes) represent an example of the boundaries. In some examples, it may be possible that one element can be designed as multiple elements, or multiple elements can be designed as one element. In some examples, an element shown as an internal component of one element can be implemented as an external component of another element, and vice versa. The following drawings are described in a non-limiting and non-exhaustive manner. The components in the figures are not necessarily drawn to scale, but the emphasis is placed on illustrating the principles. In the following detailed description, the embodiments are described only by way of illustration, since various changes and modifications may be apparent to those skilled in the art from the following detailed description.

[0014] Figure 1 Shows a system that can implement a vision model according to one or more embodiments.

[0015] Figure 2 Shows the micro-level architecture of a convolutional backbone block according to an embodiment.

[0016] Figure 3 Shows the micro-level architecture of an MBConv block according to an embodiment.

[0017] Figure 4 Shows the micro-level architecture design of a TFB block according to an embodiment.

[0018] Figure 5Shows the structure of a text transformer for processing feature maps according to an embodiment.

[0019] Figure 6 Is a flowchart showing a method 600 for training a neural network according to one or more embodiments. Detailed Description

[0020] The embodiments described herein generally relate to methods and systems for training and implementing a vision model that has a hybrid model for identifying features in an image. More specifically, the embodiments described herein relate to a hybrid model having (multiple) mobile convolutional blocks and (multiple) transformer blocks for extracting features from an input image and / or generating (multiple) feature maps based on the input image.

[0021] In the following detailed description, specific embodiments of the present disclosure are described with reference to the accompanying drawings, which form a part of this specification. In this specification and the drawings, unless the context otherwise requires, the same reference numerals represent elements that can perform the same, similar, or equivalent functions. Additionally, unless otherwise stated, the description of each subsequent drawing can refer to features from one or more of the previous drawings to provide a clearer context and a more substantial explanation for the current example embodiment. Nevertheless, the example embodiments described in the detailed description, the drawings, and the claims are not intended to be limiting. Other embodiments can be utilized and other changes can be made without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the various aspects of the present disclosure, as generally described and shown in the drawings herein, can be arranged, substituted, combined, separated, and designed in a wide variety of configurations, all of which are explicitly contemplated herein.

[0022] It should be understood that the disclosed embodiments are merely examples of the present disclosure, which can be embodied in various forms. Well-known functions or constructions are not described in detail to avoid obscuring the present disclosure with unnecessary details. Thus, the specific structural and functional details disclosed herein should not be construed as restrictive, but merely as a basis for the claims and as a representative basis for teaching one of ordinary skill in the art to employ the present disclosure in any suitable detailed structure in various ways.

[0023] Additionally, the present disclosure may be described in terms of functional block components and various processing steps. It should be understood that such functional blocks can be implemented by any number of hardware and / or software components configured to perform the specified functions.

[0024] The scope of the present disclosure should be determined by the appended claims and their legal equivalents, rather than by the examples given herein. For example, the steps recited in any method claim may be performed in any order, and are not limited to the order presented in the claim. Additionally, no element is essential to the practice of the present disclosure unless explicitly described herein as "critical" or "necessary".

[0025] As mentioned herein, "data" or "data set" is a technical term and may refer to an organized collection of data stored and accessed electronically. In an example embodiment, the data may refer to a database, a data table, a part of a database or a data table, etc. It should be understood that the data may correspond to one or more database tables, where each column of the database table represents a specific variable or field, and each row of the database table corresponds to a given record of the data set. The data may list the values of each variable in the variables and / or the values of each record of the data. It should also be understood that the data set may also or alternatively refer to a set of related data and the organization manner of the related data.

[0026] As referred to herein, "neural network" is a technical term that includes layers of nodes, which include an input layer, one or more hidden layers, and an output layer. Each node is connected to another node and has an associated weight and threshold. If the output of any single node is higher than the specified threshold, then the node is activated and data is sent to the next layer of the network. Neural networks rely on training data to learn and improve their accuracy over time.

[0027] As referred to herein, "self-attention" is a technical term in machine learning used to retain relationship information from previous training rounds so that subsequent training rounds focus on more relevant data / variables in the data presented in the model.

[0028] As referred to herein, "convolutional layer" is a technical term in machine learning that represents any computational layer that receives an input and provides an output based on mathematical operations on the input.

[0029] As referred to herein, "transformer block" is a technical term in machine learning that includes one or more layers of self-attention mechanisms. For example, in each layer, each feature split can be contextually associated with other (unmasked) features within the context window via a parallel multi-head attention mechanism, thereby allowing amplification of the signal of key features and attenuation of less important features.

[0030] As referred to herein, "convolution" is a technical term in machine learning that includes performing one or more mathematical operations on image data (i.e., patches of image data, such as 1-by-1 pixel patches, 3-by-3, etc.) through filters for optimization.

[0031] As used herein, an "activation function" is a technical term in machine learning that calculates the output of a node based on the input to the node and the weights of the respective inputs.

[0032] As used herein, a "residual block" is a technical term in machine learning that refers to a sub-network having a certain number of stacked layers.

[0033] As used herein, a "model" can be an algorithm and / or program, hardware, or firmware, or any combination thereof. By generating a model (e.g., a vision model), the weights and biases of the nodes in a neural network can be provided to, for example, one or more algorithms and / or programs.

[0034] As used herein, "spatial interaction" is a technical term that refers to interactions that strongly depend on geometric relationships, such as the Euclidean distance and relative orientation between objects in an image.

[0035] Figure 1 A system 100 is shown that can implement a vision model 101 according to one or more embodiments.

[0036] As Figure 1 shown, the system 100 includes an input device 102 that is configured to provide one or more images 110 as input. The input is provided to a vision model 101 that is configured to provide an output (e.g., a feature map) based on the one or more images 110 provided by the input. The vision model 101 can be a model generated by training a neural network that has an architecture including, for example, a convolutional backbone layer 102, a first-stage mobile convolutional block 130 (MBConv), a second-stage MBConv 140, and a transformer block (TFB) 150.

[0037] Although illustrated as discrete components, the various components can be divided into additional components, combined into fewer components, or all excluded while still being within the scope of the disclosed subject matter. Those skilled in the art will understand that each function and / or operation of these components can be implemented individually and / or jointly by various hardware, software, firmware, or any combination thereof.

[0038] The input device 102 may refer to one or more embodiments of a computing environment, which may be or may include a computer, a processing device, a microprocessor, a microcontroller, a digital signal processor, or any combination thereof. The input device 102 may be one electronic device or a combination of various electronic devices, having one or more image and / or video capture components, i.e., a camera and / or a video recorder, a display screen with audio and / or video input / output, and supporting the provision and consumption of content related to a media platform. The various electronic devices may include, but are not limited to, security / monitoring devices, smartphones, tablet computers, laptop computers, desktop computers, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4s, and / or any other suitable electronic devices. Non-limiting examples of the input device 102 as a security device may include a video doorbell, a vehicle dashboard camera, a security camera (whether continuously activated or motion-activated), etc. Other non-limiting examples of the input device 102 may include a database, a local server, a cloud-based service, a virtual reality (VR) and / or an augmented reality (AR) server, etc. Further, any algorithm or program described, recited, or proposed herein may be executed by one or more processors hosted on the input device 102.

[0039] According to at least some of the embodiments disclosed and recited herein, the image 110 may refer to one or more digital images, e.g., each digital image having a size of H (height) × W (width) (e.g., 224 pixels × 224 pixels, 336 pixels × 336 pixels, etc.). In an embodiment, for open vocabulary detection, the image size may be 896 pixels × 896 pixels, and for a segmentation task, the image size may be 1344 pixels × 1344 pixels. In an embodiment, during training and / or validation, the size of the image in the visual model of the embodiment may be adjusted (e.g., enlarged or reduced) to provide an image size that is the same as or similar to that of other operations or a comparison visual model, e.g., fine-tuning the visual model at a larger image (or input) size.

[0040] The image 110 may be transmitted from the input device 102 via a wired or wireless network or otherwise delivered to a receiving component corresponding to the visual model 101. Such a network may be regarded as a medium provided as a two-way communication link between the media platform on which the visual model 101 is hosted and the input device 102. The network may include the Internet, a local area network (LAN), a wide area network (WAN), a local interconnect network (LIN), a local cloud, etc.

[0041] The vision model 101 is an implementation of a neural network. The vision model 101 can be, for example, an algorithm and / or program, hardware, or firmware, or any combination thereof, to classify, detect, isolate, and / or localize objects and / or segments of interest in the image 110. In an embodiment, the vision model 101 can be an algorithm obtained based on the result of training a neural network using one or more data sets. In an embodiment, the vision model 101 can include a neural network having an architecture as described in this disclosure. Training of the neural network can provide a machine learning model.

[0042] Non-limiting examples of such a model can include architectural blocks (e.g., MBConv, TFB, etc.) hosted on one or more servers (ranging from approximately hundreds to thousands in number), which can be hosted on a cloud-based infrastructure. Further, the vision model 101 can be implemented by a single or multiple classical computers, mobile devices, and facilitate transmission across a single or multiple connections or channels to one or more of the input devices 102 in the input device 102.

[0043] The output of the vision model 101 can be a feature map 160, which can be any type of digital representation of the input of the image 110 and can be provided to (multiple) subsequent operations to represent the features detected from the input of the image 100. This representation can be in a human-readable format or a machine-readable format for further processing. In an embodiment, subsequent operations of the vision model (e.g., Figure 1 the vision model 101 with a hybrid architecture as shown) can be a text transformer implemented by a contrastive language-image pre-training (CLIP) architecture.

[0044] Block 120 is a convolutional backbone block configured to process the image 110. The convolutional backbone can be configured to reduce the resolution of the image to reduce noise and / or reduce the computational amount required for subsequent operations. The vision model 101 processes the image 110 through block 120 to obtain a set of image data (e.g., backbone image data), which is then processed by block 130. It should be understood that the convolutional backbone block 120 can have convolutional layers that sequentially process the received data, as described with respect to Figure 2 what is described.

[0045] Blocks 130 and 140 can be MBConv blocks, each including one or more layers for downsampling the backbone image data and determining the features present in the backbone image data provided by block 120. In an embodiment, blocks 130 and 140 downsample the backbone image data in two stages. It should be understood that blocks 130 and 140 can each have layers to sequentially process the backbone image data, as described with respect to Figure 3 what is described.

[0046] In an embodiment, block 130 can be a first-stage MBConv that receives backbone image data to provide first-stage convolutional data. Block 140 can be a second-stage MBConv that receives the first-stage convolutional data to provide second-stage convolutional data.

[0047] Block 150 can be a TFB that is configured to provide a self-attention mechanism when detecting features to generate a feature map from the second-stage convolutional data from block 140, as further described with respect to Figure 4 Further described.

[0048] Regarding the scaling rules of the blocks, in an embodiment, block 120 has C channels, and the image size processed by these channels is H / 2 (for example, the height of 224 pixels divided by 2) by W / 2. It should be understood that the number of channels C can be 64, 128, 160, etc. Block 130 has 2C channels, and the image size processed by these channels is H / 4 by W / 4, such that block 130 can have the same number of channels C as block 120.

[0049] In an embodiment, the stride of the blocks gradually increases across blocks 120 to 150. For example, the stride of block 110 can be 2, the stride of block 120 can be 4, the stride of block 140 can be 8, and the stride of block 150 can be 16. It should be understood that "stride" can be a technical term in machine learning, representing a parameter that determines the movement of a kernel or filter across input data (such as an image). When performing a convolution operation, the stride determines how many units the filter moves in each step.

[0050] In an embodiment, the image size processed by the blocks gradually decreases across blocks 120 to 150. For example, block 120 can have an image size of H / 2 × W / 2, block 130 can have an image size of H / 4 × W / 4, block 140 can have an image size of H / 8 × W / 8, and block 150 can have an image size of H / 16 × W / 16.

[0051] The number of channels in the MBConv and TFB blocks can gradually increase (for example, blocks 130, 140, 150 have C, 2C, and 6C channels respectively). In an embodiment, the convolutional backbone block (such as block 120) and the first-stage MBConv block (such as block 130) can have the same number of channels C.

[0052] In some embodiments, the number of channels C, 2C, 6C can be 64, 128, 384; 128, 256, 768; 160, 320, 1024; etc. It should be understood that relative to the number of channels C, the number of channels may not be an exact multiple of C (for example, 2C is 2 times C, 6C is 6 times C, etc.).

[0053] In some embodiments, the convolutional backbone block (e.g., block 120) can have 2 blocks. In an embodiment, the convolutional backbone block can have the same number of blocks as the first-stage MBConv block. These two stages of MBConv blocks can have a gradually increasing number of blocks (e.g., block 120 has 2 blocks, while block 130 has 4 blocks). The number of blocks in the TFB block can be greater than that of other blocks, or have the largest number of blocks. In an embodiment, the TFB block can have 14 or 31 blocks (N b ).

[0054] It should be understood that by adopting the macro architecture as Figure 1 shown, the vision model 101 of the system 100 can exhibit better data and model scalability and feature resolution than alternative models (e.g., ViT, CoAtNet, etc.). In an embodiment, the vision model 101 uses a convolutional backbone (e.g., block 110) and a three-stage network (e.g., blocks 130, 140, and 150) to process an image to obtain a feature map based on the provided image. It should be understood that by adopting a three-stage network (e.g., blocks 130, 140, and 150), the vision model 101 provides a feature map with an output stride of 16 (i.e., a factor of 16 downsampling is performed from the image data before being processed by the vision model 101).

[0055] It should be understood that by adopting this architecture, the neural network can have fewer parameters (e.g., compared to the number of parameters in a neural network using an alternative architecture), thereby allowing additional transformer blocks to be stacked into a deeper architecture and / or reducing the computational power required to train the network.

[0056] It should be further understood that by successively downsampling the image data (or the backbone image data from block 120) through two stages of MBConv (blocks 130 and 140), the vision model 101 can contain a smaller number of blocks than an alternative design (e.g., block 130 has 2 blocks, and block 140 has 4 blocks), such that the computational resources required by the vision model 101 can be lower than that of a vision model with a larger number of MBConv blocks.

[0057] In an embodiment, the TFB block (e.g., 140) is arranged as the last block to obtain the feature map to achieve superior data and model scalability, e.g., by having two residual blocks, one with self-attention and the other with a feed-forward network, e.g., where the first linear layer with GeGLU can be replaced by a gated linear unit with a 2x expansion rate. In some embodiments, such a TFB block can have fewer parameters, e.g., less than 20%, 15%, 12%, 10%, or 5%. It should be understood that the performance can be characterized based on the required computational resources, zero-shot accuracy, etc.

[0058] Figure 2Shows the micro - level architecture of block 120 according to an embodiment. As Figure 2 shown, block 120 can be a convolutional backbone block. Block 120 includes a first convolutional layer 220 and a second convolutional layer 240. In an embodiment, the first convolutional layer 220 can have a filter or kernel size of, for example, 3×3 to reduce the resolution of the image and reduce the computational resources required for subsequent layer(s) and / or block(s). The second convolutional layer 240 can have a size of, for example, 3×3 to further reduce the resolution of the image and further reduce the computational resources required for subsequent layer(s) and / or block(s). It should be understood that the size of 3×3 can be the size of the filter or kernel, which can be a technical term such that the mathematical operation on the input image can be performed on an area of a first predetermined number of pixels multiplied by a second predetermined number of pixels (e.g., 3 pixels by 3 pixels labeled as 3×3).

[0059] Figure 3 Shows the micro - level architecture of block 130 or 140 according to an embodiment. It should be understood that MBConv blocks 130 and 140 can have the same micro - level architecture. However, the number of blocks, number of channels, stride, image size, etc. of blocks 130 and 140 may be different.

[0060] As Figure 3 shown, the MBConv block includes a layer normalization (LN) layer 310, a Gaussian error linear unit (GELU) convolutional layer 320, a GELU depth - wise (DWConv) convolutional layer 340, and a resize convolutional layer 360. It should be understood that GELU can include one or more activation functions to maintain the non - linearity in the image data. DWConv can include an activation function configured to capture the spatial interaction in the image data.

[0061] The LN layer 310 includes one or more activation functions for normalizing the data from a previous layer or block based on a reference value (such as the mean, standard deviation, etc. of the data).

[0062] The GELU convolutional layer 320 includes one or more activation functions to maintain the non - linearity in the image data, retaining more information in the input data compared to, for example, ReLU (a function used in an alternative model that sets all negative values to zero). In an embodiment, layer 320 can have a filter size of 1 pixel by 1 pixel to expand the channel size.

[0063] The GELU DWConv layer 340 can be a convolutional layer with an activation function that captures spatial interaction. In an embodiment, layer 340 can have a filter size greater than 1 pixel by 1 pixel (e.g., 3 pixels by 3 pixels).

[0064] The resize convolutional layer 360 can have a filter size of 1 pixel by 1 pixel to project the data back to the original channel size, e.g., Figure 1 the channel size C of block 130, 2C of block 140, etc.

[0065] It should be understood that the input data to block 130 or 140 can be transmitted through bypass 370 to be added to the output of layer 360, e.g., to maintain the history of prior training information. Thus, the sum of the output of layer 360 and the input to block 130 can be provided as the output of block 130. In an embodiment, the sum of the output of layer 360 and the input to block 140 can be provided as the output of block 140.

[0066] It should be understood that block 130 and / or 140 are configured to have an inverted bottleneck structure such that the channel size of the data input to the block is expanded at layer 320 and returned to the original channel size at layer 360. By adopting the inverted bottleneck structure, a large number of batch normalization layers (BN), squeeze-and-excitation layers (SE) can be omitted to simplify the architecture and / or reduce the computational resources required for training a neural network adopting the architecture of the vision model 101 described in the present disclosure. It should be understood that the role of the LN layer as the first layer of the block is similar to the pre-normalization layer of the TFB in (multiple) MBConv blocks (e.g., block 130, block 140, etc.). Thus, in an embodiment, although the BN and SE layers are omitted in the vision model 101, by using the LN layer (e.g., 310, 410 (as Figure 4 shown), etc.) as the first layer of part or all of the blocks (e.g., block 130, 140, and / or 150), the performance of the vision model 101 can be similar to that of the (multiple) MBConv-BN-SE blocks in other architectures while having a simpler structure, thus requiring less computational resources and / or time during training.

[0067] Figure 4 Shows the micro-level architecture design of block 150 according to one embodiment. As Figure 4 shown, block 150 is a TFB, which is configured to extract features in the data to generate a feature map as the output. Block 150 can be a GELU gated linear unit (GeGLU). Block 150 can include a self-attention residual block and a feed-forward network (FFN) residual block to extract more relevant features from the provided data. It should be understood that the residual block can be a computational block within block 150. In an embodiment, the self-attention residual block includes the LN layer 410 and the self-attention (SA) layer 420. The FFN includes the LN layer 440, the linear layer 450, the linear layer 460, the GELU layer 470, and the linear layer 480.

[0068] The LN layers 410 and 440 can each include one or more activation functions for normalizing data from a previous layer or block based on a reference value (such as the mean, standard deviation, etc. of the data). The SA layer 420 transforms the input data and focuses on the parameters and / or features in the input data that are more relevant to generating the feature map. The linear layers 450, 460, 480 can be configured to perform linear transformations or regressions to train the weights and biases in the network. The GELU layer 470 can include one or more activation functions to maintain non-linearity through this layer, thus retaining more information from the input data.

[0069] It should be understood that the output of the LN layer 440 can be the input to the linear layer 450, which processes the data to provide the output of the linear layer 450. The output of the LN layer 440 can be used as the input to the linear layer 460, which is subsequently processed by the GELU 470. The output of the linear layer 450 and the output of the GELU 470 can be multiplied to collect the weight and bias information in the linear layer 450 and the GELU layer 470. The product of the output of the linear layer 450 and the output of the GELU 470 is processed in the linear layer 480 to further adjust the weights and biases in network training. Then the input to the LN layer 440 is added to the output of the linear layer 480 to provide a feature map as the output of the block 150. It should be further understood that in some embodiments, by adopting a doubling expansion rate (e.g., a 2x expansion rate), the GeGLU block can improve the accuracy in the FFN residual block. In some embodiments, for example, compared to the number of parameters in a neural network using an alternative TFB architecture, the GeGLU block generates fewer parameters, thus allowing additional transformer blocks to be stacked towards a deeper architecture.

[0070] Figure 5 The structure of a VLM having an image encoder (such as a vision model) according to an embodiment is shown. The VLM can include a vision model and a text transformer as discussed above. That is, the collection of the vision model 101 (as Figure 1 shown) and the text transformer 500 can be referred to as the VLM, which converts / extracts the context information in the image data that is input to the vision model 101 into an output that describes the context information as machine-readable information (such as a feature map), human-readable text, etc. In some embodiments, the VLM can be configured to process the feature map provided as the output of the vision model, where the text transformer 500 is configured to align the features in the feature map with the text output describing the image input. In some embodiments, the text output with the lowest contrast loss can be selected as the text output corresponding to the feature map.

[0071] In an embodiment, Figure 5The VLM is used to train a vision model, such as a dataset of N image-text pairs of the CLIP framework, where the output I of the vision model i can be configured to align with the output T of its corresponding text transformer i , for example, by reducing the contrastive loss of the loss function to a predetermined contrastive loss threshold or below the predetermined contrastive loss threshold for a labeled dataset. Then, the labeled dataset (e.g., a dataset with image-text pairs from the CLIP framework) can be used to train the vision model, and iterative training is performed until the loss function reaches or is below the predetermined contrastive loss threshold or is minimized.

[0072] Figure 6 is a flowchart showing a method 600 for training a neural network according to one or more embodiments. The neural network can be developed to have an architecture that provides a framework for processing input image data. At a macro level, the network architecture can include a set of blocks as described above with respect to Figures 1 to 5 . These blocks are configured to process the input image data, selectively weight and extract features from the input data of the previous layer, and output the extracted data to the next layer.

[0073] As Figure 6 shown, method 600 starts with initializing a text encoder (e.g., Figure 5 the text transformer 500) 610. The text encoder can be initialized to train the neural network to be trained using the locked text tuning (LTT) method. LTT can use one of the pre-trained models, such as a previously trained vision model 101 (as Figure 1 shown), CLIP, etc.

[0074] For example, in an embodiment, a pre-trained model (e.g., OpenCLIP) and a dataset (e.g., DataComp-1B) can be used. The training can include short-term schedule training and / or long-term schedule training. The block size of the short-term schedule training can be 8000, 16000 images, etc. 32 A100 GPUs can be used for training. The number of iterations of the image data (i.e., the number of epochs) can be 1. The training time can be 1.8 days, 3.3 days, or 5.6 days. The batch size of the long-term schedule training can be 90000, using 184 A100 GPUs, 10 epochs, and 11 days of training time. In an embodiment, the model trained by short-term schedule training is configured to benchmark the model and conduct ablation studies. The model trained by long-term schedule training is configured to obtain a vision model. In an embodiment, the sample size of the images can be 200 million visible samples of a large image size (e.g., 336 pixels by 336 pixels). Then, method 600 continues to 620.

[0075] At 620, the values in the text encoder are frozen and used to train the image encoder, such that the feature correlations, weights, and biases from the text encoder can be used to train the image encoder. For example, the visual model can be iteratively trained by adjusting the parameters, weights, and / or biases of the various layers and / or blocks of the visual model. Then, method 600 proceeds to 630.

[0076] At 630, the image encoder is randomly initialized and trained by the frozen text encoder. For example, the text encoder can be a text encoder with a CLIP framework (e.g., CLIP A-v2, OpenCLIP, etc.), Figure 5 the text transformer 500, a large multimodal model, etc.

[0077] In an embodiment, the visual model can be iteratively trained to minimize a loss function:

[0078]

[0079] In this loss function, a batch of N image-text pairs can be {(I1, T1),..., (I N , T N )}, where I i and T i represent the image and text of the i-th pair. The goal is to align the image embedding x i and the text embedding y i in each pair, where and When the loss function is minimized, e.g., reaches a predetermined value, the training ends. In an embodiment, the resulting visual model can be an image encoder, which can be an implementation of a trained neural network having the architecture shown in system 100 with visual model 101 as shown in Figure 1 .

[0080] In some embodiments, benchmarking can be used to characterize, evaluate, and / or compare the performance of neural network architectures according to different architectural designs, to identify relevant features (or blocks) from different architectures and / or generate models with the identified relevant features, e.g., to generate improved models with relevant features or blocks. According to some embodiments, benchmarking the performance of a neural network can include measuring the performance of one or more upstream tasks and / or one or more downstream tasks to perform an overall evaluation of the network, visual model, and / or vision-language model (VLM). Upstream tasks can include evaluating classification ability and / or retrieval ability, e.g., for a labeled dataset (e.g., image-text pairs). For example, retrieval can include a visual model (e.g., the visual model in a VLM) retrieving an image from a dataset based on an input (such as a text input describing the image to be retrieved). Classification can include a visual model classifying an image in a dataset into one or more of a plurality of classes. In an embodiment, benchmarking can include evaluating the performance of downstream tasks, e.g., in open-vocabulary detection and segmentation, large multi-modal models (LMM), etc.

[0081] In some embodiments, validation can be used to validate one or more visual models. Validation can include validating a visual model based on certain model capabilities (such as classification ability, retrieval ability, open-vocabulary detection ability, and large multi-modal model performance, etc.), and / or further validating data scalability, model scalability, and feature resolution.

[0082] In an embodiment, validation can include validating a visual model based on the zero-shot accuracy of some or all of the capabilities (such as classification ability and retrieval ability). In an embodiment, a visual model can be validated by minimizing the contrastive loss and / or until the contrastive loss according to a loss function is reduced to a predetermined threshold.

[0083] For example, open-vocabulary detection and segmentation can include panoramic segmentation and semantic segmentation. In an embodiment, for open-vocabulary object detection, the F-ViT (i.e., a two-stage detector baseline built based on a frozen CLIP ViT) framework can be utilized. For open-vocabulary segmentation, the FC-CLIP (i.e., shared frozen convolutional CLIP) framework can be used and zero-shot evaluation can be performed on multiple segmentation datasets.

[0084] For example, an LMM can include a visual model (as part of a VLM) according to an embodiment as a visual encoder within the LMM. The visual model / VLM can provide image embeddings that are well-aligned with text, thus bridging the gap in visual understanding of the LLM (large language model). In an embodiment, LLaVA (Large Language and Vision Assistant)-1.5 can be used as an evaluation LMM framework.

[0085] In an embodiment, the F-ViT and / or FC-CLIP framework can be used for benchmarking. In an embodiment, the VLM can be a plug-and-freeze backbone of F-ViT, FC-CLIP, etc., for separately evaluating open-vocabulary detection and / or segmentation. Features can be extracted in a sliding window manner, where the window size is equal to or similar to the pre-trained image size. In such an evaluation, the image encoder for the VLM according to the present disclosure has an accuracy on the environmental common object dataset (e.g., OV-COCO novel AP 50 ) and the pre-trained dataset (e.g., DataComp-1B) that can be 1.4% higher than that of ViT-L / 14. Further, in terms of the zero-shot evaluation of the open-vocabulary segmentation task (e.g., the accuracy of the VLM output on a dataset that has not been previously evaluated by the VLM), the VLM according to the present disclosure also outperforms the image encoders (e.g., ViT-L / 14, ConvNeXt-L, etc.) in the FC-CLIP framework trained on COCO. Additionally, LLaVA-1.5 can be used as an evaluation framework for providing image embeddings that are well-aligned with the text. When benchmarking in LLaVA-1.5, using the datasets in the VLM described in the present disclosure is superior to VLMs such as ViT-L / 14 or CLIPA-v2. It should be understood that the dataset can be ImageNet, Visual Question Answering V2.0 data, etc.

[0086] In an embodiment, a test platform for designing a vision model using the DataComp-1B dataset (i.e., a high-quality dataset provided by DataComp, for example) under the CLIP framework is provided. Specifically, two training schemes (short-term plan and long-term plan) can be adopted. The short-term plan can be used for quickly benchmarking vision models across model and data scales on DataComp-1B. The long-term plan can be used to train the best-performing vision model on DataComp-1B. Under the short-term plan, the state-of-the-art vision models can be re-benchmarked in the ImageNet setting of the VLM. It should be understood that the "test platform" can be a technical term referring to the hardware and / or software environment configured to test the performance of vision models.

[0087] From the above, it will be understood that the various embodiments of the present disclosure have been described for illustrative purposes, and various modifications can be made without departing from the scope and spirit of the present disclosure. Therefore, the various embodiments disclosed herein are not intended to be limiting, and its true scope and spirit are indicated by the appended claims.

[0088] It should be understood that the disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them.

[0089] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), stored in a single file dedicated to the program in question, or stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.

[0090] The processes and logical flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. These processes and logical flows can also be executed by dedicated logic circuitry, and the apparatus can also be implemented as dedicated logic circuitry, e.g., a field programmable gate array, an application specific integrated circuit, etc.

[0091] Processors suitable for executing computer programs include, for example, general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as, for example, magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to one or more mass storage devices for storing data to receive data therefrom or transfer data thereto, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile or non-transitory memory, media, and memory devices, including, for example, semiconductor memory devices, such as, erasable programmable read-only memory, electrically erasable programmable read-only memory, and flash memory devices; magnetic disks, such as, internal hard disks or removable disks; magneto-optical disks; and compact disk read-only memory and digital versatile disk read-only memory disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0092] It should be understood that different features, variations, and multiple different embodiments have been shown and described in various details. What is sometimes described in terms of specific embodiments in this application is for illustrative purposes only and is not intended to limit or imply that only one particular embodiment or multiple specific embodiments are contemplated. It should be understood that the present disclosure is not limited to any single specific embodiment or enumerated variation. Many modifications, variations, and other embodiments will occur to those skilled in the art, and these modifications, variations, and other embodiments are intended to and in fact are covered by the present disclosure. Indeed, the scope of the present disclosure should be determined by a proper legal interpretation and construction (including equivalents) of the present disclosure as would be understood by those skilled in the art based on the complete disclosure presented at the time of filing.

[0093] Aspect: It should be understood that any aspect among the aspects can be combined with each other.

[0094] Aspect 1. A vision system for generating a feature map from an image, the vision system comprising:

[0095] A vision model configured to process the image to generate a feature map implemented on a neural network, wherein the vision model comprises:

[0096] A first convolutional block for downsampling an image data set to obtain first-level convolutional data;

[0097] A second convolutional block for downsampling the first-level convolutional data to obtain second-level convolutional data, wherein,

[0098] One or both of the first convolutional block and the second convolutional block is a mobile convolutional block (MBConv) including the following: a first Gaussian error linear unit (GELU) layer, a depthwise convolution (DWConv) layer, and a resize convolution layer; and

[0099] A transformer block (TFB) that generates the feature map based on the second-stage convolution data.

[0100] Aspect 2. The vision system according to aspect 1, wherein the GELU layer has a first kernel size and a first channel size, and is configured to expand the first channel size to a second channel size.

[0101] Aspect 3. The vision system according to aspect 1 or 2, wherein the resize convolution layer is configured to return the second channel size to the first channel size.

[0102] Aspect 4. The vision system according to any one of aspects 1 to 3, wherein the DWConv layer has a second kernel size and is configured to capture spatial interactions.

[0103] Aspect 5. The vision system according to any one of aspects 1 to 4, wherein the vision model further includes a backbone convolutional block having two convolutional layers with the same kernel size, and the backbone convolutional block processes the image to obtain backbone image data, and the backbone image data is provided to the first convolutional block as the image dataset.

[0104] Aspect 6. The vision system according to any one of aspects 1 to 5, wherein from the first convolutional block to the second convolutional block and then to the TFB, the number of blocks and the number of channels gradually increase.

[0105] Aspect 7. The vision system according to any one of aspects 1 to 6, wherein the TFB includes a self-attention (SA) residual block and a feed-forward network (FFN) residual block.

[0106] Aspect 8. The vision system according to aspect 7, wherein the first layer in the SA residual block includes a layer normalization (LN) layer.

[0107] Aspect 9. The vision system according to aspect 7 or 8, wherein the first layer in the FFN residual block is a layer normalization (LN) layer.

[0108] Aspect 10. The vision system according to any one of aspects 7 to 9, wherein the output of the layer normalization (LN) layer of the FFN is provided to a first linear layer and a second linear layer, and the output of the second linear layer is processed by a second GELU layer.

[0109] Aspect 11. The vision system according to aspect 10, wherein the output of the first linear layer and the output of the second GELU layer are combined to provide an input to a subsequent linear layer for generating the feature map.

[0110] Aspect 12. The vision system according to any one of aspects 1 to 11, wherein the neural network is trained using a contrastive language-image pre-training (CLIP) framework.

[0111] Aspect 13. The vision system according to any one of aspects 1 to 13, wherein the neural network is trained using locked text adaptation, including:

[0112] Initializing a text encoder using a pre-trained model;

[0113] Freezing the text encoder initialized using the pre-trained model; and

[0114] Training the neural network to obtain the vision model, wherein the training includes training using an image dataset to determine the weights of nodes in the neural network until a loss function is less than or equal to a predetermined value.

[0115] Aspect 14. A method for generating a vision model according to any one of aspects 1 to 13, the method comprising:

[0116] Benchmarking a plurality of vision models in a test platform of the model, the test platform being configured to benchmark the plurality of vision models according to a short-term plan for quickly benchmarking the vision models under contrastive language-image pre-training (CLIP) and a long-term plan for determining the performance of the plurality of vision models, wherein,

[0117] The benchmarking includes analyzing the plurality of vision models for data scalability, model scalability, and feature resolution, at least with respect to classification ability, retrieval ability, open vocabulary detection ability, or large multi-modal model performance; and

[0118] Generating the vision model based on the short-term plan and the long-term plan.

[0119] Aspect 15. The method according to aspect 14, wherein the benchmarking includes analyzing the plurality of vision models with respect to one or both of the open vocabulary detection ability or the large multi-modal model performance, and the classification ability and the retrieval ability.

[0120] Aspect 16. The method according to any one of aspects 1 to 13, further comprising training the vision model by:

[0121] Initialize a text encoder using a pre-trained model;

[0122] Freeze the text encoder initialized using the pre-trained model; and

[0123] Train a randomly initialized image encoder using an image dataset having multiple image-text pairs to obtain the neural network.

[0124] Aspect 17. The vision system according to any one of Aspects 1 to 13, wherein the vision model is generated by fitting a dataset having multiple image-text pairs to determine weights of the nodes such that a contrast loss is lower than a predetermined threshold.

[0125] Aspect 18. A method for validating a vision model according to any one of Aspects 1 to 13, the method comprising:

[0126] Process an image dataset using the vision model; and

[0127] Validate the vision model based on classification ability, retrieval ability, open vocabulary detection ability, and large multi-modal model performance for data scalability, model scalability, and feature resolution.

[0128] Aspect 19. The method according to Aspect 18, wherein the validation further comprises validating the vision model based on zero-shot accuracy with respect to classification ability and retrieval ability.

[0129] Aspect 20. The vision system according to any one of Aspects 1 to 13, wherein the vision model is validated by validating the vision model until a contrast loss according to a loss function is reduced to a predetermined threshold.

[0130] The terms used in this specification are intended to describe particular embodiments and are not intended to be limiting. Unless otherwise clearly stated, the terms "a / an" and "the" are intended to include the plural forms as well. When used in this specification, the term "comprises and / or comprising" specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or components.

[0131] Regarding the foregoing description, it should be understood that changes may be made to the details, particularly the construction materials and the shape, size, and arrangement of the parts used, without departing from the scope of the present disclosure. This specification and the described embodiments are merely exemplary, and the true scope and spirit of the present disclosure are indicated by the appended claims.

Claims

1. A vision system for generating a feature map from an image, the vision system comprising: A vision model configured to process the image to generate a feature map implemented on a neural network, wherein the vision model comprises: A first convolutional block for downsampling an image dataset to obtain first-level convolutional data; A second convolutional block for downsampling the first-level convolutional data to obtain second-level convolutional data, wherein One or both of the first convolutional block and the second convolutional block is a mobile convolutional block MBConv comprising: a first Gaussian error linear unit GELU layer, a depthwise convolution DWConv layer, and a resize convolution layer; and A transformer block TFB that generates the feature map based on the second-level convolutional data.

2. The vision system according to claim 1, wherein the GELU layer has a first kernel size and a first channel size and is configured to expand the first channel size to a second channel size.

3. The vision system according to claim 2, wherein the resize convolution layer is configured to return the second channel size to the first channel size.

4. The vision system according to claim 1, wherein the DWConv layer has a second kernel size and is configured to capture spatial interactions.

5. The vision system according to claim 1, wherein the vision model further comprises a backbone convolutional block having two convolutional layers with the same kernel size, the backbone convolutional block processing the image to obtain backbone image data, the backbone image data being provided to the first convolutional block as the image dataset.

6. The vision system according to claim 1, wherein from the first convolutional block to the second convolutional block to the TFB, the number of blocks and the number of channels gradually increase.

7. The vision system according to claim 1, wherein the TFB comprises a self-attention SA residual block and a feed-forward network FFN residual block.

8. The vision system according to claim 7, wherein the first layer in the SA residual block comprises a layer normalization LN layer.

9. The vision system according to claim 7, wherein the first layer in the FFN residual block is a layer normalization LN layer.

10. The vision system according to claim 7, wherein the output of the layer normalization LN layer of the FFN is provided to both a first linear layer and a second linear layer, and the output of the second linear layer is processed by a second GELU layer.

11. The vision system according to claim 10, wherein the output of the first linear layer and the output of the second GELU layer are combined to provide an input to a subsequent linear layer for generating the feature map.

12. The vision system according to claim 1, wherein the neural network is trained using a contrastive language-image pre-training CLIP framework.

13. The vision system according to claim 1, wherein the neural network is trained using locked text adaptation, comprising: Initializing a text encoder using a pre-trained model; Freeze the text encoder initialized with the pre-trained model; And Train the neural network to obtain the visual model, where the training includes training with an image dataset to determine the weights of the nodes in the neural network until the loss function is less than or equal to a predetermined value.

14. A method for generating the visual model according to claim 1, the method comprising: Benchmarking multiple visual models in a test platform of the model, the test platform being configured to benchmark the multiple visual models according to a short-term plan for quickly benchmarking the visual models under contrastive language-image pre-training CLIP and a long-term plan for determining the performance of the multiple visual models, where The benchmarking includes analyzing the multiple visual models for data scalability, model scalability, and feature resolution, at least with respect to classification ability, retrieval ability, open vocabulary detection ability, or large multi-modal model performance; and Generating the visual model based on the short-term plan and the long-term plan.

15. The method according to claim 14, wherein the benchmarking includes analyzing the multiple visual models with respect to one or both of the open vocabulary detection ability or the large multi-modal model performance, the classification ability, and the retrieval ability.

16. The method according to claim 1, further comprising training the visual model by: Initializing a text encoder with a pre-trained model; Freezing the text encoder initialized with the pre-trained model; and Using an image dataset with multiple image-text pairs to train a randomly initialized image encoder to obtain the neural network.

17. The visual system according to claim 1, wherein the visual model is generated by fitting a dataset with multiple image-text pairs to determine the weights of the nodes such that the contrastive loss is below a predetermined threshold.

18. A method for validating the visual model according to claim 1, the method comprising: Processing an image dataset using the visual model; And Validating the visual model for data scalability, model scalability, and feature resolution, based on classification ability, retrieval ability, open vocabulary detection ability, and large multi-modal model performance.

19. The method according to claim 18, wherein the validation further includes validating the visual model with respect to classification ability and retrieval ability based on zero-shot accuracy.

20. The visual system according to claim 1, wherein the visual model is validated by validating the visual model until the contrastive loss according to the loss function is reduced to a predetermined threshold.