Apparatus and method for sharing and pruning weights of visual and language models
By using hypernetwork to generate shared and pruning vectors in multimodal models, pruning and sharing of text encoder and visual encoder weights is realized, and joint training is carried out, the problems of low efficiency and insufficient flexibility of model parameter utilization in the prior art are solved, and the reduction of model size and the maintenance of prediction quality are achieved.
Patent Information
- Application Number
- CN202380068894.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-14
- Filing Date
- 2023-09-26
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art lacks flexibility in realizing weight sharing of visual and language models, resulting in inefficient utilization of model parameters and difficulty in deploying in mobile devices and resource-constrained environments.
By introducing hypernetwork into the multimodal model, the shared vector and pruning vector are generated, the pruning and sharing of text encoder and visual encoder weights are realized, and joint training is carried out to optimize the utilization of model parameters and model size.
Improves the utilization efficiency of model parameters, reduces model size, while maintaining prediction quality, and is suitable for deployment of mobile devices and resource-constrained environments.
Smart Images

Figure CN119948490A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an apparatus and method for sharing and pruning weights of vision and language models, and in particular, to an apparatus and method for reducing the network size of vision and language models while maintaining their prediction quality. Background Art
[0002] Extraction modules have seen growing interest in processing different modalities, including language and image data. However, the large number of parameters in extraction modules may pose challenges when deploying them on real-world applications, especially on mobile devices. For example, high-level vision and language models may have millions of parameters, making parameter reduction and size optimization critical for their deployment on mobile devices and other resource-constrained environments.
[0003] Recent advances in computer vision have shown that models based on extractive modules can use architecturally similar models for cross-modal tasks involving both visual and language data. Such a setting naturally allows weight sharing across different modalities.
[0004] Weight sharing offers the advantage of encouraging weight reuse, thereby reducing the number of parameters while preserving the capacity of the model to some extent. However, existing weight sharing techniques have certain limitations, as many of them rely on manually designed sharing rules for sharing entire layers or blocks, which significantly limits the flexibility of weight sharing. Due to the potential performance degradation associated with reduced flexibility, there is a growing demand to maximize the utilization of model parameters. Summary of the invention
[0005] Technical Solution
[0006] One or more embodiments of the present disclosure provide an electronic device for operating a multimodal model including a text encoder and a visual encoder, the electronic device including: a user interface configured to receive a query; and one or more processors configured to: obtain one or more input images; input the query to the text encoder to obtain text features; input the one or more input images to the visual encoder to obtain image features; and output a response to the query based on the similarity between the text features and the image features, wherein weight vectors of the text encoder and the visual encoder are pruned and shared according to a shared vector and a pruned vector generated by a hypernetwork to compress the multimodal model, and wherein the hypernetwork and the multimodal model are jointly trained to minimize at least one of the difference between the weight vectors in the text encoder and the visual encoder, the difference between the weight vectors in different layers of the text encoder, and the number of parameters in the multimodal model.
[0007] Any one or any combination of the one or more processors can be configured to: perform weight sharing and weight pruning on the multi-head self-attention layers of the text encoder and the visual encoder.
[0008] The weight sharing may include one or both of cross-modal weight sharing between the text encoder and the visual encoder and block-by-block sharing performed within a block including multiple layers of the text encoder.
[0009] Any one or any combination of the one or more processors may be configured to perform weight sharing and weight pruning on the feed-forward network layers of the text encoder and the visual encoder.
[0010] The weight sharing may include one or both of cross-modal weight sharing between the text encoder and the visual encoder and block-by-block sharing performed within a block including multiple layers of the text encoder.
[0011] The hypernetwork can be trained to generate shared vectors and pruned vectors, so that the same weight vector is neither pruned nor shared.
[0012] When the weight vectors of the text encoder and the visual encoder are pruned and shared, weight sharing can take precedence over weight pruning.
[0013] The hypernetwork and the multimodal model can be jointly trained by freezing the weight vectors in the visual encoder and the text encoder after they are pruned and shared, and then updating the weights in the hypernetwork based on a comparison between the initial number of parameters in the multimodal model and the number of parameters remaining in the multimodal model after weight sharing and weight pruning.
[0014] When a query requires identifying a target image from one or more input images, any one or combination of the one or more processors may also be configured to: select an image from the one or more input images that has the greatest similarity to the text feature or has a similarity to the text feature that exceeds a predetermined similarity threshold, and provide the selected image as a response to the query.
[0015] When a query requires identifying a target object from a specific image among one or more images, any one or combination of the one or more processors may be further configured to: identify an object having a maximum similarity to a text feature or having a similarity to the text feature exceeding a predetermined similarity threshold in one or more input images, and visually indicate the identified object within the specific image as a response to the query.
[0016] According to another aspect of the present disclosure, a method for performing a multimodal task by using a multimodal model including a text encoder and a visual encoder may include: inputting a query to the text encoder to obtain text features from the query; inputting one or more input images to the visual encoder to obtain image features from the one or more input images; and outputting a response to the query based on the similarity between the text features and the image features, wherein weight vectors of the text encoder and the visual encoder are pruned and shared according to a shared vector and a pruned vector generated by a hypernetwork to compress the multimodal model, and wherein the hypernetwork and the multimodal model are jointly trained to minimize at least one of the difference between the weight vectors in the text encoder and the visual encoder, the difference between the weight vectors in different layers of the text encoder, and the number of parameters in the multimodal model.
[0017] The method may also include: performing weight sharing and weight pruning on the multi-head self-attention layers of the visual encoder and the text encoder.
[0018] The weight sharing may include one or both of cross-modal weight sharing between the text encoder and the visual encoder and block-by-block sharing performed within a block including multiple layers of the text encoder.
[0019] The method may also include: performing weight sharing and weight pruning on the feed-forward network layers of the visual encoder and the text encoder.
[0020] The weight sharing may include one or both of cross-modal weight sharing between the text encoder and the visual encoder and block-by-block sharing performed within a block including multiple layers of the text encoder.
[0021] The hypernetwork can be trained to generate shared vectors and pruned vectors such that the same weight vector is neither pruned nor shared.
[0022] When the weight vectors of a text encoder are pruned and shared, weight sharing can take precedence over weight pruning.
[0023] The hypernetwork and the multimodal model can be jointly trained by freezing the weight vectors in the visual encoder and the text encoder after they are pruned and shared, and then updating the weights in the hypernetwork based on a comparison between the initial number of parameters in the multimodal model and the number of parameters remaining in the multimodal model after weight sharing and weight pruning.
[0024] When a query requires identifying a target image from one or more input images, outputting a response to the query may include: selecting an image from the one or more input images that has the greatest similarity to a text feature or has a similarity to the text feature that exceeds a predetermined similarity threshold, and providing the selected image as a response to the query.
[0025] When a query requires identifying a target object from a specific image among one or more images, outputting a response to the query may include: identifying an object having a maximum similarity to a text feature or having a similarity to the text feature exceeding a predetermined similarity threshold in one or more input images, and visually indicating the identified object within the specific image as a response to the query.
[0026] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing a program is provided, wherein the program can be executed by at least one processor to perform a method for performing a multimodal task by using a multimodal model including a text encoder and a visual encoder, the method comprising: obtaining text features from a query by inputting a query into the text encoder; obtaining image features from one or more input images by inputting one or more input images into the visual encoder; and outputting a response to the query based on the similarity between the text features and the image features, wherein weight vectors of the text encoder and the visual encoder are pruned and shared according to a shared vector and a pruned vector generated by a hypernetwork to compress the multimodal model, and wherein the hypernetwork and the multimodal model are jointly trained to minimize at least one of the difference between the weight vectors in the text encoder and the visual encoder, the difference between the weight vectors in different layers of the text encoder, and the number of parameters in the multimodal model.
[0027] Additional aspects will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the presented embodiments of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The above and other aspects, features and aspects of the embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings, in which:
[0029] Figure 1 is a diagram of an electronic device including a multimodal search model according to an embodiment of the present disclosure;
[0030] Figure 2 is a diagram illustrating a method of performing weight pruning and sharing between a text model and a visual model according to an embodiment of the present disclosure;
[0031] Figure 3 A text transformer and a visual transformer according to an embodiment of the present disclosure are shown.
[0032] Figure 4A is a diagram of a multimodal search model according to an embodiment of the present disclosure, the multimodal search model performing a task of detecting an object from an input image and a task of selecting one or more images from a plurality of input images based on an input query;
[0033] Figure 4B is a diagram illustrating an example multimodal task for identifying a target object in response to an input query according to an embodiment of the present disclosure;
[0034] Figure 5 is a flowchart illustrating a method of performing weight pruning and sharing between a text model and a visual model according to an embodiment of the present disclosure;
[0035] Figure 6 is a diagram of an apparatus for performing a multi-modal task according to an embodiment.
[0036] Figure 7 According to the embodiment of the present disclosure Figure 6 a diagram of components of one or more devices;
[0037] Figure 8 An example of a task execution result according to an embodiment of the present disclosure is shown; and
[0038] Fig. 9 An example of another task execution result according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0039] Example embodiments are described in more detail below with reference to the accompanying drawings.
[0040] In the following description, the same reference numerals are used for the same elements even in different drawings. Matters defined in the description, such as detailed construction and elements, are provided to assist in a comprehensive understanding of the exemplary embodiments. However, it is apparent that the exemplary embodiments can be practiced without those specifically defined matters. In addition, well-known functions or structures are not described in detail because they would obscure the description with unnecessary detail.
[0041] Expressions such as “at least one of…” when preceding a list of elements modify the entire list of elements and do not modify the individual elements of the list. For example, the expression “at least one of A, b, and c” should be understood to include only A, only b, only c, both A and b, both A and c, both b and c, all of A, b, and c, or any variation of the foregoing.
[0042] Although terms such as "first", "second", etc. may be used to describe various elements, these elements must not be limited to the above terms. The above terms may be used only to distinguish one element from another element.
[0043] The term "module" or "component" is intended to be broadly interpreted as hardware, firmware, or a combination of hardware and software.
[0044] It is apparent that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0045] Although particular combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may be directly dependent on only one claim, the disclosure of possible implementations includes the combination of each dependent claim with every other claim in the claim set.
[0046] Unless explicitly described as such, the elements, actions or instructions used herein should not be interpreted as critical or necessary. In addition, as used herein, the articles "one" and "an" are intended to include one or more projects, and can be used interchangeably with "one or more". In addition, as used herein, the term "set" is intended to include one or more projects (e.g., related projects, unrelated projects, combinations of related projects and unrelated projects, etc.), and can be used interchangeably with "one or more". In the case of intending only one project, the term "one" or similar language is used. In addition, as used herein, the terms "have", "have", "have" etc. are intended to be open terms. In addition, unless otherwise expressly stated, the phrase "based on" is intended to mean "based at least in part on".
[0047] One or more embodiments of the present disclosure provide an apparatus and method for maximizing the utilization of model parameters in the field of neural networks and model compression by integrating cross-modal weight sharing, block-by-block cross-layer weight sharing, and weight pruning into a unified framework. Weight sharing and pruning according to an embodiment can operate at the level of weight vectors rather than entire layers or blocks to enhance the flexibility of sharing and pruning operations.
[0048] In addition, in order to remove the reliance on manually designed strategies, the apparatus and method according to the embodiment adopt an end-to-end differentiable learning method to determine the locations for sharing and pruning. This allows automatic and optimized determination of sharing and pruning points without the need for explicit human intervention or predefined rules.
[0049] In addition, one or more embodiments of the present disclosure provide apparatus and methods for generating compact but effective multimodal models through selective weight sharing, weight pruning, and hypernetwork training, as well as a trade-off between the number of shared parameters and accuracy.
[0050] A hypernetwork can be trained to identify weight vectors for pruning or sharing. Instead of considering the entire layer, each weight vector can be considered a separate element for pruning and sharing to provide a denser search space. The hypernetwork can effectively optimize the dense search space, thereby enhancing the overall effectiveness of model size reduction.
[0051] According to an embodiment of the present disclosure, model compression can be achieved through three dimensions: cross-modal sharing, block-by-block cross-layer sharing, and pruning. Weight sharing and pruning techniques can be applied to two core components of the visual and text modules: multi-head self-attention and feedforward networks. Cross-modal sharing can facilitate the sharing of corresponding weight vectors across visual and text extraction modules. Block-by-block layer sharing splits the text extraction module into multiple blocks, enabling weight sharing across layers within the same block. Pruning is used to remove weight vectors that are not used for sharing. By incorporating these three dimensions, selective weight sharing and pruning are performed to strike a balance between shared parameters and accuracy.
[0052] According to an embodiment of the present disclosure, joint training of model weights and super networks may be performed. During the joint training process, weight vectors may not be pruned or shared directly. Instead, regularization terms may be introduced to minimize conflicts between models before and after sharing and pruning. This soft alignment process significantly enhances the performance of the compressed model.
[0053] In an embodiment of the present disclosure, a method for model size reduction may include the following steps: training a hypernetwork, selective weight sharing and pruning, and joint training of model weights and the hypernetwork.
[0054] At the step of training the hypernetwork, the hypernetwork is trained to identify weight vectors for pruning or sharing. Each weight vector can be viewed as a different element in the search space.
[0055] At the step of selective weight sharing and pruning, weight sharing and pruning techniques can be applied to the multi-head self-attention and feed-forward network components of the vision and text models. Cross-modal sharing allows corresponding weight vectors to be shared across modalities, while block-by-block cross-layer sharing enables weight sharing within a specific block. Pruning removes weight vectors that are not used for sharing.
[0056] At the step of jointly training the model weights and the hypernetwork, the model weights and the hypernetwork are jointly trained via a soft alignment process that uses a regularization term to minimize conflicts between the models before and after weight sharing and pruning.
[0057] Various embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0058] Figure 1 is a diagram of a computer system including a multimodal search model according to an embodiment of the present disclosure.
[0059] The computer system may include one or more neural networks to employ artificial intelligence (AI) techniques.
[0060] like Figure 1 As shown, the computer system may include a multimodal search model 100, which is configured to receive a query from a user input and identify an image corresponding to the query from a plurality of images retrieved from a data storage device (e.g., a photo library of a mobile device or an external database) or received through an internal search. When the query requests to retrieve a specific image, the multimodal search model 100 may select candidate images corresponding to the query from the plurality of images, rank the candidate images based on the similarity between each of the candidate images and the query, and select one or more final images based on the ranking of the candidate images. When the query requests to identify a specific target object included in the image, the multimodal search model 100 may identify one or more candidate objects from the input image, rank the candidate objects based on the similarity between each of the candidate objects and the query, select at least one final object based on the ranking of the candidate objects, and display a bounding box to show the selected final object.
[0061] The multimodal search model 100 may include a text encoder 110, a visual encoder 120, a similarity function module 130, an image selection module 140, and an object recognition module 150. In addition, when an electronic device including the multimodal search model 100 updates the multimodal search model 100 based on on-device learning, the multimodal search model 100 may include a loss calculator 160. If the electronic device uses the multimodal search model 100 as a pre-trained fixed model, the loss calculator 160 may be omitted.
[0062] The text encoder 110 may receive the query via a touch screen, a keyboard, a microphone and / or a communication interface. When the query is received via a voice signal, the voice signal may be converted from speech to text to obtain text information corresponding to the voice in the voice signal.
[0063] The text encoder 110 may include a text feature extraction module 111 , a text feature transformation module 112 , a text feature weighting module 113 , and a text feature aggregation module 114 .
[0064] The text feature extraction module 111 can extract text features (e.g., word features) from one or more words included in the query (e.g., vectors representing each of the words). For example, when a query stating "a woman is throwing a Frisbee in the park" is provided to the text feature extraction module 111, the text feature extraction module 111 can identify the four words "woman", "throwing", "Frisbee", and "park" in the query, and can extract word features from each word. The word feature has a content value, which is a vector value corresponding to the contextual representation of the word in the query. The text extraction module 111 can be embodied as a transformer, and in this case, the text extraction module 111 can be incorporated into the text feature transformation module 122.
[0065] The text feature transformation module 122 may include a linear projection layer to project word features to a joint embedding space to which both image features and word features are projected. The projection layer may apply a word feature transformation function that transforms word features of n words included in the query to a joint embedding space having a constant dimension.
[0066] The text feature weighting module 113 may provide a learnable weight function that is optimized to assign higher weights to relatively more important words among the n words included in the query. The feature weighting module 113 may load a set of weights pre-stored in a memory, and may update the weights using a learnable weight function for word features according to the loss calculated by the loss calculator 160.
[0067] The text feature aggregation module 114 may apply weights to the transformed word features and the aggregated weighted word features, for example, via mean pooling.
[0068] The visual encoder 120 may include an image feature extraction module 121 , an image feature transformation module 122 , an image feature weighting module 123 , and an image feature aggregation module 124 .
[0069] The image feature extraction module 121 can extract image features from an image, which captures spatial information in the image (e.g., the appearance of an object and / or scene). The content value of the image feature can be calculated by detecting salient regions or grid cells in the image, mapping the detected salient regions or grid cells to a set of vectors, and averaging the set of vectors.
[0070] The spatial information may enable the visual encoder 120 to remove regions of the image that include uninformative scenes or objects. The image feature extraction module 121 may be embodied as a transformer, or alternatively, as a two-dimensional (2D) convolutional neural network (CNN), R-CNN, fast R-CNN, or faster R-CNN. For example, when an image capturing a dog playing with a toy is provided to the image feature extraction module 121, the image feature extraction module 121 may identify a first image region of the dog and a second image region of the toy from the image, and may extract image features (e.g., a first vector representing the first region of the image and a second vector representing the second region of the image) from each of the first image region and the second image region. The extracted image features are fed to the image feature transformation module 122 and the image feature weighting module 123, respectively. When the image feature extraction module 121 is embodied as a transformer, the image feature extraction module 121 may be merged into the image feature transformation module 122.
[0071] The image feature transformation module 122 may include a linear projection layer to project the image features into a joint embedding space where semantically similar feature points in different modalities (i.e., image and text) are closer to each other in distance. The projection layer may apply an image feature transformation function that transforms the image features of a region of the image into the joint embedding space.
[0072] The image feature weighting module 123 may provide a learnable weight function that is optimized to assign higher weights to important areas of the image. The image feature weighting module 123 may load a set of weights pre-stored in a memory, and may update the weights using a learnable weight function for image features according to the loss calculated by the loss calculator 160.
[0073] The image feature aggregation module 124 may apply weights to the transformed image features and the aggregated weighted image features, for example, via mean pooling.
[0074] During the training process, the similarity function module 130 may calculate similarity scores for matching query-image pairs and similarity scores for non-matching query-image pairs. For example, cosine similarity or negative Euclidean distance may be calculated as the similarity score.
[0075] The loss calculator 160 can calculate the triplet loss based on the similarity scores of the matched query-image pairs and the similarity scores of the non-matched query-image pairs. For training purposes, non-matching query features and non-matching image features can be randomly selected to generate random negative non-matching samples. The triplet loss can be back-propagated to the visual encoder 120 and the text encoder 110, so that the image feature weighting module 123 and the text feature weighting module 113 can respectively update the weights of the image features and the weights of the word features to minimize or converge the triplet loss. When the triplet loss has reached a predetermined minimum value or a constant value with a preset margin, it can be determined that the triplet loss is minimized or converged. The visual encoder 120 and the text encoder 110 can be jointly trained based on the triplet loss.
[0076] In the inference phase, the similarity function module 130 may calculate a similarity score between the input query and each of the plurality of input images, and may provide the similarity score to the image selection module 140 or the object recognition module 150 , depending on the task given according to the input query.
[0077] When a query requests to retrieve a specific image, the image selection module 140 can rank the input images based on the similarity scores, and can select candidate images based on the ranking. For example, a preset percentage (e.g., the top 10% or 20% images) or a preset number of images (e.g., the 4 images with the highest similarity scores) can be selected from the plurality of input images based on the ranking. Alternatively, or in conjunction with the use of the ranking, a predetermined similarity threshold can be applied to select the candidate images. For example, any image having a similarity score above a predetermined similarity threshold can be selected as the final image, or among the images selected based on the ranking, an image having a similarity score above a predetermined similarity threshold can be selected as the final image.
[0078] When a query requests identification of a specific target object from a given image, the object identification module 150 may identify one or more candidate objects from the given image, rank the candidate objects based on a similarity between each of the candidate objects and the query, select at least one final object based on the ranking of the candidate objects, and display a bounding box (or another visual indicator) to show the selected final object as a response to the input query.
[0079] Figure 2 is a diagram illustrating a method for performing weight pruning and sharing between a text transformer and a visual transformer according to an embodiment of the present disclosure. Figure 2 The text transformer and visual transformer shown may correspond to Figure 1The presented text feature transformation module 112 and image feature transformation module 122. The text transformer may include multiple layers constituting a neural network, and the multiple layers may be grouped into multiple blocks. Similarly, the visual transformer may include multiple layers constituting another neural network, and the multiple layers of the visual encoder 120 may be grouped into multiple blocks.
[0080] For ease of description, refer to Figure 2 , the weight pruning and sharing according to the embodiments of the present disclosure are described as being applied to the text transformer and the visual transformer. However, the weight pruning and sharing can be applied to any other elements within the text encoder 110 and the visual encoder 120, such as a pair of text feature extraction modules 111 and image feature extraction modules 121, and a pair of text feature weighting modules 113 and image feature weighting modules 123.
[0081] refer to Figure 2 , a hypernetwork can be used to perform weight pruning and sharing between the text transformer and the visual transformer. The hypernetwork can generate a shared vector s and a pruned vector m, so that the text transformer shares its weights with the transformer based on the shared vector s, and the text transformer and the visual transformer prune their weights according to the pruned vector m.
[0082] Each of the text transformer and the visual transformer may include a multi-head attention (MSA) layer and a feed-forward network (FFN) layer. Since the MSA and FFN layers contain an important portion of the weights in the multimodal search model 100, they can be the main targets for structural weight sharing and pruning. In order to determine which node weights should be pruned or shared, the shared vector s and pruned vector m generated by the hypernetwork are optimized during the pre-training phase. This optimization process improves the efficiency of weight utilization and improves model performance. In addition, to enhance the flexibility of the pruning and sharing process, block-by-block cross-layer sharing is applied within each block of the text transformer. This technique allows for better adaptability and fine-grained control of the weight pruning and sharing mechanism within the multimodal search model 100.
[0083] The text transformer t and the visual transformer v may have the same or substantially the same structure. When each of the text transformer and the visual transformer includes L layers in total and l=1, ..., L, The weight matrix of the lth layer of the text transformer t can be used to represent and the weight matrix of the lth layer of the text transformer v The text transformer may include a plurality of layers constituting a neural network, and the plurality of layers may be grouped into a plurality of blocks. Similarly, the visual transformer may include a plurality of layers constituting another neural network, and the plurality of layers of the visual transformer may be grouped into a plurality of blocks.
[0084] For the text transformer t and the visual transformer v, there is a weight matrix for query , weight matrix for keywords and a weight matrix for the values Three weight matrices, where , d is the embedding dimension, and N* represents the number of tokens for a given modality. Using the input tokens, the final Q*, K*, and V* for self-attention are obtained as follows:
[0085] Equation (1)
[0086] In order to make minimal changes to the original multimodal search model 100, only the weight matrix used for the query and the weight matrix for keywords The two weight matrices perform pruning and structural sharing between the text transformer t and the visual transformer v and across each block of the text transformer t for query weight matrices , weight matrix for keywords , for values The weight matrix of all three weight matrices. However, the embodiments of the present disclosure are not limited thereto, and pruning can be applied to query The weight matrix for keywords , the weight matrix for the values Any one or any combination of .
[0087] More specifically, for pruning, the binary vector The weight matrix that can be applied to the query and the weight matrix for keywords When pruning is applied to the visual transformer, the query and keywords The generation of is as follows:
[0088] Equation (2)
[0089] in, In the initial stage it is resized to have Same size, is generated element by element, and is the feature map from the previous layer.
[0090] Hypernetworks can generate shared vectors To perform cross-modal weight sharing between the text transformer and the visual transformer, and also to perform cross-layer weight sharing between different layers of the text transformer within each block of the text transformer. For sharing, the weights of the text transformer can be used to share with the visual transformer. As a result, the weight vector can be used in different layers across the text transformer and the visual transformer.
[0091] By sharing across modalities, the weight matrix for query in the visual transformer can be obtained as follows :
[0092] Equation (3)
[0093] in First expand to have The same size, and The weight vector of index i in is shared.
[0094] The weight matrix for query in the text transformer is shared across layers in a block-by-block manner. It can be obtained as follows:
[0095] Equation (4)
[0096] in, represents the weight of the lth layer in the text transformer, represents the base weight of the assignment, and is a shared vector for cross-layer sharing. As a result, the final weights of the text transformer are split into two parts: layer-specific weights and shared weights from the base layer.
[0097] In an embodiment of the present disclosure, for block-by-block cross-layer sharing, weights from a certain layer may be used as base weights for all other layers.
[0098] In other embodiments of the present disclosure, the text transformer can be divided into multiple blocks, and each block includes multiple layers and has its own base weights to enhance the flexibility of cross-layer sharing. Specifically, a set of base layers can be established as follows: All layers of the text transformer can be split into multiple blocks based on the base layer, so that the nth block contains For example, the weight of the first layer in each block can be set to the base weight. When l is the base layer ( ), you can assign its own weight to the base layer Set to 1, and the weights of the base layer can be used for other layers in the same block.
[0099] After block-wise cross-layer sharing, the base weights for cross-modal sharing can also be changed. When both block-wise cross-layer sharing and cross-modal sharing are applied, the final weights of the visual transformer are It can be expressed as follows:
[0100] Equation (5)
[0101] As shown above, in addition to the visual transformer specific weights ( ), the final weight of the visual transformer You can also include weights from the text transformer of the same layer ( ), the weights of the text transformer from the base layer ( ).
[0102] Block-wise cross-layer sharing and cross-modal sharing can be applied to the FFN layer and MSA layer of the text transformer and the visual transformer.
[0103] According to an embodiment of the present disclosure, restrictions may be applied to cross-modal sharing and pruning to resolve conflicts between pruned vectors m and shared vectors s. For example, when pruning and sharing the same weights are guided based on pruned vectors m and shared vectors s, meaningless actions of sharing pruned weights will occur. In order to resolve conflicts between pruned vectors m and shared vectors s, the following restrictions may be applied to cross-modal sharing and pruning:
[0104] Equation (6)
[0105] Where i represents the index of a weight vector to be shared or pruned. , then the text transformer and visual transformer do not share and pruned the same weight vector.
[0106] For cross-layer sharing and pruning of the text transformer, the following restrictions can be applied to resolve the conflict between the shared vector s and the pruned vector m as follows:
[0107] ,in
[0108] ,in Equation (7)
[0109] in Indicates the last element of the current block. For example, if , then the current block consists of layers ( ), and . Represents all shared elements from different layers in the current block, and these shared elements can remain in the base layer. Because there may be no conflict between shares, this restriction can be removed and Between applications.
[0110] For example, when the pruning vector and the shared vector are zero ( ==0), by Set to 1 to enable sharing of vectors Prioritizes pruning vectors , so that pruning of shared weights is not allowed according to the restrictions in Equation (6). Sharing weights can be prioritized over pruning weights to preserve model capacity through sharing.
[0111] To generate the pruned vector m and the shared vector s, the hypernetwork can be obtained as follows: and the Gumbel-Sigmoid technique parameterized by:
[0112] Equation (8)
[0113] Where z is a predetermined vector used as input to the hypernetwork. For example, z can be random noise sampled from a Gaussian distribution. The hypernetwork can include a gated recurrent unit (GRU) configured to capture inter-layer interactions and a multi-layer perceptron (MLP) used as intra-layer interactions. The hypernetwork is trained to solve the following optimization problem:
[0114] Equation (9)
[0115] The first part of Equation (9) represents the pre-training task loss for the cross-modal pre-training task. The second part of Equation (9) represents parameter regularization, which pushes the number of parameters towards a predetermined threshold p.
[0116] In equation (9), x is the input sample of an image and text pair, y is the label (true value) of the input sample, and is the original pre-training loss for multimodal models (e.g., MDETR and GLIP) with a model structure determined by m and s. Indicates a given (0, 1] regularization loss that controls the amount of parameters the model should keep. represents the control reference for controlling the strength of the positive lateralization. P(m,s) in the regularization loss represents the remaining number of parameters retained by m and s, and P totalRepresents the total number of parameters in the MSA and FFN layers of the text and visual transformers. R can be a regression loss function such as mean squared error (MSE) or mean absolute error (MAE). For weight sharing and pruning, m and s derived from the hypernetwork can be directly applied to the visual and text transformers. To address the problem of possible accuracy degradation due to significant changes in the outputs of the visual and text transformers, a selection-based regularization mechanism can be applied to gently push the selected weights closer to each other for sharing, or push them towards zero for pruning, as shown below:
[0117]
[0118] Equation (10)
[0119] In the first term (i.e., ) pushes the selected weight vectors closer to reduce the difference between the text transformer weights and the visual transformer weights, the second term (i.e., ) pushes the selected weight vectors closer to reduce the difference between different layers in the same block, and the third term (i.e., ) penalizes the weights pruned by the pruning vector m by pushing the pruned weights close to 0. With this regularization term, the weights of the multimodal model are aligned with the pruning vector m and the shared vector s, which creates a smoothing process for reducing the number of model parameters.
[0120] Given a regularized loss, the model weights W are learned by optimizing the following objective function:
[0121] Equation (11)
[0122] The first part of equation (11) represents the pre-training task loss for the cross-modal pre-training task, where x and y are the input samples and labels, and W represents the multimodal model weights. The second part of equation (11) represents the regularization loss used to align the multimodal model weights W before and after fine-tuning. In equation (11), Indicates control The control parameter of the strength of . After the pre-training process, the corresponding weights are pruned and shared to compress the multimodal model. Then, the compressed multimodal model can be used for fine-tuning on downstream tasks.
[0123] The training process of the hypernetwork and multimodal models can be performed using Algorithm 1 presented below:
[0124] _____________________________________________________________
[0125] Algorithm 1: Learning to jointly share and prune weights for grounded vision and language tasks
[0126] Input: Pre-training dataset and sub-dataset for learning structure vectors: ;
[0127] The remaining parameter rates are: ; Hyperparameters: ; Pre-training rounds: ; Models for pre-training: ;Depend on Parameterized Hypernetwork HN
[0128]
[0129] Based on Pruning and Sharing Come get .return For task-specific fine-tuning.
[0130] According to Algorithm 1, during the training of the hypernetwork, the pruned and shared vectors m and s are applied in the forward computation. However, when optimizing the multimodal model, the forward computation remains unchanged. In order to reduce the training time of the hypernetwork, a smaller subset D of the pre-training dataset D can be used. sub . The hypernetwork is used to accelerate the learning of the pruned and shared vectors m and s. In addition, the hypernetwork helps capture the complex interactions between pruning and sharing across modalities and layers. Although it is possible to directly set the pruned and shared vectors m and s as learning parameters, doing so may lead to a slowdown in the learning process. Therefore, this may potentially reduce the final performance of the multimodal model.
[0131] Figure 3 A text transformer and a visual transformer according to an embodiment of the present disclosure are shown.
[0132] The text transformer may include a normalization layer 301 configured to normalize input text features, a multi-head self-attention (MSA) layer 302 configured to capture relationships between different words in an input sequence, an adder 303 configured to add the normalized text features and the self-attention output, another normalization layer 304 configured to normalize the addition of the normalized text features and the self-attention output, a feed-forward neural network (FFN) layer 305 configured to apply a nonlinear transformation to the self-attention output, and another adder 306 configured to add the output of the FFN layer 305 to the addition of the normalized text features and the self-attention output.
[0133] The visual transformer may have the same or substantially the same structure as the text transformer. The visual transformer may include a normalization layer 311 configured to normalize input image features, a multi-head self-attention (MSA) layer 312 configured to capture the relationship between different image features in the input sequence, an adder 313 configured to add the normalized image features and the self-attention output, another normalization layer 314 configured to normalize the addition of the normalized image features and the self-attention output, a feedforward neural network (FFN) layer 315 configured to apply a nonlinear transformation to the self-attention output, and another adder 316 configured to add the output of the FFN layer 315 to the addition of the normalized image features and the self-attention output.
[0134] Figure 4A is a diagram of a multimodal search model according to an embodiment of the present disclosure, which performs the tasks of detecting objects from an input image and selecting one or more images from multiple input images based on an input query.
[0135] like Figure 2 As shown, the multimodal search model 100 may receive a second query requesting retrieval of one or more images from a plurality of input images that match the description in the query.
[0136] The text encoder 110 may identify the words included in the query and may extract word features (eg, vectors representing word features) from each of the words. When there are n words in the query, the text encoder 110 may extract a first word feature, a second word feature, ..., and an nth word feature.
[0137] The visual encoder 120 may identify regions of objects or scenes from the candidate images, and may extract region features R from the identified regions via the image feature extraction module 121. 1 , R 2 , R 3 , …, R m .
[0138] The attention module 230 can determine the region features R of the candidate image for the i-th word feature. 1 , R 2 , R 3 , ..., R m The corresponding weight w 1 、w 2 、w 3 ,....w m , where i 1, 2, ..., and n. The attention module 230 can be weighted 1 、w 2 、w 3, ..., w 4 、w m Applied to the regional features R 1 , R 2 , R 3 , …, R m , and the weighted regional feature w 1 R 1 、w 2 R 2 、w 3 R 3 ,…,w m R m The aggregated regional feature value is added to obtain the aggregated regional feature value, which is fed into the image selection module 240A as the regional feature noticed by the i-th word feature.
[0139] For example, when there are three word features extracted from three words of the query, the attention module 230 may calculate (1) the region feature R of the first candidate image with respect to the first word feature 1 , R 2 , R 3 , ..., R m The corresponding first set of weights w 11 、w 12 、w 13 ,... 11 、w 12 、w 13 ,.... 1m ; (2) Regional features R of the second word feature and the first candidate image 1 , R 2 , R 3 , ..., R m The corresponding second set of weights w 21 , w 22 , w 23 , ..., w 2m ; and (3) the regional feature R of the first candidate image for the third word feature 1 , R 2 , R 3 , ..., R m The corresponding third set of weights w 31 , w 32 , w 33 , ..., w 3m The attention module 230 may be configured to take the first set of weights w 11 、w 12 、w 13 ,....w 1m Applied to the regional features R 1 , R 2 , R3 , ..., R m , and the weighted regional feature w 11 R 1 、w 12 R 2 、w 13 R 3 , ..., w 1m R m The attention module 230 can add the second set of weights w 21 、w 22 、w 23 ,....w 2m Applied to the regional features R 1 , R 2 , R 3 , …, R m , and the weighted regional feature w 21 R 1 、w 22 R 2 、w 23 R 3 ,…,w 2m R m The attention module 230 can add the third set of weights w 31 、w 32 、w 33 ,....w 3m Applied to the regional features R 1 , R 2 , R 3 , …, R m , and the weighted regional feature w 31 R 1 、w 32 R 2 、w 33 R 3 , …, W 3m R m The third aggregation region feature value of the third word feature is obtained by adding them together.
[0140] The image selection module 240A may calculate a similarity score (e.g., cosine similarity or negative Euclidean distance) between the region features and the query features. In particular, the image selection module 240A may use a normalized similarity function to calculate a similarity score for each word feature, and may apply a mean aggregation of the similarity scores to obtain a final image-query similarity score. The final image-query similarity score may also be referred to as a "noted similarity score."
[0141] For example, the image selection module 240A can calculate a first similarity score between the first aggregated region feature and the first word feature, a second similarity score between the second aggregated region feature and the second word feature, and a third similarity score between the third aggregated region feature and the third word feature, and can calculate a weighted sum or average of the first similarity score, the second similarity score, and the third similarity score as the final image query similarity score.
[0142] The image selection module 240A can rank the candidate images based on their final image-query similarity scores, and can select at least one image based on the ranking of the candidate images. For example, a preset percentage (e.g., the top 10% or 20% images) or a preset number of images (e.g., the 100 images with the highest similarity scores) can be selected from the candidate images based on the ranking and can be presented to the user in the order of the ranking. Alternatively, or in conjunction with the use of the ranking, a predetermined similarity threshold can be applied to select the candidate images. For example, any candidate image having a similarity score above a predetermined similarity threshold can be selected, or only images having a similarity score above a predetermined similarity threshold among the candidate images selected based on the ranking are selected as a response to the query.
[0143] The multimodal search model 100 may receive another query requesting identification of a target object from an input image. In this case, the text features and the region features may be input to the object identification module 240B. The object identification module 240B may calculate similarity scores between a plurality of objects detected from the input image and the text features obtained from the input query, and identify the object having the highest similarity score with the input query as a response to the query.
[0144] Figure 4B is a diagram illustrating an example multimodal task for identifying a target object in response to an input query according to an embodiment of the present disclosure. Figure 2 The hypernetwork is used to compress Figure 4B The model weights of the text encoder 110 and the visual encoder 120 shown in FIG.
[0145] like Figure 4B As shown, a user query (e.g., “a woman holds a blow dryer, wearing protective goggles”) is input into a text encoder 110, and an image is input into a visual encoder 120. The text encoder 110 processes the user query to obtain word features P via a text feature extraction module 111 and a plurality of BERT layers 115. 1 , P 2 ,...P M, the multiple BERT layers 115 embody the text feature transformation module 112 (as well as the text feature weighting module 113 and the text feature aggregation module 114).
[0146] The visual encoder 120 processes the input image via the image feature extraction module 121 and multiple dynamic head (DyHead) modules 125 to obtain image region features. 1 , O 2 ,... N The multiple dynamic head (DyHead) modules 125 embody the image feature transformation module 122 (as well as the image feature weighting module 123 and the image feature aggregation module 124).
[0147] The object recognition module 240B can calculate the word feature P 1 , P 2 ,...P M With the image region feature O 1 , O 2 ,... N A similarity score (also called an “alignment score”) is calculated between the two images, and objects corresponding to “woman”, “blow dryer”, and “protective goggles” can be identified from the input image based on the similarity score.
[0148] Figure 5 is a flowchart illustrating a method for performing weight pruning and sharing between a text encoder and a visual encoder according to an embodiment of the present disclosure.
[0149] In operation S10 , a super network is pre-trained to generate a shared vector and a pruned vector such that both pruning and sharing are not applied to the same weight vector.
[0150] In operation S20, a shared vector and a pruned vector are obtained from a pre-trained super network.
[0151] In operation S30, the multimodal model is compressed based on the shared vector and the pruned vector. Operation S30 may include cross-modal sharing S31, block-by-block cross-layer sharing S32, and pruning S33. In detail, in operation S31, cross-modal sharing is performed between a text transformer and a visual transformer of the multimodal model, for example using equation (3). In operation S32, block-by-block cross-layer sharing is performed on the text transformer, for example using equation (4). In operation S33, pruning is performed on the visual transformer, for example using equation (2).
[0152] In operation S40, the compressed multimodal model and the pre-trained hypernetwork are jointly trained for optimization. Specifically, in operation S41, the weight vectors in the visual transformer and the text transformer are fixed. In operation S42, the weights of the hypernetwork are updated based on a comparison between the number of parameters remaining after model compression and the initial number of parameters of the multimodal model, for example, using equations (9)-(11).
[0153] Figure 6 is a diagram of an apparatus for performing a multi-modal task according to an embodiment. Figure 6 It includes user equipment 610, server 620 and communication network 630. User equipment 610 and server 620 may be interconnected via a wired connection, a wireless connection, or a combination of wired and wireless connections.
[0154] The user device 610 includes one or more devices (e.g., a processor 611 and a data storage device 612) configured to retrieve images corresponding to a search query. For example, the user device 610 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smart phone, a wireless phone, etc.), a camera device, a wearable device (e.g., a pair of smart glasses, a smart watch, etc.), a home appliance (e.g., a robot vacuum cleaner, a smart refrigerator, etc.), or the like. The data storage device 612 of the user device 610 may include the multimodal search model 100. When the multimodal search model 100 is stored in the server 602 instead of the user device 610, the user device 610 may send an input query to the server 602 and may receive a response to the query from the server 602 operating the multimodal search model 100.
[0155] The server 620 includes one or more devices (e.g., a processor 621 and a data storage device 622) configured to train the multimodal search model 100 and the refined search model 200, and / or retrieve images corresponding to a search query received from the user device 610. The data storage device 622 of the server 620 may include the multimodal search model 100.
[0156] The communication network 630 includes one or more wired and / or wireless networks. For example, the network 1300 may include a cellular network, a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, etc., and / or a combination of these or other types of networks.
[0157] Figure 6The number and arrangement of devices and networks shown in are provided as examples. Figure 6 There may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or differently arranged devices and / or networks than those shown in FIG. Figure 6 Two or more of the devices shown in may be implemented in a single device, or Figure 6 A single device shown in the may be implemented as multiple distributed devices. Additionally or alternatively, a set of devices (eg, one or more devices) may perform one or more functions described as being performed by another set of devices.
[0158] Figure 7 According to the embodiment Figure 6 A diagram of components of one or more electronic devices. Figure 7 The electronic device 1000 in the example may correspond to the user device 610 and / or the server 620.
[0159] Figure 7 This is for illustration only, and other embodiments of the electronic device 1000 may be used without departing from the scope of the present disclosure. For example, the electronic device 1000 may correspond to a client device or a server.
[0160] The electronic device 1000 includes a bus 1010 , a processor 1020 , a memory 1030 , an interface 1040 , and a display 1050 .
[0161] The bus 1010 includes a circuit for connecting the components 1020 to 1050 to each other. The bus 1010 serves as a communication system for transmitting data between the components 1020 to 1050 or between electronic devices.
[0162] The processor 1020 includes one or more of a central processing unit (CPU), a graphics processor unit (GPU), an accelerated processing unit (APU), an integrated many-core (MIC), a field programmable gate array (FPGA), or a digital signal processor (DSP). The processor 1020 is capable of controlling any one or any combination of the other components of the electronic device 1000, and / or performing operations related to communication or data processing. For example, the processor 1020 may execute Figure 5 Operations S10 - S40 are shown. Processor 1020 executes one or more programs stored in memory 1030 .
[0163] The memory 1030 may include volatile and / or non-volatile memory. The memory 1030 stores information, such as one or more of commands, data, programs (one or more instructions), applications 1034, etc., which are related to at least one other component of the electronic device 1000 and are used to drive and control the electronic device 1000. For example, the commands and / or data may prepare an operating system (OS) 1032. The information stored in the memory 1030 may be executed by the processor 1020. In particular, the memory 1030 may store the multimodal search model 100 and a plurality of images.
[0164] Application 1034 includes the above-described embodiments. These functions may be performed by a single application or multiple applications, each of which performs one or more of these functions. For example, application 1034 may include a program for executing Figure 5 Artificial intelligence (AI) model of operations S10-S40 is shown.
[0165] The display 1050 includes, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a quantum dot light emitting diode (QLED) display, a micro-electromechanical system (MEMS) display, or an electronic paper display. The display 1050 may also be a depth perception display, such as a multi-focal display. The display 1050 is capable of presenting, for example, various contents, such as text, images, videos, icons, and symbols.
[0166] Interface 1040 includes an input / output (I / O) interface 1042, a communication interface 1044, and / or one or more sensors 1046. I / O interface 1042 serves as an interface that can transmit commands and / or data between a user and / or other external devices and other components of electronic device 1000, for example.
[0167] The communication interface 1044 can realize the communication between the electronic device 1000 and other external devices via a wired connection, a wireless connection, or a combination of a wired and wireless connection. The communication interface 1044 can allow the electronic device 1000 to receive information from another device and / or provide information to another device. For example, the communication interface 1044 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc. The communication interface 1044 can receive video and / or video frames from an external device (such as a server).
[0168] The (one or more) sensors 1046 of the interface 1040 can measure physical quantities or detect the activation state of the electronic device 1000, and convert the measured or detected information into electrical signals. For example, the (one or more) sensors 1046 may include one or more cameras or other imaging sensors for capturing images of the scene. The (one or more) sensors 1046 may also include any one or any combination of a microphone, a keyboard, a mouse, and one or more buttons for touch input. The (one or more) sensors 1046 may also include an inertial measurement unit. In addition, the (one or more) sensors 1046 may include a control circuit for controlling at least one of the sensors included herein. Any of these sensors 1046 may be located within the electronic device 1000 or coupled to the electronic device 1000. The sensor 1046 may receive a text and / or voice signal containing one or more queries.
[0169] Figure 8 An example of a task execution result according to an embodiment of the present disclosure is shown.
[0170] refer to Figure 8 , the mobile device 2000 may receive a search query (e.g., "a squirrel is eating an acorn") via a microphone, a virtual keyboard, or a communication interface. The mobile device 2000 may input the search query and each image retrieved from the photo library of the mobile device 200 into a multimodal image retrieval model including the multimodal search model 100, and may output one or more images (e.g., image 1, image 2, image 3, and image 4) as search results corresponding to the search query. The one or more images are displayed in order of similarity between each of the images and the search query.
[0171] Fig. 9 An example of another task execution result according to an embodiment of the present disclosure is shown.
[0172] When an image (e.g., an image showing multiple cats) is displayed on the mobile device 2000, the mobile device 2000 may receive a search query (e.g., "largest cat") via a microphone, a virtual keyboard, or a communication interface. The mobile device 2000 may input the search query and the image into a multimodal model including the multimodal search model 100, and may display a bounding box located above the largest cat in the image as a search result corresponding to the search query.
[0173] The multimodal model may be written as a computer-executable program or instructions that may be stored in a medium.
[0174] The medium may store computer executable programs or instructions continuously, or temporarily store computer executable programs or instructions for execution or downloading. In addition, the medium may be any of a variety of recording media or storage media in which a single or multiple pieces of hardware are combined, and the medium is not limited to a medium directly connected to the electronic device 100, but may be distributed on a network. Examples of media include magnetic media (such as hard disks, floppy disks, and tapes), optical recording media (such as CD-ROMs and DVDs), magneto-optical media (such as optical magnetic floppy disks), and ROM, RAM, and flash memory configured to store program instructions. Other examples of media include recording media and storage media managed by application stores that distribute applications or by websites, servers, etc. that supply or distribute various other types of software.
[0175] The multimodal model may be provided in the form of downloadable software. The computer program product may include a product in the form of a software program that is electronically distributed through a manufacturer or an electronic marketplace (e.g., a downloadable application). For electronic distribution, at least a portion of the software program may be stored in a storage medium or may be temporarily generated. In this case, the storage medium may be a storage medium of a server or server 106.
[0176] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the embodiments.
[0177] It is apparent that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0178] Although particular combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may be directly dependent on only one claim, the disclosure of possible implementations includes the combination of each dependent claim with every other claim in the claim set.
[0179] The model related to the above-mentioned neural network can be implemented via a software module. When the model is implemented via a software module (eg, a program module including instructions), the model can be stored in a computer-readable recording medium.
[0180] In addition, the model can be integrated in the form of a hardware chip to become a part of the above-mentioned electronic device 1000. For example, the model can be manufactured in the form of a dedicated hardware chip for artificial intelligence, or can be manufactured as a part of an existing general-purpose processor (e.g., a CPU or an application processor) or a graphics-specific processor (e.g., a GPU).
[0181] In addition, the model may be provided in the form of downloadable software. The computer program product may include a product in the form of a software program that is electronically distributed through a manufacturer or an electronic market (e.g., a downloadable application). For electronic distribution, at least a portion of the software program may be stored in a storage medium or may be temporarily generated. In this case, the storage medium may be a storage medium of a server of the manufacturer or electronic market, or a relay server.
[0182] Although embodiments of the present disclosure have been described with reference to the accompanying drawings, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope defined by the following claims.
Claims
1. An electronic device for operating a multimodal model including a text encoder and a visual encoder, the electronic device comprising: a user interface configured to receive a query; as well as One or more processors configured to: obtaining one or more input images; Inputting the query into the text encoder to obtain text features; Inputting the one or more input images into the visual encoder to obtain image features; as well as outputting a response to the query based on the similarity between the text feature and the image feature, wherein the weight vectors of the text encoder and the visual encoder are pruned and shared according to the shared vectors and pruned vectors generated by the hypernetwork to compress the multimodal model, and The hypernetwork and the multimodal model are jointly trained to minimize at least one of the difference between the weight vectors in the text encoder and the visual encoder, the difference between the weight vectors in different layers of the text encoder, and the number of parameters in the multimodal model.
2. The electronic device according to claim 1, wherein: Any one or any combination of the one or more processors is configured to: Weight sharing and weight pruning are performed on the multi-head self-attention layers of the text encoder and the visual encoder.
3. An electronic device according to claim 2, wherein the weight sharing comprises one or both of cross-modal weight sharing between the text encoder and the visual encoder and block-by-block sharing performed within a block including multiple layers of the text encoder.
4. The electronic device according to claim 1, wherein: Any one or any combination of the one or more processors is configured to: Weight sharing and weight pruning are performed on the feed-forward network layers of the text encoder and the visual encoder.
5. The electronic device of claim 4, wherein the weight sharing comprises one or both of cross-modal weight sharing between the text encoder and the visual encoder and block-by-block sharing performed within a block comprising multiple layers of the text encoder. 6 . The electronic device of claim 1 , wherein the super network is trained to generate the shared vectors and the pruned vectors such that the same weight vector is neither pruned nor shared.
7. The electronic device according to claim 6, wherein: When pruning and sharing the weight vectors of the text encoder and the visual encoder, weight sharing is prioritized over weight pruning.
8. The electronic device according to claim 1, wherein the hypernetwork and the multimodal model are jointly trained in the following manner: After the weight vectors in the visual encoder and the text encoder are pruned and shared, the weight vectors in the visual encoder and the text encoder are frozen, and then the weights in the hypernetwork are updated based on a comparison between an initial number of parameters in the multimodal model and a remaining number of parameters in the multimodal model after weight sharing and weight pruning.
9. The electronic device of claim 1, wherein when the query requires identifying a target image from the one or more input images, any one or combination of the one or more processors is further configured to: selecting an image from the one or more input images that has the greatest similarity to the text feature or has a similarity to the text feature exceeding a predetermined similarity threshold, and The selected image is provided as the response to the query.
10. The electronic device of claim 1, wherein when the query requires identifying a target object from a specific image among the one or more images, any one or combination of the one or more processors is further configured to: identifying in the one or more input images an object having a maximum similarity to the text feature or having a similarity to the text feature exceeding a predetermined similarity threshold, and The identified object within the particular image is visually indicated as the response to the query.
11. A method for performing a multimodal task by using a multimodal model including a text encoder and a visual encoder, the method comprising: obtaining text features from the query by inputting the query into the text encoder; Obtaining image features from the one or more input images by inputting the one or more input images into the visual encoder; as well as outputting a response to the query based on the similarity between the text feature and the image feature, wherein the weight vectors of the text encoder and the visual encoder are pruned and shared according to the shared vectors and pruned vectors generated by the hypernetwork to compress the multimodal model, and The hypernetwork and the multimodal model are jointly trained to minimize at least one of the difference between the weight vectors in the text encoder and the visual encoder, the difference between the weight vectors in different layers of the text encoder, and the number of parameters in the multimodal model.
12. The method according to claim 11, further comprising: Weight sharing and weight pruning are performed on the multi-head self-attention layers of the visual encoder and the text encoder.
13. The method of claim 12, wherein the weight sharing comprises one or both of cross-modal weight sharing between the text encoder and the visual encoder and block-by-block sharing performed within a block comprising multiple layers of the text encoder.
14. The method according to claim 11, further comprising: Weight sharing and weight pruning are performed on the feed-forward network layers of the visual encoder and the text encoder.
15. A non-transitory computer-readable storage medium storing a program executable by at least one processor to perform a method for performing a multimodal task by using a multimodal model including a text encoder and a visual encoder, the method comprising: obtaining text features from the query by inputting the query into the text encoder; Obtaining image features from the one or more input images by inputting the one or more input images into the visual encoder; as well as outputting a response to the query based on the similarity between the text feature and the image feature, wherein the weight vectors of the text encoder and the visual encoder are pruned and shared according to the shared vectors and pruned vectors generated by the hypernetwork to compress the multimodal model, and The hypernetwork and the multimodal model are jointly trained to minimize at least one of the difference between the weight vectors in the text encoder and the visual encoder, the difference between the weight vectors in different layers of the text encoder, and the number of parameters in the multimodal model.