Efficient knowledge distillation method and micro-model generation method and system for multimodal models

By standardizing the image data and prompt information and feature extraction, combined with knowledge distillation optimization and dynamic prompt sampling technology, the problem of insufficient knowledge utilization during knowledge distillation in the existing technology is solved, efficient knowledge transfer and precise target positioning of small models are achieved, and the adaptability and performance of the model are improved.

CN119762943BActive Publication Date: 2025-05-09深圳智眸未来科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510265897.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-09
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The existing technology cannot fully utilize the knowledge of large models during knowledge distillation, resulting in small models having defects in the clarity of model parameters and information volume, and insufficient optimization of data characteristics of different scenarios, resulting in poor adaptability.

Method used

By standardizing the original image data and different types of prompt information and feature extraction, the knowledge distillation optimization method is used to transfer the knowledge of multimodal large models to the small model, and through dynamic prompt sampling technology and lightweight regional suggestions network, the mask prediction ability and target positioning accuracy of the small model are improved.

Benefits of technology

While ensuring the performance of small models, it reduces the calculation and storage costs of the model, avoids information loss and performance degradation, and improves the adaptability and accuracy of small models in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762943B_ABST
    Figure CN119762943B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer vision and the technical field of quantitative models, and specifically relates to an efficient knowledge distillation method and a micro-model generation method and system for a multimodal model, comprising: standardizing, encoding, and fusing original image data to obtain a feature map, and outputting a depth feature under a small model encoder; encoding, splicing, and re-encoding different types of prompt information to obtain a prompt feature; using a small model optimized by knowledge distillation to predict candidate regions for the prompt feature, and dynamically sampling inaccurate regions in the prediction results to generate new prompt features, and combining them with the original prompt features to output an updated mask decoder and a prompt feature combination; using a lightweight region proposal network to generate region proposals, and fusing the region proposals with the prompt features to obtain auxiliary information; combining the depth features, the prompt feature combination, the region proposals, and the auxiliary information to perform mask prediction and generate a binary mask of the region where the target object is located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and quantitative model technology, and specifically relates to an efficient knowledge distillation method and a micro-model generation method and system for a multimodal model. Background Art

[0002] In the development of artificial intelligence computer vision, the application of AI models needs to take into account accuracy, timeliness and storage efficiency. The high cost of large multimodal models limits their direct application to specific downstream tasks, and simply shrinking the model will damage the model's accuracy and generalization ability. Therefore, efficiently distilling large model knowledge into small models becomes the key.

[0003] Traditional knowledge distillation methods mainly focus on the migration of large model knowledge, including static migration of large model parameters and dynamic comparison of the output of large and small models as a means of knowledge verification. However, these methods have not been able to fully utilize the knowledge contained in the large model. Although these knowledge distillation methods ensure the accuracy of small models to a certain extent, in the process of trying to quantify and reduce the model, traditional methods often ignore the clarity of the meaning of model parameters and the amount of information contained in the parameters. Moreover, since model parameters are very sensitive, modifying them may have a significant impact on subsequent use effects. At the same time, these methods lack targeted optimization of data features in different scenarios, resulting in poor adaptability in actual application scenarios. These shortcomings greatly limit the performance and applicability of downstream models. Summary of the invention

[0004] The present invention provides an efficient knowledge distillation method and a micro-model generation method and system for a multimodal model. By processing the original image data and different types of prompt information, the consistency and standardization of the data are ensured, and the corresponding image features and prompt features are extracted. The knowledge distillation optimization method is used to realize the knowledge migration from the multimodal large model to the small model. While ensuring the performance of the small model, the calculation and storage costs of the model are reduced, and the information loss and performance degradation problems in the traditional knowledge distillation method are avoided. The small model is guided to pay attention to inaccurate areas through dynamic prompt sampling technology, the mask prediction ability is improved, and the misjudgment and missed judgment are reduced. A lightweight region proposal network is used to capture the granularity prior information of the data, and the target object is accurately positioned, thereby accurately generating a binary mask prediction.

[0005] An efficient knowledge distillation method and micro-model generation method for a multimodal model, comprising:

[0006] Standardize the original image data to obtain standard image data, encode the standard image data to obtain a feature map, fuse the feature maps with different resolutions, and output the deep features under the small model encoder;

[0007] Encode and concatenate different types of prompt information to form a prompt information matrix, and re-encode based on the prompt information matrix to obtain prompt features;

[0008] The mask output of the large model is used as a supervisory signal to perform knowledge distillation optimization on the mask decoder of the small model. The optimized small model is used to predict candidate regions for the hint features, and new hint features are generated by dynamically sampling the inaccurate regions in the prediction results. The newly generated hint features are combined with the original hint features to output the updated mask decoder and hint feature combination.

[0009] Design and train a lightweight region proposal network to generate region proposals, and fuse the region proposals with the hint features to obtain auxiliary information;

[0010] Mask prediction is performed by combining deep features, hint feature combinations, region proposals and auxiliary information to generate a binary mask of the target object area.

[0011] By processing the original image data and different types of prompt information, the consistency and standardization of the data are ensured, and the corresponding image features and prompt features are extracted; the knowledge distillation optimization method is used to realize the knowledge transfer from the multimodal large model to the small model, while ensuring the performance of the small model, reducing the calculation and storage costs of the model, and avoiding the information loss and performance degradation problems in the traditional knowledge distillation method. The dynamic prompt sampling technology is used to guide the small model to focus on inaccurate areas, improve the mask prediction ability, and reduce misjudgment and missed judgments; a lightweight region proposal network is used to capture the granular prior information of the data, to achieve accurate positioning of the target object, and then accurately generate binary mask predictions.

[0012] Furthermore, the standardization of the original image data to obtain standard image data, encoding based on the standard image data to obtain a feature map, feature fusion of feature maps with different resolutions, and output of deep features under a small model encoder include:

[0013] Get the original image data;

[0014] Based on the original image data, normalization processing is performed to obtain standard image data;

[0015] Based on the standard image data, a convolutional neural network is used for encoding processing to obtain a feature map;

[0016] Based on the original image data, calculate the pixel-level feature distillation loss between the output features of the large model image encoder and the small model image encoder;

[0017] Construct a feature pyramid network, fuse feature maps with different resolutions, and output deep features under the small model encoder.

[0018] By normalizing the original image data, we ensure the consistency and standardization of the data and avoid the performance degradation caused by inconsistent data formats. We use convolutional neural networks for encoding and extracting representative feature maps to improve the adaptability and robustness of the model to different image data, so that we can accurately capture the key information in the image and enhance the overall performance of the model. At the same time, through a reasonable distillation loss function and feature pyramid network, we can effectively transfer the rich knowledge of the large model to the small model, avoiding the information loss and performance degradation problems in traditional knowledge distillation methods.

[0019] Furthermore, the different types of prompt information are encoded and concatenated to form a prompt information matrix, and re-encoding is performed based on the prompt information matrix to obtain prompt features, including:

[0020] Obtain different types of prompt information; the prompt information includes point prompts and box prompts;

[0021] Based on the point prompt, a set of coordinate points is used to represent the coordinate information of the target object in the original image data, and a point prompt set is constructed; and the coordinate mapping of each point prompt in the point prompt set is converted into a point prompt vector, and all the point prompt vectors are encoded to form a point prompt encoding matrix;

[0022] Based on the box hint, the coordinate information of the upper left corner and the lower right corner of the rectangular bounding box corresponding to the target object is represented in the form of a bounding box, and a box hint set is constructed; and the coordinate mapping of each box hint in the box hint set is converted into a box hint vector, and all the box hint vectors are encoded to form a box hint encoding matrix;

[0023] The point prompt coding matrix and the frame prompt coding matrix are concatenated to form a prompt information matrix;

[0024] The cue information matrix is ​​re-encoded using a convolution operation to obtain cue features.

[0025] By encoding different types of prompt information and converting them into the same format to build a point prompt encoding matrix, the performance degradation caused by different formats can be avoided; and convolution operations are used to extract prompt features to provide more informative input for subsequent operations.

[0026] Furthermore, the mask output of the large model is used as a supervisory signal, the mask decoder of the small model is optimized by knowledge distillation, the optimized small model is used to predict the candidate region of the prompt feature, and the inaccurate region in the prediction result is dynamically sampled to generate a new prompt feature, and the newly generated prompt feature is combined with the original prompt feature to output an updated mask decoder and prompt feature combination, including:

[0027] Use the mask output of the large model as the supervision signal of the small model, perform knowledge distillation optimization on the mask decoder of the small model, and calculate the distillation loss of the mask decoder of the small model;

[0028] The optimized small model is used to predict candidate regions for the hint feature, and a dynamic hint sampling strategy is used to dynamically sample inaccurate regions in the small model prediction results to generate new hint features.

[0029] The newly generated cue features are combined with the original cue features to form a new cue combination, which guides the small model to learn mask generation and outputs the updated mask decoder and cue feature combination.

[0030] By using the mask output of the large model as supervision information and adopting dynamic hint sampling technology to guide the small model to focus on inaccurate areas, the mask decoder of the small model can learn more accurately, thereby improving the accuracy and generalization ability of the small model's mask prediction, thereby improving the accuracy of target object detection and segmentation, and significantly reducing misjudgments and missed judgments.

[0031] Furthermore, the lightweight region proposal network is designed and trained to generate region proposals, and the region proposals are fused with the prompt features to obtain auxiliary information, including:

[0032] Designing and training a lightweight region proposal network, generating region proposals, and calculating a lightweight region proposal network loss; the lightweight region proposal network includes a feature pyramid network and a shared detection head;

[0033] The generated region proposals are fused with the encoded hint features to obtain auxiliary information.

[0034] By introducing a lightweight region proposal network and training it on different data, it can capture the granular prior information of the data, thereby providing additional contextual information for the small model. This enables the small model to more accurately locate the target when generating mask predictions and better handle the position and boundaries of the target object, thereby enhancing the adaptability and performance of the small model in actual application scenarios.

[0035] Furthermore, it also includes optimizing and compressing the structure and parameters of the small model, including:

[0036] Select the appropriate network architecture based on the characteristics of the hardware device, use the group convolution technology to group the input feature map and convolution kernel, and perform convolution operations independently in each group to optimize the model structure of the small model;

[0037] The parameters of the small model are quantified according to the statistical information of the small model parameters in each group, and the parameters of the small model are mapped to a low-precision representation. According to the preset pruning rate and threshold, the parameters of the small model are pruned using an amplitude-based pruning method to compress the number of parameters of the small model.

[0038] By selecting a suitable network architecture and using group convolution technology for grouping, we can improve computational efficiency while ensuring feature extraction capabilities; by using parameter quantization and pruning operations, we can greatly reduce storage requirements and computing costs without significantly affecting the performance of the small model, so that the small model can be more efficiently deployed on resource-constrained devices, and the adaptability of the small model is improved by significantly reducing the number of parameters and computational complexity of the small model.

[0039] Furthermore, it also includes:

[0040] Based on the requirements of different scenarios and data characteristics, the small model is trained in a targeted manner, and the parameters of the small model are adjusted to achieve optimization of the small model in different scenarios.

[0041] By conducting targeted training on tasks in different scenarios, the performance of the small model in specific scenarios can be improved, and the parameters of the small model can be adjusted according to actual needs, which improves the flexibility and scalability of the small model. At the same time, it helps the small model to better adapt to different environments and data distributions, enhances the generalization ability of the small model, and extends the effective use cycle of the small model.

[0042] A system of an efficient knowledge distillation method and a micro-model generation method for a multimodal model, comprising:

[0043] An image data processing module is used to standardize the original image data to obtain standard image data, perform encoding processing based on the standard image data to obtain a feature map, perform feature fusion on feature maps with different resolutions, and output deep features under a small model encoder;

[0044] A prompt information processing module is used to encode different types of prompt information to construct a prompt information matrix, and re-encode based on the prompt information matrix to obtain prompt features;

[0045] The small model optimization module is used to use the mask output of the large model as a supervisory signal, perform knowledge distillation optimization on the mask decoder of the small model, use the optimized small model to predict candidate regions for the prompt features, dynamically sample inaccurate regions in the prediction results to generate new prompt features, and combine the newly generated prompt features with the original prompt features to output an updated mask decoder and prompt feature combination;

[0046] A granular prior embedding module, which is used to design and train a lightweight region proposal network to generate region proposals and fuse the region proposals with the hint features to obtain auxiliary information;

[0047] The mask generation module is used to combine deep features, hint feature combinations, region proposals and auxiliary information to generate a binary mask of the target object area.

[0048] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the above method when executing the computer program.

[0049] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.

[0050] The beneficial effects of the present invention are:

[0051] The present invention processes the original image data and different types of prompt information to ensure the consistency and standardization of the data, and extracts the corresponding image features and prompt features; adopts the knowledge distillation optimization method to realize the knowledge transfer from the multimodal large model to the small model, while ensuring the performance of the small model, reducing the calculation and storage costs of the model, and avoiding the information loss and performance degradation problems in the traditional knowledge distillation method, and guides the small model to pay attention to inaccurate areas through dynamic prompt sampling technology, improves the mask prediction ability, and reduces misjudgment and missed judgment; adopts a lightweight region proposal network to capture the granularity prior information of the data, realizes the accurate positioning of the target object, and then accurately generates binary mask predictions. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flow chart of the present invention;

[0053] Figure 2 Schematic diagram of the process of model prediction and optimization in the present invention

[0054] Figure 3 It is a schematic diagram of the system structure of the present invention;

[0055] Figure 4 This is a schematic diagram of the computer equipment structure. DETAILED DESCRIPTION

[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0057] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this device and / or practice this method.

[0058] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples, and a person of ordinary skill in the art can understand the specific meanings of the above terms in this application under specific circumstances.

[0059] Example 1

[0060] Figure 1 The method shown is an efficient knowledge distillation method and micro-model generation method for a multimodal model. By processing the original image data and different types of prompt information, the consistency and standardization of the data are ensured, and the corresponding image features and prompt features are extracted. The knowledge distillation optimization method is used to realize the knowledge transfer from the multimodal large model to the small model. While ensuring the performance of the small model, the calculation and storage costs of the model are reduced, and the information loss and performance degradation problems in the traditional knowledge distillation method are avoided. The dynamic prompt sampling technology is used to guide the small model to focus on inaccurate areas, improve the mask prediction ability, and reduce misjudgment and missed judgments. A lightweight region proposal network is used to capture the granularity prior information of the data, to accurately locate the target object, and then accurately generate binary mask predictions. The specific steps include the following:

[0061] S1: Standardize the original image data to obtain standard image data, encode the standard image data to obtain a feature map, fuse the feature maps with different resolutions, and output the deep features under the small model encoder;

[0062] S11: Get original image data ;

[0063] S12: Based on original image data , perform normalization processing to obtain standard image data ;

[0064] The normalized expression is:

[0065] ;

[0066] In the formula, Represents standard image data, whose pixel value data range is ; Represents the original image data; Indicates the minimum value; Indicates the maximum value.

[0067] S13: Based on standard image data , the convolutional neural network CNN is used for encoding processing to obtain the feature map ;

[0068] Get feature map The expression is:

[0069] ;

[0070] In the formula, Represents a feature map; Represents a convolutional neural network.

[0071] In this embodiment, when a convolutional neural network is used for encoding processing, the convolutional neural network includes a convolution layer and a pooling layer, and its expression is:

[0072] ;

[0073] In the formula, Indicates Layer feature map, position The characteristic value of represents the activation function; Indicates the size of the convolution kernel; Indicates that the convolution kernel is at position The weight value on ; represents the bias term;

[0074] In this embodiment, the activation function selects the ReLUctant function, and its expression is: .

[0075] S14: Based on original image data , calculate the pixel-level feature distillation loss between the output features of the large model image encoder and the small model image encoder ;

[0076] Among them, the mean square error As the loss function, calculate the large model image encoder and small model image encoder Pixel-level feature distillation loss between output features , whose expression is:

[0077] ;

[0078] In the formula, represents the pixel-level feature distillation loss; represents the mean square error function; represents a large model image encoder; represents a small model image encoder;

[0079] S15: Construct a feature pyramid network to fuse feature maps with different resolutions and output deep features under the small model encoder ;

[0080] Due to the large model image encoder and small model image encoder There is inconsistency in the downsampling step size and number of channels of the feature map. A feature pyramid network FPN is constructed to process the feature maps with different resolutions and fuse them to make the small model image encoder Learning Large Model Image Encoders The complex feature representation of the output is the updated small model image encoder The deep features .

[0081] Among them, the feature pyramid network FPN is a neural network structure used for computer vision tasks such as target detection and semantic segmentation. It mainly uses the representation of image features at different scales to generate multi-scale feature maps to better handle the scale change problem of the target object in the image. The feature pyramid network FPN includes the bottom-up use of convolutional neural networks to extract features from the input image data and generate a series of feature maps with gradually decreasing resolutions, and the top-down fusion of low-level feature maps by upsampling from high-level feature maps layer by layer, which can propagate high-level semantic information to low-level feature maps while retaining the spatial detail information of low-level features. The expression for the fusion operation is:

[0082] ;

[0083] In the formula, Represents the feature map after the fusion of feature maps of different resolutions; Represents an upsampling operation; Represents element addition operation;

[0084] S2: Encode and concatenate different types of prompt information to form a prompt information matrix , and based on the prompt information matrix Re-encoding process to obtain prompt features ;

[0085] S21: Get different types of prompt information ; The prompt information includes point prompts and box prompts;

[0086] S22: Based on the point prompt, a set of coordinate points is used to represent the coordinate information of the target object in the original image data, and a point prompt set is constructed. ; and point the prompt collection The coordinates of each point cue in the map are converted into a point cue vector , and encode all point cue vectors to form a point cue encoding matrix ;

[0087] S221: Using a set of point prompt coordinates to represent the coordinate information of the target object in the original image data, constructing a point prompt set ;

[0088] Among them, click prompt collection The expression is:

[0089] ;

[0090] In the formula, Indicates a point prompt collection; Indicates The coordinates of the points in the original image are shown. , Respectively represent The points indicate the horizontal and vertical coordinates in the original image. ;

[0091] It should be noted that these point prompts can be the center of the region of interest manually marked by the user, or can be the location information that may contain the target object preliminarily detected by some algorithms.

[0092] S222: Point prompt collection The coordinates of each point cue in the map are converted into a point cue vector ;

[0093] Among them, based on the point prompt set The coordinates of each point in the prompt , normalized to interval, and then use linear mapping to convert it into dimensional vector. Point hint vector The expression is:

[0094] ;

[0095] In the formula, Indicates The point hint vector of the point hints; Indicates the width of the original image; Indicates the height of the original image; express dimensional space;

[0096] S223: Encode all point hint vectors to form a point hint encoding matrix ;

[0097] Among them, the point prompt encoding matrix The expression is:

[0098] ;

[0099] In the formula, Representation point cue encoding matrix; Indicates The point hint vector of the point hints; express dimensional space;

[0100] S23: Based on the box hint, the coordinate information of the upper left corner and lower right corner of the rectangular bounding box corresponding to the target object is represented in the form of a bounding box, and a box hint set is constructed. ; and convert the coordinate map of each box hint in the box hint set into a box hint vector , and encode all box cue vectors to form a box cue encoding matrix ;

[0101] S231: Using a bounding box format to represent the coordinate information of the upper left corner and the lower right corner of the rectangular bounding box corresponding to the target object, and constructing a box prompt set;

[0102] Among them, the box prompt set The expression is:

[0103] ;

[0104] In the formula, Represents a collection of box prompts; Indicates The boxes indicate the coordinates in the original image. , Respectively represent The horizontal and vertical coordinates of the upper left corner of the rectangular bounding box corresponding to the box prompt, , Respectively represent The horizontal and vertical coordinates of the upper right corner of the rectangular bounding box corresponding to the box prompt, ;

[0105] In this embodiment, the frame prompt set is used to determine the approximate area range of the target object.

[0106] S232: Collect the frame prompts The coordinate map of each box hint in is converted to a box hint vector ;

[0107] Among them, based on the box prompt set The coordinates of each box tip in , normalized to interval, and then use linear mapping to convert it into dimensional vector.

[0108] Box Tip Vector The expression is:

[0109] ;

[0110] In the formula, Indicates box hint vector for box hints;

[0111] S233: Encode all box hint vectors to form a box hint encoding matrix ;

[0112] Among them, the box prompt encoding matrix The expression is:

[0113] ;

[0114] In the formula, Represents the box prompt encoding matrix; Indicates box hint vector for box hints; express dimensional space;

[0115] S24: Point prompt encoding matrix and box prompt encoding matrix Perform splicing to form a prompt information matrix ;

[0116] Among them, the prompt information matrix The expression is:

[0117] ;

[0118] In the formula, Represents the prompt information matrix; express dimensional space;

[0119] S25: Use convolution operation to convert the prompt information matrix Re-encoding process to obtain prompt features ;

[0120] Among them, the prompt feature The expression is:

[0121] ;

[0122] In the formula, Represents the feature vector corresponding to the prompt information; Represents the convolution operation;

[0123] It should be noted that different types of prompt information will be converted into corresponding feature vector forms, thereby providing more informative input for subsequent operations.

[0124] S3: Use the mask output of the large model as a supervisory signal, perform knowledge distillation optimization on the mask decoder of the small model, use the optimized small model to predict candidate regions for the hint feature, and dynamically sample the inaccurate regions in the prediction results to generate new hint features , and the newly generated prompt feature With the original prompt feature Combine and output the updated mask decoder and hint feature combination ;

[0125] S31: Use the mask output of the large model as the supervision signal of the small model, perform knowledge distillation optimization on the mask decoder of the small model, and calculate the distillation loss of the mask decoder of the small model ;

[0126] In this embodiment, when calculating the distillation loss of the mask decoder of the small model, a combination of Dice loss and BCE loss is used, that is, the expression of the distillation loss of the mask decoder of the small model is:

[0127] ;

[0128] In the formula, represents the distillation loss of the mask decoder; represents the mask loss function; represents a binary threshold processing function; Representing the characteristics of a large model image encoder; Represents the characteristics of the small model image encoder; Indicates a cue point; Indicates the mask shared by the large model and the small model; Indicates IoU marking;

[0129] S32: Use the optimized small model to predict the candidate regions of the prompt features, and use the dynamic prompt sampling strategy to dynamically sample the inaccurate regions in the prediction results of the small model to generate new prompt features ;

[0130] Among them, the new prompt feature The expression is:

[0131] ;

[0132] In the formula, Indicates the new cue point generated; represents the sampling function; Represents the mask prediction of the large model; represents the mask prediction of the small model;

[0133] S33: The newly generated prompt feature With the original prompt feature Combined to form a new prompt combination, guide the small model to learn mask generation, and output the updated mask decoder and prompt feature combination ;

[0134] In this embodiment, the new prompt combination is used in the decoding process of the next iteration to guide the small model to better learn mask generation, and then output the updated mask decoder and prompt feature combination .

[0135] S4: Design and train a lightweight region proposal network to generate region proposals, and fuse the region proposals with the hint features to obtain auxiliary information;

[0136] S41: Design and train a lightweight region proposal network RPN to generate region proposals while keeping the avatar encoder of the small model frozen , and calculate the lightweight region proposal network loss ;

[0137] In this embodiment, the lightweight region proposal network RPN is a lightweight network used to generate candidate region proposals, thereby providing high-quality candidate regions for subsequent classification and regression networks in target detection tasks, reducing the number of candidate boxes that the model needs to process, and improving computational efficiency. Its workflow includes: using a convolutional layer to extract a feature map of the input image data, generating a set of fixed-size candidate regions on the feature map, calculating the confidence of each candidate region, and determining whether each candidate region contains the target object, and then regressing the boundaries of the candidate regions to more accurately locate the target object. The lightweight region proposal network RPN includes a feature pyramid network FPN and a shared detection head.

[0138] Among them, the lightweight region proposal network loss The expression is:

[0139] ;

[0140] In the formula, represents the lightweight region proposal network loss; represents the Smooth L1 loss function, which is used to regress the location of the region proposal; represents the predicted region suggestion; represents the real region suggestion; represents the cross entropy loss function, which is used to classify the categories of region proposals; Represents the true category label; Represents the predicted category label;

[0141] S42: Propose the generated region and the encoded cue features Fusion to obtain auxiliary information ;

[0142] Among them, auxiliary information The expression is:

[0143] ;

[0144] In the formula, Indicates auxiliary information; represents the fusion function; Indicates regional recommendations;

[0145] S5: Combining deep features , prompt feature combination , Regional Recommendations and auxiliary information , perform mask prediction and generate a binary mask of the target object area ;

[0146] In this embodiment, the decoder network of the small model is used to generate a binary mask for mask prediction based on the input information. , whose expression is:

[0147] ;

[0148] In the formula, represents mask prediction, , Represents the size of the target object in the original image data; represents the decoder network; Represents the updated prompt collection.

[0149] In this embodiment, the structure and parameters of the small model are also optimized and compressed, including:

[0150] S61: Select the appropriate network architecture based on the characteristics of the hardware device, use the group convolution technology to group the input feature map and convolution kernel, and perform convolution operations independently in each group to optimize the model structure of the small model;

[0151] In this embodiment, when computing resources are limited, the CNN-based RepViT-M1 network architecture can be used. In practical applications, a suitable network price can be selected according to the hardware equipment.

[0152] Among them, the process of optimizing the structure of the small model using group convolution technology is expressed as:

[0153] ;

[0154] In the formula, represents the output feature map, , Represents the output feature map; Indicates the number of channels of the input feature map, Indicates the number of channels of the input feature map; represents the grouped convolution function; Represents the input feature map; Represents the grouped part of the convolution kernel; Indicates grouping, , Indicates the number of groups; represents the convolution function;

[0155] The input feature map is converted into and convolution kernel Divide The feature maps and convolution kernels of different groups are grouped so that convolution operations are performed only within the group, which reduces the number of convolution parameters from the original Reduce to , Indicates the size of the convolution kernel, which helps to reduce the computational complexity of the small model and improve the computational efficiency of the small model.

[0156] S62: quantizing the parameters of the small model according to the statistical information of the parameters of the small model in each group, mapping the parameters of the small model to a low-precision representation, and pruning the parameters of the small model using an amplitude-based pruning method according to a preset pruning rate and threshold, so as to compress the number of parameters of the small model;

[0157] S621: quantizing the parameters of the small models according to the statistical information of the parameters of the small models in each group, and mapping the parameters of the small models to low-precision representations;

[0158] Among them, the quantization operation can be expressed as:

[0159] ;

[0160] In the formula, Represents the low-precision parameters after quantization; Represents a quantization operation; Represents the original parameters; Indicates quantitative scale; Indicates zero offset;

[0161] In the actual quantization process, the grouping advantage brought by group convolution is used to quantize the parameters of different groups separately, so as to better retain the characteristics of the parameters of different groups, so that the quantized small models can better perform on different hardware platforms. That is, the quantization operation within each group can be expressed as:

[0162] ;

[0163] In the formula, Indicates The parameters after group quantization; Indicates The original parameters of the group; Indicates The quantitative scale of the group, i.e. , Indicates The maximum value of the group's parameters, Indicates The minimum value of the group's parameter, The number of bits representing the target low-precision data type; Indicates Zero offset of the group;

[0164] At the same time, in the quantization calculation process, for matrix multiplication , the quantized calculation can be expressed as:

[0165] ;

[0166] In the formula, Represents the quantized output feature map; Represents the input feature map Quantitative scale of Represents the input feature map Zero point offset; Represents the input feature map Quantitative scale of Represents the input feature map Zero point offset;

[0167] Among them, the quantization calculation uses sparse matrix calculation, and the sparse matrix multiplication can be expressed as:

[0168] ;

[0169] In the formula, Represents a sparse matrix and sparse matrices The result matrix obtained by multiplication; Represents a sparse matrix Middle Row, No. Column data, and represents non-zero elements, , Represents a sparse matrix The total number of rows and columns; Indicates the index quantity. , Indicates the total number of indexes; Represents a sparse matrix Middle Row, No. Column data;

[0170] These optimization methods can accelerate the model's calculation process, allowing small models to achieve fast reasoning on resource-constrained devices and meet real-time requirements.

[0171] S622: According to the preset pruning rate and threshold , the parameters of the small model are pruned using an amplitude-based pruning method to compress the number of parameters of the small model;

[0172] Among them, the parameters of the small model are pruned based on the amplitude-based pruning method, and the parameters after pruning are:

[0173] ;

[0174] In the formula, represents the parameters of the small model; Represents the small model parameters The amplitude measure of The absolute value of

[0175] By adopting different training methods, the information content of model parameters is increased while reducing their sensitivity, so that parameter pruning can be carried out efficiently without damaging model performance.

[0176] In this embodiment, it also includes targeted training of the small model based on different scene requirements and data characteristics, and adjusting the parameters of the small model to achieve optimization of the small model in different scenes. Through targeted training of different scene tasks, the performance of the small model in a specific scene can be improved, and the parameters of the small model can be adjusted according to actual needs, which improves the flexibility and scalability of the small model; at the same time, it is convenient for the small model to better adapt to different environments and data distributions, enhance the generalization ability of the small model, and extend the effective use period of the small model.

[0177] In practical applications, for the target object detection task in the factory scene, targeted training is carried out using image data from scenes such as factory assembly lines and real-time monitoring. During the training process, the model can optimize its own structure and parameters based on data characteristics such as object shape, lighting conditions, and background, so that the performance of the small model in a specific scene can be enhanced, and features can be extracted more accurately to make more accurate predictions. At the same time, the parameter quantization and parameter pruning strategies can be adjusted more specifically according to the characteristics of the scene data to ensure that the optimized model can achieve the best performance in this scene.

[0178] Figure 2 Schematic diagram of the model prediction and optimization process shown.

[0179] Example 2

[0180] Based on the same technical concept, Figure 3 As shown, this embodiment also provides an efficient knowledge distillation method and micro-model generation system for a multimodal model, including an image data processing module, a prompt information processing module, a small model optimization module, a granularity prior embedding module, and a mask generation module.

[0181] Specifically, the image data processing module is used to standardize the original image data to obtain standard image data, encode the standard image data to obtain a feature map, fuse the feature maps with different resolutions, and output the deep features under the small model encoder;

[0182] Specifically, the prompt information processing module is used to encode different types of prompt information to construct a prompt information matrix, and re-encode based on the prompt information matrix to obtain prompt features;

[0183] Specifically, the small model optimization module is used to use the mask output of the large model as a supervisory signal, perform knowledge distillation optimization on the mask decoder of the small model, use the optimized small model to predict candidate regions for the prompt features, dynamically sample inaccurate regions in the prediction results to generate new prompt features, and combine the newly generated prompt features with the original prompt features to output an updated mask decoder and prompt feature combination;

[0184] Specifically, the granular prior embedding module is used to design and train a lightweight region proposal network to generate region proposals, and fuse the region proposals with the prompt features to obtain auxiliary information;

[0185] Specifically, the mask generation module is used to combine deep features, hint feature combinations, region proposals and auxiliary information to generate a binary mask of the target object area.

[0186] Example 3

[0187] Based on the same technical concept, the embodiment of the present application also provides a computer device, including a memory 1 and a processor 2, such as Figure 4 As shown, the memory 1 stores a computer program, and the processor 2 implements any of the above-mentioned methods when executing the computer program.

[0188] Among them, the memory 1 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 1 can be an internal storage unit of the efficient knowledge distillation method and micro-model generation system for multimodal models, such as a hard disk. In other embodiments, the memory 1 can also be an external storage device of the efficient knowledge distillation method and micro-model generation system for multimodal models, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Further, the memory 1 can also include both the internal storage unit and the external storage device of the efficient knowledge distillation method and micro-model generation system for multimodal models. The memory 1 can not only be used to store application software and various types of data installed in the efficient knowledge distillation method and micro-model generation system for multimodal models, such as the code of the efficient knowledge distillation method and micro-model generation system program for multimodal models, but can also be used to temporarily store data that has been output or is to be output.

[0189] In some embodiments, the processor 2 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor or other data processing chip, which is used to run the program code or process data stored in the memory 1, such as executing an efficient knowledge distillation method and a micro-model generation system program for a multimodal model.

[0190] The disclosed embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to perform the steps of the method described in the above method embodiment. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0191] The computer program product of the application page content refresh method provided in the disclosed embodiment of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the steps of the method described in the above method embodiment. For details, please refer to the above method embodiment, which will not be repeated here.

[0192] The disclosed embodiment of the present invention further provides a computer program, which implements any one of the methods of the aforementioned embodiments when executed by a processor. The computer program product can be implemented in hardware, software, or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium, and in another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0193] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.

[0194] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" refers to at least two.

[0195] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention belong.

[0196] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0197] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0198] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0199] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0200] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0201] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. An efficient knowledge distillation method and micro-model generation method for a multimodal model, characterized in that: include: Standardize the original image data to obtain standard image data, encode the standard image data to obtain a feature map, fuse the feature maps with different resolutions, and output the deep features under the small model encoder; Encode and concatenate different types of prompt information to form a prompt information matrix, and re-encode based on the prompt information matrix to obtain prompt features; The mask output of the large model is used as a supervisory signal to perform knowledge distillation optimization on the mask decoder of the small model. The optimized small model is used to predict candidate regions for the hint features, and new hint features are generated by dynamically sampling the inaccurate regions in the prediction results. The newly generated hint features are combined with the original hint features to output the updated mask decoder and hint feature combination. Design and train a lightweight region proposal network to generate region proposals, and fuse the region proposals with the hint features to obtain auxiliary information; Combine deep features, hint feature combinations, region proposals, and auxiliary information to perform mask prediction and generate a binary mask of the target object area. The mask output of the large model is used as a supervisory signal, the mask decoder of the small model is optimized by knowledge distillation, the optimized small model is used to predict the candidate region of the prompt feature, and the inaccurate region in the prediction result is dynamically sampled to generate a new prompt feature, and the newly generated prompt feature is combined with the original prompt feature to output an updated mask decoder and prompt feature combination, including: Use the mask output of the large model as the supervision signal of the small model, perform knowledge distillation optimization on the mask decoder of the small model, and calculate the distillation loss of the mask decoder of the small model; The optimized small model is used to predict candidate regions for the hint feature, and a dynamic hint sampling strategy is used to dynamically sample inaccurate regions in the small model prediction results to generate new hint features. Combine the newly generated hint features with the original hint features to form a new hint combination, guide the small model to learn mask generation, and output the updated mask decoder and hint feature combination; It also includes optimization and compression of the structure and parameters of the small model, including: Select the appropriate network architecture based on the characteristics of the hardware device, use the group convolution technology to group the input feature map and convolution kernel, and perform convolution operations independently in each group to optimize the model structure of the small model; The parameters of the small model are quantified according to the statistical information of the small model parameters in each group, and the parameters of the small model are mapped to a low-precision representation. According to the preset pruning rate and threshold, the parameters of the small model are pruned using an amplitude-based pruning method to compress the number of parameters of the small model.

2. According to claim 1, the efficient knowledge distillation method and micro-model generation method of a multimodal model is characterized in that: The method of normalizing the original image data to obtain standard image data, encoding the standard image data to obtain a feature map, fusing the feature maps with different resolutions, and outputting the deep features under the small model encoder includes: Get the original image data; Based on the original image data, normalization processing is performed to obtain standard image data; Based on the standard image data, a convolutional neural network is used for encoding processing to obtain a feature map; Based on the original image data, calculate the pixel-level feature distillation loss between the output features of the large model image encoder and the small model image encoder; Construct a feature pyramid network, fuse feature maps with different resolutions, and output deep features under the small model encoder.

3. The efficient knowledge distillation method and micro-model generation method of a multimodal model according to claim 1, characterized in that: The encoding and splicing of different types of prompt information to form a prompt information matrix, and re-encoding based on the prompt information matrix to obtain prompt features include: Obtain different types of prompt information; the prompt information includes point prompts and box prompts; Based on the point prompt, a set of coordinate points is used to represent the coordinate information of the target object in the original image data, and a point prompt set is constructed; and the coordinate mapping of each point prompt in the point prompt set is converted into a point prompt vector, and all the point prompt vectors are encoded to form a point prompt encoding matrix; Based on the box hint, the coordinate information of the upper left corner and the lower right corner of the rectangular bounding box corresponding to the target object is represented in the form of a bounding box, and a box hint set is constructed; and the coordinate mapping of each box hint in the box hint set is converted into a box hint vector, and all the box hint vectors are encoded to form a box hint encoding matrix; The point prompt coding matrix and the frame prompt coding matrix are concatenated to form a prompt information matrix; The cue information matrix is ​​re-encoded using a convolution operation to obtain cue features.

4. The efficient knowledge distillation method and micro-model generation method of a multimodal model according to claim 1, characterized in that: The lightweight region proposal network is designed and trained to generate region proposals, and the region proposals are fused with the prompt features to obtain auxiliary information, including: Designing and training a lightweight region proposal network, generating region proposals, and calculating a lightweight region proposal network loss; the lightweight region proposal network includes a feature pyramid network and a shared detection head; The generated region proposals are fused with the encoded hint features to obtain auxiliary information.

5. The efficient knowledge distillation method and micro-model generation method of a multimodal model according to claim 1, characterized in that: Also includes: Based on the requirements of different scenarios and data characteristics, the small model is trained in a targeted manner, and the parameters of the small model are adjusted to achieve optimization of the small model in different scenarios.

6. A system for an efficient knowledge distillation method and a micro-model generation method for the multimodal model of claim 1, characterized in that: include: An image data processing module is used to standardize the original image data to obtain standard image data, perform encoding processing based on the standard image data to obtain a feature map, perform feature fusion on feature maps with different resolutions, and output deep features under a small model encoder; A prompt information processing module is used to encode different types of prompt information to construct a prompt information matrix, and re-encode based on the prompt information matrix to obtain prompt features; The small model optimization module is used to use the mask output of the large model as a supervisory signal, perform knowledge distillation optimization on the mask decoder of the small model, use the optimized small model to predict candidate regions for the prompt features, dynamically sample inaccurate regions in the prediction results to generate new prompt features, and combine the newly generated prompt features with the original prompt features to output an updated mask decoder and prompt feature combination; A granular prior embedding module, which is used to design and train a lightweight region proposal network to generate region proposals and fuse the region proposals with the hint features to obtain auxiliary information; The mask generation module is used to combine deep features, hint feature combinations, region proposals and auxiliary information to generate a binary mask of the target object area.

7. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Model distillation method, target detection method and related equipment

    CN115273136A

  • Self-distillation training method and device for convolutional neural network, and scalable dynamic prediction method

    WO2021023202A1