Feature extraction method and device, electronic equipment, storage medium and program product

CN122821150APending Publication Date: 2026-09-25MOORE THREADS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610975654.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

虽然这种一致的分辨率处理方式能够实现图像的统一输入,但它也限制了大型视觉语言模型在不同尺度下捕捉信息的能力,容易导致大量细节的丢失

Benefits of technology

[0022]在本公开实施例中,可在大型视觉语言模型中设置包括编码器和裁切器的特征提取模块,可通过编码器对待处理图像进行特征提取,得到初始特征图,该初始特征图的分辨率与所述待处理图像的分辨率相同,该初始特征图的尺寸是所述视觉语言预训练模型要求的输入尺寸的整数倍;通过裁切器对所述初始特征图进行切分处理,得到至少一个目标特征图,所述目标特征图的尺寸符合所述视觉语言预训练模型要求的输入尺寸。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821150A_ABST
    Figure CN122821150A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a feature extraction method and device, electronic equipment, storage medium and program product, the method comprising: performing feature extraction on a to-be-processed image by an encoder to obtain an initial feature map, the initial feature map having the same resolution as that of the to-be-processed image, and the size of the initial feature map being an integer multiple of the input size required by a visual language pre-training model; and performing cutting processing on the initial feature map by a cutter to obtain at least one target feature map, the target feature map having a size conforming to the input size required by the visual language pre-training model. The embodiments of the present disclosure can directly extract multi-scale features of the to-be-processed image, ensuring that the complete information of image details can be effectively extracted and applied to a large visual language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a feature extraction method and apparatus, electronic device, storage medium and program product. Background Technology

[0002] In the field of artificial intelligence, Large Vision Language Models (LVLMs) have unlocked more and more application scenarios by leveraging the powerful text processing capabilities of Large Language Models (LLMs) and further combining them with visual information parsing technology.

[0003] However, in related technologies, large visual language models typically encode input images at a fixed resolution. This approach is often achieved by reducing or increasing the sampling rate, or by using a "scale-then-paste" method. While this consistent resolution processing enables uniform input images, it also limits the ability of large visual language models to capture information at different scales, easily leading to the loss of a significant amount of detail. Summary of the Invention

[0004] This disclosure presents a feature extraction method and apparatus, electronic device, storage medium, and program product.

[0005] According to one aspect of this disclosure, a feature extraction method is provided, the method being applied to a feature extraction module in a large visual language model, the feature extraction module including an encoder and a cropper, the large visual language model further including a visual language pre-training model connected to the feature extraction module, the method comprising: acquiring an image to be processed; performing feature extraction on the image to be processed by the encoder to obtain an initial feature map, the resolution of the initial feature map being the same as the resolution of the image to be processed, and the size of the initial feature map being an integer multiple of the input size required by the visual language pre-training model; and performing segmentation processing on the initial feature map by the cropper to obtain at least one target feature map, the size of the target feature map conforming to the input size required by the visual language pre-training model.

[0006] In one possible implementation, the encoder includes a first convolutional layer and a second convolutional layer. The encoder extracts features from the image to be processed to obtain an initial feature map, comprising: extracting features from the image to be processed in a first dimension using the first convolutional layer to obtain a first feature map, wherein the size of the first feature map in the first dimension is an integer multiple of the input size required by the visual language pre-training model; and extracting features from the first feature map in a second dimension using the second convolutional layer to obtain the initial feature map.

[0007] In one possible implementation, the encoder further includes a first normalization layer and a second normalization layer. The encoder extracts features from the image to be processed to obtain an initial feature map, including: extracting features from the image to be processed in a first dimension using the first convolutional layer to obtain a first feature map; performing a first normalization process on the first feature map using the first normalization layer to obtain a second feature map; extracting features from the second feature map in a second dimension using the second convolutional layer to obtain a third feature map; and performing a second normalization process on the third feature map using the second normalization layer to obtain the initial feature map.

[0008] In one possible implementation, the size of the convolutional kernel in the first convolutional layer is determined based on a first size and a first input size, wherein the first size is the size of the image to be processed in the first dimension, and the first input size is the input size required by the visual language pre-training model in the first dimension; the size of the convolutional kernel in the second convolutional layer is determined based on a second size and a second input size, wherein the second size is the size of the image to be processed in the second dimension, and the second input size is the input size required by the visual language pre-training model in the second dimension.

[0009] In one possible implementation, both the convolutional kernels in the first convolutional layer and the convolutional kernels in the second convolutional layer perform convolution operations with no padding and a stride of 1.

[0010] In one possible implementation, the initial feature map is segmented by the cropper to obtain at least one target feature map, including: segmenting the initial feature map according to the input size required by the visual language pre-training model to obtain at least one first sub-map; sampling the image to be processed according to the input size required by the visual language pre-training model to obtain a sampled image; and aggregating the at least one first sub-map and the sampled image to obtain at least one target feature map.

[0011] In one possible implementation, the large visual language model further includes a fusion module and a large language model. The training method of the large visual language model includes: training the initial large visual language model with the parameters of the visual language pre-training model, the fusion module, and the large language model fixed to obtain a first large visual language model; training the first large visual language model with the parameters of the feature extraction module fixed to obtain a second large visual language model; and training the second large visual language model with all parts of the second large visual language model participating in the training to obtain a trained large visual language model.

[0012] According to one aspect of this disclosure, a feature extraction apparatus is provided, the apparatus being applied to a feature extraction module in a large visual language model, the feature extraction module including an encoder and a cropper, the large visual language model further including a visual language pre-training model connected to the feature extraction module, the apparatus comprising: an acquisition unit for acquiring an image to be processed; a feature extraction unit for extracting features from the image to be processed using the encoder to obtain an initial feature map, the resolution of the initial feature map being the same as the resolution of the image to be processed, and the size of the initial feature map being an integer multiple of the input size required by the visual language pre-training model; and a segmentation unit for segmenting the initial feature map using the cropper to obtain at least one target feature map, the size of the target feature map conforming to the input size required by the visual language pre-training model.

[0013] In one possible implementation, the encoder includes a first convolutional layer and a second convolutional layer. The feature extraction unit is configured to: extract features from the image to be processed through the first convolutional layer in a first dimension to obtain a first feature map, wherein the size of the first feature map in the first dimension is an integer multiple of the input size required by the visual language pre-training model; and extract features from the first feature map through the second convolutional layer in a second dimension to obtain the initial feature map.

[0014] In one possible implementation, the encoder further includes a first normalization layer and a second normalization layer. The feature extraction unit is configured to: extract features from the image to be processed using the first convolutional layer in a first dimension to obtain a first feature map; perform a first normalization process on the first feature map using the first normalization layer to obtain a second feature map; extract features from the second feature map using the second convolutional layer in a second dimension to obtain a third feature map; and perform a second normalization process on the third feature map using the second normalization layer to obtain the initial feature map.

[0015] In one possible implementation, the size of the convolutional kernel in the first convolutional layer is determined based on a first size and a first input size, wherein the first size is the size of the image to be processed in the first dimension, and the first input size is the input size required by the visual language pre-training model in the first dimension; the size of the convolutional kernel in the second convolutional layer is determined based on a second size and a second input size, wherein the second size is the size of the image to be processed in the second dimension, and the second input size is the input size required by the visual language pre-training model in the second dimension.

[0016] In one possible implementation, both the convolutional kernels in the first convolutional layer and the convolutional kernels in the second convolutional layer perform convolution operations with no padding and a stride of 1.

[0017] In one possible implementation, the segmentation unit is used to: segment the initial feature map according to the input size required by the visual language pre-training model to obtain at least one first sub-map; sample the image to be processed according to the input size required by the visual language pre-training model to obtain a sampled image; and aggregate the at least one first sub-map and the sampled image to obtain at least one target feature map.

[0018] In one possible implementation, the large visual language model further includes a fusion module and a large language model, and the device further includes a training unit, which is configured to: train the initial large visual language model with the parameters of the visual language pre-training model, the fusion module, and the large language model fixed, to obtain a first large visual language model; train the first large visual language model with the parameters of the feature extraction module fixed, to obtain a second large visual language model; and train the second large visual language model with all parts of the second large visual language model participating in the training, to obtain a trained large visual language model.

[0019] According to one aspect of this disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the method described above.

[0020] According to one aspect of this disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the above-described method.

[0021] According to one aspect of this disclosure, a computer program product is provided, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program implements the above-described method when executed by a processor.

[0022] In this embodiment of the disclosure, a feature extraction module including an encoder and a cropper can be set in a large visual language model. The encoder can extract features from the image to be processed to obtain an initial feature map. The resolution of the initial feature map is the same as that of the image to be processed, and the size of the initial feature map is an integer multiple of the input size required by the visual language pre-training model. The cropper performs segmentation processing on the initial feature map to obtain at least one target feature map. The size of the target feature map conforms to the input size required by the visual language pre-training model.

[0023] In this way, the loss of image details and the introduction of useless information caused by image resolution processing (such as scaling) in related technologies can be reduced or avoided. The feature extraction module can directly extract multi-scale features of the image to be processed, ensuring that the complete information of image details can be effectively extracted and applied to large-scale visual language models.

[0024] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.

[0026] Figure 1 A flowchart illustrating the feature extraction method provided in an embodiment of this disclosure is shown.

[0027] Figure 2 A schematic diagram of the structure of a large visual language model provided in an embodiment of this disclosure is shown.

[0028] Figure 3 This diagram illustrates the structure of a feature extraction module provided in an embodiment of the present disclosure.

[0029] Figure 4 This diagram illustrates the feature extraction process performed by the first convolutional layer according to an embodiment of the present disclosure.

[0030] Figure 5 This diagram illustrates the feature extraction process performed by the second convolutional layer according to an embodiment of the present disclosure.

[0031] Figure 6A block diagram of a feature extraction apparatus provided in an embodiment of this disclosure is shown.

[0032] Figure 7 This diagram illustrates a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0033] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0034] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0035] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0036] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0037] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions.

[0038] Large visual language models in related technologies (such as LLaVA-OV, InternVL2, and QwenVL2) employ image processing schemes based on dynamic resolution. These large visual language models scale images to multiple fixed-resolution sizes, cut the scaled images into multiple patches, and attach corresponding thumbnails. Each patch maintains a fixed resolution size before entering the visual encoder. During the image scaling process, padding or interpolation methods are used to address the problems caused by fixed resolution. However, while these methods avoid forcibly reducing the size of high-resolution images to minimize the loss of key information, the uncertainty of image size often leads to the introduction of useless information or the loss of effective details through padding or interpolation. For example, interpolation methods may cause blurring of image details, especially when processing high-resolution images, resulting in severe detail loss; while padding may produce unnatural image edges, affecting the accuracy of feature extraction.

[0039] In view of this, embodiments of the present disclosure provide a feature extraction method. A feature extraction module including an encoder and a cropper can be set in a large visual language model. The encoder can extract features from the image to be processed to obtain an initial feature map. The resolution of the initial feature map is the same as the resolution of the image to be processed, and the size of the initial feature map is an integer multiple of the input size required by the visual language pre-training model. The cropper performs segmentation processing on the initial feature map to obtain at least one target feature map. The size of the target feature map conforms to the input size required by the visual language pre-training model.

[0040] In this way, the loss of image details and the introduction of useless information caused by image resolution processing (such as scaling) in related technologies can be reduced or avoided. The feature extraction module can directly extract multi-scale features of the image to be processed, ensuring that the complete information of image details can be effectively extracted and applied to large-scale visual language models.

[0041] Figure 1 A flowchart illustrating the feature extraction method provided in an embodiment of this disclosure is shown. Figure 1 As shown, the method is applied to a feature extraction module in a large visual language model. The feature extraction module includes an encoder and a cropper. The large visual language model also includes a visual language pre-trained model connected to the feature extraction module. The method includes:

[0042] In step S11, the image to be processed is acquired;

[0043] In step S12, the encoder extracts features from the image to be processed to obtain an initial feature map. The resolution of the initial feature map is the same as the resolution of the image to be processed, and the size of the initial feature map is an integer multiple of the input size required by the visual language pre-training model.

[0044] In step S13, the initial feature map is segmented by the cropper to obtain at least one target feature map, the size of which conforms to the input size required by the visual language pre-trained model.

[0045] In one possible implementation, the feature extraction method can be executed by a feature extraction device, such as a terminal device, server, or other electronic device. The terminal device can be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, or wearable device, etc. In some possible implementations, the feature extraction method can be implemented by a processor calling computer-readable instructions stored in memory.

[0046] In one possible implementation, the feature extraction method provided in this disclosure can be applied to various large vision language models (LVLMs). Large vision language models are multimodal artificial intelligence models that combine visual understanding and natural language processing capabilities. They can receive multimodal inputs such as images, videos, and text, and generate related natural language outputs.

[0047] In one possible implementation, the large visual language model includes at least a feature extraction module and a visual language pre-training model (Sigmoid Loss for Language Image Pre-Training, SIGLIP). The output of the feature extraction module is connected to the input of the visual language pre-training model. This feature extraction module is used to perform lossless feature extraction on the image to be processed, using the efficiently extracted effective features as the target feature map without changing the image resolution. The visual language pre-training model is used to generate a multi-dimensional representation vector based on the target feature map.

[0048] Figure 2 A schematic diagram of the structure of a large visual language model provided in an embodiment of this disclosure is shown. Figure 2As shown, the large-scale visual language model, in addition to the feature extraction module 3 and the visual language pre-trained model 4, may also include a fusion module 5, a splicing module 6, an embedding module 7, and a large language model 8. It should be understood that the embodiments disclosed herein are only for illustrative purposes. Figure 2 As an example, it can be any large visual language model that includes feature extraction module 3 and visual language pre-trained model 4. The embodiments of this disclosure do not limit the specific structure of the large visual language model.

[0049] After processing by feature extraction module 3, lossless image features (e.g., target feature maps) can be extracted from the images to be processed. Simultaneously, feature extraction module 3 must ensure that the shape of the output image features (e.g., target feature maps) meets the input size requirements of the visual language pre-training model 4. The visual language pre-training model 4 converts the image features output by feature extraction module 3 into multi-dimensional representation vectors. Fusion module 5 fuses the multi-dimensional representation vectors of different image features belonging to the same image to be processed, obtaining a fused vector. Concatenation module 6 concatenates the fused vector with the text features output by embedding module 7 and sends the concatenation result to the large language model 8, obtaining the output of the large visual language model.

[0050] In one possible implementation, the feature extraction module 3 may include an encoder 1 and a cropper 2, wherein the encoder 1 is used to perform the lossless feature extraction task and the cropper 2 is used to perform the segmentation task.

[0051] In one possible implementation, in step S11, the acquired image to be processed can be an image of arbitrary size and resolution. Here, image size represents the number of pixels in the width and height directions of the image to be processed, while image resolution refers to the number of pixels per unit length of the image to be processed, for example, expressed in "pixels per inch". Higher image resolution results in denser pixels per unit area of ​​the image to be processed, and a more detailed image. Conversely, lower image resolution results in sparser pixels per unit area of ​​the image to be processed, and a coarser image.

[0052] The images to be processed can be video surveillance images, industrial inspection images, satellite remote sensing images, microscope images, infrared thermal imaging images, remote sensing images, medical images, etc. This disclosure does not limit the type of images to be processed.

[0053] In step S11, the image to be processed is obtained. In step S12, the encoder can extract features from the image to be processed to obtain an initial feature map. The encoder can be designed with multiple convolutional networks (e.g., a two-layer convolutional network) to dynamically adjust the size of the convolutional kernels at various specific resolutions, thus preserving the detailed features of the image to be processed. The resolution of the initial feature map output by the encoder is the same as the resolution of the image to be processed, and the size of the initial feature map is an integer multiple of the input size required by the visual language pre-training model. This allows for direct adaptation to the cropping segment in step S13 without additional padding or interpolation. The cropping process segments the initial feature map to obtain a target feature map that meets the input size requirements of the visual language pre-training model. This allows the segmented target feature map to be losslessly transferred to subsequent modules (e.g., the fusion module) via the visual language pre-training model.

[0054] In this way, the feature extraction method of the present disclosure significantly improves the expressive power of visual features by setting a feature pre-processing module, eliminates the interpolation and padding steps, and can efficiently extract effective features without changing the image resolution. It avoids the errors caused by interpolation and padding, thereby improving the performance of large visual language models. Especially in high-resolution scenes, it can ensure the integrity of feature information and greatly improve the accuracy of large visual language models.

[0055] The feature extraction method of this disclosure embodiment will be described in detail below.

[0056] After obtaining the image to be processed in step S11, the encoder can be used to extract features from the image to be processed in step S12 to obtain an initial feature map.

[0057] In one possible implementation, the encoder includes a first convolutional layer and a second convolutional layer. The encoder extracts features from the image to be processed to obtain an initial feature map, comprising: extracting features from the image to be processed in a first dimension using the first convolutional layer to obtain a first feature map, wherein the size of the first feature map in the first dimension is an integer multiple of the input size required by the visual language pre-training model; and extracting features from the first feature map in a second dimension using the second convolutional layer to obtain the initial feature map, wherein the size of the initial feature map is an integer multiple of the input size required by the visual language pre-training model in both the first and second dimensions.

[0058] For example, suppose the first dimension is the height direction, the second dimension is the width direction, the size of the image to be processed is H×W, where H represents the height of the input image to be processed, W represents the width of the input image to be processed, and the required input size of the visual language pre-training model is SH×SW, where SH represents the height of the input size required by the visual language pre-training model, and SW represents the width of the input size required by the visual language pre-training model.

[0059] The first convolutional layer extracts features from the image to be processed in the height direction, resulting in a first feature map with a height of SH×n and a width of W. The second convolutional layer then extracts features from the first feature map in the width direction, resulting in an initial feature map with a height of SH×n and a width of SW×m, where m and n are positive integers, and SH×n is less than or equal to H, and SW×m is less than or equal to W.

[0060] It should be understood that the first dimension direction and the second dimension direction are different dimension directions. For example, the first dimension direction can be the height direction and the second dimension direction can be the width direction; or, the first dimension direction can be the width direction and the second dimension direction can be the height direction. The embodiments of this disclosure do not impose specific limitations on this.

[0061] In this way, the convolutional kernels in the first and second convolutional layers can extract local features of the image to be processed, such as edges and textures, through a sliding window mechanism, which is beneficial to improving the local perception capability during the feature extraction process. Furthermore, by setting the size of the convolutional kernels in the first and second convolutional layers, the size of the image to be processed can be dynamically adjusted, ensuring that the shape of the target feature map output by the encoder, after being segmented by the cropper, conforms to the requirements of the visual language pre-trained model. In addition, the first and second convolutional layers also possess scale invariance, supporting the extraction of multi-scale features from multiple convolutional layers and enhancing robustness to inputs at different resolutions.

[0062] In one possible implementation, the size of the convolutional kernel in the first convolutional layer is determined based on a first size and a first input size, wherein the first size is the size of the image to be processed in the first dimension, and the first input size is the input size required by the visual language pre-training model in the first dimension; the size of the convolutional kernel in the second convolutional layer is determined based on a second size and a second input size, wherein the second size is the size of the image to be processed in the second dimension, and the second input size is the input size required by the visual language pre-training model in the second dimension.

[0063] It should be understood that the first and second convolutional layers also have the functions of dimensionality reduction and shape adjustment. By setting the size and stride of their respective convolutional kernels, the size of the input image to be processed can be dynamically adjusted to ensure that the output shape meets the requirements of the visual language pre-trained model. The embodiments of this disclosure do not limit the specific values ​​of the size and stride of the convolutional kernels set by the first and second convolutional layers.

[0064] The following section details the dimensions of the convolutional kernels in the first convolutional layer and the derivation of the dimensions of the convolutional kernels in the second convolutional layer. Assuming the first dimension is the height direction and the second dimension is the width direction, the convolution operation can follow the following formula:

[0065] , (1)

[0066] Where H is the first dimension, representing the height of the input image to be processed, and W is the second dimension, representing the width of the input image to be processed. Height Output represents the height of the output image after the convolution operation, which is also the height of the initial feature map output by the encoder. Width k1 represents the width of the output image after the convolution operation, which is also the width of the initial feature map output by the encoder. k2 represents the height of the convolution kernel, p represents the padding, which is the extra boundary value added around the input image to be processed, and s represents the stride, which is the jump distance when the convolution kernel slides on the input image to be processed.

[0067] If the convolution operation is padded, it may produce unnatural image edges, affecting the accuracy of feature extraction. Therefore, the convolution operation should be padded-free, while preserving as many image details as possible. In this case, the convolution kernels in the first convolutional layer and the convolution kernels in the second convolutional layer can be padded-free (e.g., p=0) and stride 1 (e.g., s=1).

[0068] This type of convolution operation can reduce or even avoid useless information caused by padding or interpolation during feature extraction, reduce the loss of image details, and facilitate lossless extraction of image features.

[0069] Substituting p=0 and s=1 into formula (1), and then reversing formula (1), we can obtain:

[0070] (2)

[0071] Where H is the first dimension, representing the height of the input image to be processed, and W is the second dimension, representing the width of the input image to be processed. HeightOutput represents the height of the output image after the convolution operation. Width k1 represents the width of the output image after the convolution operation, k2 represents the height of the convolution kernel, and k2 represents the width of the convolution kernel.

[0072] Assuming the input size required for the visual language pre-training model is SH×SW, where SH is the first input size and SW is the second input size, and since the initial feature map output by the encoder is segmented by a cropper before being fed into the visual language pre-training model, in order to match the input size required by the visual language pre-training model, the Output can be set to... Height =SH×n,Output Width =SW×m, where n and m are positive integers.

[0073] Output Height =SH×n,Output Width Substituting =SW×m into formula (2), we get:

[0074] (3)

[0075] Where H is the first dimension, representing the height of the input image to be processed; W is the second dimension, representing the width of the input image to be processed; SH is the first input dimension, representing the required dimension of the visual language pre-trained model in the height direction; SW is the second input dimension, representing the required dimension of the visual language pre-trained model in the width direction; k1 represents the height of the convolution kernel; k2 represents the width of the convolution kernel; and n and m are positive integers.

[0076] As n, m, k1, and k2 change, they can represent any positive integer H greater than or equal to SH and any positive integer W greater than or equal to SW. For high-quality image datasets, the size of the image dataset is mostly greater than SH × SW. Therefore, the feature extraction method of this disclosure embodiment can be adapted to high-quality image datasets of large visual language models.

[0077] Continuing to reverse the formula (3), we can obtain:

[0078] (4)

[0079] Where H is the first dimension, representing the height of the input image to be processed; W is the second dimension, representing the width of the input image to be processed; SH is the first input dimension, representing the required dimension of the visual language pre-trained model in the height direction; SW is the second input dimension, representing the required dimension of the visual language pre-trained model in the width direction; k1 represents the height of the convolution kernel; k2 represents the width of the convolution kernel; and n and m are positive integers.

[0080] To improve computational efficiency, the smaller the values ​​of k1 and k2, the better. We can maximize n and m, subtract the largest multiple of SH from H, and subtract the largest multiple of SW from W. In this case, we can replace H-SH×n with H mod SH, and W-SW×n with W mod SW, resulting in:

[0081] (5)

[0082] Where mod represents the modulo operation, which represents the remainder after dividing two integers; H is the first dimension, which represents the height of the input image to be processed; W is the second dimension, which represents the width of the input image to be processed; SH is the first input dimension, which represents the required dimension of the visual language pre-trained model in the height direction; SW is the second input dimension, which represents the required dimension of the visual language pre-trained model in the width direction; k1 represents the height of the convolution kernel; and k2 represents the width of the convolution kernel.

[0083] Therefore, the kernel size of the first convolutional layer is k1×1, which is responsible for processing the height dimension of the image to be processed, and the kernel size of the second convolutional layer is 1×k2, which is responsible for processing the width dimension of the image to be processed.

[0084] While convolutional operations possess strong local awareness during feature extraction, the lack of global consistency adjustment capabilities within convolutional layers means that the features generated after convolution may not perfectly align with the training distribution of the visual language pre-trained model, impacting the overall performance of the large-scale visual language model. Therefore, a normalization layer can be added to the feature extraction module to globally normalize the features generated by convolution, improving the consistency of the feature distribution and enabling it to better adapt to the training requirements of the visual language pre-trained model.

[0085] Figure 3 A schematic diagram of the feature extraction module provided in an embodiment of this disclosure is shown. Figure 3 As shown, the feature extraction module 3 may include an encoder 1 and a cropper 2. The encoder 1 includes a first convolutional layer 11, a second convolutional layer 12, a first normalization layer 13, and a second normalization layer 14. The encoder 1 extracts features from the image to be processed to obtain an initial feature map, including: extracting features from the image to be processed in the first dimension using the first convolutional layer 11 to obtain a first feature map; performing a first normalization process on the first feature map using the first normalization layer 13 to obtain a second feature map; extracting features from the second feature map in the second dimension using the second convolutional layer 12 to obtain a third feature map; and performing a second normalization process on the third feature map using the second normalization layer 14 to obtain the initial feature map.

[0086] The first normalization layer 13 and the second normalization layer 14 can employ lightweight and efficient root mean square normalization (RMSNorm). The specific structure of the first normalization layer 13 and the second normalization layer 14 is not limited in the embodiments disclosed herein.

[0087] By setting the first normalization layer 13 and the second normalization layer 14, the features generated by convolution (such as the first feature map and the third feature map) can be normalized globally, thereby dynamically adjusting the feature distribution, ensuring the consistency of the feature distribution, and enabling it to better adapt to the training requirements of the visual language pre-training model, so that the convolutional data matches the visual language pre-training model more closely.

[0088] The following example uses the visual language pre-training model SIGLIP_14 (i.e., the input size required by the visual language pre-training model is 14×14) and the size of the image to be processed is 32×30×3 to illustrate the feature extraction module 1.

[0089] It should be understood that the feature extraction method of the present disclosure embodiments can be adapted to visual language pre-trained models of any size and image to be processed of any size and resolution, and the comparison of the embodiments of the present disclosure is not limited.

[0090] The first dimension (e.g., height dimension) of the image to be processed can be adjusted using the convolution kernel of the first convolutional layer 11, so that the size of the first feature map output by the first convolutional layer 11 in the first dimension direction is an integer multiple of the input size required by the visual language pre-trained model for SIGLIP_14.

[0091] Figure 4 This diagram illustrates feature extraction performed by the first convolutional layer according to an embodiment of the present disclosure. Figure 4 As shown, for an image to be processed with a size of 32×30×3, refer to formula (5). When H=32 and SH=14, 32 mod 14 = 4 and k1=4 + 1 = 5. Then, the size of the convolution kernel of the first convolutional layer 11 is k1×1=5×1. After the convolution operation Conv1 of the first convolutional layer 11, the size of the first feature map output by the first convolutional layer 11 is 28×30×3.

[0092] The first normalization layer 13 performs a first normalization process on the first feature map output by the first convolutional layer 11 to obtain a second feature map, the size of which is the same as that of the first feature map. The feature distribution in the first feature map is adjusted by the learnable parameters of the first normalization layer 13 to make it consistent with the feature distribution of the visual language pre-trained model SIGLIP_14 (especially for the feature distribution in the height direction) to enhance adaptability.

[0093] The second feature map output from the first normalization layer 13 can be fed into the second convolutional layer 12. The second dimension (e.g., width dimension) of the second feature map is adjusted by the convolution kernel of the second convolutional layer 12, which is equivalent to adjusting the second dimension (e.g., width dimension) of the image to be processed using the convolution kernel of the second convolutional layer 12. The dimensions of the third feature map output from the second convolutional layer 12 in both the first and second dimensions are integer multiples of the input size required by SIGLIP_14 for the visual language pre-trained model.

[0094] Figure 5 This diagram illustrates feature extraction performed by the second convolutional layer according to an embodiment of the present disclosure. Figure 5 As shown, for the second feature map with a size of 28×30×3, refer to formula (5). When W=30 and SW=14, 30 mod 14 = 2 and k2=2+1=3. Then, the size of the convolution kernel of the second convolutional layer 12 is 1×k2=1×3. After the convolution operation Conv2 of the second convolutional layer 12, the size of the third feature map output by the second convolutional layer 12 is 28×28×3.

[0095] The second normalization layer 14 performs a second normalization process on the third feature map output by the second convolutional layer 12 to obtain an initial feature map, the size of which is the same as the size of the third feature map. By adjusting the feature distribution in the third feature map through the learnable parameters of the second normalization layer 14, the stability and distribution adaptability of the features can be further optimized.

[0096] In one possible implementation, the initial feature map is segmented by the cropper 2 to obtain at least one target feature map, including: segmenting the initial feature map according to the input size required by the visual language pre-training model to obtain at least one first sub-map; sampling the image to be processed according to the input size required by the visual language pre-training model to obtain a sampled image; and aggregating the at least one first sub-map and the sampled image to obtain at least one target feature map.

[0097] For example, the initial feature map output by the second normalization layer 14 can be fed into the cropper 2. The size of the initial feature map is 28×28×3. The input size required by the visual language pre-training model SIGLIP_14 is 14×14. The initial feature map after the second normalization process can be divided into 14×14 patches as the first sub-image. The original image to be processed is sampled and resized to 14×14 as the sampled image. The four first sub-images after the segmentation and one sampled image are aggregated to generate five target feature maps with a size of 14×14. These five target feature maps are then input into the visual language pre-training model SIGLIP_14.

[0098] In this way, it is possible to adapt to the input size requirements of visual language pre-trained models of various specifications, realize multi-scale feature fusion, and improve the ability of large visual language models to understand image details and global semantics.

[0099] In one possible implementation, the large visual language model further includes a fusion module and a large language model. The training method of the large visual language model includes: training the initial large visual language model with the parameters of the visual language pre-training model, the fusion module, and the large language model fixed to obtain a first large visual language model; training the first large visual language model with the parameters of the feature extraction module fixed to obtain a second large visual language model; and training the second large visual language model with all parts of the second large visual language model participating in the training to obtain a trained large visual language model.

[0100] For example, with Figure 2 Taking the large visual language model shown as an example, the training process of a large visual language model can be divided into three stages:

[0101] In the first stage, the parameters of the visual language pre-trained model 4, the fusion module 5, and the large language model 8 can be kept unchanged, while the feature extraction module 3 is trained.

[0102] In this module, by fine-tuning the weights of the convolutional kernels in the first convolutional layer 11 and the second convolutional layer 12 of the encoder 1 in the feature extraction module 3, the feature distribution of different images can be automatically adapted to extract key information. By fine-tuning the first normalization layer 13 and the second normalization layer 14 in the encoder 1 of the feature extraction module 3 (for example, by adding learnable scaling and offset parameters as learnable parameters of the first normalization layer 13 and the second normalization layer 14 on the basis of standard root mean square normalization), the convolutional image features (e.g., target feature maps) can be made consistent with the feature distribution of the visual language pre-trained model 4, thereby achieving effective adaptation of visual features. The goal of this stage is to enable the feature extraction module 3 to achieve more uniform and robust feature representation of images with different input resolutions after convolution and normalization through the adjustment of learnable parameters. (Specifically, the first convolutional layer 11 and the second convolutional layer 12 after training can accurately extract image features, and the first normalization layer 13 and the second normalization layer 14 after training can make the convolutional image features consistent with the feature distribution of the visual language pre-trained model 4.) This allows the target feature map output by the feature extraction module 3 to be adapted to the subsequent processing of the visual language pre-trained model 4. Furthermore, optimizing the first normalization layer 13 and the second normalization layer 14 in the first stage can effectively mitigate gradient oscillation problems during training, ensure the stability of feature data during transmission, and improve the convergence speed of training.

[0103] In the second stage, the visual language pre-training model 4, the fusion module 5, and the large language model 8 can be optimized while keeping the parameters of the feature extraction module 3 unchanged. Specifically, training the visual language pre-training model 4 further ensures the effectiveness of its visual feature extraction and guarantees that the image features transferred from the feature extraction module 3 can play a maximum role in the visual language pre-training model 4. Furthermore, this stage aims to optimize the collaborative work between visual features and the large language model 8, enabling the input features of the entire large visual language model to achieve optimal performance in multimodal training.

[0104] In the third stage, all parts of the large visual language model (e.g., feature extraction module 3, visual language pre-trained model 4, fusion module 5, and large language model 8) participate in training. Through joint optimization, the efficiency and performance of the entire large visual language model are further improved. The goal of this stage is to improve the accuracy and generalization ability of the large visual language model through global training, thereby enabling the large visual language model to provide more efficient feature extraction and understanding capabilities when processing real-world data.

[0105] This phased training approach allows for gradual optimization of the feature extraction module, which is well-suited to the visual language pre-trained model, ensuring lossless extraction and optimization of image features and improving the performance and efficiency of the entire large-scale visual language model.

[0106] In summary, the feature extraction method of this disclosure achieves lossless image feature extraction by setting an encoder and a cropper, effectively avoiding the loss of image details in related technologies. The normalization process introduced in the encoder ensures good adaptation between convolutional features and the visual language pre-trained model, improving the stability and training effect of large-scale visual language models. In addition, this method also trains large-scale visual language models through staged training, further improving the accuracy and efficiency of large-scale visual language models.

[0107] In addition, this disclosure also provides a feature extraction apparatus, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any of the feature extraction methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the relevant section on methods and will not be repeated here.

[0108] Figure 6 A block diagram of the feature extraction apparatus provided in an embodiment of this disclosure is shown, such as Figure 6 As shown, the device is applied to a feature extraction module in a large visual language model. The feature extraction module includes an encoder and a cropper. The large visual language model also includes a visual language pre-trained model connected to the feature extraction module. The device includes:

[0109] Acquisition unit 61 is used to acquire the image to be processed;

[0110] The feature extraction unit 62 is used to extract features from the image to be processed by the encoder to obtain an initial feature map. The resolution of the initial feature map is the same as the resolution of the image to be processed, and the size of the initial feature map is an integer multiple of the input size required by the visual language pre-training model.

[0111] The segmentation unit 63 is used to segment the initial feature map through the cropper to obtain at least one target feature map, the size of which conforms to the input size required by the visual language pre-trained model.

[0112] In one possible implementation, the encoder includes a first convolutional layer and a second convolutional layer, and the feature extraction unit 62 is configured to: extract features from the image to be processed through the first convolutional layer in a first dimension to obtain a first feature map, wherein the size of the first feature map in the first dimension is an integer multiple of the input size required by the visual language pre-training model; and extract features from the first feature map through the second convolutional layer in a second dimension to obtain the initial feature map.

[0113] In one possible implementation, the encoder further includes a first normalization layer and a second normalization layer. The feature extraction unit 62 is configured to: extract features from the image to be processed using the first convolutional layer in a first dimension to obtain a first feature map; perform a first normalization process on the first feature map using the first normalization layer to obtain a second feature map; extract features from the second feature map using the second convolutional layer in a second dimension to obtain a third feature map; and perform a second normalization process on the third feature map using the second normalization layer to obtain the initial feature map.

[0114] In one possible implementation, the size of the convolutional kernel in the first convolutional layer is determined based on a first size and a first input size, wherein the first size is the size of the image to be processed in the first dimension, and the first input size is the input size required by the visual language pre-training model in the first dimension; the size of the convolutional kernel in the second convolutional layer is determined based on a second size and a second input size, wherein the second size is the size of the image to be processed in the second dimension, and the second input size is the input size required by the visual language pre-training model in the second dimension.

[0115] In one possible implementation, both the convolutional kernels in the first convolutional layer and the convolutional kernels in the second convolutional layer perform convolution operations with no padding and a stride of 1.

[0116] In one possible implementation, the segmentation unit 63 is configured to: segment the initial feature map according to the input size required by the visual language pre-training model to obtain at least one first sub-map; sample the image to be processed according to the input size required by the visual language pre-training model to obtain a sampled image; and aggregate the at least one first sub-map and the sampled image to obtain at least one target feature map.

[0117] In one possible implementation, the large visual language model further includes a fusion module and a large language model, and the device further includes a training unit, which is configured to: train the initial large visual language model with the parameters of the visual language pre-training model, the fusion module, and the large language model fixed, to obtain a first large visual language model; train the first large visual language model with the parameters of the feature extraction module fixed, to obtain a second large visual language model; and train the second large visual language model with all parts of the second large visual language model participating in the training, to obtain a trained large visual language model.

[0118] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0119] This disclosure also proposes a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the above-described method. The computer-readable storage medium can be volatile or non-volatile.

[0120] This disclosure also proposes an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to execute the above-described method.

[0121] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.

[0122] Electronic devices can be provided as terminals, servers, or other forms of devices.

[0123] Figure 7 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server or a terminal device. (Refer to...) Figure 7The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.

[0124] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output interface 1958. Electronic device 1900 can operate on an operating system stored in memory 1932, such as a server operating system, a graphical user interface-based operating system, or a multi-user, multi-process computer operating system (Unix). TM Linux is a free and open-source Unix-like operating system. TM ), an open-source Unix-like operating system (FreeBSD) TM (or similar.)

[0125] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of an electronic device 1900 to perform the above-described method.

[0126] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.

[0127] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, (but not limited to) electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0128] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0129] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0130] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0131] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0132] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0134] The computer program product can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0135] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0136] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0137] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

[0138] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A feature extraction method, characterized in that, The method is applied to the feature extraction module of a large visual language model. The feature extraction module includes an encoder and a cropper. The large visual language model also includes a visual language pre-trained model connected to the feature extraction module. The method includes: Obtain the image to be processed; The encoder extracts features from the image to be processed to obtain an initial feature map. The resolution of the initial feature map is the same as the resolution of the image to be processed, and the size of the initial feature map is an integer multiple of the input size required by the visual language pre-training model. The initial feature map is segmented by the cropper to obtain at least one target feature map, the size of which conforms to the input size required by the visual language pre-trained model.

2. The method according to claim 1, characterized in that, The encoder includes a first convolutional layer and a second convolutional layer. The encoder extracts features from the image to be processed to obtain an initial feature map, including: The first convolutional layer extracts features from the image to be processed in the first dimension to obtain a first feature map, wherein the size of the first feature map in the first dimension is an integer multiple of the input size required by the visual language pre-training model. The first feature map is obtained by extracting features from the second feature map in the second dimension through the second convolutional layer.

3. The method according to claim 2, characterized in that, The encoder further includes a first normalization layer and a second normalization layer. The encoder extracts features from the image to be processed to obtain an initial feature map, including: The first convolutional layer extracts features from the image to be processed in the first dimension to obtain a first feature map. The first feature map is subjected to a first normalization process through the first normalization layer to obtain a second feature map; The second convolutional layer extracts features from the second feature map in the second dimension to obtain the third feature map; The third feature map is subjected to a second normalization process through the second normalization layer to obtain the initial feature map.

4. The method according to claim 2 or 3, characterized in that, The size of the convolution kernel in the first convolutional layer is determined based on a first size and a first input size, wherein the first size is the size of the image to be processed in the first dimension, and the first input size is the input size required by the visual language pre-trained model in the first dimension. The size of the convolution kernel in the second convolutional layer is determined based on a second dimension and a second input dimension, wherein the second dimension is the size of the image to be processed in the second dimension, and the second input dimension is the input dimension required by the visual language pre-trained model in the second dimension.

5. The method according to claim 4, characterized in that, Both the convolution kernels in the first convolutional layer and the convolution kernels in the second convolutional layer perform convolution operations with no padding and a stride of 1.

6. The method according to claim 1, characterized in that, The initial feature map is segmented using the cropper to obtain at least one target feature map, including: The initial feature map is segmented according to the input size required by the visual language pre-training model to obtain at least one first sub-map; The image to be processed is sampled according to the input size required by the visual language pre-training model to obtain a sampled image; The at least one first sub-image and the sampled image are aggregated to obtain at least one target feature map.

7. The method according to claim 1, characterized in that, The large-scale visual language model also includes a fusion module and a large-scale language model. The training methods for the large-scale visual language model include: With the parameters of the visual language pre-training model, the fusion module, and the large language model in the initial large visual language model fixed, the initial large visual language model is trained to obtain the first large visual language model. With the parameters of the feature extraction module in the first large-scale visual language model fixed, the first large-scale visual language model is trained to obtain the second large-scale visual language model. With all parts of the second large visual language model participating in the training, the second large visual language model is trained to obtain a trained large visual language model.

8. A feature extraction device, characterized in that, The device is applied to a feature extraction module in a large visual language model. The feature extraction module includes an encoder and a cropper. The large visual language model also includes a visual language pre-trained model connected to the feature extraction module. The device includes: The acquisition unit is used to acquire the image to be processed; The feature extraction unit is used to extract features from the image to be processed by the encoder to obtain an initial feature map. The resolution of the initial feature map is the same as the resolution of the image to be processed, and the size of the initial feature map is an integer multiple of the input size required by the visual language pre-training model. The segmentation unit is used to segment the initial feature map through the cropper to obtain at least one target feature map, the size of which conforms to the input size required by the visual language pre-trained model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

11. A computer program product comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.