Image screening method and system based on visual large model
Patent Information
- Application Number
- CN202511577290.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-10-31
AI Technical Summary
该方法虽在一定程度上实现了自动化筛选,但仍存在明显局限性:首先,图像到文本的转换过程本质上是一种信息压缩,导致图像中的细节特征和多模态信息丢失;其次,文本标签的质量严重依赖标注模型的性能,标注偏差或错误会直接影响筛选结果的准确性;此外,面对复杂、多维度或隐含特征的筛选条件时,纯文本表征难以充分表达图像语义,限制了模型在高精度场景下的应用
[0015]本申请公开了一种基于视觉大模型的图像筛选方法与系统,首先构建改进的视觉大模型,之后使用预训练数据集对所述多个视觉-语言融合器进行训练,同时冻结所述视觉编码器和所述语言解码器的参数,得到训练后的视觉大模型,构建图像筛选任务的训练数据集,并基于LoRA技术对所述训练后的视觉大模型进行微调,得到图像筛选模型,利用图像筛选模型对待筛选图像及筛选条件进行分析,输出筛选结果。改进的视觉大模型包括ViT提取多层次图像特征、多个独立MLP融合器映射特征到语言空间、LLM解码器生成输出,结合LoRA微调技术,显著提升了图像筛选的准确性和效率,有效避免了图像信息损失并增强了在复杂筛选条件下的鲁棒性。
Smart Images

Figure CN121366340B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and more specifically, to an image screening method and system based on a large visual model. Background Technology
[0002] In the field of image processing and recognition technology, image screening refers to identifying and selecting target images that meet specific conditions from a set of images, or excluding images that do not meet the requirements. Traditional image screening methods mainly rely on manual judgment, which is not only inefficient but also easily affected by subjective human factors, making it difficult to guarantee the consistency and accuracy of the screening.
[0003] With the development of Large Language Models (LLMs), the automation level of image filtering tasks has been improved to some extent. Since LLMs typically use text as input, images must first undergo labeling processing, i.e., generating descriptive text labels through image annotation techniques, and then inputting these labels into the model for judgment. While this method achieves automation filtering to some extent, it still has significant limitations: First, the image-to-text conversion process is essentially an information compression, leading to the loss of detailed features and multimodal information in the image; second, the quality of text labels heavily depends on the performance of the annotation model, and annotation bias or errors directly affect the accuracy of the filtering results; furthermore, when faced with complex, multi-dimensional, or implicit feature-rich filtering conditions, pure text representations are insufficient to fully express the image semantics, limiting the model's application in high-precision scenarios.
[0004] In recent years, the development of vision-large models has provided a new technical path for direct image processing. These models can process image input end-to-end, avoiding information loss caused by intermediate representations. However, most existing vision-large models are based on the Transformer architecture. While they can capture image features layer by layer during feature extraction, in practical applications they typically only use the feature vectors of the final output layer as image representations, neglecting the multi-scale and multi-level feature information contained in intermediate layers. This representation method struggles to fully capture the detailed texture, structural information, and semantic hierarchy of images, resulting in limited model performance in complex image selection tasks. Therefore, current technology still lacks a visual model processing method that can effectively integrate multi-level image feature representations, especially in tasks requiring high-granularity semantic understanding such as image selection. How to fully mine and utilize feature information at different levels has become a key issue in improving model performance. Summary of the Invention
[0005] The purpose of this application is to overcome the shortcomings of existing technologies and provide an image screening method and system based on a large visual model. By improving the large visual model, including ViT to extract multi-level image features, multiple independent MLP fusion machines to map features to the language space, and LLM decoder to generate output, combined with LoRA fine-tuning technology, the accuracy and efficiency of image screening are significantly improved, image information loss is effectively avoided, and robustness under complex screening conditions is enhanced.
[0006] The objective of this application is achieved through the following technical solution: Firstly, this application proposes an image filtering method based on a large visual model, the method comprising: Construct an improved visual big model, which includes a visual encoder, multiple visual-language fusion units, and a language decoder; The multiple vision-language fusion processors are trained using a pre-trained dataset, while the parameters of the visual encoder and the language decoder are frozen, resulting in a large-scale trained vision model. The pre-trained dataset contains images and corresponding text descriptions or labels. A training dataset for an image selection task is constructed, and the trained visual model is fine-tuned based on LoRA technology to obtain an image selection model. The training dataset for the image selection task includes images, selection condition text, and corresponding expected outputs. The image filtering model is used to analyze the images to be filtered and the filtering conditions, and the filtering results are output.
[0007] In one possible implementation, the visual encoder is a visual Transformer used to extract initial layer features, intermediate layer features, and final layer features of the input image; The visual-language fusion unit is a multilayer perceptron used to map corresponding visual features to a space with the same dimension as the language features. The language decoder uses a large language model to receive mapped features and generate output.
[0008] In one possible implementation, in the visual encoder, each feature layer is mapped separately through a visual-language fusion processor.
[0009] In one possible implementation, initial layer features are used to capture low-level features of the image's edges and textures, intermediate layer features are used to capture mid-level features of the image's shape and components, and final layer features are used to capture high-level global semantic features of the image.
[0010] In one possible implementation, the multilayer perceptron includes two hidden layers: a first layer for mapping visual features to an intermediate dimension, and a second layer for mapping the intermediate dimension to a space consistent with the language feature dimension.
[0011] In one possible implementation, the method further includes: The parameters of the vision-language fusion engine are adjusted by optimizing the loss function, which is calculated based on the difference between the predicted language description and the actual language description.
[0012] In one possible implementation, the step of fine-tuning the trained large visual model based on LoRA technology to obtain the image selection model includes: Add a LoRA adaptation layer to the trained large visual model; The parameters of the LoRA adaptation layer are trained using the image filtering task training dataset, while keeping the parameters of the visual encoder, visual-language fusion unit, and language decoder in the large visual model unchanged, to obtain the image filtering model. The image filtering model is used to receive the image to be filtered and the corresponding filtering condition text, and output the judgment result of whether the image meets the filtering criteria.
[0013] Secondly, this application proposes an image filtering system based on a large visual model, the system comprising: Build modules are used to construct improved large-scale visual models, which include a visual encoder, multiple visual-language fusion units, and a language decoder. The training module is used to train the multiple vision-language fusion processors using a pre-training dataset, while freezing the parameters of the visual encoder and the language decoder to obtain a trained large-scale vision model. The pre-training dataset contains images and corresponding text descriptions or labels. The fine-tuning module is used to construct the training dataset for the image selection task and to fine-tune the trained visual model based on LoRA technology to obtain the image selection model. The training dataset for the image selection task includes images, selection condition text, and corresponding expected outputs. The analysis module is used to analyze the images to be filtered and the filtering conditions using an image filtering model, and output the filtering results.
[0014] The main solution and its various further alternatives described above can be freely combined to form multiple solutions, all of which are solutions that can be adopted and are claimed in this application; furthermore, the (non-conflicting alternatives) can also be freely combined with each other and with other alternatives. Those skilled in the art, after understanding the solution of this application, will realize from the prior art and common general knowledge that there are many combinations, all of which are technical solutions to be protected in this application, and will not be exhaustively listed here.
[0015] This application discloses an image filtering method and system based on a large visual model. First, an improved large visual model is constructed. Then, multiple vision-language fusion processors are trained using a pre-trained dataset, while simultaneously freezing the parameters of the visual encoder and the language decoder, resulting in a trained large visual model. A training dataset for the image filtering task is then constructed, and the trained large visual model is fine-tuned using LoRA technology to obtain an image filtering model. This model is then used to analyze the images to be filtered and the filtering conditions, outputting the filtering results. The improved large visual model includes ViT for extracting multi-level image features, multiple independent MLP fusion processors mapping features to the language space, and an LLM decoder generating the output. Combined with LoRA fine-tuning technology, this significantly improves the accuracy and efficiency of image filtering, effectively avoids image information loss, and enhances robustness under complex filtering conditions. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 The diagram shows a flowchart of an image filtering method based on a large visual model proposed in an embodiment of this application.
[0018] Figure 2 A schematic diagram of the visual large model proposed in the embodiments of this application is shown. Detailed Implementation
[0019] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0020] Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] To address the shortcomings of existing image filtering methods, such as image information compression, label quality dependence, and the lack of multi-level image feature extraction by large visual models, this application proposes an image filtering method and system based on large visual models. This method and system can process image filtering tasks more efficiently and accurately, and can also directly process image data and extract multi-level image features, thereby providing richer and more accurate image representations. Furthermore, it constructs a dataset for image filtering tasks for fine-tuning, ensuring that more ideal filtering results can be obtained under complex filtering conditions.
[0022] Please refer to Figure 1 , Figure 1 This paper illustrates a flowchart of an image filtering method based on a large visual model, as proposed in an embodiment of this application. The method includes: Step S1: Construct an improved visual big model, which includes a visual encoder, multiple visual-language fusion units, and a language decoder.
[0023] To address the shortcomings of existing technologies in multi-level image feature extraction, this application constructs a large-scale visual model that directly processes raw image data, avoiding information compression and feature loss caused by reliance on text labels in traditional methods. This model includes a visual encoder, multiple visual-language projectors, and a language decoder. These components work together to achieve hierarchical extraction and fusion of features at the initial, intermediate, and final layers of the image, thereby providing a more refined and comprehensive feature representation for image filtering tasks.
[0024] Figure 2 The diagram illustrates the visual large-scale model proposed in this application. First, the original image is input to a visual encoder, which, based on the ViT architecture, extracts initial layer features, intermediate layer features, and final layer features. Next, each feature layer passes through an independent vision-language fusion processor, using a multilayer perceptron to map the visual features to a dimension consistent with the language feature space. Then, these mapped features, along with the filtering condition text, are input to a language decoder. This decoder, based on a large language model, fuses visual and linguistic information, ultimately outputting the result indicating whether the image meets the filtering conditions. The entire process, through multilayer feature extraction and fusion, effectively improves the accuracy and efficiency of image filtering.
[0025] By extracting and mapping image features through an improved visual large model architecture, the ability to capture multi-level features of images is enhanced, thereby better performing image screening tasks.
[0026] The visual encoder is a visual Transformer used to extract initial layer features, intermediate layer features, and final layer features of the input image. Each feature layer is mapped by a visual-language fusion machine, which is a multilayer perceptron used to map the corresponding visual features to a space with the same dimension as the language features. The language decoder uses a large language model to receive the mapped features and generate the output.
[0027] The input image is encoded using a visual encoder, transforming it into visual features. The visual encoder employs ViT (Visual Transformer). Features from its initial, intermediate, and final layers are extracted. The initial layer of ViT captures basic low-level features such as edges and textures. These low-level features are crucial for capturing detailed image information, especially in image selection tasks, helping to accurately identify the basic content of the image. The intermediate layers of ViT primarily capture higher-level features, such as the shape, color, or other complex elements of objects. These intermediate features help to deeply understand the inherent relationships within the image, improving the model's performance under more complex selection conditions and enhancing semantic understanding. The final layer of ViT focuses on the overall semantics and contextual information of the image. Features in this layer contain the most representative and important parts of the image, effectively expressing its global semantics. For image selection tasks, especially when dealing with complex content, high-level features provide a deeper and more comprehensive understanding of the image, enabling the model to more accurately determine the relevance of images and selection criteria.
[0028] After feature extraction, the next step is to map the extracted features to interface with the language feature space. For the features of each layer (initial layer, intermediate layer, and final layer) of the aforementioned visual encoder, feature mapping is performed through an independent visual-language fusion unit. In this invention, the visual-language fusion unit employs an MLP (Multilayer Perceptron).
[0029] The multilayer perceptron includes two hidden layers. The first layer maps visual features to an intermediate dimension, and the second layer maps the intermediate dimension to a space consistent with the language feature dimension.
[0030] The first hidden layer receives the raw visual features extracted by the visual encoder and maps them from the initial feature dimension to an intermediate hidden dimension. This operation aims to perform preliminary transformation and dimensionality reduction on the visual features, reducing modal differences. The second hidden layer further maps the intermediate dimension features to a space that is completely consistent with the feature dimension of the language decoder, thereby achieving dimensional alignment and semantic fusion between visual and language features.
[0031] The mapped visual features are passed as input to the language decoder, which employs an LLM (Large Language Model). By understanding the relationship between visual and linguistic information, the decoder further enhances the model's ability to understand images and provides accurate prediction results for image selection tasks.
[0032] The initial layer features are used to capture low-level features of the image's edges and textures, the intermediate layer features are used to capture mid-level features of the image's shape and components, and the final layer features are used to capture high-level global semantic features of the image. The initial layer captures low-level features, such as edges and textures. These low-level features are crucial for capturing detailed image information, especially in image filtering tasks, helping to accurately identify the basic content of the image and providing a solid foundation for further filtering. The intermediate layers extract semantic features: These layers primarily capture higher-level features, such as the shape, color, or other complex elements of objects. These mid-level features help to deeply understand the inherent relationships within the image, improving the model's performance under more complex filtering conditions and enhancing its semantic understanding. The final layer strengthens overall semantic and contextual information: This final layer focuses on the overall semantic and contextual information of the image, typically containing the most representative parts. For image filtering tasks, especially when dealing with complex content, high-level features provide a deeper and more comprehensive understanding of the image, enabling the model to more accurately determine the relevance of images and filtering criteria.
[0033] Step S2: Train the multiple vision-language fusion processors using a pre-trained dataset, while freezing the parameters of the visual encoder and the language decoder to obtain a trained large vision model. The pre-trained dataset contains images and corresponding text descriptions or labels.
[0034] Multiple vision-language fusion engines were trained using a pre-trained dataset. During training, the parameters of the visual encoder and language decoder were frozen to preserve their pre-training results. The pre-trained dataset contained images and corresponding text descriptions or labels, providing rich training materials for the vision-language fusion engines.
[0035] The Vision Encoder (ViT) and LLM components utilize open-source pre-trained models: For the Vision Encoder (ViT), an open-source pre-trained model is used for initialization. This model has been trained on large-scale image datasets and is capable of effectively extracting various visual features. For the Language Decoder (LLM), an open-source pre-trained Large Language Model (LLM) is also used. This model has been trained on massive amounts of language data and is capable of handling various language tasks. Using this pre-trained language model ensures that the model has strong capabilities in the language generation part.
[0036] The core of training the vision-language fusion unit lies in training each vision-language fusion unit to optimize the feature mapping transformation capability. During training, to avoid destroying the pre-trained knowledge already acquired by the ViT and LLM parts, the parameters of these two parts are frozen. The purpose of freezing the parameters is: 1) to prevent excessive changes in the parameters of the ViT and LLM parts during subsequent training, which could lead to a deterioration in model performance; and 2) to preserve the pre-training capability of these parts, allowing the training process to focus on optimizing the feature mapping transformation process.
[0037] The method also includes: The parameters of the vision-language fusion engine are adjusted by optimizing the loss function, which is calculated based on the difference between the predicted language description and the actual language description.
[0038] The parameters of the vision-language fusion engine are adjusted by optimizing the loss function. The loss function is calculated based on the difference between the predicted and actual language descriptions, thus quantifying the accuracy of visual feature mapping to the language feature space. Using a supervised learning mechanism, the parameters of the vision-language fusion engine are iteratively optimized according to the pre-defined loss function, making the predicted language descriptions closer to the actual language descriptions, thereby improving the accuracy of visual feature mapping and enhancing the model's performance in vision-language fusion tasks.
[0039] The training process for the vision-language fusion processor includes: First, prepare the dataset: Select a suitable training dataset, such as an open-source pre-training dataset for large-scale vision models. These datasets should contain images and their corresponding text descriptions or labels. The dataset should be able to cover diverse image categories and semantic relationships to ensure that the model can handle various scenarios in real-world applications.
[0040] Second, input data: During training, the image input passes through the feature extraction layer of ViT, generating initial visual features. These features are then passed to the corresponding MLP, which performs mapping to ultimately output a feature representation consistent with the linguistic feature dimension.
[0041] Third, target optimization: The parameters of the MLP are adjusted by optimizing the loss function (e.g., based on the difference between the predicted language description and the actual language description). The specific optimization goal is to enable the representation after visual feature mapping to more accurately express the semantic information of the image in the linguistic space.
[0042] The goal of this training process is to enable each MLP to accurately map the visual features extracted by ViT to a space consistent with the language feature space, thereby improving the ability to fuse visual and linguistic information. Through training, the model's visual features will better meet the input requirements of the subsequent language decoder, further enhancing its performance in image selection tasks.
[0043] Step S3: Construct a training dataset for the image selection task, and fine-tune the trained visual model based on LoRA technology to obtain the image selection model. The training dataset for the image selection task includes images, selection condition text, and corresponding expected outputs.
[0044] First, it's crucial to define the goal of the image filtering task: to select images from a large dataset that meet specific filtering criteria. These criteria are typically based on the direct content of the image or semantic information related to that content, and should be in textual form, such as the background, objects, or semantic descriptions of the image content. Therefore, the training dataset needs to include images, filtering criteria, and their corresponding expected filtering outputs. This allows the model to effectively filter images by learning the relationship between them and the target filtering criteria. For example, it could be structured as follows: First, image data: These images should be diverse and as rich as possible.
[0045] Second, filtering criteria: Each image has corresponding filtering criteria. These criteria can take many forms, including: those directly related to the image's content, such as "studio" or "Ragdoll cat"; and those related to the image's semantic content, such as "Does the following image match 'difficult parking in the neighborhood'?".
[0046] Third, expected output: For each image, there should also be expected output under the filtering conditions. Expected output can take various forms, such as filtering match degree (the degree of matching with the filtering conditions, e.g., "0.7"), filtering conclusion (whether it matches the filtering conditions, e.g., "compliant"), and can also include additional content such as reasons, explanations, and clarifications. Expected output can be manually annotated or semi-automatically annotated using a large visual model. Specifically, the image and filtering conditions are input into the large visual model, the model generates the output, and then the output is manually checked and corrected.
[0047] Therefore, this application constructs a training dataset specifically for image selection tasks and uses this dataset to deeply train a large-scale visual model, ultimately obtaining a well-trained large-scale visual model. This training dataset contains three key components: first, rich image resources, providing the model with the input foundation of visual information; second, detailed selection condition text, clearly defining the specific requirements and standards for image selection in text form, guiding the model to learn how to select images based on the conditions; and third, the corresponding expected output, i.e., the expected selection result of the image based on the selection conditions, used by the model to learn whether an image meets the selection conditions. By organically combining these three components, the model can continuously learn and optimize during training, accurately understanding the relationship between image content and selection conditions, thereby efficiently and accurately completing image selection tasks in practical applications.
[0048] This paper employs Low-Rank Adaptation (LoRA), a parameter-efficient fine-tuning method, to adapt a pre-trained large-scale visual model to downstream tasks without significantly increasing computational overhead. During fine-tuning, the core parameters of the visual encoder, visual-language fusion unit, and language decoder remain unchanged. Only a set of low-rank factorization matrices is inserted and trained as an adaptation layer to indirectly adjust weight updates during the model's forward propagation. By introducing the low-rank constraint, the decision boundaries and feature patterns specific to the image selection task can be effectively captured while greatly reducing the number of trainable parameters. The fine-tuning process uses the image selection training dataset constructed in the preceding steps as a supervision signal. By optimizing the task-specific loss function, the model ultimately learns to accurately match and semantically infer the input image and the selection condition text, thereby obtaining an efficient, lightweight image selection model with excellent generalization capabilities.
[0049] The steps for fine-tuning the trained large-scale visual model based on LoRA technology to obtain the image selection model include: Add a LoRA adaptation layer to the trained large visual model; The parameters of the LoRA adaptation layer are trained using the image filtering task training dataset, while keeping the parameters of the visual encoder, visual-language fusion unit, and language decoder in the large visual model unchanged, to obtain the image filtering model. The image filtering model is used to receive the image to be filtered and the corresponding filtering condition text, and output the judgment result of whether the image meets the filtering criteria.
[0050] After constructing the training dataset, the next task is to fine-tune the large visual model for image selection tasks using the LoRA (Low-Rank Adaptation) technique. LoRA is an efficient training method that optimizes large models through low-rank adaptation layers, reducing computational resource consumption while improving the model's adaptability to image selection tasks.
[0051] A LoRA adaptation layer is added to the original large visual model obtained in step two. Using the image filtering dataset constructed in step three, the parameters of the LoRA adaptation layer are trained, while keeping the parameters of other parts of the model unchanged. After training with LoRA technology, the resulting image filtering model will be able to determine whether an image meets the filtering criteria based on its content or label. The LoRA-optimized model not only achieves high accuracy in image filtering tasks but also achieves a good balance between computational efficiency and training cost.
[0052] Step S4: Analyze the images to be filtered and the filtering conditions using the image filtering model, and output the filtering results.
[0053] An image filtering model was fine-tuned and trained using the LoRA technique. The next task is to perform image filtering based on this model. For each image to be filtered, the model is input along with the image and filtering criteria text, and the model will output the filtered results. After all images have been processed, further actions can be taken based on the model's output, such as removing images that do not meet the criteria, saving images that do meet the criteria, or categorizing and summarizing the results. Ultimately, the image filtering task can be completed efficiently, and the model's judgment will help users quickly filter out images that meet the criteria.
[0054] Compared with the prior art, the embodiments of this application have the following beneficial effects: First, by utilizing the visual Transformer to extract multi-level features from images and mapping them through independent fusion machines, the model can obtain richer and more refined image representations than relying on a single final layer feature or text label. This enables the model to make judgments based on more comprehensive image information when faced with complex, ambiguous, or diverse screening conditions, thereby significantly improving the accuracy and reliability of the screening task.
[0055] Secondly, the visual-language fusion model employed can directly process image pixel data and align visual features with linguistic features in the latent space. This method avoids incomplete or biased textual summarization of image information, fundamentally solving the feature loss problem caused by information compression and ensuring the integrity of the original image information.
[0056] Third, a phased training strategy was adopted: First, the pre-trained parameters of the visual encoder and language decoder were frozen, and only a lightweight visual-language fusion module (MLP) was trained; in the final fine-tuning stage, LoRA (low-rank adaptive) technology was used, requiring only a few additional parameters to enable the model to accurately adapt to the downstream screening task. This strategy greatly reduced the demand for computing resources and training data, and shortened the training time.
[0057] The following provides a possible implementation of an image filtering system based on a large visual model, which performs the various execution steps and corresponding technical effects of the image filtering method shown in the above embodiments and possible implementations. The system includes: Build modules are used to construct improved large-scale visual models, which include a visual encoder, multiple visual-language fusion units, and a language decoder. The training module is used to train the multiple vision-language fusion processors using a pre-training dataset, while freezing the parameters of the visual encoder and the language decoder to obtain a trained large-scale vision model. The pre-training dataset contains images and corresponding text descriptions or labels.
[0058] The fine-tuning module is used to construct the training dataset for the image selection task and to fine-tune the trained visual model based on LoRA technology to obtain the image selection model. The training dataset for the image selection task includes images, selection condition text, and corresponding expected outputs. The analysis module is used to analyze the images to be filtered and the filtering conditions using an image filtering model, and output the filtering results.
[0059] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An image filtering method based on a large visual model, characterized in that, The method includes: An improved large-scale visual model is constructed, which includes a visual encoder, multiple visual-language fusion units, and a language decoder. In the visual encoder, each feature layer is mapped through a visual-language fusion unit. The visual encoder is a visual Transformer used to extract initial layer features, intermediate layer features, and final layer features from the input image. The initial layer features are used to capture low-level features of the image's edges and textures, the intermediate layer features are used to capture mid-level features of the image's shape and components, and the final layer features are used to capture high-level global semantic features of the image. The visual-language fusion unit is a multilayer perceptron used to map corresponding visual features to a space with the same dimension as the language features. The language decoder uses a large language model to receive mapped features and generate output; Multiple vision-language fusion machines are trained using a pre-trained dataset, while the parameters of the visual encoder and language decoder are frozen to obtain a large-scale vision model after training. The pre-trained dataset contains images and corresponding text descriptions or labels. A training dataset for the image selection task is constructed, and the trained visual model is fine-tuned based on LoRA technology to obtain the image selection model. The training dataset includes images, selection condition text, and corresponding expected outputs. The image filtering model is used to analyze the images to be filtered and the filtering conditions, and the filtering results are output.
2. The image filtering method as described in claim 1, characterized in that, The multilayer perceptron includes two hidden layers. The first layer maps visual features to an intermediate dimension, and the second layer maps the intermediate dimension to a space consistent with the language feature dimension.
3. The image filtering method as described in claim 1, characterized in that, The method further includes: The parameters of the vision-language fusion engine are adjusted by optimizing the loss function, which is calculated based on the difference between the predicted language description and the actual language description.
4. The image filtering method as described in claim 1, characterized in that, The steps for fine-tuning the trained large-scale visual model based on LoRA technology to obtain the image selection model include: Add a LoRA adaptation layer to the trained large visual model; The parameters of the LoRA adaptation layer are trained using the image filtering task training dataset, while keeping the parameters of the visual encoder, visual-language fusion unit, and language decoder in the large visual model unchanged, to obtain the image filtering model. The image filtering model is used to receive the image to be filtered and the corresponding filtering condition text, and output the judgment result of whether the image meets the filtering criteria.
5. An image filtering system based on a large visual model, characterized in that, The system includes: The building block is used to build an improved large visual model, which includes a visual encoder, multiple visual-language fusion units and a language decoder. In the visual encoder, each feature layer is mapped through a visual-language fusion unit. The visual encoder is a visual Transformer used to extract initial layer features, intermediate layer features, and final layer features from the input image. The initial layer features are used to capture low-level features of the image's edges and textures, the intermediate layer features are used to capture mid-level features of the image's shape and components, and the final layer features are used to capture high-level global semantic features of the image. The visual-language fusion unit is a multilayer perceptron used to map corresponding visual features to a space with the same dimension as the language features. The language decoder uses a large language model to receive mapped features and generate output; The training module is used to train the multiple vision-language fusion processors using a pre-training dataset, while freezing the parameters of the visual encoder and the language decoder to obtain a trained large-scale vision model. The pre-training dataset contains images and corresponding text descriptions or labels. The fine-tuning module is used to construct the training dataset for the image selection task and to fine-tune the trained visual model based on LoRA technology to obtain the image selection model. The training dataset for the image selection task includes images, selection condition text, and corresponding expected outputs. The analysis module is used to analyze the images to be filtered and the filtering conditions using an image filtering model, and output the filtering results.
Citation Information
Patent Citations
Image description generation method based on text hierarchical structure
CN113569932A
Label-free image screening and model training method and device, equipment and storage medium
CN116740497A