Photo album cleaning method and device, electronic equipment and medium

By initially screening images and videos in the album, and using multimodal large models to combine user personalized information to accurately identify and clean low-quality images and videos, the problem of inability to effectively clean up low-quality media in the album in the existing technology is solved, and personalized cleaning of the album and optimization of storage space is achieved.

CN120067059APending Publication Date: 2025-05-30CHINA MOBILE GROUP SICHUAN +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510126396.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology cannot effectively clean up low-quality pictures and videos in the album based on the user's personalized aesthetic, resulting in these low-quality media occupying a large amount of storage space and affecting the album's user experience.

Method used

By conducting preliminary screening of images and videos in the album, and using multimodal large models to fuse the feature vectors of images and videos, combining user personalized information, we accurately identify and clean low-quality images and videos.

Benefits of technology

It realizes personalized cleaning of images and videos in the album, improves the available storage space of the album, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067059A_ABST
    Figure CN120067059A_ABST
Patent Text Reader

Abstract

The invention provides a photo album cleaning method and device, electronic equipment and a medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: performing secondary preliminary screening on an original image set and an original video set in a photo album to obtain a first target image set formed by low-quality pictures and a first target video set formed by low-quality videos; and based on the to-be-detected image, the to-be-detected video and the user personalized information, fusing the image feature vector of the to-be-detected image and the video feature vector of the to-be-detected video through a multi-modal large model, outputting the quality evaluation of the to-be-detected image and the quality evaluation of the to-be-detected video, and determining the quality of the to-be-detected video according to the quality evaluation. Respectively determining a low-quality second target image set and a low-quality second target video set; and clearing the target image set and the target video set in the photo album. Through the technical scheme provided by the invention, the problem that low-quality images and videos in the photo album cannot be accurately cleaned based on personalized aesthetic appreciation of the user is solved, personalized cleaning of the photo album is realized, and the use experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to an album cleaning method, apparatus, electronic device, and medium. Background Art

[0002] With the popularization of smart phones, taking pictures and recording videos have become one of the most frequently used functions of users. During frequent photo-taking or recording processes, the pictures or videos accumulated in the album are also increasing day by day. Retaking or failed shooting will generate many low-quality pictures or videos, such as all-black, out-of-focus, subject occlusion, etc., as well as low-quality pictures and videos that do not meet the user's aesthetics.

[0003] The cleaning of low-quality pictures and videos in the album mainly relies on manual screening. Manually deleting a large amount of content will consume a lot of time and energy, seriously affecting the user experience. In related technologies, the expert system screening based on knowledge construction will use image algorithms to extract the basic features of images to filter images; or by collecting a large number of low-quality pictures as a data set, calculating the similarity between the current image and the data set to evaluate the picture quality; or screening black-edge images through the distribution structure of pixels; or evaluating the blur degree of images to filter out-of-focus images. Or, the above methods are combined to detect and clean low-quality images and videos. Since the above methods need to define fixed rules for low-quality images and lack generalization ability, when the data does not conform to the rules, the detection effect will be greatly reduced.

[0004] In order to improve the detection effect, a deep model is proposed to be trained based on a large number of low-quality pictures for classification and screening. This scheme needs to collect a large number of low-quality images and videos as a training set, and mix them with high-quality images to train an image quality classification model. Compared with the expert system screening method, it has a certain generalization ability, but the quality of the training data plays a decisive role in the final classification result, and the generalization ability is limited.

[0005] Furthermore, related technologies have proposed a technical solution based on the aesthetic scoring of AI models, using computers to simulate human perception of beauty and automatically evaluating the "aesthetic feeling" of images. Through a convolutional neural network, information such as image style and image content is modeled, and the model is used to score the aesthetics of the image, so as to screen out low-quality images for deletion. However, the aesthetic score is not completely equivalent to low-quality images, and the user's personalized aesthetics is closely related to the determination of low-quality pictures. To a certain extent, this technical solution cannot completely screen all low-quality images and videos, resulting in the inability to accurately clean low-quality pictures and videos based on the user's personalized aesthetics, and allowing low-quality images and videos to occupy a large amount of storage space in the album, affecting the use experience of the album. Summary of the Invention

[0006] The present disclosure provides an album cleaning method, apparatus, electronic device and medium to solve the problem that low-quality pictures and videos cannot be accurately cleaned based on the user's personalized aesthetics, and the low-quality images and videos occupy a large amount of storage space in the album, affecting the use experience of the album. By preliminarily screening the images and videos in the album and accurately identifying and cleaning the low-quality images and videos through a multimodal large model, personalized cleaning of the images and videos in the album is achieved, the available storage space of the album is increased, and the user experience of the album is improved.

[0007] According to a first aspect embodiment of the present disclosure, an album cleaning method is provided, and the method includes:

[0008] Perform a secondary preliminary screening on the original image set and the original video set in the album to obtain a first target image set formed by low-quality pictures and a first target video set formed by low-quality videos;

[0009] Based on the images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected, and the user's personalized information, a multimodal large model is used to fuse the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected, and a quality evaluation of the images to be detected and the videos to be detected is output. According to the quality evaluation, a second target image set and a second target video set of low quality in the first image set to be detected and the first video set to be detected are respectively determined, where the image feature vectors and the video feature vectors have the same length, the first image set to be detected is the difference set between the original image set and the first target image set, and the first video set to be detected is the difference set between the original video set and the first target video set;

[0010] Delete at least one of the first target image set, the first target video set, the second target image set, and the second target video set in the album.

[0011] In an embodiment of the present disclosure, based on the images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected, and the user's personalized information, a multimodal large model is used to fuse the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected, and a quality evaluation of the images to be detected and the videos to be detected is output. According to the quality evaluation, a second target image set and a second target video set of low quality in the first image set to be detected and the first video set to be detected are respectively determined, including:

[0012] Based on the images to be detected in the first image set to be detected and the videos to be detected in the first video set to be detected, a multimodal large model is used to extract features from the images to be detected and encode them into image feature vectors in the image modality, and extract features from the videos to be detected and encode them into video feature vectors in the video modality;

[0013] Based on the image feature vector and the video feature vector, determine the target query vector for aligning the visual modality to the text modality, where the visual modality includes the image modality and the video modality, and the image feature vector and the video feature vector have the same length;

[0014] According to the user's personalized information, parse the target query vector through the base large language model of the multimodal large model, and output the quality evaluation of the image to be detected and the video to be detected. Use the pictures with low-quality evaluation in the first set of images to be detected as the second target image set, and use the videos with low-quality evaluation in the first set of videos to be detected as the second target video set.

[0015] In one embodiment of the present disclosure, before extracting the features of the image to be detected through the multimodal large model and encoding them into the image feature vector of the image modality, and extracting the features of the video to be detected and encoding them into the video feature vector of the video modality, it further includes:

[0016] Obtain video samples, image samples, and fine-tuning text, where the fine-tuning text is generated based on the preset text descriptions of the video samples and image samples and the user portrait of the album;

[0017] Perform data differencing on the video samples to obtain the first sample video;

[0018] Divide the first sample video and the image samples using the same-sized sliding window, and respectively extract the first video features of the first sample video in the sliding window and the first image features of the image samples in the sliding window through the feature extraction network;

[0019] Map the first video features and the first image features into the first video feature vector and the first image feature vector respectively, and the first video feature vector and the first image feature vector have the same length;

[0020] Based on the first video feature vector, the first image feature vector, and the fine-tuning text, train the pre-trained multimodal large model by means of instruction fine-tuning to obtain the multimodal large model.

[0021] In one embodiment of the present disclosure, before parsing the target query vector through the base large language model of the multimodal large model according to the user's personalized information, it further includes:

[0022] Generate a set of learnable query vectors with a fixed length based on the image feature vector and the video feature vector;

[0023] According to the learnable query vectors, generate the query embedding representation through the self-attention mechanism and the cross-attention mechanism;

[0024] The query embedding representation is linearly projected through a fully connected layer to obtain a target query vector with the same dimension as the text embedding of the base large language model;

[0025] The target query vector, user personalization information, and the identifier of the visual modality are used as the input of the base large language model.

[0026] In one embodiment of the present disclosure, according to the user personalization information, the target query vector is parsed by the base large language model of the multimodal large model, and the quality evaluations of the images and videos to be detected are output, including:

[0027] Based on the user's historical operation data of the photo album, a user portrait is created;

[0028] According to the user portrait and the personalized requirement text, user personalization information is generated;

[0029] Using the user personalization information as the prompt of the base large language model, the target query vector is parsed by the base large language model to generate the quality evaluations of the images and videos to be detected.

[0030] In one embodiment of the present disclosure, when clearing at least one of the first target image set, the first target video set, the second target image set, and the second target video set in the photo album, it further includes:

[0031] In response to the user selecting the target images and target videos to be deleted from the first target image set, the first target video set, the second target image set, and the second target video set, the target images and target videos are deleted from the photo album.

[0032] In one embodiment of the present disclosure, the original image set and the original video set in the photo album are subjected to secondary preliminary screening to obtain a first target image set formed by low-quality pictures and a first target video set formed by low-quality videos, including:

[0033] Based on the original images to be detected in the original image set, the basic image attributes are used to screen out low-quality pictures to form a first intermediate image set, and the quality of the original images to be detected is classified according to the low-quality image classification model, and the set composed of the low-quality original images to be detected in the original image set is used as the second intermediate image set;

[0034] Based on the original videos to be detected in the original video set, the low-quality videos are detected by frame extraction of the videos to be detected to form a first intermediate video set, and the quality of the original videos to be detected is classified according to the low-quality video classification model, and the set composed of the low-quality original videos to be detected in the original video set is used as the second intermediate video set;

[0035] Take the union of the first intermediate image set and the second intermediate image set as the first target image set, and take the union of the first intermediate video set and the second intermediate video set as the first target video set.

[0036] An embodiment of the second aspect of the present disclosure provides an album cleaning device, which includes:

[0037] A preliminary screening module, configured to perform a secondary preliminary screening on the original image set and the original video set in the album to obtain a first target image set formed by low-quality pictures and a first target video set formed by low-quality videos;

[0038] A quality evaluation module, configured to, based on the images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected, and personalized information, fuse the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected through a multimodal large model, and output the quality evaluations of the images to be detected and the videos to be detected. According to the quality evaluations, respectively determine a second target image set and a second target video set with low quality in the first image set to be detected and the first video set to be detected. Among them, based on the images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected, and the user's personalized information, fuse the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected through a multimodal large model, and output the quality evaluations of the images to be detected and the videos to be detected. According to the quality evaluations, respectively determine a second target image set and a second target video set with low quality in the first image set to be detected and the first video set to be detected. Among them, the image feature vectors and the video feature vectors have the same length. The first image set to be detected is the difference set between the original image set and the first target image set, and the first video set to be detected is the difference set between the original video set and the first target video set;

[0039] A cleaning module, configured to clean at least one of the first target image set, the first target video set, the second target image set, and the second target video set in the album.

[0040] An embodiment of the third aspect of the present disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of the embodiments of the first aspect of the present disclosure.

[0041] An embodiment of the fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to cause a computer to execute the method according to the embodiments of the first aspect of the present disclosure.

[0042] An embodiment of the fifth aspect of the present disclosure provides a computer program product, which includes a computer program that, when executed by a processor, implements the method according to any one of the embodiments of the first aspect of the present disclosure.

[0043] In summary, the album cleaning method proposed according to the present disclosure can achieve the following beneficial effects:

[0044] Perform a secondary preliminary screening on the original image set and the original video set in the album to obtain a first target image set formed by low-quality pictures and a first target video set formed by low-quality videos, screening out significantly low-quality pictures and videos, thereby avoiding subjecting all the contents of the album to precise recognition by the subsequent multimodal large model, which helps improve the screening efficiency;

[0045] Based on the images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected, and the user's personalized information, the multimodal large model fuses the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected, and outputs the quality evaluations of the images to be detected and the videos to be detected. According to the quality evaluations, respectively determine the second target image set and the second target video set with low quality in the first image set to be detected and the first video set to be detected, where the image feature vectors and the video feature vectors have the same length, the first image set to be detected is the difference set between the original image set and the first target image set, and the first video set to be detected is the difference set between the original video set and the first target video set. Combine the images / videos in the album whose quality is difficult to distinguish with the user's personalized information to screen out the low-quality images / videos that do not meet the user's aesthetic standards.

[0046] Delete at least one of the first target image set, the first target video set, the second target image set, and the second target video set in the album, completing the intelligent cleaning of the low-quality images / videos in the album, realizing the personalized cleaning of the album, and improving the user experience.

[0047] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0049] Figure 1 It is a flowchart of an album cleaning method according to an embodiment of the present disclosure;

[0050] Figure 2 It is an algorithm architecture diagram of a personalized album cleaning multimodal large model according to an embodiment of the present disclosure;

[0051] Figure 3 The working flowchart of the Q-Former for the embodiments of the present disclosure;

[0052] Figure 4 Schematic diagram of the inference process of a personalized photo album cleaning multimodal large model for the embodiments of the present disclosure;

[0053] Figure 5 Schematic diagram of the structure of an album cleaning device for the embodiments of the present disclosure;

[0054] Figure 6 Block diagram of an electronic device for implementing the photo album cleaning method of the present disclosure shown according to an exemplary embodiment;

[0055] Figure 7 Schematic diagram of the structure of the chip for the embodiments of the present disclosure. Detailed implementation manners

[0056] The embodiments of the present disclosure will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, but should not be construed as a limitation of the present disclosure.

[0057] First, a brief introduction to the relevant terms in the present disclosure:

[0058] Multimodal Large Models: In the present disclosure, it refers to large artificial intelligence models that can process and fuse input data from different types (such as text, images, voice, video, etc.). Such models can simultaneously understand and generate various forms of information, not limited to a single modality (such as only text or only images). For example, the Generative Pretrained Transformer 4 (GPT-4) model of OpenAI can be used as a multimodal model to process text and image inputs. Users can not only ask it text questions but also upload images, and the model can analyze the image content and provide comprehensive answers in combination with the text. Specifically, the characteristics of multimodal large models include: Multiple input types: They can accept multiple data inputs (such as text, images, audio, etc.) and understand the relationships between them. Data fusion: They can combine information from different modalities for comprehensive processing to provide a more comprehensive and accurate understanding. Cross-modal generation ability: They can generate different forms of outputs based on different types of inputs, such as generating descriptive text from images and generating images from text.

[0059] Large Language Model (LLM): In this disclosure, it refers to a natural language processing model trained using deep learning techniques, especially models based on the Transformer architecture, with a large amount of text data. These models usually have billions to trillions of parameters and can capture the complex structures and semantic rules of language. The training process of large language models is divided into two stages: pre-training and fine-tuning. In the pre-training stage, a wide range of text data is used to learn the basic patterns of language, and in the fine-tuning stage, it is optimized for specific tasks. Large language models possess powerful generation and understanding capabilities. It can not only generate natural and coherent text but also perform language tasks such as question answering, translation, and summarization. By analyzing the context, it can understand and process complex language inputs and provide high-quality outputs.

[0060] Instruction Tuning: In this disclosure, it refers to a method for optimizing a pre-trained language model for specific tasks. It uses training data containing clear instructions and examples to help the model better understand how to respond to different types of inputs. Different from traditional fine-tuning methods, instruction tuning focuses on enabling the model to learn how to execute tasks according to specific instructions rather than just optimizing the model's performance on a specific dataset. Instruction tuning usually uses a dataset containing clear task descriptions and corresponding examples, and the model learns how to generate compliant outputs through this instruction-based data. For example, when the model sees "Please give a summary of climate change", it can understand the task requirements and generate a relevant summary. This method can significantly improve the flexibility and reliability of the model in practical applications, enabling it to show higher adaptability and task execution capabilities when facing various instructions. Instruction tuning is widely used in fields such as intelligent assistants and dialogue systems.

[0061] During the process of taking photos or recording videos, end-side devices such as mobile phones, tablets, cameras, etc. basically all have high-quality cameras. However, during the process of taking photos or videos, users have a large number of duplicate or failed shots. If not processed in a timely manner, it will occupy a large amount of storage space on end-side devices such as mobile phones, and at the same time, it will also affect the user experience when browsing photos later. During continuous and extensive shooting processes such as traveling, if users delete them manually, it will consume a large amount of time and energy, seriously affecting the user experience; how to efficiently and personalized screen out images and videos that meet the user's aesthetic and are truly "good" in quality has become a problem presented to end users.

[0062] Currently, the relatively mainstream technical solutions for screening low-quality pictures and videos mainly include manual screening, screening based on knowledge to build an expert system, classification screening using deep models trained with a large number of low-quality images, and screening based on an AI image aesthetics scoring model. The specific implementation processes are as follows:

[0063] The rule filtering technical solution based on the expert system mainly uses image algorithms and basic image detection to complete the extraction of basic features of the image, thereby filtering the image. Usually, indicators such as the mean square error and peak signal-to-noise ratio of the image are calculated for judgment, and when certain indicators are too high, they will be used as the judgment basis. At the same time, a large number of low-quality picture sets can also be collected, and then the quality of the pictures can be determined by calculating the structural similarity between the current image and the low-quality picture sets. Or the black-edge images can be screened through the distribution structure of the image pixels, or filtered according to the blur degree of the image itself. This solution also combines and cross-verifies some previous expert knowledge and the aforementioned rules to form an expert system for filtering low-quality pictures to screen low-quality pictures. Relatively speaking, the rules and methods have limited knowledge inclusion, the rules themselves are relatively fixed, and lack generalization ability. When data outside the training data distribution appears, the processing effect will be greatly reduced.

[0064] The solution for classification and screening based on a deep model trained with a large number of low-quality images. This solution requires collecting a large number of low-quality pictures and low-quality videos as the training set, and then mixing the low-quality picture set and the high-quality picture set to train a picture quality classification model. Generally, features can be extracted and a classifier can be trained using the method of Support Vector Machine (SVM), or a Convolutional Neural Network (CNN) classification model such as AlexNet, VGG, GoogLeNet, etc. This solution is simple and fast. Compared with traditional feature extraction methods, it does not require manual feature extraction and can better classify low-quality pictures, with a certain generalization ability. However, the quality of the training data plays a decisive role in the final classification result.

[0065] A technical solution based on AI model aesthetic scoring. This solution mainly uses computers to simulate human perception and cognition of beauty, automatically evaluates the "aesthetics" of images, and analyzes the aesthetic stimulation formed by the image under the influence of aesthetic factors such as composition, color, light and shadow, depth of field, and virtuality. This solution uses image recognition convolutional neural networks, and explicitly or implicitly models additional information such as image style and image content in the modified convolutional neural network, so as to use the model to score the aesthetic degree of the image and filter out low-scoring images as low-quality images. The general aesthetic scoring network model consists of three parts: feature extraction part, attention part, and classification and regression part. The feature extraction part generally uses the efficient EfficientNet, the attention part uses a combination of position attention and channel attention, and the classification and regression part is designed to classify first and then regress. The introduction of aesthetic scoring classification is to guide the aesthetic regression task with aesthetic classification. The main basis is that aesthetic classification is a weak classification, and there is no strict boundary between categories like object recognition. Therefore, the method of classification first and then regression can improve the performance of single numerical evaluation of aesthetics. The judgment of this solution is more objective, reliable and fast. However, aesthetic scores are not completely equivalent to low-quality images. To a certain extent, they cannot completely filter out all low-quality images. They are more effective by comparing obviously low-quality and poorly shot video images.

[0066] In summary, the relevant technology cannot accurately clean up low-quality pictures and videos based on the user's personalized aesthetics, causing low-quality images and videos to occupy a large amount of storage space in the album, affecting the user experience of the album.

[0067] The method proposed in the present disclosure is applied to album cleaning tasks, and its application has a wide range of scenarios and can be applied to multiple technical fields and scenarios, mainly including the following aspects: In the field of digital album management, in the personalized album cleaning scenario, this technology can be used in smart album applications to help users automatically identify and clean up low-quality or duplicate photos and videos, and improve storage efficiency. In the intelligent classification and search scenario, through the image recognition capability of the large model, more refined photo classification can be achieved, and users can more easily find photos of specific scenes, people or events.

[0068] In the field of cloud storage services, in the scenario of photo backup and management, in a cloud storage platform, using this solution can enable users to seamlessly backup and manage photos between different devices, ensuring data security and easy access. In the asset management scenario, for users who need to manage a large number of pictures and videos (such as photographers and social media content creators), this technology can provide efficient asset classification and retrieval functions. In the field of social media and sharing platforms, in the scenario of content generation and sharing, in social media applications, this technology can be used to generate personalized videos or photo albums, enhancing the user's sharing experience and attracting more interactions. In the intelligent recommendation system scenario, by analyzing the user's photo and video content, the system can provide personalized content recommendations, increasing user stickiness. In the enterprise application field, in the scenario of marketing and brand management, enterprises can use this technology to intelligently manage product pictures and promotional videos, improving the efficiency of marketing activities. In the customer relationship management scenario, in customer service, by automatically processing the picture materials uploaded by customers, customer relationships can be better maintained and service quality can be improved.

[0069] Particularly in the following specific scenarios, such as family album management: providing convenient photo sorting, cleaning, and sharing functions for family users to help them easily record the little details in life. Travel and event recording: After a travel or event, automatically sort the relevant photos and videos and generate a memory video, providing a wonderful memory experience for users. Such as education and training: In the field of education, it can be used for the sorting and display of students' portfolios to help teachers and students better manage learning materials. Such as medical image processing: In the medical industry, this technology can be used to sort patients' image materials, improving the doctor's management efficiency of cases. The application scenarios are not limited in the embodiments of the present disclosure. Through the above application fields and scenarios, it can be seen that this technical solution has wide applicability and potential, not only improving the usage experience of individual users but also bringing more efficient data management solutions to enterprises.

[0070] The following will introduce the photo album cleaning method provided by the present disclosure in detail with reference to the accompanying drawings.

[0071] Figure 1 It is a flowchart of a photo album cleaning method according to an embodiment of the present disclosure. As Figure 1 shown in the embodiment, the photo album cleaning method includes:

[0072] Step 101, perform a secondary preliminary screening on the original image set and the original video set in the photo album to obtain a first target image set formed by low-quality pictures and a first target video set formed by low-quality videos.

[0073] Among them, the original image set refers to all image files stored in the electronic album of the terminal device, and the formats include jpg, png, bmp, gif, svg, tiff, etc. The original video set refers to all video files stored in the electronic album of the terminal device, and the formats include mp4, avi, mov, wmv, mkv, flv, etc. The present disclosure does not limit the formats of images and videos in the album. Low-quality pictures and low-quality videos generally refer to media content with obvious visual distortion, blurriness, color deviation, or lack of details, including images / videos affected by overexposure, out-of-focus, jitter, low light, strong light, noise, etc., and also include images / videos that the user of the album subjectively deems not meeting their personalized aesthetics. The first target image set refers to the set formed by the low-quality pictures obtained after the second-level preliminary screening of the images in the original image set. The first target video set refers to the set formed by the low-quality videos obtained after the second-level preliminary screening of the videos in the original video set.

[0074] Further, the second-level preliminary screening includes the preliminary screening of low-quality graphics / videos caused by objective factors such as overexposure, out-of-focus, jitter, low light, strong light, noise, etc., to complete the detection of significantly low-quality content. It also includes the secondary screening based on a deep neural network through an image / video quality classification model trained on a low-quality image / video data set. The secondary screening is a supplement to the preliminary screening. Optionally, the second-level screening undergoes at least one of the preliminary screening and the secondary screening.

[0075] Step 102: Based on the images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected, and the user's personalized information, the multi-modal large model fuses the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected, and outputs the quality evaluations of the images to be detected and the videos to be detected. According to the quality evaluations, the second target image set and the second target video set with low quality in the first image set to be detected and the first video set to be detected are respectively determined. Among them, the image feature vectors and the video feature vectors have the same length. The first image set to be detected is the difference set between the original image set and the first target image set. The first video set to be detected is the difference set between the original video set and the first target video set.

[0076] In this disclosure, the first image set to be detected refers to the difference set between the original image set and the first target image set in the album, that is, among all the images in the album, the remaining images after excluding the images that have been detected as low-quality in the secondary preliminary screening. The first video set to be detected refers to the difference set between the original video set and the first target video set in the album, that is, among all the videos in the album, the remaining videos after excluding the videos that have been detected as low-quality in the secondary preliminary screening. The images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected, and the user's personalized information are used as the input of the multimodal large model. Among them, the user's personalized information includes a user portrait constructed based on the user's historical usage information or usage habits of the album, or descriptive information about the user's aesthetics of images and videos.

[0077] Extract the image features of the images to be detected and the video features of the videos to be detected through the feature extraction network in the multimodal large model. Utilize the modality fusion ability of the multimodal large model to map the image features in the image modality to image feature vectors, and map the video features in the video modality to video feature vectors. Among them, the image feature vectors and the video feature vectors have the same length, which is convenient for normalizing the features of images and videos and the corresponding language text features of the user's personalized information, and providing processed data of the same length for subsequent calculations.

[0078] Based on the image feature vectors, video feature vectors, and language text features, the multimodal large model can obtain the quality evaluation of the images to be detected and the videos to be detected. This quality evaluation can be a text description in the form of words, a quality score for the image / video, or a program interface that conforms to a specific structure. According to the detection marks for low-quality images / videos in the quality evaluation, the second target image set composed of low-quality images can be obtained from the first image set to be detected, and the second target video set composed of low-quality videos can be obtained from the first video set to be detected, so as to obtain low-quality images / videos that do not conform to aesthetics from the user's subjective perspective.

[0079] Step 103, clear at least one of the first target image set, the first target video set, the second target image set, and the second target video set in the album.

[0080] Based on the first target image set composed of low-quality images obtained through the secondary screening in the foregoing steps, the first target video set obtained from low-quality videos, as well as the second target image set composed of low-quality images and the second target video set composed of low-quality videos detected by the multimodal large model. As candidates for the images / videos that can be cleared in the album, and then automatically clear some or all of the content in the image set and video set composed of the first target images, the first target video set, the second target image set, and the second target video set.

[0081] Furthermore, other variables can be combined, such as time, number of views, file size of individual images / videos, etc., to automatically clean the content in the above low-quality image / video set. Intelligent cleaning of low-quality images / videos in the album is achieved to save storage space and enhance the usage experience of the album.

[0082] Figure 2 It is an algorithm architecture diagram of a personalized album cleaning multi-modal large model according to an embodiment of the present disclosure. As Figure 2As shown, the multi-modal large model algorithm is completed through the collaboration of the edge side (terminal) and the cloud side (cloud). Optionally, based on the X-InstructBLIP multi-modal large model architecture, the construction of X-InstructBLIP uses the framework of the LAVIS software library. On the edge side, based on the photos and videos in the terminal device's photo album, using EVA-CLIP-VIT-G / 14 as the frozen pre-trained encoder for video and image modalities, a multi-modal large model can obtain content understanding and text generation from photos and videos. In other words, this multi-modal large model can recognize and understand the objects, scenes, tasks, and their activities in photos and videos. For example, it can distinguish between a cat or a dog in a picture or understand a football game in progress in a video. For images, pixel-level semantic segmentation can be performed, that is, identifying and labeling the category to which each pixel belongs, such as sky, grass, building, etc. For videos, it can recognize the actions and activities in them, such as running, jumping, waving, etc. It can also generate text descriptions corresponding to the pictures or videos. Based on the preliminary understanding results of the basic multi-modal large model for photos and videos, the input Modality-X Q-Former is used to bridge the modality differences between the visual modality (image / video) and the language modality (text) for deeper semantic analysis and reasoning. For example, understanding the emotional state of the people in a photo or the development context of an event in a video. More context information can also be combined to better understand the details in photos and videos. For example, connecting the objects in a photo with their usage methods in a video to provide a more comprehensive explanation. It can also generate more detailed and rich text descriptions based on the preliminary descriptions of the basic multi-modal large model for photos and videos. For example, it can not only describe the scene in a photo but also add background stories, emotional colors, etc. For videos, Modality-X Q-Former can generate more complex storylines based on the existing plot and even create new plot developments. Preferably, through the output results of the basic multi-modal large model for photos and videos, that is, the visual modality features composed of the image feature vectors corresponding to the photos and the video feature vectors corresponding to the videos, more accurate cross-modal retrieval can be performed through Modality-X Q-Former. For example, finding images and videos that match the low-quality content identified by the user. The Modality-X Q-Former uses the method of instruction fine-tuning to achieve the instruction fine-tuning guidance for the discrimination ability of the multi-modal large model for low-quality images and videos, improving the model's discrimination ability for low-quality pictures / videos. Thus, the personalized photo album cleaning multi-modal large model can simultaneously have the ability to process pictures and videos. After the content of photos and videos is understood as feature vectors of equal length in the text space through the multi-modal large model on the edge side, the LLM in the cloud is called for quality-related semantic analysis to provide a judgment result on the quality of the photo or video.

[0083] In summary, according to the album cleaning method proposed in the present disclosure, a secondary preliminary screening is performed on the original image set and the original video set in the album to obtain a first target image set formed by low-quality pictures and a first target video set formed by low-quality videos, screening out significantly low-quality pictures and videos, thereby avoiding subjecting all the contents of the album to accurate recognition by the subsequent multi-modal large model, which helps to improve the screening efficiency;

[0084] Based on the images to be detected in the first set of images to be detected, the videos to be detected in the first set of videos to be detected, and the user's personalized information, the multi-modal large model fuses the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected, and outputs the quality evaluations of the images to be detected and the videos to be detected. According to the quality evaluations, the second target image set and the second target video set of low quality in the first set of images to be detected and the first set of videos to be detected are respectively determined. Among them, the image feature vectors and the video feature vectors have the same length. The first set of images to be detected is the difference set between the original image set and the first target image set, and the first set of videos to be detected is the difference set between the original video set and the first target video set. By combining the images / videos in the album that are difficult to distinguish in terms of quality with the user's personalized information, low-quality images / videos that do not meet the user's aesthetic standards are screened out.

[0085] Delete at least one of the first target image set, the first target video set, the second target image set, and the second target video set in the album, completing the intelligent cleaning of low-quality images / videos in the album, saving the storage space occupied by the album, realizing the personalized cleaning of the album, and improving the user experience.

[0086] In a possible embodiment of the present disclosure, based on the images to be detected in the first set of images to be detected, the videos to be detected in the first set of videos to be detected, and the user's personalized information, the multi-modal large model fuses the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected, and outputs the quality evaluations of the images to be detected and the videos to be detected. According to the quality evaluations, the second target image set and the second target video set of low quality in the first set of images to be detected and the first set of videos to be detected are respectively determined, which can be realized through the following steps:

[0087] Based on the images to be detected in the first set of images to be detected and the videos to be detected in the first set of videos to be detected, the multi-modal large model extracts features from the images to be detected and encodes them into image feature vectors in the image modality, and extracts features from the videos to be detected and encodes them into video feature vectors in the video modality.

[0088] Specifically, based on the images to be detected in the first set of images to be detected, the multi-modal large model's image feature extraction network extracts features from the images to be detected and encodes the image features into image feature vectors in the image modality.

[0089] Similarly, based on the videos to be detected in the first set of videos to be detected, the video feature extraction network of the multi-modal large model is used to extract features from the videos to be detected, and the video features are encoded into video feature vectors in the video modality.

[0090] Among them, the image feature vector and the video feature vector are vector representations of the same length.

[0091] Based on the image feature vector and the video feature vector, determine the target query vector for aligning the visual modality to the text modality.

[0092] In this step, according to the image feature vector and the video feature vector, the multi-modal large model aligns the visual modality to the text modality to obtain the target query vector.

[0093] It can be understood that the visual modality includes the image modality and the video modality.

[0094] According to the user's personalized information, the base large language model of the multi-modal large model is used to parse the target query vector, and the quality evaluation of the images and videos to be detected is output. The images with low quality evaluation in the first set of images to be detected are used as the second target image set, and the videos with low quality evaluation in the first set of videos to be detected are used as the second target video set.

[0095] Specifically, according to the user's personalized information, for example, the user portrait constructed based on the user's historical operation data of the photo album, user preferences, basic information such as the shooting time and location of the image / video, and multi-modal instructions composed of personalized needs, the base large language model of the multi-modal large model parses the target query vector and outputs the quality evaluation of the images and videos to be detected. Among them, the base large language model usually refers to a large-scale, pre-trained language model used to understand and generate natural language text. Through the quality evaluation information output by the base large language model, the quality of the images / videos to be detected can be determined. The images with low quality evaluation in the first set of images to be detected are used as the second target image set, and the videos with low quality evaluation in the first set of videos to be detected are used as the second target video set. The aesthetic preference of the user is obtained from the user's personalized information, and then the multi-modal large model is used to analyze the low-quality images that do not conform to the user's aesthetics in the text modality description of the image / video to form the second target image set, and the low-quality videos that do not conform to the user's aesthetics form the second target video set.

[0096] Further, through the feature extraction and comparison of images / videos in multiple modalities in this embodiment, graphics / videos with very similar or identical features can be screened out. Here, very similar means the case where the difference between the feature vectors of two or more images is less than a preset threshold. This difference can be represented by the Mean Square Error (MSE), Mean Absolute Error (MAE), Root Mean Square Error (RMSE), etc.

[0097] Further, according to the personalized requirements of the user, guidance is provided through natural language. For example, it is required to delete portrait photos with closed eyes, pictures with facial occlusion, out-of-focus images or videos in the album.

[0098] In this embodiment, by using multimodal data and user personalized information, images / videos in the album that are difficult to distinguish in terms of quality are combined with personalized information to screen out low-quality images / videos that do not meet the user's aesthetic standards. The multimodal large model for processing multimodal data can be trained through the following embodiments.

[0099] In a possible embodiment of the present disclosure, before the feature extraction of the image to be detected by the multimodal large model and encoding it into an image feature vector in the image modality, and the feature extraction of the video to be detected and encoding it into a video feature vector in the video modality, the following steps are further included:

[0100] Obtain video samples, image samples, and fine-tuning text, where the fine-tuning text is generated based on preset text descriptions of the video samples and image samples and the user portrait of the album.

[0101] Perform data differencing on the video samples to obtain the first sample video.

[0102] Divide the first sample video and the image samples using sliding windows of the same size, and respectively extract the first video features of the first sample video in the sliding window and the first image features of the image samples in the sliding window through a feature extraction network.

[0103] Map the first video features and the first image features into the first video feature vector and the first image feature vector respectively, and the first video feature vector and the first image feature vector have the same length.

[0104] Based on the first video feature vector, the first image feature vector, and the fine-tuning text, train the pre-trained multimodal large model through instruction fine-tuning to obtain the multimodal large model.

[0105] In this embodiment, based on a multi-modal large model with general capabilities, the quality evaluation ability is trained by means of instruction fine-tuning using a small number of labeled samples.

[0106] First, obtain video samples, image samples, and fine-tuning text from local or network. Optionally, the fine-tuning text is the quality description text corresponding to the video samples and image samples, or it can be the preset text description of the video samples and image samples and the user portrait generated by the user using album data.

[0107] Among them, since video data includes multiple frames of image data, in order to reduce the number of frames to be processed, the frame set with large data frame changes is extracted from the video samples using data difference, that is, the first sample video, to replace the original video samples, which can compress most of the repeated and meaningless frames.

[0108] Then, in order to unify the video and pictures for combined processing, it is designed that the image and video are uniformly subdivided into n*n blocks in units of images, where n is a positive integer. Each block is used as a minimum processing unit token for extracting text features by the feature extraction model. The video and image can both be input into the corresponding feature extraction network, only the number of tokens of the two is different. Finally, it is mapped and learned onto a fixed-length feature vector.

[0109] In this way, different videos and images are all mapped to a fixed-size feature length through the mapping learning network, obtaining the first video feature vector and the first image feature vector, and flattening the spatial mapping between the video, picture, and language text features to achieve feature normalization.

[0110] Finally, using the first video feature vector, the first image feature vector, and the fine-tuning text as the training set, the multi-modal large model with pre-trained parameters of the visual modality frozen is trained by means of instruction fine-tuning to obtain the multi-modal large model.

[0111] In one implementation manner of this embodiment, the training of the multi-modal large model is implemented through the following logical code:

[0112] Among them, the modality identifier of the image is I, the corresponding low-quality image data set is DI, the modality identifier of the video is V, and the corresponding low-quality video data set is Dv.

[0113] 1: Initialize tokenizer h, embedding layer E

[0114] 2: for each modality M in {I, V} do

[0115] 3: Initialize modality-specific pre-trained encoder EncM

[0116] 4: Initialize Q-Former module QFM with BLIP-2 stage-1 weights

[0117] 5: Initialize modality-specific linear projection layer LPM

[0118] 6: For each step in number of iterations do

[0119] 7: Sample (x, y) from SDM

[0120] 8: Sample iM from IMt where t is the task mapping x to y

[0121] 9:

[0122] 10:

[0123] 11: xLLM ← E(h(cM)) ∥ LPM(QM) ∥ E(h(iM)) ∥ E(h(x))

[0124] 12:

[0125] 13:

[0126] 14:

[0127] 15: End for

[0128] 16: End for

[0129] After training the multimodal large model, it provides an analysis engine for content understanding and parsing of image and video related quality, which helps to improve the judgment of the quality of images / videos for users' personalized aesthetic decisions in the photo album. Among them, before using the base large language model of the multimodal large model to generate text for visual multimodal information, the participation of query vectors is also required.

[0130] In a possible embodiment of the present disclosure, before parsing the target query vector through the base large language model of the multimodal large model according to the user's personalized information, the following steps are further included:

[0131] Generate a set of learnable query vectors with a fixed length based on the image feature vectors and video feature vectors.

[0132] According to the learnable query vectors, generate query embedding representations through the self-attention mechanism and cross-attention mechanism.

[0133] Linearly project the query embedding representations through a fully connected layer to obtain target query vectors with the same dimension as the text embeddings of the base large language model.

[0134] Use the target query vectors, user personalized information, and the identifier of the visual modality as the input to the base large language model.

[0135] In this embodiment, the image feature vectors and video feature vectors extracted from images / videos by the multimodal large model are used to generate a set of learnable query vectors with a fixed length by the Q-Former model of images and videos. The learnable query vectors can be used as a query library for future query vectors to represent the vector form of user queries.

[0136] Generate query embedding representations of the learnable query vectors through the self-attention mechanism and cross-attention mechanism of the Q-Former model. Among them, the query embedding representation is the result of mapping the query vectors from high-dimensional data to low-dimensional space, and the query embedding representation is a continuous vector representation to capture deeper semantic information.

[0137] Linearly project the query embedding representations through a fully connected layer to achieve feature fusion between the visual modality and the text modality, and obtain target query vectors with the same dimension as the text embeddings of the base large language model.

[0138] Then, the target query vectors, user personalized information, and the identifier of the visual modality can be used as the input to the base large language model to obtain the quality evaluation of images / videos.

[0139] Specifically, the base large language model can use LLaMA, GPT-4, Vicuna, Claude 3, Gemini, Vicuna, etc. The present disclosure does not limit the base large language model.

[0140] Figure 3 This is the flowchart of the Q-Former for the embodiments of the present disclosure. As Figure 3As shown in the figure, the image / video is passed through the feature extraction network of the multimodal large model to extract features and perform image / video encoding to obtain the image / video feature vector, and the image / video feature vector is represented through image / video embedding. Based on the image / video embedding representation and the user's instruction text, the image and video Q-Former model is used to determine the query vector of the text modality from the learned query vectors. The instruction text, the query vector, and the modality prefix "{Image / Video}" are used as the input to the base large language model to perform quality rating on the image / video. This embodiment is mainly used to solve the modality differences between the two modalities of image and video and the text modality, so that the LLM can better understand the specific content of different modalities. At the same time, when the parameters of the pre-trained model are frozen, the learnable interface is responsible for connecting different modalities. It can effectively translate visual content into text that the LLM can understand.

[0141] In a possible embodiment of the present disclosure, according to the user's personalized information, the base large language model of the multimodal large model is used to parse the target query vector, and the quality evaluation of the image to be detected and the video to be detected is output, which can be realized through the following steps:

[0142] Based on the user's historical operation data of the photo album, a user portrait is created.

[0143] According to the user portrait and the personalized requirement text, user personalized information is generated.

[0144] Using the user's personalized information as the prompt of the base large language model, the base large language model is used to parse the target query vector to generate the quality evaluation of the image to be detected and the video to be detected.

[0145] In this embodiment, according to the user's historical operation records of the photo album, for example, recording the attributes of the photos and videos that the user often views, a user portrait is created for the user. For example, if the user often views images and videos of mountains, waters, trees, clouds, etc. in the photo album, a user portrait of a natural scenery lover is created for the user.

[0146] According to the user portrait and the user's personalized requirement text, for example, the user requests to delete unclear and abnormally colored photos, then the user's personalized information is generated, for example: If you are a natural scenery lover, please analyze the photos or videos in the photo album from the perspective of a natural scenery lover and select the photos or videos that are not clearly shot or have abnormal shooting colors.

[0147] Using the user's personalized information as the prompt of the base large language model, the base large language model can be used to parse the target query vector to generate the quality evaluation of the image or video represented in the target query vector.

[0148] In a possible embodiment of the present disclosure, when clearing at least one of the first target image set, the first target video set, the second target image set, and the second target video set in the album, the following implementation manners are further included:

[0149] In response to the user selecting the target images and target videos to be deleted from the first target image set, the first target video set, the second target image set, and the second target video set, delete the target images and target videos from the album.

[0150] Specifically, in order to facilitate the user to make a final deletion decision, the low-quality pictures / videos that have been automatically screened are displayed to the user. The display order can be random, can be sorted according to file size, or can be sorted according to the quality evaluation score. Optionally, display an identification for prompting deletion for the images or videos with large file size and / or low quality score. This facilitates the user to select and delete.

[0151] In a possible embodiment of the present disclosure, performing a secondary preliminary screening on the original image set and the original video set in the album to obtain the first target image set formed by low-quality pictures and the first target video set formed by low-quality videos, includes:

[0152] Based on the original images to be detected in the original image set, use the basic image attributes to screen for low-quality pictures to form a first intermediate image set, and perform quality classification on the original images to be detected according to the low-quality image classification model. The set composed of the low-quality original images to be detected in the original image set is used as the second intermediate image set.

[0153] Based on the original videos to be detected in the original video set, detect low-quality videos by extracting frames from the videos to be detected to form a first intermediate video set, and perform quality classification on the original videos to be detected according to the low-quality video classification model. The set composed of the low-quality original videos to be detected in the original video set is used as the second intermediate video set.

[0154] Use the union of the first intermediate image set and the second intermediate image set as the first target image set, and use the union of the first intermediate video set and the second intermediate video set as the first target video set.

[0155] In this embodiment, the original image set is a set composed of all image files in the album, and the original video set is a set composed of all video files in the album. The original image to be detected refers to the image in the original image set that needs to be detected or screened for low-quality images. Optionally, the original image to be detected is a file of all image formats in the original image set, or can also be an image for which the user customizes the image quality detection in the album. Similarly, the original video to be detected refers to the video in the original video set that needs to be detected or screened for low-quality videos. Optionally, the original video to be detected is a file of all video formats in the original video set, or can also be a video for which the user independently sets the video quality detection.

[0156] In addition, the first intermediate image set and the first intermediate video set refer to the sets respectively composed of the images and videos initially detected as low-quality in the original image set and the original video set of the album. The first intermediate image set and the first intermediate video set include images and videos with obvious low quality. The first intermediate image set is a set formed by screening low-quality images from the original image set using the basic image attributes.

[0157] It can be understood that the most common low-quality photos in reality include out-of-focus pictures, relative motion blur pictures formed by object movement and camera shake, black pictures or low-illumination dark pictures caused by imperfect imaging systems such as lenses and camera photosensitive modules with poor performance, and distorted images caused by accidental factors, etc. The image attribute parameters of the images to be detected in the original image set can be analyzed according to the preset image parameters. If at least one image attribute parameter is within the abnormal range, it can be regarded as a low-quality image. Similarly, in reality, the most common low-quality videos on the edge device include out-of-focus videos, black screen videos, blurred videos, low-light videos, blocked videos, and so on. For the known low-quality videos, the abnormal parameters are determined, and then the video attribute parameters of the videos to be detected in the original video set are analyzed. If at least one video attribute parameter is within the abnormal range, it is regarded as a low-quality video.

[0158] Through the above analysis of the attributes of images / videos, low-quality images / videos with obvious abnormalities can be detected.

[0159] Low-quality images / videos can also be screened through deep learning. By using the images / videos and their corresponding quality labels as the training set, convolutional neural networks and other methods are used to train low-quality image classification models and low-quality video classification models for classifying low-quality videos / images. According to the low-quality image classification model, the quality of the original image to be detected is classified, and the set of low-quality original images to be detected in the original image set is used as the second intermediate image set. For the original videos to be detected in the original video set, low-quality videos are detected by frame extraction of the videos to be detected, forming the first intermediate video set, and according to the low-quality video classification model, the quality of the original videos to be detected is classified, and the set of low-quality original videos to be detected in the original video set is used as the second intermediate video set.

[0160] It can be understood that the two methods of screening low-quality images / videos in the secondary preliminary screening are not executed in a fixed order. It is possible to first execute the low-quality image / video classification model to achieve the initial screening of low-quality images / videos in the album, and then execute the secondary screening of low-quality images / videos according to the attributes of the images / videos.

[0161] The two screening methods can also be executed in parallel, and the low-quality images / videos obtained from the two screenings are integrated.

[0162] It is also possible to only execute any one of the two screening methods, and use the obtained screening results as the detection of significantly abnormal low-quality images / videos.

[0163] Preferably, the two screening methods are executed in parallel;

[0164] For images, in the original image set, the low-quality images detected by image attributes form the first intermediate image set, and the low-quality images detected by the low-quality image classification model form the second intermediate image set. The union of the first intermediate image set and the second intermediate image set is used as the first target image set, that is, the first target image set contains significantly low-quality images.

[0165] For videos, in the original video set, the low-quality videos detected by video attributes form the first intermediate video set, and the low-quality videos detected by the low-quality video classification model form the second intermediate video set. The union of the first intermediate video set and the second intermediate video set is used as the first target video set, that is, the first target video set contains significantly low-quality videos.

[0166] In an implementation manner of this embodiment,

[0167] For significantly low-quality images, secondary preliminary screening processing is adopted, combining traditional image processing with classification model analysis and grading filtering processing to ensure the integrity and accuracy of the processing of low-quality images.

[0168] In the first-level processing, basic detection is performed on images and videos separately. The basic detection of images includes: brightness / contrast detection, which is used to filter low-quality images and videos such as black, dark, and occluded images that are too dark or overexposed; blur detection, which is used to filter out-of-focus and blurred images; image color detection; and detection of distorted images.

[0169] Specifically, obtain basic physical statistical attributes of the image, such as parameters like objective brightness, color, resolution, entropy, signal-to-noise ratio (SNR), and peak signal-to-noise ratio (PSNR). Use the Laplacian variance method for image blur detection and similarity detection of blurred and out-of-focus images. At the same time, use the image histogram brightness threshold judgment method to delete overexposed and underexposed images. Also, perform color cast detection on the image. When there is a color cast, the mean values of a and b in the LAB channels of the image will deviate far from the origin point, and the variance will be small.

[0170] For videos, use the method of large-span frame extraction to extract frames from the video for detection and make a quick judgment. Here, only videos with obvious abnormal features such as out-of-focus videos, black screen videos, abnormal resolution, and stretching and deformation are filtered.

[0171] In the second-level processing, a low-quality image classification model and a low-quality video classification model are respectively constructed for pictures and videos, and the constructed models are used for hierarchical filtering.

[0172] For the photo processing part, a low-quality image classification model is constructed based on the Residual Neural Network (ResNet) architecture. Use low-quality images to train a fast image classification model with a simple structure and fast inference speed. It can quickly screen and mark relatively complex low-blur, out-of-focus, and cluttered images in the image.

[0173] Among them, the low-quality image classification model based on ResNet is implemented based on machine learning methods. It mainly extracts feature parameters that can reflect the image quality from images of known quality, conducts training and learning to establish an analysis model, and then extracts the features of the image to be evaluated through the network and inputs them into the analysis model to predict the quality category of the image. The evaluation results of this method are generally better than those predicted by function fitting. It has good discrimination ability for types such as image blur, image noise, JPEG compression, JPEG2000 compression, and blocking effect, and also has a certain generalization ability. The training data of this classification model uses the image quality evaluation database of Categorical Subjective Image Quality (CSIQ) developed by Oklahoma State University in the United States and its own low-quality image classification database, which are manually labeled.

[0174] For the video processing part, a low-quality video classification model based on the architecture of Inflated 3D ConvNets (I3D) is retrained using the typical low-quality video dataset collected in this case. After 4-bit quantization processing, it can run smoothly on end devices such as mobile phones. It has good classification ability for obvious defocus, accidental touch videos, pure white, pure black videos, and large-area occlusion, and can effectively and quickly mark the videos in these scenarios.

[0175] Among them, the low-quality video classification model based on the I3D network architecture mainly uses the video understanding ability of I3D to understand and extract features from video clips, and then conducts low-quality category classification. Since the common frame rate of videos on mobile phones is relatively high, in this case, the video size is compressed by evenly extracting frames and then sent to the network for classification. The I3D network in this case expands the 2D model to a 3D model, thus forming a network specifically for video understanding. At the same time, the network already designed in 2D can be reused. For example, VGG and ResNet can be directly expanded to 3D, and even the pre-trained model can be utilized in some ingenious ways. This not only simplifies the design but also saves a lot of trouble in pre-training. Its structure is similar to that of a 3D convolutional network, but it combines optical flow and extracts spatio-temporal features through 3D convolutional operations. Finally, weighted averaging is performed and sent to the classification network to obtain the probability scores of each category. It can accurately classify some obvious low-quality videos.

[0176] Figure 4 It is a schematic diagram of the inference process of a personalized album cleaning multi-modal large model for an embodiment of the present disclosure. As Figure 4As shown, the image / video to be detected extracts features through the feature extraction network of the multimodal large model and performs image / video encoding to obtain an image / video feature vector, and the image / video feature vector is represented by image / video embedding. Based on the image / video embedding representation, the image and video Q-Former model is used to determine the query vector of the text modality from the learned query vectors. The query vector is mapped through a fully connected layer to form a target query vector, and the target query vector and the prefix text including user personalized information are used as the input of the base large language model Vicuna v1.1 13b to rate the quality of the image / video. Through this multimodal large model, visual content can be effectively translated into text that the LLM can understand, and the result text Text that conforms to the prefix text is selected from the understood text. Optionally, in this embodiment, either 7b or 13b of Vicuna v1.1 can be selected as the base large language model. Each Q-Former of the model optimizes 188M trainable parameters and learns K = 32 query tokens with a hidden dimension size of 768. Through the inference mechanism of this multimodal large model, the images / videos in the photo album that are difficult to distinguish in terms of quality are combined with the user's personalized information, and low-quality images / videos that do not meet the user's aesthetic standards are screened out.

[0177] The present disclosure proposes an album cleaning method that performs a secondary preliminary screening on the original image set and the original video set in the album to obtain a first target image set formed by low-quality pictures and a first target video set formed by low-quality videos, screening out significantly low-quality pictures and videos, thereby avoiding subjecting all the content of the album to precise recognition by the subsequent multimodal large model, which helps improve the screening efficiency;

[0178] Based on the images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected, and the user personalized information, the multimodal large model fuses the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected, and outputs the quality evaluations of the images to be detected and the videos to be detected. According to the quality evaluations, the second target image set and the second target video set of low quality in the first image set to be detected and the first video set to be detected are respectively determined, where the image feature vectors and the video feature vectors have the same length. The first image set to be detected is the difference set between the original image set and the first target image set, and the first video set to be detected is the difference set between the original video set and the first target video set. The images / videos in the photo album that are difficult to distinguish in terms of quality are combined with the user's personalized information, and low-quality images / videos that do not meet the user's aesthetic standards are screened out.

[0179] Clearing at least one of the first target image set, the first target video set, the second target image set, and the second target video set in the album completes the intelligent cleaning of low-quality images / videos in the album, realizes the personalized cleaning of the album, and improves the user experience.

[0180] Corresponding to the methods provided in the above several embodiments, the present disclosure also provides an album cleaning device. Since the device provided in the embodiments of the present disclosure corresponds to the methods provided in the above several embodiments, the implementation manners of the methods are also applicable to the device provided in this embodiment and will not be described in detail in this embodiment.

[0181] Figure 5 It is a schematic structural diagram of an album cleaning device 500 according to an embodiment of the present disclosure. As Figure 5 shown, the album cleaning device includes:

[0182] A primary screening module 510, configured to perform a secondary primary screening on the original image set and the original video set in the album to obtain a first target image set formed by low-quality pictures and a first target video set formed by low-quality videos;

[0183] A quality evaluation module 520, configured to, based on the images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected, and the personalized information, fuse the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected through a multimodal large model, and output the quality evaluations of the images to be detected and the videos to be detected. According to the quality evaluations, respectively determine a second target image set and a second target video set with low quality in the first image set to be detected and the first video set to be detected, where, based on the images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected, and the user's personalized information, fuse the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected through a multimodal large model, and output the quality evaluations of the images to be detected and the videos to be detected. According to the quality evaluations, respectively determine a second target image set and a second target video set with low quality in the first image set to be detected and the first video set to be detected, where the image feature vectors and the video feature vectors have the same length, the first image set to be detected is the difference set between the original image set and the first target image set, and the first video set to be detected is the difference set between the original video set and the first target video set;

[0184] A clearing module 530, configured to clear at least one of the first target image set, the first target video set, the second target image set, and the second target video set in the album.

[0185] In some embodiments, the quality evaluation module 520 is configured to:

[0186] Based on the images to be detected in the first image set to be detected and the videos to be detected in the first video set to be detected, the multi-modal large model is used to extract features from the images to be detected and encode them into image feature vectors in the image modality, and extract features from the videos to be detected and encode them into video feature vectors in the video modality;

[0187] Based on the image feature vectors and video feature vectors, determine the target query vector that aligns the visual modality to the text modality, where the visual modality includes the image modality and the video modality, and the image feature vectors and video feature vectors have the same length;

[0188] According to the user's personalized information, the base large language model of the multi-modal large model is used to parse the target query vector, and the quality evaluation of the images to be detected and the videos to be detected is output. The pictures with low quality evaluation in the first image set to be detected are used as the second target image set, and the videos with low quality evaluation in the first video set to be detected are used as the second target video set.

[0189] In some embodiments, before the quality evaluation module 520 extracts features from the images to be detected through the multi-modal large model and encodes them into image feature vectors in the image modality, and extracts features from the videos to be detected and encodes them into video feature vectors in the video modality, it is also used for:

[0190] Obtain video samples, image samples, and fine-tuning text, where the fine-tuning text is generated based on the preset text descriptions of the video samples and image samples and the user portrait of the photo album;

[0191] Perform data difference on the video samples to obtain the first sample video;

[0192] Divide the first sample video and the image samples using sliding windows of the same size, and respectively extract the first video features of the first sample video in the sliding window and the first image features of the image samples in the sliding window through the feature extraction network;

[0193] Map the first video features and the first image features into the first video feature vectors and the first image feature vectors respectively, and the first video feature vectors and the first image feature vectors have the same length;

[0194] Based on the first video feature vectors, the first image feature vectors, and the fine-tuning text, train the pre-trained multi-modal large model by means of instruction fine-tuning to obtain the multi-modal large model.

[0195] In some embodiments, before the quality evaluation module 520 parses the target query vector through the base large language model of the multi-modal large model according to the user's personalized information, it is also used for:

[0196] Generate a set of learnable query vectors with a fixed length based on image feature vectors and video feature vectors;

[0197] Generate a query embedding representation according to the learnable query vectors through the self-attention mechanism and the cross-attention mechanism;

[0198] Linearly project the query embedding representation through a fully connected layer to obtain a target query vector with the same dimension as the text embedding of the base large language model;

[0199] Use the target query vector, user personalized information, and the identifier of the visual modality as the input to the base large language model.

[0200] In some embodiments, the quality evaluation module 520 adopts the following method to parse the target query vector through the base large language model of the multimodal large model according to the user personalized information, and outputs the quality evaluation of the image to be detected and the video to be detected:

[0201] Create a user portrait based on the user's historical operation data of the photo album;

[0202] Generate user personalized information according to the user portrait and the personalized requirement text;

[0203] Use the user personalized information as the prompt of the base large language model, and parse the target query vector through the base large language model to generate the quality evaluation of the image to be detected and the video to be detected.

[0204] In some embodiments, the clearing module 530 is further configured to:

[0205] In response to the user selecting the target image and target video to be deleted from the first target image set, the first target video set, the second target image set, and the second target video set, delete the target image and target video from the photo album.

[0206] In some embodiments, the preliminary screening module 510 is used for:

[0207] Based on the original images to be detected in the original image set, use the basic image attributes to screen out low-quality pictures to form a first intermediate image set, and classify the quality of the original images to be detected according to the low-quality image classification model, and use the set composed of the low-quality original images to be detected in the original image set as the second intermediate image set;

[0208] Based on the original videos to be detected in the original video set, detect low-quality videos by extracting frames from the videos to be detected to form a first intermediate video set, and classify the quality of the original videos to be detected according to the low-quality video classification model, and use the set composed of the low-quality original videos to be detected in the original video set as the second intermediate video set;

[0209] Take the union of the first intermediate image set and the second intermediate image set as the first target image set, and take the union of the first intermediate video set and the second intermediate video set as the first target video set.

[0210] In summary, through the album cleaning device, the original image set and the original video set in the album are subjected to a two-stage preliminary screening to obtain the first target image set formed by low-quality pictures and the first target video set formed by low-quality videos; based on the images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected, and the user's personalized information, the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected are fused through a multi-modal large model, and the quality evaluations of the images to be detected and the videos to be detected are output. According to the quality evaluations, the second target image set and the second target video set with low quality in the first image set to be detected and the first video set to be detected are respectively determined. Among them, the image feature vectors and the video feature vectors have the same length. The first image set to be detected is the difference set between the original image set and the first target image set, and the first video set to be detected is the difference set between the original video set and the first target video set; at least one of the first target image set, the first target video set, the second target image set, and the second target video set is cleared in the album. This device solves the problem that low-quality images and videos in the album cannot be accurately cleared based on the user's personalized aesthetics, realizes the personalized cleaning of the album, saves the storage space of the album, and improves the user experience.

[0211] In the above embodiments provided by the present disclosure, the methods and devices provided by the embodiments of the present disclosure are introduced. In order to implement the various functions in the methods provided by the above embodiments of the present disclosure, an electronic device may include a hardware structure and software modules, and implement the above various functions in the form of a hardware structure, a software module, or a combination of a hardware structure and a software module. A certain function among the above various functions may be executed in the form of a hardware structure, a software module, or a combination of a hardware structure and a software module.

[0212] Figure 6 It is a block diagram of an electronic device 600 for implementing the above album cleaning method shown according to an exemplary embodiment.

[0213] For example, the electronic device 600 may be a mobile phone, a computer, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0214] Refer to Figure 6 , the electronic device 600 may include one or more of the following components: a processing component 602, a memory 604, a power component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.

[0215] The processing component 602 generally controls the overall operation of the electronic device 600, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the above - mentioned methods. In addition, the processing component 602 may include one or more modules to facilitate the interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate the interaction between the multimedia component 608 and the processing component 602.

[0216] The memory 604 is configured to store various types of data to support the operation of the electronic device 600. Examples of such data include instructions for any application or method operating on the electronic device 600, contact data, phone book data, messages, pictures, videos, etc. The memory 604 can be implemented by any type of volatile or non - volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read - only memory (EEPROM), erasable programmable read - only memory (EPROM), programmable read - only memory (PROM), read - only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.

[0217] The power component 606 provides power to various components of the electronic device 600. The power component 606 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 600.

[0218] The multimedia component 608 includes a screen that provides an output interface between the electronic device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 608 includes a front - facing camera and / or a rear - facing camera. When the electronic device 600 is in an operating mode, such as a shooting mode or a video mode, the front - facing camera and / or the rear - facing camera can receive external multimedia data. Each front - facing camera and rear - facing camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0219] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 600 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 further includes a speaker for outputting audio signals.

[0220] The I / O interface 612 provides an interface between the processing component 602 and a peripheral interface module, which may be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0221] The sensor component 614 includes one or more sensors for providing status assessments of various aspects of the electronic device 600. For example, the sensor component 614 can detect the on / off state of the electronic device 600, the relative positioning of components, such as the display and keypad of the electronic device 600. The sensor component 614 can also detect a change in the position of the electronic device 600 or a component of the electronic device 600, the presence or absence of user contact with the electronic device 600, the orientation or acceleration / deceleration of the electronic device 600, and a change in the temperature of the electronic device 600. The sensor component 614 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 614 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 614 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0222] The communication component 616 is configured to facilitate communication between the electronic device 600 and other devices in a wired or wireless manner. The electronic device 600 can access a wireless network based on communication standards, such as WiFi, 2G or 3G, 4G LTE, 5G NR (New Radio), or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0223] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0224] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, and the above instructions can be executed by a processor 620 of the electronic device 600 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0225] An embodiment of the present disclosure also proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the album cleaning method described in the above embodiments of the present disclosure.

[0226] An embodiment of the present disclosure also proposes a computer program product, including a computer program, and the computer program executes the album cleaning method described in the above embodiments of the present disclosure when being executed by a processor.

[0227] Figure 7 FIG. is a schematic structural diagram of a chip 700 for implementing the above album cleaning method according to an exemplary embodiment.

[0228] Referring to Figure 7 , the chip 700 includes at least one communication interface 701 and a processor 702; the communication interface 701 is used to receive signals input to the chip 700 or signals output from the chip 700, and the processor 702 communicates with the communication interface 701 and implements the album cleaning method described in the above embodiments through logic circuits or by executing code instructions.

[0229] It should be noted that the terms "first", "second", etc. in the specification, claims, and the above drawings of the present disclosure are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0230] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0231] Any process or method description shown in the flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions may be executed in a manner not shown or discussed, including substantially concurrently according to the functions involved or in a reverse order, which should be understood by those skilled in the technical field to which the embodiments of the present disclosure pertain.

[0232] The logic and / or steps represented in the flowchart or described otherwise herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processing module, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or more wirings (control method), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0233] It should be understood that each part of the embodiments of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0234] Those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0235] In addition, in each of the embodiments of the present disclosure, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.

[0236] Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations of the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.

Claims

1. A photo album cleaning method, characterized in that: The method comprises: Performing a secondary preliminary screening on the original image set and the original video set in the album to obtain a first target image set formed by low-quality images and a first target video set formed by low-quality videos; Based on the images to be detected of the first image set to be detected, the videos to be detected of the first video set to be detected and user personalized information, the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected are fused through a multimodal large model, and quality evaluations of the images to be detected and the videos to be detected are output, and according to the quality evaluations, low-quality second target image sets and second target video sets in the first image set to be detected and the first video set to be detected are determined respectively, wherein the image feature vectors and the video feature vectors have the same length, the first image set to be detected is a difference set between the original image set and the first target image set, and the first video set to be detected is a difference set between the original video set and the first target video set; At least one of the first target image set, the first target video set, the second target image set and the second target video set is cleared from the album.

2. The method according to claim 1, characterized in that The method comprises: based on the images to be detected in the first image set to be detected, the videos to be detected in the first video set to be detected and the user personalized information, fusing the image feature vectors of the images to be detected and the video feature vectors of the videos to be detected through a multimodal large model, and outputting quality evaluations of the images to be detected and the videos to be detected, and determining low-quality second target image sets and second target video sets in the first image set to be detected and the first video set to be detected respectively according to the quality evaluations, including: Based on the images to be detected in the first set of images to be detected and the videos to be detected in the first set of videos to be detected, extracting features from the images to be detected and encoding them into the image feature vectors of the image modality through the multimodal large model, and extracting features from the videos to be detected and encoding them into the video feature vectors of the video modality; Determining a target query vector of a visual modality aligned to a text modality based on the image feature vector and the video feature vector, wherein the visual modality includes the image modality and the video modality, and the image feature vector and the video feature vector have the same length; According to the user personalized information, the target query vector is parsed by the base large language model of the multimodal large model, and the quality evaluation of the image to be detected and the video to be detected is output, and the pictures with low quality evaluation in the first image set to be detected are taken as the second target image set, and the videos with low quality evaluation in the first video set to be detected are taken as the second target video set.

3. The method according to claim 2, characterized in that Before extracting features from the image to be detected by the multimodal large model and encoding the image feature vector in image mode, and extracting features from the video to be detected and encoding the video feature vector in video mode, the method further includes: Acquire a video sample, an image sample, and a fine-tuning text, wherein the fine-tuning text is generated based on a preset text description of the video sample and the image sample and a user portrait of the album; Performing data differentiation on the video sample to obtain a first sample video; Dividing the first sample video and the image sample using sliding windows of the same size, and respectively extracting a first video feature of the first sample video in the sliding window and a first image feature of the image sample in the sliding window through a feature extraction network; Mapping the first video feature and the first image feature into a first video feature vector and a first image feature vector, respectively, wherein the first video feature vector and the first image feature vector have the same length; Based on the first video feature vector, the first image feature vector and the fine-tuning text, the pre-trained multimodal large model is trained by instruction fine-tuning to obtain the multimodal large model.

4. The method according to claim 3, characterized in that: Before parsing the target query vector by the base large language model of the multimodal large model according to the user personalized information, the method further includes: Based on the image feature vector and the video feature vector, a set of fixed-length learnable query vectors are generated; According to the learnable query vector, a query embedding representation is generated through a self-attention mechanism and a cross-attention mechanism; Linearly projecting the query embedding representation through a fully connected layer to obtain the target query vector of the same dimension as the text embedding of the base large language model; The target query vector, the user personalized information and the identifier of the visual modality are used as inputs of the base large language model.

5. The method according to claim 1, characterized in that The method of parsing the target query vector by using the base large language model of the multimodal large model according to the user personalized information and outputting the quality evaluation of the image to be detected and the video to be detected includes: Creating a user profile based on the user's historical operation data on the photo album; Generate the user personalized information according to the user portrait and the personalized requirement text; The user personalized information is used as a prompt of the base large language model, and the target query vector is parsed by the base large language model to generate a quality evaluation of the image to be detected and the video to be detected.

6. The method according to claim 1, characterized in that Clearing at least one of the first target image set, the first target video set, the second target image set, and the second target video set in the album further includes: In response to the user selecting a target image and a target video to be deleted from the first target image set, the first target video set, the second target image set, and the second target video set, the target image and the target video are deleted from the album.

7. The method according to claim 1, characterized in that The secondary preliminary screening of the original image set and the original video set in the album to obtain a first target image set formed by low-quality pictures and a first target video set formed by low-quality videos includes: Based on the original images to be detected in the original image set, low-quality images are screened by using basic image attributes to form a first intermediate image set, and the quality of the original images to be detected is classified according to a low-quality image classification model, and a set consisting of the low-quality original images to be detected in the original image set is used as a second intermediate image set; Based on the original videos to be detected in the original video set, a first intermediate video set is formed by extracting frames of the videos to be detected to detect low-quality videos, and the quality of the original videos to be detected is classified according to a low-quality video classification model, and a set consisting of the low-quality original videos to be detected in the original video set is used as a second intermediate video set; A union of the first intermediate image set and the second intermediate image set is used as the first target image set, and a union of the first intermediate video set and the second intermediate video set is used as the first target video set.

8. A photo album cleaning device, characterized in that: The device comprises: A primary screening module, used for performing secondary primary screening on the original image set and the original video set in the album to obtain a first target image set formed by low-quality images and a first target video set formed by low-quality videos; A quality evaluation module, for fusing the image feature vector of the image to be detected and the video feature vector of the video to be detected through a multimodal large model based on the image to be detected of the first image set to be detected, the video to be detected of the first video set to be detected and personalized information, and outputting the quality evaluation of the image to be detected and the video to be detected, and respectively determining the second target image set and the second target video set of low quality in the first image set to be detected and the first video set to be detected according to the quality evaluation, wherein, based on the image to be detected of the first image set to be detected, the video to be detected of the first video set to be detected and user personalized information, the image feature vector of the image to be detected and the video feature vector of the video to be detected are fused through a multimodal large model, and the quality evaluation of the image to be detected and the video to be detected is output, and respectively determining the second target image set and the second target video set of low quality in the first image set to be detected and the first video set to be detected according to the quality evaluation, wherein the image feature vector and the video feature vector have the same length, the first image set to be detected is the difference set of the original image set and the first target image set, and the first video set to be detected is the difference set of the original video set and the first target video set; A clearing module is used to clear at least one of the first target image set, the first target video set, the second target image set and the second target video set in the album.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

11. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 7.