Image quality evaluation method and device
A multi-modal image quality assessment model addresses the limitations of traditional methods by using a multimodal large language model and multilayer perceptron to provide precise qualitative and quantitative image quality evaluations, improving accuracy and efficiency.
Patent Information
- Application Number
- CN202510376851.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-15
AI Technical Summary
Existing image quality evaluation methods, especially deep learning-based methods, are difficult to effectively process complex distortion types and multimodal information, resulting in low evaluation accuracy.
The target multimodal image quality evaluation model is adopted, including multimodal large language sub-modal, visual sub-modal and multi-layer perceptual machine sub-model, and qualitative and quantitative image quality evaluation results are output through feature extraction and multi-task learning.
Accurate description and quantitative scoring of image quality are achieved, reducing evaluation costs and time, and providing a fast and automated image quality evaluation solution.
Smart Images

Figure CN120318165A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to a method and apparatus for image quality evaluation, a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Image quality assessment (IQA) is a fundamental task in the fields of image processing and computer vision, aiming to evaluate the quality of an image by analyzing its characteristics.
[0003] Traditional IQA methods mainly rely on mathematical models and algorithms, such as mean squared error (MSE) and peak signal-to-noise ratio (PSNR). These methods evaluate image quality by calculating the differences between image pixels. However, these methods often cannot fully reflect the perception ability of the human visual system, especially when dealing with complex image distortions and subjective visual experiences.
[0004] With the development of deep learning and artificial intelligence technologies, deep learning-based IQA methods have gradually become a research hotspot. These methods train neural networks to learn the feature representations of image quality, thereby improving the accuracy of evaluation.
[0005] However, the inventors found that although deep learning methods such as convolutional neural networks (CNNs) have made significant progress in IQA tasks, there are still problems in dealing with complex distortion types and multi-modal information, resulting in low evaluation accuracy.
[0006] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention
[0007] Embodiments of the present application provide a method and apparatus for image quality evaluation, a computer device, a computer-readable storage medium, and a computer program product to solve or alleviate one or more of the above technical problems.
[0008] One aspect of embodiments of the present application provides a method for image quality evaluation, the method including: Obtaining an image to be evaluated for image quality; Inputting the image to be evaluated for image quality and a pre-set target prompt word template into a pre-trained target multi-modal image quality evaluation model, where the target multi-modal image quality evaluation model includes a multi-modal large language sub-model, a visual sub-model, and a multi-layer perceptron sub-model; Extracting feature data of the image by the visual sub-model from the image to be evaluated for image quality; The multi-modal large language sub-model outputs the qualitative image quality evaluation result corresponding to the image to be evaluated for quality and the word vector associated with the quantitative image quality evaluation result based on the image feature data and the target prompt template, and the multi-layer perceptron sub-model outputs the quantitative image quality evaluation result corresponding to the image to be evaluated for quality based on the word vector.
[0009] Optionally, obtaining the image to be evaluated for quality includes: Extracting multiple video frames from the video to be evaluated for quality as the image to be evaluated for quality.
[0010] Optionally, the visual sub-model extracts features from the image to be evaluated for quality, and the obtained image feature data includes: The visual sub-model extracts features from multiple video frames respectively to obtain the image feature data corresponding to each video frame; The multi-modal large language sub-model outputs the qualitative image quality evaluation result corresponding to the image to be evaluated for quality and the word vector associated with the quantitative image quality evaluation result based on the image feature data and the target prompt template, and the multi-layer perceptron sub-model outputs the quantitative image quality evaluation result corresponding to the image to be evaluated for quality based on the word vector, including: The multi-modal large language sub-model outputs the qualitative frame image quality evaluation result corresponding to each video frame and the word vector associated with the qualitative frame image quality evaluation result based on the image feature data corresponding to each video frame and the prompt template; The multi-layer perceptron sub-model outputs the quantitative frame image quality evaluation result corresponding to each video frame based on the word vector corresponding to each video frame; Determining the qualitative image quality evaluation result corresponding to the image to be evaluated for quality based on the qualitative frame image quality evaluation results corresponding to each video frame, and determining the quantitative image quality evaluation result corresponding to the image to be evaluated for quality based on the quantitative frame image quality evaluation results corresponding to each video frame.
[0011] Optionally, determining the qualitative image quality evaluation result corresponding to the image to be evaluated for quality based on the qualitative frame image quality evaluation results corresponding to each video frame includes: Performing statistical analysis on the qualitative frame image quality evaluation results corresponding to each video frame; Taking the qualitative frame image quality evaluation result with the largest quantity as the qualitative image quality evaluation result corresponding to the image to be evaluated for quality; Determining the qualitative image quality evaluation result corresponding to the image to be evaluated for quality based on the qualitative frame image quality evaluation results corresponding to each video frame includes: Calculating the average value of the quantitative frame image quality evaluation results corresponding to each video frame, and taking the average value as the quantitative image quality evaluation result corresponding to the image to be evaluated for quality.
[0012] Optionally, the target multi-modal image quality assessment model is trained through the following operations: Obtain a qualitative training sample data set, where the qualitative training sample data set includes multiple qualitative training sample images, and each qualitative training sample image carries a qualitative attribute label; Input the qualitative training sample data set and a pre-set first prompt word template into an initial model for training to obtain an initial multi-modal image quality assessment model; Obtain a quantitative training sample data set, where the quantitative training sample data set includes multiple quantitative training sample images, and each quantitative training sample image carries multi-dimensional quantitative attribute labels, and the multi-dimensional quantitative attribute labels include at least two of a color label, a noise label, an artifact label, a blur label, and a temporal consistency label; Input the qualitative training sample data set and a pre-set second prompt word template into the initial multi-modal image quality assessment model for training to obtain the target multi-modal image quality assessment model.
[0013] Optionally, in the qualitative training sample data set, multiple pairs of qualitative sample images are used as multiple qualitative training sample images, and the difference between the two sample images in the qualitative sample image pair in the target dimension exceeds a preset value; In the quantitative training sample data set, multiple pairs of quantitative sample images are used as multiple quantitative training sample images, and the difference between the two sample images in the quantitative sample image pair in the target dimension exceeds a preset value.
[0014] Another aspect of the embodiments of the present application provides an image quality assessment device, and the device includes: An acquisition module, configured to acquire an image to be subjected to image quality assessment; An input module, configured to input the image to be subjected to image quality assessment and a pre-set target prompt word template into a pre-trained target multi-modal image quality assessment model, where the target multi-modal image quality assessment model includes a multi-modal large language sub-model, a visual sub-model, and a multi-layer perceptron sub-model; An extraction module, configured to perform feature extraction on the image to be subjected to image quality assessment through the visual sub-model to obtain image feature data; An output module, configured to output a qualitative image quality assessment result corresponding to the image to be subjected to image quality assessment and a word vector associated with the quantitative image quality assessment result based on the image feature data and the target prompt word template through the multi-modal large language sub-model, and output a quantitative image quality assessment result corresponding to the image to be subjected to image quality assessment based on the word vector through the multi-layer perceptron sub-model.
[0015] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0016] Another aspect of the embodiments of the present application provides a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions are executed by a processor, the method as described above is implemented.
[0017] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.
[0018] The embodiments of the present application adopting the above technical solutions may include the following advantages: by inputting the image to be evaluated for image quality into a pre-trained target multi-modal image quality evaluation model, the powerful image understanding ability and description ability of the target multi-modal image quality evaluation model, as well as the powerful multi-task learning ability can be utilized to accurately output the qualitative image quality evaluation result and quantitative image quality evaluation result of the image to be evaluated for image quality, realizing the qualitative description and quantitative scoring of the image quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0020] Figure 1 Schematically shows the operating environment diagram of the image quality evaluation method according to Embodiment 1 of the present application; Figure 2 Schematically shows the flowchart of the image quality evaluation method according to Embodiment 1 of the present application; Figure 3 Schematically shows Figure 2 The sub-step flowchart of step S206 in; Figure 4 Schematically shows the refined training step diagram of the target multi-modal image quality evaluation model; Figure 5 Schematically shows the block diagram of the image quality evaluation device according to Embodiment 2 of the present application; and Figure 6 Schematically shows the hardware architecture diagram of the computer device according to Embodiment 3 of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.
[0022] It should be noted that in the embodiments of the present application, the descriptions involving "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0023] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and to distinguish each step, and thus cannot be understood as a limitation to the present application.
[0024] First, the following provides the term explanations involved in the present application: Multimodal Large Language Model (MLLM): It is an emerging artificial intelligence technology that combines a large language model (LLM) with multimodal information processing and can process and understand data in multiple modalities including text, images, videos, audio, etc.
[0025] Image Quality Assessment (IQA): It is one of the basic technologies in image processing. By mainly analyzing the characteristics of images, it evaluates the quality of images, quantifies the image quality, and reflects the characteristics of images in terms of sharpness, contrast, distortion degree, etc.
[0026] Mean Opinion Score (MOS): It is a commonly used subjective evaluation method in image quality assessment. It evaluates the image quality by collecting the scores of multiple observers and calculating the average value.
[0027] No Reference (NR). NR-IQA refers to the situation where when performing image quality assessment, there is no original image as a reference, and the evaluation algorithm scores only based on the information of the distorted image itself.
[0028] Full Reference (FR). FR-IQA refers to the situation where both the reference image and the distorted image are available during image quality assessment. The evaluation algorithm scores by comparing the differences between the two.
[0029] Vision Transformer (ViT) model: A model that applies the Transformer architecture to computer vision tasks, first published at ICLR 2021. The core idea of the ViT model is to convert image data into sequence data and use the Transformer architecture to process these sequences, thereby capturing the global dependencies in the image.
[0030] Multilayer Perceptron (MLP) model, a classic feedforward neural network model, widely used in tasks such as classification, regression, and feature learning. It is one of the basic models in deep learning, consisting of multiple layers of neurons. Each neuron introduces non-linearity through a non-linear activation function, enabling the model to learn complex input-output relationships.
[0031] Secondly, to facilitate the understanding of the technical solutions provided in the embodiments of the present application by those skilled in the art, the related technologies are described below: Traditional IQA methods mainly rely on mathematical models and algorithms, such as Mean Squared Error (MSE) and Peak Signal-to-Noise Ratio (PSNR). These methods evaluate image quality by calculating the differences between image pixels. However, these methods often cannot fully reflect the perceptual ability of the human visual system, especially when dealing with complex image distortions and subjective visual experiences.
[0032] Traditional objective IQA methods also include Structural Similarity (SSIM), Gradient Magnitude Similarity Deviation (GMSD), etc. These methods have improved the simulation of human visual perception to a certain extent, but there are still limitations. Subjective evaluation relies on the scores of human observers and is generally considered the most accurate evaluation method, but it is time-consuming and costly.
[0033] With the development of deep learning and artificial intelligence technologies, deep learning-based IQA methods have gradually become a research hotspot. These methods train neural networks to learn the feature representations of image quality, thereby improving the accuracy of evaluation.
[0034] However, although deep learning methods such as Convolutional Neural Networks (CNNs) have made significant progress in IQA tasks, there are still problems in dealing with complex distortion types and multi-modal information, resulting in low evaluation accuracy.
[0035] To this end, the embodiments of the present application provide a technical solution for image quality evaluation. In this technical solution: (1) Utilize the powerful image understanding ability and description ability of the target multi-modal image quality evaluation model, as well as the powerful multi-task learning ability to accurately output the qualitative image quality evaluation result and quantitative image quality evaluation result of the image to be evaluated for quality, so as to realize the qualitative description and quantitative scoring of the image quality; (2) Provide a fast and automated image quality evaluation solution, reducing the evaluation cost and time. See the following for details.
[0036] Finally, for the convenience of understanding, an exemplary operating environment is provided below.
[0037] As Figure 1 shown, the operating environment diagram includes: a service platform 2, a network 4, and a client 6, where: The service platform 2 can be composed of a single or multiple computing devices. The multiple computing devices can include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing devices can load virtual machines based on virtual images and / or other data that define specific software (e.g., operating systems, dedicated application programs, servers) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices. A hypervisor can be implemented to manage the use of different virtual machines on the same computing device.
[0038] The service platform 2 can be configured to communicate with the client 6, etc. through the network 4. The network 4 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, and combinations thereof, or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.
[0039] The service platform 2 can provide services such as storage, reading, writing, querying, deleting, etc., such as running the target multi-modal image quality evaluation model or providing image quality evaluation services for the client.
[0040] The client 6 can be an electronic device running an operating system such as Windows, AndroidTM, or iOS, such as a smart phone, a tablet device, a laptop computer, a virtual reality device, a gaming device, a set-top box, a vehicle-mounted terminal, a smart TV. Based on the above operating systems, various application programs can be run, such as application programs for image quality evaluation.
[0041] The client 6 can provide / configure a user access page for controlling the service platform 2 or uploading objects (such as videos), etc.
[0042] Note that the above devices are exemplary, and in different scenarios or according to different requirements, the number and types of devices are adjustable.
[0043] Taking the service platform 2 or the client 6 as the execution entity, the technical solution of the present application will be introduced through multiple embodiments. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.
[0044] Embodiment 1 Figure 2 A flowchart of the image quality evaluation method according to Embodiment 1 of the present application is schematically shown.
[0045] As Figure 2 shown, the image quality evaluation method may include steps S200 to S206, where: Step S200, obtaining an image to be evaluated for image quality.
[0046] Step S202, inputting the image to be evaluated for image quality and a pre-set target prompt word template into a pre-trained target multi-modal image quality evaluation model, where the target multi-modal image quality evaluation model includes a multi-modal large language sub-model, a visual sub-model, and a multi-layer perceptron sub-model.
[0047] Step S204, extracting features from the image to be evaluated for image quality through the visual sub-model to obtain image feature data.
[0048] Step S206, outputting a qualitative image quality evaluation result corresponding to the image to be evaluated for image quality and a word vector associated with the quantitative image quality evaluation result through the multi-modal large language sub-model based on the image feature data and the target prompt word template, and outputting a quantitative image quality evaluation result corresponding to the image to be evaluated for image quality through the multi-layer perceptron sub-model based on the word vector.
[0049] The image quality evaluation method provided in this embodiment can accurately output a qualitative image quality evaluation result and a quantitative image quality evaluation result for the image to be evaluated for image quality by inputting the image to be evaluated for image quality into a pre-trained target multi-modal image quality evaluation model, thereby realizing a qualitative description and a quantitative score of the image quality by utilizing the powerful image understanding ability, description ability, and multi-task learning ability of the target multi-modal image quality evaluation model.
[0050] The following will elaborate in detail on each step in steps S200 to S206 and optional other steps in combination with Figure 2 .
[0051] Step S200 , obtaining an image to be evaluated for image quality.
[0052] In specific implementation, the image to be evaluated for image quality can be images in various formats on a network platform, or multiple video frames obtained by extracting video frames from movies, documentaries, animated videos, teaching videos, etc. on the network platform.
[0053] Among them, the network platform can be a social platform, a video sharing platform, an e-commerce platform, etc. For example, Instagram, TikTok, YouTube, Weibo, Douyin, etc.
[0054] In an alternative embodiment, obtaining the image to be evaluated for image quality includes: Extracting multiple video frames from the video to be evaluated for image quality as the image to be evaluated for image quality.
[0055] In one implementation manner, when it is necessary to evaluate the image quality of the video to be evaluated for image quality, multiple video frames will be randomly or extracted from the video to be evaluated for image quality at a preset frame extraction frequency as the image to be evaluated for image quality for subsequent evaluation.
[0056] In another implementation manner, when it is necessary to evaluate the image quality of the video to be evaluated for image quality, one video frame can also be randomly extracted from the video to be evaluated for image quality as the image to be evaluated for image quality for subsequent evaluation, or the first frame, the last frame, or the middle frame in the video to be evaluated for image quality can be used as the image to be evaluated for image quality for subsequent evaluation.
[0057] In this embodiment, by using multiple video frames as the image to be evaluated for image quality, it is convenient to output more accurate qualitative and quantitative image quality evaluation results by comprehensively considering the evaluation results of multiple video frames later.
[0058] Step S202 Input the image to be evaluated for image quality and the pre-set target prompt template into a pre-trained target multi-modal image quality evaluation model, where the target multi-modal image quality evaluation model includes a multi-modal large language sub-model, a visual sub-model, and a multi-layer perceptron sub-model.
[0059] The target prompt template (Prompt template) refers to a structured framework used to construct guiding prompts (Prompt). It helps the model better understand the task requirements and generate high-quality outputs by defining the input, context, instructions, and expected output formats of the task.
[0060] As an example, the target prompt template may include the following content: "The quality of the image has obvious noise." "The image has obvious (noise / artifacts / blur)." "Compared with the previous image, the noise in this image is (higher / lower)." "The image is clear and there is almost no blur." "The image has a relatively high noise level." "This video has an obvious frame inconsistency problem." The target multi-modal image quality evaluation model may include a multi-modal large language sub-model, a visual sub-model, and a multi-layer perceptron sub-model. These three models, as sub-models of the target multi-modal image quality evaluation model, will cooperate with each other to output the qualitative image quality evaluation result and the quantitative image quality evaluation result corresponding to the image to be evaluated for image quality.
[0061] Among them, the visual sub-model can be used to extract features from the image to be evaluated for image quality to obtain image feature data.
[0062] The multi-modal large language sub-model can be used to output the qualitative image quality evaluation result corresponding to the image to be evaluated for image quality and the word vector associated with the quantitative image quality evaluation result based on the image feature data and the prompt template.
[0063] The multi-layer perceptron sub-model can be used to output the quantitative image quality evaluation result based on the word vector.
[0064] In this embodiment, the target multi-modal image quality evaluation model is composed of a multi-modal large language sub-model, a visual sub-model, and a multi-layer perceptron sub-model, so as to accurately realize the qualitative description and quantitative scoring of the image quality.
[0065] Step S204 , extract features from the image to be evaluated for image quality through the visual sub-model to obtain image feature data.
[0066] In some embodiments, after the target multi-modal image quality evaluation model obtains the image to be evaluated for image quality, it will immediately extract features from the image to be evaluated for image quality through the visual sub-model therein to obtain image feature data.
[0067] In this embodiment, the visual sub-model can extract low-level visual features from the image to be evaluated for image quality as image feature data. For example, visual features such as color, noise, artifacts, and blur are extracted as image feature data. The visual sub-model can be a convolutional neural network model, a ResNet model, a VGG16 model, a ViT model, etc.
[0068] It should be noted that the image feature data is represented by a group of vectors.
[0069] In an alternative embodiment, the visual sub-model is preferably a ViT model.
[0070] In this embodiment, by using the ViT model as the visual sub-model, compared with other models (such as convolutional neural network models), the ViT model can perform global modeling on all parts of the image through the self-attention mechanism of Transformer. Therefore, it performs better in capturing image details, visual features, and long-range dependencies. In addition, the ViT model can adapt to input images of various sizes and forms.
[0071] In an alternative embodiment, when multiple video frames are used as the image to be evaluated for image quality, feature extraction is performed on the image to be evaluated for image quality through the visual sub-model, and the obtained image feature data includes: Feature extraction is respectively performed on multiple video frames through the visual sub-model to obtain image feature data corresponding to each video frame.
[0072] In some embodiments, after the target multi-modal image quality evaluation model obtains multiple video frames, it will perform feature extraction on the multiple video frames respectively through the visual sub-model therein to obtain image feature data corresponding to each video frame.
[0073] As an example, there are 3 video frames input to the target multi-modal image quality evaluation model, namely video frame 1, video frame 2, and video frame 3. The visual sub-model will perform feature extraction on video frame 1, video frame 2, and video frame 3 respectively, so as to obtain image feature data of video frame 1, image feature data of video frame 2, and image feature data of video frame 3.
[0074] Step S206 , the target multi-modal large language sub-model outputs the qualitative image quality evaluation result corresponding to the image to be evaluated for image quality and the word vector associated with the quantitative image quality evaluation result based on the image feature data and the target prompt word template, and the multi-layer perceptron sub-model outputs the quantitative image quality evaluation result corresponding to the image to be evaluated for image quality based on the word vector.
[0075] In some embodiments, after the image feature data is extracted by the visual sub-model, the image feature data and the prompt word template will be used together as the input data of the multi-modal large language sub-model, and the multi-modal large language sub-model will output the qualitative image quality evaluation result corresponding to the image to be evaluated for image quality and the word vector associated with the quantitative image quality evaluation result based on these data.
[0076] The qualitative image quality evaluation result refers to using qualitative descriptive words to describe the image quality evaluation result. For example, "The image color performance is excellent and relatively noise-free", "The image color is poor and there are serious artifacts", etc.
[0077] The quantitative image quality evaluation result refers to using quantitative descriptors to describe the image quality evaluation result. For example, "Quality score: 92 / 100", "Quality score: 80 / 100", etc.
[0078] Among them, the word vector is the probability of Tokens used to represent quantitative attribute words in multiple dimensions. The quantitative attribute words may include color, noise, artifacts, blur, temporal consistency, etc.
[0079] The multi-modal large language sub-model is obtained by training an open-source large language model using training sample data. Among them, the large language model can be Shusheng large model, Doubao large model, etc.
[0080] The multi-modal large language sub-model obtained in the above manner has good multi-task learning ability and strong text generation ability. Therefore, it can generate natural language descriptions related to image quality according to the input image features to represent the qualitative image quality evaluation result, and can also generate word vectors associated with the quantitative image quality evaluation result, facilitating the subsequent generation of quantitative image quality evaluation results.
[0081] After the multi-modal large language sub-model outputs the word vector associated with the quantitative image quality evaluation result, it will input the word vector into the multi-layer perceptron sub-model, so that the multi-layer perceptron sub-model can output the quantitative image quality evaluation result based on the word vector, that is, output the quantitative score of the image, such as the image quality score.
[0082] It should be noted that the multi-layer perceptron sub-model can include multiple fully connected layers and use activation functions to increase the non-linear ability of the model.
[0083] In one embodiment, in order to enable the multi-layer perceptron sub-model to output a more accurate quantitative image quality evaluation result, additional image features of the image to be evaluated for quality can also be extracted as the input data of the multi-layer perceptron sub-model. Among them, the additional image features may include feature data such as the color histogram, noise intensity, and blur of the image to be evaluated for quality.
[0084] In an alternative embodiment, after the vision sub-model obtains the image feature data corresponding to multiple video frames, referring to Figure 3 , step S206 may include: Step S300, through the multi-modal large language sub-model, based on the image feature data corresponding to each video frame and the prompt word template, output the qualitative frame image quality evaluation result corresponding to each video frame and the word vector associated with the qualitative frame image quality evaluation result.
[0085] Step S302, through the multi-layer perceptron sub-model, based on the word vector corresponding to each video frame, output the quantitative frame image quality evaluation result corresponding to each video frame.
[0086] Step S304: Determine the qualitative image quality evaluation result corresponding to the image to be evaluated based on the qualitative frame image quality evaluation results corresponding to each video frame, and determine the quantitative image quality evaluation result corresponding to the image to be evaluated based on the quantitative frame image quality evaluation results corresponding to each video frame.
[0087] In this embodiment, when performing image quality evaluation on a video to be evaluated, the target multi-modal image quality evaluation model will evaluate each video frame extracted from the video to be evaluated respectively to obtain the qualitative image quality evaluation results and quantitative image quality evaluation results of each video frame. Then, the final qualitative image quality evaluation result and quantitative image quality evaluation result can be determined by synthesizing the qualitative image quality evaluation results and quantitative image quality evaluation results of multiple frames.
[0088] In an alternative embodiment, determining the qualitative image quality evaluation result corresponding to the image to be evaluated based on the qualitative frame image quality evaluation results corresponding to each video frame includes: Conduct statistical analysis on the qualitative frame image quality evaluation results corresponding to each video frame; Use the qualitative frame image quality evaluation result with the largest quantity as the qualitative image quality evaluation result corresponding to the image to be evaluated.
[0089] In one implementation manner, when determining the qualitative image quality evaluation result corresponding to the image to be evaluated based on the qualitative frame image quality evaluation results corresponding to each video frame, statistical analysis can be conducted on the qualitative frame image quality evaluation results corresponding to each video frame to obtain the quantities of various types of qualitative frame image quality evaluation results. Then, the qualitative frame image quality evaluation result with the largest quantity can be used as the qualitative image quality evaluation result corresponding to the image to be evaluated.
[0090] In another implementation manner, it is also possible to combine the words related to image quality in various types of qualitative frame image quality evaluation results to generate the final qualitative image quality evaluation result.
[0091] In an alternative embodiment, determining the quantitative image quality evaluation result corresponding to the image to be evaluated based on the quantitative frame image quality evaluation results corresponding to each video frame includes: Calculate the average value of the quantitative frame image quality evaluation results corresponding to each video frame, and use the average value as the quantitative image quality evaluation result corresponding to the image to be evaluated.
[0092] In one implementation manner, when determining the final quantitative image quality evaluation result based on the quantitative frame image quality evaluation results of multiple frames, the average value of multiple quantitative frame image quality evaluation results (quality scores) can be used as the quantitative image quality evaluation result corresponding to the image to be evaluated.
[0093] In another embodiment, when determining the final quantitative image quality evaluation result based on the quantitative frame image quality evaluation results of multiple frames, the median of the multiple quantitative frame image quality evaluation results can also be used as the final quantitative image quality evaluation result.
[0094] In this embodiment, by evaluating the image quality of a video based on multiple video frames in the video when evaluating the image quality of the video, an accurate evaluation of the image quality of the video can be achieved.
[0095] In an alternative embodiment, refer to Figure 4 , the target multi-modal image quality evaluation model is trained through the following operations: Step S400, obtain a qualitative training sample data set, the qualitative training sample data set includes a plurality of qualitative training sample images, and each qualitative training sample image carries a qualitative attribute label.
[0096] The sources of the qualitative training sample images can be various types of UGC content, including pictures, short videos, long videos, live video replays, etc., to ensure that the qualitative training sample data set can cover various forms of expression.
[0097] In addition, to ensure the generalization ability of the model, the qualitative training sample data set needs to be large enough (tens of millions of samples) and diverse, that is, the qualitative training sample images in the qualitative training sample data set need to cover images under different shooting devices, shooting scenes, lighting conditions, user backgrounds, etc.
[0098] The qualitative attribute label can be a description word obtained by manual annotation by trained annotators to describe the quality of the image in a specific dimension (such as color, noise, etc.), such as "the left picture has more vivid colors" or "the right picture has less noise", etc.
[0099] Step S402, input the qualitative training sample data set and a pre-set first prompt word template into the initial model for training to obtain an initial multi-modal image quality evaluation model.
[0100] The initial multi-modal image quality evaluation model can output accurate qualitative image quality evaluation results.
[0101] Step S404, obtain a quantitative training sample data set, the quantitative training sample data set includes a plurality of quantitative training sample images, and each quantitative training sample image carries multi-dimensional quantitative attribute labels, and the multi-dimensional quantitative attribute labels include at least two of a color label, a noise label, an artifact label, a blur label, and a temporal consistency label.
[0102] The source of the quantitative training sample images can be various types of UGC content, including pictures, short videos, long videos, live broadcast replays, etc., ensuring that the quantitative training sample dataset can cover multiple forms of expression.
[0103] In addition, to ensure the generalization ability of the model, the quantitative training sample dataset needs to be large enough (tens of millions of samples) and diverse, that is, the quantitative training sample images in the quantitative training sample dataset need to cover images under different shooting devices, shooting scenes, lighting conditions, user backgrounds, etc.
[0104] The quantitative attribute labels in each dimension annotate the true quality score of the image in that dimension. For example, the color label is "70" or "80", etc. The noise label is "40" or "55", etc., the artifact label is "10" or "76", etc., the blur label is "54" or "77", etc., and the temporal consistency label is "30" or "90", etc.
[0105] Among them, the color label can be obtained by analyzing the color of the image through a professional color analysis tool to evaluate aspects such as color vividness, saturation, and hue balance, and finally scoring based on the combined score values of these aspects.
[0106] The noise label can be obtained by quantifying the noise value in the image through noise estimation algorithms - waveform analysis, frequency domain analysis, etc., to classify different types of noise (such as Gaussian noise, pepper noise) and score.
[0107] The artifact label can use image restoration techniques to detect and label the types of artifacts, such as compression artifacts, mosaics, blur artifacts, etc., and then use existing image restoration models to assist in judging the severity of the artifacts. Finally, the score is obtained according to the severity.
[0108] The blur label can be obtained by quantifying the blur degree of the image through methods such as edge detection and Fourier transform, and scoring according to the blur degree.
[0109] The temporal consistency label can be obtained by using methods such as optical flow method and frame difference method to evaluate the temporal consistency and scoring according to the detection results.
[0110] Step S406, input the qualitative training sample dataset and the pre-set second prompt template into the initial multi-modal image quality evaluation model for training to obtain the target multi-modal image quality evaluation model.
[0111] In this embodiment, the model is initially trained with a qualitative training sample dataset so that the obtained initial multi-modal image quality evaluation model has a more accurate evaluation ability for qualitative image quality evaluation results. Then, the model is further trained with the qualitative training sample dataset so that the finally obtained target multi-modal image quality evaluation model also has a more accurate evaluation ability for quantitative image quality evaluation results.
[0112] It should be noted that in other embodiments, the model can also be trained with a quantitative training sample dataset first and then with a qualitative training sample dataset, or the qualitative training sample dataset and the quantitative training sample dataset can be input into the model simultaneously for training.
[0113] In an alternative embodiment, to enable the target multi-modal image quality evaluation model to have a more accurate evaluation ability, in the qualitative training sample dataset, multiple pairs of qualitative sample images are used as multiple qualitative training sample images, and the difference between the two sample images in the qualitative sample image pair in the target dimension exceeds a preset value.
[0114] In the quantitative training sample dataset, multiple pairs of quantitative sample images are used as multiple quantitative training sample images, and the difference between the two sample images in the quantitative sample image pair in the target dimension exceeds a preset value.
[0115] The two sample images in the quantitative sample image pair can be selected from the images used as quantitative training sample images.
[0116] The two sample images in the qualitative sample image pair can be selected from the images used as qualitative training sample images.
[0117] Among them, the target dimension can be a color dimension, a noise dimension, an artifact dimension, a blur dimension, etc.
[0118] The preset value can be set according to the actual situation.
[0119] As an example, if the target dimension is the color dimension, the preset value can be 10%. That is to say, the difference between the two sample images in the qualitative sample image pair in the color dimension exceeds 10%.
[0120] As an example, if the target dimension is the noise dimension, the preset value can be 5%. That is to say, the difference between the two sample images in the qualitative sample image pair in the noise dimension exceeds 5%.
[0121] As an example, if the target dimension is the blur dimension, the preset value can be 3%. That is to say, the difference between the two sample images in the qualitative sample image pair in the noise dimension exceeds 3%.
[0122] As an example, if the target dimension is the color dimension, the preset value can be 7%. That is to say, the difference in the color dimension between the two sample images in the quantitative sample image pair exceeds 7%.
[0123] As an example, if the target dimension is the noise dimension, the preset value can be 4%. That is to say, the difference in the noise dimension between the two sample images in the quantitative sample image pair exceeds 4%.
[0124] As an example, if the target dimension is the blur degree dimension, the preset value can be 6%. That is to say, the difference in the noise dimension between the two sample images in the quantitative sample image pair exceeds 6%.
[0125] In this embodiment, by doping sample image pairs as training sample images in the training sample dataset, the model can accurately learn the attribute features of each dimension during the training process.
[0126] Embodiment 2 Figure 5 Schematically shows a block diagram of a picture quality evaluation device 500 according to Embodiment 2 of the present application. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 5 shown, the device 500 may include: an acquisition module 510, an input module 520, an extraction module 530, and an output module 540, where: The acquisition module 510 is configured to acquire an image to be evaluated for picture quality; The input module 520 is configured to input the image to be evaluated for picture quality and a preset target prompt word template into a pre-trained target multi-modal picture quality evaluation model. The target multi-modal picture quality evaluation model includes a multi-modal large language sub-model, a visual sub-model, and a multi-layer perceptron sub-model.
[0127] The extraction module 530 is configured to extract features from the image to be evaluated for picture quality through the visual sub-model to obtain image feature data.
[0128] The output module 540 is configured to output a qualitative picture quality evaluation result corresponding to the image to be evaluated for picture quality and a word vector associated with the quantitative picture quality evaluation result through the multi-modal large language sub-model based on the image feature data and the target prompt word template, and output a quantitative picture quality evaluation result corresponding to the image to be evaluated for picture quality through the multi-layer perceptron sub-model based on the word vector.
[0129] As an alternative embodiment, the target multi-modal image quality assessment model includes a multi-modal large language sub-model, a visual sub-model, and a multi-layer perceptron sub-model; The visual sub-model is used to extract features from the image to be quality-assessed to obtain image feature data; The multi-modal large language sub-model is used to output the qualitative image quality assessment result corresponding to the image to be quality-assessed and the word vector associated with the quantitative image quality assessment result based on the image feature data and the prompt template; The multi-layer perceptron sub-model is used to output the quantitative image quality assessment result based on the word vector.
[0130] As an alternative embodiment, obtaining the image to be quality-assessed includes: Extracting multiple video frames from the video to be quality-assessed as the image to be quality-assessed.
[0131] As an alternative embodiment, extracting features from the image to be quality-assessed by the visual sub-model to obtain image feature data includes: Extracting features from multiple video frames respectively by the visual sub-model to obtain the image feature data corresponding to each video frame; Outputting the qualitative image quality assessment result corresponding to the image to be quality-assessed and the word vector associated with the quantitative image quality assessment result by the multi-modal large language sub-model based on the image feature data and the target prompt template, and outputting the quantitative image quality assessment result corresponding to the image to be quality-assessed by the multi-layer perceptron sub-model based on the word vector includes: Outputting the qualitative frame image quality assessment result corresponding to each video frame and the word vector associated with the qualitative frame image quality assessment result by the multi-modal large language sub-model based on the image feature data corresponding to each video frame and the prompt template; Outputting the quantitative frame image quality assessment result corresponding to each video frame by the multi-layer perceptron sub-model based on the word vector corresponding to each video frame; Determining the qualitative image quality assessment result corresponding to the image to be quality-assessed based on the qualitative frame image quality assessment results corresponding to each video frame, and determining the quantitative image quality assessment result corresponding to the image to be quality-assessed based on the quantitative frame image quality assessment results corresponding to each video frame.
[0132] As an alternative embodiment, determining the qualitative image quality assessment result corresponding to the image to be quality-assessed based on the qualitative frame image quality assessment results corresponding to each video frame includes: Performing statistical analysis on the qualitative frame image quality assessment results corresponding to each video frame; Taking the qualitative frame image quality assessment result with the largest quantity as the qualitative image quality assessment result corresponding to the image to be quality-assessed; Determining the quantitative image quality evaluation result corresponding to the image to be evaluated based on the quantitative frame image quality evaluation results corresponding to each video frame includes: Calculating the average value of the quantitative frame image quality evaluation results corresponding to each video frame, and using the average value as the quantitative image quality evaluation result corresponding to the image to be evaluated.
[0133] As an alternative embodiment, the target multi-modal image quality evaluation model is trained through the following operations: Obtaining a qualitative training sample data set, where the qualitative training sample data set includes multiple qualitative training sample images, and each qualitative training sample image carries a qualitative attribute label; Inputting the qualitative training sample data set and a preset first prompt word template into an initial model for training to obtain an initial multi-modal image quality evaluation model; Obtaining a quantitative training sample data set, where the quantitative training sample data set includes multiple quantitative training sample images, and each quantitative training sample image carries multi-dimensional quantitative attribute labels, and the multi-dimensional quantitative attribute labels include at least two of a color label, a noise label, an artifact label, a blurriness label, and a temporal consistency label; Inputting the qualitative training sample data set and a preset second prompt word template into the initial multi-modal image quality evaluation model for training to obtain the target multi-modal image quality evaluation model.
[0134] As an alternative embodiment, in the qualitative training sample data set, multiple pairs of qualitative sample images are used as multiple qualitative training sample images, and the difference between the two sample images in the qualitative sample image pair in the target dimension exceeds a preset value; In the quantitative training sample data set, multiple pairs of quantitative sample images are used as multiple quantitative training sample images, and the difference between the two sample images in the quantitative sample image pair in the target dimension exceeds a preset value.
[0135] Embodiment III Figure 6 Schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing the image quality evaluation method according to Embodiment III of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers). As Figure 6As shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 10010 can be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 can also be an external storage device of the computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 can also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the image quality evaluation method. In addition, the memory 10010 can also be used to temporarily store various data that have been output or will be output.
[0136] In some embodiments, the processor 10020 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.
[0137] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal via a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0138] It should be noted that Figure 6 Only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0139] In this embodiment, the image quality evaluation method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.
[0140] Embodiment 4 The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the image quality evaluation method in the embodiments are implemented.
[0141] In this embodiment, the computer-readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), random access memories (RAM), static random access memories (SRAM), read-only memories (ROM), electrically erasable programmable read-only memories (EEPROM), programmable read-only memories (PROM), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the image quality evaluation method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various data that have been output or will be output.
[0142] Embodiment 5 The embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.
[0143] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device. Thus, they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to be implemented. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0144] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are similarly included in the patent protection scope of the present application.
Claims
1. A method for evaluating image quality, characterized in that, The method includes: Obtain the image to be evaluated for image quality; Input the image to be evaluated for image quality and a pre-set target prompt word template into a pre-trained target multi-modal image quality evaluation model, where the target multi-modal image quality evaluation model includes a multi-modal large language sub-model, a visual sub-model, and a multi-layer perceptron sub-model; Extract features from the image to be evaluated for image quality through the visual sub-model to obtain image feature data; Output the qualitative image quality evaluation result corresponding to the image to be evaluated for image quality and the word vector associated with the quantitative image quality evaluation result based on the image feature data and the target prompt word template through the multi-modal large language sub-model, and output the quantitative image quality evaluation result corresponding to the image to be evaluated for image quality based on the word vector through the multi-layer perceptron sub-model.
2. The method according to claim 1, wherein Obtain the image to be evaluated for image quality, including: Extract multiple video frames from the video to be evaluated for image quality as the image to be evaluated for image quality.
3. According to the method described in claim 2, extracting features from the image to be evaluated for image quality through the visual sub-model to obtain image feature data includes: Extract features from each of the multiple video frames through the visual sub-model to obtain the image feature data corresponding to each video frame; Output the qualitative frame image quality evaluation result corresponding to each video frame and the word vector associated with the qualitative frame image quality evaluation result based on the image feature data corresponding to each video frame and the prompt word template through the multi-modal large language sub-model, and output the quantitative frame image quality evaluation result corresponding to each video frame based on the word vector corresponding to each video frame through the multi-layer perceptron sub-model includes: Output the qualitative frame image quality evaluation result corresponding to each video frame and the word vector associated with the qualitative frame image quality evaluation result based on the image feature data corresponding to each video frame and the prompt word template through the multi-modal large language sub-model; Output the quantitative frame image quality evaluation result corresponding to each video frame based on the word vector corresponding to each video frame through the multi-layer perceptron sub-model; Determine the qualitative image quality evaluation result corresponding to the image to be evaluated for image quality based on the qualitative frame image quality evaluation results corresponding to each video frame, and determine the quantitative image quality evaluation result corresponding to the image to be evaluated for image quality based on the quantitative frame image quality evaluation results corresponding to each video frame.
4. The method according to claim 3, wherein Determining the qualitative image quality evaluation result corresponding to the image to be evaluated for image quality based on the qualitative frame image quality evaluation results corresponding to each video frame includes: Perform statistical analysis on the qualitative frame image quality evaluation results corresponding to each video frame; Use the qualitative frame image quality evaluation result with the largest quantity as the qualitative image quality evaluation result corresponding to the image to be evaluated for image quality; Determining the quantitative image quality evaluation result corresponding to the image to be evaluated for image quality based on the quantitative frame image quality evaluation results corresponding to each video frame includes: Calculate the average value of the quantitative frame image quality evaluation results corresponding to each video frame, and use the average value as the quantitative image quality evaluation result corresponding to the image to be evaluated for image quality.
5. The method according to any one of claims 1 to 4, characterized in that, The target multi-modal image quality evaluation model is obtained through the following operations for training: Obtain a qualitative training sample data set, where the qualitative training sample data set includes multiple qualitative training sample images, and each qualitative training sample image carries a qualitative attribute label; Input the qualitative training sample data set and the pre-set first prompt template into the initial model for training to obtain an initial multi-modal image quality assessment model; Obtain a quantitative training sample data set, where the quantitative training sample data set includes multiple quantitative training sample images, and each quantitative training sample image carries multi-dimensional quantitative attribute labels, and the multi-dimensional quantitative attribute labels include at least two of color labels, noise labels, artifact labels, blur labels, and temporal consistency labels; Input the qualitative training sample data set and the pre-set second prompt template into the initial multi-modal image quality assessment model for training to obtain the target multi-modal image quality assessment model.
6. The method according to claim 5, wherein In the qualitative training sample data set, multiple pairs of qualitative sample images are used as multiple qualitative training sample images, and the difference between the two sample images in the qualitative sample image pair in the target dimension exceeds a preset value; In the quantitative training sample data set, multiple pairs of quantitative sample images are used as multiple quantitative training sample images, and the difference between the two sample images in the quantitative sample image pair in the target dimension exceeds a preset value.
7. An image quality evaluation device, characterized in that, The device includes: An acquisition module for acquiring an image to be evaluated for image quality; An input module for inputting the image to be evaluated for image quality and the pre-set target prompt template into a pre-trained target multi-modal image quality assessment model, where the target multi-modal image quality assessment model includes a multi-modal large language sub-model, a visual sub-model, and a multi-layer perceptron sub-model; An extraction module for extracting feature data of the image to be evaluated for image quality through the visual sub-model to obtain image feature data; An output module for outputting a qualitative image quality assessment result corresponding to the image to be evaluated for image quality and a word vector associated with the quantitative image quality assessment result based on the image feature data and the target prompt template through the multi-modal large language sub-model, and outputting a quantitative image quality assessment result corresponding to the image to be evaluated for image quality based on the word vector through the multi-layer perceptron sub-model.
8. A computer device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein: The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to claims 1 to 6 are implemented.
Citation Information
Cited By
Video evaluation method and device based on artificial intelligence, equipment and medium
CN120852971A