A method and system for evaluating generative image videos

By combining subjective and objective evaluation methods and systems, the problem of lack of effective evaluation of the quality of generated images in the existing technology is solved, and a multi-dimensional comprehensive evaluation is achieved, which improves the scientificity and practicality of the evaluation.

CN117237296BActive Publication Date: 2025-05-27GUANGDONG BOHUA UHD INNOVATION CENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311186342.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-14
Publication Date
2025-05-27
Estimated Expiration
2043-09-14

AI Technical Summary

Technical Problem

The existing technology lacks effective subjective and objective evaluation methods to evaluate the quality and effectiveness of generated image-like videos, resulting in limited industrial applications and technological progress.

Method used

By combining subjective and objective evaluation, a method and system is designed, including a data subsystem, a subjective evaluation subsystem and an objective evaluation subsystem. The system collects and labels image video data, conducts subjective evaluation and forms a scoring data set, and uses this data set to train a feature comparison model for objective evaluation.

Benefits of technology

A multi-dimensional comprehensive evaluation of generated image-like videos is realized, which improves the scientificity and practicality of the evaluation, so that the quality and effect of image videos can be optimized according to user needs and preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117237296B_ABST
    Figure CN117237296B_ABST
Patent Text Reader

Abstract

The present invention provides a method and a system for evaluating generative image videos. The method includes: S1. Data collection and annotation; S2. Data generation: using the text description in step S1 as the input of the generative model, generating data and keeping the technical parameters of the simultaneously annotated image videos in step S1 consistent; S3. Data alignment; S4. Subjective evaluation of data; S5. Data model training: training the feature comparison model in the method using the subjective scoring data set formed in step S4; S6. Outputting the feature comparison model: outputting the trained feature comparison model and using it in the objective evaluation of the objective evaluation subsystem. This method and system can objectively and subjectively evaluate the quality and effect of generative image videos, and promote the development of generative-related technologies and industries by combining subjective evaluation and objective evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and system for evaluating generated image videos. Background Art

[0002] Content generation of images and videos (AIGC) is an important field of artificial intelligence technology. Recently, significant progress has been made in this type of technology, and the technology has gradually matured and been widely applied. This type of technology involves various types of tasks, such as generating image videos from text, generating image videos from images, generating image videos from text, and repairing text combined with image videos. These tasks all require high-quality image video outputs to meet different human visual perception needs. The main problem in the prior art is that there is currently no effective subjective and objective evaluation method to evaluate and test the quality and effect-related content of the generated image videos. On the one hand, it hinders the application and development of the generation-related industries, and at the same time, it also makes it difficult for the industrial and academic communities to compare and evaluate different methods and models, restricting the progress of the technology.

[0003] The difficulty in solving the above problems and defects is that the currently popular FID (with reference) and CLIPScore (without reference) evaluation methods are already outdated and have a large difference from human subjective values. At present, there is no mature subjective and objective evaluation method for reference, and it is difficult to solve this problem.

[0004] The significance of solving the above problems and defects is that the method and system for evaluating generated image videos proposed by the present invention can effectively provide subjective and objective evaluations and feedback for current image video generation-related methods, applications, etc., and can promote the progress and development of related algorithms and the entire industry. Summary of the Invention

[0005] The present invention provides a method and system for evaluating generated image videos, which can objectively and subjectively evaluate the quality and effect of generated image videos, and promote the development of generation-related technologies and industries by combining subjective evaluation and objective evaluation.

[0006] The technical solution of the present invention is as follows:

[0007] According to one aspect of the present invention, a method for evaluating generated images and videos is provided, comprising the following steps: S1. Data collection and annotation: collecting different types of images and videos, and describing the images and videos in text; S2. Data generation: using the text description in step S1 as the input of the generated model, generating data and keeping the data consistent with the technical parameters of the images and videos annotated at the same time in step S1; S3. Data alignment: aligning the data according to the original text description, the original image and video data, and the generated data to ensure their correspondence and consistency; S4. Subjective evaluation data: evaluating and scoring the aligned data one by one according to the main dimensions and contents of the subjective evaluation, and obtaining the scores of each evaluation to form a subjective scoring data set; S5. Feature comparison model training: using the subjective scoring data set formed in step S4 to train the feature comparison model in the objective method; S6. Outputting the feature comparison model: outputting the trained feature comparison model and using it in the objective evaluation of the objective evaluation subsystem.

[0008] Optionally, in the above-mentioned method for evaluating generated image videos, in step S1, the image video is described in text in the order of the subject in the image video, the details in the image video, the modification of the content in the image video, and the description of the style of the image video, and the resolution and encoding format of the image video are annotated to form a reference data set; at the same time, only natural language is used to describe different scenes to form a non-reference data set.

[0009] Optionally, in the above-mentioned method for evaluating generated image videos, in step S4, the steps of evaluating and scoring include: A1. preparing equipment: preparing at least two displays, a data playback device and a room with controllable light; A2. preparing data: preparing generated data and source data to be evaluated; A3. whether there is a reference: whether there is reference data in the data, if there is reference data, proceed to step A4, if there is no reference data, proceed to step A5; A4. evaluation with reference: for evaluation with reference, it is necessary to display the source data, reference data and data to be evaluated to several evaluators, and play the data at a fixed time for the evaluators to watch; A5. evaluation without reference: for evaluation without reference, it is necessary to display the source data and data to be evaluated to several evaluators, and play the data at a fixed time for the evaluators to watch; A6. selecting evaluation items: according to different requirements, selecting different evaluation options for combined evaluation; A7. evaluation and scoring: subjectively scoring the evaluation data according to the evaluation requirements and subjective feelings; A8. result calculation: summarizing the scores of all evaluators, and calculating the score of the evaluation object according to certain rules.

[0010] Optionally, in the above method for evaluating generative image videos, in step S6, the steps for the objective evaluation subsystem to conduct objective evaluation include: T1. Select evaluation content: Select the dimensions and content to be evaluated for objective evaluation; T2. Set the feature extraction network: Set different feature extraction networks and other modules for combination according to the evaluation content selected in step T1; T3. Set the feature comparison model: Set different feature comparison models for the model according to different evaluation requirements; T4. Calculate the evaluation score: Only evaluate one item at a time. After calculating the score of this item, return to step T1 to select other options for evaluation. After all the evaluated options are evaluated, summarize the individual scores of different evaluation items and calculate the overall score; T5. Output the evaluation result: The objective evaluation outputs the evaluation result to be combined and complemented with the subjective evaluation result.

[0011] Optionally, in the above method for evaluating generative image videos, in step T2, the combination method is: Combine the networks according to different evaluation requirements, and use a two-stream or multi-stream converter feature network to implement feature extraction for different data.

[0012] Optionally, in the above method for evaluating generative image videos, in step T3, design a feature comparison model to conduct different feature comparisons. The feature comparison model includes feature space transformation and feature similarity calculation. The feature space transformation projects the features of different modalities into the same feature space for similarity comparison, and design a feature comparison method with adaptive weights. According to different evaluation requirements, the feature comparison model adaptively assigns different weights to the features of different layers and different modalities, and then conducts feature similarity calculation to output relevant comparison results.

[0013] According to another aspect of the present invention, there is provided a method for evaluating generative image videos, including three subsystems: a data subsystem, a subjective evaluation subsystem, and an objective evaluation subsystem. Among them, the data subsystem is used for data collection, data generation, and data annotation; the subjective evaluation subsystem is used for subjectively evaluating the generated image videos by using the subjective evaluation method; the objective evaluation subsystem is used for training the designed feature extraction network and feature comparison model by using the relevant data of the subjective evaluation.

[0014] According to the technical solution of the present invention, the beneficial effects are:

[0015] The content generation of images and videos in the present invention belongs to an emerging hot technology field, mainly serving for human subjectivity. Compared with the previous evaluation of images and videos only focusing on the quality of images and videos, the present invention mainly evaluates the newly emerging images and videos generated by artificial intelligence technology. It not only evaluates from a single dimension of the quality of images and videos, but also conducts the determination evaluation of whether the images and videos are generated, the determination evaluation of the generated images and videos and the generation sources, and also evaluates from the aspects of the emotion and security of the generated images and videos, the data bias and aesthetics of the generated images and videos, etc., to realize a multi-dimensional, comprehensive and effective method and system for evaluating intelligent generated images and videos from subjectivity to objectivity, improve the scientificity and practicality of the comprehensive evaluation of generated images and videos, and enable the quality and effect of generated images and videos to be continuously optimized and improved according to user needs and preferences.

[0016] To better understand and illustrate the concept, working principle and invention effect of the present invention, the following will combine with the attached drawings and through specific embodiments, elaborate on the present invention in detail as follows: Brief Description of the Drawings

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the attached drawings required for use in the description of the specific embodiments or the prior art.

[0018] Figure 1 is the flowchart of the method for evaluating generated images and videos of the present invention;

[0019] Figure 2 is the schematic structural diagram of the entire objective evaluation network where the source data is only natural language;

[0020] Figure 3 is the schematic structural diagram of the entire objective evaluation network where the source data is only image data;

[0021] Figure 4 is the schematic structural diagram of the entire objective evaluation network where the source data is natural language and images;

[0022] Figure 5 is the flowchart of the method for subjective evaluation by the subjective evaluation subsystem; and

[0023] Figure 6 is the flowchart of the method for objective evaluation by the objective evaluation subsystem. Detailed Embodiments

[0024] To make the purpose, technical method and advantages of the present invention clearer, the following will further elaborate on the present invention in detail in combination with the attached drawings and specific examples. These examples are merely illustrative and not restrictive of the present invention.

[0025] The principle of the present invention is as follows: The method of the present invention includes five main contents designed to conduct subjective and objective evaluations of generated image videos. The present invention designs a subjective evaluation method; in the objective evaluation, the system of the present invention includes three subsystems: a data subsystem, a subjective evaluation subsystem, and an objective evaluation subsystem. These three subsystems are interrelated and can also be used independently.

[0026] The evaluation system for generated image videos of the present invention includes three subsystems: a data subsystem, a subjective evaluation subsystem, and an objective evaluation subsystem. These three subsystems are interrelated and can also be used independently; among them, the data subsystem is mainly for data collection, data generation, and data annotation. The subjective evaluation subsystem uses the subjective evaluation method to subjectively evaluate the generated image videos, while the objective evaluation subsystem uses the relevant data of the subjective evaluation to train the designed feature extraction network and feature comparison model.

[0027] As Figure 1 shown, the method for evaluating generated image videos of the present invention includes the following steps:

[0028] S1. Data collection and annotation: Collect different types of image videos, and use professional annotators to describe the image videos in words. Specifically, describe the image videos in the order of the main body in the image video, the details in the image video, the modification of the content in the image video, and the description of the image video style, and at the same time annotate technical parameters such as the resolution and encoding format of the image video to form a reference data set; at the same time, only use natural language to describe different scenarios to form a non-reference data set.

[0029] S2. Data generation: Use the text description in step S1 as the input of the generation model, and use a variety of different generation algorithms to generate data, and keep the technical parameters of the image videos marked simultaneously in step S1 consistent.

[0030] S3. Data alignment: Align the data according to the original text description, the original image video data, and the generated data to ensure their correspondence and consistency.

[0031] S4. Subjective evaluation data: Evaluate and score the aligned data one by one according to the main dimensions and contents of the subjective evaluation, and obtain the scores of each evaluation to form a subjective scoring data set.

[0032] Design a subjective evaluation method, such as evaluating the clarity, authenticity, and fit of the generated data. In combination with specific tasks, the following all or part of the dimensions can be selected for evaluation and assessment:

[0033] a) Quality assessment of image and video: It is divided into 5 levels. The image quality is very poor, the image quality is poor, the image is average, the image quality is good, and the image quality is very good. This standard mainly examines the visual effects such as the resolution, color, contrast, and details of the generated data, as well as whether there are defects such as noise, blur, and distortion.

[0034] b) Evaluation of the generation traces of image and video: It is divided into 5 levels. It is an intelligently generated image, may be intelligently generated, uncertain, may not be intelligently generated, and is not intelligently generated. This standard mainly examines whether the generated data conforms to the real world, whether there are phenomena violating physical laws or common sense, whether there are obvious splicing traces or unnatural transitions, and whether there are obvious traces of intelligent generation;

[0035] c) Evaluation of the association between the image and video and the generation source: It is divided into 5 levels. It is completely irrelevant, has a large difference, has a partial difference, has some small differences, and completely matches. This standard mainly examines whether the generated data has a logical or semantic connection with the input sources such as the text, audio, and image provided by the user, and whether it can reflect the information such as the theme, style, and intention of the input source;

[0036] d) Evaluation of the emotions and safety brought: It is divided into 5 levels. It makes people feel severely uneasy, makes people feel uneasy and depressed, makes people feel a small amount of negative emotions, makes people feel plain, and makes people feel happy and satisfied. This standard mainly examines whether the generated data can cause positive emotional feedback from users, whether it meets the expectations and needs of users, whether it will generate content that damages physical and mental health, and whether it may cause discomfort or trouble to users;

[0037] e) Evaluation of the data bias brought: It is divided into 5 levels. No bias is detected, slight bias is detected, moderate bias, severe bias, extremely severe bias and discrimination. This standard mainly examines whether there are phenomena such as discrimination, demeaning, misleading, and distorting certain groups, individuals, events, viewpoints, etc. in the generated data, and whether it may affect the judgment or decision-making of users;

[0038] f) Aesthetic evaluation: It is divided into 5 levels. It conforms to the aesthetic of the general public, conforms to the aesthetic of a specific group, conforms to the personal aesthetic, and does not conform to any aesthetic. This standard mainly examines whether the generated data has beauty and artistry, whether it can attract and move users, and whether it can express a certain aesthetic concept or value.

[0039] The above is the design and description of the subjective evaluation. The setting of its score levels can be set to different levels according to the actual situation. For example, 10 levels can be set.

[0040] The subjective evaluation subsystem mainly adopts a series of scientific processes and system designs, with professionals with video evaluation experience as the main evaluators. According to the evaluation items, the intelligently generated image videos that need to be evaluated are evaluated and scored, such as Figure 5 As shown, the specific steps are as follows:

[0041] A1. Prepare equipment: First prepare the monitors needed for the evaluation, at least two monitors, data playback equipment, and a room with controllable lighting;

[0042] A2. Data preparation: Prepare the generated data and source data required for evaluation;

[0043] A3. Whether there is a reference: whether there is reference data in the data. If there is reference data, proceed to step A4. If there is no reference data, proceed to step A5.

[0044] A4. Evaluation with reference: For evaluation with reference, the source data, reference data and data to be evaluated need to be displayed to several evaluators, and the data is played for a fixed duration for the evaluators to watch;

[0045] A5. Evaluation without reference: For evaluation without reference, the source data and the data to be evaluated need to be displayed to several evaluators, and the data is played for a fixed duration for the evaluators to watch;

[0046] A6. Select evaluation items: According to different requirements, select different evaluation options for combined evaluation;

[0047] A7. Evaluation and scoring: Subjectively score the evaluation data according to the evaluation requirements and subjective feelings;

[0048] A8. Result calculation: Summarize the scores of all evaluators and calculate the score of the evaluation object according to certain rules.

[0049] S5. Feature comparison model training: Use the subjective scoring data set formed in step S4 to train the feature comparison model in the objective method.

[0050] S6. Output feature comparison model: Output the trained feature comparison model and use it in the objective evaluation of the objective evaluation subsystem.

[0051] The objective evaluation subsystem mainly uses the trained Transformer feature network and feature comparison model to objectively evaluate and score different evaluation contents, such as Figure 6 As shown, the specific steps of the objective evaluation method are as follows:

[0052] T1. Select the evaluation content: select the dimensions and content to be evaluated for objective evaluation;

[0053] T2. Set the feature extraction network: Set different feature extraction networks and other modules for combination according to the content selected and evaluated in step T1. The combination method (designing and training a specific objective evaluation network for the combination) is as follows: Combine the networks according to different evaluation requirements, and use a two-stream or multi-stream Transformers feature network to achieve feature extraction for different data. If the source in the evaluation is natural language, model A (as shown in Figure 2 is used); if the source in the evaluation is an image, model B (as shown in Figure 3 is used), and if the source in the evaluation is natural language and an image, model C (as shown in Figure 4 is used).

[0054] Specifically, if the source data is only natural language, the designed network feature extraction model is: a language Transformers feature network to extract the features of the source data, a visual Transformers feature network to extract the features of the generated data, and then a feature comparison model is used to compare the features and output the results. The overall network structure of the objective evaluation is as shown in Figure 2 Similarly, if the source data is only image data, the overall network structure of the objective evaluation is as shown in Figure 3 If the source data is natural language and an image, the overall network structure of the objective evaluation is as shown in Figure 4 , and it is necessary to combine all the features of the source data and compare them with the features of the generated data.

[0055] T3. Set the feature comparison model: Set different feature comparison models according to different evaluation requirements;

[0056] Specifically, design and train the feature comparison model of the objective evaluation method: Design a feature comparison model to perform different feature comparisons. The feature comparison model includes feature space transformation and feature similarity calculation. Feature space transformation refers to projecting the features of different modalities into the same feature space so that similarity comparison can be performed. And design a feature comparison method with adaptive weights. According to different evaluation requirements, the feature comparison model can adaptively assign different weights to the features of different layers and different modalities (visual modality and natural language modality data), improve the effectiveness and accuracy of the comparison, and then perform feature similarity calculation to output the relevant comparison results.

[0057] T4. Calculate the evaluation score: Only one item of content is evaluated each time. After calculating the score of this item, return to step T1 to select other options for evaluation. After all the evaluated options are evaluated, summarize the individual scores of different evaluation items and calculate the overall score;

[0058] T5. Output evaluation results: The objective evaluation outputs the evaluation results, which can be combined and complemented with the subjective evaluation results.

[0059] Among them, the language Transformers feature network for designing and training the objective evaluation method: Based on the basic transformers network, the attention mechanism is fully used to extract the features of natural language. Before use, the model needs to be pre-trained with a large amount of natural language data. Then, the language Transformers feature extraction module is taken out and combined with other modules to form Model A and Model C respectively according to different requirements. After the overall training of Model A and Model C using subjectively collected data to achieve the best effect, the model can be deployed and applied.

[0060] The visual Transformers feature network for designing and training the objective evaluation method: Based on the basic transformers network, the attention mechanism is fully used to extract the features of images. Before use, the model needs to be pre-trained with a large amount of image data. Then, the Transformers feature extraction module of the image is taken out and combined with other modules to form Model A, Model B, and Model C respectively according to different requirements. After the overall training of Model A, Model B, and Model C using subjectively collected data to achieve the best effect, the model can be deployed and applied.

[0061] The present invention provides a comprehensive and effective evaluation system for generated image and video, which is beneficial to the standardization, normalization, and optimization of generated-related technologies and industries. The present invention provides a set of effective methods and systems for subjective evaluation and objective evaluation for intelligent-generated images and videos. Combining the objective evaluation results and the subjective evaluation results, a comprehensive evaluation result of the generated image and video is given and fed back to the user. On the one hand, it can effectively evaluate intelligent-generated images and videos from multiple dimensions, and on the other hand, the loss function can be continuously optimized according to the evaluation results to improve the generation effect of the image and video quality.

[0062] The above description is the best embodiment according to the concept and working principle of the invention. The above embodiments should not be construed as limiting the protection scope of the present claims. Combinations of other implementation manners and implementation modes according to the concept of the present invention all belong to the protection scope of the present invention.

Claims

1. A method for evaluating generative image videos, characterized in that, it includes the following steps: S1. Data collection and annotation: Collect different types of image videos, describe the image videos in text, and describe the image videos in the order of the main body in the image video, the details in the image video, the modification of the content in the image video, and the description of the image video style. At the same time, annotate the resolution and coding format of the image video to form a reference data set; at the same time, only use natural language to describe different scenarios to form a non-reference data set; S2. Data generation: Use the text description in step S1 as the input of the generative model to generate data and keep it consistent with the technical parameters of the image video simultaneously annotated in step S1; S3. Data alignment: Align the data according to the original text description, original image video data, and generated data to ensure their correspondence; S4. Subjective evaluation data: Evaluate and score the aligned data item by item according to the main dimensions and content of subjective evaluation, and obtain the scores of each evaluation to form a subjective scoring data set; S5. Feature comparison model training: Use the subjective scoring data set formed in step S4 to train the feature comparison model in the objective method; S6. Output the feature comparison model: Output the trained feature comparison model and use it in the objective evaluation of the objective evaluation subsystem. Specifically, the steps for the objective evaluation subsystem to conduct objective evaluation include: T1. Select evaluation content: Select the dimensions and content to be evaluated for objective evaluation; T2. Set the feature extraction network: Set different feature extraction networks and other modules in combination according to the evaluation content selected in step T1; T3. Set the feature comparison model: Set different feature comparison models for the model according to different evaluation requirements; T4. Calculate the evaluation score: Only evaluate one item at a time, calculate the score of this item and then return to step T1 to select other options for evaluation. After all the evaluation options are evaluated, summarize the individual scores of different evaluation items and calculate the overall score; T5. Output the evaluation result: The objective evaluation outputs the evaluation result to be combined and complemented with the subjective evaluation result In step T3, design the feature comparison model to perform different feature comparisons. The feature comparison model includes feature space transformation and feature similarity calculation. The feature space transformation projects the features of different modalities into the same feature space for similarity comparison, and designs a feature comparison method with adaptive weights. According to different evaluation requirements, the feature comparison model adaptively assigns different weights to the features of different layers and different modalities, and then performs feature similarity calculation to output relevant comparison results.

2. The method for evaluating generative image videos according to claim 1, characterized in that, in step S4, the steps of evaluation and scoring include: A1. Prepare equipment: Prepare at least two monitors, a data playback device, and a room with controllable lighting; A2. Data preparation: Prepare the generated data and source data required for evaluation; A3. Whether there is a reference: whether there is reference data in the data, if there is reference data, proceed to step A4, if there is no reference data, proceed to step A5; A4. Evaluation with reference: For evaluation with reference, the source data, reference data and data to be evaluated need to be displayed to several evaluators, and the data is played for a fixed duration for the evaluators to watch; A5. Evaluation without reference: For evaluation without reference, the source data and the data to be evaluated need to be displayed to several evaluators, and the data is played for a fixed duration for the evaluators to watch; A6. Select evaluation items: According to different requirements, select different evaluation options for combined evaluation; A7. Evaluation and scoring: Subjectively score the evaluation data according to the evaluation requirements and subjective feelings; A8. Result calculation: Summarize the scores of all evaluators and calculate the score of the evaluation object according to certain rules.

3. The method for evaluating generated image videos according to claim 1, It is characterized in that In step T2, the combination method is: combining networks according to different evaluation requirements, and using a dual-stream or multi-stream converter feature network to achieve feature extraction of different data.

4. A system for evaluating generated image videos, used to implement the method for evaluating generated image videos as claimed in any one of claims 1 to 3, It is characterized in that It includes three subsystems: data subsystem, subjective evaluation subsystem and objective evaluation subsystem, among which: The data subsystem is used for data collection, data generation and data annotation; The subjective evaluation subsystem is used to perform subjective evaluation on the generated image video using a subjective evaluation method; The objective evaluation subsystem is used to train the designed feature extraction network and feature comparison model using the relevant data of the subjective evaluation.

Citation Information

Patent Citations

  • Subjective and objective ultrasonic medical image quality evaluation method and system

    CN113628174A