Multi-modal understanding and generation evaluation method and system oriented to Chinese context

By constructing an image and video evaluation dataset oriented towards Chinese elements, generating reference JSON using the GPT-4V reference model, and performing structured field alignment and dynamic weighted evaluation, the shortcomings of existing multimodal models in Chinese context evaluation are addressed, enabling fine-grained evaluation and performance analysis of the model.

CN120932044APending Publication Date: 2025-11-11CHINA UNIV OF MINING & TECH (BEIJING)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511038074.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing multimodal model evaluation methods lack semantic interpretation and structured analysis capabilities for Chinese elements in the Chinese context, making it difficult to evaluate the performance of models on complex and variable tasks in open environments.

Method used

We construct an image and video evaluation dataset with Chinese elements as its features, generate reference JSON using the GPT-4V reference model, and evaluate the model output, including the similarity between the understanding task and the generation task, through structured field alignment and dynamic weighting strategies.

Benefits of technology

It enables fine-grained evaluation of multimodal models in Chinese context, provides an interpretable automated evaluation process, reveals the model's inference path and performance bottlenecks, and adapts to the complex semantic challenges of Chinese-characteristic elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932044A_ABST
    Figure CN120932044A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal understanding and generation evaluation method for Chinese context, and relates to the technical field of large models, and the method comprises the steps: constructing an image and video evaluation data set for Chinese element features; constructing a reference answer text to form a reference description set; a reference JSON is constructed; generating understanding task test description; generating task test description; carrying out structured field alignment on the understanding task test description by utilizing a GPT-4 model so as to construct an understanding task test JSON (JavaScript Object Notation); carrying out structured field alignment on the generated task test description by utilizing a GPT-4 model so as to construct a generated task test JSON (JavaScript Object Notation); calculating the similarity between the structured fields of the understanding task test JSON and the reference JSON, and generating the similarity between the structured fields of the task test JSON and the reference JSON; and respectively introducing a dynamic weighting strategy and calculating respective total scores. The invention further discloses a multi-modal understanding and generation evaluation system oriented to the Chinese context. According to the method and the system, a reproducible standardized evaluation process with context awareness is established.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large model technology, and in particular to a multimodal understanding and generation evaluation method and system for Chinese context. Background Technology

[0002] In model evaluation, semantic alignment technology aims to achieve accurate semantic correspondence between different modalities (such as images, text, and videos); multimodal understanding and generation tasks cover obtaining information from multiple sources such as images, text, and videos, and further performing modal transformation to output different modal elements, which is one of the key directions in model evaluation; model evaluation assesses the multidimensional capabilities of a series of models by constructing a scientific, systematic, and comprehensive evaluation system.

[0003] Traditional benchmarks (such as COCO (Lin TY, Maire M, Belongie S, et al., Microsoft COCO: Common Objects in Context, ECCV 2014) and VQAv2 (Teney D, Anderson P, He X, et al., Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge, CVPR 2018)) are widely used due to their popularity and stability. COCO is primarily used for evaluating image description tasks, employing automated metrics such as BLEU and CIDEr to measure text generation quality; VQAv2 tests the model's visual understanding and reasoning abilities through question-answering tasks. However, these benchmarks are mainly task-specific and lack a comprehensive evaluation of large multimodal models in complex reasoning and knowledge integration, making it difficult to reflect the model's performance in real-world scenarios.

[0004] In recent years, researchers have proposed a variety of more targeted multimodal evaluation benchmarks. MME (Fu C, Zhang YF, Yin S, et al., MME-survey: A comprehensive survey on evaluation of multimodal LLMs, arXiv:2411.15296, 2024) designed a comprehensive benchmark including 14 perceptual and cognitive tasks, systematically evaluating 30 large multimodal models through manually constructed instruction-answer pairs; MMBEC (Liu Y, Duan H, Zhang Y, et al., MMBEC: Is your multi-modal model an all-around player?, ECCV 2024:216–233) introduced a fine-grained annotation mechanism for 20 capability dimensions, and combined dataset annotation with the CircularEval strategy, significantly improving the comprehensiveness and robustness of the evaluation by addressing the problems of insufficient granularity and unstable indicators in previous assessments; SEED-Bench (Li B, Ge Y, Ge Y, et al.) The benchmark (al., SEED-Bench: Benchmarking Multimodal Large Language Models, CVPR 2024: 13299–13308) consists of 19,000 precisely manually annotated multiple-choice questions, covering 12 evaluation dimensions, including image and video modal understanding, spatiotemporal reasoning, etc., and systematically evaluates the capabilities of 18 multimodal large models. Although the above benchmark can effectively test the perception and reasoning capabilities of multimodal large models through a question-and-answer format with pre-defined task boundaries, it also limits the free exploration of multimodal large models in open environments.

[0005] For Chinese multimodal tasks, CMMU (He Z, Wu X, Zhou P, et al., CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning, arXiv:2401.14011, 2024) proposed a Chinese context-oriented evaluation framework, which designs three question types—multiple choice, multiple answer, and fill-in-the-blank—around text-image comprehension to evaluate understanding and reasoning abilities in Chinese multimodal tasks. MMMU (Yue X, Ni Y, Zhang K, et al., MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI, CVPR 2024:9556–9567) simulates expert-level multidisciplinary task scenarios, conducting experiments on 16 mainstream multimodal large models to evaluate their interdisciplinary reasoning abilities and potential for general artificial intelligence.

[0006] Meanwhile, researchers have also gradually focused on the performance of multimodal large models in real-world open tasks, proposing a series of evaluation methods oriented towards practical scenarios. For example, LLaVA-Bench (Liu H, Li C, Wu Q, et al., Visual Instruction Tuning, NeurIPS 2024, 36) proposed quantitative evaluation metrics, using the GPT model to evaluate the relevance and accuracy of multimodal large model outputs, and verified the positive impact of image-instruction-annotation diversity on performance; BLINK (Fu X, Hu Y, Li B, et al., BLINK: Multimodal Large Language Models Can See But Not Perceive, ECCV 2024: 148–166) designed an evaluation scheme including 14 visual perception tasks, systematically exploring the capability boundaries of multimodal large models in visual understanding. However, these evaluation methods are still task-oriented and do not truly simulate the complex and ever-changing task environments in open worlds.

[0007] Chinese elements, rich in symbolism, aesthetic traditions, and metaphorical meanings, represent an ideal yet challenging domain for evaluating the capabilities of multimodal models. Evaluating multimodal large language models in domains rich in Chinese characteristics (such as Chinese tradition and art) presents complex semantic challenges: models must be able to distinguish between textual and visual imagery, understand metaphorical meanings, and integrate multimodal information. For example, recognizing a "festival atmosphere" requires integrating visual cues such as lanterns, traditional clothing, and group behavior.

[0008] While multimodal models have made significant progress in understanding and generating tasks across image, text, and video modalities, existing evaluation methods still have significant shortcomings in semantic interpretability, adaptability to Chinese element features, and structured analysis capabilities. On the one hand, existing evaluation systems often rely on single metrics or coarse-grained scoring, lacking fine-grained evaluation of models at specific semantic dimensions, making it difficult to reveal the model's inference path and performance bottlenecks. On the other hand, existing evaluation datasets primarily consist of general content, lacking a systematic examination of the graphic imagery, metaphorical meanings, and multimodal symbol cognition within Chinese-specific element features. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a multimodal understanding and generation evaluation method and system for Chinese context. The present invention constructs an image and video evaluation dataset for Chinese element features and establishes a standardized evaluation process that is context-aware and reproducible.

[0010] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0011] A multimodal understanding and generation evaluation method for Chinese context proposed in this invention includes:

[0012] Step 1: Construct an image and video evaluation dataset featuring Chinese elements;

[0013] Step 2: Based on the image and video evaluation dataset featuring Chinese elements, construct reference answer text to form a reference description set;

[0014] The reference answer text in the reference description set is processed using the GPT-4V reference model to construct reference JSON, which is a reference JSON in a uniform structured format based on key contextual elements.

[0015] Step 3: Using the image and video evaluation dataset with Chinese element features, generate the output description of the model to be tested based on the image and video understanding task. This output description is the test description of the understanding task.

[0016] Using a reference description set, an output description of the model under test is generated based on the image and video generation task. This output description is a test description of the generation task.

[0017] Step 4: Use the GPT-4 model to perform structured field alignment on the comprehension task test description to construct the comprehension task test JSON.

[0018] The generated task test description is then used to perform structured field alignment using a GPT-4 model to construct the generated task test JSON.

[0019] Calculate the similarity of structured fields between the task test JSON and the reference JSON, and generate the similarity of structured fields between the task test JSON and the reference JSON;

[0020] Step 5: Based on Step 4, calculate the similarity of the structured fields of the understanding task test JSON and the reference JSON, and the similarity of the structured fields of the generated task test JSON and the reference JSON. Introduce dynamic weighting strategies and calculate their respective total scores.

[0021] As a further optimization of the multimodal understanding and generation evaluation method for Chinese context described in this invention, step 1 includes:

[0022] Step 1.1: Collect raw data from the data source. Based on the raw data, obtain multiple images and multiple video clips. The multiple images and multiple video clips constitute the initial sample pool. The data source includes image-text pairs and video and accompanying text pairs.

[0023] Step 1.2: Automated screening of samples in the initial sample pool, filtering images and videos with resolutions below a preset resolution threshold and those whose content is irrelevant to Chinese element features; further screening removes samples with ambiguous semantic features of Chinese elements, forming the final image and video sample set for evaluation of Chinese element features, which is the image and video evaluation dataset for Chinese element features.

[0024] As a further optimization of the multimodal understanding and generation evaluation method for Chinese context described in this invention, the image and video sample set includes categories such as Chinese food, scenery, clothing, art, festivals, and customs.

[0025] As a further optimization of the multimodal understanding and generation evaluation method for Chinese context described in this invention, step 2 includes:

[0026] Step 2.1: For the image and video sample set with Chinese element features in Step 1.2, use the GPT-4V reference model to generate descriptive text, and perform text filtering based on semantic consistency to obtain the reference description set of the image and video sample set;

[0027] Step 2.2: Perform semantic analysis on the reference description set in Step 2.1 to extract key contextual elements, including object, background, text, and element features. Then, use the GPT-4 model to process the reference answer text in the reference description set to construct a reference JSON with a unified structured format based on the key contextual elements.

[0028] As a further optimization of the multimodal understanding and generation evaluation method for Chinese context described in this invention, the field structure of the reference JSON is consistent with the key contextual elements, including objects, background, text, and element features.

[0029] The objects include entities and their core attributes presented in images or videos. The object fields are divided into name, features, and appearance subfields.

[0030] Background information includes the space and environmental context that describes the scene, as well as recognizable Chinese landmark elements.

[0031] The text includes text elements with Chinese characteristics;

[0032] The elements include Chinese cuisine, scenery, clothing, art, festivals, and customs.

[0033] As a further optimization of the multimodal understanding and generation evaluation method for Chinese context described in this invention, the subfields of the object field are used to capture fine-grained object semantic information, Chinese element landmarks are used to emphasize their role in scene positioning and style expression, and text is used to evaluate the ability to recognize and understand text content embedded in images and videos.

[0034] As a further optimization of the multimodal understanding and generation evaluation method for Chinese context described in this invention, in step 3...

[0035] Generate an output description of the model under test based on the image and video understanding task, including:

[0036] The image and video sample sets of Chinese element features from step 1.2 are sequentially input into the model to be tested, and the corresponding text descriptions generated are recorded as the comprehension task test descriptions.

[0037] Generate an output description of the model under test based on the image and video generation task; including:

[0038] The reference description set is input into the model under test, and images or videos corresponding to the reference description set are generated. Then, the GPT-4V reference model is used to analyze the images or videos generated by the model under test, and the visual content is transformed into a generation task test description.

[0039] As a further optimization of the multimodal understanding and generation evaluation method for Chinese context described in this invention, step 4 includes:

[0040] Step 4.1: Extract the four fields of "understanding task test description" and "generating task test description" one by one, including "object", "background", "text" and "element features". Use the GPT-4 model to perform structured field alignment to ensure that the four fields are completely consistent, and obtain the "understanding task test JSON" and "generating task test JSON".

[0041] Step 4.2: Input the reference JSON and the comprehension task test JSON, as well as the reference JSON and the generation task test JSON, into the GPT-4 model. Perform a one-to-one semantic comparison of the four fields (object, background, text, and element features) of the reference JSON and the comprehension task test JSON, and perform a one-to-one semantic comparison of the four fields (object, background, text, and element features) of the reference JSON and the generation task test JSON. The GPT-4 model outputs the semantic similarity scores corresponding to the four fields (object, background, text, and element features) of the reference JSON and the comprehension task test JSON, which are the similarity scores of the structured fields of the comprehension task test JSON and the reference JSON, and the similarity scores of the structured fields of the generation task test JSON and the reference JSON.

[0042] As a further optimization of the multimodal understanding and generation evaluation method for Chinese context described in this invention, step 5 includes:

[0043] Step 5.1: Assign weight coefficients to the four semantic similarity scores corresponding to the four fields of object, background, text, and element features in Step 4.2, and generate a weight vector;

[0044] Step 5.2: The semantic similarity scores of the four fields (object, background, text, and element features) are weighted and summed with their corresponding weight vectors to obtain the total score, which serves as the overall evaluation result of the model under test.

[0045] A multimodal understanding and generation evaluation system for Chinese context includes:

[0046] The dataset building module is used to construct image and video evaluation datasets that feature Chinese elements.

[0047] The JSON module is used to construct reference answer text and form a reference description set based on image and video evaluation datasets featuring Chinese elements.

[0048] The reference answer text in the reference description set is processed using the GPT-4V reference model to construct reference JSON, which is a reference JSON in a uniform structured format based on key contextual elements.

[0049] The output description module is used to generate an output description of the model under test based on the image and video understanding task using an image and video evaluation dataset featuring Chinese elements. This output description is a test description of the understanding task.

[0050] Using a reference description set, an output description of the model under test is generated based on the image and video generation task. This output description is a test description of the generation task.

[0051] The similarity calculation module is used to align the structured fields of the comprehension task test description using the GPT-4 model to construct the comprehension task test JSON.

[0052] The generated task test description is then used to perform structured field alignment using a GPT-4 model to construct the generated task test JSON.

[0053] Calculate the similarity of structured fields between the task test JSON and the reference JSON, and generate the similarity of structured fields between the task test JSON and the reference JSON;

[0054] The scoring module is used to calculate the total score by introducing a dynamic weighting strategy based on the similarity of the structured fields of the understanding task test JSON and the reference JSON, and the similarity of the structured fields of the generated task test JSON and the reference JSON.

[0055] Compared with the prior art, the present invention, employing the above technical solution, has the following technical effects:

[0056] This invention constructs an image and video evaluation dataset oriented towards Chinese elements, and uses this dataset to construct reference answer text, forming a reference description set. It then designs a unified understanding task and generation evaluation framework, extracts structured contextual semantics, aligns structured fields to generate corresponding JSON, and introduces the GPT4 evaluation model as an interpretable automated evaluator. Based on the similarity of structured fields, a dynamic weighting strategy is introduced to calculate individual scores and the total score. The above content is systematically validated on image and video understanding and generation tasks to establish a context-aware and reproducible standardized evaluation process. Attached Figure Description

[0057] Figure 1 This is a flowchart of the evaluation framework;

[0058] Figure 2 These are evaluation examples;

[0059] Figure 3 This is a flowchart of the present invention. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] like Figure 3 As shown, a multimodal understanding and generation evaluation method for Chinese context includes:

[0062] Step 1: Construct an image and video evaluation dataset featuring Chinese elements; details are as follows:

[0063] Step 1.1: Collect raw data from the data source. Based on the raw data, obtain multiple images and multiple video clips. The multiple images and multiple video clips constitute the initial sample pool. The data source includes image-text pairs and video and accompanying text pairs.

[0064] Step 1.2: Automated screening of samples in the initial sample pool, filtering images and videos with resolutions below a preset resolution threshold and those whose content is irrelevant to Chinese element features; further screening removes samples with ambiguous semantic features of Chinese elements, forming the final image and video sample set for evaluation of Chinese element features, which is the image and video evaluation dataset for Chinese element features.

[0065] The image and video sample set includes categories such as Chinese food, scenery, clothing, art, festivals, and customs.

[0066] Step 2: Based on the image and video evaluation dataset featuring Chinese elements, construct reference answer text to form a reference description set;

[0067] The reference answer text in the reference description set is processed using the GPT-4V reference model to construct reference JSON. Reference JSON is a uniformly structured reference JSON based on key contextual elements; specifically as follows:

[0068] Step 2.1: For the image and video sample set with Chinese element features in Step 1.2, use the GPT-4V reference model to generate descriptive text, and perform text filtering based on semantic consistency to obtain the reference description set of the image and video sample set;

[0069] Step 2.2: Perform semantic analysis on the reference description set in Step 2.1 to extract key contextual elements, including object, background, text, and element features. Then, use the GPT-4 model to process the reference answer text in the reference description set to construct a reference JSON with a unified structured format based on the key contextual elements.

[0070] The field structure of the reference JSON is consistent with the key contextual elements, including object, background, text, and element characteristics.

[0071] The objects include entities and their core attributes presented in images or videos. The object fields are divided into name, features, and appearance subfields.

[0072] Background information includes the space and environmental context that describes the scene, as well as recognizable Chinese landmark elements.

[0073] The text includes text elements with Chinese characteristics;

[0074] The elements include Chinese cuisine, scenery, clothing, art, festivals, and customs.

[0075] The subfields of the object field are used to capture fine-grained semantic information of the object, the Chinese element landmarks are used to emphasize their role in scene positioning and style expression, and the text is used to evaluate the ability to recognize and understand text content embedded in images and videos.

[0076] Step 3: Using the image and video evaluation dataset with Chinese element features, generate the output description of the model to be tested based on the image and video understanding task. This output description is the test description of the understanding task.

[0077] Using a reference description set, an output description of the model under test is generated based on the image and video generation task. This output description serves as a test description for the generation task; specifically as follows:

[0078] Generate an output description of the model under test based on the image and video understanding task, including:

[0079] The image and video sample sets of Chinese element features from step 1.2 are sequentially input into the model to be tested, and the corresponding text descriptions generated are recorded as the comprehension task test descriptions.

[0080] Generate an output description of the model under test based on the image and video generation task; including:

[0081] The reference description set is input into the model under test, and the corresponding image or video output is generated. Then, the GPT-4V reference model is used to analyze the image or video generated by the model under test, and the visual content is converted into a generation task test description.

[0082] Step 4: Use the GPT-4 model to perform structured field alignment on the comprehension task test description to construct the comprehension task test JSON.

[0083] The generated task test description is then used to perform structured field alignment using a GPT-4 model to construct the generated task test JSON.

[0084] Calculate the similarity of structured fields between the task test JSON and the reference JSON, and calculate the similarity of structured fields between the generated task test JSON and the reference JSON; details are as follows:

[0085] Step 4.1: Extract the four fields of "understanding task test description" and "generating task test description" one by one, including "object", "background", "text" and "element features". Use the GPT-4 model to perform structured field alignment to ensure that the four fields are completely consistent, and obtain the "understanding task test JSON" and "generating task test JSON".

[0086] Step 4.2: Input the reference JSON and the comprehension task test JSON, as well as the reference JSON and the generation task test JSON, into the GPT-4 model. Perform a one-to-one semantic comparison of the four fields (object, background, text, and element features) of the reference JSON and the comprehension task test JSON, and perform a one-to-one semantic comparison of the four fields (object, background, text, and element features) of the reference JSON and the generation task test JSON. The GPT-4 model outputs the semantic similarity scores corresponding to the four fields (object, background, text, and element features) of the reference JSON and the comprehension task test JSON, which are the similarity scores of the structured fields of the comprehension task test JSON and the reference JSON, and the similarity scores of the structured fields of the generation task test JSON and the reference JSON.

[0087] Step 5: Based on Step 4, calculate the similarity of structured fields between the understanding task test JSON and the reference JSON, and the similarity of structured fields between the generated task test JSON and the reference JSON. Introduce a dynamic weighting strategy for each and calculate their respective total scores; details are as follows:

[0088] Step 5.1: Assign weight coefficients to the four semantic similarity scores corresponding to the four fields of object, background, text, and element features in Step 4.2, and generate a weight vector;

[0089] Step 5.2: The semantic similarity scores of the four fields of object, background, text, and element features are weighted and summed with their corresponding weight vectors to obtain the total score, which is used as the overall evaluation result of the model under test.

[0090] It also includes semantic similarity scores and similarity score explanations for four fields: GPT-4 output object, background, text, and element features, which facilitate subsequent result analysis and model optimization.

[0091] This invention also discloses a multimodal understanding and generation evaluation system for Chinese context, comprising:

[0092] The dataset building module is used to construct image and video evaluation datasets that feature Chinese elements.

[0093] The JSON module is used to construct reference answer text and form a reference description set based on image and video evaluation datasets featuring Chinese elements.

[0094] The reference answer text in the reference description set is processed using the GPT-4V reference model to construct reference JSON, which is a reference JSON in a uniform structured format based on key contextual elements.

[0095] The output description module is used to generate an output description of the model under test based on the image and video understanding task using an image and video evaluation dataset featuring Chinese elements. This output description is a test description of the understanding task.

[0096] Using a reference description set, an output description of the model under test is generated based on the image and video generation task. This output description is a test description of the generation task.

[0097] The similarity calculation module is used to align the structured fields of the comprehension task test description using the GPT-4 model to construct the comprehension task test JSON.

[0098] The generated task test description is then used to perform structured field alignment using a GPT-4 model to construct the generated task test JSON.

[0099] Calculate the similarity of structured fields between the task test JSON and the reference JSON, and generate the similarity of structured fields between the task test JSON and the reference JSON;

[0100] The scoring module is used to calculate the total score by introducing a dynamic weighting strategy based on the similarity of the structured fields of the understanding task test JSON and the reference JSON, and the similarity of the structured fields of the generated task test JSON and the reference JSON.

[0101] The multimodal understanding and generation evaluation method proposed in this invention can perform automated evaluation of multimodal understanding and generation tasks from multiple perspectives and with semantic drive, and output quantifiable scores and explanations.

[0102] In this evaluation method, the invention first collects raw data from a data source, including image-text pairs and video pairs with accompanying text. Based on the raw data, 24,455 images and 12,048 video clips are obtained to form an initial sample pool, constructing an image and video evaluation dataset oriented towards Chinese element features. Based on this, reference answer text is constructed, forming a reference description set. The reference answer text in the reference description set is processed using the GPT-4V reference model to construct reference JSON. The reference JSON is a unified structured format based on key contextual elements. To achieve a structured evaluation of the model's capabilities, this invention introduces a set of contextual elements based on key contextual elements, including four elements: object, background, text, and element features. Based on the different semantic structures of the understanding and generation tasks, key information in the model output is extracted and aligned, resulting in test descriptions for the understanding and generation tasks.

[0103] This invention utilizes the GPT-4 model to align the structured fields of the task comprehension test description and the task generation test description, thereby constructing comprehension test JSON and generation test JSON. It then performs field-level semantic comparisons of four elements: object, background, text, and element features. The similarity of the structured fields between the comprehension test JSON and the reference JSON, and between the generation test JSON and the reference JSON, is calculated, ultimately outputting a similarity score from 0 to 100. Furthermore, a dynamic weighting mechanism is introduced to assign weight coefficients, generate a weight vector, and generate the model's total score based on the weighted average.

[0104] like Figure 1 As shown, in the evaluation process of the multimodal understanding and generation evaluation method for Chinese context, a reference description set for image and video sample sets is first established. Next, according to the specific task type, the image and video sample sets are sequentially input into the model under test to generate corresponding understanding task test descriptions and generation task test descriptions, collectively referred to as the test description set. Then, both the test description set generated by the model under test and the reference description set are converted into structured understanding task test JSON and generation task test JSON using a unified structured format method, and aligned with four predefined elements: object, background, text, and element features. Finally, similarity comparisons are performed on these structured representations, and a similarity score is calculated for each element to quantitatively evaluate the model's performance.

[0105] Image or video understanding tasks:

[0106] Evaluation methods in image or video understanding tasks systematically assess a model's comprehensive semantic understanding of images or videos by generating natural language descriptions of the overall content. Specifically, the reference model GPT-4V first processes the input image or video sample set sequentially and generates text-based reference descriptions to express key information. Based on this, the text descriptions are parsed to extract contextual elements, including four types: objects, background, text, and element features. These are then organized into a standardized JSON structure, called reference JSON.

[0107] Upon receiving the same multimodal input, the evaluation model also generates a corresponding comprehension task test description and a structured comprehension task test JSON, with the same structure as the reference JSON. Subsequently, the evaluation model (in this invention, the GPT-4 model) compares the four element fields of the reference JSON and the comprehension task test JSON, and assigns a similarity score between 0 and 100 to each corresponding field.

[0108] The final total score is then weighted and aggregated using a dynamic weighting strategy that adjusts the weights of different elements based on their relative importance in a specific task. Furthermore, the evaluation process provides explanatory information for the scores of the four element fields, and the output includes both fine-grained scores for each element and the aggregated overall score.

[0109] Image or video generation tasks:

[0110] The generation task aims to generate corresponding images or videos based on given text descriptions. The text originates from a set of reference descriptions, which also serve as reference text in image or video understanding tasks. The corresponding visual reference content is directly taken from the images or videos within these text-image pairs. Specifically, the model under test generates visual output based on text prompts, and then the GPT-4V reference model converts the visual output into textual descriptions for the generation task. Simultaneously, the GPT-4V reference model also generates textual reference descriptions from the corresponding reference images or videos; both are used for subsequent comparison and evaluation.

[0111] Next, similar to the comprehension task, the GPT-4 model parses the generated task test description and reference description according to contextual elements, constructing a structured generated task test JSON. Subsequently, the GPT-4 evaluation model compares the similarity of each contextual field between the generated task test JSON and the reference JSON, assigning a score between 0 and 100. A dynamic weighting strategy, consistent with the comprehension task, is then used to aggregate the scores of each field to obtain the final score. This method ensures consistency in the evaluation process between the generated task and the comprehension task.

[0112] Unified Structured Formatting Method Prompt:

[0113] The two obtained image (video) segments are described with text: one is a reference description, and the other is a test description generated by the model. The next step is to extract and structure key information from the test description, including: background, objects, element features, and text content, and output it in the specified JSON format.

[0114] Extract key information:

[0115] Identify and extract key objects, including their appearance, actions, and location descriptions.

[0116] Identify the background in the description.

[0117] Extract any element features reflected in the description.

[0118] Extract the text content that appears in the description, and enclose it in quotation marks (if any).

[0119] Strictly adhere to JSON format:

[0120] You must fill in the form strictly according to the given JSON structure.

[0121] All fields are required, including: "Object", "Background", "Text", and "Element Characteristics".

[0122] Speculation or additions are prohibited:

[0123] It can only extract, understand, or generate information explicitly mentioned in the task test description.

[0124] No speculation or additional content may be added.

[0125] Output format:

[0126] {

[0127] "object":[{"name":"<…>",

[0128] "Feature":{"Feature 1":"<...>","Feature 2":"<...>"}}],

[0129] "Appearance":{"Appearance1":"<...>","Appearance2":"<...>"}}],

[0130] Background: "<...>",

[0131] "Text":"<...>",

[0132] "Element characteristics":"<...>"

[0133] }

[0134] The present invention provides an embodiment, a multi-modal understanding and generation benchmark MUGC in the context of Chinese culture:

[0135] To meet the unique requirements of Chinese characteristic element content, MUGC designs four types of domain-specific hierarchical evaluation elements in the unified evaluation framework method, which are as follows:

[0136] Object: Identify the main objects presented in the image or video and their core attributes. This element is further divided into sub-fields of name, feature, and appearance to capture fine-grained object semantic information.

[0137] Background: Describe the spatial and environmental context of the scene, with a focus on recognizable Chinese element landmarks such as temples, traditional gardens, etc., emphasizing their role in scene positioning and style expression.

[0138] Text: Evaluate the ability of the model to recognize and understand the text content embedded in the figure, including text elements with Chinese characteristics such as Chinese characters (e.g., the character "Fu"), slogans, or Spring Festival couplets.

[0139] Element Feature: Evaluate the model's ability to understand Chinese element features from two dimensions: one is the recognition ability of feature symbols (such as lanterns, dragons, etc.); the other is the ability to accurately grasp and express the overall element atmosphere.

[0140] The invention designs prompt words based on the four elements of object, background, text, and element feature in MUGC and inputs them into the GPT-4 model for evaluation. The GPT-4 model first extracts a structured JSON expression, and then compares the reference JSON with the understanding task test JSON or the generation task test JSON to generate a similarity score between 0 and 100. This process ensures item-by-item comparison and quantitative evaluation among the elements at the structured level.

[0141] Based on the above evaluation framework method, data collection is carried out. A total of 24,455 images and 12,048 videos are collected. Low-quality or content-irrelevant samples are automatically screened out and further manually reviewed to remove samples lacking Chinese characteristic element features. Finally, 120 representative images and 120 videos covering a variety of different Chinese aesthetic elements are selected from the screened sample pool. For each image and video sample, a descriptive text is generated using the GPT-4V reference model. On the premise of considering the input token limit problem in image generation and video generation tasks, the invention refines the generation results, removing ambiguous expressions, redundant information, and irrelevant details. After processing, while retaining the core element content, the average length of the Chinese description is reduced from the original 72 words to 48 words, and the English description is reduced from 135 words to 97 words.

[0142] Similarity comparison method suggestion words:

[0143] You are given two images with formatted descriptive text. The reference JSON is the reference result, and the test JSON is the result generated by the model under test. Your task is to score the accuracy of the model's generated result based on the reference result.

[0144] Comparison content:

[0145] Compare the similarity of three aspects: background, object, and element features.

[0146] Rate each attribute individually, ranging from 0 to 100, and enter the score in the corresponding score field. Text field:

[0147] If the standard answer contains a text field, then compare and score it;

[0148] If the text field does not exist, the rating is skipped.

[0149] Overall rating:

[0150] Based on the importance distribution (priority) of the current sample task attributes, corresponding weights are established, and the weights are summed to obtain an overall score.

[0151] Reasons for rating:

[0152] Briefly explain the reasons for each score;

[0153] Summarize the main similarities and differences to explain the basis for the overall score.

[0154] Experimental model:

[0155] This invention evaluates multiple multimodal large language models on understanding tasks (image understanding, video understanding) and generation tasks (image generation, video generation) to verify their performance in understanding and generating content with Chinese characteristics.

[0156] For image understanding and video understanding tasks, the invention selected mainstream open-source multimodal visual language models to ensure that the evaluation is comprehensive and representative.

[0157] Image understanding models include: Qwen-vl-chat, Qwen2-vl-7B-Instruct, Qwen2.5-vl-7B-Instruct, Onevision-qwen2-7B, Llava-v1.6-mistral-7B, Llama3-llava-next-8B, InternVL2-8B, Internlm-x composer2-vl-7B, Aquila-vl-2B, Deepseek-vl-7B, MiniCPMV-2.6, Molmo-7B, Monkey-Chat, mPLUG-Owl3-7B, Janus-Pro-7B, Emu3-Chat, GLM4v-9B, Cogvlm2-llama3-chinese-19B.

[0158] The video understanding models include: Qwen2-vl-7B-Instruct, Qwen2.5-vl-7B-Instruct, Llava-Next-Video-7B, Llava-Video-7B-Qwen2, MiniCPMV-2.6, mPLUG-Owl-7B, mPLUG-Owl3-7B, VideoLLaMA2-7B, and InternVL2-8B.

[0159] The models used for the image generation task include: Stable Diffusion-v1.4 (Sdv1.4), Stable Diffusion-v1.5 (Sdv1.5), Stable Diffusion-v3.5 (Sdv3.5), Sdxl-lightning, Sdxl-turbo, Stable Diffusion XL base 1.0 (Sdxl), Dreamlike-photoreal, Emu3-Gen, Flux, Kandinsky3, Openjourney, Playgroundv2.5, and Janus-Pro-7B.

[0160] The models used in the video generation task include: Dream Machine, Gen2, Hailuo, Kling, Pika, Pixverse, Tongyi, VIDU, and Zhipu.

[0161] Experimental results and analysis:

[0162] Table 1. Evaluation results of the image understanding model based on the evaluation framework of this invention.

[0163] Model background object Element characteristics text Total Score Qwen-vl-chat 52.41 60.32 56.05 11.3 58.55 Qwen2-vl-7B-Instruct 66.80 65.11 60.32 26.7 66.70 Qwen2.5-vl-7B-Instruct 70.99 69.35 70.92 60.3 71.77 Onevision-qwen2-7B 71.97 66.61 62.63 24.6 69.04 Llava-v1.6-mistral-7B 51.91 53.74 54.04 16.7 55.26 Llama3-Ilava-next-8B 61.05 59.09 57.19 15.3 61.31 InternVL2-8B 63.60 59.21 59.52 38.3 62.43 Internlm-xcomposer2-vl-7B 54.43 57.15 51.10 12.7 56.59 Aquila-vl-2B 59.39 54.84 50.13 10.3 55.72 Deepseek-vl-7B 69.96 63.91 61.88 55.3 67.16 MiniCPMV-2.6 73.99 67.33 65.26 9 70.19 Molmo-7B 64.30 57.02 57.63 39.3 60.02 Monkey-Chat 57.15 60.60 59.98 33.9 61.27 mPLUG-Owl3-7B 55.39 55.50 56.05 15.7 57.45 Janus-Pro-7B 69.43 68.28 62.59 36.3 69.51 Emu3-Chat 71.01 62.44 60.57 65 66.32 GLM4v-9B 65.70 63.73 61.84 40.7 65.31 Cogvlm2-llama3-chinese-19B 64.21 65.18 64.34 29.3 66.28

[0164] Table 2 Evaluation results of the video understanding model based on the evaluation framework of this invention.

[0165] Model background object Element characteristics text Total Score Qwen2-vl-7B-Instruct 30.00 27.71 29.50 8 30.03 Qwen2.5-vl-7B-Instruct 32.93 31.02 31.63 23.7 34.09 Llava-Next-Video-7B 18.29 19.88 21.71 13.3 20.98 Llava-Video-7B-Qwen2 33.58 31.26 30.96 25 33.27 MiniCPMV-2.6 35.98 32.48 34.69 36.3 34.70 mPLUG-Owl-7B 20.42 17.39 19.33 0 19.59 mPLUG-Owl3-7B 27.17 25.50 25.75 5.3 27.04 VideoLLaMA2-7B 26.29 24.50 25.88 30 26.72 InternVL2-8B 33.75 28.83 34.38 24.7 33.24

[0166] The evaluation results for the image understanding and video understanding tasks are analyzed, as shown in Tables 1 and 2. Since this benchmark particularly emphasizes the understanding of Chinese element features, "element features" are one of the key focuses of the analysis. The reference results were generated by GPT-4V, with its output assigned a perfect score of 100, and used as the evaluation benchmark.

[0167] In the image understanding task, the overall performance of different models showed a clear distribution: scores ranged from 55 to 72, with Qwen2.5-VL-7B-Instruct performing best with an overall score of 71.77 and a significant element feature score of 70.92. Overall, most models outperformed those that understood element features in recognizing objects and backgrounds. Among the 18 participating models, the average background score was 63.54, the average object score was 61.63, and the average element feature score was 59.56. Although some models deviated from this general order, the overall trend remained clear. Notably, Lava-v1.6-mistral-7B and mPLUG-Owl3-7B had higher element feature scores (54.04 and 56.05, respectively) than their background (51.91 and 55.39) and object scores (53.74 and 55.50), but these results indicate that their overall visual recognition ability was relatively weak compared to all other models.

[0168] It can be observed that text recognition is the most challenging part of image understanding. The average text score is only 30.04, indicating a significant limitation of the model in scene text understanding. Since only about one-eighth of the reference answers contain text content, the text score, although low, has a limited impact on the overall score.

[0169] In contrast, performance on video understanding tasks declined significantly across all dimensions. Video models scored an average of 28.71 for background, 26.51 for object, and 28.20 for element features, with an overall average score of only 28.85. Even the best-performing video understanding model, MiniCPMV-2.6, scored more than 30 points lower than the average for image tasks in background (35.98), object (32.48), and element features (34.69). This decline highlights the complexity of temporal reasoning in video data and the higher demands placed on attention and memory mechanisms when processing spatiotemporal information. Text recognition in videos further deteriorated, with an average score of only 18.48, again demonstrating that scene-based text understanding remains a major weakness for the models.

[0170] Figure 2 This is a sample evaluation. "Reference JSON" represents the reference output generated by GPT-4V, while "Test JSON" corresponds to the output generated by the test model 1lava-onevision-qwen2-7B. Both contain object, background, and element feature elements. Below are the corresponding comparison scores and explanations, including evaluations of individual elements and the overall score.

[0171] Overall, while current visual language models have demonstrated good basic visual perception capabilities, they still face significant challenges in knowledge reasoning, temporal information processing, and text recognition.

[0172] Table 3 Evaluation results of the image generation model based on the evaluation framework of this invention.

[0173] Model background object Element characteristics text Total Score Sdv1.4 52.21 53.60 49.63 11.3 52.76 Sdv1.5 54.58 54.81 56.54 1.3 56.50 Sdv3.5 74.63 73.40 71.79 14.7 74.37 Sdxl-lightning 58.63 57.99 62.25 10.7 61.08 Sdxl-turbo 65.19 64.63 69.63 4.7 66.79 Sdxl 59.88 60.09 61.73 10 61.73 Dreamlike-photoreal 57.08 54.54 59.21 3.3 57.73 Emu3-Gen 67.71 65.33 65.21 29.7 67.11 Flux 74.54 73.18 69.81 46 72.96 Kandinsky3 70.92 67.38 66.28 6 69.88 Openjourney 51.96 55.94 58.29 5.3 56.35 Playgroundv2.5 64.67 64.17 61.71 8 65.88 Janus-Pro-7B 75.67 72.79 69.83 40 73.90

[0174] Table 4 Evaluation results of the video generation model based on the evaluation framework of this invention

[0175]

[0176] The evaluation results for image generation and video generation tasks are shown in Tables 3 and 4. Unlike comprehension tasks, generation tasks differ fundamentally in the setting of reference answers: the reference answers in comprehension tasks are manually refined to approximate the "ideal answer"; while generation tasks inherently do not have a single correct output, and the reference images and videos used are merely derived from the dataset and serve as "reference answers." Therefore, the scores of generation tasks are essentially relative evaluation results, and their absolute values ​​are usually higher than the scores judged based on the ideal answer, causing a certain degree of score inflation.

[0177] Despite score inflation, the performance gap between image generation and video generation remains significant. In image generation tasks, model scores range from 52 to 75, with the best-performing model, Sdv3.5, scoring 74.37. However, in video generation tasks, scores are significantly lower, ranging only from 46 to 62, with the highest-scoring model, Dream Machine, only achieving 61.36. This difference indicates that generating video content with temporal coherence and rich context remains more challenging than generating static images. Unlike images, video requires models to maintain spatial consistency across frames while accurately modeling temporal dynamics, placing higher demands on large multimodal models.

[0178] In image generation models, Sdv3.5 and Janus-Pro-7B perform well and are relatively balanced across dimensions such as element features, objects, and background. According to the technical report, this performance may be attributed to their progressive training improvement and guided distillation strategy based on a teacher-student mechanism. In contrast, even the best-performing video generation models still have significant shortcomings in element features. For example, Dream Machine's element feature score is only 57.75, far below the reference score of 100. Furthermore, although the VG model can restore object and background details in some frames, maintaining consistency between element features and context throughout the entire video remains a major bottleneck, a problem also reflected in the limited upper limit of its overall score.

[0179] Text generation remains a persistent challenge in image and video generation. The average text score for video generation is 42.73, while the average score for image generation is even lower, at only 14.69. Although some visual generation models possess some text generation capabilities, this is not their core design goal. Performance in this area often depends on the model's optimization direction, training strategy, and the characteristics of the data used. Therefore, both image and video generation models have significant limitations in terms of accuracy, fidelity, and contextual integration in text generation, a problem particularly pronounced in video tasks.

[0180] Overall, current multimodal large models have shown great potential in visual content generation, especially in static image generation tasks, but they still face many challenges in element connotation analysis, temporal reasoning, and text generation.

[0181] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A multimodal understanding and generation evaluation method for Chinese context, characterized in that, include: Step 1: Construct an image and video evaluation dataset featuring Chinese elements; Step 2: Based on the image and video evaluation dataset featuring Chinese elements, construct reference answer text to form a reference description set; The reference answer text in the reference description set is processed using the GPT-4V reference model to construct reference JSON, which is a reference JSON in a uniform structured format based on key contextual elements. Step 3: Using the image and video evaluation dataset with Chinese element features, generate the output description of the model to be tested based on the image and video understanding task. This output description is the test description of the understanding task. Using a reference description set, an output description of the model under test is generated based on the image and video generation task. This output description is a test description of the generation task. Step 4: Use the GPT-4 model to perform structured field alignment on the comprehension task test description to construct the comprehension task test JSON. The generated task test description is then used to perform structured field alignment using a GPT-4 model to construct the generated task test JSON. Calculate the similarity of structured fields between the task test JSON and the reference JSON, and generate the similarity of structured fields between the task test JSON and the reference JSON; Step 5: Based on Step 4, calculate the similarity of the structured fields of the understanding task test JSON and the reference JSON, and the similarity of the structured fields of the generated task test JSON and the reference JSON. Introduce dynamic weighting strategies and calculate their respective total scores.

2. The multimodal understanding and generation evaluation method for Chinese context as described in claim 1, characterized in that, Step 1 includes: Step 1.1: Collect raw data from the data source. Based on the raw data, obtain multiple images and multiple video clips. The multiple images and multiple video clips constitute the initial sample pool. The data source includes image-text pairs and video and accompanying text pairs. Step 1.2: Automated screening of samples in the initial sample pool, filtering images and videos with resolutions below a preset resolution threshold and those whose content is irrelevant to Chinese element features; further screening removes samples with ambiguous semantic features of Chinese elements, forming the final image and video sample set for evaluation of Chinese element features, which is the image and video evaluation dataset for Chinese element features.

3. The multimodal understanding and generation evaluation method for Chinese context as described in claim 2, characterized in that, The image and video sample set includes categories such as Chinese food, scenery, clothing, art, festivals, and customs.

4. The multimodal understanding and generation evaluation method for Chinese context as described in claim 2, characterized in that, Step 2 includes: Step 2.1: For the image and video sample set with Chinese element features in Step 1.2, use the GPT-4V reference model to generate descriptive text, and perform text filtering based on semantic consistency to obtain the reference description set of the image and video sample set; Step 2.2: Perform semantic analysis on the reference description set in Step 2.1 to extract key contextual elements, including object, background, text, and element features. Then, use the GPT-4 model to process the reference answer text in the reference description set to construct a reference JSON with a unified structured format based on the key contextual elements.

5. The multimodal understanding and generation evaluation method for Chinese context as described in claim 4, characterized in that, The field structure of the reference JSON is consistent with the key contextual elements, including object, background, text, and element characteristics. The objects include entities and their core attributes presented in images or videos. The object fields are divided into name, features, and appearance subfields. Background information includes the space and environmental context that describes the scene, as well as recognizable Chinese landmark elements. The text includes text elements with Chinese characteristics; The elements include Chinese cuisine, scenery, clothing, art, festivals, and customs.

6. The multimodal understanding and generation evaluation method for Chinese context as described in claim 5, characterized in that, The subfields of the object field are used to capture fine-grained semantic information of the object, the Chinese element landmarks are used to emphasize their role in scene positioning and style expression, and the text is used to evaluate the ability to recognize and understand text content embedded in images and videos.

7. The multimodal understanding and generation evaluation method for Chinese context as described in claim 5, characterized in that, In step 3, Generate an output description of the model under test based on the image and video understanding task, including: The image and video sample sets of Chinese element features from step 1.2 are sequentially input into the model to be tested, and the corresponding text descriptions generated are recorded as the comprehension task test descriptions. Generate an output description of the model under test based on the image and video generation task; including: The reference description set is input into the model under test, and images or videos corresponding to the reference description set are generated. Then, the GPT-4V reference model is used to analyze the images or videos generated by the model under test, and the visual content is transformed into a generation task test description.

8. The multimodal understanding and generation evaluation method for Chinese context as described in claim 7, characterized in that, Step 4 includes: Step 4.1: Extract the four fields of "understanding task test description" and "generating task test description" one by one, including "object", "background", "text" and "element features". Use the GPT-4 model to perform structured field alignment to ensure that the four fields are completely consistent, and obtain the "understanding task test JSON" and "generating task test JSON". Step 4.2: Input the reference JSON and the comprehension task test JSON, as well as the reference JSON and the generation task test JSON, into the GPT-4 model. Perform a one-to-one semantic comparison of the four fields (object, background, text, and element features) of the reference JSON and the comprehension task test JSON, and perform a one-to-one semantic comparison of the four fields (object, background, text, and element features) of the reference JSON and the generation task test JSON. The GPT-4 model outputs the semantic similarity scores corresponding to the four fields (object, background, text, and element features) of the reference JSON and the comprehension task test JSON, which are the similarity scores of the structured fields of the comprehension task test JSON and the reference JSON, and the similarity scores of the structured fields of the generation task test JSON and the reference JSON.

9. The multimodal understanding and generation evaluation method for Chinese context as described in claim 8, characterized in that, Step 5 includes: Step 5.1: Assign weight coefficients to the four semantic similarity scores corresponding to the four fields of object, background, text, and element features in Step 4.2, and generate a weight vector; Step 5.2: The semantic similarity scores of the four fields (object, background, text, and element features) are weighted and summed with their corresponding weight vectors to obtain the total score, which serves as the overall evaluation result of the model under test.

10. A multimodal understanding and generation evaluation system for Chinese context, characterized in that, include: The dataset building module is used to construct image and video evaluation datasets that feature Chinese elements. The JSON module is used to construct reference answer text and form a reference description set based on image and video evaluation datasets featuring Chinese elements. The reference answer text in the reference description set is processed using the GPT-4V reference model to construct reference JSON, which is a reference JSON in a uniform structured format based on key contextual elements. The output description module is used to generate an output description of the model under test based on the image and video understanding task using an image and video evaluation dataset featuring Chinese elements. This output description is a test description of the understanding task. Using a reference description set, an output description of the model under test is generated based on the image and video generation task. This output description is a test description of the generation task. The similarity calculation module is used to align the structured fields of the comprehension task test description using the GPT-4 model to construct the comprehension task test JSON. The generated task test description is then used to perform structured field alignment using a GPT-4 model to construct the generated task test JSON. Calculate the similarity of structured fields between the task test JSON and the reference JSON, and generate the similarity of structured fields between the task test JSON and the reference JSON; The scoring module is used to calculate the total score by introducing a dynamic weighting strategy based on the similarity of the structured fields of the understanding task test JSON and the reference JSON, and the similarity of the structured fields of the generated task test JSON and the reference JSON.

Citation Information

Patent Citations

  • Alignment evaluation method for Chinese large language model

    CN117633225A

  • Structured video understanding method and device based on multi-modal large model, computer equipment and readable storage medium

    CN119478769A