A problem-aware visual image saliency enhancement method and system
Patent Information
- Application Number
- CN202610920855.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-06-25
AI Technical Summary
然而,这类图像通常承载大量信息,当用户需要基于图像回答特定分析问题时,往往依赖手动定位相关视觉元素、逐一识别数据关系,过程耗时且认知负荷较高,尤其在图像存在视觉杂乱的情况下,效率进一步降低
在本发明中,首先识别静态可视化图像中与自由形式问题相关的答案区域,生成对应的目标区域掩码,并利用多模态大语言模型对该区域进行显著性增强,获得答案区域被显著凸显的编辑图像;在此基础上,进一步生成与该编辑图像中目标区域相对应的文本答案,并根据文本答案与编辑图像之间的一致性进行修正,输出最终的最优编辑图像。由此,最终输出的编辑图像能够直接显示答案所依据的视觉区域,且与最终文本答案形成相互验证的图文对,用户既可获得明确的答案文本,又能从图像中直观确认答案的视觉证据。本发明能够在无需依赖图表底层数据、规格文件或配对训练数据的前提下,直接对静态可视化图像进行问答感知型视觉编辑,通过自动定位答案相关区域并增强其视觉显著性,大幅降低用户认知负荷,提升图表问答效率与准确性。
Smart Images

Figure CN122473308B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image editing technology, and in particular relates to a method and system for enhancing the saliency of visual images with problem awareness. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Visualized images (such as charts and infographics) have become an important medium for conveying data information due to their ability to intuitively present complex data relationships. However, these images typically carry a large amount of information, and when users need to answer specific analytical questions based on the images, they often have to manually locate relevant visual elements and identify data relationships one by one. This process is time-consuming and cognitively demanding, and efficiency is further reduced, especially when the images are visually cluttered.
[0004] In existing technologies, automated methods often directly generate text answers but fail to provide visual evidence corresponding to the answers. Users still need to manually check the charts to confirm the reasonableness of the answers. In traditional visualization design, researchers control saliency by adjusting visual encoding, but once a static visualization image is generated, its visual saliency cannot be adaptively adjusted to match the query needs of different users.
[0005] Multimodal large language models (MLLMs) have demonstrated strong capabilities in chart question answering, data reasoning, and chart generation. However, existing chart editing methods based on MLLMs either only support style and layout editing without targeting answer-related areas, or rely on structured data rather than static images. Furthermore, MLLMs often struggle to accurately locate task-related areas, the reasoning process is opaque, and the consistency between visual editing and textual answers is difficult to guarantee, thus limiting their application in question-answering-aware visual editing. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, the present invention provides a method and system for enhancing the saliency of problem-aware visualization images. It can directly edit static visualization images without relying on the underlying data or specifications of the charts, enhance the visual saliency of answer-related areas, ensure the consistency between visual editing and text answers, guide users to quickly locate answers, and at the same time preserve the structural integrity of the original static visualization image.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for enhancing the saliency of problem-aware visualization images, comprising: Based on the acquired static visualization image to be processed and the corresponding analysis question, identify the answer-related region corresponding to the analysis question and generate a target region mask; Editing operations are performed on the static visualization image based on the target region mask to generate at least one preliminary edited image. Each preliminary edited image and the static visualization image are then input into a multimodal large language model to enhance the visual saliency of the target region and generate an optimized edited image. The analysis question, static visualization image, and optimized edited image are input into a multimodal large language model to generate a text answer corresponding to the target region in the edited image; Based on the consistency between the text answer and the optimized edited image, the text answer and the optimized edited image are corrected to obtain the optimal edited image and the final text answer.
[0008] Secondly, the present invention provides a problem-aware visualization image saliency enhancement system, comprising: The target region masking module is configured to: based on the acquired static visualization image to be processed and the corresponding analysis question, identify the answer-related region corresponding to the analysis question and generate a target region mask; The optimization editing module is configured to: perform editing operations on the static visualization image based on the target region mask, generate at least one preliminary edited image, input each preliminary edited image and the static visualization image into a multimodal large language model to enhance the visual saliency of the target region, and generate an optimized edited image; The text answer generation module is configured to: input the analysis question, static visualization image, and optimized edited image into a multimodal large language model to generate a text answer corresponding to the target region in the edited image; The correction module is configured to correct the text answer and the optimized edited image based on their consistency, so as to obtain the optimal edited image and the final text answer.
[0009] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0010] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0011] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0012] The above one or more technical solutions have the following beneficial effects: In this invention, the answer region related to the free-form question in a static visualization image is first identified, and a corresponding target region mask is generated. A multimodal large language model is then used to enhance the saliency of this region, resulting in an edited image where the answer region is significantly highlighted. Based on this, a text answer corresponding to the target region in the edited image is generated, and adjustments are made based on the consistency between the text answer and the edited image, outputting the final optimal edited image. Thus, the final output edited image directly displays the visual region upon which the answer is based, forming a mutually verifying image-text pair with the final text answer. Users can obtain both clear answer text and visual evidence of the answer directly from the image. This invention enables question-and-answer perceptual visual editing of static visualization images without relying on underlying chart data, specification documents, or paired training data. By automatically locating answer-related regions and enhancing their visual saliency, it significantly reduces the user's cognitive load and improves the efficiency and accuracy of chart-based question-and-answer systems.
[0013] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0014] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0015] Figure 1 This is a flowchart illustrating the problem-aware visualization image saliency enhancement method in Embodiment 1 of the present invention. Detailed Implementation
[0016] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0017] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0018] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0019] Example 1 like Figure 1 As shown in the figure, this embodiment discloses a problem-aware visualization image saliency enhancement method, including: Based on the acquired static visualization image to be processed and the corresponding analysis question, identify the answer-related region corresponding to the analysis question and generate a target region mask; Editing operations are performed on the static visualization image based on the target region mask to generate at least one preliminary edited image. Each preliminary edited image and the static visualization image are then input into a multimodal large language model to enhance the visual saliency of the target region and generate an optimized edited image. The analysis question, static visualization image, and optimized edited image are input into a multimodal large language model to generate a text answer corresponding to the target region in the edited image; Based on the consistency between the text answer and the optimized edited image, the text answer and the optimized edited image are corrected to obtain the optimal edited image and the final text answer.
[0020] This embodiment generates a target region mask by identifying answer-related areas and performs editing operations on static visualization images. It also utilizes a multimodal large language model to enhance saliency. This allows it to automatically locate question-related areas in an image for any free-form question and generate an edited image with significantly enhanced answer areas. Users no longer need to manually search and compare data elements in complex visualization images; they can quickly locate the answer simply by observing the edited image, significantly reducing the cognitive load and time cost in the chart-based question-and-answer process.
[0021] Based on the generated optimized edited image, a text answer corresponding to the target region in the edited image is further generated. Corrections are made based on the consistency between the text answer and the edited image. The final optimal edited image directly identifies the visual region related to the answer, forming a mutually verifying image-text pair with the final text answer. Thus, users receive both clear textual answers and visual evidence supporting the answers, solving the problems of existing automated methods that only output textual answers and lack interpretability and visual evidence.
[0022] This embodiment relies solely on the static visualization image itself, requiring no original data, specification files, or paired training data for the charts. Region localization, saliency enhancement, and answer generation are all based on the general visual language understanding capabilities of a multimodal large language model, capable of adapting to various chart types such as bar charts, line charts, pie charts, and scatter plots, as well as analysis problems using any form of natural language. Compared to existing methods that rely on manual rules or limited datasets, this approach offers stronger generalization and practical application value.
[0023] The following is a detailed description of a problem-aware visualization image saliency enhancement method proposed in this embodiment: Step 1: Based on the acquired static visualization image to be processed and the corresponding analysis question, the region locator is used to identify the answer-related region corresponding to the analysis question and generate the target region mask.
[0024] Step 1 specifically includes the following sub-steps: Step 1.1: Perform content parsing on the static visualization image to generate text descriptions and segmentation hints. The specific implementation method is as follows: Use a large vision-language model to perform visual understanding and structured parsing on the input static visualization image.
[0025] Existing large-scale visual models have significant limitations in segmenting charts and images. They exhibit poor adaptability to semantic understanding and target segmentation of visualized charts, generally suffering from insufficient segmentation accuracy, inability to generate accurate and effective image masks, and susceptibility to generating invalid segmentation regions, making it difficult to perform fine-grained segmentation of complex charts and images. To address these shortcomings, this embodiment proposes a targeted improvement scheme: The visualized image is input into a pre-trained large-scale visual-language model. Through multi-level feature extraction and semantic reasoning, the model accurately identifies key information such as the chart type (line chart, bar chart, heatmap, scatter plot, etc.), data dimensions, coordinate axis parameters, legend, and data distribution, and automatically organizes this information into standardized structured text descriptions. Furthermore, different types of charts have significantly different visual compositions, key segmentation targets, and analytical focuses. A single general prompt cannot meet the segmentation needs of all charts, easily leading to mask generation errors and invalid segmentation. To address this, this embodiment pre-designs and includes multiple sets of exclusive and refined segmentation prompts for mainstream chart types such as line charts, bar charts, heatmaps, and scatter plots, based on the common structural features and segmentation task requirements of various chart types, forming a standardized optional prompt word library. Simultaneously, to prevent chaotic matching of multiple prompts and further avoid invalid segmentation, this embodiment sets priority ranking rules for prompts of different chart types within the word library. The model can automatically filter, match, and call the optimal segmentation prompts according to the predetermined priority based on the specific chart type identified and the parsed structured text semantic information, ultimately generating exclusive text segmentation guidance instructions adapted to the current image features, thereby achieving accurate segmentation for multiple chart types. These improvements effectively overcome the shortcomings of traditional methods, such as lack of semantic adaptation and generic instructions. By combining structured text with a priority prompt word matching mechanism based on chart type, the problem of invalid segmentation is avoided at its root, achieving refined and accurate segmentation of chart visual elements. Furthermore, the generated structured text can provide support for subsequent intent understanding, task matching, and natural language interaction, significantly improving the segmentation accuracy and effect of chart images.
[0026] Step 1.2: Use an image segmentation model (such as the SAM model) to segment the static visualization image to obtain binary masks of multiple visual elements. Each binary mask corresponds to an independent visual element in the image.
[0027] Step 1.3: The segmented binary mask is filtered using a multimodal large language model driven by question answering. The specific implementation process is as follows: First, since the segmented binary mask is only the result of image structure segmentation, it cannot autonomously adapt to the user's specific analysis question, resulting in a large number of redundant masks unrelated to text analysis needs. Therefore, a multimodal large language model is needed to complete accurate filtering. While existing multimodal large language models are not ideal for generating accurate masks for charts and images and for non-intrusive editing, they possess excellent cross-modal reasoning capabilities, accurately capturing visual differences before and after image editing, and deeply associating image visual features with text semantic information. Based on this, this embodiment performs region-by-region color replacement preprocessing on the original static visualization image for each independent binary mask obtained from segmentation, and highlights and edits the target regions of each mask to generate multiple color-replaced images corresponding to different segmented regions. Subsequently, the three sets of information—a single color-replaced preprocessed image, the original unedited image, and the user-defined analysis question—are input into the existing multimodal large language model in batches. The multimodal large language model locates the target region corresponding to each segmentation mask by accurately comparing the visual differences between the images before and after editing. Simultaneously, it combines text semantics to complete visual-text joint reasoning, judging the matching correlation between each mask region and the user's analysis question, filtering out all irrelevant redundant masks, and accurately selecting all effective binary masks highly relevant to the analysis question and task answer. Finally, all the selected effective masks are merged to generate a complete and unfragmented final target region mask, which can then be used to complete corresponding image editing and analysis tasks. The calculation formula for the target region mask is as follows:
[0028] in, This represents the set of mask indices related to the analysis problem. Indicates the first A binary mask related to the problem. This represents the target region mask after merging.
[0029] Step 2: Based on the target region mask and the static visualization image, input them into the visualization editor to edit the static visualization image and generate at least one preliminary edited image. Input each preliminary edited image and the static visualization image into the multimodal large language model to enhance the visual saliency of the target region and generate the optimized edited image.
[0030] Step 2 specifically includes the following sub-steps: Step 2.1: Deterministic editing.
[0031] Based on the target region mask, three complementary editing operations are applied to the original static visualization image to generate three preliminary edited images: (1) Color highlighting: Replace the color of visual elements in the target area with a more visually salient color to highlight features; (2) Darken the background: Reduce the brightness or saturation of non-target areas to suppress visual interference; (3) Bounding box annotation: Draw a bounding box around the target area to provide spatial hints.
[0032] Step 2.2: Model optimization and refinement.
[0033] Each pre-edited image and the original static visualization image are input into a multimodal large language model for optimization. The pre-editing task prompts and rule constraint prompts are integrated into a unified text prompt, which is then input into the multimodal large language model. The multimodal large language model combines the text prompt, the original static image, and the pre-edited image to complete the fine optimization of the image.
[0034] Optimization is achieved by combining initial editing prompts with cue word constraints to correct visual artifacts introduced during editing. This process uses the original image as a baseline, ensuring that the edited image retains key information such as axis labels and data values, and avoids introducing visual elements not present in the original static visualization image. The optimized edited image... The calculation formula is as follows:
[0035] in, The optimization function for a multimodal large language model is expressed in the following form:
[0036] in, Indicates the first The image generated by the initial editing operation, Represents the original static visualization image. Constraints are provided for prompts such as coordinate axes, data labels, and structure fidelity; This represents the fusion of text and image features in a multimodal model. Indicates the optimized first... One edited image, These correspond to three editing operations: color highlighting, background darkening, and bounding box annotation. This represents the visual understanding mapping of a multimodal large language model.
[0037] Step 3: Input the analysis question, static visualization image, and optimized edited image into the multimodal large language model to generate the text answer corresponding to the target region in the edited image; The analysis question, the original static visualization image, and the optimized edited image are input into a multimodal large language model. The multimodal large language model uses the analysis question as the reasoning guide, the original static visualization image as the data benchmark, and the highlighted target area in the optimized edited image as the basis. It jointly understands visual features and text semantics, locates the key data and chart relationships corresponding to the question, and generates a text answer corresponding to the target area in the edited image.
[0038] Initial text answer The calculation formula is as follows:
[0039] in, This represents the answer generation function, specifically in the form of:
[0040] in, This indicates the analysis of the problem. Represents the original static visualization image. This indicates the optimized edited image. For image visual feature extraction operations, For multimodal feature fusion operations, Mapping semantic reasoning and answer generation for multimodal large language models. This represents the initial text answer.
[0041] In this embodiment, after generating the initial edited images, each initial edited image is further input into a multimodal large language model along with the original static visualization image for optimization. This ensures that editing operations (such as color highlighting, background darkening, etc.) do not destroy key information in the original chart, such as axis labels, data values, and legends, nor introduce false elements not present in the original image. Therefore, the enhanced image highlights the answer area while fully preserving the authenticity of the original data and the readability of the chart.
[0042] Step 4: Using the cross-modal optimization module, the text answer and the optimized edited image are corrected based on their consistency to obtain the optimal edited image and the final text answer.
[0043] Step 4 specifically includes the following sub-steps: Step 4.1: Optimize the answer.
[0044] By combining the analysis of the question, the original static visualization image, and all optimized edited images, the initial text answer is verified and corrected: First, the data features of the highlighted areas of the optimized edited image are compared with those of the original static visualization image using a multimodal large model to verify whether the initial text answer matches the question and the true values of the chart; then, based on the cross-verification of multiple optimized edited images, deviations and errors in the initial text answer are corrected, and finally, the final text answer is output, which is highly consistent with the question and chart structure.
[0045] The calculation formula is as follows:
[0046] in, This represents the answer optimization function, specifically in the form of:
[0047] in, For multimodal feature fusion, For multimodal large model answer verification and correction mapping, This represents the final text answer. This represents a set of three optimized edited images; This indicates the problem being analyzed; Image feature extraction; This indicates the analysis of the problem. Represents the original static visualization image. Indicates the optimized first... One edited image, These correspond to three editing operations: color highlighting, background darkening, and bounding box annotation.
[0048] Step 4.2: Image optimization.
[0049] The system evaluates the visual quality and structural integrity of each optimized edited image. It uses a multimodal large language model to determine the effectiveness of the editing operation. Combining human visual preferences, color contrast, and other visual characteristics, and based on preset prompts, the system compares the optimized edited image with the original static visualization image to determine whether the editing operation effectively improves the saliency of the target region. At the same time, it verifies whether the edited image retains key information such as coordinate axis labels and data values of the original image, and whether it introduces visual elements that are not present in the original image. Finally, it generates targeted image optimization feedback.
[0050] The calculation formula is as follows:
[0051] in, The image optimization function is represented in the following form:
[0052] in, This indicates the analysis of the problem. Represents the original static visualization image. Indicates the optimized first... One edited image, These correspond to three editing operations: color highlighting, background darkening, and bounding box annotation. Image feature extraction; To ensure structural fidelity, a set of natural language constraint hints, such as coordinate axes, data labels, and no artifacts, is constructed using manually constructed prior rules. To constrain regional saliency, a set of natural language constraint prompts, such as color contrast and visual preference, is constructed using artificial prior rules. For multimodal feature fusion; This is a mapping for evaluating the visual quality and effectiveness of multimodal large models. The image optimization feedback is in text form, used to describe the editing effect and structural fidelity, and to guide the next round of cross-modal iterative optimization.
[0053] Step 4.3: Answer - Image Consistency Optimization.
[0054] Check the consistency between the final text answer and the edited image, determine whether the prominent target areas in the edited image correspond to the information in the text answer, and generate consistency feedback. The calculation formula is as follows:
[0055] in, The answer-image consistency optimization function has the following form:
[0056] in For the final text answer, In order to analyze the problem, Represents the original static visualization image; To optimize the edited image, For image visual feature extraction, For multimodal feature fusion, The answer is a region matching consistency constraint hint word. Perform consistency checks and optimizations on multimodal large models, and output consistency feedback in text form to guide iterative corrections; This indicates consistent feedback.
[0057] Step 4.4: Iterative convergence.
[0058] Based on semantic evaluation (consistency feedback) and saliency evaluation, it is determined whether the edited image and text answer meet the requirements. If they do, the iteration terminates and the optimal edited image and final text answer are output. If they do not meet the requirements, steps 2-4 are re-executed based on the optimization feedback until convergence.
[0059] Among them, significance assessment The calculation formula is as follows:
[0060] in, Represents the original static visualization image. Indicates the optimized first... One edited image, These correspond to three editing operations: color highlighting, background darkening, and bounding box annotation. Represents the target region mask related to the problem; The pre-trained saliency prediction network can be represented by the TempSal saliency detection network, which outputs a saliency probability prediction map. This indicates that the average value of the pixels within the target region mask is taken; A positive value indicates that the visual salience of the target area is higher than that of the original image after editing.
[0061] In this embodiment, S≥0.02 is preset, and consistency feedback is used. When the result is "consistent", it is considered convergent, the iteration is terminated, and the optimal edited image (selected from the three edited images) is output. The image with the highest value and best visual effect, and the final text answer; if the convergence condition is not met, then the image optimization feedback is used. and consistency feedback Repeat steps 2-4 until convergence, with an upper limit of 5 iterations to avoid infinite iteration.
[0062] Among them, the region locator, the visualization editor, and the cross-modal optimization module are all implemented based on a multimodal large language model, and adopt a capability-oriented model allocation strategy to adapt to the functional requirements of each module; the deterministic editing of the visualization editor includes three complementary operations: color highlighting, background darkening, and bounding box annotation; the region locator uses the SAM series image segmentation model to achieve visual element segmentation.
[0063] This embodiment explicitly corrects the initial text answer based on the consistency between the text answer and the optimized edited image after generation. Through cross-modal iterative optimization, it mandates that significantly enhanced regions in the edited image maintain consistency with the information described in the final text answer. This overcomes the shortcomings of existing multimodal large language models in chart editing tasks, such as inaccurate localization, opaque reasoning, and mismatch between visual editing and text answers, significantly improving the reliability and accuracy of question-answer-driven visual editing results.
[0064] To demonstrate the effectiveness of the solution in this embodiment, specific experiments were conducted: This experiment uses two datasets for evaluation: a portion of the ChartQA dataset and corresponding data for 16 chart types covered in the ChartX dataset, to assess the generalization ability of the proposed solution.
[0065] The evaluation metrics used in the experiment included: task saliency (measuring the correlation between the saliency distribution of the edited image and the question-guided attention map), image saliency (measuring the degree of saliency improvement after editing the target region), structural similarity (SSIM, MS-SSIM, measuring the structural consistency between the edited image and the original image), and answer accuracy (using semantic answer similarity (SAS) and natural language inference (NLI) to measure the correctness of the text answer).
[0066] The experiment compared the proposed solution with existing chart editing methods (ChartOptimiser, ChartLlama) and four mainstream multimodal large language models (GPT-4o-Image, Gemini-3-Pro, Qwen-2.5VL, Doubao-SEEDream-5.0). The results are shown in Table 1. Table 1: Comparison of the Invention's Solution with Existing Chart Editing Methods and Four Mainstream Multimodal Large Language Models
[0067] As shown in Table 1, on both datasets, the task saliency and image saliency of the present invention are both good, significantly better than existing comparison methods; the structural similarity remains at 1.0000, ensuring the integrity of the original image structure; the answer accuracy (SAS, NLI) remains at a high level, proving that the editing operation of the present invention can effectively support accurate question-answering reasoning.
[0068] The technical solution of this embodiment can directly perform question-and-answer-aware visual editing on static visualization images without relying on the underlying data of the chart, specification documents, or paired training data. By automatically locating the answer-related area and enhancing its visual salience, it solves the pain point of existing methods that cannot adapt to free queries and require manual location of related areas, greatly reducing the user's cognitive load and improving the efficiency and accuracy of chart question-and-answer. The technical solution of this embodiment introduces a dual mechanism that combines deterministic editing and model optimization. It adopts three complementary editing operations: color highlighting, background darkening, and bounding box annotation. At the same time, it corrects visual artifacts through a multimodal large language model to ensure that the edited image fully retains the structural integrity and data authenticity of the original image, avoiding the problems of strong intrusion and damage to chart structure that exist in traditional editing methods. The technical solution in this embodiment achieves iterative verification and consistency alignment between edited images and text answers through a cross-modal optimization module. Combined with multi-dimensional evaluation indicators such as task saliency and image saliency, it solves the problems of opaque reasoning and mismatch between visual guidance and text answers in existing multimodal large model chart editing, thereby improving the interpretability and reliability of question-and-answer reasoning.
[0069] Example 2 The purpose of this embodiment is to provide a problem-aware visualization image saliency enhancement system, including: The target region masking module is configured to: based on the acquired static visualization image to be processed and the corresponding analysis question, identify the answer-related region corresponding to the analysis question and generate a target region mask; The optimization editing module is configured to: perform editing operations on the static visualization image based on the target region mask, generate at least one preliminary edited image, input each preliminary edited image and the static visualization image into a multimodal large language model to enhance the visual saliency of the target region, and generate an optimized edited image; The text answer generation module is configured to: input the analysis question, static visualization image, and optimized edited image into a multimodal large language model to generate a text answer corresponding to the target region in the edited image; The correction module is configured to correct the text answer and the optimized edited image based on their consistency, so as to obtain the optimal edited image and the final text answer.
[0070] In further embodiments, the following is also provided: An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When executed by the processor, the computer instructions perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0071] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0072] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0073] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.
[0074] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0075] A computer program product includes a computer program that, when executed by a processor, implements the method described in Embodiment 1.
[0076] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0077] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0078] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0079] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0080] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for enhancing the saliency of problem-aware visualization images, characterized in that, include: Based on the acquired static visualization image to be processed and the corresponding analysis question, identify the answer-related region corresponding to the analysis question and generate a target region mask; Specifically: Perform content parsing on static visualization images to generate text descriptions and segmentation hints for the images; An image segmentation model is used to segment a static visualization image to obtain binary masks for multiple visual elements; each binary mask corresponds to an independent visual element in the static visualization image. The segmented binary mask is filtered using a multimodal large language model, retaining the valid binary mask related to the answer, and the target region mask is obtained based on the valid binary mask. Editing operations are performed on the static visualization image based on the target region mask to generate at least one preliminary edited image. Each preliminary edited image and the static visualization image are then input into a multimodal large language model to enhance the visual saliency of the target region and generate an optimized edited image. The analysis question, static visualization image, and optimized edited image are input into a multimodal large language model to generate a text answer corresponding to the target region in the edited image; The text answer and the optimized edited image are corrected to obtain the optimal edited image and the final text answer; specifically: By combining the analysis of the problem, the original static visualization image, and all optimized edited images, the initial text answer is verified and corrected to obtain the verified and corrected text answer; The visual quality and structural integrity of each optimized edited image are evaluated. The effectiveness of the editing operation is judged by calling a multimodal large language model. Combined with visual characteristics, the optimized edited image is compared with the original static visualization image to determine whether the editing operation effectively improves the saliency of the target region. At the same time, it is verified whether the optimized edited image retains the key information of the original static visualization image and does not introduce visual elements that are not present in the original static visualization image, and generates targeted image optimization feedback. Check the consistency between the corrected text answer and the optimized edited image, determine whether the prominent target area in the optimized edited image corresponds to the information in the corrected text answer, and generate consistency feedback; Based on saliency assessment and consistency feedback, it is determined whether the optimized edited image and the verified corrected text answer meet the requirements. If they do, the optimal edited image and the final text answer are obtained; if they do not, the optimization process is re-executed until convergence.
2. The problem-aware visualization image saliency enhancement method as described in claim 1, characterized in that, Based on the target region mask and the static visualization image, three editing operations are performed on the static visualization image: color highlighting, background darkening, and bounding box annotation, to generate the corresponding preliminary edited image.
3. The problem-aware visualization image saliency enhancement method as described in claim 1, characterized in that, The analysis question, static visualization image, and optimized edited image are input into a multimodal large language model to generate a text answer corresponding to the target region in the edited image. Specifically, the multimodal large language model uses the analysis question as the reasoning guide, the original static visualization image as the data benchmark, and the highlighted target region in the optimized edited image to jointly understand visual features and text semantics, locate the key data and chart relationships corresponding to the question, and generate a text answer corresponding to the target region in the edited image.
4. The problem-aware visualization image saliency enhancement method as described in claim 1, characterized in that, Significance assessment The calculation formula is as follows: in, This represents a pre-trained saliency prediction network. This indicates that the average value of the pixels within the target region mask is taken; Represents the target region mask related to the problem; A positive value indicates that the visual saliency of the target region is higher than that of the original image after editing; The original static visualization image; Represents the optimized first... One edited image, These correspond to three editing operations: color highlighting, background darkening, and bounding box annotation.
5. A problem-aware visualization image saliency enhancement system, characterized in that, include: The target region masking module is configured to: based on the acquired static visualization image to be processed and the corresponding analysis question, identify the answer-related region corresponding to the analysis question and generate a target region mask; Specifically: Perform content parsing on static visualization images to generate text descriptions and segmentation hints for the images; An image segmentation model is used to segment a static visualization image to obtain binary masks for multiple visual elements; each binary mask corresponds to an independent visual element in the static visualization image. The segmented binary mask is filtered using a multimodal large language model, retaining the valid binary mask related to the answer, and the target region mask is obtained based on the valid binary mask. The optimization editing module is configured to: perform editing operations on the static visualization image based on the target region mask, generate at least one preliminary edited image, input each preliminary edited image and the static visualization image into a multimodal large language model to enhance the visual saliency of the target region, and generate an optimized edited image; The text answer generation module is configured to: input the analysis question, static visualization image, and optimized edited image into a multimodal large language model, and generate a text answer corresponding to the target region in the edited image; The correction module is configured to correct the text answer and the optimized edited image to obtain the optimal edited image and the final text answer; specifically: By combining the analysis of the problem, the original static visualization image, and all optimized edited images, the initial text answer is verified and corrected to obtain the verified and corrected text answer; The visual quality and structural integrity of each optimized edited image are evaluated. The effectiveness of the editing operation is judged by calling a multimodal large language model. Combined with visual characteristics, the optimized edited image is compared with the original static visualization image to determine whether the editing operation effectively improves the saliency of the target region. At the same time, it is verified whether the optimized edited image retains the key information of the original static visualization image and does not introduce visual elements that are not present in the original static visualization image, and generates targeted image optimization feedback. Check the consistency between the corrected text answer and the optimized edited image, determine whether the prominent target area in the optimized edited image corresponds to the information in the corrected text answer, and generate consistency feedback; Based on saliency assessment and consistency feedback, it is determined whether the optimized edited image and the verified corrected text answer meet the requirements. If they do, the optimal edited image and the final text answer are obtained; if they do not, the optimization process is re-executed until convergence.
6. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-4.
8. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-4.
Citation Information
Patent Citations
Visual question and answer data enhancement method and device, equipment and storage medium
CN119128118A
Bidirectional Transform-based multi-modal video description generation method
CN120544093A