Visualization method supporting comparison of text-image pairs
By modeling the text-image pair data into a change map and generating the final view, the problem of difficulty in efficiently comparing text-image pairs in the prior art is solved, and the intuitive display of the impact on text modification and the improvement of image generation quality are achieved.
Patent Information
- Application Number
- PCT/CN2023/136318
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-05
AI Technical Summary
The prior art lacks a method to support efficient comparison of text-image pairs, which makes it difficult for users to understand the impact of text prompt word modification on image generation results and affects image generation quality.
By modeling the text-image data as a change graph, the node position is calculated using image feature projection, and the difference between text is represented as edges, combining the edge weight metric model and edge bundling algorithm, the final view is generated to visually display the image-text difference.
It realizes efficient comparison of text-image data, helping users understand the impact of text modification on image generation results, so as to facilitate targeted modification and adjustment of images generated by the image generation model and improve image generation quality.
Smart Images

Figure CN2023136318_05062025_PF_FP_ABST
Abstract
Description
A visualization method to support comparison of text-image pairs Technical Field
[0001] The present invention belongs to the field of visualization, and in particular relates to a visualization method supporting comparison of text-image pairs. Background Art
[0002] Image collections with textual captions are a common type of data. Existing visualization methods utilize image and text features to compute projections, aiming to help users understand the semantic distribution of large image collections. Recently, text-to-image generation models (text-graph models) have emerged that support image generation based on textual cues. Generating high-quality images using text-graph models typically requires repeated modification of textual cues. However, due to the lack of efficient methods for comparing text-image pairs, users struggle to understand the impact of cues on generated images. This results in users being unable to effectively modify and adjust generated images, impacting image quality.
[0003] Summary of the Invention
[0004] To address the shortcomings of the prior art, the present invention aims to provide a visualization method that supports the comparison of text-image pairs. Given text-image pair data, the method models it as a variation graph, with images as nodes and differences between texts as edges between corresponding image nodes. By screening and clustering the edges and nodes in the variation graph, the positional relationships between nodes and edges in the final visualization graph are obtained, and a final visualization graph is generated based on this. In the final visualization graph, the positions of image nodes are calculated based on image feature projections, visually displaying image-text differences. That is, node positions indicate image differences, while edges directly reveal text differences. Furthermore, through point-edge connections, the image and text are tightly coupled, reflecting the association between text and image differences. By calculating inter-text differences at the word level, an edge weight measurement model and an edge bundling algorithm are proposed to reveal the associations between differences between multiple image-text pairs and the varying impacts of different words on image differences.
[0005] To achieve the above objectives, the present invention adopts a technical solution: a visualization method for supporting comparison of text-image pairs, the method comprising the following steps:
[0006] S1. Modeling the differences in the text-image pair data to be analyzed as an original change graph, wherein the images are represented as nodes in the change graph, the text differences are represented as edges between the corresponding nodes, and the image differences are represented as node position relationships;
[0007] S2. Preprocess the text-image pair data to be analyzed, and process the edges in the change graph according to the preprocessing results to obtain the positional relationship between the edges and the graph in the final visible graph;
[0008] S3. Layout and render the image and edges according to the preprocessing results to generate the final visible graph.
[0009] Furthermore, step S2 includes preprocessing the text and image separately, obtaining text difference edges based on the text preprocessing results, aggregating the edges in the change graph in combination with the image preprocessing results, and then calculating the edge weights and filtering and merging the edges based on the weights.
[0010] Further, step S2 includes the following sub-steps:
[0011] S21. Preprocess the text in the text-image pair data to be analyzed, use the obtained text differences as text difference edges, and add differences to the edges in the change graph according to the text difference edges;
[0012] S22, pre-processing the images in the text-image pair data to be analyzed, and clustering the image nodes in the two-dimensional plane;
[0013] S23, aggregating the edges in the change graph based on the image preprocessing results;
[0014] S24. Calculate the weight of the aggregated edge, and further filter and merge the edges in the change graph according to the weight of the aggregated edge.
[0015] Furthermore, step S21 includes calculating the edit distance between any two texts at the word granularity to obtain a text distance matrix; screening out text pairs whose distance is not greater than the distance threshold according to a preset or user-entered distance threshold; selecting several comparable text pairs; for the selected text pairs, comparing them pairwise to obtain text differences, and using the obtained text differences as text difference edges.
[0016] Furthermore, the difference in step S21 includes two attributes: word and operation.
[0017] Furthermore, step S22 includes encoding the image into a high-dimensional vector using the CLIP model, projecting the high-dimensional image vector onto a two-dimensional plane using the t-SNE algorithm, and clustering the image nodes within the two-dimensional plane using a hierarchical agglomerative clustering method.
[0018] Furthermore, step S23 includes aggregating edges having the same preset attribute values for the change graph based on the text preprocessing and image preprocessing results, where the preset attributes include words, operations, source clusters, and target clusters.
[0019] Further, step S24 includes the following sub-steps:
[0020] S241, uniformly assign weights to the original edges that have not been aggregated, so that the sum of the original edge weights between every two image nodes is 1;
[0021] S242. Calculate the aggregated edge weight based on the original edge weights, where the aggregated edge weight is the sum of the original edge weights.
[0022] S243. According to a preset or user-entered weight threshold, aggregated edges with weights greater than the weight threshold are retained; and source clusters, target clusters, and aggregated edges with equal weights are merged.
[0023] Furthermore, in step S3, the image position in the final visible graph is calculated according to the projection result in the image preprocessing, the word anchor point position is determined based on the image position, the edge in the final visible graph is determined according to the image position and the word anchor point position, and the final visible graph is generated accordingly.
[0024] Further, step S3 includes the following sub-steps:
[0025] S31, based on the t-SNE projection result in the image preprocessing, scaling the image to the screen size to generate the image layout in the final visual graph;
[0026] S32, performing aggregation edge word anchor layout based on image positions, where the word anchor coordinates are the average of the coordinates of all source images and target images in the aggregation edge;
[0027] S33. Use force-directed algorithm to fine-tune the layout of image and aggregated edge word anchors;
[0028] S34. Draw edges in the final visible graph, where each edge is a Bezier curve starting from the source image node, passing through the word anchor point, and reaching the target image node.
[0029] The beneficial technical effect of the present invention is that: a visualization method for supporting comparison of text-image pairs disclosed in the present invention is adopted to support efficient comparison of text-image pairs, that is, through change graph modeling and final visual graph layout, image differences are represented as node position relationships, text differences are directly represented as edges, and the correlation between text differences and image differences is revealed through edge weight measurement, which can help users efficiently compare text-image pair data. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] FIG1 is a flow chart of a visualization method for supporting comparison of text-image pairs disclosed in Embodiment 1 of the present invention;
[0031] FIG2 is a flowchart of a method for preprocessing text-image pair data to be analyzed in a step of a visualization method for supporting comparison of text-image pairs disclosed in Embodiment 1 of the present invention;
[0032] FIG3 is an effect diagram of generating a final change graph using a visualization method supporting comparison of text-image pairs disclosed in the first embodiment of the present invention. DETAILED DESCRIPTION
[0033] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0034] Example 1
[0035] As shown in FIG1 , an embodiment of the present invention provides a visualization method for supporting comparison of text-image pairs, the method comprising the following steps:
[0036] S1. Model the data differences of the text-image pair to be analyzed as a change graph.
[0037] In the change graph, the nodes are images, and the edges between corresponding image nodes are text differences.
[0038] For any two text-image pairs, calculating the difference between the two texts yields several distinct words, corresponding to several edges between the two image nodes. Edges between different node pairs may also correspond to the same word. However, not all edges between any two text-image pairs need to be presented in the entire text-image dataset. For example, when the differences between the texts are significant, it is not necessary to compare the differences between each specific word. Alternatively, when there are multiple distinct words between the texts, some words may not significantly contribute to the distribution differences between the corresponding images. Therefore, the edges displayed in the final graph need to be weighted and filtered.
[0039] S2. Preprocess the text-image pair data to be analyzed, and process the edges in the change graph according to the preprocessing results to obtain the edges and graph in the final visible graph.
[0040] As shown in Figure 2, for the text-image pair data to be analyzed, the text and image are first preprocessed separately. The text difference edges are obtained based on the text preprocessing results, and the edges are aggregated based on the image preprocessing results. Then, the edge weights are calculated and filtered and merged according to the weights to reduce visual clutter.
[0041] Step S2 includes the following sub-steps:
[0042] S21 . Preprocess the text in the text-image pair data to be analyzed, use the obtained text differences as text difference edges, and adjust the edges in the change graph according to the text difference edges.
[0043] The edit distance between any two texts is calculated at the word level to generate a text distance matrix. Based on a preset or user-entered distance threshold, text pairs with a distance no greater than the threshold are filtered out. This allows for the selection of relatively similar and comparable text pairs.
[0044] For the selected text pairs, the text differences are obtained by pairwise comparison and the obtained text differences are used as text difference edges.
[0045] Each difference consists of two attributes: a word and an operation. Operations are divided into three categories: add, delete, and move. Using the edit distance algorithm, we can identify the add and delete operations. Then, by matching the same words in the added and deleted words, we can identify the move operation.
[0046] The edges in the change graph are adjusted according to the text difference edges, so that the edges of the adjusted change graph have terms and operations.
[0047] S22. Preprocess the images in the text-image pair data to be analyzed.
[0048] The CLIP model is used to encode the image into a high-dimensional vector, the t-SNE algorithm is used to project the high-dimensional image vector onto a two-dimensional plane, and the hierarchical agglomerative clustering method is used to cluster the image nodes in the two-dimensional plane.
[0049] S23. Aggregate the edges in the change graph based on the image preprocessing results.
[0050] The edges in the change graph have the following four attributes: word, operation, source cluster, and target cluster. Based on the text preprocessing and image preprocessing results, edges with the same values for these four attributes are clustered.
[0051] S24. Calculate the weight of the aggregated edge, and further filter and merge the edges according to the weight of the aggregated edge.
[0052] The edge weight measures the degree of correlation between text differences and image distribution differences. That is, if the source cluster of the edge representing the addition of a word does not contain the word, but the target cluster does, then the text difference is closely related to the image difference. Conversely, if the word has similar distribution density in the source and target clusters, then there is no significant correlation between the text difference and the image difference.
[0053] Step S24 includes the following sub-steps:
[0054] S241 , assigning original edge weights: At this time, consider the original edges that have not been aggregated, set the sum of the edge weights between every two image nodes to 1, and evenly distribute them to each edge between the nodes.
[0055] S242. Calculate the aggregated edge weight based on the original edge weight: The weight of the aggregated edge is the sum of the original edge weights. For each original edge between two image nodes, the weight is redistributed according to the aggregated edge weight to which it belongs, so that the sum of the edge weights between the two image nodes is still 1. The aggregated edge weight is updated accordingly.
[0056] S243. Further filter and merge the edges according to the weights of the aggregated edges.
[0057] Based on the preset or user-entered weight threshold, edges with weights greater than the weight threshold are retained. Edges with equal weights in the source cluster, target cluster, and weights are merged.
[0058] S3. Layout and render the image and edges according to the preprocessing results to obtain the final visible graph.
[0059] The image positions in the final visible graph are calculated based on the projection results from image preprocessing. For edges in the final visible graph, the positions of the difference word anchors are first calculated. Each edge starts from the source image, passes through the word anchor, and reaches the target image. A force-directed algorithm is then used to fine-tune the layout to reduce occlusions.
[0060] Step S3 includes the following sub-steps:
[0061] S31. Layout the images in the final visual graph based on the projection results from image preprocessing: Based on the t-SNE projection results from image preprocessing, scale the images proportionally to the screen size. The images are scaled proportionally to a preset or user-specified size, and overlapping images are simplified to small rectangles.
[0062] S32. Perform clustered edge anchor point layout: Each clustered edge corresponds to a word anchor point, bundling the clustered edges together. The word anchor point coordinates are the average of the coordinates of all source and target images in the clustered edge.
[0063] S33. Fine-tune the layout of image and aggregated edge word anchors: Use a force-directed algorithm to disperse overlapping image and word anchors with a small repulsive force.
[0064] S34. Draw edges in the final visible graph: each edge is a Bezier curve starting from the source image node, passing through the word anchor point, and reaching the target image node, wherein the word anchor point control point is parallel to the line connecting the source image node and the target image node.
[0065] The final visual graph is obtained by laying out and rendering images and edges. In the final visual graph, image differences are represented as node position relationships, and text differences are directly represented as edges. The edge weight measurement reveals the correlation between text differences and image differences, which can help users efficiently compare text-image data.
[0066] It can be seen from the above embodiments that the visualization method disclosed in the present invention that supports the comparison of text-image pairs can construct visualizations for multiple groups of text-image pair data, and display text differences, image differences and their connections. It can help users efficiently compare text-image pair data, facilitate users to make targeted modifications and adjustments to images generated using image generation models, and improve image generation quality.
[0067] The method described in the present invention is not limited to the embodiments described in the specific implementation manner. Those skilled in the art may derive other implementation manners based on the technical solution of the present invention, which also fall within the scope of the technical innovation of the present invention.
Claims
1. A visualization method for supporting the comparison of text-image pairs, the method comprises the following steps: S1. Model the data differences of the text-image pairs to be analyzed as an original change graph, where the images are represented as nodes in the change graph, the text differences are represented as edges between the corresponding nodes, and the image differences are represented as the positional relationships between the nodes; S2. Preprocess the data of the text-image pairs to be analyzed, and process the edges in the change graph according to the preprocessing results to obtain the edges in the final visual graph and the positional relationships between the graphs; S3. Layout and render the images and edges according to the preprocessing results to generate the final visual graph.
2. A visualization method for supporting the comparison of text-image pairs according to claim 1, characterized in that: In step S2, it includes preprocessing the text and the images respectively, obtaining text difference edges according to the text preprocessing results, aggregating the edges in the change graph in combination with the image preprocessing results, and then calculating the edge weights and screening and merging the edges according to the weights.
3. A visualization method for supporting the comparison of text-image pairs according to claim 2, characterized in that, step S2 includes the following sub-steps: S21. Preprocess the text in the data of the text-image pairs to be analyzed, take the obtained text differences as text difference edges, and add differences to the edges in the change graph according to the text difference edges; S22. Preprocess the images in the data of the text-image pairs to be analyzed, and cluster the image nodes in the two-dimensional plane; S23. Aggregate the edges in the change graph in combination with the image preprocessing results; S24. Calculate the aggregated edge weights, and further screen and merge the edges in the change graph according to the weights of the aggregated edges.
4. A visualization method for supporting the comparison of text-image pairs according to claim 3, characterized in that: Step S21 includes calculating the edit distance between any two texts at the word granularity to obtain a text distance matrix; according to a preset or user-input distance threshold, screening out text pairs with a distance not greater than the distance threshold; to select several comparable text pairs; for the selected text pairs, compare them pairwise to obtain text differences, and take the obtained text differences as text difference edges.
5. A visualization method for supporting the comparison of text-image pairs according to claim 3, characterized in that: The differences in step S21 include two attributes: words and operations.
6. A visualization method for supporting the comparison of text-image pairs according to claim 3, characterized in that: Step S22 includes encoding the images into high-dimensional vectors using the CLIP model, projecting the high-dimensional image vectors onto a two-dimensional plane using the t-SNE algorithm, and clustering the image nodes in the two-dimensional plane using the hierarchical agglomerative clustering method.
7. A visualization method for supporting the comparison of text-image pairs according to claim 5, characterized in that: Step S23 includes, for the change graph, aggregating the edges with the same preset attribute values based on the text preprocessing and image preprocessing results, and the preset attributes include words, operations, source clusters, and target clusters.
8. A visualization method for supporting the comparison of text-image pairs according to claim 3, characterized in that, step S24 includes the following sub-steps: S241. Uniformly assign weights to the original edges that have not been aggregated, and make the total weight of the original edges between every two image nodes equal to 1; S242. Calculate the weights of the aggregated edges according to the weights of the original edges, and the weight of the aggregated edge is the sum of the weights of the original edges; S243. According to the preset or user-input weight threshold, retain the aggregated edges with weights greater than the weight threshold; merge the aggregated edges with equal source clusters, target clusters, and weights.
9. A visualization method for supporting comparison of text image pairs as claimed in claim 1, characterized in that: In step S3, calculate the image positions in the final visual graph according to the projection results in the image preprocessing, determine the word anchor positions based on the image positions, determine the edges in the final visual graph according to the image positions and the word anchor positions, and generate the final visual graph accordingly.
10. A visualization method for supporting comparison of text image pairs as claimed in claim 1, characterized in that, step S3 includes the following sub-steps: S31. Based on the t-SNE projection results in the image preprocessing, scale the image proportionally to the screen size to generate the image layout in the final visual graph; S32. Perform the layout of the word anchors of the aggregated edges based on the image positions, and the word anchor coordinates are the average values of the coordinates of all the source images and target images in the aggregated edges; S33. Use the force-directed algorithm to finely adjust the layout of the images and the word anchors of the aggregated edges; S34. Draw the edges in the final visual graph, and each edge is a Bezier curve starting from the source image node, passing through the word anchor, and reaching the target image node.
Citation Information
Patent Citations
Correspondence probability map driven visualization
CN107111881A
Dimension reduction method for carrying out visual comparison on multiple pieces of high-dimensional data
CN113537281A
Image difference analysis method and device
CN115375653A
Processing system and method for generating image based on user text cue word
CN116680425A
Automatic document image extraction and comparison
US20120082372A1