A scatter plot overlapping removal method based on templates and quadtrees

The quadtree-based scatter plot method effectively addresses overplotting by optimizing data layout and density adjustment, ensuring uniform distribution and readability with interactive customization.

CN119579766BActive Publication Date: 2025-07-15COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410953446.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2025-07-15
Estimated Expiration
2044-07-16

AI Technical Summary

Technical Problem

With the exponential growth of data scale, scatter plots face problems of over-drawing, resulting in visual interference and data pattern identification difficulties. The existing technology methods have data losses and biases when dealing with large-scale data sets.

Method used

Using a template-based and quadtree method, the data space is adaptively divided through the quadtree, the data points are mapped to the pre-designed non-overlapping template, and the node radius is adjusted through grayscale values to optimize the data point layout.

Benefits of technology

It realizes uniform distribution and efficient data representation, improves the readability and information expression capabilities of scatter plots, and enhances data analysis efficiency and user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119579766B_ABST
    Figure CN119579766B_ABST
Patent Text Reader

Abstract

The present invention provides a method for removing overlap in scatter plots based on templates and quadtrees, belonging to the technical field of data processing, including: 1. According to the number of data points in the grid divided by the quadtree, and determining the similarity between the scatter plot and the template as well as the similar template based on the number of data points in the grid, and optimizing the layout of the data points in the grid through the similar template; 2. Calculating the data density of each grid, mapping the calculation result of the data density to the gray space, and adjusting the node radius between the data points in the grid according to the distribution of the gray values of the grid in the gray space and the area of the grid; 3. Constructing a density distribution adjustment tool based on gray scale, and when the user conducts data interaction, using the density distribution adjustment tool to refine the layout of the grid of the scatter plot at different zoom levels. The present invention solves the problems of visual chaos caused by over-drawing of large-scale scatter plots and the inability to effectively observe data patterns, and can also enhance the visibility of the pattern through the interactive enhancement mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method for removing overlap in scatter plots based on templates and quadtrees. Background Art

[0002] With the exponential growth of data scale, scatter plots face an important problem - over plotting. When a large number of data points are superimposed in a limited viewport, visual interference occurs, making it difficult to identify data patterns, thus weakening the effectiveness of scatter plots.

[0003] To solve the above technical problems, existing technical solutions often reduce overlap by operating on the data points themselves and capture local features of the data by adjusting the sampling density. Data sampling methods can significantly alleviate over plotting, but they cause a certain degree of data loss and data bias due to discarding some data points and pursuing different sampling characteristics respectively.

[0004] To address the above technical problems, the present invention proposes a method for removing overlap in scatter plots based on templates and quadtrees. The quadtree is used to efficiently divide the data space to ensure that the number of data points in each grid is effectively controlled, and then through the template layout technology, the data points are mapped onto a pre-designed non-overlapping template. Summary of the Invention

[0005] The object of the present invention is to provide a method for removing overlap in scatter plots based on templates and quadtrees.

[0006] To solve the above technical problems, the present invention provides a method for removing overlap in scatter plots based on templates and quadtrees, which is characterized in that it specifically includes:

[0007] Construct a quadtree, set the maximum number of data points in each grid, and adaptively divide the data space according to the maximum number of data points and the quadtree. Divide the data points of the scatter plot into multiple grids according to the division result of the data space;

[0008] Determine the similarity between the scatter plot and the template and the similar template according to the number of data points in the grid, and optimize the layout of the data points in the grid through the similar template;

[0009] Calculate the data density of each grid, map the calculation result of the data density to the grayscale space, and adjust the node radius between the data points in the grid according to the distribution of the grayscale values of the grid in the grayscale space and the area of the grid;

[0010] Construct a density distribution adjustment tool based on grayscale. When the user performs data interaction, use the density distribution adjustment tool to refine the layout of the grids of the scatter plot at different zoom levels.

[0011] The present invention solves the problems of visual chaos caused by over-drawing of large-scale scatter plots and the inability to effectively observe data patterns, and can also enhance the visibility of patterns through an interactive enhancement mode. The beneficial effects of the present invention are specifically described as follows:

[0012] Adaptive spatial partitioning. Adaptive spatial partitioning is achieved through a quadtree, which is a method of recursively partitioning a two-dimensional space. In this process, each node represents a specific spatial region. If a node contains more data points than a preset maximum value, this region will be further subdivided into four sub-regions, each represented by a new node. This partitioning strategy allows the algorithm to adaptively adjust the size and shape of each region according to the actual distribution of data points, thus achieving the following purposes: First, to achieve uniform distribution: by controlling the number of data points within the grid, over-aggregation or sparsity of data points is avoided, and a more uniform data distribution is achieved; Second, to be more efficient and accurate: Adaptive partitioning improves space utilization, reduces unnecessary calculations, and also improves the accuracy of representing data distribution. In addition, by setting parameters m, min_level, and max_level, the algorithm can automatically adjust the partitioning strategy according to the specific data distribution characteristics, further improving the efficiency and effect of spatial partitioning.

[0013] Data display optimized through templates. The template layout technique is a key method for optimizing data display through predefined templates. These templates contain nodes with optimized distributions, and each template is applicable to a certain number of data points. When mapping data points to the template, the node distribution within the template ensures that each data point can be effectively laid out in the visible space without overlap, effectively avoiding the overlap of data points and improving the readability of the scatter plot. At the same time, the node distribution in the template is carefully designed, making the visual representation of the data both beautiful and practical. Figure 4 The change curves of the K-nearest neighbor index, density preservation index, displacement minimization, and shape preservation index are given respectively. It can be seen from the curves that as the partitioning level deepens, the displacement and K-nearest neighbor index of the scatter plot increase significantly. By the third level, it has basically approached the original layout of the data set. For example, for the classic large-scale data set Hathi_trust_library, at level zero, because there are many nodes within the grid after partitioning, the local structure is damaged, and the K-nearest neighbor is only 0.22. However, after two layout refinements, it will change to 0.75, and after three times, it will rise to above 0.9, basically maintaining the original local semantics. The same is true for the minimum displacement index. At level zero, it is about 0.11 for both large data sets, but it will quickly decay to about 0.02 at the third level, basically maintaining the original position.

[0014] Dynamic node radius adjustment. The adjustment of the dynamic node radius is based on the density calculation of data points within each grid. After mapping the density to grayscale values, the radius of the nodes is dynamically adjusted according to the grayscale values. This method allows the scatter plot to more realistically reflect the distribution density of the data, especially providing a clearer visual distinction between different density regions, enabling users to intuitively perceive the density distribution of data points through the node size, improving the data analysis efficiency, while strengthening the visual contrast between high-density and low-density regions and enhancing the information expression ability of the scatter plot.

[0015] Interactive layout optimization. By providing an interactive user interface, users can adjust the layout parameters of the scatter plot according to their needs, such as node size, color or other visual attributes, and use the grayscale adjustment tool to refine the layout at different zoom levels. Users can customize the scatter plot according to specific analysis needs and personal preferences, enhancing the user experience. It is also possible to adjust the layout at different zoom levels, enabling the scatter plot to maintain the best visual quality and information transmission effect from different viewing angles.

[0016] A further technical solution lies in adaptively partitioning the data space according to the number of the maximum data points and the quadtree, specifically including:

[0017] Determining the maximum division level and the minimum division level respectively according to the minimum side length of the grid and the maximum value of the allowable displacement of the data points;

[0018] With the maximum division level, the minimum division level and the number of the maximum data points as constraint conditions, all data points in the scatter plot are inserted into the quadtree one by one until the minimum division level is reached, and each data point is assigned to the corresponding leaf node or internal node in the tree according to its coordinate position;

[0019] Realize the adaptive partitioning of the data of the scatter plot into the data space according to the coordinate positions of different data points.

[0020] A further technical solution lies in the method for determining the similarity between the scatter plot and different templates:

[0021] Determine the deviation amount of the number of grids between the scatter plot and the template according to the grid distribution of the scatter plot and the template, and determine the deviation amount of the grid parameters between the scatter plot and the template in combination with the deviation amount of the number of grids in different grid area intervals;

[0022] Determine the deviation amount of the number of grids and the deviation amount of the grid area in different division regions through the grid distribution of the scatter plot and the template, and determine the deviation amount of the grid distribution in different division regions in combination with the deviation situation of the grid distribution in different division regions;

[0023] Determine the similarity between the scatter plot and the template based on the grid distribution deviation and grid parameter deviation within different divided regions.

[0024] Furthermore, the technical solution lies in that the similarity value between the scatter plot and the template ranges from 0 to 1. When the similarity between the template and the scatter plot meets the requirements, the template is determined as a similar template.

[0025] Furthermore, the technical solution lies in that the layout of different grids of the scatter plot is refined at different zoom levels, specifically including:

[0026] Statistically calculate the cumulative distribution function after the gray values of all grids are counted, and recalculate the gray values according to the function;

[0027] Screen the low-density grids according to the recalculated gray values, and flatten the left end of the mapping curve of histogram equalization to 1 in the low-density grids, so that the low-density area is adjusted to a gray level of 1, improving the visibility of the low-density area.

[0028] Other features and advantages will be described in the following specification, and partly will become obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification and the drawings.

[0029] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Brief Description of the Drawings

[0030] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other features and advantages of the present invention will become more obvious.

[0031] Figure 1 It is a flowchart of a scatter plot overlapping removal method based on a template and a quadtree.

[0032] Figure 2 It is a flowchart of a method for determining the similarity between a scatter plot and different templates.

[0033] Figure 3 It is a schematic diagram of a non-overlapping scatter plot layout algorithm based on a quadtree and a template.

[0034] Figure 4 It is a schematic diagram of the change curves of the K-nearest neighbor index, density preservation index, displacement minimization and shape preservation index.

[0035] Figure 5 It is a schematic diagram of the calculation pipeline from density to node radius.

[0036] Figure 6 It is a schematic diagram of dynamic layout and density distribution adjustment. DETAILED DESCRIPTION

[0037] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that the present invention will be comprehensive and complete and fully convey the concepts of the example embodiments to those skilled in the art. The same reference numerals in the figures represent the same or similar structures, and thus their detailed description will be omitted.

[0038] The terms "a", "an", "the", and "said" are used to indicate the presence of one or more elements / components / etc.; the terms "comprising" and "having" are used to express an open-ended inclusive meaning and mean that additional elements / components / etc. may be present in addition to the listed elements / components / etc.

[0039] In the rich picture of data visualization, scatter plots occupy a pivotal position for their ability to intuitively display the distribution characteristics of data points. It provides strong visual support for data analysis by marking the positions of data points on a two-dimensional plane, revealing the correlation, clustering trends and outliers between data. However, with the exponential growth of data size, scatter plots face an important problem - overplotting. When a large number of data points are superimposed in a limited field of view, visual interference occurs, making it difficult to identify data patterns, thereby weakening the effectiveness of scatter plots. For example, in financial market analysis, overplotting may cause investors to miss important price trends; in social network analysis, overlapping data points may obscure key social connections; in medical research, overlap may make it difficult to discover the correlation between symptoms. Therefore, scatter plot de-overlapping can improve the quality of data visualization.

[0040] Existing scatter plot de-overlapping methods mainly focus on three aspects: data space methods, visual space methods and hybrid space methods.

[0041] The data space method reduces overlap by operating on the data points themselves. For example, the non-uniform sampling method proposed by Bertini and Santucci captures local features of the data by adjusting the sampling density. The data sampling method can significantly alleviate over-drawing, but it causes a certain degree of data loss and data bias due to discarding some data points and pursuing different sampling characteristics respectively. The visual space method reduces visual interference by changing the visual attributes of data points, such as size, shape, and transparency. However, these methods often have difficulty dealing with large-scale data sets and may cause color confusion and loss of individual information of data points when processing multi-class scatter plots. The hybrid space method refers to a method in which there are core operations in both the data space and the visual space, and reduces over-drawing by changing the visual unit from a node to a region. The existing hybrid space methods are good at quickly presenting the macroscopic distribution of data, but a large amount of details will be lost because individual data points are discarded.

[0042] In view of the limitations of the existing methods, the present invention proposes a scatter plot de-overlapping method based on a template and a quadtree. The quadtree is used to efficiently divide the data space to ensure that the number of data points in each grid is effectively controlled. Then, through the template layout technology, the data points are mapped onto a pre-designed template without overlap. In addition, considering that the densities of different grids may vary, this method introduces a density-aware adjustment mechanism based on grayscale, and optimizes the user's perception of the density distribution by dynamically adjusting the node radius. And the dynamic layout optimization technology is used to allow the user to gradually reduce the node displacement when zooming in on the view, so as to approximate the true state of the original data distribution. Embodiment 1

[0043] To solve the above technical problems, as Figure 1 shown, the present invention provides a scatter plot de-overlapping method based on a template and a quadtree, which is characterized in that it specifically includes:

[0044] Construct a quadtree, set the maximum number of data points in each grid, and adaptively divide the data space according to the maximum number of data points and the quadtree, and divide the data points of the scatter plot into multiple grids according to the division result of the data space;

[0045] Specifically, adaptively dividing the data space according to the maximum number of data points and the quadtree specifically includes:

[0046] Determine the maximum division level and the minimum division level respectively according to the minimum side length of the grid and the maximum value of the allowable displacement of the data points;

[0047] With the maximum division level, minimum division level, and the number of maximum data points as constraints, all data points in the scatter plot are inserted into the quadtree one by one until the minimum division level is reached. Each data point is assigned to the corresponding leaf node or internal node in the tree according to its coordinate position.

[0048] The data of the scatter plot is adaptively divided into the data space according to the coordinate positions of different data points.

[0049] Construction of the quadtree and spatial division

[0050] Create a root node of a PR quadtree (Point Region Quadtree), which represents the entire two-dimensional space region where the data points are located. The PR quadtree is a tree-shaped data structure where each node can represent a spatial region and can be further divided into four child nodes. All data points are inserted into the PR quadtree one by one. Each data point is assigned to the corresponding leaf node or internal node in the tree according to its coordinate position. This process involves searching for a suitable node in the tree from top to bottom and storing the data point in that node.

[0051] The input of the quadtree construction process includes the original 2D data point set , where N is the number of data points in the dataset, represents the set of all points in the two-dimensional space with positive real coordinates

[0052] The output of the algorithm is the root node of the quadtree, including seven attributes: points are the data points within the node, nodes are the subtrees of the node, which are also quadtree node classes, x, y, width, height, and level respectively represent the x and y coordinates of the upper left corner of the grid represented by the current node, the width and height of the grid, and the level where the current node is located.

[0053] m, min_l eve l and max_l eve These three parameters, m, min_level, and max_level, are straightforward for constructing the quadtree but not for users. Especially, m is the minimum number of data points included in each square in the quadtree. It is difficult to select a reasonable m without understanding the data distribution characteristics. Therefore, a method for automatically setting m is proposed, and based on the actual observation experience, min_level and max_level are transformed into intuitive variables bound to pixels.

[0054] min_l eve l : min_l evel is deeply bound to the max_offset that measures the degree of displacement distortion. This variable is intuitive and of great concern to users. Therefore, users can set min_l by considering the maximum allowable displacement under a given canvas size. eve l , That is .

[0055] Among them ceil denotes ceil function

[0056] max_l eve l and min_l eve l is similar. max_l can be indirectly set from the perspective of actual observation. eve l. Given that at the macroscopic scale, the size of the observable grid has an upper limit, max_l can be set by setting the minimum side length min_length of the grid. eve l .

[0057] m: The setting of parameter m is closely related to the data distribution, especially the distribution in high-density regions. If m is set too small, most regions will be divided too deeply; if it is too large, the high-density regions will not be further divided, resulting in the "curse of high dynamic range". Therefore, a suitable parameter setting can prevent most regions from being divided to max_l eve l, while finding extremely high-density regions so that they can be further divided, that is, being able to separate medium- and low-density regions from extremely high-density regions. According to the method of automatically setting m, a suitable m value is given based on the original data point set and max_l eve l. Specifically, the data space is divided into equally spaced grids according to the grid side length corresponding to max_l eve l, and the number of data points in each grid is counted and sorted. Finally, the number corresponding to the fixed quantile value (here the empirical value 0.995 is used) is taken as m. When the data distribution is relatively uniform, that is, there are only very few extremely high-density regions, the problem of a small m result will occur in this algorithm. Therefore, finally, the maximum value of the above-obtained m and 1 / 4 of the maximum number of nodes is taken as the final m. Such a setting meets the need to disperse extremely high-density regions, while strictly controlling the tree depth and ensuring the calculation efficiency.

[0058] Compared with equally spaced grids, the spatial division based on quadtrees realizes adaptive division for density and detail, and the characterization resolution is more controllable under the parameters min_l eve l and max_l eve l. Just increase these two l evel can obtain more accurate resolution in different areas. Although the equidistant grid can also adjust the depiction resolution by controlling the size of the grid, it can only be adjusted globally and uniformly, and each adjustment requires global re-division. Compared with the equidistant grid, the quadtree grid has a significantly smaller number of grids under the condition of having the same maximum resolution. Furthermore, adaptive grid divisions of different sizes and without obvious rules are not easy to cause perceptual misunderstandings to people. Since all the grid sides of the equidistant grid are the same length, if the grid side length is large, the user will mistakenly perceive the non-existent "grid boundary", and the operation of increasing the resolution will increase its spatial complexity and reduce its time efficiency by a square factor.

[0059] Determine the similarity between the scatter plot and the template and the similar template according to the number of data points in the grid, and optimize the layout of the data points of the grid by using the similar template;

[0060] Specifically, as shown in FIG2 , the method for determining the similarity between the scatter plot and different templates is:

[0061] Determine the deviation of the number of grids between the scatter plot and the template according to the grid distribution of the scatter plot and the template, and determine the grid parameter deviation of the scatter plot and the template in combination with the grid number deviation in different grid area intervals; determine the deviation of the number of grids and the deviation of grid areas in different divided areas according to the grid distribution of the scatter plot and the template, and determine the grid distribution deviation in different divided areas in combination with the deviation of the grid distribution in different divided areas; determine the similarity between the scatter plot and the template based on the grid distribution deviation and grid parameter deviation in different divided areas.

[0062] Furthermore, the similarity between the scatter plot and the template ranges from 0 to 1, wherein when the similarity between the template and the scatter plot meets the requirement, the template is determined to be a similar template.

[0063] In another embodiment, the method for determining the similarity between the scatter plot and different templates is:

[0064] Determine the deviation of the number of grids of the scatter plot and the template according to the distribution of the grids of the scatter plot and the template, and judge whether the deviation of the number of grids of the scatter plot and the template meets the requirement, if so, proceed to the next step, if not, determine that the template does not belong to the similar template;

[0065] Determine whether the template is a similar template based on the mesh number deviation in different mesh area intervals, if yes, proceed to the next step, if no, determine that the template is not a similar template;

[0066] Determine the grid parameter deviation between the scatter plot and the template based on the deviation of the number of grids between the scatter plot and the template and the deviation of the number of grids in different grid area intervals, and determine whether the grid parameter deviation between the scatter plot and the template meets the requirements. If so, proceed to the next step; if not, determine that the template does not belong to the similar template;

[0067] Determine the deviation of the number of grids and the deviation of the grid area in different divided regions based on the grid distribution of the scatter plot and the template, and determine the grid distribution deviation in different divided regions by combining the deviation of the grid distribution in different divided regions. Determine whether the number of divided regions where the grid distribution deviation does not meet the requirements meets the requirements. If so, proceed to the next step; if not, determine that the template does not belong to the similar template;

[0068] Determine the similarity between the scatter plot and the template based on the grid distribution deviation and the grid parameter deviation in different divided regions.

[0069] Specifically, optimize the layout of data points in different grids through the similar template, which specifically includes:

[0070] Set an upper limit for the number of template nodes. In each 1×1 square area, use a preset sampling method to generate a template containing a sampling points, and divide the number of data points in different grids according to the template;

[0071] Count the number of data points in different divided grids to generate a sampling point combination, then randomly map the data point set in the scatter plot and the sampling point set to obtain a mapping result, and optimize the layout of data points in different grids based on the mapping result.

[0072] It should be noted that as shown in Figure 3, it is a schematic diagram of a non-overlapping scatter plot layout algorithm based on a quadtree and a template. The sample points of the template need to be evenly distributed in the template space. Evenly means that the spacing between sample points is almost the same. On the one hand, even distribution is beneficial to perceiving the regional density. It is a necessary condition for perceiving the regional density based on the area ratio of colored pixels that the colored pixels are evenly distributed in space; on the other hand, since changing the radius to the same extent does not damage the uniform property, the density reflected by the template can be adjusted by controlling the radius of the node and the adjustable range of the node radius can be maximized.

[0073] Among many sampling methods, blue noise sampling has the characteristics of random and uniform node distribution. The mainstream blue noise sampling methods include Lolyd's method, Poisson disk sampling, etc. Considering performance and special requirements for uniformity, Lolyd's d method is selected. After sampling with this method, the node layout is regular and beautiful, and a proper distance is maintained from the boundary, so that node overlap will not occur at the tile joints. In contrast, the randomness brought by Poisson disk sampling not only destroys the above two advantages but also may make users mistakenly think that the current reflects the real data distribution. In the specific implementation, a template node upper limit (set to 200 according to experience) is specified for each template node value a from 1 to , a template containing a sampling points is generated within a 1×1 square area using Lolyd's method. Then, the number of data points in the divided grids is counted to match the corresponding templates, and the data point set and the sampling point set are randomly mapped.

[0074] Calculate the data density of each grid, map the calculation result of the data density to the gray space, and adjust the node radius between the data points of the grid according to the distribution of the gray values of the grid in the gray space and the area of the grid;

[0075] Specifically, the specific method for adjusting the node radius between the data points of the grid is: calculate the node radius of the grid when the template gray value is 1, that is, the template node radius;

[0076] Determine the gray density of the grid based on the distribution of the gray scale of the grid, and construct a preset formula based on the gray density, the template node radius, and the ratio of the actual grid to the template side length. Determine the node radius between the data points of the grid through the preset formula;

[0077] Adjust the node radius between the data points of the grid based on the node radius. Further, determining the gray density of the grid based on the distribution of the gray scale of the grid specifically includes: determining the average gray value of different regions of the grid based on the distribution of the gray scale of the grid, and determining the gray density of the grid according to the average gray value of different regions.

[0078] Specifically, the specific method for adjusting the node radius between the data points of the grid is:

[0079] Determine the average gray value of the grid based on the distribution of the gray values of the grid in the gray space, and determine the recommended node adjustment radius of the grid in combination with the area of the grid;

[0080] Determine the grayscale attention area based on the distribution of grayscale values of the grid in the grayscale space, and determine the compensation amount of the recommended node adjustment radius according to the number of the grayscale attention areas, the average grayscale values of different grayscale attention areas, and the distances between different grayscale attention areas;

[0081] Determine the node radius between data points of different grids according to the compensation amount of the node adjustment radius and the recommended node adjustment radius of the grid, and use the node radius to adjust the node radius between data points of different grids.

[0082] In another embodiment, the specific method for adjusting the node radius between data points of the grid is as follows:

[0083] S31 Determine the average grayscale value of the grid based on the distribution of grayscale values of the grid in the grayscale space, and determine the recommended node adjustment radius of the grid in combination with the area of the grid;

[0084] S32 Determine whether there is a grayscale attention area by the distribution of grayscale values of the grid in the grayscale space. If so, go to the next step; if not, adjust the node radius between data points of the grid based on the recommended node adjustment radius of the grid;

[0085] S33 Determine whether the number of the grayscale attention areas is less than the preset area number. If so, go to the next step; if not, go to step S35;

[0086] S34 Determine whether there is a grayscale attention area whose grayscale value is not within the preset grayscale interval. If so, go to the next step; if not, adjust the node radius between data points of the grid based on the recommended node adjustment radius of the grid;

[0087] S35 Determine the compensation amount of the recommended node adjustment radius according to the number of the grayscale attention areas, the average grayscale values of different grayscale attention areas, and the distances between different grayscale attention areas, determine the node radius between data points of different grids according to the compensation amount of the node adjustment radius and the recommended node adjustment radius of the grid, and use the node radius to adjust the node radius between data points of different grids.

[0088] For a grid containing n data points, only need to find a template with a potential of n and then randomly map the set of data points and the set of template sample points to complete the layout of the data points within the grid. However, the grids obtained based on the quadtree are of different sizes, and grids with the same number of data points may also have different sizes. At this time, it is necessary to introduce amplitude modulation halftoning by adjusting the node radius to smooth out the density perception differences caused by the grid size. Figure 5 is a schematic diagram of the calculation pipeline from density to node radius. As shown in Figure 5, in actual calculation, first calculate the data density of the grid. The formula is the number of data points in the grid divided by the grid area, that is where sizeG is the side length of the grid, density which can be abbreviated as d, and its value range is , in a specific embodiment, the data density density can be linearly mapped to the grayscale grey , that is , this mapping ensures that the global maximum grayscale is 1. Through the above method, the grayscale of all grids has been successfully calculated grey , and the size range of the grayscale value is , where is preset, and the default value can be 0.2. The grayscale grey can be abbreviated as g . At this time, the problem to be solved becomes how to represent this value in the scatter plot. The present invention adjusts the regional perception density by setting the node radius radius . Here, the node radius radius can be abbreviated as r , and its value range is This method is based on a conjecture: the coverage ratio of colored pixels

[0089] is proportional to the perceived density intensity. In the extreme case, a region completely filled with pixels is considered to have a regional density of 1 in perception. Since the square of the node radius within the grid is proportional to the grayscale, the node radius when the template grayscale is 1 can be calculated, that is, the template node radius , so the calculation formula for the inner radius of the template can be obtained . In the formula, is the ratio of the side length of the actual grid to the side length of the template. Templates with different potentials have different The mapping from grayscale to node radius is realized through the square root function , that is, is realized. It should be noted that the grayscale corresponding to the upper limit of the template radius is 1. In theory, it should fill the entire square template space, but to avoid node overlap, the upper limit can only reach the size where the sample points are tangent to each other. In specific calculation, the average value of the distance between each sample point and its nearest sample point is used as this value.

[0090] Construct a density distribution adjustment tool based on grayscale. When the user conducts data interaction, the density distribution adjustment tool is used to refine the layout of the grid of the scatter plot at different zoom levels.

[0091] Specifically, refining the layout of different grids of the scatter plot at different zoom levels specifically includes: calculating the cumulative distribution function after statistically analyzing the grayscale values of all grids, and recalculating the grayscale values according to the function; screening low-density grids based on the recalculated grayscale values, and flattening the left end of the mapping curve of histogram equalization to 1 in the low-density grids, so that the low-density area is adjusted to a grayscale of 1, enhancing the visibility of the low-density area.

[0092] To enhance the user's perception of the data density distribution, a method based on grayscale value adjustment is proposed. This method allows the user to adjust the visual display of data points through flexible interaction methods to more accurately reflect the true distribution of the data.

[0093] Specifically, construct a mapping from the original grayscale to the adjusted grayscale. The domain of the mapping is the grayscale range of all grids within the current drawing frame , and the range is , and then recalculate using the mapped grayscale . In the unadjusted state, it can be regarded as a proportional function with a slope of 1. Therefore, the original grayscale reflects the density distribution without modification. When there is a need to enhance visibility, by default, histogram equalization is used to perform HDR (High Dynamic Range) for the scatter plot, that is, calculating the cumulative distribution function after statistically analyzing the grayscale values of all grids and recalculating the grayscale values according to the function to achieve the effect of "spreading out" the grayscale of all grids. Such processing is very effective when dealing with the situation where the grayscale values of the grids are relatively concentrated. The grayscale of the grids originally confined to a very small range is evenly extended to nearly the [0,1] range, increasing the visible range, but the effect on the low-density area is not so satisfactory. Because histogram equalization is a monotonic mapping, the grayscale values in the low-density area are still relatively low after mapping, so the low-density area is still not easily observable, and even outliers may be ignored. To solve this problem, additional visual enhancement for the low-density area is required to increase the target grayscale, that is, flattening the left end of the mapping curve of histogram equalization to 1, so that the low-density area is adjusted to a grayscale of 1, greatly increasing the visibility of outliers and even the low-density area. The present invention uses parameters to control the grayscale flattening point at the left end. The final grayscale curve is a piecewise function. When , , when , , where g is the original grayscale, For the previous obtained grayscale value. A notable point is that to ensure that the grayscale value of the outliers within the current bounding box is 1, thus it is based on .

[0094] Based on the above embodiments, the following technical effects are achieved:

[0095] Adaptive spatial partitioning. Adaptive spatial partitioning is implemented through a quadtree, which is a method of recursively partitioning a two-dimensional space. In this process, each node represents a specific spatial region. If a node contains more data points than a preset maximum value, this region will be further divided into four sub-regions, each represented by a new node. This partitioning strategy allows the algorithm to adaptively adjust the size and shape of each region according to the actual distribution of data points, thereby achieving the following purposes: First, to achieve uniform distribution: by controlling the number of data points within the grid, over-aggregation or sparsity of data points is avoided, and a more uniform data distribution is achieved; Second, to be more efficient and accurate: Adaptive partitioning improves space utilization, reduces unnecessary calculations, and also improves the accuracy of representing data distribution. In addition, by setting parameters m, min_level, and max__level, the algorithm can automatically adjust the partitioning strategy according to the specific data distribution characteristics, further improving the efficiency and effect of spatial partitioning

[0096] Data display optimized by templates. The template layout technology is a key method for optimizing data display through predefined templates. These templates contain nodes with optimized distributions, and each template is applicable to a certain number of data points. When mapping data points to the templates, the node distributions within the templates ensure that each data point can be effectively laid out in the visual space without overlap, effectively avoiding the overlap of data points, improving the readability of the scatter plot. At the same time, the node distributions in the templates are carefully designed to make the visual representation of the data both beautiful and practical. Figure 4 shows the change curves of the K-nearest neighbor index, density preservation index, displacement minimization, and shape preservation index respectively. It can be seen from the curves that as the division level deepens, the displacement of the scatter plot and the K-nearest neighbor index increase significantly. By the third level, it is basically close to the original layout of the dataset. For example, for the classic large-scale dataset Hathi_trust_library, at level zero, because there are many nodes in the grid after division, the local structure will be damaged, and the K-nearest neighbor is only 0.22. But after two layout refinements, it will change to 0.75, and after three times, it will rise above 0.9, basically maintaining the original local semantics. The same is true for the minimum displacement index. At level zero, it is about 0.11 for both large datasets, but it will quickly decay to about 0.02 at the third level, basically maintaining the original position.

[0097] Dynamic node radius adjustment. The adjustment of the dynamic node radius is based on the density calculation of data points within each grid. After mapping the density to the gray value, the radius of the node is dynamically adjusted according to the gray value. This method allows the scatter plot to more realistically reflect the distribution density of the data, especially providing a clearer visual distinction between different density regions, enabling users to intuitively perceive the density distribution of data points through the node size, improving the data analysis efficiency, and at the same time strengthening the visual contrast between high-density and low-density regions, enhancing the information expression ability of the scatter plot.

[0098] Interactive layout optimization. By providing an interactive user interface, users can adjust the layout parameters of the scatter plot according to their own needs, such as node size, color, or other visual attributes, and use the gray value adjustment tool to refine the layout at different zoom levels. Users can customize the scatter plot according to specific analysis needs and personal preferences, improving the user experience. It is also possible to adjust the layout at different zoom levels to make the scatter plot maintain the best visual quality and information transmission effect from different viewing angles. Figure 6 shows a schematic diagram of dynamic layout and density distribution adjustment.

[0099] In the description of this specification, the description of terms such as "one embodiment" and "one preferred embodiment" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or instance. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0100] The above are only the preferred embodiments of the embodiments of the present invention and are not used to limit the embodiments of the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the embodiments of the present invention.

Claims

1. A scatter plot de - overlapping method based on templates and quad - trees, characterized in that Specifically, it includes: Construct a quadtree, set the maximum number of data points in each grid, adaptively divide the data space according to the maximum number of data points and the quadtree, and divide the data points of the scatter plot into multiple grids according to the division result of the data space; Determine the similarity between the scatter plot and a preset template, as well as the similar template, and optimize the layout of the data points in the grid through the similar template; stipulate the node upper limit max of the preset template a , for each template node value a from 1 to max m , use Lloyd's method to generate a preset template containing a sampling points within a 1×1 square area The determination of the similarity between the scatter plot and the preset template and the similar template is specifically as follows: Determine the deviation amount of the number of grids between the scatter plot and the preset template according to the grid distribution of the scatter plot and the preset template, and determine the grid parameter deviation amount between the scatter plot and the preset template by combining the grid number deviation amounts in different grid area intervals; Determine the deviation amount of the number of grids and the deviation amount of the grid area in different division regions through the grid distribution of the scatter plot and the preset template, and determine the grid distribution deviation amount in different division regions by combining the deviation conditions of the grid distribution in different division regions; Determine the similarity between the scatter plot and the preset template based on the grid parameter deviation amount between the scatter plot and the preset template and the grid distribution deviation amount in different division regions; When the similarity between the preset template and the scatter plot meets the requirements, determine the preset template as the similar template; The layout optimization of the data points in the grid by the similar template is specifically as follows: Randomly map the data point set in the scatter plot and the sampling point set of the similar template to obtain a mapping result; Calculate the data density of each grid, map the calculation result of the data density to the gray space, and adjust the node radius between the data points of the grid according to the gray value distribution of the grid in the gray space and the area of the grid; Construct a density distribution adjustment tool based on gray scale. When the user performs data interaction, use the density distribution adjustment tool to refine the layout of the grids of the scatter plot at different zoom levels.

2. The method for removing overlap in a scatter plot based on a template and a quadtree according to claim 1, characterized in that Adaptive division of the data space according to the maximum number of data points and the quadtree specifically includes: Determine the maximum division level and the minimum division level according to the minimum side length of the grid and the maximum allowable displacement of the data points respectively; With the maximum division level, the minimum division level and the maximum number of data points as constraints, insert all the data points in the scatter plot into the quadtree one by one until the minimum division level is reached, and each data point is assigned to the corresponding leaf node or internal node in the tree according to its coordinate position; Realize the adaptive division of the data of the scatter plot into the data space according to the coordinate positions of different data points.

3. A scatter plot overlap removal method based on templates and quadtrees as described in claim 1, characterized in that, The value range of the similarity between the scatter plot and the preset template is between 0 and 1.

4. A scatter plot overlapping removal method based on a template and a quadtree as claimed in claim 1, wherein, The determination of the gray density of the grid based on the gray distribution of the grid specifically includes: Determine the average gray value of different regions of the grid based on the gray distribution of the grid, and determine the gray density of the grid according to the average gray values of different regions.

5. A method for removing overlap in scatter plots based on templates and quadtrees, as described in claim 1, wherein The refinement of the layout of different grids of the scatter plot at different zoom levels specifically includes: After counting the grayscale values of all grids, calculate the cumulative distribution function and recalculate the grayscale values according to the function; Screen the low-density grids according to the recalculated grayscale values, and flatten the left end of the mapping curve of histogram equalization to 1 in the low-density grids, so that the low-density area is adjusted to a grayscale of 1, improving the visibility of the low-density area.

Citation Information

Patent Citations

  • Method for displaying three-dimensional scatter diagram in browser and system thereof

    CN106971417A

  • Evaluation method for effectiveness of scatter diagram over-drawing solution

    CN116432059A