Visual disturbance degree evaluation method and related device
By acquiring saliency maps and gaze heatmaps of scene images and combining them with correlation index data, the problem of inaccurate visual interference assessment in existing technologies has been solved, and a more accurate visual interference assessment has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN UNIV
- Filing Date
- 2026-01-23
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, visual interference assessment mainly relies on subjective evaluation methods, which leads to inaccurate results and fails to accurately reflect the user's true visual response in complex scenarios.
By acquiring saliency maps and eye-tracking data of scene images, processing the images using a pre-defined saliency model and generating gaze heatmaps, and combining several correlation index data, including linear correlation coefficient, standardized scan path saliency, area under the curve, target similarity, and transportation cost, the degree of visual interference is evaluated.
It provides a more accurate assessment of visual interference, which matches the user's real visual reaction in complex scenarios, thus improving the accuracy of the assessment.
Smart Images

Figure CN121564531B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method and related apparatus for evaluating visual interference. Background Technology
[0002] In complex architectural spaces such as high-speed rail stations, museums, and commercial complexes, as well as in information-dense scenarios like web pages and terminal interfaces, users need to quickly extract key information from a mixture of text and images within a limited time. If the layout is poorly designed, non-critical information areas will excessively attract attention, creating visual clutter and reducing navigation efficiency and readability. Therefore, before the official release of a webpage or terminal interface, or before construction begins on buildings like high-speed rail stations, a visual clutter assessment of the layout design should be conducted, and adjustments made based on the assessment results.
[0003] Currently, the assessment of visual interference mainly relies on subjective evaluation methods such as questionnaires and interviews. The results are greatly affected by the subjective feelings of the subjects, resulting in inaccurate determination of the degree of visual interference. Summary of the Invention
[0004] In view of this, this application provides a method and related apparatus for evaluating visual interference, which determines the gaze heatmap by using the saliency map of the scene image and eye-tracking data, and then obtains a more accurate visual interference level of the scene image based on the concentration of salient regions in the saliency map and several correlation index data.
[0005] This application provides a method for evaluating visual interference, applied to electronic devices, the method comprising the following steps:
[0006] Acquire the scene image to be evaluated and the eye movement data of the subject during the observation of the scene image;
[0007] The scene image is processed using a preset saliency model to obtain a saliency map of the scene image, and a gaze heatmap is obtained using the eye-tracking data;
[0008] Determine several correlation index data between the saliency map and the gaze heatmap;
[0009] The visual interference degree of the scene image is determined based on the concentration of salient regions within the saliency map and the data of the several correlation indicators.
[0010] In one possible embodiment, the plurality of correlation index data includes linear correlation coefficient, standardized scan path saliency, area under the curve, target similarity, and transportation cost; determining the visual interference degree of the scene image based on the concentration of salient regions within the saliency map and the plurality of correlation index data includes:
[0011] The correspondence between several preset scene categories and several weighted reassemblies is queried to obtain the target weighted reassembly corresponding to the scene category to which the scene image belongs. The target weighted reassembly includes the concentration degree, the linear correlation coefficient, the significance of the standardized scan path, the area under the curve, the target similarity, and the weighted weight of the transportation cost.
[0012] Based on the target weighting, the concentration degree, the linear correlation coefficient, the significance of the standardized scanning path, the area under the curve, the target similarity, and the transportation cost are weighted and fused to obtain the visual interference degree.
[0013] In one possible embodiment, before querying the correspondence between several preset scene categories and several weighted reorganizations, the method further includes:
[0014] Obtain several sets of sample images under various preset scene categories. Each set of sample images includes a sample saliency map and a sample gaze heatmap.
[0015] For each preset scene category, perform the following steps:
[0016] Obtain the reference concentration degree of the salient region within the salient map of each sample image group, the reference linear correlation coefficient between the sample salient map and the sample gaze heatmap, the reference normalized scan path saliency, the reference area under the curve, the reference target similarity, and the reference transportation cost, and obtain the index dataset under the preset scene category; and,
[0017] Principal component analysis is performed on the index dataset to obtain the weighted reorganizations corresponding to the preset scene categories.
[0018] In one possible embodiment, the indicator dataset includes a matching factor, a discriminant factor, and a distribution concentration factor for each group of sample images. The matching factor includes the reference linear correlation coefficient and the reference target similarity. The discriminant factor includes the reference standardized scan path significance and the reference area under the curve. The distribution concentration factor includes the reference transportation cost and the reference concentration degree.
[0019] The step of performing principal component analysis on the index dataset to obtain the weighted reorganizations corresponding to the preset scene categories includes:
[0020] Based on the index dataset, an interference factor matrix is determined for each group of sample images. The interference factor matrix includes the factor values of the matching degree factor, the discrimination factor, and the distribution concentration factor.
[0021] Clustering is performed on each of the interference factor matrices to obtain the visual interference pseudo-labels for each group of sample images;
[0022] Based on the interference factor matrix of each group of sample images and the visual interference pseudo-label, the weighted weights of the matching factor, the discrimination factor and the distribution concentration factor under the preset scene category are determined.
[0023] The weighted reorganization corresponding to the preset scene category is determined based on the weighted weights of the matching factor, the distinguishing factor, and the distribution concentration factor.
[0024] In one possible embodiment, the step of weighting and fusing the concentration degree, the linear correlation coefficient, the significance of the standardized scanning path, the area under the curve, the target similarity, and the transportation cost according to the target weighting and recombining to obtain the visual interference degree includes:
[0025] The concentration level and each of the correlation indicators are normalized respectively.
[0026] The visual interference degree is obtained by weighting and fusing the normalized linear correlation coefficient, the standardized scanning path significance, the area under the curve and the target similarity, the concentration degree after polarity conversion, and the transportation cost according to the target weighting.
[0027] In one possible embodiment, determining several correlation index data between the saliency map and the gaze heatmap includes:
[0028] The saliency map and the gaze heatmap are flattened to obtain a first vector corresponding to the saliency map and a second vector corresponding to the gaze heatmap. The linear correlation coefficient is obtained based on the sum of the covariances between the first and second vectors.
[0029] The saliency map is standardized to obtain a standardized saliency map, and the standardized scan path saliency is obtained based on the pixel value of the corresponding pixel in the standardized saliency map for each gaze point in the gaze heatmap; and,
[0030] Based on the correspondence between the distribution of salient regions in the salientity map and the distribution of each fixation point in the fixation heatmap, the area under the curve is determined; and,
[0031] Obtain the pixel-level similarity between the saliency map and the gaze heatmap as the target similarity; and,
[0032] Obtain the first probability distribution matrix of the saliency map and the second probability distribution matrix of the gaze heatmap, and obtain the transportation cost based on the distance matrix between the first probability distribution matrix and the second probability distribution matrix.
[0033] In one possible embodiment, obtaining a gaze heatmap using the eye-tracking data includes:
[0034] The eye-tracking data is cleaned.
[0035] The gaze heatmap is obtained based on the frequency of occurrence of several gaze points contained in the cleaned eye movement data.
[0036] In one possible embodiment, the eye-tracking data includes the coordinates of several fixation points, eye movement states, and velocities. The data cleaning of the eye-tracking data includes:
[0037] Based on the coordinates, eye movement state, and velocity of each fixation point, valid fixation points and abnormal fixation points are determined.
[0038] The coordinates, eye movement state, and velocity of abnormal fixation points are removed from the eye movement data.
[0039] This application provides a visual interference assessment device for use in electronic devices, the visual interference assessment device comprising:
[0040] The data acquisition unit is used to acquire the scene image to be evaluated and the eye movement data of the subject during the observation of the scene image;
[0041] The data processing unit is used to process the scene image using a preset saliency model to obtain a saliency map of the scene image, and to obtain a gaze heatmap using the eye-tracking data.
[0042] A correlation index acquisition unit is used to determine several correlation index data between the saliency map and the gaze heatmap;
[0043] The interference evaluation unit is used to determine the visual interference of the scene image based on the concentration of salient regions in the saliency map and the data of several correlation indicators.
[0044] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps in any of the above-described visual interference evaluation methods.
[0045] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the visual interference evaluation methods described above.
[0046] This application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the above-described visual interference evaluation methods.
[0047] In the visual interference assessment method and related apparatus provided in this application, in the process of determining the visual interference of the scene image to be evaluated, not only is the concentration of salient regions in the saliency map of the scene image obtained, but also several correlation index data between the saliency map and the gaze heatmap corresponding to the eye movement data of the observer during the observation of the scene image are obtained. Then, based on the concentration of salient regions in the saliency map and the several correlation index data, the visual interference of the scene image is obtained. Compared with the visual interference determined by means of questionnaires or interviews, this solution not only considers the saliency distribution of the scene image itself, but also introduces real gaze behavior for calibration, so that the determined visual interference is more accurate, that is, the determined visual interference is more consistent with the user's real visual reaction in complex scenes. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 This is one of the flowcharts illustrating a visual interference assessment method provided in this application embodiment;
[0050] Figure 2 This is a schematic diagram of the saliency map provided in the embodiments of this application;
[0051] Figure 3 This is a schematic diagram of the gaze heatmap provided in the embodiments of this application;
[0052] Figure 4 This is a second schematic flowchart of a visual interference assessment method provided in the embodiments of this application;
[0053] Figure 5 This is the third flowchart illustrating a visual interference assessment method provided in this application embodiment;
[0054] Figure 6This is one of the functional unit block diagrams of a visual interference evaluation device provided in the embodiments of this application;
[0055] Figure 7 This is the second functional unit block diagram of a visual interference evaluation device provided in the embodiments of this application;
[0056] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0057] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0058] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0059] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0060] See also Figure 1 , Figure 1 This is one of the flowcharts illustrating a visual interference assessment method provided in this application embodiment. The visual interference assessment method is applied to electronic devices. The method includes the following steps:
[0061] S101, Acquire the scene image to be evaluated and the eye movement data of the subject during the observation of the scene image.
[0062] Scene images can be images taken of a physical scene, or design drawings of the scene, web pages, terminal interfaces, etc.
[0063] The subject can be a human being. Eye movement data during the subject's observation of the scene image can be collected by an eye tracker worn by the subject during this observation. This eye movement data can be obtained by the electronic device from the original eye movement file (e.g., a CSV file) received by the eye tracker when the eye tracker is connected to the electronic device. For example, it can include relevant data segments of the fixation point extracted from the original eye movement file, and invalid headers and explanatory information can be removed from the original eye movement file.
[0064] S102, the scene image is processed using a preset saliency model to obtain a saliency map of the scene image, and a gaze heatmap is obtained using eye-tracking data.
[0065] The pre-defined saliency model can be a pre-trained neural network model. The input to the saliency model is a scene image, and the output is a saliency map of the scene image. For example, the saliency map output by the saliency model can be as follows: Figure 2 As shown, Figure 2 In a significance graph, the closer the color is to red, the higher the significance; the closer the color is to blue, the lower the significance.
[0066] One way to obtain a gaze heatmap using eye-tracking data is to first obtain multiple gaze points and their coordinates using eye-tracking data, and then obtain the gaze heatmap based on the frequency of occurrence of each gaze point. The gaze heatmap obtained by determining the gaze heatmap from eye-tracking data collected during the subject's observation of scene images can be as follows: Figure 3 As shown, Figure 3 The background is dark blue, and the dark blue areas represent areas that are not noticed or are rarely observed. The focal point is... Figure 3 The greater the difference between the color of the bright spot and the dark blue, the higher the frequency of the bright spot's appearance. Figure 3 In the middle, the most frequently occurring focus is on Figure 3 The bottom right side is a warm orange color.
[0067] S103, determine several correlation index data between the saliency map and the gaze heatmap.
[0068] Several correlation metrics may include at least one of the following: linear correlation coefficient (CC), standardized scan path significance (NSS), area under the curve (AUC), target similarity (SIM), and transportation cost (EMD).
[0069] The linear correlation coefficient can be the Pearson linear correlation coefficient, which measures the degree of linear correlation between pixel values between the saliency map and the gaze heatmap. The closer the value is to 1, the stronger the linear correlation, and the closer it is to 0, the less obvious the linear correlation.
[0070] Normalized scan path saliency is used to characterize the degree of matching between salient regions in the saliency map and fixation points in the fixation heatmap. The higher the value of normalized scan path saliency, the higher the degree of matching.
[0071] The area under the curve is used to characterize the ability of a saliency model to distinguish between salient and non-salient regions in a scene image, that is, the accuracy of distinguishing between salient and non-salient regions in the saliency map.
[0072] Target similarity can be cosine similarity; the higher the target similarity, the more similar the saliency map and the fixation heatmap are.
[0073] Transportation cost is used to characterize the difference between the distribution of salient regions in the salientity map and the distribution of fixation points in the fixation heatmap. The smaller the transportation cost, the more similar the distribution of the salientity map and the fixation heatmap.
[0074] S104. Based on the concentration of salient regions within the saliency map and several correlation index data, determine the visual interference level of the scene image.
[0075] Among them, the concentration level and several correlation index data can be weighted and fused, and then the visual interference degree of the scene image can be obtained based on the weighted fusion result.
[0076] In the above scheme, in determining the visual interference of the scene image to be evaluated, not only is the concentration of salient regions in the saliency map of the scene image obtained, but also several correlation index data between the saliency map and the gaze heatmap corresponding to the eye movement data of the observer during the observation of the scene image are obtained. Then, based on the concentration of salient regions in the saliency map and several correlation index data, the visual interference of the scene image is obtained. Compared with the visual interference determined by means of questionnaires or interviews, this scheme not only considers the saliency distribution of the scene image itself, but also introduces real gaze behavior for calibration, so that the determined visual interference is more accurate, that is, the determined visual interference is more consistent with the user's real visual reaction in complex scenes.
[0077] In one possible embodiment, the method of obtaining the gaze heatmap using eye-tracking data in S102 above may include: data cleaning of the eye-tracking data. For example, abnormal gaze points in the eye-tracking data may be removed. In one possible embodiment, the eye-tracking data includes the coordinates, eye movement states, and velocities of several gaze points. The eye-tracking data includes multiple eye-tracking records, each including a gaze point and its coordinates, eye movement state, and velocity. The above-mentioned method of data cleaning of the eye-tracking data may include: determining valid gaze points and abnormal gaze points based on the coordinates, eye movement states, and velocities of each gaze point; and removing the coordinates, eye movement states, and velocities of abnormal gaze points from the eye-tracking data.
[0078] For example, fixations with anomalous coordinates and velocities, as well as fixations where the eye movement state is blinking or empty (missing) can be considered anomalous fixations. Anomalous coordinates refer to fixations falling in physically impossible locations (e.g., beyond the resolution of the eye tracker screen) or exhibiting significant spatial jumps compared to surrounding points. Anomalous velocities refer to instantaneous eye movements far exceeding the subject's physiological limits. During blinking, the eyes are closed, and no effective visual information is received; therefore, blinking data can be discarded when analyzing fixation behavior.
[0079] Then, a gaze heatmap is obtained based on the frequency of occurrence of several gaze points contained in the cleaned eye-tracking data.
[0080] The cleaned eye-tracking data is converted into a standardized tabular format (e.g., standardized coordinate scale and field naming) to obtain normalized data. For example, this normalized data includes multiple records, each containing an identifier for the scene image, an identifier for the subject, and the x-coordinate and y-coordinate of the fixation point. Here, the x-coordinate and y-coordinate of the fixation point can be coordinates in the coordinate system constructed by the eye tracker's screen, or they can be the x-coordinate and y-coordinate of the pixel corresponding to that fixation point in the scene image.
[0081] In this study, considering that eye-tracking data may contain fixations that occur frequently and fixations that occur less frequently, and that fixations that occur less frequently may be due to the subject’s occasional eye movements, these fixations that occur less frequently may not be included in the construction of the fixation heatmap.
[0082] For example, fixation points that appear more or less frequently than a preset frequency are selected as target fixation points, and then the fixation heatmap is constructed based on the coordinates of each target fixation point.
[0083] In the above scheme, the gaze heatmap is constructed after cleaning the eye movement data, making the obtained gaze heatmap more accurate.
[0084] In one possible embodiment, several relevance metrics include linear correlation coefficient (CC), normalized scan path significance (NSS), area under the curve (AUC), target similarity (SIM), and transportation cost (EMD). S103 above may include the following steps:
[0085] The saliency map and gaze heatmap are flattened to obtain a first vector corresponding to the saliency map and a second vector corresponding to the gaze heatmap. The linear correlation coefficient is obtained based on the sum of the covariances between the first and second vectors. This linear correlation coefficient can be the Pearson linear correlation coefficient, used to measure the degree of linear correlation between pixel values in the saliency map and gaze heatmap; the closer the value is to 1, the stronger the linear correlation, and close to 0, there is no significant linear correlation. The saliency map is flattened to obtain a one-dimensional first vector, and the gaze heatmap is flattened to obtain a one-dimensional second vector. The linear correlation coefficient can be obtained by taking the mean of the first and second vectors and calculating the sum of their covariances. The sum of squared deviations from the mean of the first and second vectors is then calculated. The square root of the product of the two sums of squared deviations is used as the denominator, the sum of the covariances as the numerator, and the fractional value is used as the linear correlation coefficient.
[0086] The saliency map is standardized to obtain a standardized saliency map. Based on the pixel values of the corresponding pixels of each gaze point in the gaze heatmap within the standardized saliency map, the standardized scan path saliency is obtained. The standardized scan path saliency characterizes the degree of matching between salient regions in the saliency map and gaze points in the gaze heatmap; a higher standardized scan path saliency value indicates a higher degree of matching. Specifically, the mean pixel value of the corresponding pixels of each gaze point in the standardized saliency map is used as the standardized scan path saliency.
[0087] The area under the curve (AUC) is determined based on the correspondence between the distribution of salient regions in the saliency map and the distribution of gaze points in the gaze heatmap. The AUC characterizes the ability of the saliency model to distinguish between salient and non-salient regions in a scene image, i.e., the accuracy of distinguishing between salient and non-salient regions in the saliency map. Pixels corresponding to gaze points in the saliency map are used as positive samples, and pixels corresponding to non-gaze points are used as negative samples. All pixel values in the saliency map are sequentially selected as saliency thresholds. Each saliency threshold is used to classify each pixel in the saliency map, determining the category of each pixel under each threshold. The category can predict saliency or non-saliency. The true positive rate and false positive rate are calculated for each saliency threshold. Then, a curve is plotted with the true positive rate on the horizontal axis and the false positive rate on the vertical axis. The area enclosed by the curve and the horizontal axis is the AUC.
[0088] Pixel-level similarity between the saliency map and the fixation heatmap is used as the target similarity. Cosine similarity between the saliency map and the fixation heatmap is also used as the target similarity. The higher the target similarity, the closer the distributions of the saliency map and the fixation heatmap are.
[0089] Obtain the first probability distribution matrix of the saliency map and the second probability distribution matrix of the gaze heatmap. Based on the distance matrix between the first and second probability distribution matrices, the transportation cost is derived. The transportation cost characterizes the difference between the distribution of salient regions in the saliency map and the distribution of gaze points in the gaze heatmap. The smaller the transportation cost, the more similar the distributions of the saliency map and the gaze heatmap. The saliency map and gaze heatmap are normalized to obtain the first and second probability distribution matrices, respectively, where the sum of all elements in both matrices is 1. A distance matrix is constructed based on the distance between any two points in the first and second probability distribution matrices, where each element represents the distance between two points in each matrix. A flow matrix and a minimum cost optimization model are then constructed based on the distance matrix, the first probability distribution matrix, and the second probability distribution matrix. The cost optimization model is essentially a linear programming problem, and the optimal flow matrix can be calculated using a linear programming solver (such as the simplex method). The total workload corresponding to the flow matrix is the transportation cost.
[0090] In one possible embodiment, considering that the types of interference factors in scene images under different scene categories are different, the degree of influence of the determined partial correlation indicators on the visual interference degree of scene images of different scene types may vary. For example, scene images of website images: mainly interactive components, icons, and text layouts, with interference factors mostly being pop-up ads and complex navigation bars; scene images of physical building designs such as buildings: mainly building structures and spatial layouts, with interference factors mostly being redundant lines and significant area blurring caused by monotonous colors; scene images of newspaper images: mainly text blocks and image interspersed, with interference factors mostly being visual conflicts between dense text and images. Therefore, in order to improve the accuracy of the visual interference degree assessment of scene images, the above-mentioned S104 may include Figure 4 The following steps:
[0091] S201, query the correspondence between several preset scene categories and several weighted reassemblies, and obtain the target weighted reassembly corresponding to the scene category to which the scene image belongs.
[0092] The target weighting includes weights for concentration, linear correlation coefficient, significance of standardized scanning path, area under the curve, target similarity, and transportation cost.
[0093] In this process, a unique weighting structure can be pre-established for each scenario category, and the weighting structures for different scenario categories can be different.
[0094] In one possible embodiment, before performing S201, the method further includes:
[0095] Several sets of sample images are obtained under various preset scene categories. Each set of sample images includes a sample saliency map and a sample gaze heatmap. For example, each scene category may include multiple sets of sample images. The sample saliency map in each set of sample images is obtained by processing the sample scene images using a saliency model, while the sample gaze heatmap is obtained based on the eye movement data of the subjects observing the sample scene.
[0096] For each preset scene category, the following steps are performed: The reference concentration of salient regions within the sample salient map in each sample image group, the reference linear correlation coefficient between the sample salient map and the sample gaze heatmap, the reference standardized scan path saliency, the reference area under the curve, the reference target similarity, and the reference transportation cost are obtained, resulting in an indicator dataset for the preset scene category. Then, principal component analysis is performed on the indicator dataset to obtain the weighted reassemblies corresponding to the preset scene categories.
[0097] The reference concentration of salient regions within the salientity map of each sample group can be represented by the Shannon entropy of the salientity map. The method for obtaining the reference linear correlation coefficient between the sample salientity map and the sample fixation heatmap is the same as described above, and will not be repeated here. The method for obtaining the reference standardized scan path saliency is the same as described above, and will not be repeated here. The method for obtaining the reference area under the curve is the same as described above, and will not be repeated here. The method for obtaining the reference target similarity is the same as described above, and will not be repeated here. The method for obtaining the reference transportation cost is the same as described above, and will not be repeated here.
[0098] The indicator dataset may include the reference concentration, reference linear correlation coefficient, reference standardized scan path significance, reference area under the curve, reference target similarity, and reference transportation cost for each sample image group.
[0099] The above method of performing principal component analysis on the indicator dataset to obtain the weighted reorganization corresponding to the preset scene category can be as follows: based on the amount of information (the greater the variance, the greater the information content) carried by each type of indicator (reference concentration, reference linear correlation coefficient, reference standardized scan path significance, reference area under the curve, reference target similarity, and reference transportation cost) in different sample image groups and its correlation with other indicators, the weighted weight of each indicator is determined, and the weighted reorganization is obtained.
[0100] Furthermore, considering that treating the above indicators as isolated mathematical variables and relying solely on the statistical characteristics of the data for weight allocation might fail to explain the correlation between each indicator and interference in each scene category, this embodiment categorizes the indicators to obtain a matching degree factor, a discrimination factor, and a distribution concentration factor. Then, based on the correlation between the matching degree factor, discrimination factor, and distribution concentration factor and visual interference degree in each scene category, the weighted weights of each indicator in each scene category are determined.
[0101] In one possible embodiment, the indicator dataset includes a matching factor, a discriminant factor, and a distribution concentration factor for each group of sample images. The matching factor includes a reference linear correlation coefficient and a reference target similarity; the discriminant factor includes a reference standardized scan path significance and a reference area under the curve; and the distribution concentration factor includes a reference transportation cost and a reference concentration degree. The method of performing principal component analysis on the indicator dataset to obtain the weighted reorganization corresponding to the preset scene category can include, for example... Figure 5 The following steps are shown:
[0102] S301, based on the indicator dataset, determine the interference factor matrix corresponding to each group of sample graphs.
[0103] The interference factor matrix includes the factor values of matching degree factor, discrimination factor, and distribution concentration factor.
[0104] The matching factor can be determined by first normalizing the reference linear correlation coefficient and reference target similarity corresponding to the sample image group, and then using the mean of the normalized reference linear correlation coefficient and reference target similarity as the factor value.
[0105] The method for determining the factor value of the discrimination factor can be as follows: first, normalize the significance of the reference standardized scan path and the area under the reference curve corresponding to the sample image group, and then use the mean of the normalized reference standardized scan path significance and the area under the reference curve as the factor value.
[0106] The method for determining the factor value of the distribution concentration factor can be as follows: first, normalize the reference transportation cost and reference concentration degree corresponding to the sample map group; then, perform a polarity conversion on the normalized reference transportation cost and reference concentration degree; and take the mean of the reference transportation cost and reference concentration degree after the polarity conversion as the factor value.
[0107] S302, cluster each interference factor matrix to obtain the visual interference pseudo-labels for each group of sample images.
[0108] The number of clusters can be determined based on the classification of visual interference levels. For example, if visual interference needs to be classified into low, medium, and high levels, then three clusters will be obtained. Each interference factor matrix belongs to one cluster. Then, the visual interference level corresponding to each cluster is determined based on the distance between each cluster center and the ideal point. The ideal point can be a pre-defined interference factor matrix where the matching factor, discriminant factor, and distribution concentration factor are all 1.
[0109] Specifically, the pseudo-label for the visual interference degree of cluster centers belonging to high-interference clusters is 1, that of cluster centers belonging to low-interference clusters can be 0, and that of cluster centers belonging to medium-interference clusters can be 0.5. The pseudo-label for the visual interference degree of each point (interference factor matrix) within each cluster is determined by the distance between that point and the cluster center, thus obtaining the pseudo-label for the visual interference degree of each group of sample images.
[0110] S303, based on the interference factor matrix of each group of sample images and the visual interference pseudo-label, determine the weighted weights of the matching factor, the discrimination factor and the distribution concentration factor under the preset scene category.
[0111] For example, a mapping model is constructed based on the visual interference pseudo-label, matching factor, discrimination factor, and distribution concentration factor of each group of sample images. This mapping model can be shown in formula (1):
[0112] Formula (1);
[0113] in, Indicates the first A sample image group, Indicates the first The factor values of the matching factor for each sample plot group Indicates the first Factor values of the discrimination factor for each sample group of plots Indicates the first The factor value of the distribution concentration factor for each sample plot group The weighted average of the matching factor. The weighted weights of the discrimination factors are indicated. This represents the weighted weight of the distribution concentration factor. Indicates the first The visual interference pseudo-labels for each sample image group can range from [0,1]. , and The sum is 1.
[0114] This can be solved using the least squares method. , as well as .
[0115] S304. Based on the weighted weights of the matching factor, the discrimination factor, and the distribution concentration factor, determine the weighted reorganization corresponding to the preset scene category.
[0116] Specifically, the weighted weights of the reference linear correlation coefficient and the reference target similarity are determined based on the weighted weights of the matching degree factor. Furthermore, the weighted weights of the reference standardized scanning path significance and the area under the reference curve are determined based on the discrimination factor. Finally, the weighted weights of the reference transportation cost and the reference concentration degree are determined based on the distribution concentration factor.
[0117] For example, the weighted weights of the reference linear correlation coefficient and the reference target similarity are determined based on their information contribution, wherein the sum of the weighted weights of the reference linear correlation coefficient and the reference target similarity is equal to the weighted weight of the matching factor.
[0118] For example, the weights of the reference normalized scan path significance and the area under the reference curve are determined based on their information contribution. The sum of the weights of the reference normalized scan path significance and the area under the reference curve is equal to the weight of the discrimination factor.
[0119] For example, the information contribution of reference transportation costs and reference concentration can be used to determine their respective weights, where the sum of the weights of reference transportation costs and reference concentration is equal to the weight of the distribution concentration factor.
[0120] This yields a weighted reorganization for each scenario category. Specifically, for each scenario category, the weights of reference concentration, reference linear correlation coefficient, reference standardized scan path significance, reference area under the curve, reference target similarity, and reference transportation cost are used as the weights of concentration, linear correlation coefficient, standardized scan path significance, area under the curve, target similarity, and transportation cost in the weighted reorganization.
[0121] In the above scheme, by classifying the data of each correlation index, the relationship between different correlation index data can be better explored, making the determined weight reassembly more accurate, and thus the accuracy of the visual interference obtained by using the corresponding weight reassembly to evaluate the scene image is higher.
[0122] S202, based on target weight reorganization, weighted fusion of concentration degree, linear correlation coefficient, standardized scanning path significance, area under the curve, target similarity and transportation cost is performed to obtain visual interference degree.
[0123] Among them, the weighted fusion result can be directly used as the visual interference degree.
[0124] In the above scheme, the concentration degree, linear correlation coefficient, standardized scan path significance, area under the curve, target similarity and transportation cost are weighted and fused by using the weighted reorganization corresponding to the scene category to which the scene image to be evaluated belongs, so that the determined visual interference degree is more accurate.
[0125] In one possible embodiment, the concentration level includes the Shannon entropy of the saliency map. The smaller the Shannon entropy of the saliency map, the more concentrated the saliency regions in the saliency map. S202 above may include the following steps: normalizing the concentration level and each correlation index respectively; weighting and fusing the normalized linear correlation coefficient, standardized scan path saliency, area under the curve and target similarity, as well as the concentration level after polarity conversion and transportation cost according to target weight reorganization to obtain the visual interference level.
[0126] As mentioned above, the linear correlation coefficient can be the Pearson linear correlation coefficient, which measures the degree of linear correlation between pixel values between the saliency map and the gaze heatmap. The closer the value is to 1, the stronger the linear correlation, and the closer it is to 0, the less obvious the linear correlation.
[0127] Normalized scan path saliency is used to characterize the degree of matching between salient regions in the saliency map and fixation points in the fixation heatmap. The higher the value of normalized scan path saliency, the higher the degree of matching.
[0128] The area under the curve is used to characterize the ability of a saliency model to distinguish between salient and non-salient regions in a scene image, that is, the accuracy of distinguishing between salient and non-salient regions in the saliency map.
[0129] Target similarity can be cosine similarity; the higher the target similarity, the more similar the saliency map and the fixation heatmap are.
[0130] Transportation cost is used to characterize the difference between the distribution of salient regions in the salientity map and the distribution of fixation points in the fixation heatmap. The smaller the transportation cost, the more similar the distribution of the salientity map and the fixation heatmap.
[0131] The degree of concentration includes the Shannon entropy of the saliency map. The smaller the Shannon entropy of the saliency map, the more concentrated the salient regions in the saliency map.
[0132] Therefore, after normalizing transportation costs and concentration levels, a polarity reversal is performed, resulting in a negative correlation between the magnitude of transportation costs and concentration levels and the values after the polarity reversal.
[0133] In the above scheme, the visual interference degree is more accurate by weighting and fusing the normalized linear correlation coefficient, the significance of the standardized scanning path, the area under the curve and the target similarity, as well as the concentration degree after polarity conversion and the transportation cost according to the target weight reorganization.
[0134] Optionally, if there are multiple subjects, the visual interference level corresponding to each subject can be averaged to obtain the final visual interference level of the scene image.
[0135] If there are multiple scene images, some of which belong to different scene categories, and there are multiple subjects, the visual interference level can be obtained from the eye-tracking data collected for each subject observing each scene image. The visual interference levels are listed according to scene number, scene category, and subject's code. The average interference level for the same scene or scene category can be calculated, and this average interference level can be used to compare the interference levels of different scene images or different scene categories.
[0136] In particular, if the scene image has high interference, the saliency concentration area and the fixation point concentration area in the saliency map are often deviated from the key information area or scattered by a large number of irrelevant elements; if the scene image has low interference, the above areas are more concentrated in the key information area.
[0137] The visual interference assessment method provided in this application can be applied to the visual interference assessment of a large number of scene images with multiple subjects and multiple scene categories.
[0138] The following describes a visual interference assessment device provided in this application. The visual interference assessment device described below corresponds to the method of the visual interference assessment device described above.
[0139] This application also provides a visual interference assessment device 500 for ultrasonic endoscopes, applied to electronic devices. Please refer to [link to relevant documentation]. Figure 6The visual interference assessment device 500 includes: a data acquisition unit 501, a data processing unit 502, a correlation index acquisition unit 503, and an interference assessment unit 504. The data acquisition unit 501 is used to acquire the scene image to be assessed and eye movement data of the subject observing the scene image; the data processing unit 502 is used to process the scene image using a preset saliency model to obtain a saliency map of the scene image, and to obtain a gaze heatmap using the eye movement data; the correlation index acquisition unit 503 is used to determine several correlation index data between the saliency map and the gaze heatmap; the interference assessment unit 504 is used to determine the visual interference of the scene image based on the concentration of salient regions within the saliency map and the several correlation index data.
[0140] In one possible embodiment, the plurality of correlation index data includes linear correlation coefficient, standardized scan path saliency, area under the curve, target similarity, and transportation cost; the interference evaluation unit 504 determines the visual interference degree of the scene image based on the concentration of salient regions within the saliency map and the plurality of correlation index data, including:
[0141] The correspondence between several preset scene categories and several weighted reassemblies is queried to obtain the target weighted reassembly corresponding to the scene category to which the scene image belongs. The target weighted reassembly includes the concentration degree, the linear correlation coefficient, the significance of the standardized scan path, the area under the curve, the target similarity, and the weighted weight of the transportation cost.
[0142] Based on the target weighting, the concentration degree, the linear correlation coefficient, the significance of the standardized scanning path, the area under the curve, the target similarity, and the transportation cost are weighted and fused to obtain the visual interference degree.
[0143] In one possible embodiment, before querying the correspondence between several preset scene categories and several weighted reassemblies, the interference evaluation unit 504 is further configured to:
[0144] Obtain several sets of sample images under various preset scene categories. Each set of sample images includes a sample saliency map and a sample gaze heatmap.
[0145] For each preset scene category, perform the following steps:
[0146] Obtain the reference concentration degree of the salient region within the salient map of each sample image group, the reference linear correlation coefficient between the sample salient map and the sample gaze heatmap, the reference normalized scan path saliency, the reference area under the curve, the reference target similarity, and the reference transportation cost, and obtain the index dataset under the preset scene category; and,
[0147] Principal component analysis is performed on the index dataset to obtain the weighted reorganizations corresponding to the preset scene categories.
[0148] In one possible embodiment, the indicator dataset includes a matching factor, a discriminant factor, and a distribution concentration factor for each group of sample images. The matching factor includes the reference linear correlation coefficient and the reference target similarity. The discriminant factor includes the reference standardized scan path significance and the reference area under the curve. The distribution concentration factor includes the reference transportation cost and the reference concentration degree.
[0149] The interference evaluation unit 504 performs principal component analysis on the index dataset to obtain the weighted reassemblies corresponding to the preset scene categories, including:
[0150] Based on the index dataset, an interference factor matrix is determined for each group of sample images. The interference factor matrix includes the factor values of the matching degree factor, the discrimination factor, and the distribution concentration factor.
[0151] Clustering is performed on each of the interference factor matrices to obtain the visual interference pseudo-labels for each group of sample images;
[0152] Based on the interference factor matrix of each group of sample images and the visual interference pseudo-label, the weighted weights of the matching factor, the discrimination factor and the distribution concentration factor under the preset scene category are determined.
[0153] The weighted reorganization corresponding to the preset scene category is determined based on the weighted weights of the matching factor, the distinguishing factor, and the distribution concentration factor.
[0154] In one possible embodiment, the concentration degree includes the Shannon entropy of the saliency map. The interference evaluation unit 504 performs a weighted fusion of the concentration degree, the linear correlation coefficient, the standardized scan path saliency, the area under the curve, the target similarity, and the transportation cost according to the target weight reassembly to obtain the visual interference degree, including:
[0155] The concentration level and each of the correlation indicators are normalized respectively.
[0156] The visual interference degree is obtained by weighting and fusing the normalized linear correlation coefficient, the standardized scanning path significance, the area under the curve and the target similarity, the concentration degree after polarity conversion, and the transportation cost according to the target weighting.
[0157] In one possible embodiment, the correlation index acquisition unit 503 determines several correlation index data between the saliency map and the gaze heatmap, including:
[0158] The saliency map and the gaze heatmap are flattened to obtain a first vector corresponding to the saliency map and a second vector corresponding to the gaze heatmap. The linear correlation coefficient is obtained based on the sum of the covariances between the first and second vectors.
[0159] The saliency map is standardized to obtain a standardized saliency map, and the standardized scan path saliency is obtained based on the pixel value of the corresponding pixel in the standardized saliency map for each gaze point in the gaze heatmap; and,
[0160] Based on the correspondence between the distribution of salient regions in the salientity map and the distribution of each fixation point in the fixation heatmap, the area under the curve is determined; and,
[0161] Obtain the pixel-level similarity between the saliency map and the gaze heatmap as the target similarity; and,
[0162] Obtain the first probability distribution matrix of the saliency map and the second probability distribution matrix of the gaze heatmap, and obtain the transportation cost based on the distance matrix between the first probability distribution matrix and the second probability distribution matrix.
[0163] In one possible embodiment, the data processing unit 502 uses the eye-tracking data to obtain a gaze heatmap, including:
[0164] The eye-tracking data is cleaned.
[0165] The gaze heatmap is obtained based on the frequency of occurrence of several gaze points contained in the cleaned eye movement data.
[0166] In one possible embodiment, the eye-tracking data includes the coordinates of several fixation points, eye movement states, and velocities. The data processing unit 502 performs data cleaning on the eye-tracking data, including:
[0167] Based on the coordinates, eye movement state, and velocity of each fixation point, valid fixation points and abnormal fixation points are determined.
[0168] The coordinates, eye movement state, and velocity of abnormal fixation points are removed from the eye movement data.
[0169] It is understood that since the method embodiments and the device embodiments are different presentations of the same technical concept, the content of the method embodiment section in this application should be adapted to the device embodiment section in a synchronous manner, and will not be repeated here.
[0170] In the case of using integrated units, please refer to Figure 7 , Figure 7 This is the second functional unit block diagram of a visual interference assessment device provided in this application embodiment. The visual interference assessment device is applied to electronic devices. Figure 7 The visual interference assessment device 500 includes a processing module 512 and a communication module 511. The processing module 512 controls and manages the operation of the visual interference assessment device 500, for example, executing the steps of the data acquisition unit, data processing unit, correlation index acquisition unit, and interference assessment unit, and / or performing other processes described herein. The communication module 511 is used for interaction between the visual interference assessment device 500 and other devices. Figure 7 As shown, the visual interference assessment device 500 may also include a storage module 513, which is used to store the program code and data of the visual interference assessment device 500.
[0171] The processing module 512 can be a processor or controller, such as a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The communication module 511 can be a transceiver, RF circuitry, or a communication interface, etc. The storage module 513 can be a memory.
[0172] All relevant content for each scenario involved in the above method embodiments can be referenced from the functional descriptions of the corresponding functional modules, and will not be repeated here. The above-mentioned visual interference evaluation device 500 can execute the above-mentioned visual interference evaluation method.
[0173] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 8As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute the aforementioned visual interference evaluation method.
[0174] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0175] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the visual interference evaluation method provided in the above embodiments.
[0176] This application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the visual interference evaluation methods described above.
[0177] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0178] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes an electronic device.
[0179] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0180] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0181] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0182] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0183] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0184] If the aforementioned integrated units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0185] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0186] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
[0187] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
Claims
1. A method of visual disturbance assessment, characterized by, Applied to electronic devices, the method includes the following steps: Acquire the scene image to be evaluated and the eye movement data of the subject during the observation of the scene image; The scene image is processed using a preset saliency model to obtain a saliency map of the scene image, and a gaze heatmap is obtained using the eye-tracking data; Determine several correlation index data between the saliency map and the gaze heatmap, including linear correlation coefficient, standardized scan path saliency, area under the curve, target similarity, and transportation cost; The visual interference degree of the scene image is determined based on the concentration of salient regions within the saliency map and the several correlation index data, including: querying the correspondence between several preset scene categories and several weighted reassemblies to obtain the target weighted reassembly corresponding to the scene category to which the scene image belongs; and weighting and fusing the concentration degree and the several correlation index data according to the target weighted reassembly to obtain the visual interference degree. The method further includes: Obtain several sets of sample image groups under various preset scene categories. Each set of sample image groups includes a sample saliency map and a sample gaze heatmap. For each preset scene category, perform the following steps: Obtain an indicator dataset under a preset scene category. The indicator dataset includes a matching factor, a discriminant factor, and a distribution concentration factor for each group of sample images. The matching factor includes a reference linear correlation coefficient and a reference target similarity between the sample saliency map and the sample gaze heatmap. The discriminant factor includes a reference normalized scan path saliency and a reference area under the curve between the sample saliency map and the sample gaze heatmap. The distribution concentration factor includes a reference transportation cost between the sample saliency map and the sample gaze heatmap, as well as a reference concentration degree of saliency regions within the sample saliency map. Based on the index dataset, an interference factor matrix is determined for each group of sample images. The interference factor matrix includes the factor values of the matching degree factor, the discrimination factor, and the distribution concentration factor. Clustering is performed on each of the interference factor matrices to obtain the visual interference pseudo-labels for each group of sample images; Based on the interference factor matrix of each group of sample images and the visual interference pseudo-label, the weighted weights of the matching factor, the discrimination factor and the distribution concentration factor under the preset scene category are determined. The weighted reorganization corresponding to the preset scene category is determined based on the weighted weights of the matching factor, the distinguishing factor, and the distribution concentration factor.
2. The method of claim 1, wherein, The step of weighting and fusing the concentration degree, the linear correlation coefficient, the significance of the standardized scanning path, the area under the curve, the target similarity, and the transportation cost according to the target weighting and recombining to obtain the visual interference degree includes: The concentration degree and each of the correlation indicators are normalized respectively, and the polarity of the normalized concentration degree and the transportation cost are reversed. The visual interference degree is obtained by weighting and fusing the normalized linear correlation coefficient, the standardized scanning path significance, the area under the curve, the target similarity, the concentration degree after polarity conversion, and the transportation cost according to the target weighting.
3. The method of claim 2, wherein, The determination of several correlation index data between the saliency map and the gaze heatmap includes: The saliency map and the gaze heatmap are flattened to obtain a first vector corresponding to the saliency map and a second vector corresponding to the gaze heatmap. The linear correlation coefficient is obtained based on the sum of the covariances between the first and second vectors. The saliency map is standardized to obtain a standardized saliency map, and the standardized scan path saliency is obtained based on the pixel value of the corresponding pixel in the standardized saliency map for each gaze point in the gaze heatmap; and, Based on the correspondence between the distribution of salient regions in the salientity map and the distribution of each fixation point in the fixation heatmap, the area under the curve is determined; and, Obtain the pixel-level similarity between the saliency map and the gaze heatmap as the target similarity; and, Obtain the first probability distribution matrix of the saliency map and the second probability distribution matrix of the gaze heatmap, and obtain the transportation cost based on the distance matrix between the first probability distribution matrix and the second probability distribution matrix.
4. The method according to claim 1 or 2, characterized in that, The process of obtaining a gaze heatmap using the eye-tracking data includes: The eye-tracking data is cleaned. The gaze heatmap is obtained based on the frequency of occurrence of several gaze points contained in the cleaned eye movement data.
5. The method of claim 4, wherein, The eye-tracking data includes the coordinates of several fixation points, eye movement states, and velocities. The data cleaning process for the eye-tracking data includes: Based on the coordinates, eye movement state, and velocity of each fixation point, valid fixation points and abnormal fixation points are determined. The coordinates, eye movement state, and velocity of abnormal fixation points are removed from the eye movement data.
6. A visual disturbance evaluation device, characterized by, Applied to electronic devices, including: The data acquisition unit is used to acquire the scene image to be evaluated and the eye movement data of the subject during the observation of the scene image; The data processing unit is used to process the scene image using a preset saliency model to obtain a saliency map of the scene image, and to obtain a gaze heatmap using the eye-tracking data. The correlation index acquisition unit is used to determine several correlation index data between the saliency map and the gaze heatmap. The several correlation index data include linear correlation coefficient, standardized scan path saliency, area under the curve, target similarity, and transportation cost. An interference evaluation unit is used to determine the visual interference degree of a scene image based on the concentration of salient regions within the saliency map and several correlation index data. This includes: querying the correspondence between several preset scene categories and several weighted reassemblies to obtain the target weighted reassembly corresponding to the scene category to which the scene image belongs; and weighting and fusing the concentration and several correlation index data based on the target weighted reassembly to obtain the visual interference degree. The interference evaluation unit is further used for: Obtain several sets of sample image groups under various preset scene categories. Each set of sample image groups includes a sample saliency map and a sample gaze heatmap. For each preset scene category, perform the following steps: Obtain an indicator dataset under a preset scene category. The indicator dataset includes a matching factor, a discriminant factor, and a distribution concentration factor for each group of sample images. The matching factor includes a reference linear correlation coefficient and a reference target similarity between the sample saliency map and the sample gaze heatmap. The discriminant factor includes a reference normalized scan path saliency and a reference area under the curve between the sample saliency map and the sample gaze heatmap. The distribution concentration factor includes a reference transportation cost between the sample saliency map and the sample gaze heatmap, as well as a reference concentration degree of saliency regions within the sample saliency map. Based on the index dataset, an interference factor matrix is determined for each group of sample images. The interference factor matrix includes the factor values of the matching degree factor, the discrimination factor, and the distribution concentration factor. Clustering is performed on each of the interference factor matrices to obtain the visual interference pseudo-labels for each group of sample images; Based on the interference factor matrix of each group of sample images and the visual interference pseudo-label, the weighted weights of the matching factor, the discrimination factor and the distribution concentration factor under the preset scene category are determined. The weighted reorganization corresponding to the preset scene category is determined based on the weighted weights of the matching factor, the distinguishing factor, and the distribution concentration factor.
7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the visual interference evaluation method as described in any one of claims 1-5.