Page analysis method and system, computer equipment and storage medium

By combining eye-tracking data prediction models and language understanding models, we can automatically analyze page design issues and provide optimization suggestions, solving the problems of high cost and low efficiency, enabling non-professionals to optimize pages, and expanding the applicable scenarios of the analysis.

CN121349865APending Publication Date: 2026-01-16MICRO INSURANCE AGENCY LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511371320.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing eye-tracking data acquisition is costly, page analysis methods are inefficient and have limited applicability, professional analysis results are difficult to understand, and non-professionals cannot optimize page design.

Method used

It uses an eye-tracking data prediction model to generate eye-tracking data images corresponding to page images, and uses a language understanding model to convert professional eye-tracking indicators into page analysis conclusions that non-experts can understand. Combined with a visual analysis module and a rule-based reasoning layer, it automatically analyzes page design problems and provides optimization suggestions.

Benefits of technology

It reduces the cost of acquiring eye-tracking data, improves the efficiency of page analysis and the understandability of results, allows non-professionals to optimize page design, and expands the applicable scenarios for analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121349865A_ABST
    Figure CN121349865A_ABST
Patent Text Reader

Abstract

The invention relates to a page analysis method and system, computer equipment and a storage medium. The method comprises the following steps: automatically generating an eye movement data picture corresponding to a page picture by using an eye movement data prediction model, and automatically analyzing the page picture and an eye movement index corresponding to the eye movement data picture, so that an eye tracker experiment is not needed, the acquisition cost of the eye movement index (namely eye movement data) is saved, and the user experience is improved. The eye movement data prediction model can mine eye movement indexes having internal relation with image features on the basis of eye movement data pictures, and then the language understanding large model is used for outputting page analysis conclusions corresponding to the eye movement indexes, so that professional index numerical values are converted into languages which can be understood by non-professionals; according to the method, non-design professional personnel can be assisted to obtain the page evaluation of the to-be-analyzed page and subsequent optimization suggestions, so that the problems that the existing eye movement data acquisition cost is relatively high, and the existing page analysis method is relatively low in efficiency and relatively few in applicable scenes can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a page analysis method, system, computer device, and storage medium. Background Technology

[0002] Currently, eye-tracking data is used in web development for testing and analysis to optimize web design and improve website efficiency. However, existing eye-tracking data is obtained through testing with eye-tracking instruments. In practical applications, eye-tracking testing experiments are costly, requiring specialized eye-tracking equipment and significant investment of manpower and time. Even when eye-tracking data is obtained, current processing methods require professionals to analyze it from a single perspective, resulting in low page analysis efficiency. Furthermore, the results from professional analysis are difficult to understand, as non-professionals cannot grasp the problems or optimization directions of the page. This makes it unsuitable for scenarios where professionals are optimizing page design, limiting the applicability of existing page analysis methods. Summary of the Invention

[0003] This application provides a page analysis method, system, computer device, and storage medium to address the problems of high cost of existing eye-tracking data acquisition, low efficiency of existing page analysis methods, and limited applicable scenarios.

[0004] To solve the above-mentioned technical problems, or at least partially solve them, this application provides a page analysis method, system, computer device, and storage medium.

[0005] On the one hand, this application provides a page analysis method, the method further comprising: When the page image of the page to be analyzed is obtained, the eye-tracking data prediction model is used to generate the eye-tracking data image corresponding to the page image and the eye-tracking index corresponding to the image to be identified, wherein the image to be identified includes the page image and the eye-tracking data image; The language understanding model is used to generate page analysis conclusions corresponding to the eye movement indicators.

[0006] On the other hand, this application provides a page analysis system, the page analysis system comprising: The visual analysis module is used to generate eye-tracking data images corresponding to the page images and eye-tracking indicators corresponding to the images to be identified when the page images of the page to be analyzed are obtained using an eye-tracking data prediction model. The images to be identified include the page images and the eye-tracking data images. The language interpretation module is used to generate page analysis conclusions corresponding to the eye-tracking indicators using a large language understanding model.

[0007] On the other hand, a computer device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor implements the page analysis method described above when executing programs stored in memory.

[0008] On the other hand, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the page analysis method as described above.

[0009] On the other hand, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the page analysis method described above.

[0010] The technical solutions provided in this application have the following advantages compared with the prior art: The method provided in this application uses an eye-tracking data prediction model to automatically generate eye-tracking data images corresponding to page images, and automatically analyzes the page images and the eye-tracking indicators corresponding to the eye-tracking data images. This eliminates the need for eye-tracking experiments, saving the cost of acquiring eye-tracking indicators (i.e., eye-tracking data). Furthermore, the eye-tracking data prediction model can mine eye-tracking indicators that are intrinsically related to image features based on the eye-tracking data images. Then, a language understanding model is used to output page analysis conclusions corresponding to the eye-tracking indicators. This language understanding model can convert professional indicator values ​​into language that non-professionals can understand, assisting non-design professionals in obtaining page evaluations and subsequent optimization suggestions for the page to be analyzed. Compared to manual analysis by professionals, this method improves page analysis efficiency and the comprehensibility of the page analysis results, helping non-professionals optimize page design. Therefore, it can solve the problems of high eye-tracking data acquisition costs, low efficiency of existing page analysis methods, and limited applicable scenarios. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1This is a schematic diagram of the structure of a page analysis system provided in an embodiment of this application; Figure 2 A flowchart illustrating a page analysis method provided in an embodiment of this application; Figure 3 A page image corresponding to the page to be analyzed is provided in an embodiment of this application; Figure 4 This application provides a focus image corresponding to the page to be analyzed in an embodiment of the present application; Figure 5 A heatmap corresponding to a page to be analyzed is provided in an embodiment of this application; Figure 6 A contrast image corresponding to a page to be analyzed is provided in an embodiment of this application; Figure 7 This application provides a page analysis display interface corresponding to the page to be analyzed in an embodiment of the present application; Figure 8 A page image corresponding to the page to be analyzed is provided in an embodiment of this application; Figure 9 A page comparison diagram corresponding to the page to be analyzed is provided for an embodiment of this application; Figure 10 A page comparison diagram corresponding to the page to be analyzed is provided for an embodiment of this application; Figure 11 A page comparison diagram corresponding to the page to be analyzed is provided for an embodiment of this application; Figure 12 A page comparison diagram corresponding to the page to be analyzed is provided for an embodiment of this application; Figure 13 A flowchart illustrating a page analysis method provided in an embodiment of this application; Figure 14 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] Figure 1 This is a diagram illustrating the application environment of a page analysis method in one embodiment. (Refer to...) Figure 1 This page analysis method is applied to a page analysis system. This page analysis system comprises a three-tier architecture, as described above. Figure 13 The first layer is the visual perception layer, which is used to receive page images that have not been tested by actual users, i.e., to receive uploaded design drawings, and generate a heat map representing the probability density of the attention area, a focus map representing the binarized mask corresponding to the visual focus area, and a contrast map marking the color contrast violation area based on the eye-tracking data prediction model. The eye-tracking data prediction model is used to completely replace traditional physical eye-tracking devices and realize simulated eye-tracking data analysis without hardware dependence.

[0016] Reference Figure 13 The second level is the rule-based reasoning layer, a configurable analysis engine. Its core consists of a decision rule system based on a directed acyclic graph (DAG). Specifically, the focus image verification module, sharpness scoring module, and AOI contrast analysis module are stored using a DAG structure, and keypoint tracking algorithms and region recognition models are integrated as auxiliary analysis tools. The processing strictly follows the analysis workflow corresponding to the page analysis method. Using the heatmap, focus image, and contrast image output from the previous level as input, the final output includes a multi-dimensional diagnostic vector containing key element coverage, sharpness score, attention focus, attention percentage, and contrast, used to transform professional design evaluation into an automated quantitative analysis process.

[0017] Reference Figure 13 The third level is the semantic interaction layer, which acts as a cross-modal alignment translator. It receives the multi-dimensional diagnostic vectors output from the second level as input. The core adopts a multi-modal large model with a visual-language alignment architecture and integrates a historical case database with similarity index as a case support system. The processing includes three key steps: mapping the multi-dimensional diagnostic vectors to natural language prompts, generating a complete analysis description through the large model decoder, and retrieving similar optimization schemes from the historical case database. Finally, it outputs natural language optimization suggestion text designed for non-professionals and comparative data on the effects of historical revision cases with empirical reference value, which is used to achieve efficient conversion between professional technical indicators and user-operable decision guidance.

[0018] The visual analysis module 110 includes a first-level visual perception layer and a second-level rule reasoning layer, while the language interpretation module 120 includes a third-level semantic interaction layer. The page analysis system, through the deep application of eye-tracking data prediction models, enables it to learn from a large amount of historical data and cases, understand the complex relationships between data, predict user behavior patterns based on input page images, judge the rationality of page design, and provide targeted optimization suggestions. Furthermore, the eye-tracking data prediction model has the ability to continuously learn and optimize; as new data is continuously input, it can automatically update its internal parameters and knowledge graph, improving the accuracy and adaptability of analysis and calculation tasks, effectively lowering the design evaluation threshold. By using eye-tracking data prediction models and large-scale language understanding models, professional eye-tracking analysis is transformed into an automated technical process, enabling non-designers to independently complete design verification and improving the accuracy of users' understanding of the page analysis system's output suggestions. The visualized analysis path ensures the technical verifiability of each page analysis conclusion.

[0019] In one embodiment, Figure 2 This is a flowchart illustrating a page analysis method in one embodiment, with reference to... Figure 2 This paper provides a page analysis method. This embodiment primarily applies this method to the above-mentioned... Figure 1 Taking a page analysis system as an example, the specific steps of this page analysis method are as follows: Step S210: When the page image of the page to be analyzed is obtained, the eye-tracking data prediction model is used to generate the eye-tracking data image corresponding to the page image and the eye-tracking index corresponding to the image to be identified, wherein the image to be identified includes the page image and the eye-tracking data image.

[0020] Specifically, the page to be analyzed can be a website page, a mobile page, or a page embedded in a mini-program or application. The page image is an image that shows the entire page to be analyzed. The page image is uploaded to the page analysis system by testers, who can be web design professionals or non-professionals. The page image can be a photograph taken by the tester of the page to be analyzed, or a screenshot taken by the tester of the page to be analyzed.

[0021] The eye-tracking data prediction model is trained with a domain-knowledge-enhanced visual attention prediction model to replace traditional physical eye-tracking devices. It generates multiple eye-tracking data images for page images and simultaneously outputs quantitative evaluation metrics, i.e., multiple eye-tracking metrics. The eye-tracking data images are used to reflect the visual attention prediction results of the page images, i.e., to display the predicted eye-tracking data. The eye-tracking data includes the distribution of gaze points, the duration of gaze point dwell, and the gaze point movement path, while the eye-tracking metrics indicate the analysis results of the eye-tracking data images and image features.

[0022] Furthermore, eye-tracking data prediction models have applications beyond web design, potentially playing a significant role in advertising, marketing, product packaging design, and exhibition design. By deeply analyzing users' eye-tracking behavior, businesses and designers can better understand user needs and preferences, thereby developing more effective marketing strategies and design solutions, and enhancing the competitiveness of their products and services.

[0023] Step S220: Use the language understanding big data model to generate the page analysis conclusions corresponding to the eye movement indicators.

[0024] Specifically, the language understanding big data model is used to convert multiple eye-tracking metrics into corresponding natural language prompts based on the mapping relationship between eye-tracking metrics and natural language prompts. The big data model decoder generates a complete analysis description and searches the historical case database to find similar optimization solutions that match the natural language prompts. The final page analysis conclusion is then output, which includes at least the analysis descriptions corresponding to multiple eye-tracking metrics, natural language optimization suggestion text corresponding to similar optimization solutions, and historical redesign case effect comparison data corresponding to similar optimization solutions.

[0025] Based on the above method, an eye-tracking data prediction model is used to automatically generate eye-tracking data images corresponding to page images, and automatically analyze the page images and the corresponding eye-tracking indicators. This eliminates the need for eye-tracking experiments, saving the cost of acquiring eye-tracking indicators (i.e., eye-tracking data). Furthermore, the eye-tracking data prediction model can mine eye-tracking indicators that are intrinsically related to image features based on the eye-tracking data images. Then, a language understanding model is used to output the page analysis conclusions corresponding to the eye-tracking indicators. The language understanding model can convert professional indicator values ​​into language that non-professionals can understand, which can help non-design professionals obtain page evaluations and subsequent optimization suggestions for the page to be analyzed. Compared with manual analysis by professionals, this improves the efficiency of page analysis and the comprehensibility of page analysis results, and can help non-professionals optimize page design. Therefore, it can solve the problems of high cost of acquiring eye-tracking data, low efficiency of existing page analysis methods, and limited applicable scenarios.

[0026] In one embodiment, generating eye-tracking data images corresponding to the page images and eye-tracking metrics corresponding to the images to be identified using an eye-tracking data prediction model includes: An eye-tracking data prediction model is used to generate a focus image, a heatmap, and a contrast image corresponding to the page image, as well as eye-tracking indices corresponding to the page image, the focus image, the heatmap, and the contrast image. The eye-tracking data image includes the focus image, the heatmap, and the contrast image.

[0027] Specifically, Figure 3For the page image to be analyzed, refer to... Figure 4 The focus map is used to display the predicted distribution of gaze points through two display states: highlighted and hidden. The focus map uses a gradient of transparency to show what content a user will see and miss within a preset timeframe when first viewing the page being analyzed. Highlighted content indicates that the user has seen it, while hidden content indicates that the user has not seen it. (See reference...) Figure 5 Heatmaps are used to display the predicted level of attention at different gaze points using different colors. This is based on the predicted distribution of gaze points and the duration of gaze at each point, determining the level of attention for each gaze point. A visual gradient of colors is used to show the areas of content that users are focusing on; warm colors indicate more user attention, and cool colors indicate less user attention. (See reference...) Figure 6 Contrast charts are used to show the degree of contrast between different important elements on a page being analyzed, using different colors. For example, green indicates sufficient contrast, while blue indicates insufficient contrast. Contrast charts help create visually appealing and easy-to-understand designs. Analysis of contrast charts can improve user experience, increase conversion rates, and ensure that content is inclusive.

[0028] The eye-tracking data prediction model evaluates and calculates page images, focus maps, heatmaps, and contrast maps separately, thereby generating eye-tracking metrics corresponding to page images, focus maps, heatmaps, and contrast maps.

[0029] For webpage images, eye-tracking metrics can accurately reflect the attractiveness of different parts of the image to the user. By analyzing eye-tracking metrics, it's possible to determine which areas of the image quickly capture the user's attention and which areas are easily overlooked. Based on this, images can be optimized during page design, such as highlighting highly attractive areas or adjusting less attractive areas, to improve the overall visual appeal and information delivery of the image.

[0030] For featured images, eye-tracking metrics can help refine the settings for their highlighted and hidden states. By analyzing eye-tracking data from different user groups, more precise preset durations can be determined, allowing featured images to more accurately represent the distribution of users' gaze during actual browsing. Furthermore, the parameters for transparency gradients can be adjusted based on eye-tracking metrics, making the featured image display more intuitive and clear, thus better assisting page designers in understanding users' focus areas.

[0031] Eye-tracking metrics from heatmaps provide strong evidence for optimizing page layout. Based on the different colors representing varying levels of attention, designers can adjust the position and size of various elements on the page. For example, placing content that users are more interested in in more prominent positions or appropriately increasing its size can further enhance user focus on important content. Furthermore, long-term monitoring and analysis of heatmap eye-tracking metrics can reveal trends in user browsing habits, allowing for timely dynamic adjustments to the page layout.

[0032] Eye-tracking metrics in contrast charts play a crucial role in enhancing the visual experience and inclusivity of a webpage. By continuously analyzing these metrics, the criteria for different colors representing contrast levels can be refined, allowing contrast charts to more accurately reflect the contrast between different elements on a page. For areas with insufficient contrast, designers can make targeted adjustments based on eye-tracking metrics, such as changing the color, brightness, or saturation of elements, to ensure that the page content is both visually appealing and easy to understand.

[0033] In one embodiment, generating eye-tracking metrics corresponding to the page image, the focus image, the heatmap, and the contrast image using an eye-tracking data prediction model includes: Using the eye-tracking data prediction model, key element coverage and key element sharpness are generated for the focal map; The eye-tracking data prediction model is used to generate page clarity for the page image; Using the eye-tracking data prediction model, attention focus and attention percentage of each key element are generated for the heatmap. Using the eye-tracking data prediction model, the contrast of each key element is generated for the contrast map. The multiple eye-tracking indicators include the coverage of the key elements, the clarity of the key elements, the clarity of the page, the focus of attention, the attention percentage of each key element, and the contrast of each key element.

[0034] Specifically, eye-tracking data prediction models are used to analyze the distribution of fixation points in the focus image, thereby calculating key element coverage and key element clarity, converting the image display effect into quantitative indicators. Key element coverage indicates the exposure range of key elements; a higher coverage rate indicates a larger exposure range and a larger area related to the product's selling points. Key element clarity indicates the duration and frequency of user attention to key elements; higher clarity indicates a longer duration and / or higher frequency of user attention to key elements.

[0035] The eye-tracking data prediction model uses a sharpness algorithm to comprehensively analyze the amount of text, font size, text contrast, color richness, number of images, and image size in page images, and finally obtains a comprehensive score, which is the page sharpness. Different page sharpness scores are used to reflect the readability of key elements in page images. Assuming the page clarity score ranges from 0 to 100, a score of 0-29 indicates extremely difficult visibility of key elements, meaning the page images are cluttered and the key elements are hard for users to see. A score of 30-56 indicates moderate difficulty. A score of 57-94 indicates easy visibility, meaning the page images are clean and clear, with 57-94 being the optimal range for page design optimization. A score of 95-100 indicates overly easy visibility, meaning the page images are too simplistic and straightforward, potentially negatively impacting the user experience.

[0036] The eye-tracking data prediction model uses a keypoint tracking algorithm to analyze heatmaps to determine the duration and / or frequency of attention at each fixation point. It then combines this with a region recognition model to calculate attention focus and the area of ​​interest (AOI) for each key element. Attention focus indicates the degree of user concentration; a higher focus indicates that the user's attention is concentrated in a few key areas, while a lower focus indicates more scattered attention, making it harder for the user to quickly identify important content. Key elements refer to product-related elements on the page being analyzed, or other elements defined by the testers. Examples of key elements include product-related logos, titles, subtitles, call-to-action (CTA) controls, images, text, menus, search functions, banners, and prices. The area of ​​interest for key elements indicates the degree of user attention to these elements; a higher area of ​​interest is better, as it helps users focus on important product-related content.

[0037] The eye-tracking data prediction model calculates the contrast of each key element based on the contrast of different fixation points in the contrast map. The contrast of key elements is used to indicate the prominence of key elements. Ideally, the contrast of key elements should be as high as possible so that users can quickly see the prominent key elements.

[0038] The eye-tracking data prediction model uses preset algorithms based on focus maps, heatmaps, contrast maps, and page images to calculate multiple eye-tracking metrics, including key element coverage, key element clarity, page clarity, attention focus, attention share of each key element, and contrast of each key element. These metrics help understand user focus and page strengths and weaknesses during testing. (Reference) Figure 13 As shown, the preset algorithms corresponding to the focus image are stored in the focus image verification module, the sharpness algorithms corresponding to the page image are stored in the sharpness scoring module, and the preset algorithms corresponding to the heat map and the contrast map are stored in the AOI contrast analysis module. That is, the preset algorithms corresponding to different types are stored in a directed acyclic graph structure.

[0039] Reference Figure 7 It also features a dedicated user interface that intuitively and clearly integrates and displays the generated heatmaps, focus maps, and analysis suggestions and conclusions generated by the language understanding model.

[0040] To further improve the accuracy and practicality of eye-tracking data prediction models, dynamic upgrades can be implemented. A real-time data update mechanism can be introduced, allowing the model to adjust the calculation results of various eye-tracking indicators in real time based on the latest eye-tracking data. For example, during the actual display of the page being analyzed, as the user's browsing behavior changes, eye-tracking indicators such as key element coverage, clarity, attention focus, and contrast are updated in real time, ensuring that the optimization of the page being analyzed keeps pace with the user's real-time feedback.

[0041] Regarding multi-device adaptation, the eye-tracking data prediction model is specifically optimized to account for differences in screen size, resolution, and display ratio across various devices (such as mobile phones, tablets, and computers). For different device types, the eye-tracking data prediction model establishes corresponding device parameter adjustment rules to ensure accurate calculation of eye-tracking metrics that reflect real-world conditions across various devices. For example, on small-screen devices, it may be necessary to appropriately increase the contrast and sharpness thresholds of key elements to ensure that these elements can still be clearly recognized by the user within the limited screen space.

[0042] Furthermore, machine learning algorithms can be combined to deeply mine and analyze massive amounts of eye-tracking data. By training eye-tracking data prediction models to learn the browsing habits and preferences of different user groups, personalized page optimization suggestions can be provided. For example, younger users may prefer page designs with rich colors and strong dynamic effects. Based on this characteristic, the eye-tracking data prediction model can adjust the layout and style of key elements in the page display for this group, improving user attention and experience.

[0043] In one embodiment, generating page sharpness for the page image using the eye-tracking data prediction model includes: The eye-tracking data prediction model is used to generate gradient magnitude, local contrast, and information entropy for the page image. The page clarity is determined based on the weighted sum of the gradient magnitude, the local contrast, and the information entropy.

[0044] Specifically, the eye-tracking data prediction model uses a sharpness algorithm to calculate the gradient magnitude, local contrast, and information entropy of the page image. For each pixel in the page image, its gradient values ​​in the horizontal and vertical directions are calculated. The formula for calculating the gradient magnitude is as follows: ,in and These are the gradient values ​​in the horizontal and vertical directions, respectively.

[0045] Simultaneously, for each pixel in the page image, a pixel window is defined with that pixel as the center and a preset size of n*n. Local contrast is then calculated for each pixel window using the following formula: , The standard deviation of multiple pixel values ​​within a pixel window. It is the average of multiple pixel values ​​within the pixel window.

[0046] Furthermore, considering that the information entropy of page images also reflects their clarity, the higher the information entropy, the more information the page image contains, and the higher its clarity. The formula for calculating information entropy is: ,in Is the grayscale value The probability of a pixel appearing. This refers to the grayscale level of the image on the page.

[0047] Page image sharpness is determined by a weighted sum of three factors: gradient magnitude, local contrast, and information entropy. The page sharpness value ranges from [0, 100], with higher values ​​indicating sharper images. The weights of the three parameters can be customized based on the specific application scenario, allowing for flexible application to different types of pages and images. Whether it's a static or dynamic page, or a high-resolution or standard-resolution image, adjusting the weights can achieve the best sharpness calculation results.

[0048] Calculating page image sharpness using eye-tracking data prediction models provides accurate and scientific data for page design and optimization. During the page design phase, testers can use the page sharpness calculated by this model to adjust elements such as contrast, color, and detail, thereby improving the overall visual appeal and user experience. Improved image sharpness allows users to more easily identify content within images, reducing visual fatigue and difficulty in information retrieval caused by blurry images. This is particularly important for e-commerce pages, news pages, and other scenarios that rely on images to display information, increasing user attention and dwell time, ultimately improving conversion rates and user satisfaction.

[0049] In one embodiment, generating attention focus using the eye-tracking data prediction model for the heatmap includes: The heatmap is binarized using the eye-tracking data prediction model to determine the region of interest within the heatmap. Determine the total area of ​​each region of interest in the heat map and the location of the region's centroid. The area ratio is determined based on the ratio between the sum of the areas of each region of interest and the area of ​​the heatmap image. The target distance is determined based on the distance between the centroid of the region and the center of the key region in the heat map; The weighted difference between the area ratio of the region and the target distance is used as the attention focus.

[0050] Specifically, the eye-tracking data prediction model binarizes the heatmap, identifying regions with binarized values ​​exceeding a preset threshold as regions of interest, and summing the areas of all regions of interest to obtain the total area. The ratio of the sum of the areas of the regions to the area of ​​the image is obtained by dividing the sum of the areas of the regions. ,in The area is the image area.

[0051] The location of the centroid of all regions of interest is determined based on the distribution of each region of interest. The key region in the heat map is used to indicate the region containing multiple key elements. The distance between the center of the key region and the centroid of the region is the target distance d.

[0052] Attention focus is determined by a weighted difference between the region area ratio and the target distance; the formula for calculating attention focus is as follows: ,in and For example, weighting coefficients ,or ,or Wait, in this embodiment, let The weighting coefficients are optimized and determined through extensive experimental data, or can be customized according to different application scenarios. This flexibility allows the model to be widely applied in various fields, such as rehabilitation training monitoring in the medical field and user experience evaluation in game design. In different scenarios, users can flexibly adjust the weighting coefficients according to their specific needs and concerns, so that the attention focus can more accurately reflect the user's concentration on key areas in that scenario, improving the model's adaptability and practicality.

[0053] By comprehensively analyzing the area ratio and target distance, the model can not only quantify attention focus but also gain deep insights into user behavior patterns. The area ratio reflects the breadth of the user's attention area, while the target distance reflects the proximity of the user's attention area to key areas. Combining these two factors, researchers can understand how users distribute their attention when viewing images or interfaces—for example, whether users tend to browse broadly or focus more on key areas. This in-depth user behavior insight helps optimize product design, interface layout, and information presentation to better meet user needs and improve user experience.

[0054] Since attention focus can accurately reflect the user's level of concentration on key areas, the range of attention focus values ​​is: A higher numerical value indicates a higher level of focus, and this intuitive numerical representation makes the results easy to understand and interpret. Both professional researchers and ordinary users can quickly determine a user's level of focus on key areas based on the attention focus value. At the same time, this intuitive presentation of results facilitates data comparison and analysis in different scenarios, helping to uncover potential problems and patterns.

[0055] Eye-tracking data prediction models accurately quantify a user's attention focus using scientific calculation methods. Traditional eye-tracking data analysis often presents vague assessments of attention, making precise numerical measurement difficult. This eye-tracking data prediction model, however, binarizes heatmaps, calculates region area ratios and target distances, and combines these with a weighted difference method to determine attention focus, resulting in more accurate and objective attention quantification.

[0056] In one embodiment, the eye-tracking data prediction model is used to generate the attention percentage of each key element in the heatmap, including: The eye-tracking data prediction model is used to segment multiple key element regions in the heat map, and the number of fixation points and fixation duration in each key element region are counted. The attention level of each key element region is obtained by weighting and summing the number of fixations and fixation duration in each key element region. The ratio between the attention given to each of the key element regions and the sum of the attention given to all the key element regions is determined as the attention percentage of each of the key element regions.

[0057] Specifically, the eye-tracking data prediction model uses image segmentation technology to accurately segment predefined key element regions (AOI regions) from the heatmap. An AOI region is the smallest region containing a single key element; each key element corresponds to one AOI. The model then counts the number of fixations and fixation duration within each AOI region. Specifically, it uses the mapping relationship between eye-tracking data and the heatmap to count the number of fixations (N) falling within the AOI region, and the fixation duration is calculated by summing the duration of each fixation point within the AOI region. The result, i.e., the fixation duration, is .

[0058] The attention level of key element regions is determined by weighted summation of the number of fixations and fixation duration. The formula for calculating attention level is as follows: ,in and For example, weighting coefficients ,or ,or Wait, in this embodiment, let This can be determined by analyzing and optimizing eye-tracking data from different types of pages and user groups, or by customizing settings according to different business scenarios.

[0059] Divide the attention level of the current key element region by the sum of the attention levels of all key element regions to obtain the attention percentage of the current key element region. The formula for calculating the attention percentage is as follows: ,in, For the attention given to the i-th key element region, The AOI value is the sum of attention received by all key element areas. This ensures that the AOI value accurately reflects the importance and user attention of key element areas. A higher AOI value indicates that the key element area receives more attention. If the attention share of the current key element area is lower than the average attention, it is recommended to optimize the design of the current key element.

[0060] The success rate of a transaction can be predicted by analyzing the attention share of each key element area. For example, the click-through rate of a button can be predicted based on the attention share of the area containing a navigation control, thus predicting the probability that a user will click the "transaction" button and complete the transaction. Comparing the AOI values ​​of multiple elements within a single page helps understand the prominence of each key element, assisting testers in optimizing these elements based on product and design goals. Comparing the AOI values ​​of similar key elements across multiple pages helps testers determine which page better achieves the design objectives.

[0061] In one embodiment, generating the focus map, heatmap, and contrast map corresponding to the page image using an eye-tracking data prediction model includes: An eye-tracking data prediction model is used to generate an initial focus image, an initial thermal image, and an initial contrast image corresponding to the page image. The initial focus image, the initial heat map, and the initial contrast image are modified according to preset constraints to obtain the corresponding focus image, the heat map, and the contrast image. The preset constraints include at least one of the following: standardized design conditions, user visual habit constraints, and industry compliance conditions.

[0062] Specifically, the eye-tracking data prediction model generates initial focus maps, initial heatmaps, and initial contrast maps based on the eye-tracking data predicted from page images. The eye-tracking data prediction model is equipped with a configuration rule engine, which, after structured processing, forms a computable knowledge graph. The configuration rule engine includes preset constraints, which must include at least one of the following: standardized design conditions, user visual habit constraints, and industry compliance conditions. Standardized design conditions include element salience constraints, accessibility compliance constraints, and information hierarchy constraints. Element salience constraints specify the minimum attention threshold for key functional areas in the generated heatmap. When the actual output value fails to meet the standard, the system automatically triggers a correction procedure to increase color saturation or size. Accessibility compliance constraints strictly adhere to design guidelines, requiring a specific contrast ratio between text and background colors. The system highlights non-compliant areas by scanning the contrast map in real time. Information hierarchy constraints clarify the minimum exposure area ratio of core content in the visual focus map and set a spatial position weight matrix to ensure that important elements are distributed in the visual center area, such as the visual weight allocation of premium amounts and exclusion clauses. The above standardized design conditions are stored in the form of a structured data table. Each record contains four elements: the category of the constraint object, the name of the metric, the threshold range value, and a description of the correction strategy.

[0063] User visual habit constraints include reading flow models that conform to visual habits, minimum size thresholds for interactive elements, and baseline data on user attention dwell time. In conditional generative adversarial networks, the preset constraints extracted from the aforementioned knowledge graph are encoded as dynamic control vectors, which run through the generation process of the eye-tracking data prediction model. For example, when the eye-tracking data prediction model recognizes the presence of a legal text box in a page image, it automatically triggers minimum font size and contrast enhancement conditions; when it detects the insurance button area, it forces that area to present 30%-50% visual weight in the heatmap. This conditional mechanism essentially transforms industry knowledge into control signals that the algorithm can recognize, and continuously corrects the output of the eye-tracking data prediction model through adversarial training between the generator and the discriminator.

[0064] Therefore, based on the trained eye-tracking data prediction model, the initial focus map, initial heatmap, and initial contrast map are dynamically adjusted according to preset constraints, thereby outputting focus maps, heatmaps, and contrast maps that meet the preset constraints. For example, when the highlighted area in the initial heatmap exceeds the functional area coordinate boundary, the eye-tracking data prediction model automatically applies gradient penalties and initiates an inverse reinforcement mechanism for necessary elements with insufficient weights. After iterative training, the heatmap output by the eye-tracking data prediction model can strictly follow industry standards, eliminate invalid noise, and accurately locate the focus.

[0065] Image information (heatmaps, focus maps, contrast maps) generated by a large eye-tracking data model, along with pre-defined analytical methodologies, is transformed into natural language text descriptions. This text is then input into a high-performance language understanding model, which comprehensively and deeply interprets the importance of key regions in the focus map, the meaning of AOI region scores, and user browsing behavior reflected in the heatmap trajectory. This multi-model collaborative approach enables in-depth analysis from both user attention distribution and semantic understanding perspectives, overcoming the limitations of single-model analysis in existing technologies and providing users with more comprehensive and in-depth image analysis results.

[0066] In one specific embodiment, the page image of the page to be analyzed is as follows: Figure 8 As shown, the eye-tracking data prediction model outputs a focus image based on the page image, as follows: Figure 9 The old version of the focus map and heat map in the image are as follows: Figure 10 As shown in the older version, the heatmap with attention focus is as follows: Figure 11 As shown in the old version, the heatmap with attention percentage is as follows: Figure 12As shown in the old version, the eye-tracking data prediction model outputs multiple eye-tracking indicators based on multiple eye-tracking data images. The language understanding model outputs page analysis conclusions based on focus maps, heatmaps, and contrast maps (not shown here). The page analysis conclusions are: the selling point area is poorly exposed, the focus brightness of the selling point area is dim, the selling points are noticed by users less frequently and for less time, and the focus is relatively scattered, resulting in low attention concentration. The page clarity is low, and the attention share of the key element area corresponding to the selling points is low. Therefore, important content on the page to be analyzed is not easily seen by users, and the page design needs to be optimized. The page design suggestions are to expand the selling point exposure area, concentrate the selling points, and improve page clarity. Based on the page analysis conclusions, non-professionals can easily understand the design problems and optimization solutions of the page to be analyzed, and then optimize the page. Specifically, the page to be analyzed can be optimized manually, or an artificial intelligence model can be used to output an optimized page based on the page analysis conclusions and the page to be analyzed. The optimized page can be referenced. Figure 9 The new version in [the text].

[0067] Then, using eye-tracking data prediction models, new focus maps and heatmaps are output for the optimized page. The new focus map is based on... Figure 9 The new focus map and heat map are referenced in the text. Figure 10 , Figure 11 , Figure 12 The new version in China is based on Figure 9 The new featured image reveals more of the selling points, meaning they are more easily seen by users. The brighter focus area of ​​the selling points indicates that they are being noticed by users more frequently and for longer periods. Figure 10 The results showed that the new heatmap had higher page clarity than the old version, indicating that the optimized page was clearer and easier to understand, and had a higher focus, meaning the new version was more focused and clearer. Figure 11 It was found that the key elements of the new homepage are more focused, and attention is concentrated on the core selling points, which is the desired effect. Based on Figure 12 The analysis revealed that the new version of the page prioritizes key selling points, aligning better with the design requirement of focusing user attention on these points, thus improving conversion rates. Furthermore, the order of emphasis on key selling points on the page also conforms to design expectations. In conclusion, the optimized new page is superior to the old page under analysis, demonstrating that page optimization can improve user viewing experience and conversion rates.

[0068] Figure 2 This is a flowchart illustrating a page analysis method in one embodiment. It should be understood that, although... Figure 2 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0069] In one embodiment, such as Figure 1 As shown, a page analysis system is provided, including: The visual analysis module 110 is used to generate an eye-tracking data image corresponding to the page image and an eye-tracking index corresponding to the image to be identified when the page image of the page to be analyzed is obtained, using an eye-tracking data prediction model. The image to be identified includes the page image and the eye-tracking data image. The language interpretation module 120 is used to generate page analysis conclusions corresponding to the eye movement indicators using a large language understanding model.

[0070] In one embodiment, the visual perception layer in the visual analysis module 110 is used for: An eye-tracking data prediction model is used to generate a focus image, a heatmap, and a contrast image corresponding to the page image, as well as eye-tracking indices corresponding to the page image, the focus image, the heatmap, and the contrast image. The eye-tracking data image includes the focus image, the heatmap, and the contrast image.

[0071] In one embodiment, the rule reasoning layer in the visual analysis module 110 is used for: Using the eye-tracking data prediction model, key element coverage and key element sharpness are generated for the focal map; The eye-tracking data prediction model is used to generate page clarity for the page image; Using the eye-tracking data prediction model, attention focus and attention percentage of each key element are generated for the heatmap. Using the eye-tracking data prediction model, the contrast of each key element is generated for the contrast map. The multiple eye-tracking indicators include the coverage of the key elements, the clarity of the key elements, the clarity of the page, the focus of attention, the attention percentage of each key element, and the contrast of each key element.

[0072] In one embodiment, the rule reasoning layer is further used for: The eye-tracking data prediction model is used to generate gradient magnitude, local contrast, and information entropy for the page image. The page clarity is determined based on the weighted sum of the gradient magnitude, the local contrast, and the information entropy.

[0073] In one embodiment, the rule reasoning layer is further used for: The heatmap is binarized using the eye-tracking data prediction model to determine the region of interest within the heatmap. Determine the total area of ​​each region of interest in the heat map and the location of the region's centroid. The area ratio is determined based on the ratio between the sum of the areas of each region of interest and the area of ​​the heatmap image. The target distance is determined based on the distance between the centroid of the region and the center of the key region in the heat map; The weighted difference between the area ratio of the region and the target distance is used as the attention focus.

[0074] In one embodiment, the rule reasoning layer is further used for: The eye-tracking data prediction model is used to segment multiple key element regions in the heat map, and the number of fixation points and fixation duration in each key element region are counted. The attention level of each key element region is obtained by weighting and summing the number of fixations and fixation duration in each key element region. The ratio between the attention given to each of the key element regions and the sum of the attention given to all the key element regions is determined as the attention percentage of each of the key element regions.

[0075] In one embodiment, the visual perception layer is further used for: An eye-tracking data prediction model is used to generate an initial focus image, an initial thermal image, and an initial contrast image corresponding to the page image. The initial focus image, the initial heat map, and the initial contrast image are modified according to preset constraints to obtain the corresponding focus image, the heat map, and the contrast image. The preset constraints include at least one of the following: standardized design conditions, user visual habit constraints, and industry compliance conditions.

[0076] like Figure 14 As shown, this application provides a computer device including a processor 711, a communication interface 712, a memory 713, and a communication bus 714, wherein the processor 711, the communication interface 712, and the memory 713 communicate with each other through the communication bus 714. Memory 713 is used to store computer programs; The processor 711, when executing the program stored in the memory 713, implements the page analysis control method provided in any of the foregoing method embodiments.

[0077] Those skilled in the art will understand that Figure 14 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0078] In one embodiment, the page analysis system provided in this application can be implemented as a computer program, which can be implemented in various ways, such as... Figure 14 It runs on the computer device shown. The computer device's memory can store the various program modules that make up the page analysis system, for example, Figure 1 The visual analysis module 110 and the language interpretation module 120 are shown. The computer program, composed of these various program modules, causes the processor to execute the steps in the page analysis methods of the various embodiments of this application described in this specification.

[0079] Figure 14 The computer device shown can be used as follows Figure 1 In the page analysis system shown, the visual analysis module 110, upon acquiring a page image of the page to be analyzed, uses an eye-tracking data prediction model to generate an eye-tracking data image corresponding to the page image, and an eye-tracking index corresponding to the image to be identified. The image to be identified includes both the page image and the eye-tracking data image. The computer device can use the language interpretation module 120 to execute a language understanding model to generate page analysis conclusions corresponding to the eye-tracking indexes.

[0080] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0081] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method of page analysis, characterized by, The method further comprises: when the page picture of the page to be analyzed is acquired, generating, by using an eye movement data prediction model, an eye movement data picture corresponding to the page picture and an eye movement index corresponding to a picture to be recognized, wherein the picture to be recognized comprises the page picture and the eye movement data picture; generating, by using a language understanding large model, a page analysis conclusion corresponding to the eye movement index.

2. The method of claim 1, wherein, The generating, by using the eye movement data prediction model, of the eye movement data picture corresponding to the page picture and the eye movement index corresponding to the picture to be recognized comprises: generating, by using the eye movement data prediction model, a focal point map, a heat map, and a contrast map corresponding to the page picture, and generating an eye movement index corresponding to the page picture, the focal point map, the heat map, and the contrast map, wherein the eye movement data picture comprises the focal point map, the heat map, and the contrast map.

3. The method of claim 2, wherein, The generating, by using the eye movement data prediction model, of the eye movement index corresponding to the page picture, the focal point map, the heat map, and the contrast map comprises: generating, by using the eye movement data prediction model, key element coverage and key element clarity for the focal point map; generating, by using the eye movement data prediction model, page clarity for the page picture; generating, by using the eye movement data prediction model, attention focus degree and attention proportion of each key element for the heat map; generating, by using the eye movement data prediction model, contrast of each key element for the contrast map, wherein the plurality of eye movement indexes comprise the key element coverage, the key element clarity, the page clarity, the attention focus degree, the attention proportion of each key element, and the contrast of each key element.

4. The method of claim 3, wherein, The generating, by using the eye movement data prediction model, of the page clarity for the page picture comprises: generating, by using the eye movement data prediction model, gradient amplitude, local contrast, and information entropy for the page picture; determining the page clarity according to a weighted sum result of the gradient amplitude, the local contrast, and the information entropy.

5. The method of claim 3, wherein, The generating, by using the eye movement data prediction model, of the attention focus degree for the heat map comprises: performing binarization processing on the heat map by using the eye movement data prediction model to determine attention regions in the heat map; determining a total area sum and a region center position of each of the attention regions in the heat map; determining a region area ratio value according to a ratio between the total area sum of each of the attention regions and an image area of the heat map; determining a target distance according to a distance between the region center position and a center position of a key region in the heat map; taking a weighted difference result of the region area ratio value and the target distance as the attention focus degree.

6. The method of claim 3, wherein, The generating, by using the eye movement data prediction model, of the attention proportion of each key element for the heat map comprises: segmenting, by using the eye movement data prediction model, a plurality of key element regions in the heat map, and counting a number of fixation points and a fixation time length in each of the key element regions; The number of gaze points in each of the key element regions and the gaze duration are weighted and summed to obtain the attention degree of each of the key element regions; The ratio between the attention degree of each of the key element regions and the sum of the attention degrees of all the key element regions is determined as the attention proportion of each of the key element regions.

7. The method of claim 2, wherein, The eye movement data prediction model is used to generate the focus map, the heat map, and the contrast map corresponding to the page picture, including: The eye movement data prediction model is used to generate the focus initial map, the heat initial map, and the contrast initial map corresponding to the page picture; The focus initial map, the heat initial map, and the contrast initial map are respectively modified according to preset constraint conditions to obtain the corresponding focus map, heat map, and contrast map, wherein the preset constraint conditions include at least one of standardization design conditions, user visual habit constraint conditions, and industry compliance conditions.

8. A page analysis system characterized by, The page analysis system includes: A visual analysis module is configured to, when a page picture of a page to be analyzed is obtained, generate eye movement data pictures corresponding to the page picture by using an eye movement data prediction model, and generate eye movement indicators corresponding to the to-be-recognized pictures, wherein the to-be-recognized pictures include the page picture and the eye movement data pictures; A language interpretation module is configured to generate a page analysis conclusion corresponding to the eye movement indicators by using a language understanding large model.

9. A computer device, comprising: The computer program product includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory are in communication with each other through the communication bus; The memory is configured to store a computer program; The processor is configured to execute the program stored in the memory to implement the page analysis method of any one of claims 1-7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the page analysis method of any one of claims 1-7.

11. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to implement the page analysis method of any one of claims 1-7.