Method and system for simulating attention focus area of person and predicting page churn rate

By using machine vision technology to simulate attention focus areas and predict page drop-off rates, the problem of time-consuming and labor-intensive eye trackers is solved, and rapid design suggestions are provided to improve visual design effects.

CN121523531APending Publication Date: 2026-02-13NANJING CHANGQI MIND TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411341246.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-12
Filing Date
2024-09-24
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

In existing technologies, using eye trackers to measure the viewer's attention focus area is time-consuming and laborious, and cannot serve the design process in real time, making it difficult to predict the attention focus area and page drop-off rate in visual design.

Method used

By classifying and recognizing visual input content using machine vision technology, simulating the dynamic changes of the attention focus area, and combining human-computer interaction information to calculate attention values ​​and dwell time, the dynamic change process of the attention focus area and page drop-off rate prediction are generated.

Benefits of technology

It enables rapid, real-time assessment of attention focus areas and page drop-off rates in visual designs, reduces reliance on expensive eye trackers, and provides design suggestions to improve design effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121523531A_ABST
    Figure CN121523531A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of human-computer interaction, in particular to a human attention focus area simulation and page churn rate prediction method and system, and the method comprises the steps: obtaining the visual input content and human-computer interaction information of a user; classifying and identifying the visual input content by using machine vision, and distinguishing at least one of identifiable content and unidentifiable content; simulating at least one attention focus area of a person according to the recognizable and unrecognizable contents, and calculating an attention value of the at least one attention focus area; according to the attention value, the classification result and the man-machine interaction information, the staying time and the display sequence of each attention focus area are calculated; and according to the display sequence, the retention time and the human-computer interaction information of each attention focus area, generating and outputting an attention focus area dynamic change process of the viewer, an attention focus area with relatively long accumulated retention time, an unconcerned attention focus area, an attention focus area which is not understood by the viewer, a page churn rate, viewing time and design suggestions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of human-computer interaction, and particularly relates to a method and system for simulating a focus area of attention of a person and predicting a page abandonment rate. BACKGROUND

[0002] In the era of digital information explosion, various visual information designs, product designs, or designs and evaluations of human-computer systems, including advertisements, posters, e-commerce products, web pages, APP pages, live broadcasts, industrial designs, human-computer system and user interface designs, cockpit designs, safety designs, and the like, the key to visual content design is to attract the attention of the viewer and let the viewer pay attention to the key information and not let the viewer abandon the current page. Which information will be noticed by the viewer, which information is not noticed by the viewer, where is the focus area of attention, and how to predict the page abandonment rate are problems that designers are very concerned about.

[0003] In the field of human-computer interaction, it is generally necessary to use experimental equipment such as an eye tracker to measure the changes in the focus area of attention of the viewer during the process of watching pictures or videos, but the eye tracker device is expensive, time-consuming and laborious, and cannot provide real-time service to the design process. SUMMARY

[0004] The present application provides a method and system for simulating a focus area of attention of a person and predicting a page abandonment rate, which can simulate the focus area of attention of a person when watching pictures or videos and its dynamic changes, generate a video marking the dynamic changes of the focus area of attention, a picture marking the focus area of attention including the focus area of attention of the viewer, a region that may not be noticed by the viewer, a region that may not be understood by the viewer, a predicted page abandonment rate, a predicted viewing time, and design suggestions. The viewer refers to a person who watches pictures or videos and hears possible sounds in the videos. The user of the system refers to a person using the system, including designers of videos and pictures and the like. The system simulates the changes in the focus area of attention of the viewer and the attention results, predicts the page abandonment rate and provides design suggestions, and the user of the system modifies the pictures or videos according to the information. In addition, the above results predicted by the present application are for the general viewer or a viewer with characteristics such as male viewers, rather than the prediction of the behavior results of a specific person. In the present application, the page is broadly defined to include the current page content presented to the viewer, including pictures and videos.

[0005] The first aspect of this application provides a method for simulating a user's attention focus area and predicting page drop-off rate, comprising the following steps: acquiring user visual input content and optional human-computer interaction information; classifying and recognizing the visual input content using machine vision, and distinguishing at least one identifiable content and non-identifiable content; the classification result of identifiable content includes: special interest subclasses, nameable object subclasses, and text and symbol subclasses; wherein the special interest subclasses include people, faces, human body parts, or other materials that are more likely to arouse the viewer's interest; simulating a user's attention focus area based on at least one identifiable content and non-identifiable content, and based on the identifiable content... Calculate the attention value for each attention focus area based on at least one unidentifiable content, classification results, and human-computer interaction information; calculate the dwell time and display order for each attention focus area based on the attention value, classification results, and human-computer interaction information; and generate the system output based on the display order, dwell time, and optional human-computer interaction information for each attention focus area: the dynamic simulation of the change process of the viewer's attention focus area to the visual input content, the attention focus areas with longer cumulative dwell time predicted by the system, attention focus areas that may not be noticed by the viewer, attention focus areas that the viewer may not understand, page drop-off rate, viewing time, and design suggestions.

[0006] Optionally, the attention value of each attention focus area is calculated based on at least one identifiable and non-identifiable content, the classification result, and human-computer interaction information, including: extracting features of identifiable targets and features of non-identifiable targets respectively; if the visual input content includes identifiable content, then the attention value of each attention focus area is calculated based on the features of identifiable targets, the classification result, and human-computer interaction information; if the visual input content includes non-identifiable content, then the attention value of each attention focus area is calculated based on the features of non-identifiable targets and human-computer interaction information; if the visual input content includes both identifiable and non-identifiable content, then the attention value of each attention focus area is calculated based on the features of identifiable targets, the features of non-identifiable targets, and human-computer interaction information.

[0007] Optionally, the attention value of each attention focus area is calculated based on the features of identifiable targets and non-identifiable targets, classification results, and human-computer interaction information, including: calculating the attention value of each attention focus area based on the features of identifiable targets, classification results, and human-computer interaction information; and calculating the attention value of each attention focus area based on the attraction value calculated based on the features of non-identifiable targets and human-computer interaction information.

[0008] Optionally, the features of the identifiable target include at least one of text symbol features, human body features, and object features. Based on the features of the identifiable target, the classification result, and human-computer interaction information, the attention value of each attention focus area is calculated, including: if the features of the identifiable target include text symbol features, human body features, or object features, then based on multiple text symbol features, human body features, object features, classification result, and human-computer interaction information, the attention value of each attention focus area is calculated.

[0009] The object characteristics include one or more of the following: size, viewing angle, color, sharpness, position, nameable outcome, word frequency of the nameable outcome, and flickering or movement. Position includes the current target's location and the location of the previous focus area. The text / symbol characteristics include at least one or more of the following: size, viewing angle, color, position, contrast, word frequency, and flickering or movement. Based on these characteristics, the attractiveness of the object or text / symbol is calculated, considering factors such as material sharpness, color vibrancy, flickering or movement, satisfaction of the viewer's needs, consistency between the material's language and the viewer's language, consistency between the material's knowledge and culture, age group, and the viewer's knowledge and culture, age group, consistency between the material and the viewer's preferences, novelty, aesthetic score, use of sound such as music and speech, sound clarity and quality, coordination between sound and visual stimuli, composition of the visual material, frequency of uncommon words or abbreviations in the text, and comprehensibility of the material. The factors considered include: the credibility of the material, the product price, the consistency between written and non-written materials, and whether and to what extent an authority figure is used; human characteristics including: size, viewing angle, position, color, gender, age group, race, aesthetic appeal of body parts, degree of occlusion of body parts, clarity, posture, movement, flickering or movement; based on human characteristics, the attractiveness of a particular interest subclass is calculated, taking into account factors such as the clarity and aesthetic appeal of the visual material, the degree to which the material meets the viewer's needs, the degree to which the face or body is not occluded or displayed, the age of the person, the compatibility of the person with decorations and the surrounding background, the gender, face, body, race, and the degree of consistency with the viewer's aesthetics and expectations, novelty, the use of sound, the clarity and quality of sound, the compatibility between sound and visual stimuli, the consistency between written materials and people including gender, age, occupation, expression, and movement, the credibility of the relevant written language of the person, whether and to what extent an authority figure is used, and the composition of the visual material;

[0010] Based on the features of the unidentifiable target, the classification results, and multiple human-computer interaction information, the attention value of each attention focus area is calculated, including: the features of the unidentifiable target include one or more of the following: size, position, color, outline, shape, blinking or movement; the position includes the position information of the current target and the position information of the previous attention focus area.

[0011] Based on the characteristics of unidentifiable targets, their attractiveness is calculated. Factors considered in the calculation include their clarity, the vibrancy of the colors of the visual material, the flickering or movement of the visual material, the degree to which the viewer's needs are met, the consistency between the knowledge and culture and age group involved in the material and the viewer's knowledge and culture and age group, the consistency between the material and the viewer's preferences, the degree of novelty, the aesthetic score, the use of sound such as music and speech, the clarity and quality of sound, the coordination between sound and visual stimuli, the composition of the visual material, the comprehensibility of the material, the credibility of the material, and the consistency between non-textual and textual materials.

[0012] Optionally, the dwell time of each attention focus area is calculated based on the attention value, classification result, and human-computer interaction information, including: identifying the target type within each attention focus area to obtain a classification result; determining the basic attention time of the attention focus area based on the classification result; calculating the display weight of each attention focus area based on the attention value; and determining the dwell time of each attention focus area based on its display weight, basic attention time, and attractiveness.

[0013] Optionally, based on the dwell time of each attention focus area, the number of attention focus areas, the number of non-attention focus areas, and the upper limit of browsing time defined by the viewing mode, the viewing time is calculated with and without considering viewing mode and viewing interest. This includes: for images, calculating the image viewing time without considering viewing mode and viewing interest based on the total dwell time of all attention focus areas, the unit saccade time, and the number of attention focus areas; and calculating the image viewing time considering viewing mode and viewing interest based on the current time not exceeding the upper limit of browsing time defined by the viewing mode, the number of attention focus areas, the number of non-attention focus areas, the total attractiveness of the image, and the viewing interest threshold.

[0014] For videos, the viewing time is obtained based on the video's duration without considering viewing modes and viewing interests. The video is divided into N segments, and the viewing time considering viewing modes and viewing interests is calculated based on the following factors: the current time does not exceed the maximum browsing time defined by the viewing mode, the number of attention focus areas and the number of non-attention focus areas in the current video segment, the attractiveness of the current video segment, the cumulative attractiveness of video segments viewed by the viewer before the current video segment, and the viewing interest threshold.

[0015] Optionally, before generating the output based on the display order, dwell time, classification results, and multiple human-computer interaction information for each attention focus area, the method further includes: for images, determining the first batch of unattended focus areas based on material size, clarity, position, and contrast information; for videos, determining the first batch of unattended focus areas based on material size, clarity, position, contrast information, and duration; determining the second batch of unattended focus areas based on the browsing time limit set by the viewing mode in the human-computer interaction information and considering the image viewing time or video viewing time under the viewing mode and viewing interests; and marking the first batch of unattended focus areas and the second batch of unattended focus areas respectively to obtain the unattended focus areas to be marked and displayed.

[0016] Optionally, before generating the output based on the display order, dwell time, classification results, and multiple human-computer interaction information for each attention focus area, the process further includes: determining the comprehension difficulty based on factors such as word frequency, sentence length, sentence comprehension difficulty, whether it is an abbreviation, the viewer's age and education level, and viewing mode; identifying attention focus areas that the viewer may not understand; and marking these potentially incomprehensible focus areas to obtain the target marked display of potentially incomprehensible focus areas.

[0017] Optionally, in non-purchase scenarios, the basic churn rate of the page can be predicted based on the number of unnoticed focus areas, the number of noticed focus areas, the overall attractiveness of the images or videos, whether the material includes the target content the viewer is looking for and the clarity and size of that target content, the viewer's current time, the influence of viewing patterns, and the degree of brand awareness or authority usage. Alternatively, the basic churn rate can be estimated by dividing the viewing time considering viewing patterns and interests by the viewing time without considering viewing patterns and interests. The higher the ratio of the viewing time considering viewing patterns and interests to the viewing time without considering viewing patterns and interests, the lower the basic churn rate of the page.

[0018] Optionally, in a purchase scenario, the page churn rate can be predicted based on the basic churn rate, the price displayed on the current page, the price of similar products, the viewer's economic income, other benefits obtained by the viewer, and the effort, time, or other losses required by the viewer.

[0019] A second aspect of this application provides a system for simulating a person's attention focus area and predicting page drop-off rate, comprising: an acquisition module for acquiring a user's visual input content and optional human-computer interaction information; a recognition module for recognizing at least one identifiable and unidentifiable content in the visual input content, wherein the classification result of the identifiable content includes: a special interest subclass, a nameable object subclass, and a text and symbol subclass; a calculation module for simulating a person's attention focus area based on at least one identifiable and unidentifiable content, and calculating the attention value of each attention focus area based on at least one identifiable and unidentifiable content, the classification result, and the human-computer interaction information; and a result presentation module for generating system output: the dynamic change process of the viewer's attention focus area for the visual input content, attention focus areas with a long cumulative dwell time, attention focus areas that may not have been noticed by the viewer, attention focus areas that the viewer may not understand, page drop-off rate, viewing time, and design suggestions.

[0020] Optionally, the result presentation module is further used for: marking the dwell time and order of attention focus areas; marking or overlaying attention focus areas on user-input images or videos using semi-transparent shapes such as circles; marking these areas on the dynamic video or images output by the system, and not disappearing after dwelling; the longer the cumulative dwell time, the thicker the line of the circle marking shape; the presentation order is marked with numbers at the target location of the marked circle; the system records the number of times each attention focus area is repeatedly presented, and the number of repetitions and dwell time constitute the cumulative dwell time; the vividness of the semi-transparent overlay color of the attention focus area is determined by the cumulative dwell time; the specific output of the result presentation module includes: marking attention focus areas. Videos showing dynamic changes in a designated area; slow-motion videos showing dynamic changes in a marked focus area; first image or video: an image or video marking the viewer's focus area with a note indicating the order of attention; second image or video: an image or video marking the focus area where the viewer's cumulative dwell time is relatively long; third image or video: an image or video marking areas not noticed by the viewer and focus areas that are difficult for the viewer to understand; predicted viewing time of images or videos without considering viewing patterns and viewing interests; predicted viewing time of images or videos considering viewing patterns and viewing interests; predicted page drop rate; design suggestions generated based on the system's calculation module.

[0021] Therefore, this application has the following beneficial effects:

[0022] This application, based on different visual input content including images, videos, and other human-computer interaction information, generates dynamic changes in the viewer's attention focus area, areas with longer cumulative dwell time, areas that may not be noticed by the viewer, areas that the viewer may not understand, page drop-off rate, viewing time, and design suggestions. According to these outputs of the invention, system users, especially designers, can promptly modify the visual input content, including images and videos. Specifically, this includes modifying the design to encourage the viewer to focus on the information or materials the designer wants the viewer to notice first, modifying materials that may not be noticed or understood by the viewer, reducing page drop-off rate, and modifying images or videos according to design suggestions with standards and principles. This invention can partially replace expensive eye trackers and their time-consuming and labor-intensive experiments, quickly evaluating various visual designs, or serving other systems or personnel who need to predict human attention focus or page drop-off rate.

[0023] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0024] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0025] Figure 1 This is a flowchart illustrating a method for simulating human attention focus area and predicting page drop-off rate according to an embodiment of this application;

[0026] Figure 2 This is a first form of an image or video showing the order of highlighted focus areas according to an embodiment of this application;

[0027] Figure 3 This is a second form of an image or video showing the order of highlighted focus areas according to an embodiment of this application;

[0028] Figure 4 This is a form of image or video showing a region of attention with a relatively long cumulative dwell time, according to an embodiment of this application;

[0029] Figure 5 This is one form of image or video provided according to an embodiment of the present application, which includes annotations of areas not noticed or understood by the viewer.

[0030] Figure 6 This is a flowchart illustrating a specific method for controlling the attention focus area and output page drop-off rate of a simulated human according to an embodiment of this application;

[0031] Figure 7 This is a block diagram of a system for simulating human attention focus areas and predicting page drop-off rates according to an embodiment of this application. Detailed Implementation

[0032] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0033] The following describes a method and system for simulating human attention focus areas and predicting page drop-off rate according to embodiments of this application, with reference to the accompanying drawings. Addressing the problems mentioned in the background section, this application provides a method for simulating human attention focus areas and predicting page drop-off rate. In this method, based on the user's visual input content and human-computer interaction information, the visual input content is classified and identified using machine vision. The dwell time and display order of each attention focus area are calculated. Based on the display order, dwell time, and optional human-computer interaction information of each attention focus area, the following are generated: dynamic changes in the viewer's attention focus area for the visual input content; attention focus areas with longer cumulative dwell times; attention focus areas that may not be noticed by the viewer; attention focus areas that the viewer may not understand; page drop-off rate; viewing time; and design suggestions. Based on the results of this method, users, especially designers, can promptly modify the visual input content, including images and videos. Specifically, the modified design can encourage viewers to focus on information or materials the designer wants them to notice first; materials that are not noticed or not understood by the viewer can be modified; page drop-off rate can be reduced; and images or videos can be modified according to design suggestions with standards and principles. This invention can partially replace expensive eye trackers and their time-consuming experiments, quickly evaluating various visual designs, or serving other systems or personnel that need to predict a person's attention focus or page drop-off rate.

[0034] Specifically, Figure 1 This is a flowchart illustrating a method for simulating human attention focus area and predicting page drop-off rate according to an embodiment of this application.

[0035] like Figure 1 As shown, the method for simulating human attention focus areas and predicting page drop-off rate includes the following steps:

[0036] In step S101, the user's visual input content and optional human-computer interaction information are obtained.

[0037] The visual input can be images, videos, etc., which are collected and preprocessed. During actual execution, error messages are displayed for images or videos that do not meet system format requirements. If the visual input is a video containing audio, this embodiment will automatically transcribe the audio into text for further analysis of the impact of the audio content on attention.

[0038] Furthermore, embodiments of this application can also obtain optional user input, as follows:

[0039] 1) Text materials: such as WORD files, PPT files, etc.

[0040] 2) Optional definitions for images, viewers, and viewing modes include: the dimensions of the image or video, the approximate distance of the viewer from the screen, the viewer's gender (general, male, female), age group (general, young and middle-aged, elderly, minors), race, the viewer's viewing mode (quick browsing, careful reading, emergency, unlimited time), the viewer's familiar language and cultural level, whether the image or video is used for purchasing (yes, no), the viewer's approximate economic status (only applicable to purchasing scenarios), and the current time the viewer is viewing the image or video. Viewing modes include: Quick browsing is suitable for situations where the viewer can roughly view the image or video in a non-emergency situation, such as entertainment, online shopping, or quickly finding content of interest in the image or video. Careful reading refers to situations where the viewer wants to obtain detailed information or learn from the image or video in a non-emergency situation. Emergency refers to situations where the viewer may be under time pressure or have an urgent need to view the image, scene, or video; this can be used to assess whether the viewer can quickly find the target control in the visual scene presented in the current image under emergency conditions. For example, it can assess whether a driver can quickly locate the hazard light button on a car's human-machine interface with numerous control buttons. System users can also modify the time limit for viewers to view images or videos in each viewing mode. Unlimited time means there is no time limit for viewers to view images or videos.

[0041] Furthermore, if the user inputs a large amount of information, this application embodiment can provide default values ​​for the above optional inputs. For example, the default value for the viewer's gender is the overall value, and the default value for the viewer's age group is the overall value. Users can modify the default values ​​as needed.

[0042] In step S102, at least one of the identifiable and unidentifiable content in the visual input content is identified, that is, divided into two major categories: identifiable content and unidentifiable content. The identifiable content objects are further divided into three subcategories: the special interest subcategory includes people, faces, human body parts or other materials that are more interesting to the viewer, the nameable object subcategory, and the text and symbol subcategory.

[0043] For identifiable objects within identifiable content, information includes their position, size (pixel value), color (RGB value), naming result (e.g., if there's a chair in the image, the object is named "chair"), word frequency of the naming result, and any flickering or movement. For people, faces, or other materials that are of interest to the viewer, the system identifies the person's gender (male, female, unclear), age group (middle-aged / elderly, young, child), race, and body parts including the aesthetic appeal, color, degree of occlusion, clarity, posture, and any flickering or movement. For unidentifiable content, information includes size, color, outline, shape, position, and any flickering or movement.

[0044] In step S103, the attention value of each attention focus area is calculated based on at least one simulated human attention focus area of ​​identifiable and unidentifiable content, and based on at least one identifiable and unidentifiable content, the classification result, and human-computer interaction information.

[0045] In one embodiment of this application, calculating the attention value of each attention focus area based on at least one identifiable and non-identifiable content, a classification result, and human-computer interaction information includes: extracting features of identifiable targets and features of non-identifiable targets respectively; if the visual input content includes identifiable content, calculating the attention value of each attention focus area based on the features of identifiable targets, a classification result, and human-computer interaction information; if the visual input content includes non-identifiable content, calculating the attention value of each attention focus area based on the features of non-identifiable targets and human-computer interaction information; if the visual input content includes both identifiable and non-identifiable content, calculating the attention value of each attention focus area based on the features of identifiable and non-identifiable targets, a classification result, and human-computer interaction information.

[0046] It is understood that embodiments of this application can utilize machine vision to extract features of each identifiable object in an image, including at least one of text symbol features, human body features, and object features. For features of unidentifiable targets, embodiments of this application can use methods such as edge detection, contour recognition, or local texture analysis to extract them.

[0047] In one embodiment of this application, the attention value of each attention focus area is calculated based on the features of the identifiable target, the classification result, and the human-computer interaction information. This includes: if the features of the identifiable target include text symbol features, human body features, and object features, then the attention value of each attention focus area is calculated based on the text symbol features, human body features, object features, the classification result, and the human-computer interaction information; if the features of the identifiable target include text symbol features, then the attention value of each attention focus area is calculated based on the text symbol features and the human-computer interaction information.

[0048] The features of textual symbols can include: size, viewing angle, color, position, contrast, word frequency, sentence comprehension difficulty, whether it is an abbreviation, and whether it flickers or moves. Human body features can include: size, viewing angle, position, gender, age group, race, aesthetic appeal of body parts including face and hair, degree of occlusion of body parts, sharpness, posture, and whether it flickers or moves. Object features can include: size, viewing angle, color, sharpness, position, nameable results, word frequency of nameable results, interaction with the human body or other objects, and whether it flickers or moves.

[0049] The embodiments of this application can be divided into the following three cases based on the size of the recognition result:

[0050] Scenario 1: For named object subclasses and text and symbol subclasses, the attention value (A score) of these target attention focus areas is updated as follows:

[0051] A1 = A + W × Loc × ATNH × g1

[0052] Where W represents the area of ​​the nameable object subclass material or text and symbol subclass material, Loc is the position fraction, ATNH is the attraction value, and g is a constant that can be adjusted based on training samples. A1 is the updated attention value after calculation, and A is the attention value of the object or text symbol before calculation (initial value is 0). It should be noted that multiple lowercase letters in subsequent embodiments (such as p, q, h, i, m, and g, etc.) are all constants and will not be explained further.

[0053] The position fraction (Loc) in the formula is determined by position sub-fraction 1 (Loc1) and position sub-fraction 2 (Lo2): Loc = p1 × Loc1 + p2 × Loc2.

[0054] In the formula, the position sub-score 1 (Loc1) is determined by the current target's position within the entire image or video. The closer the target is to the center of the image or video, or to the top or left of the image or video, the higher the position sub-score 1 (Loc1). The position sub-score 2 (Loc2) is the current target's position coordinates (X). i Y i ) and the position coordinates (X) of the previous focus area of ​​attention i-1 Y i-1 The reciprocal of the distance between the current focus of attention and the previous focus of attention, i.e., the smaller the distance between the coordinates of the current focus of attention and the previous focus of attention, the higher the value of Loc2. The specific formula for calculation is:

[0055] Loc2=1 / (|X i -X i-1 |+|Y i -Yi-1 |)

[0056] In the formula, the attractiveness (ATNH) of a nameable object or word and symbol is determined by the clarity of the material (R), the vividness of the colors of the visual material (C), the flickering or movement of the visual material (BM), the degree to which it meets the viewer's needs (E, especially whether the material includes the target content the viewer is looking for, such as objects, words, or similar materials, and the clarity and size of these target contents), the consistency between the language of the material and the viewer's language (LC), the consistency between the knowledge, culture, and age group involved in the material and the viewer's knowledge, culture, and age group (CC), the consistency between the material and the viewer's preferences (PC), the degree of novelty (SW), the aesthetic score (B), and the sound, such as music and... The following factors determine the effectiveness of visual stimuli: the use of speech (S), clarity and quality of sound (SQ), coordination between sound and visual stimuli (SV), composition of visual material (ST) (whether and to what extent six common compositional techniques—symmetry, central composition, leading lines, rule of thirds, framing, and diagonal—are used), sentence length of text (SL), frequency of uncommon words or abbreviations in text (LF), comprehensibility of material (EU), credibility of material (TR), product price (PR) (in a purchasing scenario, if the viewer's income is average and the price is low, the h value is positive), consistency between textual and non-textual materials (TNC), and whether and to what extent authority is used (P). The specific formula is:

[0057] ATNH=h1×R+h2×C+h3×BM+h4×E+h5×LC+h6×CC+h7×PC+h8×SW+h9×B+h10×S+h11×SQ+h12×SV+h13×ST-h14×SL-h15×LF+h16×EU+h17×TR-h18×PR+h19×TNC+h20×P

[0058] For example, without considering other factors, objects or words that are large, brightly colored, or located in the upper half of an image or video will receive a higher score (A).

[0059] Scenario 2: For the subcategories of particular interest, including people, faces, body parts, or other materials that are of particular interest to the viewer, the updated value of Score A for people, faces, or other materials of particular interest to the viewer is:

[0060] A1 = A + T × Loc × ATH × g2

[0061] Where T is the area size of the material of the special interest subclass; Loc is the position (Loc), which includes the position information of the current special interest subclass and the position information of the previous attention focus area. The specific calculation method is the same as in case one.

[0062] ATH is the attractiveness of a particular interest subcategory. ATH is determined by the clarity (R) of the visual material, the vibrancy of the colors (C), the flickering or movement of the visual material (BM), the aesthetic appeal (B, which can be negative), the degree to which the material meets the viewer's needs (E), the degree to which the face or body is not obscured or displayed (N), the age of the characters (A: adults and young people score higher), the harmony between the characters and decorations and the surrounding background (M), the novelty (SW), the use of sound such as music and speech (S), the clarity and quality of sound (SQ), the harmony between sound and visual stimuli (SV), the composition of the visual material (ST) (whether and to what extent the six common composition techniques of symmetry, central composition, leading lines, rule of thirds, framing, and diagonal are used), the consistency between the textual material and the characters, including gender, age, occupation, expression, and actions (THC), the credibility of the related text or language of the characters (TR), and whether an authoritative figure is used and the degree of authority of the figure (P).

[0063] ATH=h1×R+h2×C+h3×BM+h4×B+h5×E+h6×N-h7×A+h8×M+h9×SW+h10×S+h11×SQ+h12×SV+h13×ST+h14×THC+h15×TR+h16×P

[0064] Scenario 3: For unidentifiable objects, the formula for calculating the updated value of their A score is:

[0065] A1=A+W2×Loc×ATNH2×g3

[0066] Where W2 is the area size of the unidentifiable object material. Loc is the location (Loc), which includes the current location information of the unidentifiable object and the location information of the previous attention focus area. The specific calculation method is the same as in Case 1. The attractiveness of the unidentifiable object ATNH2 is determined by its clarity (R), the color vibrancy of the visual material (C), the flickering or movement of the visual material (BM), the degree to which it meets the viewer's needs (E, especially whether the material includes the target content that the viewer wants to find), the consistency between the knowledge, culture, and age level involved in the material and the viewer's knowledge, culture, and age level (CC), the consistency between the material and the viewer's preferences (PC), the degree of novelty (SW), the aesthetic score (B), the use of sound such as music and speech (S), the clarity and quality of sound (SQ), the coordination between sound and visual stimuli (SV), the composition of the visual material (ST) (whether and to what extent the six common composition techniques of symmetry, central composition, leading lines, rule of thirds, framing, and diagonal are used), the comprehensibility of the material (EU), the credibility of the material (TR), and the consistency between non-textual and textual materials (TNC). The formula is:

[0067] ATNH2=t1×R+t2×C+t3×BM+t4×E+t5×CC+t6×PC+t7×SW+t8×B+t9×S+t10×SQ+t11×SV+t12×ST+t13×EU+t14×TR+t15×TNC.

[0068] In step S104, based on the attention value, classification results, and human-computer interaction information, the dwell time of each attention focus area is calculated. Based on the display order and dwell time of each attention focus area, the following are generated: dynamic change process of the viewer's attention focus area for visual input content, attention focus areas with longer cumulative dwell time, viewing time, attention focus areas that may not be noticed by the viewer, attention focus areas that the viewer may not understand, page drop-off rate, and design suggestions.

[0069] It is understood that, based on attention values, classification results, and human-computer interaction information, this application's embodiments calculate the dwell time of each attention focus area, and generate, according to the display order and duration, the dynamic change process of the viewer's attention focus area for visual input content, attention focus areas with longer cumulative dwell times, attention focus areas that may not be noticed by the viewer, attention focus areas that the viewer may not understand, page drop-off rate, viewing time, and design suggestions. This enables page design to more effectively attract the viewer's attention to where the designer wants to attract attention, thereby reducing page drop-off rate and improving conversion rate.

[0070] In one embodiment of this application, the dwell time of each attention focus area is calculated based on the attention value, classification result, and human-computer interaction information, including: identifying the target type in each attention focus area to obtain a classification result; determining the basic attention time of the attention focus area based on the classification result; calculating the display weight of each attention focus area based on the attention value, and then determining the dwell time of each attention focus area based on its display weight, basic attention time, and attraction level.

[0071] It is understood that the embodiments of this application identify the target type within each attention focus area to obtain a classification result, such as text, image, video, person, or interactive element. Different classification results may have different base attention times. The display weight of each attention focus area is calculated based on the attention value. The higher the attention value, the greater the display weight, meaning that the area is more likely to be given priority attention by the viewer.

[0072] Furthermore, embodiments of this application can also determine the display order. Based on attention values ​​and machine recognition classification results, this application categorizes attention focus areas into priority groups and non-priority groups. The display order of attention focus areas is typically that those in the priority group are displayed first, followed by those in the non-priority group. There can be multiple priority groups, and multiple non-priority groups. All attention focus areas within the same group are displayed from left to right and from top to bottom according to their coordinate positions. This includes determining when and for how long each area is displayed to simulate the viewer's attention focus area changes and attention outcomes, including which areas receive long-term attention and which are not noticed.

[0073] according to Figure 6 Step S104 of the embodiments of this application is described in detail in sections S201-S208 as follows:

[0074] Step 1: Identify the first batch of unnoticed focal areas.

[0075] For each attention focus area obtained from step S103, if the user-input visual content is an image, the size, clarity, position, and contrast of the visual content in that attention focus area are used to determine whether it should be included in the first batch of unattended attention focus areas. If the visual content area of ​​an attention focus area is too small, the contrast (text or characters) is too low, or the material is unclear (e.g., the height pixel value of an object, text, or outline is less than a certain pixel threshold or the viewing angle is less than a certain angle (viewing angle threshold), then these attention focus areas will not proceed to the processing in steps 2 through 8 of step S104. The dynamic video output by this system and Figures 2-4 These areas of focus are not displayed in the output. However, these areas of focus that are not noticed by the viewer, referred to as unnoticed areas of focus, will be displayed in the system's output. Figure 5 The above is displayed and labeled. Unnoticed areas entering the output of this system... Figure 5 The images are marked with bright circles and text such as "not noticed," "possibly not noticed," or "viewer may not have noticed." The pixel value threshold and viewing angle threshold for older viewers are greater than those for other age groups.

[0076] If the user inputs a video, the first sub-step follows the same logic as for images, determining whether a particular attention focus area should be included in the first batch of unattended attention focus areas based on its size, clarity, position, and contrast. The second sub-step considers the duration of each attention focus area's visual content throughout the video. If the duration of a particular attention focus area's visual content is less than the basic attention duration, this attention focus area will be included in the first batch of unattended attention focus areas and output as system output. Figure 5The annotation is then performed. These are categorized into the first batch of unnoticed focus areas, and the focus areas are not processed in steps 2 through 8 of this S104 step.

[0077] Step 2: Group the remaining focus areas from Step 1.

[0078] Based on the attention score (A score) calculated by the computation module, priority groups and non-priority groups are formed by setting a difference threshold or using a clustering algorithm (such as K-means). A priority group represents one or more attention focal areas that the viewer prioritizes. It should be noted that: 1) There can be multiple priority groups, and also multiple non-priority groups. The following analysis uses one priority group and one non-priority group as an example. 2) Similar materials should be grouped together whenever possible. For example, text of similar size and position is generally grouped together. The same method applies to image or video input; however, for videos, the logic algorithm for each of the N segments divided in step 1 is executed.

[0079] Step 3: Based on the grouping results in Step 2, the attention value (A score), and the material classification results, determine the dwell time of each attention focus area on the dynamic video output by this system.

[0080] The dwell time of each attention focus area on text and symbols is referred to as the text dwell time; this dwell time is also the time interval at which the mark of that attention focus area is presented on the dynamic video output by this system. The dwell time of each attention focus area on a nameable object is referred to as the nameable object dwell time. The naming method for the dwell time of other materials is the same as above.

[0081] There are two specific methods, as follows:

[0082] Method 1:

[0083] 1) Text dwell time = Priority group weight × (A score of the current focus area / the largest A score of the focus area in this group) × (its base attention time (td seconds) + the time determined by its attractiveness (j1 × ATNH) ± residual variation ε). For example, priority group weight = 1, non-priority group weight = 0.5, td = 0.25 seconds, and ε varies within the range of 0-0.1 seconds. For example, the priority group weight is 1, and the non-priority group weight is 0.5.

[0084] 2) Nameable object dwell time = Priority group weight × (Current focus area A score / Maximum focus area A score in this group) × (Base attention time (od seconds) + Time determined by its attraction level (j2 × ATNH) ± ε). For example, od = 0.3 seconds. The priority group weight can be 1, and the non-priority group weight can be 0.5, or it can be set according to the actual situation.

[0085] 3) The dwell time on a face, person, or other material that is more interesting to the viewer = priority group weight × (current focus area A score / the highest A score of the focus area in this group) × (its basic attention time (hd seconds) + the time determined by its attractiveness (j3 × ATH) ± ε). Where hd = 0.6 seconds, which can also be set according to the actual situation.

[0086] 4) Dwell time for unrecognizable objects = Priority group weight × (A score of the current focal area / the largest A score of the focal area in this group) × (its base dwell time (ud) seconds + the time determined by its attraction level (j4×ATNH2) ± ε seconds). Where ud = 0.3 seconds, which can also be set according to the actual situation, as shown in the example in Table 1.

[0087] It should be noted that the dwell time in the focal area should not be less than its baseline attention time. For example, the minimum dwell time for text is 0.2 seconds ± residual variation ε; for named objects, it's 0.2 seconds ± ε; for faces, people, or other materials that are more interesting to the viewer, it's 0.2 seconds ± ε; and for unidentifiable objects, it's 0.1 seconds ± ε. See Table 1 for examples below. These minimum dwell times can be adjusted based on the priority group, the viewer's age group, and education level.

[0088] Table 1

[0089]

[0090]

[0091] Method 2:

[0092] 1) Text dwell time = Priority group weight × (its base attention time (td seconds) - (the ranking of the target A score in this group - 1) × decreasing constant δ + the time determined by its attractiveness (j1 × ATNH) ± residual variation ε). For example, td = 0.25 seconds, ε varies in the range of 0-0.1 seconds. For example, the weight of the priority group is 1, and the weight of the non-priority group is 0.5. The decreasing constant δ determined by the ranking of A scores (e.g., 0.01 seconds): If the focus area is ranked 1st in this priority group, then this decreasing constant is (1-1) × δ = 0; if it is ranked 2nd, then (2-1) × decreasing constant is δ = δ (δ for example, 0.01 seconds); if it is ranked 3rd, then it is (3-1) × δ = 2δ, and so on.

[0093] 2) The basic dwell time of a nameable object = priority group weight × (its basic attention time (od seconds) - (the ranking of the target A score in this group A score - 1) × decreasing constant δ + the time determined by its attractiveness (j2 × ATNH) ± ε). For example, od = 0.3 seconds. For example, the priority group weight is 1, and the non-priority group weight is 0.5.

[0094] 3) The dwell time on a face, person, or other material that is more interesting to the viewer = priority group weight × (its basic attention time (hd seconds) - (the ranking of the target A score in the group A score - 1) × decreasing constant δ + the time determined by its attractiveness (j3 × ATH) ± ε). For example, hd = 0.6 seconds.

[0095] 4) Dwell time for unidentifiable objects = Priority group weight × (its base dwell time (ud) seconds - (the ranking of the target A score in this group A score - 1) × decreasing constant δ + time determined by its attractiveness (j4×ATNH2) ± ε seconds). For example, ud = 0.3 seconds.

[0096] The input of images or videos follows the same method as above, except that videos need to be divided into N segments according to step 1, and the logical algorithm of this step is run for each segment.

[0097] In addition, because the human eye's actual attention span is very short, in order to allow users of this system to clearly see the dynamic changes in the focus area of ​​attention, the video that marks the dynamic changes in the focus area of ​​attention can have a slow motion option.

[0098] Step 4: Determine the dynamic presentation order of the focus areas.

[0099] All focus areas within the same priority group are displayed sequentially from left to right and from top to bottom according to their coordinate positions. The magnitude of the attention value only affects the dwell time and does not affect the dynamic presentation order, as shown in the example in Table 1.

[0100] In addition, based on the above steps including the operation of Table 1, the embodiments of this application also include the following four logics and designs:

[0101] 1) Marking the focus area and its duration and sequence: This refers to marking or overlaying the focus area on the user-input image or video. Markers such as semi-transparent circles can be used, remaining on the system-output dynamic video or image and not disappearing after the focus period. The longer the focus period, the thicker the circle. The presentation order is marked with 1, 2, 3, 4, etc., at a specific location on the circle, such as the upper right corner. Figure 2 As shown. Optionally, the circles can be connected by arrows, as shown. Figure 3 As shown.

[0102] 2) If the viewer is an elderly person, the above presentation time will be appropriately extended.

[0103] 3) If the word frequency of the text or the word frequency of the words naming the namable objects is low, the dwell time will be appropriately extended. 4) For a certain segment of a picture or video, if the current time does not exceed the browsing time limit determined by the viewing mode or it is predicted that the user has not jumped out due to loss of viewing interest, the system will repeatedly run the calculation of the dwell time and display order of each attention focus area, including jumping between each attention focus area, until the current time reaches the browsing time limit determined by the viewing mode, it is predicted that the user has jumped out due to loss of viewing interest or the video segment ends. The system will record the number of times each attention focus area is repeatedly presented.

[0104] Accumulated dwell time = number of repetitions of being noticed × dwell time

[0105] The accumulated dwell time determines the vividness of the semi-transparent overlay color of the attention focus area (the higher the number of repetitions × dwell time, the more vivid the color), approaching the heat map of the eye movement experiment results, as Figure 4 shown.

[0106] 5) The input of pictures or videos follows the same method above, except that for videos, according to the N segments divided in step 1, each segment runs the logical algorithm of this step.

[0107] Step 5, calculate the viewing time.

[0108] The viewing time of the viewer for pictures and videos is predicted through the following logic. The dwell time of the two batches of non-attention focus areas is not calculated and not included in the viewing time.

[0109] Step 5.1, obtain the browsing time limit from the viewing mode (quick browsing, careful reading, emergency situation, and time unlimited) of the viewer obtained from the human-computer interaction information. Among them, the browsing time limits for quick browsing, careful reading, and emergency situation are ST, CT, and ET respectively. Generally, ET < ST < CT. There is no upper limit in the time unlimited mode. The specific time can be set as the default value by the system or defined by the user.

[0110] Step 5.2, calculate the viewing time without considering the viewing mode and viewing interest, that is, the browsing time limit:

[0111] Without considering the viewing mode and viewing interest, that is, the browsing time limit,

[0112] The viewing time of a picture = total dwell time of all attention focus areas (see step 3) + unit saccade time (0.03 - 0.05 seconds) × number of attention focus areas.

[0113] The viewing time of a video = the time of the video itself.

[0114] Step 5.3, calculate the viewing time considering viewing mode and viewing interests, i.e., the maximum viewing time limit:

[0115] Taking into account viewing modes and viewing interests, i.e., the maximum browsing time limit.

[0116] 5.3.1 Calculation logic for image viewing time:

[0117] If condition 1 (i.e., the current time has not exceeded the maximum browsing time defined by the viewing mode) is met, and condition 2 is also met, where condition 2 is at1 × number of attention focus areas / (number of non-attention focus areas + number of attention focus areas) + at2 × total image attraction (w1 × ∑ATH + w2 × ∑ATNH + w3 × ∑ATNH2) >= viewing interest threshold, then the above steps continue. If either or both of the above conditions are not met, the loop is exited, i.e., viewing stops, and the exit time is recorded. This exit time is the image viewing time considering viewing mode and viewing interest as output by the system.

[0118] 5.3.2 Calculation logic for video viewing time:

[0119] Based on the system's division of the video into N segments in step 1, the system will start from the first segment (K=1) and make judgments based on the following four sub-cases using condition 1 and the new condition 2. For each video segment:

[0120] 1) The viewer has time to watch, but is not interested in the current segment; that is, condition 1 is met, i.e., the current time has not exceeded the browsing time limit defined by the viewing mode, but new condition 2 is not met:

[0121] at1 × Number of attention-focused regions in the current video segment / (Number of non-attention-focused regions + Number of attention-focused regions)

[0122] +at2 × the attractiveness of the current video segment (w1 × ∑ATH + w2 × ∑ATNH + w3 × ∑ATNH2) +at3 × the cumulative attractiveness of the video segments previously viewed by the viewer ∑(w1 × ∑ATH + w2 × ∑ATNH + w3 × ∑ATNH2 < viewing interest threshold, then the system can immediately exit (i.e., new condition 2 is not met for the first time), or jump to another video segment, until condition 1 is not met (time limit expires), new condition 2 is not met for the Mth time, or the last video segment finishes running. The exit time is the video viewing time considering viewing mode and viewing interest output by the system. Depending on the viewer's situation, M can be set to 1, 2, or 3, etc., with a default value of 1, meaning the system will exit if the viewer is not interested in the first segment.

[0123] 2) The viewer is interested in the current segment, but the time limit has expired; that is, when condition 1 is not met, but new condition 2 is met, the upper limit of the browsing time defined by the viewing mode is the video viewing time output by the system that takes into account the viewing mode and viewing interest.

[0124] 3) When the viewer is not interested in the current segment and the time limit has expired; that is, condition 1 is not met and new condition 2 is not met, the system predicts that the viewer will choose to skip the video. The upper limit of the browsing time defined by the viewing mode is the video viewing time output by the system that takes into account the viewing mode and viewing interest.

[0125] 4) If the viewer is interested in the current segment and the time has not yet expired; that is, when condition 1 is met and new condition 2 is also met, the system will move on to the next segment of the video and repeat steps 1)-4 above.

[0126] The viewing time, both considering and not considering viewing modes and viewing interests, will be used as the system's output.

[0127] Step 6: Identify the second batch of unnoticed focus areas (those that were not noticed due to time constraints in the viewing mode or because the viewer lost interest).

[0128] Based on the time constraints of the viewing mode and the viewing time, this step identifies areas that were missed due to the time constraints of the viewing mode, or areas that were overlooked because the viewer lost interest. These unnoticed focus areas are output as the second batch of unnoticed focus areas. The specific algorithm is as follows:

[0129] Step 6.1, based on step 5, obtain the viewing time of images and videos considering viewing modes and viewing interests.

[0130] Step 6.2: Identify the second batch of unnoticed focus areas.

[0131] 6.2.1 Images

[0132] In the quick browsing mode, the system takes the minimum of the following two values: Value 1: the maximum browsing time (ST) in this viewing mode and Value 2: the image viewing time obtained in step 5.2 considering the viewing mode and viewing interests, i.e.:

[0133] Minimum value = Min(ST, viewing time considering viewing mode and viewing interests)

[0134] The system operates according to step 3 (as shown in Table 1). When the system reaches the minimum value, the attention focus area automatically stops, and these attention focus areas not reached by the system are classified into the second batch of unattended attention focus areas. The system needs to record whether the current attention focus is due to ST reaching its minimum value, or whether the minimum value is reached considering the viewing mode and viewing interest. The system then uses bright red circles and similar text labels in the output, such as "unattended," "possibly unattended," "viewer may not have paid attention due to viewing mode browsing time limitations," "viewer may not have paid attention due to loss of interest," and "viewer did not pay attention due to viewing mode browsing time limitations or loss of interest." Figure 5 (As shown).

[0135] The other two viewing modes use the same logic, only the maximum viewing time is different.

[0136] 6.2.2 Video

[0137] The same logic applies to each video segment. If the system exits at segment Q (meaning the viewer stops watching the entire video at segment Q), then the second batch of unnoticed focus areas includes all attentional focus areas within the segments from segment Q onwards until the end of the video. For example, if a video has 200 segments (N=200), and the system predicts that the viewer will exit at segment 35 due to the viewing mode's time limit being reached or insufficient viewing interest (the system needs to record the reason for stopping), then all attentional focus areas within segments 36-200 will be considered the system's second batch of unnoticed focus areas. In the system's output video, these are marked with bright red circles and similar text such as "unnoticed," "possibly unnoticed," "viewer may have been unnoticed due to viewing mode's time limit," "viewer may have been unnoticed due to loss of interest," or "viewer may have been unnoticed due to viewing mode's time limit or loss of interest." Figure 5 (As shown).

[0138] Step 7: Identify the areas of focus that viewers may not understand.

[0139] For text or symbols, the difficulty of comprehension will be determined based on factors such as word frequency, sentence length, sentence comprehension difficulty, whether it is an abbreviation, the viewer's age and education level, and viewing mode. If the difficulty of comprehension exceeds a certain threshold, this embodiment of the application can mark it as incomprehensible and circle it, such as... Figure 5 As shown, the circled area can be a dark red or similar color, along with text such as "may not understand", "does not understand", or "viewers may not understand".

[0140] For non-textual objects or outlines, the difficulty of understanding will be determined based on the word frequency of the object's naming result (in the case where it can be named), the similarity between the object or outline and real-world objects or existing system objects (in the case where it cannot be named), the viewer's age, education level, and viewing mode. If the difficulty of understanding exceeds a certain threshold, these areas can also be labeled in this application embodiment.

[0141] The input of images or videos follows the same method as above, except that videos need to be divided into N segments according to step 1, and the logical algorithm of this step is run for each segment.

[0142] Step 8: Calculate the page drop-off rate.

[0143] Based on the above steps, the system calculates the page churn rate according to the following logic.

[0144] 1) In non-purchase scenarios

[0145] Method 1:

[0146] The basic drop rate of a page is calculated as follows: DropRateBasic = dr1 × (number of unnoticed focus areas / (number of unnoticed focus areas + number of noticed focus areas)) - dr2 × total attractiveness of images or videos (w1 × ∑ATH + w2 × ∑ATNH + w3 × ∑ATNH2) + dr3 × whether the material includes the target content that the viewer is looking for, and the clarity and size of this target content + dr4 × (number of noticed focus areas - threshold value; and total attractiveness < a certain threshold) + dr5 × the influence of the viewer's current time + dr6 × the influence of viewing mode - dr7 × the degree of use of brand awareness or authority.

[0147] Wherein: Number of focus areas / (Number of non-focus areas + Number of focus areas): This calculates the proportion of non-critical information in the image or video. A higher proportion indicates more non-critical information, which will increase the page drop rate; ∑ATH, ∑ATNH, and ∑ATNH2 represent the sum of the attractiveness of the focus areas of these three types of materials.

[0148] Does the material include the content the viewer is looking for, and what is the clarity and size of that content? If the material includes the content the viewer is looking for, and that the content is clear and large, then its impact on the basic page drop-off rate is negative. Conversely, it is positive.

[0149] Note that the number of focus areas minus the threshold value, and the overall attractiveness < a certain threshold: If there are too many focus areas, the amount of information per unit area is too large, but the overall attractiveness of the images or videos is not high, it will increase the page drop-off rate.

[0150] The impact of the viewer's current time: If the viewer's current time is during the weekend, holiday, or after get off work, the impact of the viewer's current time on the basic page churn rate is negative.

[0151] The impact of viewing modes: For fast browsing and emergency modes, the impact on the basic churn rate is positive; for careful browsing and unlimited time modes, the impact on the churn rate is negative.

[0152] Method 2:

[0153] The page's basic churn rate (DropRateBasic) = (1 - dra × (viewing time considering viewing patterns and interests / viewing time without considering viewing patterns and interests)) + drb × the degree of influence of the viewer's current time.

[0154] For example, assuming dra=1 and drb=0, for a certain image, the viewing time considering viewing mode and viewing interest is 2.1 seconds, and the viewing time without considering viewing mode and viewing interest is 3 seconds. This indicates that the image received a lot of attention from the viewer, and the basic churn rate of the page is 1 - 1 × (2.1 / 3) = 30%. Conversely, if the viewing time considering viewing mode and viewing interest is only 0.3 seconds, and the viewing time without considering viewing mode and viewing interest is 3 seconds, this indicates that the image received a lot of attention from the viewer, and the basic churn rate of the page is 1 - 1 × (0.3 / 3) = 90%. The calculation method for video is the same as above, and will not be elaborated here.

[0155] 2) In the purchasing scenario:

[0156] Sub-scenario 1: If the product price information in the image or video enters the focus area of ​​attention.

[0157] DropRatePurchase = DropRateBasic + p1 × (current page displayed price / similar product price) + p2 × (1 / viewer's economic income) - p3 × other benefits gained by the viewer + p4 × effort, time or other losses required by the viewer.

[0158] Sub-case 2: If the product price information in the image or video does not enter the focus area, DropRatePurchase = DropRateBasic.

[0159] Step 8: Import design standards and principles and output design suggestions.

[0160] Based on the design standards and principles, and the results of the preceding steps, output relevant design adjustment suggestions. For example:

[0161] Example 1: Please check the predicted priority focus area for the viewer. If the system's predicted area matches the area you want the viewer to focus on, then the design is effective. If they don't match, you need to consider adjusting the size, contrast, and color vibrancy of text or objects in these focus areas. You can also adjust their positions, such as moving them from the bottom to the top of an image or video, from the right side to the center, or from a corner to closer to the center. After making changes, you can run the system again to check the results.

[0162] Example 2: Some areas of an image or video may go unnoticed by viewers in the current browsing mode. This could be because the text, icons, or objects in these areas are too small or located at the bottom of the image or video. To make these elements more visible, increase the font size, enlarge the objects, or adjust their positions. After making these changes, run the program again to see the results.

[0163] Example 3: The design standards for the size of some text in images or videos are inconsistent. It is recommended to increase the font size of these texts.

[0164] Example 4: The image contains too many colors. If the image is a user interface, it may not be consistent with design standards. It is recommended to reduce the number of colors in the design.

[0165] Example 5: A high page drop-off rate for images may be due to: A) Viewers may not have noticed the key information being presented. Please check the predicted priority focus areas for viewers. If the system's predicted area matches the area you want viewers to focus on, the design is effective. If not, consider adjusting the size, contrast, and color vibrancy of text or objects in these focus areas. You can also adjust their positions, such as moving them from the bottom to the top of the image or video, from the right to the center, or from a corner to closer to the center. After modification, run the system again to check the results. B) A high page drop-off rate for the image or video may be due to insufficient material to attract viewers' attention, too much information displayed, or a lack of aesthetic appeal. Purchase scenario: The price in the current design is significantly higher than similar products or significantly higher than the viewer's purchasing power, or the price information in the current design is not being noticed by viewers. It is recommended to increase the font size of the price information, adjust its position, or use more vibrant colors. C) Analysis results of other factors in the above algorithms.

[0166] Example 6: The current system predicts that viewers will only watch the uploaded video for 2.3 seconds, which is less than the 10-second duration of the video itself. This indicates that viewers may have lost interest or experienced time constraints in their chosen viewing mode from the beginning of the video. You can change the viewing mode to unlimited time and run the system again. If the system's predicted viewing time is still less than the video's actual duration, you can further modify the video, including its beginning, by adding more engaging content and information, especially in the first half of the video.

[0167] Example 7: The current system predicts that viewers will only watch the uploaded image for 0.1 seconds, indicating that viewers are not very interested in the image, the content is relatively short, or the selected viewing mode has a time limit. If the image contains only this much content and the purpose is for users to complete the viewing in a short time, no further modification is needed. If this is not the case, you can change the viewing mode to unlimited time and run the system again. If the system's predicted viewing time is still very short, it is recommended to further modify the image, adding more content that can attract viewers' interest and increasing information that can grab their attention.

[0168] The method for simulating human attention focus areas and predicting page drop-off rate proposed in this application can classify and identify visual input content using machine vision based on user visual input content and human-computer interaction information, calculate the dwell time and display order of each attention focus area, and generate the dynamic change process of the viewer's attention focus area, attention focus areas with longer cumulative dwell time, attention focus areas that may not be noticed by the viewer, attention focus areas that the viewer may not understand, page drop-off rate, viewing time, and design suggestions based on the display order, dwell time, and optional human-computer interaction information of each attention focus area. Based on the results of this method, users of the system, especially designers, can promptly modify visual input content, including images and videos. Specifically, the modified design can encourage viewers to pay attention to information or materials that the designer wants them to notice first, modify materials that have not been noticed by the viewer or that the viewer may not understand, reduce page drop-off rate, and modify images or videos according to design suggestions with standards and principles.

[0169] Next, referring to the accompanying drawings, a system for simulating human attention focus areas and predicting page drop-off rates according to embodiments of this application is described.

[0170] Figure 7 This is a block diagram illustrating a human attention focus area simulation and page drop-off rate prediction system according to an embodiment of this application.

[0171] like Figure 7As shown, the system for simulating a person's attention focus area and predicting page drop-off rate includes: an acquisition module 10, an identification module 20, a calculation module 30, and a result presentation module 40.

[0172] The acquisition module 10 is used to acquire the user's visual input content and optional human-computer interaction information. The visual input content includes images or videos. The optional human-computer interaction information includes: the dimensions of the image or video, the approximate distance between the viewer and the screen, the viewer's gender, age, ethnicity, familiar language and cultural level, the viewer's viewing mode (including quick browsing, careful reading, emergency, and unlimited viewing time, which the system user can modify), whether the image or video is used for a purchase scenario, the viewer's approximate economic status, and the current time the viewer is viewing the image or video. The recognition module 20 is used to recognize at least one identifiable and non-identifiable content in the visual input content. The classification results of the identifiable content include: special interest subclasses, nameable object subclasses, and text and symbol subclasses. The special interest subclasses include people, faces, human body parts, or other materials that are of interest to the viewer. The calculation module 30 is used to simulate the user's attention focus area based on at least one identifiable and non-identifiable content; and to calculate based on at least one identifiable and non-identifiable content... The system calculates the attention value for each attention focus area based on multiple classification results and human-computer interaction information. It then calculates the dwell time for each attention focus area based on the attention value, classification results, and human-computer interaction information. Finally, it generates output results based on the display order, dwell time, classification results, and / or human-computer interaction information for each attention focus area. The output results include at least one of the following: the dynamic change process of the user's attention focus area to the visual input content; attention focus areas with longer cumulative dwell times; attention focus areas not noticed by the viewer; attention focus areas that the viewer does not understand; page drop-off rate; viewing time; and design suggestions. For a segment of an image or video, if the current time has not exceeded the maximum viewing time determined by the viewing mode, or if it is predicted that the user will not leave due to loss of viewing interest, the system will repeatedly run the calculations for the dwell time and display order of each attention focus area, including jumping between attention focus areas, and accumulating the dwell time for each attention focus area until the current time reaches the maximum viewing time determined by the viewing mode, the user is predicted to leave due to loss of viewing interest, or the video segment ends.

[0173] The results presentation module 40 is also used to mark the dwell time and sequence of attention focus areas. It marks or overlays attention focus areas on user-input images or videos using semi-transparent shapes such as circles, which remain on the dynamic video or images output by the system and do not disappear after the dwell time. The longer the cumulative dwell time, the thicker the line of the circle. The presentation order is marked with numbers on the target location of the circle. The system records the number of times each attention focus area is repeatedly presented. The number of repetitions and the dwell time are the cumulative dwell time. The vividness of the semi-transparent overlay color of the attention focus area is determined by the cumulative dwell time.

[0174] The specific outputs of the 40-module result presentation include: a video showing dynamic changes in the marked focus area; a slow-motion video showing the dynamic changes in the marked focus area; a first image or video: an image or video with markers indicating the viewer's focus area and attention sequence; a second image or video: an image or video showing the focus area where the viewer's cumulative dwell time is relatively long; a third image or video: an image or video showing areas not noticed by the viewer and focus areas that are difficult for the viewer to understand; predicted viewing time of the image or video without considering viewing patterns and viewing interests; predicted viewing time of the image or video considering viewing patterns and viewing interests; predicted page drop rate; and design suggestions generated based on the system's calculation module.

[0175] It should be noted that the explanations and descriptions of the above-mentioned embodiments of the method for simulating human attention focus area and predicting page churn rate also apply to the human attention focus area simulation and page churn rate prediction system of this embodiment, and will not be repeated here.

[0176] The system for simulating human attention focus areas and predicting page drop-off rate, as proposed in this application, can classify and identify visual input content using machine vision based on user visual input and human-computer interaction information. It calculates the dwell time and display order of each attention focus area. Based on the display order, dwell time, and optional human-computer interaction information of each attention focus area, it generates the dynamic change process of the viewer's attention focus area, attention focus areas with longer cumulative dwell times, attention focus areas that may not be noticed by the viewer, attention focus areas that the viewer may not understand, page drop-off rate, viewing time, and design suggestions. Based on these outputs, users, especially designers, can promptly modify visual input content, including images and videos. Specifically, modified designs can encourage viewers to focus on information or materials the designer wants them to notice first, modify materials that may not be noticed or understood by the viewer, reduce page drop-off rate, and modify images or videos according to design suggestions with standards and principles. This invention can partially replace expensive eye trackers and their time-consuming experiments, quickly evaluate various visual designs, or serve other systems or personnel who need to predict human attention focus or page drop-off rate.

[0177] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0178] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0179] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0180] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (FPGAs), field-programmable gate arrays (FPGAs), etc.

[0181] Those skilled in the art will understand that all or part of the steps of the methods implementing the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0182] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A method for simulating human attention focus areas and predicting page churn rate, characterized in that, Includes the following steps: Acquire user visual input and optional human-computer interaction information; Identify at least one of the identifiable and non-identifiable content in the visual input content, wherein the classification result of the identifiable content includes: special interest subclass, nameable object subclass, and text and symbol subclass; Based on at least one simulated human attention focus area of ​​the identifiable content and the unidentifiable content, and based on at least one of the identifiable content and the unidentifiable content, the classification result of the identifiable content, and multiple human-computer interaction information, the attention value of each attention focus area is calculated; Based on the attention value, the classification result, and multiple human-computer interaction information, the dwell time of each attention focus area is calculated; based on the display order of each attention focus area, the dwell time, the classification result, and multiple human-computer interaction information, an output result is generated; wherein, the output result is at least one of the following: the dynamic change process of the user's attention focus area to the visual input content, attention focus areas with longer cumulative dwell time, attention focus areas not noticed by the viewer, attention focus areas not understood by the viewer, page drop-off rate, viewing time, and design suggestions.

2. The method for simulating human attention focus areas and predicting page drop-off rate according to claim 1, characterized in that, Based on at least one of the identifiable content and the non-identifiable content, the classification result of the identifiable content, and multiple human-computer interaction information, the attention value of each attention focus area is calculated, including: Extract features from identifiable targets and features from non-identifiable targets respectively; If the visual input content includes the identifiable content, then the attention value of each attention focus area is calculated based on the features of the identifiable target, the classification result, and multiple human-computer interaction information. If the visual input content includes unrecognizable content, then the attention value of each attention focus area is calculated based on the characteristics of the unrecognizable target and multiple human-computer interaction information. If the visual input content includes both identifiable and non-identifiable content, then the attention value of each attention focus area is calculated based on the features of the identifiable target, the features of the non-identifiable target, the classification result, and multiple human-computer interaction information.

3. The method for simulating human attention focus areas and predicting page drop-off rate according to claim 2, characterized in that, The attention value of each attention focus area is calculated based on at least one of the features of the identifiable target, including text symbol features, human body features, and object features, the features of the unidentifiable target, the classification result, and multiple human-computer interaction information, including: If the features of the identifiable target include text symbol features, human body features, or object features, then the attention value of each attention focus area is calculated based on multiple of the text symbol features, the human body features, the object features, the classification results, and the human-computer interaction information. The object features include one or more of the following: size, viewing angle, color, sharpness, position, nameable result, word frequency of the nameable result, and blinking or movement; the position includes the position information of the current target and the position information of the previous attention focus area; the text symbol features include at least one or more of the following: size, viewing angle, color, position, contrast, word frequency, and blinking or movement. The attractiveness of an object, text, or symbol is calculated based on its characteristics. Factors considered in the calculation include: clarity of the material; vibrancy of the colors of the visual material; flickering or movement of the visual material; degree to which the viewer's needs are met; consistency between the language of the material and the viewer's language; consistency between the knowledge, culture, and age group involved in the material and the viewer's knowledge, culture, and age group; consistency between the material and the viewer's preferences; novelty; aesthetic score; use of sound such as music and speech; clarity and quality of sound; coordination between sound and visual stimuli; composition of the visual material; frequency of rare words or abbreviations in the text; ease of understanding of the material; credibility of the material; product price; consistency between textual and non-textual materials; and whether authority is used and the degree of authority. The human body characteristics include one or more of the following: size, viewing angle, position, color, gender, age group, race, aesthetic quality of body parts, degree of occlusion of body parts, clarity, posture, movement, flickering or movement. Based on the aforementioned human characteristics, the attractiveness of a particular interest subclass is calculated. Factors considered in the calculation include the clarity of the visual material, its aesthetic appeal, the degree to which the material meets the viewer's needs, the degree to which the face or body is not obscured or displayed, the age of the person, the compatibility of the person with decorations and the surrounding background, the gender of the person, and the degree of consistency between the face, body, race and the viewer's aesthetics and expectations, the degree of novelty, the use of sound, the clarity and quality of sound, the compatibility between sound and visual stimuli, the consistency between textual material and the person, including gender, age, occupation, expression, and actions, the credibility of the person's related written language, whether an authoritative figure is used and the degree of authority of the figure, and the composition of the visual material. Based on the features of the unidentifiable target, the classification result, and multiple pieces of human-computer interaction information, the attention value of each attention focus area is calculated, including: The characteristics of the unidentifiable target include one or more of the following: size, position, color, outline, shape, flickering or movement. The position includes the current target's position information and the position information of the previous attention focus area. Based on the characteristics of the unidentifiable target, its attractiveness is calculated. Factors considered in the calculation include its clarity, the vibrancy of the visual material's colors, the flickering or movement of the visual material, the degree to which it meets the viewer's needs, the consistency between the knowledge and culture and age group involved in the material and the viewer's knowledge and culture and age group, the consistency between the material and the viewer's preferences, novelty, aesthetic score, the use of sound such as music and speech, the clarity and quality of sound, the coordination between sound and visual stimuli, the composition of the visual material, the comprehensibility of the material, the credibility of the material, and the consistency between non-textual and textual materials.

4. The method for simulating human attention focus areas and predicting page drop-off rate according to claim 1, characterized in that, Based on the attention value, the classification result, and multiple sets of human-computer interaction information, the dwell time for each attention focus area is calculated, including: The display weight of each attention focus area is calculated based on the attention value, and the dwell time of each attention focus area is determined based on the display weight, the basic attention time, and the attraction level; wherein, the basic attention time is determined based on the classification result, and the dwell time of each focus attention area is greater than or equal to the basic attention time of its category.

5. The method for simulating human attention focus areas and predicting page drop-off rate according to claim 1, characterized in that, Based on the dwell time in each attention focus area, the number of attention focus areas, the number of non-attention focus areas, and the upper limit of browsing time defined in the viewing mode in the human-computer interaction information, the viewing time is calculated with and without considering viewing mode and viewing interest, including: For images, the image viewing time is calculated without considering viewing mode and viewing interest, based on the total dwell time of all attention focus areas, the unit saccade time, and the number of attention focus areas; the image viewing time considering viewing mode and viewing interest is calculated based on the current time not exceeding the browsing time limit defined by the viewing mode, the number of attention focus areas, the number of unattended focus areas, the total attractiveness of the image, and the viewing interest threshold. For videos, the viewing time is obtained based on the video's duration without considering viewing modes and viewing interests. The video is divided into N segments, and the viewing time considering viewing modes and viewing interests is calculated based on the following factors: the current time does not exceed the maximum browsing time defined by the viewing mode, the number of attention focus areas and the number of non-attention focus areas in the current video segment, the attractiveness of the current video segment, the cumulative attractiveness of video segments viewed by the viewer before the current video segment, and the viewing interest threshold.

6. The method for simulating human attention focus areas and predicting page drop-off rate according to claim 1, characterized in that, Before generating the output result based on the display order of each attention focus area, the dwell time, the classification result, and multiple pieces of human-computer interaction information, the process further includes: For images, the first batch of unnoticed focus areas are determined based on material size, sharpness, location, and contrast information; For video, the first batch of unnoticed focus areas are determined based on material size, sharpness, location, contrast information, and duration; Based on the upper limit of browsing time set by the viewing mode in the human-computer interaction information and the image viewing time or video viewing time under the consideration of viewing mode and viewing interest, the second batch of unnoticed focus areas are determined. The first batch of unnoticed focus areas and the second batch of unnoticed focus areas are marked respectively to obtain the unnoticed focus areas displayed by the target marking.

7. The method for simulating human attention focus areas and predicting page drop-off rate according to claim 1, characterized in that, Based on the word frequency, sentence length, sentence comprehension difficulty, whether it is an abbreviation, the viewer's age and education level, and multiple comprehension difficulty factors of the viewing mode, determine the focus area of ​​attention that the viewer does not understand. The viewer marks the incomprehensible focus area, resulting in the target marked incomprehensible focus area.

8. The method for simulating human attention focus area and predicting page drop-off rate according to claim 1, characterized in that, In non-purchase scenarios, the basic churn rate of a page can be predicted based on the number of unnoticed focus areas, the number of noticed focus areas, the overall attractiveness of images or videos, whether the material includes the target content the viewer is looking for, the clarity and size of the target content, the viewer's current time, the influence of viewing patterns, and the degree of brand awareness or authority usage. Alternatively, the basic churn rate can be estimated by dividing the viewing time considering viewing patterns and interests by the viewing time not considering viewing patterns and interests, along with the viewer's current time. Assuming the viewer's current time has no impact on the page churn rate, the higher the ratio of the viewing time considering viewing patterns and interests to the viewing time not considering viewing patterns and interests, the lower the basic churn rate of the page. In a purchase scenario, the churn rate of a page in a purchase scenario is predicted based on the basic churn rate, the price displayed on the current page, the prices of similar products, the viewer's economic income, other benefits the viewer receives, and the effort or time the viewer needs to expend.

9. A system for simulating human attention focus areas and predicting page drop-off rate, characterized in that, include: The acquisition module is used to acquire the user's visual input content and multiple optional human-computer interaction information. The visual input content includes images or videos. The multiple optional human-computer interaction information specifically include: the length and width of the image or video, the approximate distance between the viewer and the screen, the viewer's gender, age group, race, familiar language and cultural level, the viewer's viewing mode (including quick browsing, careful reading, emergency, and unlimited time, and the system user can modify the viewer's viewing time), whether the viewer's image or video is used for a purchase scenario, the viewer's approximate economic income, and the current time the viewer is viewing the image or video. The recognition module is used to recognize at least one of the identifiable and unidentifiable content in the visual input content. The classification results of the identifiable content include: special interest subclass, nameable object subclass, and text and symbol subclass. The special interest subclass includes people, faces, human body parts, or other materials that are more likely to arouse the viewer's interest. A calculation module is configured to: simulate a human's attention focus area based on at least one of the identifiable and unidentifiable content; calculate an attention value for each attention focus area based on at least one of the identifiable and unidentifiable content, a classification result, and multiple human-computer interaction information; calculate the dwell time for each attention focus area based on the attention value, the classification result, and multiple human-computer interaction information; and generate an output result based on the display order of each attention focus area, the dwell time, the classification result, and / or multiple human-computer interaction information; wherein the output result is at least one of the following: the dynamic change process of the user's attention focus area to the visual input content, attention focus areas with longer cumulative dwell times, attention focus areas not noticed by the viewer, attention focus areas not understood by the viewer, page drop-off rate, viewing time, and design suggestions; For a segment of an image or video, if the current time has not exceeded the maximum browsing time determined by the viewing mode or if it is predicted that the user will not leave due to loss of viewing interest, the system will repeatedly run the calculation of the dwell time and display order of each attention focus area, including jumping between each attention focus area, and accumulating the dwell time of each attention focus area until the current time reaches the maximum browsing time determined by the viewing mode, the user is predicted to leave due to loss of viewing interest, or the video segment ends. The results presentation module is used to display at least one of the following: the dynamic changes in the viewer's focus area of ​​visual input content, the focus areas with the longest cumulative dwell time, the focus areas that were not noticed by the viewer, the focus areas that the viewer did not understand, the page drop-off rate, the viewing time, and design suggestions.

10. The system for simulating human attention focus areas and predicting page drop-off rates according to claim 9, characterized in that, The result presentation module is also used for: Pay attention to the marking of the focus area's dwell time and order. Mark or overlay the focus area on the user-input image or video, using a semi-transparent shape such as a circle to mark it, and leave it on the dynamic video or image output by the system. After leaving it, it will not disappear. The longer the cumulative dwell time, the thicker the circle-shaped line becomes, and the order of appearance is marked with numbers on the target location marked by the circle. The system records the number of times each attention focus area is repeatedly presented. The number of repetitions and the dwell time are the cumulative dwell time. The vividness of the semi-transparent overlay color of the attention focus area is determined by the cumulative dwell time. The specific output of the result presentation module includes: Mark videos where the focus area changes dynamically; Present slow-motion video showing the dynamic changes in the marked focus area; The first image or video: an image or video that marks the viewer's focal area and includes markers indicating the order of attention; The second image or video: an image or video that marks the area of ​​focus of the viewer's attention for a longer period of time; The third image or video: an image or video that highlights areas not noticed by the viewer and areas of focus that are difficult for the viewer to understand; Predicted viewing time for images or videos without considering viewing patterns and viewing interests; The predicted viewing time for images or videos takes into account viewing patterns and viewing interests; Predicted page churn rate; Design suggestions generated based on the system's computation module.