Brand visual element-oriented interest region division and gaze analysis method and system

CN122799088APending Publication Date: 2026-09-22BEIJING JIANSHU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611043219.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

现有技术多依赖于事后手动修正兴趣区域边界,再依据固定阈值判定用户是否关注特定区域,致使分析结论与用户实际认知路径存在偏差

Benefits of technology

[0061]利用视觉语义分割自动识别品牌视觉内容中的关键区域,结合目标用户多维特征画像与历史眼动数据的映射,形成基于用户认知偏好的个性化兴趣区域划分方案。该方案突破传统划分局限,使品牌标识、产品展示等核心元素根据用户注视习惯动态聚焦,显著提升兴趣区域与实际关注点的匹配精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799088A_ABST
    Figure CN122799088A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing and eye movement analysis, and particularly relates to a brand visual element-oriented interest region division and gaze analysis method and system. The method obtains image data of brand visual content, user multi-dimensional feature portrait and historical eye movement data set, performs visual semantic segmentation on the image to generate a candidate interest region set, establishes an association mapping model of the user feature portrait and the candidate interest region to predict personalized gaze probability and filter to form a personalized interest region set, extracts matched gaze trajectory samples to perform time series mining to obtain a typical gaze mode, and generates a brand visual optimization strategy based on the personalized interest region set and the typical gaze mode. The present application can individualize predict user gaze behavior and provide decision support for layout adjustment of brand visual elements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of image processing and eye-tracking analysis technology, and in particular to a method and system for interest region segmentation and gaze analysis of brand visual elements. Background Technology

[0002] In the field of brand visual communication, traditional methods of dividing regions of interest and analyzing gaze typically rely on fixed rules or manual annotation. Designers statically divide visual content into regions based on pre-defined categories such as brand logos, product images, or text information. This division method lacks a deep understanding of the semantic connotations of images and is difficult to adapt to the dynamic changes of brand visual elements in complex backgrounds. After eye-tracking data is collected, analysts often use statistical averaging or clustering algorithms to process gaze coordinates and duration, generating global heatmaps or gaze density maps. While these methods can reflect commonalities within a group, they fail to capture the differences in gaze behavior among individual users due to differences in age, gender, consumption preferences, or brand awareness. Historical eye-tracking data is often treated as an independent sample set, used only to calculate the average gaze duration or number of jumps in a region, without establishing a systematic correlation between the user's multidimensional characteristics and their corresponding gaze patterns. Existing technologies often rely on manually correcting the boundaries of regions of interest afterward and then determining whether a user is paying attention to a specific region based on fixed thresholds, leading to discrepancies between the analysis conclusions and the user's actual cognitive path. Furthermore, brand visual optimization strategies are typically based on rules of thumb or A / B testing results, lacking quantitative support regarding users' actual gaze shift patterns and dwell times. Optimization efforts are often limited to superficial parameters such as adjusting color contrast or element size. This crude analytical approach fails to provide differentiated layout optimization suggestions for different user groups, and the optimization iteration cycle is lengthy, making it difficult to meet the urgent needs of precise communication and personalized reach in brand marketing. Summary of the Invention

[0003] This invention provides a method and system for dividing interest regions and analyzing gaze patterns for brand visual elements, which can solve the problems in the prior art.

[0004] A first aspect of this invention provides a method for segmenting interest regions and performing gaze analysis on brand visual elements, comprising:

[0005] Acquire image data of brand visual content, multidimensional feature profile data of target users, and historical eye-tracking datasets;

[0006] The image data is subjected to visual semantic segmentation to identify the brand logo area, product display area and auxiliary information area, and a set of candidate interest areas and their spatial location descriptions are generated.

[0007] Based on the fixation point coordinate sequence and dwell time sequence in the historical eye-tracking dataset, an association mapping model between user feature profile data and candidate interest region set is established.

[0008] The multidimensional feature profile data is input into the association mapping model to predict the personalized gaze probability distribution of the target user for each candidate region of interest, and a set of personalized regions of interest is formed by filtering according to the gaze probability threshold.

[0009] The gaze trajectory samples that match the multidimensional feature profile data are extracted from the historical eye-tracking dataset. Temporal sequence mining is performed on the gaze trajectory samples to identify the shifting patterns and dwell patterns of the gaze points within the personalized interest region set, and to generate a typical gaze pattern description of the target user.

[0010] Based on the personalized interest region set and the typical gaze pattern description, a brand visual optimization strategy is generated for the target user. The optimization strategy specifies the layout adjustment scheme of each brand visual element in the image data.

[0011] Visual semantic segmentation is performed on the image data to identify brand logo areas, product display areas, and auxiliary information areas, generating a set of candidate regions of interest and their spatial location descriptions, including:

[0012] Multi-scale convolutional feature extraction is performed on image data to obtain a multi-level feature pyramid containing spatial detail features and semantic abstract features;

[0013] Based on the multi-level feature pyramid, a cross-scale feature fusion representation is constructed, and the spatial detail features and the semantic abstract features are weighted and fused in the channel dimension to generate a fused feature map.

[0014] Pixel-level classification prediction is performed on the fused feature map, and semantic labels of brand identification category, product display category, auxiliary information category or background category are assigned to each pixel to generate a semantic segmentation mask. Morphological connectivity analysis is performed on the semantic segmentation mask to identify pixel sets with the same semantic labels and spatial connectivity, and the bounding boundary of each pixel set is extracted to form a preliminary region boundary set.

[0015] For regions with irregular boundary shapes in the initial region boundary set, boundary refinement is performed based on the gradient direction consistency constraint of the boundary pixels to align the boundary with the actual outline of the brand visual elements, thereby generating a refined region boundary set.

[0016] For each region within the refined region boundary set, calculate the vertex coordinates of the minimum bounding rectangle, the pixel coordinates of the region centroid, and the proportion of the region area to the total area of ​​the image data. Combine the vertex coordinates of the minimum bounding rectangle, the pixel coordinates, and the proportion to form a candidate region of interest set and its spatial location description.

[0017] Pixel-level classification prediction is performed on the fused feature map, and a semantic label is assigned to each pixel, which can be categorized as brand identity, product display, auxiliary information, or background. A semantic segmentation mask is generated, and morphological connectivity analysis is performed on the semantic segmentation mask to identify a set of pixels with the same semantic label and spatial connectivity, including:

[0018] Pixel-level feature vector decoding is performed on the fused feature map to generate a four-dimensional category response vector for each pixel, which includes brand identity category, product display category, auxiliary information category and background category;

[0019] The four-dimensional category response vector is subjected to cross-category competitive suppression processing. By calculating the relative intensity difference between the response values ​​of each category, the category with the highest response value is strengthened and the responses of other categories are suppressed, thus generating a category response vector after competitive suppression.

[0020] Based on the maximum response class in the category response vector after competition suppression, a corresponding semantic label is assigned to each pixel, and the semantic labels are organized according to the pixel spatial location to form a semantic segmentation mask;

[0021] The semantic segmentation mask is grown by connected component growth based on eight-neighbor topology. An initial seed pixel is selected from the semantic segmentation mask. Pixels with the same semantic label and spatial eight-neighbor adjacency as the initial seed pixel are iteratively added to the same connected component until no new pixels can be added, thus forming the first connected component.

[0022] The connected component growth operation is repeatedly performed on the remaining pixels in the semantic segmentation mask that are not marked as the first connected component, generating multiple non-overlapping connected components in sequence. The pixels in each connected component have the same semantic label and satisfy the spatial connectivity constraint. The set of pixel position coordinates in each connected component is taken as the set of pixels with the same semantic label and spatial connectivity.

[0023] Based on the fixation point coordinate sequence and dwell time sequence in the historical eye-tracking dataset, the association mapping model between user feature profile data and candidate interest region set is established as follows:

[0024] Extract the gaze point coordinate sequence and dwell time sequence of each user from the historical eye-tracking dataset, determine the spatial region affiliation of the gaze point coordinate sequence, map each gaze point coordinate to the corresponding region in the candidate region of interest set, and generate a gaze point and region affiliation table.

[0025] Based on the attribution table and the dwell time sequence, calculate the total dwell time and gaze shift frequency of each user in the brand identification area, product display area and auxiliary information area, and combine the total dwell time and gaze shift frequency to form a user gaze behavior feature vector;

[0026] Temporal dependency modeling is performed on the user gaze behavior feature vector to extract the dynamic change trend of the jump path pattern and dwell time between different regions of the gaze point, and to generate a user behavior pattern representation containing temporal context information.

[0027] Demographic features and consumption behavior features in user profile data are jointly embedded with the user behavior pattern representation to construct a unified feature space that integrates static user attributes and dynamic eye-tracking behavior.

[0028] In the unified feature space, a multi-output regression mapping relationship is established from user feature profile data to a set of candidate interest regions. The multi-output regression mapping relationship takes user feature profile data as input and outputs the user's gaze probability distribution and expected dwell time distribution for each region in the set of candidate interest regions, forming an association mapping model.

[0029] The multidimensional feature profile data is input into the association mapping model to predict the personalized gaze probability distribution of the target user for each candidate region of interest, and a set of personalized regions of interest is formed by filtering based on the gaze probability threshold, including:

[0030] The multidimensional feature profile data is processed by feature dimension alignment and numerical scaling transformation to generate standardized feature vectors that are compatible with the input interface of the association mapping model.

[0031] The standardized feature vector is input into the association mapping model. The standardized feature vector is then subjected to nonlinear transformation and feature space projection through the multi-layer mapping structure of the association mapping model. The original gaze tendency score of the target user for each region in the candidate interest region set is then output.

[0032] The original gaze tendency score is adjusted globally by calculating the variance and kurtosis of the original gaze tendency score across all candidate regions of interest. When the variance exceeds a preset dispersion range, the score is corrected by centralization. When the kurtosis exceeds a preset centralization range, the score is corrected by broadening, thus generating a gaze score with adjusted distribution.

[0033] The distributed fixation scores are subjected to probability transformation mapping, and the distributed fixation scores of each region are converted into personalized fixation probability distributions that satisfy probability constraints through a monotonically increasing transformation relationship.

[0034] The cumulative distribution function is calculated based on the personalized gaze probability distribution. A probability value corresponding to a preset cumulative probability is selected on the cumulative distribution function as the gaze probability threshold. The gaze probability threshold is used to perform saliency screening on each region in the personalized gaze probability distribution, and regions with gaze probabilities higher than the gaze probability threshold are retained to form a personalized interest region set.

[0035] From the historical eye-tracking dataset, gaze trajectory samples matching the multidimensional feature profile data are extracted. Temporal sequence mining is performed on these gaze trajectory samples to identify the shifting patterns and dwell patterns of gaze points within the personalized region of interest set. A typical gaze pattern description of the target user is generated, including:

[0036] A subset of historical users is selected from the historical eye-tracking dataset, where the user features and multidimensional feature profiles meet the similarity matching conditions in different dimensions. The time-series records of gaze point coordinates and dwell time corresponding to the historical user subset are extracted and combined to form gaze trajectory samples.

[0037] The region attribution of the gaze point coordinate time sequence records in the gaze trajectory sample is determined, and each gaze point is mapped to the corresponding region or non-interest background region in the personalized interest region set to generate a region identification time sequence.

[0038] A state transition matrix is ​​constructed for the time sequence of the region identifiers, the number of gaze jumps from one region to another is counted, the transition probability between each region pair is calculated, and an inter-regional transition probability matrix is ​​formed as a quantitative representation of the transition pattern.

[0039] Based on the time-series records of dwell time, the statistical distribution of dwell time of the gaze point in each region within the set of personalized interest regions is calculated, and the mode value of dwell time and the dwell time dispersion index of each region are extracted.

[0040] The mode value of dwell time and the dispersion index of dwell time are used to perform pattern clustering to identify two types of dwell behavior features: stable dwell pattern and fluctuating dwell pattern. The stable dwell pattern and the fluctuating dwell pattern are then associated with the corresponding regions in the set of personalized interest regions.

[0041] The inter-regional transition probability matrix and the dwell behavior features associated with each region are structured and organized to construct a composite gaze behavior descriptor that includes the region jump probability distribution and the region dwell pattern type, which serves as a typical gaze pattern description for the target user.

[0042] Based on the personalized interest region set and the typical gaze pattern description, a brand visual optimization strategy is generated for the target user. The optimization strategy specifies the layout adjustment scheme for each brand visual element in the image data, including:

[0043] Spatial coordinates and area parameters of each region are extracted from the set of personalized interest regions. The inter-regional transition probability matrix and the statistical distribution of dwell time in each region are extracted from the description of typical gaze patterns to construct a fusion representation of regional spatial attributes and gaze behavior characteristics.

[0044] Based on the inter-regional transition probability matrix, the gaze attractiveness index of each region is calculated, and the regions in the personalized interest region set are prioritized according to the gaze attractiveness index to generate a gaze value hierarchy structure.

[0045] Identify brand logo elements, product display elements, and auxiliary information elements in image data, and extract the current spatial layout coordinates and visual hierarchy attributes of each brand's visual elements;

[0046] Construct a spatial matching metric between brand visual elements and the gaze value hierarchy. Calculate the Euclidean distance between the current spatial layout coordinates of each brand visual element and the spatial center point of each area in the gaze value hierarchy. Combine the visual hierarchy attribute of the brand visual element with the gaze attraction index of the corresponding area to generate a layout coordination score.

[0047] For brand visual elements whose layout coordination score is lower than the coordination threshold, a layout reconstruction is initiated. The area with the highest priority in the gaze value hierarchy and which is not occupied is selected as the target layout area. The position adjustment vector from the current spatial layout coordinates to the target layout area is calculated.

[0048] The position adjustment vectors of each brand's visual elements are combined with the spatial constraint parameters of the target layout area to form a layout adjustment scheme, which is then used as a brand visual optimization strategy.

[0049] A second aspect of this invention provides a system for segmenting and analyzing interest regions and gaze patterns for brand visual elements, comprising:

[0050] The image acquisition unit is used to acquire image data of brand visual content, multi-dimensional feature profile data of target users, and historical eye-tracking datasets.

[0051] The semantic segmentation unit is used to perform visual semantic segmentation on the image data, identify the brand logo area, product display area and auxiliary information area, and generate a set of candidate interest areas and their spatial location descriptions.

[0052] The mapping modeling unit is used to establish a correlation mapping model between user feature profile data and candidate interest region set based on the fixation point coordinate sequence and dwell time sequence in the historical eye movement dataset.

[0053] The prediction and filtering unit is used to input the multidimensional feature profile data into the association mapping model, predict the personalized gaze probability distribution of the target user for each candidate interest region, and filter to form a set of personalized interest regions based on the gaze probability threshold.

[0054] The pattern mining unit is used to extract gaze trajectory samples that match the multidimensional feature profile data from the historical eye-tracking dataset, perform time-series sequence mining on the gaze trajectory samples, identify the shifting patterns and dwell patterns of gaze points within the personalized interest region set, and generate a typical gaze pattern description of the target user.

[0055] An optimization generation unit is used to generate a brand visual optimization strategy for the target user based on the personalized interest region set and the typical gaze pattern description. The optimization strategy specifies the layout adjustment scheme of each brand visual element in the image data.

[0056] A third aspect of the present invention provides an electronic device, comprising:

[0057] processor;

[0058] Memory used to store processor-executable instructions;

[0059] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0060] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0061] By leveraging visual semantic segmentation to automatically identify key areas in brand visual content, and combining this with mapping of multi-dimensional feature profiles of target users with historical eye-tracking data, a personalized interest region segmentation scheme based on user cognitive preferences is formed. This scheme breaks through the limitations of traditional segmentation, enabling core elements such as brand logos and product displays to dynamically focus according to user gaze habits, significantly improving the matching accuracy between interest regions and actual points of attention.

[0062] By constructing a personalized set of interest regions through gaze probability thresholding, only areas highly correlated with the target user's gaze tendency are retained, effectively eliminating redundant visual information and reducing cognitive load. Based on the gaze probability distribution prediction results, differentiated brand visual layouts are customized for different users, realizing a shift from passive observation to proactive guidance, and improving the targeting and efficiency of brand information delivery.

[0063] By performing time-series mining on gaze trajectory samples matched with user profiles, we can extract the shift patterns and dwell patterns of gaze points between personalized interest areas, forming typical gaze patterns that reflect user browsing habits. This pattern reveals the flow and dwell points of user attention, providing data-driven optimization for the arrangement order and visual weight allocation of brand visual elements, ensuring that key elements are in advantageous positions along the user's natural eye path.

[0064] The generated brand visual optimization strategy directly outputs specific layout adjustment schemes for image elements, including optimization suggestions for position, size, and order. This ensures that the presentation of brand logos, product displays, and auxiliary information precisely matches the gaze behavior characteristics of the target users. This strategy significantly improves brand awareness speed and retention rate, reduces attention loss in visual communication, and achieves adaptive optimization of brand visual content for different user groups. Attached Figure Description

[0065] Figure 1 This is a flowchart illustrating the interest region segmentation and gaze analysis method for brand visual elements according to an embodiment of the present invention.

[0066] Figure 2 This is a flowchart illustrating the generation process of a brand visual optimization strategy based on eye-tracking data, as described in an embodiment of the present invention. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0068] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0069] Figure 1 This is a flowchart illustrating the interest region segmentation and gaze analysis method for brand visual elements according to an embodiment of the present invention.

[0070] Methods for segmenting interest regions and analyzing gaze patterns for brand visual elements include:

[0071] Acquire image data of brand visual content, multidimensional feature profile data of target users, and historical eye-tracking datasets;

[0072] The image data is subjected to visual semantic segmentation to identify the brand logo area, product display area and auxiliary information area, and a set of candidate interest areas and their spatial location descriptions are generated.

[0073] Based on the fixation point coordinate sequence and dwell time sequence in the historical eye-tracking dataset, an association mapping model between user feature profile data and candidate interest region set is established.

[0074] The multidimensional feature profile data is input into the association mapping model to predict the personalized gaze probability distribution of the target user for each candidate region of interest, and a set of personalized regions of interest is formed by filtering according to the gaze probability threshold.

[0075] The gaze trajectory samples that match the multidimensional feature profile data are extracted from the historical eye-tracking dataset. Temporal sequence mining is performed on the gaze trajectory samples to identify the shifting patterns and dwell patterns of the gaze points within the personalized interest region set, and to generate a typical gaze pattern description of the target user.

[0076] Based on the personalized interest region set and the typical gaze pattern description, a brand visual optimization strategy is generated for the target user. The optimization strategy specifies the layout adjustment scheme of each brand visual element in the image data.

[0077] In one optional implementation, visual semantic segmentation is performed on the image data to identify brand logo areas, product display areas, and auxiliary information areas, generating a set of candidate regions of interest and their spatial location descriptions, including:

[0078] Multi-scale convolutional feature extraction is performed on image data to obtain a multi-level feature pyramid containing spatial detail features and semantic abstract features;

[0079] Based on the multi-level feature pyramid, a cross-scale feature fusion representation is constructed, and the spatial detail features and the semantic abstract features are weighted and fused in the channel dimension to generate a fused feature map.

[0080] Pixel-level classification prediction is performed on the fused feature map, and semantic labels of brand identification category, product display category, auxiliary information category or background category are assigned to each pixel to generate a semantic segmentation mask. Morphological connectivity analysis is performed on the semantic segmentation mask to identify pixel sets with the same semantic labels and spatial connectivity, and the bounding boundary of each pixel set is extracted to form a preliminary region boundary set.

[0081] For regions with irregular boundary shapes in the initial region boundary set, boundary refinement is performed based on the gradient direction consistency constraint of the boundary pixels to align the boundary with the actual outline of the brand visual elements, thereby generating a refined region boundary set.

[0082] For each region within the refined region boundary set, calculate the vertex coordinates of the minimum bounding rectangle, the pixel coordinates of the region centroid, and the proportion of the region area to the total area of ​​the image data. Combine the vertex coordinates of the minimum bounding rectangle, the pixel coordinates, and the proportion to form a candidate region of interest set and its spatial location description.

[0083] When performing visual semantic segmentation on image data, a multi-scale convolutional feature extraction strategy is adopted, using the outputs of multiple layers of the convolutional neural network as feature sources. Shallow convolutional layers have smaller receptive fields, preserving spatial details such as texture edges and color transitions in the image; deep convolutional layers have larger receptive fields, capturing semantic abstract features such as the overall structure of brand logos and product outlines. By extracting feature maps at different depths of the network, a multi-level feature pyramid with resolution ranging from high to low and semantic hierarchy from shallow to deep is obtained. This feature pyramid contains a complete information hierarchy from pixel-level texture to region-level semantics, providing a sufficient feature foundation for subsequent cross-scale fusion. In practical processing, typically 3 to 5 feature layers of different depths are selected, with each layer having different spatial resolution and number of channels. Shallow feature maps have high resolution but fewer channels, while deep feature maps have low resolution but more channels, complementing each other.

[0084] In constructing cross-scale feature fusion representations, feature maps of different resolutions need to be unified to the same spatial size. Bilinear interpolation upsampling is performed on low-resolution deep semantic feature maps to align their spatial dimension with that of high-resolution shallow feature maps. At the channel level, learnable weight coefficients are assigned to spatial detail features and semantic abstract features, and a fused feature map is generated through weighted summation. The weight coefficients are automatically adjusted based on the semantic segmentation supervision signals during the training phase, enabling the fused feature map to retain spatial localization accuracy while possessing semantic discrimination capabilities for brand identification areas, product display areas, and auxiliary information areas. The number of channels in the fused feature map is reduced and compressed using 1×1 convolutions to decrease the computational load of subsequent pixel-level classification while retaining key discriminative information.

[0085] When performing pixel-level classification prediction on the fused feature map, a fully convolutional network structure outputs a class probability vector for each pixel location. The class set includes four semantic categories: brand identity, product display, auxiliary information, and background. The category with the highest probability is taken as the semantic label for that pixel, thus generating a semantic segmentation mask of the same size as the original image. In the semantic segmentation mask, each pixel carries a unique class label, and the labels of adjacent pixels together constitute the spatial distribution map of the region. After obtaining the semantic segmentation mask, morphological connectivity analysis is performed on it. The spatial connectivity relationship between pixels is determined by the 8-connected neighborhood rule. Pixels with the same semantic label and that are connected to each other are grouped into the same connected component, i.e., a candidate region. For each connected component, a contour tracking algorithm is used to extract its outer boundary pixel sequence, and the region is bounded by a minimum axis aligned rectangle to form a preliminary set of region boundaries. Connected components with too small an area are filtered out as noise regions to avoid fragmented pixels interfering with subsequent analysis.

[0086] In the initial set of region boundaries, some regions exhibit jagged or irregular shapes, due to inherent quantization errors in pixel-level classification, making it difficult to accurately reflect the true contours of brand visual elements. To address these regions, a boundary refinement method based on the consistency of boundary pixel gradient directions is employed. Specifically, the directional distribution of image gradients within the neighborhood of boundary pixels is calculated. When the pixel gradient direction of a boundary segment is highly consistent with the tangent direction of that segment, it is considered aligned with the true edge in the image and is retained. When the deviation between the gradient direction and the tangent direction exceeds a set angle threshold, the boundary pixel position is locally corrected using the gradient direction as a guide, snapping the boundary pixels to the position with the largest image gradient amplitude, thus ensuring the boundary conforms to the actual contours of the brand visual elements. After iterative correction segment by segment, a refined set of region boundaries is obtained, where the smoothness and contour alignment accuracy of each region boundary are significantly better than the initial boundary. The boundary refinement process is performed only within the local neighborhood of the boundary pixels, without changing the overall semantic category of the region.

[0087] The spatial location of each region within the refined region boundary set is quantified. For each region, the coordinates of the four vertices of its minimum bounding rectangle are calculated, with the top left corner of the image as the origin and the coordinates extending horizontally to the right as... The positive direction of the axis is vertically downward. Establish a pixel coordinate system along the positive axis, and set the vertex coordinates as follows: The format is recorded. The pixel coordinates of the centroid of the region are obtained by averaging the coordinates of all pixels within the region. Let the total number of pixels in the region be... , No. The coordinates of each pixel are Then the centroid coordinates satisfy , The proportion of the region area to the total image area. Defined as ,in and These represent the pixel width and height of the image data, respectively. The vertex coordinates, centroid pixel coordinates, and area ratio of the minimum bounding rectangle are then used. The vectors are concatenated to form a structured description vector, which serves as a spatial location description of the candidate region.

[0088] Finally, all candidate regions filtered by area, along with their corresponding semantic category labels and spatial location description vectors, are aggregated to form a candidate region of interest (ROI) set. Each element in this set contains the region's semantic category, the coordinates of the vertex of its minimum bounding rectangle, the coordinates of its centroid pixel, and its area ratio, fully describing the region's location, scale, and category attributes within the image. This structured representation provides a standardized region description input for subsequent association mapping models based on historical eye-tracking data, making candidate ROIs comparable across different images and supporting cross-image gaze probability prediction and personalized analysis.

[0089] In one optional implementation, pixel-level classification prediction is performed on the fused feature map, and a semantic label is assigned to each pixel as a brand identifier category, product display category, auxiliary information category, or background category to generate a semantic segmentation mask. Morphological connectivity analysis is then performed on the semantic segmentation mask to identify a set of pixels with the same semantic label and spatial connectivity, including:

[0090] Pixel-level feature vector decoding is performed on the fused feature map to generate a four-dimensional category response vector for each pixel, which includes brand identity category, product display category, auxiliary information category and background category;

[0091] The four-dimensional category response vector is subjected to cross-category competitive suppression processing. By calculating the relative intensity difference between the response values ​​of each category, the category with the highest response value is strengthened and the responses of other categories are suppressed, thus generating a category response vector after competitive suppression.

[0092] Based on the maximum response class in the category response vector after competition suppression, a corresponding semantic label is assigned to each pixel, and the semantic labels are organized according to the pixel spatial location to form a semantic segmentation mask;

[0093] The semantic segmentation mask is grown by connected component growth based on eight-neighbor topology. An initial seed pixel is selected from the semantic segmentation mask. Pixels with the same semantic label and spatial eight-neighbor adjacency as the initial seed pixel are iteratively added to the same connected component until no new pixels can be added, thus forming the first connected component.

[0094] The connected component growth operation is repeatedly performed on the remaining pixels in the semantic segmentation mask that are not marked as the first connected component, generating multiple non-overlapping connected components in sequence. The pixels in each connected component have the same semantic label and satisfy the spatial connectivity constraint. The set of pixel position coordinates in each connected component is taken as the set of pixels with the same semantic label and spatial connectivity.

[0095] After constructing the fused feature map, a feature vector is extracted for each pixel location, and this feature vector is mapped into a four-dimensional category response vector by a decoding network. The four components of this four-dimensional vector correspond to the response intensity of the brand identity category, product display category, auxiliary information category, and background category, respectively. The value of each component reflects the original activation level of the current pixel belonging to the corresponding category. The decoding network typically consists of several cascaded convolutional layers and upsampling layers, and its output resolution remains consistent with the original fused feature map, thus ensuring that each pixel location obtains a complete four-dimensional category response vector without loss of spatial information.

[0096] After obtaining the four-dimensional category response vector, cross-category competitive suppression is applied. The core idea of ​​this process is that different categories are mutually exclusive, meaning a pixel can only belong to a unique semantic category. Therefore, an explicit competitive mechanism is needed to strengthen the response of the optimal category and suppress interference from other categories. Specifically, the relative intensity difference between the components in the four-dimensional vector is calculated, and suppression weights are applied to other components based on the component with the largest response. Let the four-dimensional category response vector of a pixel be... ,in Corresponding brand logo category, Corresponding product display categories Corresponding auxiliary information category, The maximum response value is corresponding to the background category. For the first Each category, response value after competition suppression Calculated as follows:

[0097] ;

[0098] in To suppress the intensity coefficient and control the extent to which non-optimal categories are suppressed, For category indexing. When and The greater the gap, The stronger the suppression, the more prominent the relative advantage of the optimal category becomes. This competition suppression mechanism can effectively reduce the fuzzy response at category boundaries and improve the accuracy of segmentation masks in the edge areas of visual elements, which is particularly important for the fine boundary division between brand logos and product display areas.

[0099] After completing the competition suppression process, the four-dimensional class response vector after competition suppression is obtained. The category corresponding to the component with the largest response value is used as the semantic label for that pixel. This label assignment operation is performed on all pixels in the image one by one, and the semantic labels of each pixel are organized and arranged according to their spatial position in the image coordinate system, ultimately forming a semantic segmentation mask of the same size as the original image. Each position in the mask stores the category number of the corresponding pixel; for example, integers 1, 2, 3, and 4 are used to encode four semantic labels: brand logo, product display, auxiliary information, and background, respectively, facilitating efficient indexing and querying in the subsequent connected component analysis stage.

[0100] After the semantic segmentation mask is generated, it undergoes connected component growth analysis based on eight-neighbor topological relationships. Eight-neighbor topological relationships refer to the fact that for any pixel position in the mask, its neighborhood includes four orthogonal directions (top, bottom, left, right) and four diagonal directions (top left, top right, bottom left, bottom right), for a total of eight adjacent pixel positions. Compared to the four-neighbor connectivity definition, eight-neighbor connectivity can identify the continuity in diagonal directions, which is closer to how the human eye perceives the continuity of visual regions, helping to avoid incorrect region segmentation caused by diagonal contact.

[0101] Connected component growth begins with the selection of an initial seed pixel. Following a left-to-right, top-to-bottom scanning order, the first pixel in the semantic segmentation mask that has not yet been assigned to any connected component is found and used as the initial seed pixel, and its semantic label is recorded. Starting with the seed pixel, establish a queue to be expanded, add the seed pixel to the queue and mark it as visited. Iteratively remove the current pixel from the queue and traverse all its eight neighboring pixels: if the semantic label of the neighboring pixel is... If a pixel is identical and has not yet been visited, it is added to the queue and marked as visited, and simultaneously included in the current connected component. This process is repeated until the queue is empty, meaning no new pixels can be added. At this point, the growth of the first connected component is complete, denoted as […]. , It contains the set of pixel position coordinates that are identical to the initial seed pixel semantic label and are spatially connected in the sense of eight-neighbor topology.

[0102] First connected domain After growth is complete, continue scanning the remaining unlabeled pixels in the semantic segmentation mask, select the next unvisited pixel as the new seed pixel, and repeat the above connected component growth operation to generate the second connected component. This process is repeated iteratively to generate... ,common The process of growing each connected component involves creating non-overlapping connected components until all pixels in the mask are assigned to a particular connected component. Since the growth process of each connected component strictly requires semantic label consistency and eight-neighbor spatial connectivity, pixels of different semantic categories will not be assigned to the same connected component, and groups of pixels of the same semantic category but not spatially adjacent will also be divided into different connected components, thus ensuring that each connected component has cohesion in both semantic and spatial dimensions.

[0103] Connected components ( The set of pixel coordinates within a region constitutes a set of pixels with the same semantic label and satisfying spatial connectivity constraints, serving as the basic unit of candidate regions of interest. For connected components with the semantic label "brand identity," their pixel set corresponds to the spatial distribution range of the brand identity in the image; for connected components with the "product display" category, their pixel set describes the visually occupied area of ​​the product; and for connected components with the "auxiliary information" category, they correspond to the distribution area of ​​auxiliary visual elements such as text descriptions and price tags. Connected components of the "background" category are usually large in area and do not have significance for brand visual analysis; they can be filtered out in the subsequent candidate region of interest screening stage based on area thresholds or category labels. The set of pixel coordinates of each non-background connected component and its corresponding semantic category label are output together to form a structured description of candidate regions of interest, providing basic input for the subsequent generation of spatial location descriptions and prediction of gaze probability distribution.

[0104] In one optional implementation, establishing a correlation mapping model between user feature profile data and candidate interest region set based on the fixation point coordinate sequence and dwell time sequence in the historical eye-tracking dataset includes:

[0105] Extract the gaze point coordinate sequence and dwell time sequence of each user from the historical eye-tracking dataset, determine the spatial region affiliation of the gaze point coordinate sequence, map each gaze point coordinate to the corresponding region in the candidate region of interest set, and generate a gaze point and region affiliation table.

[0106] Based on the attribution table and the dwell time sequence, calculate the total dwell time and gaze shift frequency of each user in the brand identification area, product display area and auxiliary information area, and combine the total dwell time and gaze shift frequency to form a user gaze behavior feature vector;

[0107] Temporal dependency modeling is performed on the user gaze behavior feature vector to extract the dynamic change trend of the jump path pattern and dwell time between different regions of the gaze point, and to generate a user behavior pattern representation containing temporal context information.

[0108] Demographic features and consumption behavior features in user profile data are jointly embedded with the user behavior pattern representation to construct a unified feature space that integrates static user attributes and dynamic eye-tracking behavior.

[0109] In the unified feature space, a multi-output regression mapping relationship is established from user feature profile data to a set of candidate interest regions. The multi-output regression mapping relationship takes user feature profile data as input and outputs the user's gaze probability distribution and expected dwell time distribution for each region in the set of candidate interest regions, forming an association mapping model.

[0110] Extracting the gaze coordinate sequence and dwell time sequence for each user from historical eye-tracking datasets requires preprocessing the raw eye-tracking records. This includes removing abnormal frames caused by blinking or tracking loss, and clustering consecutive sampling points based on velocity thresholds to merge discrete sampling coordinates into stable gaze events with start and end timestamps. For each gaze event, its center coordinates are recorded. Corresponding duration The core of spatial region attribution determination lies in determining whether the coordinates of the gaze point fall within the pixel mask range of a certain region in the candidate region of interest set. For each candidate region of interest, its spatial location has already been provided by the pixel-level mask during the visual semantic segmentation stage; therefore, attribution determination can be achieved by querying the coordinates in the mask bitmap. The assignment is completed using the label value at the location. If the coordinate falls within the overlapping boundary of multiple areas, the assignment is performed according to the following priority order: brand identification area first, product display area second, and auxiliary information area third. After assignment is completed, a table of the relationship between gaze points and areas is generated. Each record in the table includes the user identifier, gaze event number, assigned area category label, and corresponding dwell time.

[0111] Based on the attribution table and dwell time sequence, the total dwell time of each user in the brand logo area, product display area, and auxiliary information area is calculated separately. Let user... In the The total dwell time in the type of area is ,in These correspond to three different regions. (Gaze shift frequency) Indicates user From the Jump to the next class region The frequency of each region is calculated by traversing the region label pairs of adjacent gaze events in the attribution table. and splicing together to form a user gaze behavior feature vector Specifically, Includes 3 total stay duration components and the maximum The frequency of transitions between regions comprises 12 dimensions. To eliminate the dimensional impact of differences in viewing time among different users, the total dwell time component is normalized by dividing it by the sum of the total duration of all gaze events for that user, thus converting it into the duration proportion of each region; the frequency of transitions component is divided by the total number of transitions for that user, thus converting it into an estimated transition probability value.

[0112] When modeling the temporal dependency of gaze behavior feature vectors, a gated recurrent unit network is used to encode the gaze point attribution sequence for each user. The gaze event sequence is arranged chronologically. The input at each time step is a concatenated vector of the region category one-hot encoding and the normalized dwell time of the current gaze event, denoted as . ,in Indexed by time step. The gated loop unit updates the hidden state at each time step. This captures the dynamic trends in transition path patterns and dwell times within the gaze sequence. After sequence encoding is complete, the hidden state at the final time step is retrieved. As the behavioral pattern representation vector of this user, denoted as . Implicit in the data is temporal contextual information such as the rhythm of user gaze region switching, lingering tendency in specific areas, and the overall trend of the gaze path during viewing, compared to static statistical feature vectors. It has stronger sequential semantic representation capabilities. To further enhance the robustness of the representation, a random time-step mask is applied to the input sequence during the training phase, enabling the model to generate stable behavioral pattern representations even when some gaze events are missing.

[0113] Demographic and consumer behavior features in user profile data are numerically encoded to obtain static attribute vectors. Demographic characteristics include age group, gender, and geographic stratum; consumer behavior characteristics include purchase frequency of product categories, average order value range, and brand preference intensity scores within the past 90 days, all of which have been normalized. With behavioral pattern representation The data is fed into a joint embedding network, which consists of two independent fully connected projection layers. and Dimensionality reduction mapping is performed, then the two mapping results are concatenated, and finally a fusion fully connected layer is passed through to output a fused representation vector in a unified feature space. The training objective of the joint embedding network is to make the representations of static attributes and dynamic eye-tracking behaviors from the same user as similar as possible in a unified feature space, while maintaining the distinguishability of representations from different user groups. This is optimized using a weighted combination of contrastive learning loss and task supervision loss. The contrastive learning loss selects eye-tracking data from the same user in different viewing scenarios as positive sample pairs and data from different users as negative sample pairs, thus promoting the fusion of representations. It can reliably reflect individual user differences.

[0114] In a unified feature space, with fusion representation As input, a multi-output regression mapping relationship is established. The output is divided into two parallel branches: a fixation probability distribution branch and an expected dwell time distribution branch. The fixation probability distribution branch outputs the user's... The fixation probability vector of each region in the candidate region of interest set Each component Indicates user gaze at the first The probability of each candidate region is calculated, and softmax normalization is applied to all components to satisfy the probability constraints. The expected dwell time distribution branch outputs the user. Vector of expected dwell time in each candidate region , of which components For users In the The estimated dwell time for each candidate region is calculated using ReLU activation to ensure non-negative output. The two branches are each implemented using a two-layer fully connected network, sharing the fusion representation from the underlying layer. As input, during training, the gaze probability distribution branch uses cross-entropy loss, with the frequency distribution of the user's actual gaze region in historical eye-tracking data as the supervision label; the expected dwell time branch uses mean squared error loss, with the mean of historical dwell time as the supervision label. The losses of the two branches are weighted and summed according to the hyperparameter weights, and then jointly backpropagated to update the entire network parameters, ultimately forming a complete correlation mapping model from user feature profile data to the gaze probability distribution and expected dwell time distribution of candidate interest regions.

[0115] In one optional implementation, the multidimensional feature profile data is input into the association mapping model to predict the personalized gaze probability distribution of the target user for each candidate region of interest, and a personalized region of interest set is formed by filtering based on the gaze probability threshold, including:

[0116] The multidimensional feature profile data is processed by feature dimension alignment and numerical scaling transformation to generate standardized feature vectors that are compatible with the input interface of the association mapping model.

[0117] The standardized feature vector is input into the association mapping model. The standardized feature vector is then subjected to nonlinear transformation and feature space projection through the multi-layer mapping structure of the association mapping model. The original gaze tendency score of the target user for each region in the candidate interest region set is then output.

[0118] The original gaze tendency score is adjusted globally by calculating the variance and kurtosis of the original gaze tendency score across all candidate regions of interest. When the variance exceeds a preset dispersion range, the score is corrected by centralization. When the kurtosis exceeds a preset centralization range, the score is corrected by broadening, thus generating a gaze score with adjusted distribution.

[0119] The distributed fixation scores are subjected to probability transformation mapping, and the distributed fixation scores of each region are converted into personalized fixation probability distributions that satisfy probability constraints through a monotonically increasing transformation relationship.

[0120] The cumulative distribution function is calculated based on the personalized gaze probability distribution. A probability value corresponding to a preset cumulative probability is selected on the cumulative distribution function as the gaze probability threshold. The gaze probability threshold is used to perform saliency screening on each region in the personalized gaze probability distribution, and regions with gaze probabilities higher than the gaze probability threshold are retained to form a personalized interest region set.

[0121] After obtaining the output capability of the association mapping model, it is necessary to transform the multidimensional feature profile data of the target user into a standardized input form that can be directly processed by the model. Multidimensional feature profile data typically covers heterogeneous features such as user demographic attributes, consumer behavior preference tags, category interest weights, and historical interaction frequency. These features differ significantly in scale, value range, and distribution. To eliminate the interference caused by inconsistencies in scale to the model input layer, feature dimension alignment and numerical scaling transformation must be performed on each dimension. For continuous numerical features, a standardization transformation based on training set statistics is used, subtracting the mean from each feature and dividing by the standard deviation to transform it into a standard normal distribution with a mean of 0 and a standard deviation of 1. For categorical features, one-hot encoding is used to expand them into binary vectors, which are then concatenated to the end of the numerical feature vector, ensuring strict alignment between the feature space dimension and the dimension declared by the association mapping model input interface. After the above processing, the generated standardized feature vector... It has a unified numerical scale and a fixed dimensional structure, and can be directly fed into the model for inference.

[0122] Standardize the feature vector After inputting the association mapping model, the model performs a layer-by-layer nonlinear transformation and feature space projection through a multi-layer mapping structure. Each mapping layer consists of a linear transformation matrix and an activation function. The activation function is a Corrected Linear Unit (ReLU) to retain positive activation signals and suppress negative noise. Batch normalization is introduced in the intermediate layers to ensure that the intermediate feature representations maintain a stable numerical range within each batch, preventing gradient vanishing or exploding phenomena from affecting inference accuracy. The final output layer corresponds to each region in the candidate region of interest set, outputting a set of real-valued scores, denoted as the original gaze tendency score vector. , of which Each component Indicates the target user's view on the first The original gaze tendency intensity of each candidate region of interest. Since it has not undergone probability normalization, there is no probability constraint relationship between the components, and its distribution pattern exhibits abnormal patterns of excessive concentration or excessive dispersion due to the extreme values ​​of user characteristics.

[0123] To ensure the rationality of subsequent probability transformation mapping, the original gaze tendency score vector needs to be... Perform global distribution pattern adjustment. Calculation. Variance index across all candidate regions of interest With kurtosis index The variance index reflects the dispersion of the scores, while the kurtosis index reflects the sharpness or flatness of the score distribution. The preset dispersion interval is... The preset concentration interval is .when Exceeding the upper bound of the preset dispersion interval This indicates that the score distribution is too scattered, with some areas scoring extremely high while others score extremely low, requiring a centralization correction: Each component in the middle contracts towards the mean, that is, the corrected score is... ,in The mean of all components. The centralization contraction coefficient is determined based on... The degree to which the upper bound is exceeded is adaptively determined; when Below the preset lower bound of the dispersion interval This indicates that the scores are too concentrated and need to be broadened to make them more concentrated. ,in This is the broadening factor. For the kurtosis index... When it exceeds the upper limit of the preset concentration range When the distribution tails are too thick, truncation is used to cut off extreme scores exceeding three standard deviations from the mean to the boundary values, then the mean is recalculated and aligned. Below the lower bound of the preset concentration interval When the distribution is too flat, a power transformation is applied to the scores. This enhances the central tendency of the distribution. After the above conditional correction, the fixation score vector with distribution adjustment is obtained. Its distribution pattern falls within a reasonable range of dispersion and concentration.

[0124] The gaze score vector after distribution adjustment A probability transformation mapping is performed to convert the data into a personalized gaze probability distribution that satisfies probability constraints. The softmax function is used as a monotonically increasing transformation relation. Each component in the output is subjected to exponentialization followed by normalization, ensuring that all output components are non-negative and their sum is 1, thus satisfying the basic constraints of the probability distribution. Specifically, the first... Personalized fixation probability corresponding to each candidate region of interest Depend on Give, all regions Construct a personalized gaze probability distribution vector The softmax function naturally preserves the relative magnitude of the scores, meaning that regions with higher scores still have higher probability values ​​after transformation, thus ensuring that the probability distribution faithfully reflects the original gaze tendency scores.

[0125] Based on personalized gaze probability distribution vector Calculate its cumulative distribution function (CDF). Then, sort all candidate regions of interest according to... Sort the regions from smallest to largest, and then sum the probability values ​​of each region sequentially to obtain a cumulative probability sequence. Preset cumulative probability. The cumulative probability is found on the cumulative distribution function. The probability value corresponding to the time is used as the fixation probability threshold. By dynamically determining the threshold based on the cumulative distribution function, the problem of over-filtering or under-filtering that occurs with a fixed threshold under different users or image content scenarios can be avoided, allowing the filtering results to adapt to the overall shape of the current probability distribution. right Saliency screening was performed on each region, and those that met the criteria were retained. All candidate regions of interest are considered, and the set of these regions is defined as the personalized region of interest set. .gather Each region in the model represents the spatial area where the target user is predicted to exhibit significant gaze behavior while browsing brand visual content. This provides precise spatial constraints for subsequent gaze trajectory sample matching and typical gaze pattern mining, and also provides a user-centric personalized basis for generating brand visual optimization strategies.

[0126] In one optional implementation, gaze trajectory samples matching the multidimensional feature profile data are extracted from the historical eye-tracking dataset. Temporal sequence mining is performed on the gaze trajectory samples to identify the shifting patterns and dwell patterns of the gaze point within the personalized region of interest set. A typical gaze pattern description of the target user is generated, including:

[0127] A subset of historical users is selected from the historical eye-tracking dataset, where the user features and multidimensional feature profiles meet the similarity matching conditions in different dimensions. The time-series records of gaze point coordinates and dwell time corresponding to the historical user subset are extracted and combined to form gaze trajectory samples.

[0128] The region attribution of the gaze point coordinate time sequence records in the gaze trajectory sample is determined, and each gaze point is mapped to the corresponding region or non-interest background region in the personalized interest region set to generate a region identification time sequence.

[0129] A state transition matrix is ​​constructed for the time sequence of the region identifiers, the number of gaze jumps from one region to another is counted, the transition probability between each region pair is calculated, and an inter-regional transition probability matrix is ​​formed as a quantitative representation of the transition pattern.

[0130] Based on the time-series records of dwell time, the statistical distribution of dwell time of the gaze point in each region within the set of personalized interest regions is calculated, and the mode value of dwell time and the dwell time dispersion index of each region are extracted.

[0131] The mode value of dwell time and the dispersion index of dwell time are used to perform pattern clustering to identify two types of dwell behavior features: stable dwell pattern and fluctuating dwell pattern. The stable dwell pattern and the fluctuating dwell pattern are then associated with the corresponding regions in the set of personalized interest regions.

[0132] The inter-regional transition probability matrix and the dwell behavior features associated with each region are structured and organized to construct a composite gaze behavior descriptor that includes the region jump probability distribution and the region dwell pattern type, which serves as a typical gaze pattern description for the target user.

[0133] like Figure 2 As shown, the method includes:

[0134] When selecting a subset of historical users from the historical eye-tracking dataset that matches the multidimensional feature profile data of the target users, similarity matching conditions must be met simultaneously across multiple feature dimensions, including age group, gender, consumption preference category, and brand awareness. Specifically, for continuous feature dimensions, similarity is determined by the absolute difference not exceeding a preset tolerance threshold; for discrete feature dimensions, the values ​​must be completely identical. Only historical users who meet the above conditions across all participating dimensions are included in the historical user subset. After selection, the time-series records of gaze coordinates and dwell time generated by each historical user in the subset on the target image or similar brand visual content are extracted. These records are then aligned by timestamp and combined to form a gaze trajectory sample. This sample retains complete temporal sequence information, providing a foundation for subsequent temporal sequence mining.

[0135] When performing region attribution determination on the time-series records of gaze point coordinates in the gaze trajectory sample, for each gaze point, it is determined whether its coordinates fall within the personalized region of interest set. Within a specific pixel area. If the coordinates of a gaze point fall within the bounding boxes of multiple regions simultaneously, the semantic segmentation region containing those coordinates is taken as the belonging region, based on the pixel-level precise mask; if the gaze point does not belong to any of these regions, the semantic segmentation region is taken as the belonging region. Any region within the sequence is then labeled as a non-interest background region. Following these rules, the attribution of each gaze point in the time-series record is determined sequentially, and the region identifiers are arranged chronologically to generate a time-series sequence of region identifiers. This sequence transforms the original two-dimensional coordinate information into a discrete region state sequence, thereby supporting subsequent state transition analysis.

[0136] Based on the generated time-series sequence of region identifiers, a state transition matrix between regions is constructed, and a set of personalized interest regions is defined. The CCP After adding non-interest background regions, the state space contains a total of [number] interest regions. There are several states. The time-series sequence of region identifiers is scanned step by step, and statistics are collected from each state. Jump to status The number of times is recorded as the jump count. ,in and These are all region identifiers in the state space. For each initial state... Jump to the sum of the counts of all target states to obtain the state. Total number of jumps from the start Then from the state Transition to state The transition probability is .when When the value is zero, the transition probability of the corresponding row is set to zero to avoid division by zero anomalies. The transition probabilities of all state pairs are organized into a matrix to obtain the inter-region transition probability matrix, which serves as a quantitative representation of the target user's gaze-jumping pattern. In this matrix, the diagonal elements... It reflects the probability that the fixation point remains in the same area without jumping, while off-diagonal elements describe the intensity of the tendency to jump between different areas.

[0137] For the time-series records of stay duration, the stay duration values ​​are grouped according to the region identifier and then summarized separately. The observed dwell time values ​​corresponding to each region of interest are used to form a dwell time sample set for each region. For each region's dwell time sample set, the mode of dwell time and the dwell time dispersion index are calculated. The mode of dwell time is obtained by kernel density estimation of the sample set. The dwell time value corresponding to the peak point of the kernel density estimation curve is taken as the typical dwell time representative value for that region, denoted as . The dispersion index of dwell time is measured using the interquartile range, which is the difference between the 75th percentile and the 25th percentile of the sample set, denoted as _____. This indicator is robust to extreme outliers and can accurately reflect the dispersion of dwell time without being affected by individual extreme gaze events.

[0138] Based on the length of stay in each area Dispersion index of dwell time A two-dimensional feature vector is constructed, and pattern clustering analysis is performed on all regions of interest. The clustering objective is to classify dwell behavior into two categories: stable dwell patterns and fluctuating dwell patterns. Stable dwell patterns correspond to... Smaller areas indicate that users' gaze duration in those areas is concentrated and consistent, demonstrating a strong gaze anchoring effect; fluctuating gaze patterns correspond to... Larger regions indicate that users' gaze duration in those areas is dispersed, with significant individual differences, reflecting the uncertainty of gaze behavior. A K-means clustering algorithm (K=2) was used to cluster the two-dimensional feature vectors. The effectiveness of the clustering results was verified using the silhouette coefficient, ensuring that the distinguishability between the two types of gaze behavior reached an acceptable level. After clustering, the gaze pattern category label of each region of interest was associated with that region, forming a correspondence between regions and gaze behavior features.

[0139] The inter-region transition probability matrix and the dwell behavior characteristics of each region are structured to construct a composite gaze behavior descriptor. This descriptor consists of two parts: one is the region jump probability distribution, i.e., the complete inter-region transition probability matrix, recording the jump tendency between all state pairs; the other is the dwell pattern type label for each region, containing the dwell pattern category corresponding to each region of interest and its corresponding... and Numerical data. The two parts of information are stored in a structured data format: the region jump probability distribution is saved in matrix form, and the region dwell pattern type is saved in key-value pairs, where the key is the region number and the value is the dwell pattern category label and statistical indicators. This composite gaze behavior descriptor comprehensively describes the target user's gaze path preferences and dwell stability in each region on brand visual content, forming a typical gaze pattern description for the target user, providing a quantitative basis for generating subsequent brand visual optimization strategies. In practical applications, if the historical user subset is too small, resulting in insufficient statistical stability, Laplace smoothing can be applied to the transition probability matrix, adding a small smoothing constant to each count value to avoid zero values ​​in the probability estimation of sparse transition paths, thereby improving the generalization reliability of the typical gaze pattern description.

[0140] In one optional implementation, based on the personalized region of interest set and the typical gaze pattern description, a brand visual optimization strategy for the target user is generated. The optimization strategy specifies layout adjustment schemes for each brand visual element in the image data, including:

[0141] Spatial coordinates and area parameters of each region are extracted from the set of personalized interest regions. The inter-regional transition probability matrix and the statistical distribution of dwell time in each region are extracted from the description of typical gaze patterns to construct a fusion representation of regional spatial attributes and gaze behavior characteristics.

[0142] Based on the inter-regional transition probability matrix, the gaze attractiveness index of each region is calculated, and the regions in the personalized interest region set are prioritized according to the gaze attractiveness index to generate a gaze value hierarchy structure.

[0143] Identify brand logo elements, product display elements, and auxiliary information elements in image data, and extract the current spatial layout coordinates and visual hierarchy attributes of each brand's visual elements;

[0144] Construct a spatial matching metric between brand visual elements and the gaze value hierarchy. Calculate the Euclidean distance between the current spatial layout coordinates of each brand visual element and the spatial center point of each area in the gaze value hierarchy. Combine the visual hierarchy attribute of the brand visual element with the gaze attraction index of the corresponding area to generate a layout coordination score.

[0145] For brand visual elements whose layout coordination score is lower than the coordination threshold, a layout reconstruction is initiated. The area with the highest priority in the gaze value hierarchy and which is not occupied is selected as the target layout area. The position adjustment vector from the current spatial layout coordinates to the target layout area is calculated.

[0146] The position adjustment vectors of each brand's visual elements are combined with the spatial constraint parameters of the target layout area to form a layout adjustment scheme, which is then used as a brand visual optimization strategy.

[0147] When extracting the spatial coordinates and area parameters of each region from the personalized region of interest set, it is necessary to uniformly incorporate the centroid coordinates, bounding box coordinates, and pixel area proportions of each region of interest generated in the previous steps into the fusion representation system. Simultaneously, the inter-region transition probability matrix and the statistical distribution of dwell time for each region are extracted from the description of typical gaze patterns, including the mode and interquartile range of dwell time for each region. These spatial attributes and gaze behavior features are concatenated to form the fusion representation vector for each region, serving as the input basis for subsequent gaze attractiveness index calculation and layout optimization. This fusion representation simultaneously preserves the physical spatial information of the region and the user's actual gaze behavior information, ensuring that subsequent optimization strategies possess both spatial operability and behavioral drive.

[0148] When calculating the gaze attractiveness index of each region based on the inter-region transition probability matrix, the transition probability matrix is ​​treated as a random walk problem on a directed weighted graph. For the th region in the personalized interest region set... Each region, its attraction index It consists of two parts: first, the sum of the ingress transfer probabilities of this area being the target of transfer by all other areas, reflecting the "convergence ability" of this area for the gaze stream; second, the normalized weight of the modulo of the duration of stay in this area, reflecting the user's tendency to stay in this area deeply. Specifically, The calculation method is as follows:

[0149] ;

[0150] in From state Transfer to the region The transition probability, For the region The duration of stay is a numerical value. This represents the maximum value among all values ​​for the duration of stay in all areas. and The weighting coefficients are configurable and satisfy... .according to All areas in the personalized interest area set are arranged in descending order from high to low to form a hierarchy of attention value. The highest priority area corresponds to the strongest user attention attraction and is the first choice for the placement of brand visual elements.

[0151] When identifying brand logo elements, product display elements, and auxiliary information elements in image data, the pixel-level category labeling results generated in the visual semantic segmentation stage are reused to extract the current spatial layout coordinates and visual hierarchy attributes of each brand visual element. The visual hierarchy attributes adopt a three-level encoding method: the brand logo element is assigned the highest level value. Product display elements are given a medium level of value. Auxiliary information elements are assigned values ​​to the basic level. This hierarchical coding reflects the differences in the strategic importance of various visual elements in brand communication practices. Elements with higher hierarchical levels should be prioritized and placed in areas with the highest attention attraction index to maximize brand exposure.

[0152] When constructing a spatial matching metric between brand visual elements and the hierarchy of attentional value, for each brand visual element... The current centroid coordinates are The visual hierarchy attribute is , and the first in the value hierarchy structure One region (the coordinates of the spatial center point are...) Attention attraction index is ), calculate Euclidean distance :

[0153] ;

[0154] Based on this, elements are generated by combining visual hierarchy attributes and gaze attraction index. Relative to region Layout coordination score :

[0155] ;

[0156] in This is the distance attenuation coefficient, used to control the severity of the penalty imposed by spatial distance on the coordination score. The higher the value, the more important it is for the brand's visual elements. With the region The better the spatial matching, meaning the higher-level element is currently positioned close to a highly attractive area. For each brand visual element... The final layout coordination score for an element is determined by taking its coordination score relative to the current or nearest region. and with preset coordination threshold Compare them.

[0157] Layout coordination score Below the coordination threshold The brand's visual elements were used to initiate a layout restructuring process. Within the attention value hierarchy, areas not yet occupied by other brand visual elements were searched sequentially from highest to lowest priority. The highest priority area meeting the criteria was identified as the target layout area, with its spatial center coordinates marked as follows: Position adjustment vector Defined as the direction vector pointing from the current centroid coordinates to the center point of the target layout region:

[0158] ;

[0159] This vector directly describes the amount of translation required for a brand visual element within the image plane, allowing image editing engines to directly read and execute pixel-level displacement operations. When multiple brand visual elements need to be reconstructed simultaneously, the visual hierarchy attribute is used for... The target layout areas are allocated sequentially from high to low to ensure that brand logo elements occupy the most eye-catching areas first, followed by product display elements, and then auxiliary information elements, thereby forming a visual hierarchy at the overall layout level that is highly consistent with the user's eye behavior.

[0160] Adjust the position vector of each brand's visual elements Combined with the spatial constraint parameters of the target layout area, a complete layout adjustment scheme is formed. Spatial constraint parameters include the bounding box range of the target layout area, a match check between the area area and the element's own size, and minimum spacing constraints between adjacent elements. If the size of a brand visual element exceeds the available area of ​​the target layout area, a proportional scaling operation is triggered, scaling the element to its maximum size that can just fit within the target area. The scaling ratio is recorded in the layout adjustment scheme for subsequent rendering reference. Finally, the position adjustment vectors, size scaling parameters, and spatial constraint parameters of all brand visual elements are summarized to form a structured brand visual optimization strategy output. This strategy is presented in the form of element-level operation instructions, which can directly drive image compositing or design tools to complete the automated rearrangement of brand visual content, ensuring a high degree of synergy between the final visual layout and the target user's personalized gaze behavior patterns.

[0161] The method further includes:

[0162] The interest region segmentation and gaze analysis method for brand visual elements initiates the entire processing flow by acquiring image data of brand visual content, multidimensional feature profile data of target users, and historical eye-tracking datasets. Image data is acquired using the standard HTTP protocol, loaded from a specified URL, supporting common formats such as JPEG, PNG, and WebP. Image resolution must meet a minimum width of 800 pixels and a height of at least 600 pixels. After loading, image data undergoes preprocessing, including color space conversion to RGB format and size normalization to a unified baseline scale, typically scaling the longer side to 1920 pixels while maintaining the aspect ratio. Multidimensional feature profile data of target users is stored using structured fields, including an age dimension field with integer age values, a gender dimension field with enumerated values, and a consumption preference dimension field that records the preference intensity scores for each category in vector form, with scores ranging from 0 to 10 (floating-point numbers). The historical eye-tracking dataset is organized in JSON format. Each record contains a participant identifier, trial identifier, time series of fixation coordinates, time series of dwell time, a snapshot of the participant's feature profile data, and experimental configuration parameters. Fixation coordinates are represented in the screen pixel coordinate system, and timestamps are recorded in the Unix timestamp format with millisecond precision.

[0163] When performing visual semantic segmentation on image data, a segmentation model based on a deep convolutional neural network is employed. This model outputs pixel-level classification results consistent with the input image size. The segmentation model is pre-trained to address the characteristics of brand visual content and can identify three main categories: brand logo areas, product display areas, and auxiliary information areas. The output of the segmentation model is the category label for each pixel and a confidence score for that label, with the confidence score ranging from 0 to 1 (a floating-point number). To generate a set of candidate regions of interest (ROIs), connectivity analysis is performed on the segmentation results, aggregating spatially contiguous pixel groups with the same category label into independent region objects. The connectivity analysis uses an eight-neighbor connectivity rule, and fragmented regions with an area less than 100 square pixels are filtered out to eliminate noise. The spatial description of each candidate ROI includes the coordinates of the top-left corner of its bounding rectangle, its width, height, centroid coordinates, area, category label, and average confidence score.

[0164] When establishing the association mapping model between user feature profile data and candidate interest region sets, effective training samples are first selected from the historical eye-tracking dataset. The selection criteria require that the tracking quality score of the samples be higher than 0.7 and the number of fixations be no less than 50. For each training sample, fixation point coordinate sequences and dwell time sequences are extracted, and the candidate interest region to which each fixation point belongs is determined through spatial mapping. The spatial mapping uses a point-within-a-polygon determination algorithm; if the fixation point coordinates fall within the bounding rectangle of a candidate interest region, the fixation point is considered to belong to that region. The number of fixations, cumulative dwell time, average single dwell time, first fixation time, and number of fixations within each candidate interest region are statistically analyzed; these statistical indicators constitute the fixation behavior feature vector for that region. The multidimensional feature profile data of participants is paired with the corresponding fixation behavior feature vectors to form training data pairs. The association mapping model adopts a gradient boosting decision tree architecture. During model training, the maximum tree depth is set to 6, the learning rate to 0.1, and the number of iterations to 100 epochs. Five-fold cross-validation is used to evaluate the model's generalization performance. The trained model can receive multi-dimensional feature profile data of any user as input and output the predicted gaze behavior feature vector of the user for each candidate region of interest.

[0165] After inputting the multidimensional feature profile data of the target user into the association mapping model, the predicted gaze behavior feature vectors of each candidate region of interest are obtained. The cumulative dwell time prediction value is extracted from this vector and normalized to the range of 0 to 1 as the personalized gaze probability. Normalization is achieved by dividing the cumulative dwell time prediction value of each region by the sum of the cumulative dwell time prediction values ​​of all regions, ensuring that the sum of the personalized gaze probabilities of all regions equals 1. A gaze probability threshold of 0.05 is set. For candidate regions of interest with a personalized gaze probability lower than this threshold, a filtering operation is performed, retaining regions with gaze probabilities reaching or exceeding the threshold to form a personalized region of interest set. If the number of regions in the personalized region of interest set after filtering is less than 3, the gaze probability threshold is lowered to 0.03 and the filtering process is repeated.

[0166] When extracting gaze trajectory samples that match the multidimensional feature profile data of the target user from historical eye-tracking datasets, a feature similarity matching strategy is employed. The similarity for the age dimension is calculated using the absolute value of the age difference: a difference less than 5 years results in a similarity score of 1, a difference between 5 and 10 years results in a similarity score of 0.5, and a difference greater than 10 years results in a similarity score of 0. The similarity for the gender dimension uses a perfect match rule: a score of 1 for the same gender and 0 for different genders. The similarity for the consumption preference dimension is obtained by calculating the cosine similarity between two preference vectors and linearly mapping it to the range of 0 to 1 as the similarity score for that dimension. The overall similarity is obtained by weighted averaging of the similarity scores for the three dimensions, with a weight of 0.3 for the age dimension, 0.2 for the gender dimension, and 0.5 for the consumption preference dimension. Historical records with an overall similarity score higher than 0.6 are selected, and the time-series sequences of gaze point coordinates and dwell time are extracted from these records and combined to form a gaze trajectory sample set.

[0167] When performing temporal sequence mining on gaze trajectory samples, the first step is to determine region affiliation, mapping each gaze point coordinate to a corresponding region in the personalized region of interest set. The original gaze point coordinate sequence is converted into a region identifier sequence, with each element recording the region identifier to which the gaze point belongs at that moment and the category label of that region. A state transition matrix is ​​constructed based on the region identifier sequence, where the rows and columns of the matrix correspond to various regions in the personalized region of interest set, and the matrix elements record the number of jumps from the region corresponding to the row to the region corresponding to the column. The region identifier sequence is traversed, and when the region identifier changes between two consecutive moments, it is considered that a region jump has occurred, and the count is accumulated at the corresponding position in the state transition matrix. After the state transition matrix is ​​statistically analyzed, the element value of each row is divided by the sum of the values ​​in that row to obtain the transition probability of jumping from that region to other regions. The transition probability value ranges from 0 to 1, and the sum of each row equals 1.

[0168] When identifying dwell patterns, the dwell time data of the gaze point in each region within the personalized region of interest set is extracted. The region identifier sequence is traversed; if the region identifier remains unchanged for multiple consecutive moments, the gaze point is considered to be dwelling in that region, and the dwell time is calculated by accumulating the corresponding time intervals. Dwell time values ​​for all dwell events are collected for each region, forming dwell time distribution data. The mode of the dwell time distribution is calculated, i.e., the dwell time value with the highest frequency, by constructing a histogram of dwell time and selecting the center value of the interval with the highest frequency as the mode. The dispersion index of the dwell time distribution is calculated, using the coefficient of variation as a measure, defined as the standard deviation divided by the mean. Pattern clustering is performed on the mode and coefficient of variation of dwell time for each region, using a two-class clustering algorithm to divide the regions into stable dwell patterns and fluctuating dwell patterns. The clustering rule is that regions with a coefficient of variation less than 0.4 are classified as stable dwell patterns, and regions with a coefficient of variation greater than or equal to 0.4 are classified as fluctuating dwell patterns. Each region is associated with its corresponding dwell pattern type label: stable dwell patterns are labeled as string S, and fluctuating dwell patterns are labeled as string V.

[0169] When generating typical gaze pattern descriptions for target users, the inter-region transition probability matrix and the dwell behavior features associated with each region are structured. A composite gaze behavior descriptor data structure is constructed, which includes a region list field, a transition matrix field, and a dwell feature field. The region list field records the identifier, category label, spatial location description, and personalized gaze probability of all regions in the personalized interest region set. The transition matrix field stores the complete numerical data of the inter-region transition probability matrix, organized in the form of a two-dimensional array. The dwell feature field is a dictionary structure, with the region identifier as the key and the value being an object containing the mode of dwell duration, the coefficient of variation of dwell duration, and the dwell pattern type label. Typical gaze pattern descriptions are serialized and stored in JSON format.

[0170] When generating brand visual optimization strategies based on personalized interest region sets and typical gaze pattern descriptions, the spatial coordinates and area parameters of each region are extracted from the personalized interest region set, and the inter-regional transfer probability matrix and the statistical distribution of dwell time for each region are extracted from the typical gaze pattern descriptions. A gaze attractiveness index is calculated for each region, which comprehensively considers the cumulative transfer probability of the region as a transfer target and the mode of dwell time for the region. The cumulative transfer probability is obtained by summing all elements in the corresponding column of the transfer probability matrix for that region. The gaze attractiveness index is calculated by multiplying the cumulative transfer probability by a weighting factor of 0.6 and adding the normalized value of the mode of dwell time multiplied by a weighting factor of 0.4. Regions in the personalized interest region set are then sorted in descending order according to the gaze attractiveness index, and the sorting results constitute a gaze value hierarchy.

[0171] When identifying brand visual elements in image data, the results of visual semantic segmentation are reused, and the segmented brand logo area, product display area, and auxiliary information area are respectively mapped to brand logo elements, product display elements, and auxiliary information elements. The current spatial layout coordinates of each brand visual element are extracted, i.e., the coordinates of the top-left and bottom-right corners of the element's bounding rectangle. The visual hierarchy attribute of each brand visual element is extracted, which is calculated by analyzing the element's area proportion, color salience, and positional salience. The visual hierarchy attribute value is obtained by weighted averaging of the three indicators: area proportion weighted at 0.4, color salience weighted at 0.3, and positional salience weighted at 0.3, with attribute values ​​ranging from 0 to 1.

[0172] When constructing a spatial matching metric between brand visual elements and the gaze value hierarchy, the Euclidean distance between the current spatial layout coordinates of each brand visual element and the spatial center point of each region in the gaze value hierarchy is calculated. For each brand visual element, the nearest region is identified as its associated region, and the gaze attractiveness index of this associated region is extracted. The layout harmony score is calculated by multiplying the inverse value of the normalized Euclidean distance by a weighting factor of 0.5, and then adding the product of the visual hierarchy attribute and the gaze attractiveness index multiplied by a weighting factor of 0.5. The normalization of the Euclidean distance is achieved by dividing by the length of the image diagonal. The layout harmony score ranges from 0 to 1.

[0173] For brand visual elements with a layout harmony score below the harmony threshold of 0.6, a layout reconstruction is initiated. The unoccupied area with the highest gaze attractiveness index in the gaze value hierarchy is selected as the target layout area. A position adjustment vector is calculated from the current spatial layout coordinates of the brand visual element to the spatial center point of the target layout area. This vector is a two-dimensional vector obtained by subtracting the current centroid coordinates of the element from the centroid coordinates of the target area. The position adjustment vectors of each brand visual element are combined with the spatial constraint parameters of the target layout area to form a layout adjustment scheme. The spatial constraint parameters include the coordinates of the bounding rectangle of the target area, the area of ​​the area, and the suggested element scaling ratio. The scaling ratio is calculated based on the ratio of the element area to the target area area. If the element area is greater than 0.8 times the target area area, it is recommended to reduce it to 0.75 times the target area area; if the element area is less than 0.5 times the target area area, it is recommended to enlarge it to 0.6 times the target area area. The layout adjustment scheme is output in a structured data format, including element identifiers, current coordinates, target coordinates, position adjustment vectors, scaling ratios, and associated target area identifiers.

[0174] A second aspect of this invention provides a system for segmenting and analyzing interest regions and gaze patterns for brand visual elements, comprising:

[0175] The image acquisition unit is used to acquire image data of brand visual content, multi-dimensional feature profile data of target users, and historical eye-tracking datasets.

[0176] The semantic segmentation unit is used to perform visual semantic segmentation on the image data, identify the brand logo area, product display area and auxiliary information area, and generate a set of candidate interest areas and their spatial location descriptions.

[0177] The mapping modeling unit is used to establish a correlation mapping model between user feature profile data and candidate interest region set based on the fixation point coordinate sequence and dwell time sequence in the historical eye movement dataset.

[0178] The prediction and filtering unit is used to input the multidimensional feature profile data into the association mapping model, predict the personalized gaze probability distribution of the target user for each candidate interest region, and filter to form a set of personalized interest regions based on the gaze probability threshold.

[0179] The pattern mining unit is used to extract gaze trajectory samples that match the multidimensional feature profile data from the historical eye-tracking dataset, perform time-series sequence mining on the gaze trajectory samples, identify the shifting patterns and dwell patterns of gaze points within the personalized interest region set, and generate a typical gaze pattern description of the target user.

[0180] An optimization generation unit is used to generate a brand visual optimization strategy for the target user based on the personalized interest region set and the typical gaze pattern description. The optimization strategy specifies the layout adjustment scheme of each brand visual element in the image data.

[0181] A third aspect of the present invention provides an electronic device, comprising:

[0182] processor;

[0183] Memory used to store processor-executable instructions;

[0184] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0185] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0186] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for dividing interest regions and analyzing gaze patterns based on brand visual elements, characterized in that, include: Acquire image data of brand visual content, multidimensional feature profile data of target users, and historical eye-tracking datasets; The image data is subjected to visual semantic segmentation to identify the brand logo area, product display area and auxiliary information area, and a set of candidate interest areas and their spatial location descriptions are generated. Based on the fixation point coordinate sequence and dwell time sequence in the historical eye-tracking dataset, an association mapping model between user feature profile data and candidate interest region set is established. The multidimensional feature profile data is input into the association mapping model to predict the personalized gaze probability distribution of the target user for each candidate region of interest, and a set of personalized regions of interest is formed by filtering according to the gaze probability threshold. The gaze trajectory samples that match the multidimensional feature profile data are extracted from the historical eye-tracking dataset. Temporal sequence mining is performed on the gaze trajectory samples to identify the shifting patterns and dwell patterns of the gaze points within the personalized interest region set, and to generate a typical gaze pattern description of the target user. Based on the personalized interest region set and the typical gaze pattern description, a brand visual optimization strategy is generated for the target user. The optimization strategy specifies the layout adjustment scheme of each brand visual element in the image data.

2. The method according to claim 1, characterized in that, Visual semantic segmentation is performed on the image data to identify brand logo areas, product display areas, and auxiliary information areas, generating a set of candidate regions of interest and their spatial location descriptions, including: Multi-scale convolutional feature extraction is performed on image data to obtain a multi-level feature pyramid containing spatial detail features and semantic abstract features; Based on the multi-level feature pyramid, a cross-scale feature fusion representation is constructed, and the spatial detail features and the semantic abstract features are weighted and fused in the channel dimension to generate a fused feature map. Pixel-level classification prediction is performed on the fused feature map, and semantic labels of brand identification category, product display category, auxiliary information category or background category are assigned to each pixel to generate a semantic segmentation mask. Morphological connectivity analysis is performed on the semantic segmentation mask to identify pixel sets with the same semantic labels and spatial connectivity, and the bounding boundary of each pixel set is extracted to form a preliminary region boundary set. For regions with irregular boundary shapes in the initial region boundary set, boundary refinement is performed based on the gradient direction consistency constraint of the boundary pixels to align the boundary with the actual outline of the brand visual elements, thereby generating a refined region boundary set. For each region within the refined region boundary set, calculate the vertex coordinates of the minimum bounding rectangle, the pixel coordinates of the region centroid, and the proportion of the region area to the total area of ​​the image data. Combine the vertex coordinates of the minimum bounding rectangle, the pixel coordinates, and the proportion to form a candidate region of interest set and its spatial location description.

3. The method according to claim 2, characterized in that, Pixel-level classification prediction is performed on the fused feature map, and a semantic label is assigned to each pixel, which can be categorized as brand identity, product display, auxiliary information, or background. A semantic segmentation mask is generated, and morphological connectivity analysis is performed on the semantic segmentation mask to identify a set of pixels with the same semantic label and spatial connectivity, including: Pixel-level feature vector decoding is performed on the fused feature map to generate a four-dimensional category response vector for each pixel, which includes brand identity category, product display category, auxiliary information category and background category; The four-dimensional category response vector is subjected to cross-category competitive suppression processing. By calculating the relative intensity difference between the response values ​​of each category, the category with the highest response value is strengthened and the responses of other categories are suppressed, thus generating a category response vector after competitive suppression. Based on the maximum response class in the category response vector after competition suppression, a corresponding semantic label is assigned to each pixel, and the semantic labels are organized according to the pixel spatial location to form a semantic segmentation mask; The semantic segmentation mask is grown by connected component growth based on eight-neighbor topology. An initial seed pixel is selected from the semantic segmentation mask. Pixels with the same semantic label and spatial eight-neighbor adjacency as the initial seed pixel are iteratively added to the same connected component until no new pixels can be added, thus forming the first connected component. The connected component growth operation is repeatedly performed on the remaining pixels in the semantic segmentation mask that are not marked as the first connected component, generating multiple non-overlapping connected components in sequence. The pixels in each connected component have the same semantic label and satisfy the spatial connectivity constraint. The set of pixel position coordinates in each connected component is taken as the set of pixels with the same semantic label and spatial connectivity.

4. The method according to claim 1, characterized in that, Based on the fixation point coordinate sequence and dwell time sequence in the historical eye-tracking dataset, the association mapping model between user feature profile data and candidate interest region set is established as follows: Extract the gaze point coordinate sequence and dwell time sequence of each user from the historical eye-tracking dataset, determine the spatial region affiliation of the gaze point coordinate sequence, map each gaze point coordinate to the corresponding region in the candidate region of interest set, and generate a gaze point and region affiliation table. Based on the attribution table and the dwell time sequence, calculate the total dwell time and gaze shift frequency of each user in the brand identification area, product display area and auxiliary information area, and combine the total dwell time and gaze shift frequency to form a user gaze behavior feature vector; Temporal dependency modeling is performed on the user gaze behavior feature vector to extract the dynamic change trend of the jump path pattern and dwell time between different regions of the gaze point, and to generate a user behavior pattern representation containing temporal context information. Demographic features and consumption behavior features in user profile data are jointly embedded with the user behavior pattern representation to construct a unified feature space that integrates static user attributes and dynamic eye-tracking behavior. In the unified feature space, a multi-output regression mapping relationship is established from user feature profile data to a set of candidate interest regions. The multi-output regression mapping relationship takes user feature profile data as input and outputs the user's gaze probability distribution and expected dwell time distribution for each region in the set of candidate interest regions, forming an association mapping model.

5. The method according to claim 1, characterized in that, The multidimensional feature profile data is input into the association mapping model to predict the personalized gaze probability distribution of the target user for each candidate region of interest, and a set of personalized regions of interest is formed by filtering based on the gaze probability threshold, including: The multidimensional feature profile data is processed by feature dimension alignment and numerical scaling transformation to generate standardized feature vectors that are compatible with the input interface of the association mapping model. The standardized feature vector is input into the association mapping model. The standardized feature vector is then subjected to nonlinear transformation and feature space projection through the multi-layer mapping structure of the association mapping model. The original gaze tendency score of the target user for each region in the candidate interest region set is then output. The original gaze tendency score is adjusted globally by calculating the variance and kurtosis of the original gaze tendency score across all candidate regions of interest. When the variance exceeds a preset dispersion range, the score is corrected by centralization. When the kurtosis exceeds a preset centralization range, the score is corrected by broadening, thus generating a gaze score with adjusted distribution. The distributed fixation scores are subjected to probability transformation mapping, and the distributed fixation scores of each region are converted into personalized fixation probability distributions that satisfy probability constraints through a monotonically increasing transformation relationship. The cumulative distribution function is calculated based on the personalized gaze probability distribution. A probability value corresponding to a preset cumulative probability is selected on the cumulative distribution function as the gaze probability threshold. The gaze probability threshold is used to perform saliency screening on each region in the personalized gaze probability distribution, and regions with gaze probabilities higher than the gaze probability threshold are retained to form a personalized interest region set.

6. The method according to claim 1, characterized in that, From the historical eye-tracking dataset, gaze trajectory samples matching the multidimensional feature profile data are extracted. Temporal sequence mining is performed on these gaze trajectory samples to identify the shifting patterns and dwell patterns of gaze points within the personalized region of interest set. A typical gaze pattern description of the target user is generated, including: A subset of historical users is selected from the historical eye-tracking dataset, where the user features and multidimensional feature profiles meet the similarity matching conditions in different dimensions. The time-series records of gaze point coordinates and dwell time corresponding to the historical user subset are extracted and combined to form gaze trajectory samples. The region attribution of the gaze point coordinate time sequence records in the gaze trajectory sample is determined, and each gaze point is mapped to the corresponding region or non-interest background region in the personalized interest region set to generate a region identification time sequence. A state transition matrix is ​​constructed for the time sequence of the region identifiers, the number of gaze jumps from one region to another is counted, the transition probability between each region pair is calculated, and an inter-regional transition probability matrix is ​​formed as a quantitative representation of the transition pattern. Based on the time-series records of dwell time, the statistical distribution of dwell time of the gaze point in each region within the set of personalized interest regions is calculated, and the mode value of dwell time and the dwell time dispersion index of each region are extracted. The mode value of dwell time and the dispersion index of dwell time are used to perform pattern clustering to identify two types of dwell behavior features: stable dwell pattern and fluctuating dwell pattern. The stable dwell pattern and the fluctuating dwell pattern are then associated with the corresponding regions in the set of personalized interest regions. The inter-regional transition probability matrix and the dwell behavior features associated with each region are structured and organized to construct a composite gaze behavior descriptor that includes the region jump probability distribution and the region dwell pattern type, which serves as a typical gaze pattern description for the target user.

7. The method according to claim 1, characterized in that, Based on the personalized interest region set and the typical gaze pattern description, a brand visual optimization strategy is generated for the target user. The optimization strategy specifies the layout adjustment scheme for each brand visual element in the image data, including: Spatial coordinates and area parameters of each region are extracted from the set of personalized interest regions. The inter-regional transition probability matrix and the statistical distribution of dwell time in each region are extracted from the description of typical gaze patterns to construct a fusion representation of regional spatial attributes and gaze behavior characteristics. Based on the inter-regional transition probability matrix, the gaze attractiveness index of each region is calculated, and the regions in the personalized interest region set are prioritized according to the gaze attractiveness index to generate a gaze value hierarchy structure. Identify brand logo elements, product display elements, and auxiliary information elements in image data, and extract the current spatial layout coordinates and visual hierarchy attributes of each brand's visual elements; Construct a spatial matching metric between brand visual elements and the gaze value hierarchy. Calculate the Euclidean distance between the current spatial layout coordinates of each brand visual element and the spatial center point of each area in the gaze value hierarchy. Combine the visual hierarchy attribute of the brand visual element with the gaze attraction index of the corresponding area to generate a layout coordination score. For brand visual elements whose layout coordination score is lower than the coordination threshold, a layout reconstruction is initiated. The area with the highest priority in the gaze value hierarchy and which is not occupied is selected as the target layout area. The position adjustment vector from the current spatial layout coordinates to the target layout area is calculated. The position adjustment vectors of each brand's visual elements are combined with the spatial constraint parameters of the target layout area to form a layout adjustment scheme, which is then used as a brand visual optimization strategy.

8. A system for segmenting and analyzing interest regions based on brand visual elements, used to implement the method as described in any one of claims 1-7, characterized in that, include: The image acquisition unit is used to acquire image data of brand visual content, multi-dimensional feature profile data of target users, and historical eye-tracking datasets. The semantic segmentation unit is used to perform visual semantic segmentation on the image data, identify the brand logo area, product display area and auxiliary information area, and generate a set of candidate interest areas and their spatial location descriptions. The mapping modeling unit is used to establish a correlation mapping model between user feature profile data and candidate interest region set based on the fixation point coordinate sequence and dwell time sequence in the historical eye movement dataset. The prediction and filtering unit is used to input the multidimensional feature profile data into the association mapping model, predict the personalized gaze probability distribution of the target user for each candidate interest region, and filter to form a set of personalized interest regions based on the gaze probability threshold. The pattern mining unit is used to extract gaze trajectory samples that match the multidimensional feature profile data from the historical eye-tracking dataset, perform time-series sequence mining on the gaze trajectory samples, identify the shifting patterns and dwell patterns of gaze points within the personalized interest region set, and generate a typical gaze pattern description of the target user. An optimization generation unit is used to generate a brand visual optimization strategy for the target user based on the personalized interest region set and the typical gaze pattern description. The optimization strategy specifies the layout adjustment scheme of each brand visual element in the image data.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.