Visual element automatic layout optimization method and system for commercial advertisement design
By constructing a user cognitive preference model and analyzing the semantic correlation of visual elements, a personalized commercial advertising layout scheme is generated. This solves the problem that individual user preferences and element correlation are ignored in existing technologies, improves the targeting of advertisements and the attractiveness of visual element layout, and realizes the dynamic and continuous optimization of advertising content.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNION UNIVERSITY
- Filing Date
- 2026-01-19
- Publication Date
- 2026-04-17
AI Technical Summary
Existing commercial advertising visual element layout technologies lack in-depth modeling of individual user cognitive preferences, cannot provide personalized layout solutions, and ignore the semantic relationships and content dependencies between elements, affecting the effective delivery and conversion of advertising information.
By acquiring multi-dimensional behavioral characteristic data of target users, a user cognitive preference model is constructed, the semantic correlation of visual elements is analyzed, the layout tension field distribution is calculated, a multi-level visual guidance path is generated, and adaptive spatial allocation and presentation sequence arrangement are performed to generate a personalized layout scheme, which is then optimized in combination with user interaction feedback data.
It achieves a precise match between advertising design and users' personalized needs, improves the targeting and user acceptance of advertisements, enhances the logic and attractiveness of visual element layout, dynamically presents advertising content, and forms a closed-loop optimization mechanism to continuously improve the layout scheme.
Smart Images

Figure CN121544323B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to advertising design technology, and more particularly to a method and system for automatically optimizing the layout of visual elements in commercial advertising design. Background Technology
[0002] With the rapid development of digital marketing, the visual design and layout of commercial advertisements have become key factors influencing user attention and conversion rates. Visual element layout in commercial advertising design refers to the rational arrangement of visual components such as text, images, icons, and logos within a limited advertising space to achieve optimal visual communication and marketing objectives. Traditional advertising layout relied primarily on the experience and aesthetic sense of designers for manual design. However, with the development of big data and artificial intelligence technologies, automated visual element layout based on user behavior characteristics and preferences has become possible.
[0003] Current commercial advertising visual element layout technologies mainly include template-based layout systems, rule-based automatic layout algorithms, and machine learning-based adaptive layout methods. Template-based layout systems quickly generate advertising content using preset design templates, offering the advantage of ease of use. Rule-based automatic layout algorithms arrange visual elements according to predefined design principles and constraints. Machine learning-based adaptive layout methods automatically optimize the arrangement of visual elements to improve advertising effectiveness by analyzing user behavior data and historical interaction records.
[0004] Existing technologies lack in-depth modeling of individual user cognitive preferences and mostly use coarse-grained optimization based on group statistical characteristics. They cannot provide truly personalized layout solutions for different users' visual perception habits and information processing methods, resulting in limited appeal and persuasiveness of advertising content to specific users.
[0005] Existing typography techniques typically treat visual elements as independent entities for layout, ignoring the semantic connections and content dependencies between elements. This fails to construct a coherent information delivery path, resulting in a lack of clear visual guidance for users when perceiving advertising information, thus affecting the effective delivery and conversion of advertising information. Summary of the Invention
[0006] This invention provides a method and system for automatically optimizing the layout of visual elements in commercial advertising design, which can solve the problems in the prior art.
[0007] A first aspect of the present invention provides a method for automatically optimizing the layout of visual elements in commercial advertising design, comprising:
[0008] Acquire multi-dimensional behavioral feature data and historical interaction records of target users, and extract the set of visual elements for the advertisement to be delivered; construct a user cognitive preference model based on the multi-dimensional behavioral feature data, and map the user cognitive preference model to a visual perception weight distribution;
[0009] Semantic association analysis is performed on each visual element in the set of visual elements to construct a dependency graph between visual elements. Based on the visual perception weight distribution and the dependency graph, the layout tension field distribution of each visual element in the page space is calculated.
[0010] Based on the layout tension field distribution and page size adaptation rules, a multi-level visual guidance path is generated, and the visual elements are adaptively allocated in space and arranged in presentation sequence based on the multi-level visual guidance path to generate a personalized layout scheme.
[0011] The personalized layout scheme is rendered as advertising content and delivered to target users. User interaction feedback data is collected after delivery. Visual dwell trajectory and conversion behavior features are extracted from the user interaction feedback data. The visual perception weight distribution and the dependency relationship graph are jointly updated using the visual dwell trajectory and the conversion behavior features.
[0012] Based on the aforementioned multi-dimensional behavioral feature data, a user cognitive preference model is constructed. Mapping this user cognitive preference model to a visual perception weight distribution includes:
[0013] From the multi-dimensional behavioral feature data, extract the user's browsing behavior sequence, interaction operation sequence, and content consumption sequence, and perform temporal feature encoding on the browsing behavior sequence, the interaction operation sequence, and the content consumption sequence to obtain a set of behavioral feature vectors;
[0014] Based on the set of behavioral feature vectors, the cognitive response patterns of users under different content types are identified through multimodal semantic aggregation, and a user cognitive preference model including cognitive sensitivity distribution and information acceptance tendency parameters is constructed.
[0015] The cognitive sensitivity distribution and information acceptance tendency parameters in the user cognitive preference model are projected onto the visual element feature space to establish a mapping relationship between the cognitive sensitivity distribution, the information acceptance tendency parameters, and visual element attributes.
[0016] Based on the mapping relationship, the cognitive influence weight coefficients corresponding to different types of visual elements are calculated, and a visual perception weight distribution reflecting the user's personalized cognitive characteristics is generated according to the cognitive influence weight coefficients.
[0017] Based on the aforementioned set of behavioral feature vectors, a user cognitive response pattern under different content types is identified through multimodal semantic aggregation. A user cognitive preference model, including parameters for cognitive sensitivity distribution and information acceptance tendency, is constructed, comprising:
[0018] Each behavioral feature vector in the behavioral feature vector set is labeled with a content type. Based on the content type labeling, the behavioral feature vector set is divided into multiple content type subsets, and each content type subset corresponds to a specific content presentation format.
[0019] For each content type subset, the user's interaction response intensity features and attention duration features are extracted to construct the correlation matrix between content type and cognitive response. Then, the correlation matrix of different content type subsets is semantically aligned and features are aggregated through a cross-modal feature fusion mechanism to obtain a unified multimodal cognitive response feature representation.
[0020] Based on the multimodal cognitive response feature representation, the cognitive sensitivity distribution curve of users to different visual stimulus intensities is calculated, and the information acceptance tendency parameter is extracted according to the information acceptance behavior pattern of users under different content types. The cognitive sensitivity distribution curve and the information acceptance tendency parameter are encapsulated in a structured manner to construct a user cognitive preference model.
[0021] Semantic correlation analysis is performed on each visual element in the set of visual elements to construct a dependency graph between visual elements. Based on the visual perception weight distribution and the dependency graph, the layout tension field distribution of each visual element in the page space is calculated, including:
[0022] Extract the semantic tags and functional attributes of each visual element in the set of visual elements, calculate the semantic similarity between the semantic tags and the functional attributes, and obtain the semantic association strength that represents the degree of content association between visual elements.
[0023] Based on the semantic association strength, a dependency graph between visual elements is constructed, and corresponding dependency weight values are assigned to the edges connecting different visual elements in the dependency graph according to the semantic association strength.
[0024] The visual perception weight distribution and the dependency graph are fused and mapped. For each visual element in the set of visual elements, the attractive potential energy and repulsive potential energy of each visual element in the page space are calculated based on its weight coefficient in the visual perception weight distribution and its topological connection relationship in the dependency graph.
[0025] By performing spatial field superposition calculations on the attractive potential energy and the repulsive potential energy, a layout tension field distribution is generated to guide the spatial arrangement of visual elements.
[0026] Based on the layout tension field distribution and page size adaptation rules, a multi-layered visual guidance path is generated. Then, based on this multi-layered visual guidance path, visual elements are adaptively allocated spatially and their presentation sequence is arranged to generate a personalized layout scheme, including:
[0027] Obtain the page size parameters and display characteristic parameters of the target projection device, and formulate page size adaptation rules based on the page size parameters and display characteristic parameters; identify the extreme value region of the tension field gradient in the page space according to the layout tension field distribution, and divide the extreme value region of the tension field gradient into multiple visual levels in combination with the page size adaptation rules;
[0028] For each visual level, visual guidance nodes are planned in descending order of tension field intensity, and a multi-level visual guidance path containing hierarchical jump relationships is generated by connecting the visual guidance nodes.
[0029] Based on the multi-level visual guidance path, a visual hierarchy and path position are assigned to each visual element, and the visual elements are adaptively allocated in combination with the content size requirements of each visual element and the space occupancy constraints in the layout tension field distribution.
[0030] Based on the hierarchical jump relationship in the multi-level visual guidance path and the user's cognitive rhythm parameters, the presentation sequence of each visual element is set, and the result of the adaptive spatial allocation is integrated with the presentation sequence to generate a personalized layout scheme that includes the spatial position information and timing control information of the visual elements.
[0031] For each visual level, visual guidance nodes are planned in descending order of tension field intensity. A multi-level visual guidance path containing hierarchical jump relationships is generated by connecting these visual guidance nodes, including:
[0032] Extract the layout tension field distribution data within the layout space area corresponding to each visual level, and identify multiple tension field intensity peak points within the visual level by performing intensity peak detection on the layout tension field distribution data.
[0033] The peak points of the tension field intensity are sorted in descending order according to the tension field intensity value. Based on the result of the descending sort, the peak points of the tension field intensity are selected as visual guidance nodes in sequence, and each visual guidance node is assigned a hierarchical priority identifier that represents the importance of the node.
[0034] Based on the priority identifier within the hierarchy, adjacent visual guidance nodes are connected sequentially according to the principle of minimum visual span, in order of priority from high to low, to generate an intra-hierarchical guidance path that represents the guidance order within the visual hierarchy.
[0035] For the guidance path within the level, identify the starting visual guidance node and the ending visual guidance node of each guidance path within the level, and establish a cross-level connection between the ending visual guidance node and the starting visual guidance node.
[0036] The intra-level guidance path and the cross-level connection are topologically combined to generate a multi-level visual guidance path containing complete hierarchical jump relationships.
[0037] The personalized layout scheme is rendered as advertising content and delivered to target users. User interaction feedback data is collected after the delivery, and visual dwell time and conversion behavior features are extracted from the user interaction feedback data, including:
[0038] Based on the spatial location information and timing control information of the visual elements in the personalized layout scheme, the graphics rendering engine is invoked to perform pixel-level drawing and compositing of each visual element, generating advertising display content that conforms to the personalized layout scheme, and the advertising display content is encapsulated into a delivery data package;
[0039] The data package is pushed to the target user's receiving terminal for display, and user interaction feedback data, including user gaze coordinate sequence, interaction operation timestamp sequence and conversion behavior record, is collected in real time during the ad display process;
[0040] Spatiotemporal correlation analysis is performed on the gaze coordinate sequence in the user interaction feedback data to identify the duration and order of the user's gaze on each visual element, and to construct a visual dwell trajectory that represents the user's visual cognition process.
[0041] The interaction operation timestamp sequence is matched and analyzed with the conversion behavior record to identify the complete operation link from the user's contact with the advertising content to the completion of the conversion behavior. Combined with the visual dwell trajectory, the interaction nodes and conversion trigger conditions in the complete operation link are extracted to construct conversion behavior features that characterize the user's response behavior.
[0042] A second aspect of the present invention provides an automatic layout optimization system for visual elements in commercial advertising design, comprising:
[0043] Module 1 is used to acquire multi-dimensional behavioral feature data and historical interaction records of target users, and extract the set of visual elements for the advertisement to be delivered; based on the multi-dimensional behavioral feature data, a user cognitive preference model is constructed, and the user cognitive preference model is mapped to a visual perception weight distribution;
[0044] Module 2 is used to perform semantic association analysis on each visual element in the set of visual elements, construct a dependency graph between visual elements, and calculate the layout tension field distribution of each visual element in the page space based on the visual perception weight distribution and the dependency graph.
[0045] Module 3 is used to generate multi-level visual guidance paths based on the layout tension field distribution and page size adaptation rules, and to adaptively allocate and arrange the presentation sequence of visual elements based on the multi-level visual guidance paths to generate personalized layout schemes.
[0046] Module 4 is used to render the personalized layout scheme into advertising content and deliver it to target users, collect user interaction feedback data after delivery, extract visual dwell trajectory and conversion behavior features from the user interaction feedback data, and use the visual dwell trajectory and conversion behavior features to jointly update the visual perception weight distribution and the dependency relationship graph.
[0047] A third aspect of the present invention provides an electronic device, comprising:
[0048] processor;
[0049] Memory used to store processor-executable instructions;
[0050] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0051] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0052] The beneficial effects of this application are as follows:
[0053] By acquiring multi-dimensional behavioral characteristic data of target users and constructing a user cognitive preference model, user preferences are mapped to visual perception weight distribution, achieving precise matching between advertising design and users' personalized needs, and improving the targeting and user acceptance of advertising.
[0054] Based on the semantic correlation analysis of visual elements, a dependency graph is constructed. Combined with the visual perception weight distribution, the layout tension field distribution is calculated, which solves the problem of the lack of scientific basis for the arrangement of visual elements in traditional advertising design, making the layout of visual elements more logical and attractive.
[0055] By generating multi-layered visual guidance paths and performing adaptive spatial allocation and presentation sequence arrangement, the dynamic presentation of advertising content is achieved, enhancing the smoothness of the user's visual experience and the effectiveness of information acquisition.
[0056] By collecting visual dwell time and conversion behavior characteristics from user interaction feedback data, and jointly updating the visual perception weight distribution and dependency relationship map, a closed-loop optimization mechanism is formed, enabling the ad layout scheme to continuously iterate and improve itself, thereby enhancing the ad placement effect. Attached Figure Description
[0057] Figure 1 A flowchart illustrating the automatic layout optimization method for visual elements in commercial advertising design, as described in an embodiment of the present invention.
[0058] Figure 2 This is a flowchart illustrating the user cognitive feature-driven visual perception weight allocation process according to an embodiment of the present invention. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0061] Figure 1 This is a flowchart illustrating the automatic layout optimization method for visual elements in commercial advertising design according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0062] Acquire multi-dimensional behavioral feature data and historical interaction records of target users, and extract the set of visual elements for the advertisement to be delivered; construct a user cognitive preference model based on the multi-dimensional behavioral feature data, and map the user cognitive preference model to a visual perception weight distribution;
[0063] Semantic association analysis is performed on each visual element in the set of visual elements to construct a dependency graph between visual elements. Based on the visual perception weight distribution and the dependency graph, the layout tension field distribution of each visual element in the page space is calculated.
[0064] Based on the layout tension field distribution and page size adaptation rules, a multi-level visual guidance path is generated, and the visual elements are adaptively allocated in space and arranged in presentation sequence based on the multi-level visual guidance path to generate a personalized layout scheme.
[0065] The personalized layout scheme is rendered as advertising content and delivered to target users. User interaction feedback data is collected after delivery. Visual dwell trajectory and conversion behavior features are extracted from the user interaction feedback data. The visual perception weight distribution and the dependency relationship graph are jointly updated using the visual dwell trajectory and the conversion behavior features.
[0066] In one optional implementation, constructing a user cognitive preference model based on the multi-dimensional behavioral feature data, and mapping the user cognitive preference model to a visual perception weight distribution includes:
[0067] From the multi-dimensional behavioral feature data, extract the user's browsing behavior sequence, interaction operation sequence, and content consumption sequence, and perform temporal feature encoding on the browsing behavior sequence, the interaction operation sequence, and the content consumption sequence to obtain a set of behavioral feature vectors;
[0068] Based on the set of behavioral feature vectors, the cognitive response patterns of users under different content types are identified through multimodal semantic aggregation, and a user cognitive preference model including cognitive sensitivity distribution and information acceptance tendency parameters is constructed.
[0069] The cognitive sensitivity distribution and information acceptance tendency parameters in the user cognitive preference model are projected onto the visual element feature space to establish a mapping relationship between the cognitive sensitivity distribution, the information acceptance tendency parameters, and visual element attributes.
[0070] Based on the mapping relationship, the cognitive influence weight coefficients corresponding to different types of visual elements are calculated, and a visual perception weight distribution reflecting the user's personalized cognitive characteristics is generated according to the cognitive influence weight coefficients.
[0071] like Figure 2 As shown, the method includes:
[0072] The collection of multi-dimensional behavioral feature data is achieved through a behavior monitoring module deployed on the client side. This module captures user actions such as page visits, clicks, scrolling, and dwell time in an event-driven manner. The extraction of browsing behavior sequences relies on timestamp-sorted page visit records. Each record includes fields such as the access URL, entry time, exit time, and source page. The data structure uses JSON format for storage, and the maximum number of records per session is set to 1000 to prevent memory overflow. Interaction operation sequences are generated by capturing mouse click coordinates, keyboard input events, and touch gestures. Each interaction event carries attributes such as operation type identifier, target element ID, and operation duration. The caching strategy uses a sliding window mechanism to retain operation records from the most recent 7 days. Content consumption sequences are constructed based on the user's consumption depth of different media types, including metrics such as text reading progress, video playback completion rate, and image viewing time. These are collected in real-time using event tracking technology and aggregated statistically by content ID.
[0073] Temporal feature encoding employs an attention-based sequence encoder to handle the temporal dependencies of behavioral data. The encoder input layer receives standardized behavioral sequences, with the sequence length uniformly truncated or padded to 512 time steps, and the padded values represented by zero vectors. Positional encoding is generated using sine and cosine functions, with a frequency parameter set to 10000, and the encoding dimension is consistent with the input feature dimension at 256 dimensions. The multi-head self-attention layer uses eight attention heads for parallel processing, with each attention head having a 32-dimensional query, key, and value matrix, and a dropout rate of 0.1 to prevent overfitting. The feedforward neural network layer contains two fully connected layers, with 1024 neurons in the intermediate layer, using ReLU activation, and the output layer dimension matches the encoder hidden state dimension. Layer normalization is applied to the output of each sub-layer, and residual connections ensure effective gradient propagation. During encoding, missing behavioral data is padded with the mean, and outliers are truncated using a rule of three times the standard deviation.
[0074] The behavioral feature vector set is generated by performing pooling operations on the encoded sequence to obtain a fixed-dimensional representation. Global average pooling extracts the overall features of the sequence, max pooling captures the most significant behavioral patterns, and attention-weighted pooling assigns weights according to importance. The results of these three pooling operations are concatenated to form a 768-dimensional comprehensive feature vector. The feature vector is subjected to L2 normalization to ensure numerical stability and reduced to 256 dimensions through principal component analysis to reduce computational complexity and storage overhead. Cosine similarity is used to calculate vector similarity, with a threshold set to 0.8 to identify similar behavioral patterns.
[0075] Multimodal semantic aggregation employs a hierarchical clustering algorithm to identify users' cognitive response patterns across different content types. It calculates the Euclidean distance between behavioral feature vectors to construct a distance matrix, initially treating each sample as an independent cluster. The clustering process uses the Ward connection criterion, merging the two clusters that minimize the increase in the sum of squares within each cluster, iterating until the number of clusters reaches a preset value or a distance threshold. The optimal number of clusters is determined using the elbow rule, calculating the sum of squares within clusters under different cluster numbers and selecting the cluster number corresponding to the inflection point where the decreasing trend significantly slows down. Clustering analysis is performed separately for different content types such as text, images, and videos, with the clustering results for each content type stored and managed independently.
[0076] The cognitive sensitivity distribution is calculated based on the differences in user response intensity to different types of content. Response intensity is obtained by weighted summation of behavioral indicators such as dwell time, interaction frequency, and repeat visits. The weighting coefficients are set according to the correlation between the indicators and user satisfaction: dwell time has a weight of 0.4, interaction frequency has a weight of 0.3, and repeat visits have a weight of 0.3. Sensitivity values are obtained by mapping response intensity to the 0-1 range, and the mapping function uses the Sigmoid function to ensure the reasonableness of the numerical range. The distribution is represented by a multi-peak Gaussian mixture model, and the model parameters are estimated using the expectation-maximization algorithm. The number of components is set to 3 to 5 to balance model complexity and fitting accuracy.
[0077] Information acceptance bias parameters reflect users' preferences for different information presentation methods. Parameter extraction relies on users' historical selection behavior, including dimensions such as content format preference, information density preference, and presentation order preference. Format preference is calculated by statistically analyzing the proportion of users selecting different content formats such as text, images, and videos. Density preference is based on assessing users' preference for concise and detailed content. Order preference is determined by analyzing users' browsing paths and attention distribution patterns. The parameter values for each dimension are normalized to the range of 0 to 1, and an exponentially weighted moving average update strategy is used to adapt to dynamic changes in user preferences. A smoothing coefficient is set to 0.1 to maintain parameter stability.
[0078] The user cognitive preference model is constructed using a deep neural network architecture. The input layer receives behavioral feature vectors, and the hidden layer contains three fully connected layers with 512, 256, and 128 neurons respectively, all using ReLU activation functions. The output layer has two branches: one outputs parameters of the cognitive sensitivity distribution, and the other outputs parameters of information acceptance tendency. The model is trained using the mean squared error loss function, and the optimizer uses the Adam algorithm. The initial learning rate is set to 0.001, and a cosine annealing scheduling strategy is used for dynamic adjustment. The training data is divided into training, validation, and test sets in an 8:1:1 ratio, with a batch size of 32 and 100 training epochs. An early stopping strategy is used to stop training if the validation set loss fails to improve for 10 consecutive epochs.
[0079] The visual element feature space is defined by including basic visual attribute dimensions such as color, shape, size, position, and texture. Color attributes are represented using the HSV color space, with hue ranging from 0 to 360 degrees and saturation and brightness ranging from 0 to 100. Shape attributes are encoded using geometric descriptors, including features such as aspect ratio, roundness, and convex hull ratio, each normalized to the range of 0 to 1. Size attributes record the element's pixel dimensions and convert them into a proportional value relative to the screen resolution. Position attributes use a relative coordinate system, with the screen center as the origin, and coordinate values normalized to the range of -1 to +1. Texture attributes are extracted using the gray-level co-occurrence matrix to obtain statistical features such as contrast, correlation, energy, and homogeneity.
[0080] The projection mapping process is achieved by establishing a non-linear relationship between cognitive preference parameters and visual attributes. The mapping function employs a multilayer perceptron network, with the input layer dimension matching the output dimension of the cognitive preference model and the output layer dimension matching the feature space dimension of visual elements. The hidden layer contains two fully connected layers with 128 and 64 neurons respectively, and the Tanh activation function is used to ensure the boundedness of the output. Supervised learning is used for network training. Label data is obtained through user preference ratings for different visual designs, ranging from 1 to 5 points. Cross-validation is used to evaluate the model's generalization performance.
[0081] The mapping relationship was established by analyzing the correlation between cognitive sensitivity distribution and information acceptance tendency parameters and visual element attributes. The Pearson correlation coefficient was used for correlation calculation, with a threshold set at 0.3 to filter significantly correlated attribute dimensions. For highly correlated attribute pairs, linear or nonlinear mapping functions were established. Linear mappings were fitted using the least squares method, while nonlinear mappings used multinomial regression or radial basis function networks. Mapping accuracy was evaluated using root mean square error, with an error threshold set at 0.05. Mapping relationships exceeding the threshold underwent retraining or parameter tuning.
[0082] The calculation of cognitive influence weight coefficients is based on the strength and stability of the mapping relationship. The weight coefficients are represented by the normalized absolute values of the correlation coefficients, ensuring that the sum of all coefficients is 1. For cases where multiple visual attributes influence the same cognitive dimension, multiple linear regression analysis is used to determine the independent contribution of each attribute, with the regression coefficients serving as the basis for weight allocation. An incremental learning strategy is employed for weight updates, with a weight of 0.2 for new data and 0.8 for historical data, ensuring a balance between model adaptability and stability.
[0083] The generation of the visual perception weight distribution applies cognitive influence weight coefficients to the evaluation of specific visual elements. The distribution is represented by a probability density function, and the support covers the entire range of visual attribute values. Weight calculation is performed by weighted summation of the cognitive influence weight coefficients and visual element attribute values, with normalization ensuring the weight sum is 1. Distribution updates employ a Bayesian inference framework, dynamically adjusting the weight parameters based on prior knowledge and observational data.
[0084] In one optional implementation, based on the set of behavioral feature vectors, a user cognitive response pattern under different content types is identified through multimodal semantic aggregation, and a user cognitive preference model including cognitive sensitivity distribution and information acceptance tendency parameters is constructed, including:
[0085] Each behavioral feature vector in the behavioral feature vector set is labeled with a content type. Based on the content type labeling, the behavioral feature vector set is divided into multiple content type subsets, and each content type subset corresponds to a specific content presentation format.
[0086] For each content type subset, the user's interaction response intensity features and attention duration features are extracted to construct the correlation matrix between content type and cognitive response. Then, the correlation matrix of different content type subsets is semantically aligned and features are aggregated through a cross-modal feature fusion mechanism to obtain a unified multimodal cognitive response feature representation.
[0087] Based on the multimodal cognitive response feature representation, the cognitive sensitivity distribution curve of users to different visual stimulus intensities is calculated, and the information acceptance tendency parameter is extracted according to the information acceptance behavior pattern of users under different content types. The cognitive sensitivity distribution curve and the information acceptance tendency parameter are encapsulated in a structured manner to construct a user cognitive preference model.
[0088] Each behavioral feature vector in the behavioral feature vector set is labeled with its content type. This set consists of behavioral data generated by users interacting with content, obtained through the preceding steps. Content type labeling is based on the content's presentation format, structural characteristics, and information delivery method. For example, content types can be categorized as mixed text and images, plain text, short videos, long videos, and audio. Each behavioral feature vector is labeled according to its corresponding content type attribute, forming a behavioral feature vector with a content type tag. The behavioral feature vector set is then divided into multiple content type subsets based on these content type labels, each subset corresponding to a specific content presentation format. For example, all behavioral feature vectors generated from interactions with short video content are grouped into the short video content type subset, and those generated from interactions with text and images are grouped into the text and images content type subset.
[0089] For each content type subset, user interaction response intensity and attention duration features are extracted. Interaction response intensity features include metrics such as user click frequency, dwell time, and interaction depth, reflecting the intensity of user interest in that content type. For example, for the short video content type subset, user completion rate, likes, comment frequency, and sharing behavior can be extracted as interaction response intensity features; for mixed text and image content, metrics such as reading depth, image click-through rate, and article completion rate can be extracted. Attention duration features primarily measure the user's attention allocation pattern across different content types, including single session duration, attention decay rate, and revisit frequency. After extracting these features, a correlation matrix between content type and cognitive response is constructed, where each element represents a specific cognitive response feature value for a particular content type.
[0090] A cross-modal feature fusion mechanism is used to semantically align and aggregate the association matrices of different content type subsets, unifying the feature dimensions and mapping features from different content types to a unified semantic space. For example, while the "completion viewing rate" of short videos and the "reading completion rate" of long articles differ in form, they are similar in cognitive meaning and can therefore be aligned to the unified dimension of "content completion" through semantic mapping. Subsequently, an attention mechanism is used for cross-modal feature fusion, assigning different weights to features of different content types based on the importance of each type of content in the user's historical interaction data. Finally, a weighted aggregation method is used to fuse the features of each content type to obtain a unified multimodal cognitive response feature representation. This feature representation is a high-dimensional vector containing the user's cognitive response features across various content types.
[0091] Based on multimodal cognitive response feature representation, we calculate the cognitive sensitivity distribution curves of users to different visual stimulus intensities. Visual stimulus intensity can be categorized into levels based on dimensions such as information density, visual complexity, and dynamics. For example, for the information density dimension, content can be divided into low, medium, and high levels based on the amount of information contained per unit area. By analyzing users' interaction patterns with content of different stimulus intensities, particularly changes in attention allocation and engagement, we plot the cognitive sensitivity distribution curves. These curves reflect users' acceptance and preference for visual stimuli of different intensities. The peak range of the curve represents the range of stimulus intensities most suitable for users, while the steepness of the curve reflects the user's sensitivity to changes in stimulus intensity.
[0092] Information acceptance preference parameters are extracted based on users' information reception behavior patterns across different content types. These parameters include: information acquisition rate preference, reflecting users' preferred information presentation speed; information depth preference, indicating users' choice of content depth; information structure preference, reflecting users' preference for linear or non-linear content organization; and multimedia element acceptance, indicating users' acceptance of multimedia elements such as images, audio, and video. By analyzing users' interactive behavior characteristics across different content types, the values of the above parameters are calculated to form a set of user information acceptance preference parameters.
[0093] A user cognitive preference model is constructed by structurally encapsulating the cognitive sensitivity distribution curve and information acceptance tendency parameters. This model adopts a hierarchical structure: the top layer consists of comprehensive cognitive preference features; the middle layer is divided into two main categories: cognitive sensitivity features and information acceptance features; and the bottom layer contains various specific parameter indicators. The model is serialized and converted into a unified data structure for easy system access and analysis. This user cognitive preference model can be applied to scenarios such as personalized content recommendation and user experience optimization, providing a basis for subsequent content adaptation and push strategies.
[0094] In practical applications, taking a news platform as an example, when user A's cognitive preference model shows that they have low cognitive sensitivity to high information density content but a high information acceptance tendency for video content, the platform will prioritize pushing video content with moderate information density and reduce the frequency of pushing text-intensive content. Through this content adaptation based on cognitive preferences, the platform improves the user's content consumption experience and platform stickiness.
[0095] In one optional implementation, semantic association analysis is performed on each visual element in the visual element set to construct a dependency graph between visual elements. Based on the visual perception weight distribution and the dependency graph, the layout tension field distribution of each visual element in the page space is calculated, including:
[0096] Extract the semantic tags and functional attributes of each visual element in the set of visual elements, calculate the semantic similarity between the semantic tags and the functional attributes, and obtain the semantic association strength that represents the degree of content association between visual elements.
[0097] Based on the semantic association strength, a dependency graph between visual elements is constructed, and corresponding dependency weight values are assigned to the edges connecting different visual elements in the dependency graph according to the semantic association strength.
[0098] The visual perception weight distribution and the dependency graph are fused and mapped. For each visual element in the set of visual elements, the attractive potential energy and repulsive potential energy of each visual element in the page space are calculated based on its weight coefficient in the visual perception weight distribution and its topological connection relationship in the dependency graph.
[0099] By performing spatial field superposition calculations on the attractive potential energy and the repulsive potential energy, a layout tension field distribution is generated to guide the spatial arrangement of visual elements.
[0100] Semantic label extraction for visual element sets is achieved by encoding the descriptive text of each visual element using a pre-trained natural language processing model. The label extraction module receives textual information such as the visual element's description, Alt attribute, and class name, and generates a 768-dimensional semantic vector representation using the word embedding layer of the BERT model. The preprocessing stage performs word segmentation, stop word removal, and stemming on the input text, with the vocabulary size limited to 50,000 to control memory usage. For visual elements lacking textual descriptions, labels are automatically generated using a visual recognition model, with a confidence threshold set to 0.7; labels below this threshold are marked as unknown. Semantic labels are organized using a hierarchical classification system. The top layer includes basic categories such as text, images, controls, and decorations, while the bottom layer is subdivided into specific functional labels such as buttons, links, titles, and body text. The hierarchy depth is limited to three levels to avoid excessive complexity.
[0101] Functional attribute extraction categorizes visual elements based on their interactivity and information delivery within the interface. The attribute extractor analyzes the element's DOM attributes, CSS styles, event binding, and other technical features to identify functional dimensions such as clickability, editability, navigation, and decorative properties. Clickability is determined by detecting click events, link address attributes, cursor styles, and other features, with a value ranging from 0 to 1 representing the probability of clickability. Editability is based on attributes such as input, text areas, and content editability, also represented by a probability value from 0 to 1. Navigability is evaluated by analyzing features such as link targets, menu structure, and breadcrumb placement. Decorative properties are calculated comprehensively based on the element's visual weight, positional importance, and content relevance. Each functional attribute dimension is scored independently and then concatenated to form a 32-dimensional functional attribute vector. Vector normalization ensures numerical stability.
[0102] Semantic similarity calculation uses cosine similarity to measure the relevance between semantic tag vectors. During the calculation, L2 normalization is applied to the semantic vectors to eliminate the influence of vector length on similarity. Batch calculation is used to improve efficiency in constructing the similarity matrix, with the number of visual elements per batch set to 256 to avoid memory overflow. For functional attribute similarity calculation, a weighted Euclidean distance metric is used, with weights allocated according to the importance of each attribute dimension: clickability 0.3, editability 0.25, navigation 0.25, and decorative 0.2. The similarity between semantic tags and functional attributes is weighted and fused to obtain the final semantic association strength, with semantic tags having a weight of 0.6 and functional attributes a weight of 0.4.
[0103] The semantic association strength threshold is used to filter meaningful element associations. A strong association threshold is set to 0.8, a medium association threshold to 0.5, and a weak association threshold to 0.3. Element pairs with association strength below the weak association threshold are not connected. The association strength calculation considers the spatial distance factor between elements, with the distance weight using a Gaussian decay function and the standard deviation parameter set to 10% of the screen diagonal length. The time decay factor is used to handle dynamically added or deleted visual elements. The initial association strength of new elements is set to 70% of the historical average, gradually adjusted to the true value as interaction data accumulates.
[0104] The dependency graph uses a directed graph data structure to store the relationships between visual elements. Nodes represent visual elements, containing basic information such as element ID, location coordinates, size, semantic tags, and functional attributes. Edges represent dependencies between elements, with attributes including semantic association strength, dependency type, and directionality. The graph is stored using an adjacency list, and node indexes use a hash table to achieve constant-time lookup complexity. Dynamic updates support adding and deleting nodes and edges; deleting a node automatically cleans up related edge connections, and adding a node automatically establishes new connections based on semantic association strength. The graph is persisted using JSON format and supports version management and incremental update mechanisms.
[0105] The assignment of dependency weights is based on the normalization of semantic association strength. The weight mapping uses a piecewise linear function: strong association intervals are mapped to 0.8 to 1.0, moderate association intervals to 0.5 to 0.8, and weak association intervals to 0.3 to 0.5. Weight values are retained to three decimal places to meet computational precision requirements. For multiple dependencies, a weight stacking strategy is used, with the total weight not exceeding an upper limit of 1.0. Dynamic adjustment of weights is based on user interaction feedback: positive feedback increases the weight by 0.05, and negative feedback decreases the weight by 0.03. A decay factor of 0.95 is set for the adjustment magnitude to avoid excessive fluctuations.
[0106] The fusion mapping module combines the visual perception weight distribution with the dependency graph. The mapping algorithm employs a graph convolutional network architecture. The input layer receives the visual perception weights and adjacency information of nodes. The convolutional layer updates the representation of the current node by aggregating the features of neighboring nodes. The aggregation function uses attention-weighted summation, and the attention weights are calculated based on the dependency weights of the edges. The network contains three graph convolutional layers with hidden layer dimensions of 128, 64, and 32, respectively. The activation function is LeakyReLU, and the negative slope parameter is set to 0.01. The output layer generates the fused element weight representation, with dimensions consistent with the input visual perception weight dimensions.
[0107] The analysis of topological connectivity relationships uses a graph traversal algorithm to identify direct and indirect dependencies between elements. Breadth-first search is used to calculate the shortest path length between nodes, and the weight of indirect dependencies with a path length exceeding 3 is reduced by 50%. Centrality analysis uses degree centrality, proximity centrality, and betweenness centrality to evaluate the importance of nodes in the graph. Degree centrality counts the number of connecting edges to a node, proximity centrality is calculated based on the average distance from a node to all other nodes, and betweenness centrality counts the number of shortest paths through that node. The three centrality indicators are weighted and fused to form a comprehensive centrality score, with weights of 0.4, 0.3, and 0.3 respectively.
[0108] The calculation of attractive potential energy is based on the visual perception weights and graph centrality scores of visual elements. A gravity model is used for potential energy calculation, where the strength of the attraction is directly proportional to the product of the weight coefficients and the centrality score, and inversely proportional to the square of the spatial distance. The effective range of the potential energy is set to 5 times the element size; the attraction strength decays to zero beyond this range. The potential energy field is discretized using a two-dimensional grid with a grid resolution set to one-quarter of the screen pixel resolution. The potential energy value of a single grid is calculated by superimposing the potential energy of all influencing elements. The potential energy values are normalized to the interval between 0 and 1 using a maximum-minimum normalization method.
[0109] The calculation of repulsive potential energy considers the competitive relationship and spatial conflict between visual elements. The intensity of the repulsive force is calculated based on the functional similarity of the elements; elements with similar functions generate a stronger repulsive force to avoid functional confusion. The repulsive force model uses an exponential decay function, with the decay constant set according to the element type: 0.1 for text elements, 0.15 for image elements, and 0.2 for control elements. The range of the repulsive force is affected by the element size; larger elements have a wider repulsive range. Boundary repulsive force handles the constraints of screen edges on element layout; the intensity of the repulsive force near the boundary increases linearly with distance, and the repulsive force is activated when the distance to the boundary is less than 50 pixels.
[0110] The spatial field superposition operation merges attractive and repulsive potential energy through vector summation. During superposition, different types of potential energy are balanced using weighting coefficients: attractive force weight is 0.6, and repulsive force weight is 0.4. These weighting ratios can be adjusted according to layout preferences. A Gaussian filter is used for smoothing the potential energy field, with a kernel size of 5x5 pixels and a standard deviation of 1.5 pixels, eliminating noise and abrupt changes in the field distribution. The Sobel operator is used to calculate the gradient of the field intensity to determine the optimal movement direction of the elements. The gradient magnitude indicates the drastic change in the field, and the gradient direction indicates the direction of the fastest decrease in potential energy.
[0111] The tension field distribution in the layout is visualized using a heatmap. The heatmap's color mapping uses a gradient from blue to red, with blue representing low-tension areas and red representing high-tension areas. Tension values are graded using a quantile method, dividing the tension field into 10 levels, each corresponding to a different color depth. Real-time updates of the tension field support interactive adjustments; user actions trigger a recalculation of the field distribution, with update latency controlled within 100 milliseconds to ensure a smooth interactive experience. The field distribution data caching strategy uses an LRU algorithm, with the cache size set to the 50 most recently used field distribution states, maintaining a hit rate above 80%.
[0112] In one optional implementation, a multi-layered visual guidance path is generated based on the layout tension field distribution and page size adaptation rules. Then, based on the multi-layered visual guidance path, visual elements are adaptively allocated spatially and their presentation sequence is arranged to generate a personalized layout scheme, including:
[0113] Obtain the page size parameters and display characteristic parameters of the target projection device, and formulate page size adaptation rules based on the page size parameters and display characteristic parameters; identify the extreme value region of the tension field gradient in the page space according to the layout tension field distribution, and divide the extreme value region of the tension field gradient into multiple visual levels in combination with the page size adaptation rules;
[0114] For each visual level, visual guidance nodes are planned in descending order of tension field intensity, and a multi-level visual guidance path containing hierarchical jump relationships is generated by connecting the visual guidance nodes.
[0115] Based on the multi-level visual guidance path, a visual hierarchy and path position are assigned to each visual element, and the visual elements are adaptively allocated in combination with the content size requirements of each visual element and the space occupancy constraints in the layout tension field distribution.
[0116] Based on the hierarchical jump relationship in the multi-level visual guidance path and the user's cognitive rhythm parameters, the presentation sequence of each visual element is set, and the result of the adaptive spatial allocation is integrated with the presentation sequence to generate a personalized layout scheme that includes the spatial position information and timing control information of the visual elements.
[0117] The target device's layout size parameters are obtained through a device detection interface, which returns key parameters such as screen physical size, logical resolution, and pixel density. Screen width and height are recorded in pixels, ranging from 320 to 4096 pixels with integer pixel precision. Pixel density is expressed in dots per inch (dpi), with a common range of 72 to 400, affecting the clarity of text and image rendering. Device orientation detection supports switching between landscape and portrait modes, dynamically updating size parameters by listening to orientation events. An exception handling mechanism includes default value settings for device information acquisition failures: a default width of 1024 pixels, a default height of 768 pixels, and a default pixel density of 96.
[0118] Display performance parameters include technical indicators that affect visual presentation, such as color gamut support, brightness range, contrast ratio, and refresh rate. Color gamut support is obtained by detecting the device's color profile and supports standard color gamuts such as sRGB, DCI-P3, and Rec2020, used to adjust color rendering strategies. Brightness parameters are recorded in nits, with a typical range of 400 to 1000 nits for mobile devices and 250 to 400 nits for desktop monitors. Contrast ratio affects text readability and image depth, with a standard dynamic contrast ratio range of 1000:1 to 5000:1. Refresh rate determines animation smoothness, with common values including 60Hz, 120Hz, and 144Hz. Parameter caching strategies use device fingerprinting to avoid redundant detection and resource consumption.
[0119] The layout size adaptation rules are based on device type classification and size threshold division. For mobile devices with a width less than 768 pixels, a single-column layout is used, with the content area occupying 95% of the screen width and left and right margins each occupying 2.5%. For tablet devices with a width of 768 to 1024 pixels, a two-column layout is used, with the main content area occupying 70% of the width, the sidebar occupying 25%, and spacing occupying 5%. For desktop devices with a width greater than 1024 pixels, a multi-column layout is used, supporting 3 or 4 columns, with a fixed column spacing of 20 pixels. Vertical adaptation considers the concept of a safe area: 44 pixels are reserved for the top status bar, 88 pixels for the bottom operation area, and the remaining space is used as the usable content area. Font size adaptation uses relative units; the basic font size is calculated based on screen density: 14 pixels for low-density screens, 16 pixels for medium-density screens, and 18 pixels for high-density screens.
[0120] Tension field gradient extremum region identification is achieved through gradient analysis of the layout tension field distribution. Gradient calculation uses the Sobel operator to calculate the partial derivative of the two-dimensional distribution of the tension field, obtaining the gradient components in the horizontal and vertical directions. The gradient magnitude is calculated by taking the square root of the sum of the squares of the two components, representing the drastic change in tension. Extremum detection employs a non-maximum suppression algorithm, comparing gradient magnitudes within a 3x3 neighborhood and retaining local maxima as gradient extrema. Extremum region division uses a clustering algorithm to group adjacent extremum points into the same region, with a clustering distance threshold set at 50 pixels to ensure spatial continuity of the regions. Extremum regions with excessively small areas are merged, with the minimum region area threshold set at 2% of the screen area.
[0121] Visual hierarchy is implemented using a combination of tension field intensity distribution and layout size adaptation rules to achieve layered management. The number of layers is determined by the device type: 3 layers for mobile devices, 4 layers for tablets, and 5 layers for desktop devices. The hierarchy is divided using the quantile method of tension field intensity, dividing the intensity distribution into different layers based on percentiles. The first layer corresponds to the 90th to 100th percentile, the second layer to the 70th to 90th percentile, and so on. Each layer is allocated a fixed proportion of space: the first layer occupies 40% of the available area, the second layer occupies 30%, the third layer occupies 20%, and the remaining layers share the remaining space equally. A soft boundary strategy is used for layer boundaries, allowing for a 10% overlap area for smooth transitions. Layer priority affects the rendering order and interactive response of elements; higher-level elements have higher Z-axis coordinate values.
[0122] Visual guidance node planning determines node positions based on the decreasing order of tension field intensity. The node planning algorithm traverses the tension field distribution within each visual level, sorting key positions by intensity value from largest to smallest. Node spacing control ensures the continuity of visual guidance, with a minimum spacing of 150 pixels and a maximum spacing limit of 300 pixels. Node density is adjusted according to the importance of the level, with the highest node density in the first level and an average spacing of 180 pixels. Lower levels have appropriately increased node spacing to reduce visual burden. Node position optimization considers aesthetic principles of layout, avoiding nodes being too concentrated or scattered, and uses a force-guided algorithm to fine-tune point coordinates. Node attributes include coordinate position, level affiliation, intensity value, and influence radius, with the influence radius determined based on the attenuation characteristics of the tension field.
[0123] Multi-level visual guidance path generation is achieved by connecting visual guidance nodes at different levels. The path connection algorithm employs shortest path search, comprehensively considering factors such as distance between nodes, level jump cost, and visual continuity. Within the same level, a greedy algorithm is used to connect nodes, selecting the closest and visually continuous nodes. Cross-level connections establish jump relationships between levels; the jump probability is related to the difference in level importance, with a jump probability of 0.8 between adjacent levels and decreasing by 20% across levels. Path smoothing uses Bézier curve fitting to connect straight lines, with curve control points determined based on the tangent direction of the nodes to ensure visual smoothness. Path conflict detection avoids the intersection and overlap of multiple paths; conflict resolution strategies include path offset, node adjustment, and level reallocation.
[0124] Visual hierarchy assignment determines the hierarchy and path position for each visual element. The assignment algorithm is based on a comprehensive evaluation of the element's visual perception weight and functional importance. Elements with a weight higher than 0.7 are assigned to the first level, weights between 0.4 and 0.7 to the second level, weights between 0.2 and 0.4 to the third level, and weights lower than 0.2 to lower levels. Path position assignment uses the Hungarian algorithm to find the optimal match. The objective function comprehensively considers factors such as element-node fit, path load balancing, and visual coordination. Fit is calculated based on the degree of matching between element type and node characteristics; text elements are suitable for assignment to nodes with high stability, and interactive elements are suitable for assignment to nodes with high reachability. Load balancing ensures that the number of elements carried by each path is relatively even, avoiding overload on some paths and affecting the user experience.
[0125] Adaptive space allocation achieves dynamic layout by combining the content size requirements of visual elements with the space occupancy constraints of the layout tension field. Size requirement calculation is based on the element's content type and display requirements; text elements have their width calculated based on the number of characters and font size, image elements maintain aspect ratio constraints, and control elements adhere to the minimum size requirements of interface design specifications. Space occupancy constraints are derived from the intensity gradient of the tension field distribution; high-intensity areas allow for greater space occupancy, while low-intensity areas restrict space usage. The allocation algorithm employs heuristic search, optimizing space utilization and visual harmony through simulated annealing. Conflict resolution mechanisms handle situations of insufficient space or overlap, using strategies such as element scaling, position adjustment, and hierarchy shifting. Boundary handling ensures that elements do not exceed the available area, automatically cropping or scrolling out any excess portions.
[0126] The presentation sequence is determined by the hierarchical jump relationship of the multi-level visual guidance path and user cognitive rhythm parameters to determine the element display order. Cognitive rhythm parameters include user-personalized characteristics such as attention duration, information processing speed, and visual switching frequency. Attention duration affects the display duration of a single element, ranging from 1 to 8 seconds, with a default value of 3 seconds. Information processing speed determines the switching interval between elements: 0.5 seconds for fast-processing users, 1 second for average users, and 2 seconds for slow-processing users. The sequence arrangement adopts a hierarchical priority strategy, with first-level elements displayed first, followed by elements within the same level according to the path order. Additional transition time is added when jumping between levels; the transition duration is determined based on the level difference: 0.3 seconds for adjacent level transitions and an additional 0.2 seconds for cross-level transitions. Animation effects configuration supports multiple methods such as fade-in / fade-out, slide-in, and zoom display; the animation duration is uniformly set to 0.4 seconds to ensure smoothness.
[0127] Personalized layout schemes integrate adaptive space allocation results and presentation timing information to generate complete layout configurations. The scheme data structure includes fields such as element identifier, spatial coordinates, size parameters, hierarchy information, and timing configuration. Spatial coordinates use relative positioning, with the top-left corner of the container as the origin, and coordinate values are in percentage form to adapt to size variations across different devices. Size parameters include constraints such as width, height, minimum size, and maximum size, supporting responsive adjustments. Hierarchy information records the element's z-index value and visibility status, while timing configuration includes control parameters such as display time, animation type, and trigger conditions. Scheme serialization uses JSON format, supporting compressed storage and fast parsing, with a single scheme file size kept under 100KB. Version control supports incremental updates and rollback operations, retaining the 10 most recent versions for comparative analysis.
[0128] In one optional implementation, for each visual level, visual guidance nodes are planned in descending order of tension field intensity, and a multi-level visual guidance path containing hierarchical jump relationships is generated by connecting the visual guidance nodes, including:
[0129] Extract the layout tension field distribution data within the layout space area corresponding to each visual level, and identify multiple tension field intensity peak points within the visual level by performing intensity peak detection on the layout tension field distribution data.
[0130] The peak points of the tension field intensity are sorted in descending order according to the tension field intensity value. Based on the result of the descending sort, the peak points of the tension field intensity are selected as visual guidance nodes in sequence, and each visual guidance node is assigned a hierarchical priority identifier that represents the importance of the node.
[0131] Based on the priority identifier within the hierarchy, adjacent visual guidance nodes are connected sequentially according to the principle of minimum visual span, in order of priority from high to low, to generate an intra-hierarchical guidance path that represents the guidance order within the visual hierarchy.
[0132] For the guidance path within the level, identify the starting visual guidance node and the ending visual guidance node of each guidance path within the level, and establish a cross-level connection between the ending visual guidance node and the starting visual guidance node.
[0133] The intra-level guidance path and the cross-level connection are topologically combined to generate a multi-level visual guidance path containing complete hierarchical jump relationships.
[0134] The extraction of layout tension field distribution data corresponding to the visual hierarchy of the page space is achieved through a region clipping algorithm. The region definition module receives the hierarchy boundary coordinate parameters, with the boundaries described by rectangles or polygons, and the coordinate precision retaining two decimal places. The data extractor traverses the complete two-dimensional array of tension field distributions, filtering out tension field values whose coordinates fall within the specified regions. Tension field data is stored as floating-point numbers, ranging from 0 to 1, with a precision of 0.001. Data points outside the regions are marked as invalid values to avoid interfering with subsequent calculations. The memory optimization strategy uses sparse matrix storage, saving only non-zero and boundary values, typically achieving a compression rate of over 60%. The region caching mechanism stores the extracted distribution data in memory-mapped files, supporting fast reading and updating, with the size of a single region data file controlled within 2MB.
[0135] Peak intensity detection employs a local maximum search algorithm to identify significant peak points in the tension field intensity. The detection window is set to a 5x5 pixel square neighborhood, and the tension field intensity of the center pixel must be strictly greater than that of all other pixels in the neighborhood to be considered a peak point. Boundary processing uses a mirror filling strategy, with the neighborhood of boundary pixels mapped through mirroring. The peak intensity threshold is set to the mean tension field intensity within the region plus twice the standard deviation, filtering out noise and weak peaks. The spatial distribution density of peak points is controlled through a minimum spacing constraint; the distance between two peak points must not be less than 30 pixels, and peak points that are too close are retained as the one with the greater intensity. The peak detection results include attributes such as coordinate position, intensity value, neighborhood range, and confidence score. The confidence score is calculated based on the sharpness of the peak and the surrounding intensity gradient.
[0136] The descending sorting of peak tension field intensity points utilizes a quicksort algorithm to process the intensity value sequence. The sorting key is the tension field intensity value of the peak point, maintaining a precision of 0.001 decimal places to ensure sorting stability. Peak points with the same intensity value use their coordinate values as a secondary sorting key, prioritizing points with smaller x-coordinates to guarantee the uniqueness and reproducibility of the sorting results. The sorted peak points are stored in a linear array, with array indices directly corresponding to the intensity ranking, and index 0 representing the highest intensity peak point. The time complexity of the sorting process is O(n log n), where n is the number of peak points. Typical processing capability supports sorting 1000 peak points within 10 milliseconds. The sorting results are persisted using binary format storage, containing a compact representation of peak point coordinates, intensity values, and rankings.
[0137] The selection of visual guidance nodes is based on a descending sorting result and a node number limitation strategy. The number of nodes is determined according to the importance of the visual hierarchy: a maximum of 8 nodes for the first hierarchy, a maximum of 6 nodes for the second hierarchy, a maximum of 4 nodes for the third hierarchy, and a limit of 2 nodes for lower levels. The selection algorithm starts from the top of the sorted list of peak points and sequentially selects peak points that meet the spatial distribution requirements as guidance nodes. Spatial distribution checks ensure that the node spacing is not less than 80 pixels to avoid overly dense nodes affecting the visual guidance effect. Node coverage analysis calculates the effective influence area of the selected nodes, requiring coverage of more than 80% of the spatial area of the hierarchy; additional nodes are added if insufficient coverage is achieved. Node quality evaluation comprehensively considers factors such as intensity value, spatial location, and neighborhood characteristics; nodes with a quality score below 0.6 are not selected.
[0138] Within each hierarchy, a priority identifier is assigned to each visual guidance node, representing its importance. Priority calculation uses a normalized intensity ranking, with the highest-intensity node having a priority of 1.0, and the priorities of other nodes decreasing linearly according to their intensity. Priority precision is retained to three decimal places to ensure accurate identification of priority differences between different nodes. The identifier is in string format, containing information such as the hierarchy number, node sequence number, and priority value, in the format "L{level}N{sequence number}P{priority}", for example, "L1N3P0.750" indicates that the third node in the first hierarchy has a priority of 0.750. A dynamic priority adjustment mechanism corrects node importance based on user interaction feedback: positive feedback increases priority by 0.05, and negative feedback decreases priority by 0.03, with an adjustment decay coefficient of 0.9 to prevent excessive fluctuations. Priority persistent storage supports fast querying and batch updates, using a key-value database for management, with query latency controlled within 1 millisecond.
[0139] The minimum visual span principle guides the connection strategy for adjacent visual guide nodes. The visual span is defined as a weighted combination of the Euclidean distance and the difference in visual weights between two nodes, with a distance weight of 0.7 and a weight difference of 0.3. The connection algorithm employs a greedy strategy, starting with the highest priority node and sequentially selecting the next node that satisfies the minimum span condition for connection. The span threshold is set to 25% of the diagonal length of the hierarchical space; connections exceeding this threshold are considered to have excessive visual spans, requiring the insertion of intermediate nodes or path replanning. Connection constraint checks prevent path intersections and loops, and a line segment intersection detection algorithm verifies the validity of new connections. Path smoothing uses B-spline curve fitting to connect straight lines, with control points determined based on the local gradient direction of the nodes. The curve order is set to 3 to ensure smoothness.
[0140] Within a hierarchy, guided paths are generated by connecting visually guided nodes in a series to form an ordered access sequence. The path data structure uses a directed graph, where nodes are visually guided nodes, edges represent connections between nodes, and edge weights are the visual span values. The path traversal algorithm uses depth-first search, starting from the highest priority node and visiting subsequent nodes sequentially according to their connections until the terminal node is reached. Path uniqueness is guaranteed through topological sorting to ensure the path is a directed acyclic graph structure. Path length is limited to no more than 150% of the total number of nodes within a single hierarchy; excessively long paths are simplified by removing redundant nodes and connections. Path quality is evaluated using metrics such as total path length, average span, and coverage. Paths with a quality score below 0.7 trigger a replanning mechanism.
[0141] The initial visual guidance node identification determines the path entry point by analyzing the topology of the guidance path within the hierarchy. An initial node is defined as a node with an in-degree of zero, meaning it has no connections from other nodes. In the case of multiple initial nodes, the node with the highest priority is selected as the primary initial node, and the remaining nodes are marked as secondary initial nodes. Initial node attributes include coordinates, priority identifier, number of connections, and influence range. The influence range is a circular area centered on the node, with a radius equal to the node's strength value multiplied by 100 pixels. Initial node stability analysis assesses the degree of change in node position during dynamic updates; nodes with changes exceeding 20 pixels are marked as unstable nodes, affecting the establishment of cross-level connections.
[0142] The termination visual guidance node determines the path exit by identifying the terminal node of the guidance path within the hierarchy. A termination node is defined as a node with an out-degree of zero, i.e., a node with no connections to other nodes. The selection of the termination node considers its spatial proximity to the starting node of the next level, and Manhattan distance is used for distance calculation to avoid the complexity of diagonal connections. The transition preparation for the termination node includes visual cue settings, attention guidance configuration, and hierarchy switching animation parameters. The transition animation duration is set to 0.6 seconds, and animation types support fade-out, slide, and zoom effects. Animation parameters are adjusted according to the hierarchy importance. The termination node's buffering mechanism delays hierarchy switching under high load; the buffering time is dynamically adjusted based on the processing queue length, with a maximum buffering time limit of 2 seconds.
[0143] Cross-level connections establish bridging relationships between terminating visual guidance nodes and the starting visual guidance nodes of the next level. Connection weights are calculated based on factors such as node distance, priority difference, and visual coherence: distance weight is 0.4, priority weight is 0.3, and coherence weight is 0.3. A connection filtering mechanism filters connections that are too far apart or visually discontinuous, with a distance threshold set to 40% of the screen diagonal length and a coherence threshold set to 0.5. In the case of multiple candidate connections, a comprehensive scoring ranking is used to select the optimal connection, with the scoring function considering all weighted factors. Connection redundancy checks avoid duplicate and circular connections, and graph theory algorithms are used to verify connection validity. Dynamic connection maintenance supports automatic updates when node positions change; updates are triggered when node displacement exceeds 30 pixels or priority change exceeds 0.1.
[0144] Topology combination processing integrates intra-level guidance paths and cross-level connections into a complete multi-level visual guidance path. The combination algorithm employs a graph merging strategy, connecting multiple independent intra-level path graphs through cross-level edges to form a unified directed graph structure. Path integrity checks verify the existence of a connected path from the highest-level starting node to the lowest-level ending node; if not connected, supplementary connections are added or node layouts are adjusted. Path optimization removes redundant connections and simplifies complex paths, aiming to minimize the total path length and maximize visual fluency. Global path quality evaluation uses comprehensive indicators such as path efficiency, coverage integrity, and visual consistency; paths with a quality score below 0.8 are replanned.
[0145] In one optional implementation, the personalized layout scheme is rendered as advertising content and delivered to target users. User interaction feedback data is collected after the delivery, and the visual dwell time and conversion behavior characteristics extracted from the user interaction feedback data include:
[0146] Based on the spatial location information and timing control information of the visual elements in the personalized layout scheme, the graphics rendering engine is invoked to perform pixel-level drawing and compositing of each visual element, generating advertising display content that conforms to the personalized layout scheme, and the advertising display content is encapsulated into a delivery data package;
[0147] The data package is pushed to the target user's receiving terminal for display, and user interaction feedback data, including user gaze coordinate sequence, interaction operation timestamp sequence and conversion behavior record, is collected in real time during the ad display process;
[0148] Spatiotemporal correlation analysis is performed on the gaze coordinate sequence in the user interaction feedback data to identify the duration and order of the user's gaze on each visual element, and to construct a visual dwell trajectory that represents the user's visual cognition process.
[0149] The interaction operation timestamp sequence is matched and analyzed with the conversion behavior record to identify the complete operation link from the user's contact with the advertising content to the completion of the conversion behavior. Combined with the visual dwell trajectory, the interaction nodes and conversion trigger conditions in the complete operation link are extracted to construct conversion behavior features that characterize the user's response behavior.
[0150] The spatial location information of visual elements in personalized layout schemes is parsed through a scheme file parser. The parser reads the JSON-formatted layout configuration file and extracts the coordinates, size, and layer position parameters of each visual element. The coordinate system uses relative positioning, with the top left corner of the canvas as the origin. The coordinate values are in percentage format, with precision retained to three decimal places. The size parameters include pixel values or percentage values for width and height, and support minimum and maximum size constraints. Layer information records the Z-axis coordinates and rendering priority of the element, with a value range of 0 to 100, where a larger value indicates a higher layer. Parsing exception handling includes format errors, out-of-bounds values, and missing required fields. In case of exceptions, the default configuration is used or the erroneous element is ignored and processing continues. The parsing results are cached in an in-memory hash table, with the element identifier as the key and the position attribute structure as the value. The query time complexity is O(1).
[0151] The timing control information extraction identifies the display time, animation parameters, and interaction trigger conditions for each visual element. Display time includes start time and duration, measured in milliseconds with a precision of 1 millisecond. The start time is calculated relative to the page load time. Animation parameters include animation type, duration, easing function, and keyframe data. Animation types support basic transformations such as fade-in / fade-out, sliding, scaling, and rotation. The easing function uses Bézier curves, with control point coordinates ranging from 0 to 1, and presets common curves such as linear, ease-in, ease-out, and ease-in / ease-out. Interaction trigger conditions include trigger mechanisms such as mouse events, keyboard events, scroll events, and time delays. Event binding uses the observer pattern for decoupling. The timing manager maintains a global timeline, supporting pause, resume, and jump operations, with timing precision controlled within 16 milliseconds to ensure a smooth 60 frames per second.
[0152] The graphics rendering engine achieves cross-platform compatibility through a rendering interface abstraction layer. This layer encapsulates the differences in the underlying graphics APIs and supports various rendering backends such as Canvas 2D, WebGL, DirectX, and OpenGL. The rendering context is initialized with parameters such as canvas size, color depth, and anti-aliasing level. The canvas size is determined based on the target device resolution, and the color depth defaults to 32-bit RGBA format. Anti-aliasing employs multi-sampling technology, with configurable sampling levels of 2x, 4x, and 8x. Higher levels result in better quality but also higher performance costs. The rendering queue manages the submission and execution of drawing tasks, using double buffering to avoid screen flicker. Buffer switching occurs during vertical synchronization. Rendering performance monitoring records metrics such as frame rate, drawing time, and memory usage. Performance optimization strategies are triggered when the frame rate drops below 30fps.
[0153] Pixel-level rendering achieves precise visual element rendering through primitive assembly and a rasterization pipeline. Primitive assembly converts the geometric description of visual elements into rendered primitives, with basic primitives such as points, lines, and triangles supporting the combined representation of complex shapes. Vertex transformation applies a model-view projection transformation matrix to achieve geometric transformations such as translation, rotation, scaling, and perspective of elements. The transformation matrix uses 4x4 homogeneous coordinates and 32-bit floating-point precision for calculation. The rasterization stage converts geometric primitives into pixel fragments, applying interpolation to determine the color, texture coordinates, normal vector, and other attributes of each pixel. Fragment shaders handle pixel-level effects such as texture sampling, lighting calculations, and color blending, and support custom shader programs to achieve special visual effects.
[0154] The compositing process enables the hierarchical combination of multiple visual elements and the generation of the final image. The compositor synthesizes each element sequentially from the bottom layer to the top layer, employing an alpha blending algorithm to handle semi-transparent effects. Blending modes support various effects such as normal, multiplicative, screen, overlay, and soft light. The blending formula calculates the final result based on the color values of the source and target pixels. Color space management ensures correct conversion between different color formats, supporting standard color gamuts such as sRGB, Adobe RGB, and DCI-P3. Gamma correction applies appropriate gamma values to correct color display, typically 2.2 or 2.4. The compositing result is output in bitmap format, supporting compression formats such as PNG, JPEG, and WebP, with compression quality configurable from 0 to 100.
[0155] The ad display content generation integrates rendering results and interactive logic to form a complete display unit. The content packager packages rendered bitmap data, animation sequences, interactive scripts, style sheets, and other resources into a single file. Resource compression uses lossless compression algorithms to reduce file size, typically achieving a compression rate of 40% to 60%. Metadata records content version information, creation time, target device type, compatibility identifiers, and other attributes for delivery strategies and cache management. Content verification checks rendering quality, file integrity, format compatibility, and other metrics; failure to verify triggers regeneration or downgrade processing. The generated content file uses a digital signature to ensure integrity and source trustworthiness; the signature algorithm uses RSA 2048-bit or ECDSA P-256.
[0156] The data packet encapsulation combines the ad display content and delivery control information into a transmission unit. The packet header includes fields such as version identifier, content type, compression method, encryption identifier, and checksum, with a fixed header length of 64 bytes. The payload stores the compressed display content and supports chunked transmission to handle large files, with a single chunk size limit of 64KB. Delivery control information includes constraints such as target user identifier, display time window, frequency limit, and geographic restrictions. Data packet compression uses the zlib algorithm, with a compression level set to 6 to balance compression ratio and speed. Multiple transmission protocols are supported, including HTTP / HTTPS, TCP, and UDP, with the appropriate protocol selected based on network conditions and latency requirements.
[0157] The push data packet delivery system transmits content to the target user's receiving terminal via a distribution network. The push server maintains the connection status of user devices and the push queue, supporting both real-time push and offline caching modes. Connection management employs persistent connections and a heartbeat mechanism to maintain channel availability, with a heartbeat interval of 30 seconds and a connection timeout of 60 seconds. The push queue is implemented using a priority queue, pushing higher-priority content first, and the queue length is limited to 1000 entries to prevent memory overflow. A retry mechanism handles push failures, employing an exponential backoff strategy with a maximum of 3 retries. Push statistics record metrics such as success rate, latency, and bandwidth usage for network optimization and fault diagnosis.
[0158] User interaction feedback data collection utilizes client-side event tracking technology to gather user behavior information in real time. Eye-tracking employs a front-facing camera and eye-tracking algorithms to identify the coordinates of the user's gaze point on the screen. The gaze point coordinates are measured in screen pixels with a precision of 1 pixel, and a sampling frequency of 30Hz ensures trajectory continuity. The eye-tracking algorithm is based on pupil detection and geometric calculations. Pupil detection uses a Haar feature classifier, and geometric calculations employ corneal reflectance to estimate the gaze direction. Interaction monitoring includes mouse clicks, swipes, double-clicks, and long presses. Event timestamps are recorded using a high-precision timer with a precision of 1 millisecond. Operation coordinates are recorded relative to the advertising content, facilitating cross-device data analysis.
[0159] The conversion behavior tracking system tracks the entire process from user exposure to ad engagement to completion of a target action. Behavior definitions include quantifiable user actions such as clicks, visits, registrations, purchases, and sharing. Each behavior is assigned a unique identifier and weight value. Behavior weights reflect business value: clicks are weighted at 1 point, visits at 3 points, registrations at 10 points, and purchases at 50 points. Behavior timestamps record the occurrence time with millisecond precision, facilitating time-series analysis and funnel calculations. Behavior attributes record contextual information such as triggering conditions, completion status, and related elements, stored in key-value pair format. Data transmission employs an asynchronous batch upload strategy to reduce network overhead, with a batch size of 100 records and an upload interval of 5 seconds.
[0160] The spatiotemporal correlation analysis of gaze coordinate sequences is achieved through trajectory reconstruction and region mapping algorithms. Trajectory reconstruction connects discrete gaze points into continuous gaze movement paths, and spline interpolation is used to fill sampling gaps. The interpolation algorithm uses cubic splines to ensure path smoothness, and natural boundaries are used for endpoint boundary conditions. Region mapping maps gaze coordinates to visual element regions of the advertising content. The mapping uses bounding box detection and region overlap calculation. The bounding box defines the rectangular boundary of each visual element, and the overlap is calculated as the proportion of the intersection area between the gaze point and the element region. Dwell time is calculated by tracking the cumulative time of gaze within each element region. A continuous dwell threshold is set at 100 milliseconds; dwell times below this threshold are ignored. Access order is determined by timestamp sorting to determine the order in which users browse elements; multiple visits to the same element are recorded as duplicate visits.
[0161] Visual dwell trajectory construction integrates dwell time and access sequence information to represent the user's visual cognitive process. The trajectory data structure adopts a directed graph model, where nodes represent visual elements, edges represent gaze shifts, and edge weights are shift frequencies. A trajectory simplification algorithm removes brief dwell times and random jumps, retaining the main visual flow. Simplification thresholds include a minimum dwell time of 200 milliseconds and a minimum shift frequency of 2. Trajectory clustering analysis identifies similar browsing patterns using sequence similarity metrics and hierarchical clustering algorithms. Similarity calculation is based on a dynamic time warping algorithm to handle comparisons of sequences of different lengths. Trajectory visualization uses heatmaps and flowcharts; the heatmap shows dwell intensity, and the flowchart shows shift relationships.
[0162] Time-series matching analysis associates interaction timestamp sequences with conversion behavior records. The matching algorithm is based on time windows and event sequence pattern recognition. The time window size is configurable from 1 second to 300 seconds, with a default of 30 seconds. Event sequence matching uses regular expressions and supports both exact and fuzzy matching, allowing for partial variations in the event sequence. Conversion path identification identifies the complete operation chain from a user's first encounter with the ad to the completion of the conversion. Path length is calculated by counting the number of events included, and path duration is calculated by the time difference between the first and last events. Path filtering removes paths with abnormal lengths and durations; the length threshold is set to 50 events, and the time threshold is set to 3600 seconds. Funnel analysis calculates the conversion rate and churn rate at each stage, identifying key conversion nodes and optimization opportunities.
[0163] The conversion behavior feature construction is achieved by extracting interaction nodes and conversion triggering conditions from the complete operation chain. Interaction nodes identify key links in the user operation set, and their importance is evaluated based on indicators such as operation frequency, dwell time, and conversion contribution. Node importance scoring uses a weighted summation, with operation frequency weighted at 0.3, dwell time weighted at 0.3, and conversion contribution weighted at 0.4. Conversion triggering condition analysis identifies key factors that drive conversion, including interactions with specific elements, dwell time thresholds, and access path patterns. Condition confidence calculates the conversion probability when a condition occurs, with a confidence threshold set at 0.6 to filter valid conditions. Feature vector encoding represents conversion behavior features as numerical vectors with a 128-dimensional dimension, using one-hot encoding and numerical normalization. Feature importance ranking identifies the main factors affecting conversion, implemented using information gain and feature selection algorithms.
[0164] A second aspect of the present invention provides an automatic layout optimization system for visual elements in commercial advertising design, comprising:
[0165] Module 1 is used to acquire multi-dimensional behavioral feature data and historical interaction records of target users, and extract the set of visual elements for the advertisement to be delivered; based on the multi-dimensional behavioral feature data, a user cognitive preference model is constructed, and the user cognitive preference model is mapped to a visual perception weight distribution;
[0166] Module 2 is used to perform semantic association analysis on each visual element in the set of visual elements, construct a dependency graph between visual elements, and calculate the layout tension field distribution of each visual element in the page space based on the visual perception weight distribution and the dependency graph.
[0167] Module 3 is used to generate multi-level visual guidance paths based on the layout tension field distribution and page size adaptation rules, and to adaptively allocate and arrange the presentation sequence of visual elements based on the multi-level visual guidance paths to generate personalized layout schemes.
[0168] Module 4 is used to render the personalized layout scheme into advertising content and deliver it to target users, collect user interaction feedback data after delivery, extract visual dwell trajectory and conversion behavior features from the user interaction feedback data, and use the visual dwell trajectory and conversion behavior features to jointly update the visual perception weight distribution and the dependency relationship graph.
[0169] A third aspect of the present invention provides an electronic device, comprising:
[0170] processor;
[0171] Memory used to store processor-executable instructions;
[0172] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0173] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0174] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automatically optimizing the layout of visual elements in commercial advertising design, characterized in that, include: Acquire multi-dimensional behavioral characteristic data and historical interaction records of target users, and extract the set of visual elements for the advertisements to be delivered; A user cognitive preference model is constructed based on the multi-dimensional behavioral feature data, and the user cognitive preference model is mapped to a visual perception weight distribution. Semantic association analysis is performed on each visual element in the set of visual elements to construct a dependency graph between visual elements. Based on the visual perception weight distribution and the dependency graph, the layout tension field distribution of each visual element in the page space is calculated, including: Extract the semantic tags and functional attributes of each visual element in the set of visual elements, calculate the semantic similarity between the semantic tags and the functional attributes, and obtain the semantic association strength that represents the degree of content association between visual elements. Based on the semantic association strength, a dependency graph between visual elements is constructed, and corresponding dependency weight values are assigned to the edges connecting different visual elements in the dependency graph according to the semantic association strength. The visual perception weight distribution and the dependency graph are fused and mapped. For each visual element in the set of visual elements, the attractive potential energy and repulsive potential energy of each visual element in the page space are calculated based on its weight coefficient in the visual perception weight distribution and its topological connection relationship in the dependency graph. By performing spatial field superposition calculations on the attractive potential energy and the repulsive potential energy, a layout tension field distribution is generated to guide the spatial layout of visual elements. Based on the layout tension field distribution and page size adaptation rules, a multi-layered visual guidance path is generated. Then, based on this multi-layered visual guidance path, visual elements are adaptively allocated spatially and their presentation sequence is arranged to generate a personalized layout scheme, including: Obtain the page size parameters and display characteristic parameters of the target projection device, and formulate page size adaptation rules based on the page size parameters and display characteristic parameters; Based on the layout tension field distribution, identify the extreme value regions of the tension field gradient in the layout space, and combine the layout size adaptation rules to divide the extreme value regions of the tension field gradient into multiple visual levels; For each visual level, visual guidance nodes are planned in descending order of tension field intensity, and a multi-level visual guidance path containing hierarchical jump relationships is generated by connecting the visual guidance nodes. Based on the multi-level visual guidance path, a visual hierarchy and path position are assigned to each visual element, and the visual elements are adaptively allocated in combination with the content size requirements of each visual element and the space occupancy constraints in the layout tension field distribution. Based on the hierarchical jump relationship in the multi-level visual guidance path and the user's cognitive rhythm parameters, the presentation sequence of each visual element is set, and the result of the adaptive spatial allocation is integrated with the presentation sequence to generate a personalized layout scheme that includes the spatial position information and timing control information of the visual elements. The personalized layout scheme is rendered as advertising content and delivered to target users. User interaction feedback data is collected after delivery. Visual dwell trajectory and conversion behavior features are extracted from the user interaction feedback data. The visual perception weight distribution and the dependency relationship graph are jointly updated using the visual dwell trajectory and the conversion behavior features.
2. The method according to claim 1, characterized in that, Based on the aforementioned multi-dimensional behavioral feature data, a user cognitive preference model is constructed. Mapping this user cognitive preference model to a visual perception weight distribution includes: From the multi-dimensional behavioral feature data, extract the user's browsing behavior sequence, interaction operation sequence, and content consumption sequence, and perform temporal feature encoding on the browsing behavior sequence, the interaction operation sequence, and the content consumption sequence to obtain a set of behavioral feature vectors; Based on the set of behavioral feature vectors, the cognitive response patterns of users under different content types are identified through multimodal semantic aggregation, and a user cognitive preference model including cognitive sensitivity distribution and information acceptance tendency parameters is constructed. The cognitive sensitivity distribution and information acceptance tendency parameters in the user cognitive preference model are projected onto the visual element feature space to establish a mapping relationship between the cognitive sensitivity distribution, the information acceptance tendency parameters, and visual element attributes. Based on the mapping relationship, the cognitive influence weight coefficients corresponding to different types of visual elements are calculated, and a visual perception weight distribution reflecting the user's personalized cognitive characteristics is generated according to the cognitive influence weight coefficients.
3. The method according to claim 2, characterized in that, Based on the aforementioned set of behavioral feature vectors, a user cognitive response pattern under different content types is identified through multimodal semantic aggregation. A user cognitive preference model, including parameters for cognitive sensitivity distribution and information acceptance tendency, is constructed as follows: Each behavioral feature vector in the behavioral feature vector set is labeled with a content type. Based on the content type labeling, the behavioral feature vector set is divided into multiple content type subsets, and each content type subset corresponds to a specific content presentation format. For each content type subset, the user's interaction response intensity features and attention duration features are extracted to construct the correlation matrix between content type and cognitive response. Then, the correlation matrix of different content type subsets is semantically aligned and features are aggregated through a cross-modal feature fusion mechanism to obtain a unified multimodal cognitive response feature representation. Based on the multimodal cognitive response feature representation, the cognitive sensitivity distribution curve of users to different visual stimulus intensities is calculated, and information acceptance tendency parameters are extracted according to the information acceptance behavior patterns of users under different content types. The cognitive sensitivity distribution curve and the information acceptance tendency parameter are structurally encapsulated to construct a user cognitive preference model.
4. The method according to claim 1, characterized in that, For each visual level, visual guidance nodes are planned in descending order of tension field intensity. A multi-level visual guidance path containing hierarchical jump relationships is generated by connecting these visual guidance nodes, including: Extract the layout tension field distribution data within the layout space area corresponding to each visual level, and identify multiple tension field intensity peak points within the visual level by performing intensity peak detection on the layout tension field distribution data. The peak points of the tension field intensity are sorted in descending order according to the tension field intensity value. Based on the result of the descending sort, the peak points of the tension field intensity are selected as visual guidance nodes in sequence, and each visual guidance node is assigned a hierarchical priority identifier that represents the importance of the node. Based on the priority identifier within the hierarchy, adjacent visual guidance nodes are connected sequentially according to the principle of minimum visual span, in order of priority from high to low, to generate an intra-hierarchical guidance path that represents the guidance order within the visual hierarchy. For the guidance path within the level, identify the starting visual guidance node and the ending visual guidance node of each guidance path within the level, and establish a cross-level connection between the ending visual guidance node and the starting visual guidance node. The intra-level guidance path and the cross-level connection are topologically combined to generate a multi-level visual guidance path containing complete hierarchical jump relationships.
5. The method according to claim 1, characterized in that, The personalized layout scheme is rendered as advertising content and delivered to target users. User interaction feedback data is collected after the delivery, and visual dwell time and conversion behavior features are extracted from the user interaction feedback data, including: Based on the spatial location information and timing control information of the visual elements in the personalized layout scheme, the graphics rendering engine is invoked to perform pixel-level drawing and compositing of each visual element, generating advertising display content that conforms to the personalized layout scheme, and the advertising display content is encapsulated into a delivery data package; The data package is pushed to the target user's receiving terminal for display, and user interaction feedback data, including user gaze coordinate sequence, interaction operation timestamp sequence and conversion behavior record, is collected in real time during the ad display process; Spatiotemporal correlation analysis is performed on the gaze coordinate sequence in the user interaction feedback data to identify the duration and order of the user's gaze on each visual element, and to construct a visual dwell trajectory that represents the user's visual cognition process. The interaction operation timestamp sequence is matched and analyzed with the conversion behavior record to identify the complete operation link from the user's contact with the advertising content to the completion of the conversion behavior. Combined with the visual dwell trajectory, the interaction nodes and conversion trigger conditions in the complete operation link are extracted to construct conversion behavior features that characterize the user's response behavior.
6. An automatic layout optimization system for visual elements in commercial advertising design, used to implement the method of any one of claims 1-5, characterized in that it comprises: Module 1 is used to acquire multi-dimensional behavioral feature data and historical interaction records of target users, and extract the set of visual elements for the advertisements to be delivered. A user cognitive preference model is constructed based on the multi-dimensional behavioral feature data, and the user cognitive preference model is mapped to a visual perception weight distribution. Module 2 is used to perform semantic association analysis on each visual element in the set of visual elements, construct a dependency graph between visual elements, and calculate the layout tension field distribution of each visual element in the page space based on the visual perception weight distribution and the dependency graph. Module 3 is used to generate multi-level visual guidance paths based on the layout tension field distribution and page size adaptation rules, and to adaptively allocate and arrange the presentation sequence of visual elements based on the multi-level visual guidance paths to generate personalized layout schemes. Module 4 is used to render the personalized layout scheme into advertising content and deliver it to target users, collect user interaction feedback data after delivery, extract visual dwell trajectory and conversion behavior features from the user interaction feedback data, and use the visual dwell trajectory and conversion behavior features to jointly update the visual perception weight distribution and the dependency relationship graph.
7. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Page structure optimization method and system for PowerPoint
CN120493879A
Advertisement information flow management method and platform applied to multivariate visual design
CN120707213A