Method and system for optimizing page structure of presentation
By combining NeRF and ray tracing algorithms with quantum annealing algorithms to optimize the presentation page structure, the problem of low efficiency in traditional presentation production is solved, automatic typesetting and intelligent layout of visual elements are realized, and the optimization efficiency of the page structure and the information transmission effect are improved.
Patent Information
- Application Number
- CN202510979723.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-16
AI Technical Summary
Traditional presentation production relies on manual operations, resulting in a large amount of time and energy being consumed in non-creative labor. It lacks automated typesetting tools and is particularly inefficient in high-frequency usage scenarios.
The three-dimensional neural radiation field model and ray tracing algorithm built by NeRF, combined with the quantum annealing algorithm and user narrative preferences, automatically identify and optimize the visual element layout of presentation pages, and optimize the page structure by calculating the visual focus heat map and information density energy function.
It realizes the automatic layout of presentation pages, improves the optimization efficiency of page structure, helps improve the work efficiency of personnel in related positions, and improves the efficiency and visual appeal of information transmission.
Smart Images

Figure CN120493879B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence. More specifically, the embodiments of the present application relate to a method and system for optimizing the page structure of a presentation. Background Art
[0002] At present, the production of traditional presentations (i.e. PPT) is highly dependent on manual operations, and its core pain point is the excessive consumption of professional human resources on non-creative labor.
[0003] In the past, professionals spent considerable time and energy on mechanical tasks like copying and pasting text, adjusting layout spacing, and drawing basic charts when creating presentations. This inefficiency is particularly prominent in high-frequency PPT usage scenarios like financial analysis and academic presentations. Clearly, there is a lack of automated typesetting tools to assist professionals in optimizing layouts and improving the intelligence of PPT production. Summary of the Invention
[0004] In this context, the embodiments of the present application hope to provide a method and system for optimizing the page structure of presentations, which can realize automatic layout of presentation pages, optimize page structure, improve layout efficiency of presentation pages, and assist in improving the work efficiency of personnel in related positions.
[0005] In a first aspect of the embodiments of the present application, a method for optimizing the page structure of a presentation is provided, comprising:
[0006] Get the presentation page entered by the user;
[0007] Identify visual elements in a presentation page and obtain visual structure information corresponding to the presentation page;
[0008] Analyze the logical relationships between visual elements to obtain the logical relationship information between visual elements in the presentation page;
[0009] Based on the 3D neural radiation field model constructed by NeRF, the visual structure information is mapped to the 3D implicit field of the page to obtain the 3D voxels corresponding to the visual elements. Using the ray tracing algorithm, the visual focus heat map of the presentation page is calculated, and the 3D voxels corresponding to the target visual elements are assigned to the visual sensitive area to obtain the first layout information.
[0010] A quantum annealing algorithm is used to predict the global optimal layout of the presentation page based on the positional relationship between each visual element in the visual structure information and the first layout information, combined with an information density energy function, to obtain the second layout information;
[0011] Predicting the narrative path in the presentation based on the logical relationship information and the user's narrative preferences marked based on the user's historical data to obtain a probability distribution of the page's narrative path; optimizing the second layout information based on the page's narrative path probability distribution to obtain third layout information so that the spatial arrangement of visual elements in the presentation page conforms to the narrative logic preferred by the user;
[0012] A secondary layout of the presentation page is performed based on the third layout information to optimize the page structure of the presentation page.
[0013] In a second aspect of the embodiments of the present application, a system for optimizing the page structure of a presentation is provided, comprising:
[0014] The acquisition module is used to obtain the presentation page input by the user;
[0015] A recognition module is used to identify visual elements in a presentation page and perform logical relationship analysis on the visual elements to obtain visual structure information corresponding to the presentation page;
[0016] The parsing module is used to parse the logical relationships between visual elements and obtain the logical relationship information between visual elements in the presentation page;
[0017] The first optimization module is used to map the visual structure information into the 3D implicit field of the page based on NeRF to obtain the 3D voxels corresponding to the visual elements. The ray tracing algorithm is used to calculate the visual focus heat map of the presentation page and assign the 3D voxels corresponding to the target visual elements to the visual sensitive area to obtain the first layout information.
[0018] a second optimization module, configured to use a quantum annealing algorithm to predict a global optimal layout of the presentation page based on the positional relationship between the visual elements in the visual structure information and the first layout information in combination with an information density energy function, thereby obtaining second layout information;
[0019] a third optimization module configured to predict the narrative path in the presentation based on the logical relationship information and the user's narrative preferences marked based on the user's historical data, to obtain a probability distribution of the page's narrative path; and optimize the second layout information based on the page's narrative path probability distribution to obtain third layout information, so that the spatial arrangement of visual elements in the presentation page conforms to the narrative logic preferred by the user;
[0020] An execution module is used to perform secondary layout of the presentation page based on the third layout information to optimize the page structure of the presentation page.
[0021] In a third aspect of the implementation of the present application, a terminal device is provided, comprising: at least one processor, a memory, and an input / output unit; wherein the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the page structure optimization method of a presentation described in any one of the first aspects.
[0022] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which includes instructions that, when executed on a computer, enable the computer to execute the page structure optimization method for a presentation described in any one of the first aspects.
[0023] In a fifth aspect of the embodiments of the present application, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the page structure optimization method of a presentation described in any one of the first aspects.
[0024] According to the page structure optimization method and system of the presentation of the embodiment of the present application, it is possible to obtain the presentation page input by the user. Identify the visual elements in the presentation page, and perform logical relationship analysis on the visual elements to obtain the visual structure information corresponding to the presentation page. Perform logical relationship analysis on the visual elements to obtain the logical relationship information between the visual elements in the presentation page. Then, based on the three-dimensional neural radiation field model constructed by NeRF, the visual structure information is mapped to the three-dimensional implicit field of the page to obtain the three-dimensional voxels corresponding to the visual elements; using the ray tracing algorithm, the visual focus heat map in the presentation page is calculated, and the three-dimensional voxels corresponding to the target visual elements are assigned to the visual sensitive area to obtain the first layout information; wherein, the target visual element includes one of the following: page title, page, key data, key chart, key image; the first layout information is used to indicate the optimal visual reception position of the target visual element matching the key content information in the presentation page. Then, a quantum annealing algorithm is used to predict the global optimal layout in the presentation page based on the positional relationship between each visual element in the visual structure information and the first layout information, combined with the information density energy function, to obtain the second layout information. Then, based on the logical relationship information and the user's narrative preferences marked based on the user's historical data, the narrative path in the presentation is predicted to obtain a probability distribution of the page narrative path; based on the probability distribution of the page narrative path, the second layout information is optimized to obtain third layout information, so that the spatial arrangement relationship of the visual elements in the presentation page conforms to the narrative logic of the user's preferences. Finally, based on the third layout information, a secondary layout of the presentation page is performed to optimize the page structure of the presentation page. According to the implementation mode of the present application, it is possible to realize the automatic layout of the presentation page, optimize the page structure, improve the layout efficiency of the presentation page, and assist in improving the work efficiency of personnel in related positions. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A flowchart of a method for optimizing the page structure of a presentation provided in one embodiment of the present application;
[0026] Figure 2 A schematic diagram of the structure of a presentation page structure optimization system provided in one embodiment of the present application;
[0027] Figure 3 The structural diagram of a medium in an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0028] Reference below Figure 1 , Figure 1This is a flow chart of a method for optimizing the page structure of a presentation provided in an embodiment of the present application. It should be noted that the implementation of the present application can be applied to any applicable presentation production scenario.
[0029] Figure 1 The process of the method for optimizing the page structure of a presentation provided in one embodiment of the present application includes:
[0030] Step S101: Obtain a presentation page input by a user.
[0031] In the embodiments of this application, a presentation is a dynamic document that uses visual elements such as text, charts, and animations to organize content. Its primary purpose is to convey and display information. Presentations can play an important role in different scenarios. Through diverse formats and flexible designs, they can make information presentation clearer, more intuitive, and more engaging.
[0032] Presentations are diverse in file formats. Common formats include Microsoft's ".ppt" and ".pptx," WPS's ".dps," and Apple's ".key." ".pptx," the default format since PowerPoint 2007, supports multimedia embedding and cloud collaboration. ".key" files are designed specifically for Keynote, offering enhanced animation rendering on Apple devices. These different formats are compatible with diverse software ecosystems.
[0033] Specifically, the content structure exhibits a modular nature. A typical presentation consists of multiple modules: the cover page should highlight the theme, subtitle, and author information, with a simple and eye-catching design; the table of contents is used to outline the content framework and guide audience expectations; transition pages can connect different chapters and enhance logical coherence; the content page uses images, charts, and case studies to develop the core argument; and the back cover is used to summarize key points and provide acknowledgments, strengthening the closing impression. Some professional designs also include appendices to supplement data or interactive pages that link to external resources via QR codes.
[0034] Dynamic interactivity is also a key feature of presentations. It allows for step-by-step presentations through animations like fade-in and fly-in, as well as transitions like push-pull and dissolve. For example, complex processes can be dynamically demonstrated using animated paths, while hyperlinks and action buttons support non-linear transitions between pages, significantly increasing audience engagement.
[0035] Presentations are widely used in business and the workplace. During work reports, performance metrics can be displayed using data charts and timelines, while project progress can be presented. During product launches, 3D models and videos can be used to demonstrate product features, while comparison tables can be used to highlight competitive advantages. During meeting presentations, using the presenter view to display notes simultaneously ensures smooth presentation.
[0036] Education and academia are also inseparable from presentations. In classroom teaching, abstract concepts, such as the principles of physics and mechanics, can be explained through step-by-step animations. Academic presentations can embed dynamic graphs of experimental data, and references can be linked via QR codes. In online courses, interactive multiple-choice questions designed with triggers can enhance the interactivity of remote learning.
[0037] In terms of technical implementation, mainstream authoring tools like PowerPoint support master layouts, enabling a unified visual style. Cloud-based collaboration features like Teams integration allow for real-time editing by multiple people. Advanced features also exist, such as data visualization that dynamically links to Excel spreadsheets for real-time chart updates. AI-assisted tools can automatically generate content outlines or optimize color schemes.
[0038] In short, thanks to these characteristics, presentations have become a core tool for cross-disciplinary information transmission. They not only carry content but also enhance communication efficiency and cognitive depth through visual language, making information transmission more efficient and in-depth.
[0039] It should be noted that a presentation page is one or more pages in a presentation. For ease of description, the following embodiments are described using one page as an example.
[0040] Step S102 : identifying visual elements in the presentation page and performing logical relationship analysis on the visual elements to obtain visual structure information corresponding to the presentation page.
[0041] Specifically, step S102 aims to parse the visual structure of the presentation page, and provide a basis for subsequent analysis and optimization by identifying visual elements and sorting out their logical relationships.
[0042] The principle of step S102 is based on the cross-application of computer vision and natural language processing. First, the presentation page can be regarded as a two-dimensional visual scene, in which the visual elements such as text, pictures, charts, shapes, etc. carry specific information. The computer uses image recognition technology to separate the various elements on the page from the background. Subsequently, the logical relationship between the elements is analyzed using the spatial position, size, color, font and other features of the elements, as well as the semantic information of the text content. For example, adjacent elements with similar styles may belong to the same theme, and the title text has a hierarchical relationship with the content text below. In this way, the visual structure information of the page is gradually constructed.
[0043] In one optional example, the image recognition component uses a convolutional neural network (CNN) to identify and classify visual elements. This CNN can learn the characteristic patterns of elements and accurately distinguish between different types of elements, such as text boxes, images, and charts. For text information, semantic analysis models in natural language processing, such as the Transformer architecture, can be used to understand text content and analyze the semantic associations between elements. In addition, graph neural networks (GNNs) are also often used to model the relationships between elements. Each visual element is considered a node in the graph, and the relationships between elements are considered edges. The complex logical connections between elements are mined through graph calculations, thereby better analyzing the visual structure.
[0044] In another optional example, during the logical relationship parsing stage, the system constructs a multi-granularity semantic graph, maps the detected visual elements into graph nodes, and encodes the spatial layout rules (such as horizontal alignment, vertical stacking) and semantic associations (such as the reference relationship between data charts and corresponding explanatory text) between elements through edge relationships. This process combines the graph convolutional network (GCN) and the bidirectional long short-term memory network (Bi-LSTM). The former is used to capture the local topological structure between elements, and the latter is good at processing serialized semantic information. For example, when the arrow direction in the flowchart is identified, the model will establish the temporal logic between the steps through the relational reasoning module and verify the logical consistency in combination with the context text. For complex layouts, the system adopts a multi-stage parsing strategy: first, the semantic blocks are divided through the region proposal network, and then the element-level relationship analysis is performed within the block. Finally, a global visual structure description including the hierarchical structure, visual flow path and semantic network is constructed.
[0045] From a technical perspective, it can accurately disassemble presentation pages. It can not only identify various visual elements on the page, but also deeply analyze logical information such as the primary and secondary relationships, hierarchical relationships, and association relationships between elements. In this way, the visual structure information ultimately obtained can intuitively present the layout logic and information organization of the page. This helps optimize page design in subsequent production and improve the readability and visual appeal of the presentation. It can also be used for automatic review to detect problems such as unreasonable element relationships or unclear information communication on the page, thereby improving the overall quality of the presentation.
[0046] Step S103 , performing logical relationship analysis on the visual elements to obtain logical relationship information between the visual elements in the presentation page.
[0047] Specifically, in step S103, the core goal of analyzing the logical relationships of visual elements is to reveal the inherent correlation patterns between different elements in the presentation page through multi-dimensional analysis technology.
[0048] First, semantic classification is performed based on the type of visual element (such as text paragraphs, charts, images, shapes, etc.), and their attribute features are extracted, such as keywords in text content, data types in charts, and color patterns in images. Then, spatial layout analysis techniques, combined with proximity principles (such as distance between elements and degree of overlap) and alignment methods (horizontal / vertical alignment, grid constraints), determine whether the physical proximity of elements suggests a logical connection. For example, closely spaced bar charts of the same color scheme may indicate a comparative data relationship, while sub-items arranged radially around the title may form a total score structure.
[0049] At the semantic level, natural language processing techniques are used to parse the text, identifying the hierarchical relationship between titles and body text, the parallel nature of list items, and causal connectives (such as "therefore" and "however"). These linguistic features are then mapped to corresponding visual elements. For chart elements, the algorithm extracts metadata such as axis labels and legends. Combined with predefined chart logic templates (e.g., flowcharts for process relationships, timelines for temporal relationships), a logical mapping is established between the chart and the text. Image recognition technology is also used to analyze the symbolic meaning of illustrations (e.g., an upward arrow indicates a growth trend) and verify their consistency with the surrounding text.
[0050] To enhance the accuracy of the analysis, a context-aware model is constructed, comprehensively considering the overall structure of the page (such as chapter divisions and the presence of transition elements) and cross-page logical clues (such as the recurrence of chapter titles and the directionality of navigation icons). For example, if the keyword "market share" appears in a chart on a certain page and the text on the previous and next pages, it is determined that the chart has a strong correlation with the theme of the full text. In addition, the algorithm will introduce the layout preferences in the user's historical data as a correction factor. For example, for users who are accustomed to using contrasting arguments, the system will prioritize identifying pairs of elements with strong color contrast on the page and assign a higher confidence level to their contrast relationship. Ultimately, the results of the logical relationship analysis are output in the form of a structured graph, which includes direct relationships between elements (such as "parallel" and "cause and effect"), indirect association paths (such as logical chains transmitted through intermediate elements), and hierarchical depth (such as the nested relationship between the main title and subtitles), providing executable logical constraints for subsequent layout optimization.
[0051] Step S104: Based on the three-dimensional neural radiation field model constructed by NeRF, the visual structure information is mapped into the three-dimensional implicit field of the page to obtain three-dimensional voxels corresponding to the visual elements.
[0052] In step S105 , a ray tracing algorithm is used to calculate a heat map of visual focus in the presentation page, and three-dimensional voxels corresponding to target visual elements are allocated to visually sensitive areas to obtain first layout information.
[0053] In the embodiment of the present application, the target visual element includes one of the following: page title, page, key data, key chart, key image. Further, the first layout information is used to indicate the optimal visual reception position of the target visual element matching the key content information in the presentation page.
[0054] For example, in the cover design of a listed company's annual report, the original page title was centered in a regular font size (24pt). However, due to the dark blue gradient background, the contrast between the text and the background was insufficient, requiring viewers to focus for a long time to recognize the core information. Through this technical solution, the system first identifies the title text as a key visual element, uses NeRF to build a three-dimensional spatial model, maps the title area into a high-density voxel cluster, and applies a gradient gold light effect to enhance the three-dimensional effect. When the ray tracing algorithm calculates the visually sensitive area, it finds that the center area of the page has the highest thermal value due to the natural focusing characteristics of the human eye. Therefore, the title three-dimensional voxels are dynamically offset to this area, and the font size is increased to 36pt and a 0.5px drop shadow effect is added. After optimization, the title recognition speed is improved, and the audience's first fixation time is shortened, which is in line with the "fovea priority" principle of visual cognition.
[0055] For example, a key data chart (a bar chart) in a medical research PPT was originally placed in the lower right corner of the page, surrounded by a large amount of supporting text, which could easily lead viewers to overlook the core conclusions. The system identified the chart as a key data element through visual structure analysis and, using 3D voxel modeling, elevated it to 1.2 times its height on the z-axis, creating a suspended, three-dimensional chart. A ray tracing algorithm detected that the chart's peak data point (e.g., the 85% cure rate) had the highest color contrast (red bar against white background) and automatically assigned voxel density to the high-sensitivity region, creating a peak light intensity at that data point in 3D space. Simultaneously, through spatial layout optimization, the conclusion text ("Significant Efficacy") was arranged along the visual flow path (extending from the chart to the upper left), forming a logical "data → conclusion" flow. Actual testing showed that viewers' recall of the core conclusions improved, confirming the effectiveness of the visual flow path design.
[0056] The core principle of step S104 is to use Neural Radiance Fields (NeRF) to map 2D visual structural information into a continuous voxel representation in a 3D implicit space. The system first converts identified visual elements (such as title boxes and chart coordinates) into geometric descriptions in a 3D coordinate system, including position, size, and hierarchical relationships. This structural information serves as input to NeRF's neural network to construct a radiance field for the scene: each voxel's 3D coordinates and viewing direction are mapped to color (RGB) and volume density (σ). The color encodes the element's semantic properties (such as the font color of the title text), while the volume density reflects the element's opacity and visual weight in space. By applying multi-view consistency constraints, NeRF can fuse discrete visual elements into a continuous 3D field, supporting the simulation of visual focus from any angle. For example, the title area is represented in 3D space as a high-density voxel cluster, with its color gradient and font gradient effects learned via an MLP network.
[0057] For example, an improved NeRF model is adopted, hierarchical feature encoders (such as MLP and hash grid) are introduced to accelerate training, and dynamic resolution adjustment is supported to adapt to different page complexities.
[0058] For example, the element bounding boxes and semantic labels output by the visual structure parsing module are used as conditional inputs, and the voxel generation process is constrained by the conditional batch normalization layer to ensure the spatial correlation between the title and the chart.
[0059] In step S105, based on the generated three-dimensional voxel field, a differentiable ray tracing algorithm is used to simulate the human visual perception mechanism and calculate a heat map of the visually sensitive areas of the page. Light is emitted from a virtual viewpoint and attenuates according to the volume density when passing through the voxel field, and the color contribution is accumulated. Through the importance sampling strategy, the algorithm prioritizes tracking high-contrast areas (such as the boundary between the title and the chart) and semantic key points (such as the peak points in the data chart), generating a heat map to reflect the distribution of visual attention. Subsequently, the three-dimensional voxels of the target visual element are dynamically adjusted according to the gradient of the heat map: the voxel density of the highly sensitive area is enhanced to increase the visual weight, and the voxels of the low-sensitivity area are optimized to the edge position through gradient descent, ultimately forming a typesetting scheme that conforms to human visual cognition.
[0060] Exemplarily, a path tracing algorithm is used in combination with a spatial acceleration structure (such as a BVH tree) to reduce invalid intersection calculations between rays and voxels.
[0061] For example, a pre-trained visual attention model (such as a Transformer-based saliency detection network) is integrated to perform secondary weighting on the heat map to strengthen the priority of titles and key data.
[0062] In this way, NeRF's continuous field representation breaks through the discrete limitations of traditional raster typesetting and supports the non-rigid layout of elements in three-dimensional space. For example, charts can be extended along the Z-axis to form a three-dimensional visual structure, and titles can achieve a gradual floating effect through density gradients, significantly enhancing the page's sense of hierarchy. Ray tracing algorithms simulate the human eye's retinal imaging and visual cortical processing mechanisms, and heat maps accurately locate visually sensitive areas (such as the fovea). As a result, title recognition speed is improved and the duration of gaze on chart data points is reduced, in line with cognitive load theory. By jointly training NeRF and ray tracing networks, end-to-end optimization of layout parameters is achieved. The joint optimization of three-dimensional voxels and two-dimensional visual elements solves the problem of cross-platform rendering distortion. For example, PDF-exported pages maintain smooth element spacing and color transitions when scaled, and edge jaggedness is reduced.
[0063] As an optional embodiment, in step S105, a ray tracing algorithm is used to calculate a visual focus heat map in the presentation page, and three-dimensional voxels corresponding to target visual elements are assigned to visually sensitive areas to obtain first layout information, including:
[0064] A beam of light is emitted from the virtual viewpoint of the presentation page, penetrating each voxel in the three-dimensional implicit field. The spatial distribution density of the visual elements is determined by solving the intersection equation of the light and the scene geometry. A high refractive index property is given to the target visual element to enhance the deflection intensity of the light on the surface of the target visual element in the three-dimensional implicit field, simulating the focusing effect of the human eye lens. Differentiated material parameters in the three-dimensional implicit field are defined according to the type of visual element. Among them, the title adopts a high specular reflection coefficient, and the chart sets the diffuse reflection coefficient to simulate the information carrying characteristics. An ambient light occlusion algorithm is introduced to calculate the visual occlusion relationship between the visual elements in the three-dimensional implicit field, enhancing the stereoscopic perception of the focus area. Through Monte Carlo light A line sampling method is used to count the number of light hits per unit area in a three-dimensional implicit field to generate a visual sensitivity distribution map; in the visual sensitivity distribution map, Gaussian blur processing is performed on areas where the hit rate exceeds the hit rate threshold to form a smoothly transitioned focus heat map, in which the thermal value of the core focus area reaches a peak; based on the direction of change of the thermal gradient of the focus heat map, the three-dimensional voxels of the target visual element are offset along the line of sight by a preset multiple of the field of view angle radius to ensure that the high-priority target visual elements are located in the core focus area with optimal retinal imaging; non-maximum suppression is performed on the visual elements corresponding to overlapping voxels to retain the layout scheme of the visual elements with the highest information density to obtain the first layout information.
[0065] Specifically, the core goal of step S105 is to accurately locate key visual elements (such as titles and charts) in the presentation to the optimal area for retinal imaging by simulating the visual perception mechanism of the human eye.
[0066] Specifically, based on ray tracing technology, a beam of light is emitted from a virtual viewpoint through voxels in the 3D implicit field. By solving the intersection equation between the light and the scene geometry, the spatial distribution density of visual elements is quantified. Furthermore, a high refractive index property is assigned to target elements to simulate the focusing effect of the human eye lens, causing light to be deflected at the element's surface, enhancing the visual weight of the focal area. Furthermore, differentiated material parameters are defined based on element type: titles use a high specular reflectance coefficient to simulate visual attraction, while charts use a diffuse reflectance coefficient to reflect their information-carrying properties. An ambient occlusion algorithm further calculates the occlusion relationship between elements, enhancing the three-dimensional perception of focus through shadows. A Monte Carlo ray sampling method counts the number of ray hits per unit area to generate a visual sensitivity distribution map. A Gaussian blur is then used to smooth the heat map, creating a natural gradient between the core focal area and the transition zone. Finally, the element voxel positions are adjusted based on the direction of the thermal gradient, and non-maximum suppression is used to optimize layout conflicts, ensuring that high-priority elements occupy the optimal visual reception position.
[0067] The above steps integrate multiple computer vision and graphics algorithms. Specifically, the core ray tracing algorithm uses a path tracing framework combined with spatial acceleration structures (such as the BVH tree) to reduce invalid ray-voxel intersection calculations and improve ray sampling efficiency. The material optical model, based on physical-based rendering (PBR) theory, defines bidirectional reflectance distribution functions (BRDFs) for different elements. High specular reflections in titles are simulated using the Fresnel equations, while diffuse reflections in charts use the Lambertian model. Ambient occlusion calculations use the screen-space ambient occlusion (SSAO) algorithm, combined with depth buffer data, to quantify occlusion relationships between elements and enhance the sense of depth in focal areas. Probabilistic sampling optimization uses a Monte Carlo method with an importance sampling strategy to prioritize sampling density in visually sensitive areas (such as high-contrast edges), improving heatmap accuracy. Gaussian filtering is used to smooth heatmap noise, and the non-maximum suppression (NMS) algorithm eliminates redundant layout solutions by detecting local extrema.
[0068] In the above steps, illustratively, a density function is constructed based on the number of ray hits: .in, Indicates the page location Distribution of visual sensitivity at .
[0069] is the attenuation coefficient, which is used to simulate the sensitivity attenuation characteristics of the human retina to visual stimuli. The larger the value, the faster the visual sensitivity decreases with distance, which is consistent with the high-resolution characteristics of the fovea of the human eye. For example, the fovea only covers about 2° of visual field, and the sensitivity of the peripheral area decreases exponentially. In PPT layout, It is usually dynamically adjusted based on the display device DPI (such as 96dpi) and the visual acuity of the human eye (about 1 arc minute) to ensure that high-density areas (such as titles) are perceived first. represents an exponential decay term.
[0070] It is a spatial coordinate point in the three-dimensional implicit field (such as the position of a title or chart), representing the position of the geometry that the light may interact with during its propagation. In density field sampling, the volume integral is represented by Sampling is performed (e.g. sampling every 1px), and each The density value at and through Cumulative contribution.
[0071] This is a density field correction term. For example, by amplifying density differences through the refractive index, the density ratio of the title to the background can be increased from 1:1 to 1.5:1, enhancing visual hierarchy.
[0072] Furthermore, the refractive index weighting mechanism can be expressed as follows: ; The refraction index for targeted visual elements (such as titles) can be set to 1.5-1.8. is the density field correction term before weighting. For page location The offset of .
[0073] The Monte Carlo integration domain covers the entire page viewport and can be dynamically adjusted to: a local focus mode that integrates only the current viewport area (such as the visible window when reading a PowerPoint presentation); a global analysis mode that covers all page cells for cross-page layout optimization. Adaptive meshing is used, with a 100μm resolution for high-density areas (such as charts) and a 1mm resolution for low-density areas (such as white space), reducing computational effort by up to 70%.
[0074] Based on the above formula and the meaning of the above data, a smoothed focus heat map is generated through Gaussian blur. : ;in, is the Gaussian kernel, standard deviation =2.
[0075] For example, in a financial data PPT, the title cell ( '=1.8) and the chart area ( =1.2) =0.15 attenuation, the title's contribution outside 300px still remains 2.3 times that of the chart area. =2 Gaussian blur creates a gradient transition zone (approximately 12px wide) around the title, highlighting the core information while avoiding visual abruptness. This parameter combination increases the probability of viewers first fixating on the title by 68% (eye tracking experiment data).
[0076] Specifically, step S105 simulates the human eye's focusing mechanism for objects at varying distances by dynamically adjusting the refractive index properties of the target visual element. While traditional ray tracing typically uses a fixed refractive index, this embodiment assigns differentiated refractive index values based on element type (e.g., title, chart). The title area has a significantly higher refractive index than the background, causing light to be more strongly deflected on its surface, creating a "visual lens"-like effect.
[0077] For example, when the virtual viewpoint focuses on a title, the Fresnel equations are used to calculate the ratio of reflection and refraction at different angles of incidence. The light path is then dynamically adjusted so that the light density in the title area forms a Gaussian distribution peak at the retinal imaging plane. This process not only enhances the visual salience of the title but also simulates the depth of field changes that occur when the human eye naturally focuses by deflecting light. This improves title recognition accuracy and reduces visual fatigue (based on eye tracker test data).
[0078] A layered material model was defined to address the functional differences between different visual elements. Titles use a high specular reflectance coefficient (Specular Reflectance >0.8) and low roughness (Roughness <0.1). The bidirectional reflectance distribution function (BRDF) simulates the directional reflectance characteristics of a smooth surface, allowing light to form bright specular highlights in the title area, enhancing visual appeal. Charts use a diffuse reflectance coefficient (Diffuse Reflectance 0.6-0.8) combined with an anisotropic microsurface model to ensure uniform lighting in the information-bearing area while enhancing information readability through micro-reflections in surface details (such as bar chart ticks). At the implementation level, a pre-trained material classification network automatically identifies element types and dynamically loads the corresponding BRDF parameter set to ensure accurate mapping of material properties to semantic labels.
[0079] The SSAO (Surrounding Light Occlusion) algorithm is used to quantify the spatial occlusion relationship between visual elements, thereby enhancing the three-dimensional hierarchy of the focus area. Traditional SSAO only calculates the occlusion degree of local pixels, but the embodiment of the present application extends it to the three-dimensional voxel space: through the joint analysis of the screen space depth buffer and the three-dimensional implicit field, the system can accurately calculate the occlusion volume between elements. For example, when the title voxel is partially occluded by a chart, the algorithm will automatically generate a shadow cone and calculate the shadow intensity distribution through ray tracing integrals. This process not only enhances the visual depth of the page, but also guides the direction of visual flow through occlusion relationships. Combined with the SSAO heat map, the audience's recognition of the association between the chart and the conclusion can be improved.
[0080] Monte Carlo ray sampling implements a dynamic weight allocation mechanism in the embodiment of the present application. Traditional methods use uniform sampling, while the embodiment of the present application generates an initial sampling density map through a pre-trained saliency detection network (such as the Transformer-based SALNet), assigning higher sampling weights to high-contrast boundaries and color mutation areas (such as red data points). During the ray tracing process, the system dynamically adjusts the sampling step size based on the cumulative number of hits of the current voxel: adaptive subdivision sampling is used for high-frequency change areas (such as the edges of title text), and the sampling density is reduced for low-frequency areas (such as solid color backgrounds). This hybrid sampling strategy shortens rendering time while ensuring accuracy, and at the same time reduces the entropy value of the heat map (a measure of information uncertainty), which is significantly better than the traditional fixed-step sampling scheme.
[0081] To eliminate Monte Carlo sampling noise and create a focus transition consistent with human visual perception, the system employs a multi-stage post-processing pipeline. First, a guided filter is used to smooth the heatmap using edges, preserving details in high-gradient areas. A spatially aware Gaussian blur is then applied, dynamically adjusting the blur kernel size based on the type of visual element (kernel size <3 pixels for titles, 5-7 pixels for charts), creating a naturally transitioning visual sensitivity gradient. Finally, contrast-limited adaptive histogram equalization (CLAHE) is used to enhance global contrast, ensuring focus legibility in low-light environments.
[0082] To address the overlap of multiple visual elements, this embodiment of the present application proposes a non-maximum suppression (NMS) optimization algorithm based on local extrema detection. While traditional NMS uses a fixed threshold to filter out regions of low response, this embodiment introduces a dynamic threshold mechanism: First, the thermal gradient direction of each candidate voxel is calculated. A search radius (typically 1.5 times the angular radius of the field of view) is then extended along the gradient direction to detect voxels with higher response values within this region. If so, the current voxel is suppressed; otherwise, the layout remains valid. Furthermore, an energy optimization function is introduced, comprehensively considering element importance weights (e.g., a weight coefficient of 0.9 for titles and 0.6 for auxiliary charts), a spatial distance attenuation factor (Gaussian kernel σ = 0.3), and visual flow continuity constraints, to determine the optimal layout using gradient descent. In practical applications, this algorithm can reduce the overlap of key elements while maintaining the coherence of the visual path (a 53% improvement, as verified by reading efficiency tests).
[0083] By synergizing the refractive index property with Monte Carlo sampling, the system identifies the "foveal region" (approximately 2.5 meters from the viewpoint, with a 35° viewing angle) where the human eye naturally focuses. This allows the peak thermal values of titles and key charts to fall within this region, shortening initial fixation times and improving information capture efficiency. The combination of an ambient occlusion algorithm and the high refractive index property enables three-dimensional voxels to form a depth gradient within the visually sensitive area. For example, title voxels create a "floating" effect through refraction and deflection, while chart data points appear layered due to occlusion. User surveys show a stereoscopic perception score (MOS) of 4.6 / 5, surpassing the 3.2 of traditional 2D layout solutions. Voxel positions are adjusted based on the direction of the thermal gradient, resolving multi-element conflicts. In complex pages containing more than 10 charts, a non-maximum suppression algorithm reduces overlap between key elements while maintaining visual flow. The layout solution supports dynamic resolution adaptation. When the page is scaled to a mobile device, the voxel density automatically adjusts to maintain the visual weight distribution. Tests show that the visual comfort of the mobile page (assessed by the NASA-TLX scale) is improved compared to a fixed layout. Through physically interpretable optical simulation and probabilistic optimization, these steps achieve the transition from "empirical typography" to "cognitive-driven layout," providing a new methodological framework for intelligent presentation design.
[0084] In step S106, a quantum annealing algorithm is used to predict the global optimal layout of the presentation page based on the positional relationship between the visual elements in the visual structure information and the first layout information, combined with the information density energy function, to obtain the second layout information.
[0085] As an optional embodiment, in step S106, a quantum annealing algorithm is used to predict the global optimal layout of the presentation page based on the positional relationship between each visual element in the visual structure information and the first layout information in combination with the information density energy function to obtain the second layout information, including:
[0086] Each visual element is mapped into a quantum bit system, and the positional relationship between each visual element is encoded as a spin variable of the Ising model based on the element correlation degree in combination with the first layout information; by analyzing the spatial correlation of visual elements in adjacent distances and visual levels, a coupling strength matrix between visual elements is dynamically generated and loaded into a quantum annealing machine; the quantum annealing evolution process is directly driven by the Ising model and the quantum annealing machine, and the initial Hamiltonian used to represent the random distribution of visual elements is converted into a final Hamiltonian that complies with the encoding layout constraints through adiabatic evolution; wherein the encoding layout constraints are dynamically configured based on the page optimization requirements of the presentation page; during the annealing process, the dual objective functions of information density and aesthetic score are integrated, and a variational quantum eigensolver is used to generate a Pareto optimal solution set of the information density energy function under the Hamiltonian evolution process; the quantum state information in the Pareto optimal solution set that complies with the encoding layout constraints is decoded into the position coordinate parameters of the visual elements; in combination with the device adaptation rules, a hierarchical layout configuration of the visual elements is generated based on the position coordinate parameters, and the hierarchical layout configuration is used as the second layout information.
[0087] It is understandable that step S106 aims to achieve the global optimal layout prediction of the presentation page with the help of the quantum annealing algorithm. Its principle integrates the characteristics of quantum mechanics and page layout requirements, and achieves efficient optimization through a specific algorithm model.
[0088] Specifically, in step S106, based on the core characteristics of the quantum annealing algorithm, the process of annealing a quantum system from high temperature to low temperature is simulated to find the lowest energy state, thereby obtaining the optimal solution. In the optimization of the presentation page layout, the visual elements are mapped to a quantum bit system, and each element corresponds to a spin variable. This mapping transforms the page layout problem into an energy optimization problem in the quantum system. By analyzing the spatial correlation between elements to generate a coupling strength matrix, the logical relationship between the elements can be quantified. The Ising model is used to describe the interaction between quantum bits, and the quantum annealer is combined to drive the system evolution. From the initial random distribution state, through adiabatic evolution, the final state that meets the layout constraints is gradually reached. In the above process, the dual objective functions of information density and aesthetic scoring are integrated, and by continuously adjusting the system energy, the quantum system is allowed to find the optimal layout solution, that is, the lowest energy state, during the annealing process.
[0089] The embodiments of this application mainly involve the Ising model and the variational quantum eigensolver (VQE). The Ising model, as a classic statistical mechanics model, is used to describe the interaction between quantum bit spins. By encoding the positional relationship and correlation of visual elements as spin variables and coupling strengths, a quantum model of the page layout problem is constructed. The variational quantum eigensolver plays a key role in the quantum annealing process. It continuously adjusts parameters to find the minimum value of the energy function under the dual objective function (such as information density and aesthetic score), generating a Pareto optimal solution set to balance the relationship between different objectives. At the same time, the quantum annealer, as the hardware carrier for implementing the quantum annealing process, uses its quantum properties to accelerate the search of the solution space.
[0090] Specifically, the core objective of step S106 is to leverage the physical properties of quantum annealing algorithms to transform the positional relationships and aesthetic constraints of visual elements into an energy optimization problem for a quantum system, thereby achieving a globally optimal layout. The principle behind step S106 is based on a combination of quantum tunneling and classical optimization algorithms. It is achieved through the following steps: First, visual elements (such as titles and charts) are mapped into a quantum bit system. Each element corresponds to a spin variable, whose value (+1 or -1) represents the element's relative position on the page. By analyzing the spatial correlations between elements in the visual structure (such as proximity and hierarchical relationships), a coupling strength matrix is dynamically generated, quantifying the strength of interactions between elements. For example, a title and its corresponding chart are assigned a high coupling strength due to their logical connection, while unrelated elements have lower coupling strengths. This coupling matrix is encoded into the Ising model, forming the interaction Hamiltonian between the quantum bits, where the permutation of the spin variables directly corresponds to the page layout scheme.
[0091] A quantum annealer drives a quantum system through an adiabatic evolution process, driving it from a high temperature (random distribution) to a low temperature (low energy state). The initial Hamiltonian describes the high-energy state with randomly distributed elements, while the final Hamiltonian encodes layout constraints (such as element non-overlapping and alignment rules). During the evolution process, quantum tunneling allows the system to traverse energy barriers and explore low-energy regions in the solution space. For example, when a title and a chart form a strong correlation due to high coupling strength, the annealing process prioritizes placing them in adjacent regions, while dynamically adjusting the annealing rate to avoid falling into local optima.
[0092] During the annealing process, the dual objective functions of information density and aesthetics are integrated. The information density function measures the compactness and logical coherence of page element distribution, while the aesthetics score is based on design principles such as visual balance and color contrast. The variational quantum eigensolver (VQE) generates trial wave functions using parameterized quantum circuits. Combined with a classical optimizer, the solver adjusts parameters to find the Pareto-optimal solution set under the dual objective functions. For example, the VQE may generate multiple layout solutions, some with high information density but low aesthetics scores, while others have the opposite. Ultimately, the solution that satisfies device adaptation rules (such as mobile display constraints) is selected.
[0093] The quantum states output by the quantum annealer are decoded and converted into positional coordinate parameters for visual elements. For example, a spin variable of +1 might correspond to a right-aligned title, while -1 corresponds to a left-aligned title. These coordinate parameters, combined with device adaptation rules (such as screen resolution and reading direction), generate a hierarchical layout configuration: key graphics might be assigned to visually sensitive areas (at the golden ratio from the viewpoint), while auxiliary text is arranged along the visual flow path. The final secondary layout information contains element coordinates, hierarchical relationships, and dynamically adjusted parameters, ensuring consistent rendering across platforms.
[0094] By combining the Ising model with VQE, the system achieves a complementary advantage: the global search capabilities of quantum annealing and the fine-tuned control of classical algorithms. At a scale of 200 qubits, layout optimization efficiency is improved compared to traditional genetic algorithms, and the coverage of the Pareto front solution set is enhanced. The coupling strength matrix is dynamically generated based on visual hierarchical relationships. For example, the coupling strength between a main title and subtitles is higher than between elements at the same level. This adaptive mechanism accelerates layout convergence and reduces logical errors for complex pages (containing over 50 elements). By incorporating a device adaptation rule base, the system automatically adjusts layout parameters to suit different display devices. For example, mobile devices prioritize compressing chart width and increasing line spacing, while retaining high-precision 3D effects on desktops. Testing shows that this solution achieves a cross-device typographic consistency score (MOSS) of 4.7 / 5, an improvement over fixed template solutions. Zero-noise extrapolation (ZNE) and matrix product state pre-training effectively mitigate the decoherence effects of the quantum annealer. Tests on quantum processors demonstrate improved fidelity of the layout solution and reduced computational time.
[0095] Through the deep integration of quantum physics properties and design aesthetics, it provides a breakthrough solution for intelligent typesetting. Its global optimization capabilities and adaptive characteristics mark the paradigm shift of page layout from experience-driven to quantum intelligence-driven.
[0096] In this way, step S106 realizes efficient and accurate optimization of the presentation page layout. With the help of the parallelism and quantum superposition characteristics of the quantum annealing algorithm, a large number of candidate layout schemes can be evaluated in a very short time, which greatly improves the computational efficiency compared to the classical algorithm. By integrating the dual objective functions of information density and aesthetic scoring, not only the efficient communication of page content is guaranteed, but also the visual aesthetics is taken into account. The Pareto optimal solution set finally generated provides users with a variety of high-quality layout options. The quantum state information is decoded into the position coordinate parameters of the visual elements, and a hierarchical typesetting configuration is generated in combination with the device adaptation rules, so that the optimized layout can be well adapted to different display devices, thereby improving the visual expression and practicality of the presentation in various scenarios, and truly realizing the prediction and application of the global optimal layout.
[0097] Step S107 : predicting the narrative path in the presentation based on the logical relationship information and the user narrative preference marked based on the user historical data, so as to obtain a probability distribution of the page narrative path.
[0098] Step S108: Based on the probability distribution of the page narrative path, the second layout information is optimized to obtain third layout information, so that the spatial arrangement relationship of the visual elements in the presentation page conforms to the narrative logic preferred by the user.
[0099] Exemplarily, in step S107, a multi-dimensional narrative path prediction model is constructed by integrating logical relationship information with historical user behavior data. First, based on the logical associations between visual elements on the page (such as the hierarchical relationship between the title and the text, and the correspondence between charts and data), the algorithm builds a structured knowledge graph, converting the logical relationships between elements (such as causality, parallelism, and progression) into a network model of nodes and edges. For example, if a page contains the title "Problem Analysis" and compares it with a subsequent bar chart, the system will extract the title keywords through natural language processing and match them with the chart labels, forming a logical link from "question → data verification."
[0100] At the same time, the algorithm leverages historical user data on layout preferences and interaction behaviors, such as frequent use of timeline layouts and duration of stay in a chart area, to generate a user preference vector. This vector is then fed into the prediction model as a weight parameter, dynamically adjusting the priority of logical relationships. For example, if a user's historical preference for data-driven narratives is evident, the algorithm will increase the transition probability of key chart nodes, giving them a higher weight in path prediction.
[0101] The prediction process uses a dynamic probabilistic model (such as a hidden Markov model or graph neural network) to simulate the user's visual movement and logical jumps on the page. The model generates a probability distribution of all possible narrative paths by calculating the transition probability matrix between nodes. For example, for a page containing "Market Status," "Competitive Analysis," and "Strategic Recommendations," the model may output a high-probability linear path of "Current Situation → Analysis → Recommendations" and a low-probability branching path of "Current Situation → Recommendations → Analysis." To optimize prediction accuracy, the system introduces a reinforcement learning mechanism that adjusts model parameters based on real-time user feedback on recommended layouts (such as clicks and dwell time), gradually improving the matching degree of path recommendations.
[0102] For example, step S108 converts the predicted probability distribution of narrative paths into spatial layout constraints to guide the optimization of the second layout information. The system first maps nodes of high-probability paths (such as core conclusions and key data) to visually sensitive areas of the page, such as the "golden triangle" in the upper right corner or the information focus band in the lower middle section, ensuring that the user's gaze naturally flows to the core content. For auxiliary information of low-probability paths (such as supplementary notes and references), the algorithm compresses it to the edge or uses foldable design to reduce visual clutter.
[0103] During the optimization process, the global search capabilities of the quantum annealing algorithm are combined to dynamically adjust the position and hierarchy of elements. For example, if the predicted path requires visual contrast between the "problem statement" and the "solution," the algorithm automatically assigns them to symmetrical areas on the left and right of the page, reinforcing the logical connection through color contrast or arrow guidance. At the same time, the system introduces a dynamic weighting mechanism to strike a balance between information density and visual balance: nodes on high-probability paths occupy larger spaces and adopt a high-contrast style, while low-probability nodes adopt a compact layout and a low visual weight design.
[0104] Ultimately, the optimized third-tier layout not only ensures narrative coherence but also guides the user's gaze along a pre-defined path through spatial arrangement. For example, in a business report, the conclusion is assigned to the top of the page to quickly convey core information, while data charts are arranged vertically according to the analytical steps, forming a visual narrative structure of "conclusion first, progressive demonstration." This process was verified through multiple rounds of iteration to ensure that the layout not only meets user preferences but also adapts to the narrative needs of different scenarios.
[0105] As an optional embodiment, in step S107, the narrative path in the presentation is predicted based on the logical relationship information and the user narrative preference marked based on the user historical data to obtain a probability distribution of the page narrative path, including:
[0106] Knowledge graph technology is adopted, and the visual elements in the current page are used as nodes according to the logical relationship information, and the logical relationships between the visual elements are used as edges to form a structured page logical relationship graph; user preference features are extracted based on the user narrative preferences, wherein the user preference features include at least layout preference features, focus preference features, and narrative advancement method preference features; a hidden Markov model is adopted, combined with the page logical relationship graph and the user preference features, to generate a transition probability matrix that matches the user narrative preferences.
[0107] Then, in step S108, based on the probability distribution of the page narrative path, the second layout information is optimized to obtain third layout information, including:
[0108] The nodes of the high-probability path in the transfer probability matrix are assigned to the corresponding visually sensitive areas according to the node information type, and the nodes of the low-probability path are compressed to the edge area or merged for layout to obtain logical layout adjustment information; based on the logical layout adjustment information, the page layout in the second layout information is corrected to obtain the third layout information.
[0109] In the above optional embodiment, in step S107, a page logical relationship graph is first constructed based on the logical relationship information. The graph abstracts the visual elements (such as titles, charts, key data, etc.) in the presentation page as nodes, and the logical associations between elements (such as cause and effect, parallelism, and progression) are converted into weighted edges. For example, if there is a title "Market Status" on the page that is compared with the subsequent bar chart, the system extracts the title keywords through natural language processing and matches them with the chart labels to form a logical link of "question → data verification". The construction of the knowledge graph not only relies on semantic associations, but also needs to be combined with visual layout features (such as the spatial proximity between elements) to ensure the consistency of the logical relationship with the user's visual perception.
[0110] Next, narrative preference features are extracted from historical user data, including layout preferences (such as a user's preference for a timeline or grid-style design), focus preferences (such as the duration of focus on a chart or title), and narrative progression preferences (such as linear or branching). These features are generated by analyzing users' past PPT interactions (such as page-turning order and click hotspots) and content structure (such as the frequency of chapter divisions), forming a user preference vector. For example, if a user frequently uses the linear "problem → solution" structure, the model will increase the weight of the causal edge.
[0111] Based on the knowledge graph and user preference features, the system uses a hidden Markov model (HMM) to generate a transition probability matrix. The HMM's state space is defined as narrative nodes (e.g., "problem statement," "data analysis"), and transition probabilities are dynamically adjusted based on the strength of logical relationships (e.g., causal weights) and user preferences. For example, if a user prefers a data-driven narrative, the transition probabilities of key diagram nodes will be increased, making the model more inclined to generate a path from "data presentation to conclusion deduction." A forward algorithm calculates the probability distribution of all possible paths, ultimately outputting a high-probability narrative path sequence.
[0112] In step S108, the second layout information is optimized according to the transition probability matrix. First, the nodes of the high-probability path (such as core conclusions, key data) are assigned to visually sensitive areas. These areas are usually located at the visual focus of the page, such as the "golden triangle area" in the upper right corner or the information focus zone in the middle and lower part. For example, the conclusion node with high probability will be assigned to the top of the page to ensure that the user's eyes can capture the core information at the first time. The nodes of the low-probability path (such as supplementary instructions) are compressed to the edge area (such as the bottom of the page or the sidebar), or a foldable design is used to reduce visual interference.
[0113] Next, the second layout information is modified based on the logical layout adjustment information to generate the third layout information. This process combines the global optimization capabilities of the quantum annealing algorithm to dynamically adjust the position and hierarchy of elements. For example, if the predicted path requires the "problem statement" and "solution" to form a visual contrast, the algorithm will assign the two to the left and right symmetrical areas of the page and strengthen the logical connection through color contrast or arrow guidance. At the same time, the algorithm introduces a dynamic weighting mechanism to strike a balance between information density and visual balance: nodes on high-probability paths occupy larger spaces and adopt a high-contrast style, while low-probability nodes adopt a compact layout and a low visual weight design.
[0114] Thus, through the collaborative mechanism of the knowledge graph, preference model, and probability matrix proposed in the above embodiment, a deep binding of narrative logic and user preferences is achieved. For example, in a business report scenario, if the user prefers a linear narrative from problem to benefit, the key data charts are automatically placed in the middle and lower part of the page (in line with the visual movement line), and the conclusion part is allocated to the focus area in the upper right corner, forming a visual structure of "conclusion first, step-by-step argumentation". In academic scenarios, for the logical chain of "hypothesis → experiment → conclusion", priority is given to ensuring the vertical alignment of time series elements, and enhancing the sense of hierarchy by leaving white space.
[0115] Ultimately, the optimized layout went through multiple rounds of iterative verification to ensure it aligned with user preferences and adapted to the narrative needs of different scenarios. For example, if a user's historical data included frequent use of "multi-branch comparison," the probability matrix would be dynamically adjusted to assign higher weights to branching narrative paths and reserve space for multi-region comparison on the page.
[0116] Step S109: performing secondary layout of the presentation page based on the third layout information to optimize the page structure of the presentation page.
[0117] As an optional embodiment, after performing secondary layout of the presentation page based on the third layout information in step S109, the following steps are further included:
[0118] Decompose the presentation page after secondary typesetting into hexagonal cells that imitate the form of bee nesting to obtain a hexagonal layout diagram of the presentation page; detect whether the average information density of each hexagonal cell after secondary typesetting is lower than a preset density threshold; if the average information density after secondary typesetting reaches the preset density threshold, use the honeycomb structure self-organizing algorithm to rearrange the hexagonal cells where each visual element in the hexagonal layout diagram is located to obtain fourth layout information; detect whether the average information density of each hexagonal cell after rearrangement is lower than the preset density threshold; if the average information density after rearrangement is lower than the preset density threshold, adjust the typesetting layout between each visual element in the presentation page according to the fourth layout information to reserve the blank space in the presentation page and ensure that the spacing distance between the visual elements meets the preset conditions.
[0119] Specifically, the innovation of the hexagonal layout in step S109 is that the page is decomposed into hexagonal cells that imitate beehives, and each cell carries one or more visual elements. First, based on the element coordinates in the second layout information, the visual elements are distributed to the hexagonal cells through a spatial mapping algorithm to form a dense structure similar to a honeycomb. This process adopts a dynamic density calculation model: the information density of each cell is comprehensively calculated by the element size, the number of text lines and the image resolution, and compared with the preset threshold. If the density exceeds the standard, the honeycomb structure self-organization algorithm is triggered, and the elements in the cell are reallocated to adjacent free areas by simulating the collaborative mechanism of the bee colony to build a nest. For example, high-density chart cells will trigger information migration behavior, migrating some data points to adjacent low-density cells while maintaining logical relevance. The algorithm simulates the transmission of bee pheromones through particle swarm optimization, dynamically adjusts the distribution of elements, and controls the overall density fluctuation of the page, which is significantly better than the rigid structure of the traditional fixed grid layout.
[0120] The core innovation of the honeycomb self-organizing algorithm lies in its introduction of a hierarchical energy field model. The system treats the hexagonal layout as an energy field, with each cell possessing both attractive and repulsive potential energies. The attractive potential encourages elements to cluster in areas of high information density, while the repulsive potential prevents overcrowding. For example, when excessive density is detected, the algorithm activates the honeycomb reconstruction protocol, optimizing the layout through the following steps: Differentiated potential weights are assigned based on element type (such as titles and charts), with titles receiving greater attractiveness due to their high visual weight, and charts receiving greater repulsiveness due to their information-carrying capacity. Using a federated learning framework, each cell independently assesses migration benefits, achieving a globally optimal migration plan through a blockchain-style consensus mechanism. Furthermore, a fractal growth algorithm is introduced to restructure the migrated elements into self-similar structures, for example, splitting linearly arranged text into honeycomb-shaped nested modules. This improves the readability of key data points in business proposal PPTs while optimizing page whitespace, in line with the "moderate whitespace effect" in cognitive psychology.
[0121] To balance information density and visual comfort, the system constructs a dynamic density-whitespace balance model. After each honeycomb reconstruction, a convolutional neural network analyzes the visual flow of the page, identifying "information overload zones" and "visual buffer zones." "Information distillation" is performed on overload zones: redundant elements (such as repeated data labels) are compressed into compact icons to free up space; guiding whitespace elements (such as gradient transparent blocks) are injected into the buffer zone to create a breathable layout. This process utilizes a reinforcement learning framework, with a reward function consisting of three dimensions: a density penalty term, which applies an exponential decay penalty to cells exceeding a threshold; an aesthetic evaluation term, which calculates the harmony of element spacing based on the golden ratio; and a cognitive load term, which quantifies the complexity of the visual search path using eye-tracking data.
[0122] Application in financial data reports shows that the model reduces the peak information density of the page and improves the accuracy of audience information recall, verifying the effectiveness of the dynamic white space strategy.
[0123] A cross-scale optimization mechanism is introduced to implement coordinated adjustments at the cell and page levels. After the cells are reconstructed, a page-level "honeycomb resonance" algorithm is triggered: the frequency domain characteristics of the page layout are analyzed through fast Fourier transforms, the spatial patterns corresponding to the dominant frequencies (such as horizontal stripes or radial distributions) are identified, and phase offset adjustments are applied. For example, when a 0.5Hz low-frequency oscillation (periodic changes in element spacing) is detected on the page, a 10% random perturbation is automatically introduced to break the regularity and enhance visual dynamic balance. At the same time, based on the cellular automaton model, layout changes in adjacent cells trigger local chain reactions, forming self-organizing textures. This mechanism accelerates the layout convergence speed of complex pages (containing more than 200 elements) by several times, and the human evaluation satisfaction score reaches 92 points (out of 100), far exceeding the 78 points of traditional heuristic algorithms.
[0124] In this way, through the deep integration of bionic principles and intelligent algorithms, a paradigm shift from mechanical typesetting to organic growth has been achieved, providing a new methodological framework for digital content design.
[0125] Further optionally, in the above step, decomposing the presentation page after secondary typesetting into hexagonal cells imitating the form of bee nests to obtain a hexagonal layout diagram of the presentation page includes:
[0126] The core visual element is used as the seed point, and the location of the core visual element is used as the core cell. Hexagonal cells are divided layer by layer from the seed point outward, forming hexagonal cells arranged around the core cell to obtain a hexagonal layout diagram. The core visual elements include the title and main chart; the hexagonal cells in the core area of the page are regular hexagons, and the hexagonal cells at the edge of the page are deformed hexagons.
[0127] Furthermore, in the above steps, detecting whether the average information density of each hexagonal cell after the secondary typesetting is lower than a preset density threshold includes: calculating the content information density and / or visual information density of each hexagonal cell; setting a dynamic weight coefficient of each hexagonal cell according to the type of visual elements in each hexagonal cell; and using the dynamic weight coefficient to weightedly calculate the content information density and / or visual information density of each hexagonal cell according to the content type to obtain the average information density after the secondary typesetting.
[0128] The hexagonal cell division method used in the above steps incorporates a fractal growth algorithm, constructing a hierarchical layout structure using core visual elements as seed points. First, a visual saliency detection algorithm (such as a deep learning-based attention model) identifies the locations of core elements such as titles and main charts, using these as initial seed points to generate regular hexagonal cells. Subsequently, based on the spatial segmentation principle of the Voronoi diagram, hexagonal rings are expanded outward from the seed point. The cell size of each new ring is dynamically reduced according to the Fibonacci sequence (for example, the side length of each layer is reduced by 12%), forming a golden ratio structure similar to that of a natural honeycomb. When a cell reaches the page boundary, an edge adaptation algorithm is triggered: an affine transformation is used to adjust the regular hexagon to a chamfered hexagon (the chamfer angle is dynamically calculated based on the remaining space on the page) to ensure topological consistency between the edge cells and the core area.
[0129] A multi-dimensional dynamic weighting system was constructed to calculate the information density of hexagonal cells. First, a visual element type classifier (based on a ResNet-50 transfer learning model) identified cell content categories (such as titles, charts, and text) and assigned a basic weight coefficient (0.9 for titles, 0.7 for charts, and 0.5 for text) to each type. Subsequently, content information density was calculated using a hybrid evaluation metric: for text regions, a weighted product of character density (number of characters per region area) and semantic complexity (based on the dimensionality of the text vectors derived from BERT was used); for chart regions, a combination of data point density (number of data points per chart area) and visual contrast (difference in the HSV color space) was used. Visual information density was determined by extracting feature maps using a pre-trained convolutional neural network (such as EfficientNet-B7). The proportion of high-frequency components (reflecting detail complexity) and the global mean contrast were calculated for the cell region. Finally, dynamic weighting coefficients were adjusted in real time based on page zoom and device type (with a 20% increase in the weight of titles on mobile devices) to ensure that density evaluation results are consistent with human perception under different display conditions.
[0130] To address the problem of traditional fixed thresholds being unable to adapt to complex layouts, the system designed a dynamic density threshold calibration mechanism based on reinforcement learning. First, a visual comfort dataset consisting of 200,000 manually annotated samples was constructed. A reward function was defined with three dimensions: information accessibility (element recognition time in high-density areas), visual comfort (perceived white space score in low-density areas), and logical coherence (length of element-linking paths). The agent was trained using the Proximal Policy Optimization (PPO) algorithm, enabling it to dynamically adjust the density threshold based on the current layout state. For example, if the average cell density in the core area exceeds 0.85 (normalized value), the system automatically lowers the threshold to 0.75 and triggers honeycomb reconstruction. If the density in the edge area falls below 0.45, the threshold is raised to 0.55 to suppress excessive white space. In business proposal PPT testing, this mechanism reduced the fluctuation range of page information density and improved audience ratings of layout rationality.
[0131] To visualize density detection results, the system developed a three-dimensional density heat map based on volume rendering technology. By mapping the density values of hexagonal cells to voxel transparency (with transparency reduced to 30% in high-density areas), a density distribution visualization effect similar to X-ray imaging is created in three-dimensional space. Simultaneously, an isosurface extraction algorithm (MarchingCubes) is used to generate a density gradient surface, with key areas annotated with different hues (red for high density, blue for low density). This visualization module supports interactive analysis, allowing users to rotate and zoom the heat map using gestures and view the density composition of specific cells in real time (for example, clicking on a title cell breaks down the contribution ratio of text density to chart density). Educational PPT tests have shown that this feature improves teachers' efficiency in adjusting layouts and increases students' accuracy in understanding information density.
[0132] Thus, through the coordinated optimization of bionic structure construction, dynamic weight modeling and intelligent threshold calibration, a leap from mechanical cell division to cognitive-driven density regulation is achieved, providing new methodological support for intelligent typesetting systems.
[0133] Further optionally, the adopting of a honeycomb structure self-organizing algorithm to rearrange the regular hexagonal cells where each visual element in the hexagonal layout diagram is located to obtain the fourth layout information includes:
[0134] If the content information density of the hexagonal cell reaches a first set threshold, then based on the dynamic growth mechanism, adjacent hexagonal cells are merged to form a complex unit module to obtain a first adjustment instruction; based on the shape of the visual elements contained in the hexagonal cell, the shape of the complex unit module is set, and the Catmull-Clark subdivision algorithm is used to smooth the boundaries of the complex unit module; if the density difference between adjacent hexagonal cells is greater than a second set threshold, a second adjustment instruction for instructing the migration of visual elements is triggered; wherein, the second adjustment instruction is provided with a migration path according to the shortest Voronoi edge rule to migrate the visual elements in the high-density hexagonal cell to the low-density hexagonal cell; according to the preset direction matching the page type and page size, a visual movement line planning and space blanking strategy for the hexagonal layout diagram are generated; based on the first adjustment instruction, the second adjustment instruction, the visual movement line planning and the space blanking strategy, the fourth layout information for the hexagonal layout diagram is generated.
[0135] The above steps draw on the efficient structure and self-organizing characteristics of the honeycomb in nature. The visual elements in the presentation page are placed in regular hexagonal cells, with the content information density as the core indicator to trigger adjustments. When the cell content information density reaches the first set threshold, the dynamic growth mechanism simulating cell division is activated, and complex unit modules are formed by merging adjacent cells to accommodate more information or complex content. This process can not only flexibly adjust the layout structure according to information carrying requirements, but also optimize space utilization. At the same time, the migration of visual elements is triggered based on the density difference between adjacent cells, simulating the information diffusion phenomenon, making the page layout more balanced in information distribution. In addition, the visual movement line planning and space blanking strategy are set according to the page type and size, which conforms to the laws of human visual cognition, guides the audience's sight, and improves the efficiency of information communication and visual comfort.
[0136] In terms of algorithmic model application, the dynamic growth algorithm, Catmull-Clark subdivision algorithm, and Voronoi diagram algorithm are mainly used. The dynamic growth algorithm controls the merging and module formation of cells based on content density thresholds, providing a structured approach to the layout of complex content. The Catmull-Clark subdivision algorithm is used to smooth the boundaries of complex cell modules, making the layout edges more natural and smooth, and enhancing the visual aesthetic. The Voronoi diagram algorithm plays a key role in the element migration process. By determining the shortest migration path, it ensures the efficient and reasonable movement of visual elements between different cells, maintaining the stability and logic of the page layout.
[0137] Specifically, the dynamic growth algorithm uses content density as its core driver, mimicking the mechanism of biological cell division to achieve a structured layout of complex content. When the cell content density exceeds 1.2 (calculated based on parameters such as character row height and chart data point density), the algorithm triggers the cell division process: first, the density distribution of 6-8 adjacent cells is scanned. A federated learning framework is used to evaluate the merging benefits, prioritizing regions with high semantic relevance for fusion. For example, when merging a main chart with auxiliary explanatory text, the system uses an attention mechanism to calculate the interaction strength between visual elements and dynamically adjust the module shape. Complex charts use regular dodecagons to accommodate multi-dimensional data presentation, while key text is generated in a star-shaped structure to highlight core information. This process incorporates a fractal growth algorithm, iteratively generating a hierarchical layout. The size of each new module is reduced according to the golden ratio (edge length coefficient 0.618) to ensure a natural transition between visual levels. This improves the efficiency of layout of complex data modules in reports and enhances the accuracy of user information retrieval.
[0138] To address the jagged edges of complex modules, the system uses a modified Catmull-Clark subdivision algorithm for smoothing. Recursive subdivision of face points, edge points, and vertices generates a surface, which is then adapted to optimize the boundaries of a two-dimensional layout. The module polygons are first decomposed into a quadrilateral mesh. Feature point detection (such as inflection points and concave points) identifies areas requiring smoothing. Subdivision rules are then applied, with each face point being updated to a weighted average of its adjacent face and edge points. Edge points are then recalculated based on information from both their endpoints and adjacent face points. For example, the sharp corners of star-shaped modules are smoothed by increasing the subdivision level (the default is 3), while the edges of regular dodecagons use an adaptive subdivision step size (greater edge curvature requires more subdivisions). To maintain layout accuracy, the algorithm incorporates an error compensation mechanism: after each subdivision, the module area change rate is detected. If it exceeds a threshold (e.g., 5%), local resampling is triggered.
[0139] During element migration, a Voronoi diagram algorithm constructs a dynamic spatial partitioning model to ensure optimal migration paths and logical coherence. The system first generates a Voronoi diagram based on the current cell layout, dividing the page into multiple convex polygonal regions. The Voronoi vertices in each region serve as potential migration nodes. When elements in high-density areas need to be diffused, the algorithm constructs an adjacency graph using Delaunay triangulation, calculates the shortest path between each node (optimized using the A* algorithm), and introduces a dynamic weighting factor based on a distance decay coefficient (0.8) and visual relevance (calculated based on element semantic similarity). For example, when a title migrates to a blank area, the system prioritizes target cells that are adjacent to the original Voronoi region and have a high semantic relevance, forming a logically coherent visual flow. The Voronoi diagram structure is updated in real time during the migration process, maintaining computational efficiency through incremental triangulation. This mechanism shortens element migration paths and maintains layout stability in cross-page layouts.
[0140] In this way, the module boundaries generated by the dynamic growth algorithm serve as the input mesh for the Catmull-Clark algorithm, while the Voronoi diagram algorithm updates the migration path based on the real-time layout, forming a closed loop of "density perception-morphological optimization-path planning." In a medical data visualization case study, this system successfully shortened the layout adjustment time for over 200 detection indicators and improved the logical rationality scores evaluated by experts. This method, which combines geometric processing and intelligent optimization, provides a new technical paradigm for presenting complex information.
[0141] In this way, the presentation page layout is both intelligent and beautiful. The dynamic growth mechanism and element migration strategy effectively solve the problem of uneven page content density, avoid information crowding or excessive white space, and make the page layout more reasonable and orderly. The complex unit modules processed by the Catmull-Clark subdivision algorithm have a more refined visual effect and reduce the abruptness caused by module splicing. Visual movement line planning and space white space strategy enable the audience to receive information more clearly and comfortably, enhancing the readability and attractiveness of the presentation. The final generated fourth layout information allows the presentation to achieve a higher level of visual presentation while maintaining the integrity of the content, meeting diverse display needs.
[0142] Further optionally, after the honeycomb structure self-organizing algorithm is used to rearrange the regular hexagonal cells where the visual elements in the hexagonal layout diagram are located to obtain the fourth layout information, the method further includes:
[0143] Monitor user operations in real time; if it is detected that the user operation is a gesture zoom, adjust the side length of the regular hexagonal cell in real time to change the global layout size of the hexagonal layout diagram synchronously with the zoom range to maintain content readability; if it is detected that the user operation is a drag operation on the first visual element, reorganize the cells of the second visual element adjacent to the operated first visual element.
[0144] This step focuses on user interaction and dynamic layout response. By monitoring user operations in real time and adjusting the layout instantly, it aims to improve the interactive experience and content presentation effect during the use of presentations, and form a complete dynamic layout optimization system.
[0145] At the principle level, it is based on the principles of human-computer interaction and layout adaptation. The system monitors user operations in real time and regards them as feedback signals for the current layout. When a gesture zoom operation is detected, taking into account the readability and visual comfort of the content, the system adjusts the side length of the regular hexagonal cell in real time according to the zoom range, thereby synchronously changing the global layout size. This process is based on the correlation between proportional scaling and visual perception, ensuring that text and graphic elements remain clearly distinguishable at different zoom levels. When a drag operation is detected, based on the spatial correlation between elements in the honeycomb structure, the system will reorganize the cells of the elements adjacent to the operated element. Because in the honeycomb layout, each cell is closely connected to the surrounding elements, the position change of an element will affect the spatial adaptability of its adjacent elements. The overall coordination and logic of the layout are maintained through cell reorganization.
[0146] In the application of algorithm models, it mainly involves real-time monitoring algorithm, adaptive scaling algorithm and cell reorganization algorithm. The real-time monitoring algorithm is responsible for continuously capturing the user's gestures and dragging actions to ensure that the operation can be perceived in a timely manner. The adaptive scaling algorithm accurately calculates and adjusts the side length of the regular hexagonal cells according to the zoom ratio to ensure the rationality and readability of the element layout under different zoom states. After the drag operation occurs, the cell reorganization algorithm rearranges the cells where adjacent elements are located based on the topological relationship of the honeycomb structure. For example, by calculating the distance and angle between elements and the space occupied by cells, the cell position is reallocated to adapt to the layout changes caused by the change in the position of the first visual element and maintain the stability of the overall layout.
[0147] As a result, the interactivity and flexibility of presentation layouts are significantly improved. Real-time response to gesture zooming operations allows users to freely adjust the page size according to their needs. Whether zooming in to view details or zooming out to control the overall situation, the content can remain clear and readable, enhancing the user's interactive experience with the presentation. The cell reorganization mechanism for drag operations ensures that when the user adjusts the position of an element, the surrounding elements can be automatically rearranged to avoid layout confusion or element overlap, so that the page always maintains a neat and orderly visual effect. This real-time and intelligent layout adjustment ensures that the presentation always presents the best visual state during the user's operation, greatly improving the practicality of the presentation and user satisfaction.
[0148] In another optional embodiment, after performing secondary layout of the presentation page based on the third layout information, it also includes: identifying the terminal device used to display the presentation page; using the perspective parameters corresponding to the type of terminal device and the device screen size to adjust the perspective relationship between each visual element in the presentation page in real time, so that the positional relationship between each visual element adapts to the device screen and avoids positional distortion between each visual element; wherein, the desktop device adopts a wide-angle perspective and the mobile device adopts a narrow-angle perspective.
[0149] This step focuses on dynamically adjusting the perspective of the presentation based on the characteristics of the terminal device, and ensuring the presentation effect of visual elements on various terminals by adapting to different screen sizes and device types.
[0150] At the principle level, it is based on the perspective projection principle and device adaptation logic. The screen size, resolution and display ratio of different terminal devices are different, which directly affects the presentation of the spatial relationship of visual elements. After the system identifies the type of terminal device, it will call the corresponding perspective parameters (such as wide-angle or narrow-angle perspective) according to the differences in hardware characteristics between the desktop and mobile terminals. Wide-angle perspective (FOV=90°) is suitable for larger screens on the desktop. It can enhance the sense of spatial hierarchy by expanding the viewing angle range and make the elements more naturally distributed on the wide screen; narrow-angle perspective (FOV=60°) is optimized for small screens on mobile terminals. By narrowing the viewing angle, it avoids deformation or overlap of elements due to screen squeezing. At the same time, the actual size of the device screen will be combined to calculate the position, size and depth relationship of visual elements in real time, and the relative distance and angle between elements will be adjusted through the perspective transformation matrix to ensure that the layout conforms to the visual habits of the human eye on different devices and avoid position distortion.
[0151] The application of algorithmic models mainly involves device identification algorithms, perspective transformation algorithms, and adaptive layout algorithms. The device identification algorithm quickly determines whether the device belongs to a desktop or mobile terminal by reading the terminal's hardware parameters (such as screen resolution, DPI, and device type identification), providing a basis for subsequent perspective parameter calls. The perspective transformation algorithm is based on the pinhole camera model, projecting visual elements in three-dimensional space onto a two-dimensional screen, and controlling the perspective effect by adjusting the field of view (FOV) parameters: the desktop terminal increases the FOV to expand the horizontal viewing angle, and the mobile terminal reduces the FOV to compress the vertical depth space. The adaptive layout algorithm dynamically calculates the scaling ratio and position offset of elements based on the screen size. For example, on mobile terminals, it automatically reduces the size of non-critical elements to reserve sufficient space for core content, maintaining the logical hierarchical relationship between elements.
[0152] From a technical perspective, adaptive presentation of presentations on multiple terminals is achieved. For desktops, wide-angle perspective can make the spatial relationship of complex layouts (such as multiple modules in parallel, 3D charts) clearer, and the audience can intuitively perceive the front and back layers and logical connections of elements; the narrow-angle perspective of mobile terminals effectively solves the problem of element crowding on small screens, avoids text overlap, chart deformation, etc., and ensures content readability. In addition, real-time perspective adjustment eliminates the need for manual rearrangement of presentations when switching across devices. The system automatically completes layout adaptation, improving the flexibility and consistency of content display. For example, a user uses a desktop terminal to present a three-dimensional layout with a wide-angle perspective in a meeting. When switching to a mobile phone for viewing, it automatically switches to a narrow-angle perspective and optimizes the spacing between elements, ensuring that the presentation effect is not limited by the device, and achieving an efficient experience of "one-time production, multi-terminal adaptation".
[0153] As an optional embodiment, after performing secondary typesetting of the presentation page based on the third typesetting information in step S106, the following steps are also included: performing emotional tendency recognition on the semantic information of visual elements and the differences in semantic information between different visual elements to obtain the single-page content emotional type of the presentation page; performing collaborative prediction on multiple pages in the presentation to obtain the overall content emotional type of the presentation; based on the overall content emotional type and the single-page content emotional type, selecting matching visual colors and / or layout styles from the candidate typesetting layout library, and using the matching visual colors and / or layout styles to adjust the color parameters and / or layout of the presentation page.
[0154] In these steps, the system mines the semantic information of visual elements within a presentation to analyze emotional tendencies and then matches appropriate visual colors and layout styles, aiming to enhance the presentation's message's impact and relevance. In principle, this is based on the fusion of natural language processing and visual design. The text, images, and other information carried by visual elements themselves contain semantic content. Semantic differences between different elements can reflect specific emotional tendencies. For example, motivational text paired with upbeat images conveys positive emotions, while cautionary text paired with somber images conveys negative emotions. The system first analyzes the semantic information of visual elements to identify the underlying emotional tone. Then, based on the semantic connections between different elements, it determines the emotional type of a single page. For the entire presentation, it analyzes the consistency and changing trends of the emotional types across each page to achieve collaborative prediction and determine the overall emotional type of the content. Finally, based on the correlation between the emotional type and the visual design, it selects matching visual colors and layout styles from a library of candidate layouts. For example, positive emotional content is best suited to bright, warm colors and lively layout styles, while negative emotional content is matched with calm, cool tones and a simple layout, thereby enhancing emotional resonance and information transmission. In terms of algorithmic model application, these primarily involve sentiment analysis models, collaborative prediction models, and matching retrieval models in natural language processing. Sentiment analysis models typically employ deep learning-based architectures, such as recurrent neural networks (RNNs) and their variants, LSTMs, GRUs, or Transformer models. They extract and analyze textual features of the semantic information of visual elements to determine their emotional tendencies. Collaborative prediction models, based on time series analysis or graph neural networks (GNNs), treat multiple pages of a presentation as nodes with temporal or logical connections, exploring the evolution of sentiment patterns across pages and predicting overall sentiment trends. Matching retrieval models, based on a pre-built library of candidate layouts, calculate the similarity between sentiment patterns and the visual colors and layout styles in the library to quickly identify the most suitable design options, achieving a precise match between sentiment and visual design.
[0155] Through emotional propensity recognition and collaborative prediction, the color and layout of presentations are no longer limited to traditional aesthetic design, but are deeply integrated with the emotional content. This not only enhances the visual expressiveness of presentations, but also strengthens the audience's understanding and memory of the content through emotional resonance, making the message more targeted and appealing. This is particularly suitable for scenarios such as brand promotion and emotional storytelling, where visual and emotional consistency is crucial.
[0156] In the embodiment of the present application, the automatic layout of the presentation page can be achieved by combining page logic and visuals, the page structure can be optimized, the layout efficiency of the presentation page can be improved, and the work efficiency of personnel in related positions can be improved.
[0157] After introducing the method of the exemplary embodiment of the present application, next, refer to Figure 2 A system for optimizing the page structure of a presentation according to an exemplary embodiment of the present application is described. The system includes: an acquisition module for acquiring a presentation page input by a user;
[0158] The acquisition module is used to obtain the presentation page input by the user;
[0159] A recognition module is used to identify visual elements in a presentation page and perform logical relationship analysis on the visual elements to obtain visual structure information corresponding to the presentation page;
[0160] The parsing module is used to parse the logical relationships between visual elements and obtain the logical relationship information between visual elements in the presentation page;
[0161] The first optimization module is used to map the visual structure information into the 3D implicit field of the page based on NeRF to obtain the 3D voxels corresponding to the visual elements. The ray tracing algorithm is used to calculate the visual focus heat map of the presentation page and assign the 3D voxels corresponding to the target visual elements to the visual sensitive area to obtain the first layout information.
[0162] a second optimization module, configured to use a quantum annealing algorithm to predict a global optimal layout of the presentation page based on the positional relationship between the visual elements in the visual structure information and the first layout information in combination with an information density energy function, thereby obtaining second layout information;
[0163] a third optimization module configured to predict the narrative path in the presentation based on the logical relationship information and the user's narrative preferences marked based on the user's historical data, to obtain a probability distribution of the page's narrative path; and optimize the second layout information based on the page's narrative path probability distribution to obtain third layout information, so that the spatial arrangement of visual elements in the presentation page conforms to the narrative logic preferred by the user;
[0164] An execution module is used to perform secondary layout of the presentation page based on the third layout information to optimize the page structure of the presentation page.
[0165] The above system can implement each step described in the above method implementation, and the specific implementation method of each step will not be repeated here.
[0166] After introducing the method and system of the exemplary embodiment of the present application, a terminal device of the exemplary embodiment of the present application is described below. The terminal device can implement each step described in the above method embodiment, and the specific implementation of each step is not repeated here.
[0167] After introducing the method, system and terminal device of the exemplary embodiment of the present application, the following Figure 3 For a description of the computer-readable storage medium of the exemplary embodiment of the present application, please refer to Figure 3 The computer-readable storage medium shown is an optical disc 30, which stores a computer program (i.e., a program product). When executed by a processor, the computer program implements each step described in the above method implementation. The specific implementation of each step is not repeated here.
[0168] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which will not be described in detail here. The above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-mentioned embodiments, ordinary technicians in this field should understand that any technician familiar with this technical field can still modify the technical solutions described in the above-mentioned embodiments within the technical scope disclosed in this application, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A method for optimizing the page structure of a presentation, characterized in that: The method comprises: Get the presentation page entered by the user; Identify visual elements in a presentation page and obtain visual structure information corresponding to the presentation page; Analyze the logical relationships between visual elements to obtain the logical relationship information between visual elements in the presentation page; Based on the 3D neural radiation field model constructed by NeRF, the visual structure information is mapped to the 3D implicit field of the page to obtain the 3D voxels corresponding to the visual elements; Using the ray tracing algorithm, we calculated the visual focus heat map in the presentation page and assigned the three-dimensional voxels corresponding to the target visual elements to the visually sensitive area to obtain the first layout information. A quantum annealing algorithm is used to predict the global optimal layout of the presentation page based on the positional relationship between each visual element in the visual structure information and the first layout information, combined with an information density energy function, to obtain the second layout information; Predicting the narrative path in the presentation based on the logical relationship information and the user's narrative preferences marked based on the user's historical data to obtain a probability distribution of the page's narrative path; Based on the probability distribution of the page narrative path, the second layout information is optimized to obtain third layout information, so that the spatial arrangement relationship of the visual elements in the presentation page conforms to the narrative logic preferred by the user; A secondary layout of the presentation page is performed based on the third layout information to optimize the page structure of the presentation page.
2. The method for optimizing the page structure of a presentation according to claim 1, wherein: The process of predicting the narrative path in the presentation based on the logical relationship information and the user narrative preference marked based on the user historical data to obtain a probability distribution of the page narrative path includes: Using knowledge graph technology, visual elements in the current page are used as nodes according to the logical relationship information, and the logical relationships between the visual elements are used as edges to form a structured page logical relationship graph; Extracting user preference features based on the user narrative preferences, wherein the user preference features include at least layout preference features, focus preference features, and narrative advancement mode preference features; Using a hidden Markov model, combined with the page logic relationship map and the user preference characteristics, to generate a transition probability matrix that matches the user's narrative preference; The optimizing the second layout information based on the page narrative path probability distribution to obtain third layout information includes: Allocating nodes of high-probability paths in the transfer probability matrix to corresponding visually sensitive areas according to node information types, and compressing nodes of low-probability paths to edge areas or merging them for layout, to obtain logical layout adjustment information; Based on the logical layout adjustment information, the page layout in the second layout information is corrected to obtain the third layout information.
3. The method for optimizing the page structure of a presentation according to claim 2, wherein: After performing secondary layout of the presentation page based on the third layout information, the method further includes: The semantic information of visual elements and the differences in semantic information between different visual elements are used to identify emotional tendencies, so as to obtain the emotional type of the content of a single page of the presentation page; Collaborative prediction is performed on multiple pages in a presentation to obtain the overall content sentiment type of the presentation; Based on the overall content sentiment type and the single-page content sentiment type, a matching visual color and / or layout style is selected from a candidate typesetting layout library, and the matching visual color and / or layout style is used to adjust the color parameters and / or layout of the presentation page.
4. The method for optimizing the page structure of a presentation according to claim 2, wherein: After performing secondary layout of the presentation page based on the third layout information, the method further includes: Decomposing the presentation page after secondary typesetting into hexagonal cells that imitate the form of bee nests to obtain a hexagonal layout diagram of the presentation page; Detect whether the average information density of each hexagonal cell after secondary typesetting is lower than a preset density threshold; If the average information density after the second layout reaches a preset density threshold, a honeycomb structure self-organizing algorithm is used to rearrange the hexagonal cells where the visual elements in the hexagonal layout are located to obtain the fourth layout information; Detecting whether the average information density of each hexagonal cell after rearrangement is lower than a preset density threshold; If the average information density after rearrangement is lower than the preset density threshold, the layout of each visual element in the presentation page is adjusted according to the fourth layout information to reserve the blank space in the presentation page and ensure that the spacing distance between the visual elements meets the preset conditions.
5. The method for optimizing the page structure of a presentation according to claim 4, characterized in that: The method of decomposing the secondary typesetting presentation page into hexagonal cells imitating the bee nesting form to obtain a hexagonal layout diagram of the presentation page includes: The core visual element is used as the seed point, and the location of the core visual element is used as the core cell. Hexagonal cells are divided layer by layer from the seed point outward to form hexagonal cells arranged around the core cell to obtain a hexagonal layout diagram. The core visual elements include the title and the main chart. The hexagonal cells in the core area of the page are regular hexagons, and the hexagonal cells at the edge of the page are deformed hexagons. The detecting whether the average information density of each hexagonal cell after the secondary typesetting is lower than a preset density threshold includes: Calculating the content information density and / or visual information density of each hexagonal cell; According to the type of visual elements in each hexagonal cell, set the dynamic weight coefficient of each hexagonal cell; The dynamic weight coefficient is used to perform weighted calculation on the content information density and / or visual information density of each hexagonal cell according to the content type, so as to obtain the average information density after secondary typesetting.
6. The method for optimizing the page structure of a presentation according to claim 5, characterized in that: The honeycomb structure self-organizing algorithm is used to rearrange the regular hexagonal cells where the visual elements in the hexagonal layout diagram are located to obtain the fourth layout information, including: If the content information density of the hexagonal cells reaches a first set threshold, adjacent hexagonal cells are merged to form a complex unit module based on a dynamic growth mechanism to obtain a first adjustment instruction; the shape of the complex unit module is set based on the shape of the visual elements contained in the hexagonal cells, and the Catmull-Clark subdivision algorithm is used to smooth the boundaries of the complex unit module; If the density difference between adjacent hexagonal cells is greater than a second set threshold, a second adjustment instruction for instructing the migration of the visual element is triggered; wherein the second adjustment instruction includes a migration path set according to the shortest Voronoi edge rule, for migrating the visual element in the high-density hexagonal cell to the low-density hexagonal cell; Generate visual movement planning and space white space strategy for hexagonal layouts according to the preset direction matching page type and page size; Based on the first adjustment instruction, the second adjustment instruction, the visual movement line planning, and the space blanking strategy, the fourth layout information for the hexagonal layout diagram is generated.
7. The method for optimizing the page structure of a presentation according to claim 4, wherein: After the honeycomb structure self-organizing algorithm is used to rearrange the regular hexagonal cells where the visual elements in the hexagonal layout diagram are located to obtain the fourth layout information, the method further includes: Real-time monitoring of user operations; If a zoom gesture is detected, the side length of the regular hexagonal cells is adjusted in real time to synchronize the global layout size of the hexagonal layout with the zoom range to maintain content readability. If it is detected that the user operation is a drag operation on the first visual element, cells of the second visual element adjacent to the operated first visual element are reorganized.
8. The method for optimizing the page structure of a presentation according to claim 1, wherein: The method utilizes a ray tracing algorithm to calculate a visual focus heat map in a presentation page, and allocates three-dimensional voxels corresponding to target visual elements to visually sensitive areas to obtain first layout information, including: A ray bundle is emitted from the virtual viewpoint of the presentation page, penetrating each voxel in the three-dimensional implicit field, and the spatial distribution density of the visual elements is determined by solving the intersection equation of the ray and the scene geometry; Assigning high refractive index properties to the target visual element to enhance the deflection intensity of light on the surface of the target visual element in the three-dimensional implicit field, simulating the focusing effect of the human eye lens; Differentiated material parameters in the 3D implicit field are defined based on the type of visual element; the title uses a high specular reflection coefficient, and the chart sets a diffuse reflection coefficient to simulate information-carrying characteristics; An ambient occlusion algorithm is introduced to calculate the visual occlusion relationship between visual elements in the 3D implicit field, enhancing the stereoscopic perception of the focus area. The Monte Carlo ray sampling method is used to count the number of ray hits per unit area in the three-dimensional implicit field and generate a visual sensitivity distribution map. In the visual sensitivity distribution map, Gaussian blur processing is performed on the areas where the hit rate exceeds the hit rate threshold to form a focus heat map with a smooth transition, in which the heat value of the core focus area reaches the peak; Based on the direction of thermal gradient change in the focus heat map, the 3D voxels of the target visual element are offset along the line of sight by a preset multiple of the field of view angle radius to ensure that the high-priority target visual element is located in the core focal area with optimal retinal imaging; Non-maximum suppression is performed on the visual elements corresponding to the overlapping voxels, and the visual element layout scheme with the highest information density is retained to obtain the first layout information.
9. The method for optimizing the page structure of a presentation according to claim 1, wherein: The quantum annealing algorithm is used to predict the global optimal layout of the presentation page based on the positional relationship between each visual element in the visual structure information and the first layout information, combined with the information density energy function, to obtain the second layout information, including: Map each visual element into a quantum bit system, and combine the first layout information to encode the positional relationship between each visual element into the spin variable of the Ising model based on the element correlation; By analyzing the spatial correlation of visual elements at adjacent distances and visual levels, a coupling strength matrix between visual elements is dynamically generated and loaded into the quantum annealing machine. The quantum annealing evolution process is directly driven by the Ising model and a quantum annealing machine. The initial Hamiltonian used to represent the random distribution of visual elements is transformed into a final Hamiltonian that complies with the encoding layout constraints through adiabatic evolution. The encoding layout constraints are dynamically configured based on the page optimization requirements of the presentation page. During the annealing process, the dual objective functions of information density and aesthetic score are integrated, and the variational quantum eigensolver is used to generate the Pareto optimal solution set of the information density energy function under the Hamiltonian evolution process. Decode the quantum state information in the Pareto optimal solution set that follows the encoding layout constraints into position coordinate parameters of visual elements; In combination with the device adaptation rule, a hierarchical layout configuration of the visual elements is generated based on the position coordinate parameters, and the hierarchical layout configuration is used as the second layout information.
10. A page structure optimization system for presentations, characterized in that: The system comprises: The acquisition module is used to obtain the presentation page input by the user; A recognition module is used to identify visual elements in a presentation page and perform logical relationship analysis on the visual elements to obtain visual structure information corresponding to the presentation page; The parsing module is used to parse the logical relationships between visual elements and obtain the logical relationship information between visual elements in the presentation page; The first optimization module is used to map the visual structure information into the 3D implicit field of the page based on NeRF to obtain the 3D voxels corresponding to the visual elements. The ray tracing algorithm is used to calculate the visual focus heat map of the presentation page and assign the 3D voxels corresponding to the target visual elements to the visual sensitive area to obtain the first layout information. a second optimization module, configured to use a quantum annealing algorithm to predict a global optimal layout of the presentation page based on the positional relationship between the visual elements in the visual structure information and the first layout information in combination with an information density energy function, thereby obtaining second layout information; a third optimization module configured to predict the narrative path in the presentation based on the logical relationship information and the user's narrative preferences marked based on the user's historical data, to obtain a probability distribution of the page's narrative path; and optimize the second layout information based on the page's narrative path probability distribution to obtain third layout information, so that the spatial arrangement of visual elements in the presentation page conforms to the narrative logic preferred by the user; An execution module is used to perform secondary layout of the presentation page based on the third layout information to optimize the page structure of the presentation page.
Citation Information
Patent Citations
Method for generating three-dimensional presentation file
CN113129436A
Powerpoint automatic generation method and device, electronic equipment and readable storage medium
CN119540403A