Immersive Interactive Display Method for Three-Fold Screen Space Architectural 3D Models

CN122569872APending Publication Date: 2026-08-14SHANGHAI BUILDING DECORATION ENG GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

左屏与中屏、右屏与中屏相接的转折处,模型表面在相邻屏幕上的投影会产生位置跳变,观察者在移动视角时会感知到明显的断裂缝隙,破坏了建筑空间的整体感

Benefits of technology

在夹角过渡带内提取左屏模型或右屏模型相对于中屏模型的深度错位点云,依据该深度错位点云生成带有非均匀缝合权重的过渡网格序列。非均匀缝合权重沿夹角过渡带从中屏模型向左屏模型或右屏模型方向逐渐减小,即靠近中屏的顶点保留更多中屏模型的位置信息,靠近侧屏的顶点保留更多侧屏模型的位置信息。按照过渡网格序列对左屏模型与中屏模型、右屏模型与中屏模型执行视口缝合,边界顶点坐标根据非均匀缝合权重进行加权平均,在夹角转折处形成平滑过渡的连续三折幕拼接模型。该缝合方式使转折区域的网格曲面从侧屏向中屏逐步过渡,避免了硬性分割带来的深度跳变,观察者在三折幕空间内移动观看时,转折处的建筑模型表面呈现自然连续的视觉形态。实时追踪观察者在三折幕空间内的头部旋转向量,根据头部旋转向量动态调整连续三折幕拼接模型在左屏、中屏和右屏上的切片渲染视场角。头部旋转向量中的水平偏航角映射为切片窗口在连续三折幕拼接模型水平方向上的起始切割列和终止切割列,垂直俯仰角映射为垂直方向上的起始切割行和终止切割行,仅渲染落入该矩形窗口内的网格面片。当观察者头部转向侧屏时,侧屏的渲染区域相应偏移,显示观察者期望看到的模型侧面内容;当头部回正时,渲染区域归位至原始视场角。视场角随头部运动同步变化,使三折幕各屏幕的显示内容始终与观察者的注视方向匹配,增强了空间沉浸感和观看真实感。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122569872A_ABST
    Figure CN122569872A_ABST
Patent Text Reader

Abstract

This invention discloses an immersive interactive display method for a three-dimensional architectural model with a three-fold screen, belonging to the field of three-dimensional model display technology. The method includes acquiring multi-source survey data of the target building and constructing an initial three-dimensional mesh model; decomposing the initial three-dimensional mesh model into a left-screen model, a middle-screen model, and a right-screen model based on the folding angle parameters of the three-fold screen space, with the left and right screen models forming angled transition zones with the middle screen model; extracting depth misalignment point clouds within the angled transition zones to generate a transition mesh sequence with non-uniform stitching weights; stitching the left and right screen models with the middle screen model to obtain a continuous three-fold screen spliced ​​model; and dynamically adjusting the slice rendering field angle of the continuous three-fold screen spliced ​​model on the left, middle, and right screens by tracking the observer's head rotation vector in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of 3D model display technology, specifically to an immersive interactive display method for 3D models of three-fold spatial architecture. Background Technology

[0002] In the design review and public presentation of large-scale architectural projects, immersive spaces formed by multi-screen splicing to present 3D models can provide observers with a truly immersive visual experience. Triple-screen spaces, consisting of left, center, and right projection screens spliced ​​together at a certain angle, are a widely used immersive display environment. To address the need for displaying 3D architectural models in such spaces, existing methods typically divide the complete 3D model directly into three independent parts, projecting each part onto one of the three screens. However, existing model segmentation and display schemes have significant flaws. Firstly, model segmentation only involves rigid division at the geometric level according to screen boundaries, failing to consider the parallax and depth discontinuities that exist in the angled areas of a triple-screen display. At the transitions between the left and center screens, and between the right and center screens, the projection of the model surface onto adjacent screens will exhibit positional jumps, causing noticeable gaps that the observer will perceive when moving their viewpoint, disrupting the overall sense of the architectural space. Secondly, most existing solutions employ fixed rendering field angles, meaning the displayed content remains unchanged regardless of how the observer's head rotates. This method does not conform to the perspective rules of human eyes when observing real space. When the observer turns their head towards the side screen, the side screen still displays the content according to the preset perspective, resulting in a discrepancy between vision and physical sensation, which reduces the realism of the immersive experience.

[0003] Therefore, how to eliminate the sense of model breakage at the corner of the three-fold screen, and how to make the rendering field of view change in real time with the observer's head movement, are problems that urgently need to be solved in this field. Summary of the Invention

[0004] The purpose of this invention is to provide an immersive interactive display method for three-dimensional architectural models with a three-fold screen. During the model splicing and stitching process, the method eliminates depth misalignment and visual breaks caused by screen folds, and dynamically adjusts the rendering field of view of each screen according to the observer's head movement, thereby achieving a coherent and consistent immersive interactive display of three-fold architectural models.

[0005] To achieve the above objectives, the present invention provides the following technical solution: The present invention provides an immersive interactive display method for a three-dimensional architectural model of a three-fold screen space, comprising: acquiring multi-source surveying data of the target building and constructing an initial three-dimensional mesh model; decomposing the initial three-dimensional mesh model into a left screen model, a middle screen model, and a right screen model according to the folding angle parameters of the three-fold screen space, wherein the left screen model and the right screen model form an angled transition zone with the middle screen model respectively; extracting the depth misalignment point cloud of the left screen model or the right screen model relative to the middle screen model within the angled transition zone, wherein the depth misalignment point cloud is used to define the left screen. The overlapping area of ​​the projection of the model or right screen model to the middle screen model; a transition mesh sequence with non-uniform stitching weight is generated based on the depth misalignment point cloud, and the non-uniform stitching weight gradually decreases from the middle screen model to the left screen model or right screen model along the angled transition zone; the left screen model and right screen model are stitched with the middle screen model according to the transition mesh sequence to obtain a continuous three-fold screen splicing model; the head rotation vector of the observer in the three-fold screen space is tracked in real time, and the slice rendering field angle of the continuous three-fold screen splicing model on the left screen, middle screen and right screen is dynamically adjusted according to the head rotation vector.

[0006] This invention constructs an initial model through the fusion of multi-source mapping data, giving the model both a detailed geometric structure and rich surface attribute information, providing a high-quality data foundation for subsequent disassembly and assembly. The model is disassembled based on the folding angle parameters of the three-fold screen space, ensuring that each screen model precisely corresponds to the projection area of ​​the physical screen, avoiding content misalignment and distortion caused by direct projection.

[0007] As a technical solution of this invention, when constructing the initial 3D mesh model, the image point cloud acquired by UAV oblique photography and the structural point cloud scanned by ground-based lidar are spatially aligned. Using the fused point cloud as input, a Poisson surface reconstruction algorithm is employed to generate an initial 3D mesh model with normal information. Each vertex stores the texture color value of the original image and the reflection intensity value of the structural point cloud. This method achieves seamless fusion of aerial and ground-based point clouds, resulting in a model surface that is continuous and possesses accurate geometric foundations and rich visual and physical properties.

[0008] When disassembling the model, the first included angle value between the left and middle screens and the second included angle value between the middle and right screens are read. Using the normal direction of the middle screen as a reference, a projection cutting plane for the left screen is drawn on the horizontal plane according to the first included angle value, and a projection cutting plane for the right screen is drawn on the horizontal plane according to the second included angle value. Boolean segmentation operations are then performed on the initial 3D mesh model using the left and right screen projection cutting planes to obtain the left, middle, and right screen models that precisely correspond to the three-fold screen space. This disassembly method ensures strict consistency between the model division and the physical screen boundaries in spatial geometry.

[0009] To address the potential geometric discontinuities within the transition zone, this invention traverses the set of boundary vertices in adjacent regions, calculates the first depth offset from the left-screen model boundary vertex to the corresponding projection position of the middle-screen model, and the second depth offset from the right-screen model boundary vertex to the corresponding projection position of the middle-screen model. Boundary vertices exceeding a preset depth threshold are extracted as depth misalignment point clouds. The preset depth threshold is dynamically calculated based on the product of the sine of the bend angle and the average side length of the model's surface. This process automatically identifies depth break regions caused by viewpoint changes and locates the projection overlap areas that require focused processing during the stitching process.

[0010] When generating the transition mesh sequence, vertices in the depth misalignment point cloud are sorted according to their distance from the boundary line. A non-uniform stitching weight sequence, from largest to smallest, is then assigned to the sorted vertices. After binding the stitching weights to the vertices, they are inserted at the corresponding model boundaries to form the transition mesh sequence. The non-uniform stitching weights decrease from high to low values ​​along an exponential decay curve in the direction perpendicular to the boundary line, ensuring that vertices closer to the center screen retain more center screen coordinates, while vertices closer to the left or right screen retain more side screen coordinates, achieving a smooth spatial transition.

[0011] During viewport stitching, the stitching weight of each vertex in the transition mesh sequence is read. Based on this weight, the spatial coordinates of the boundary vertices of the side screen model and the corresponding boundary vertices of the middle screen model are weighted and averaged to obtain the stitched vertex coordinates, which replace the original boundary vertex coordinates. After stitching, Laplacian smoothing is performed on the common boundaries. This stitching method eliminates geometric jumps at the corner joints, ensuring the visual continuity and smoothness of the three-fold screen splicing model. At the same time, the smoothing process further eliminates the minor deformations introduced by the stitching, improving the smoothness of the model boundaries.

[0012] For observer head tracking, multiple infrared cameras deployed at the apex of the triple-screen space capture depth images of the observer's head. Quaternion representations of the head orientation are extracted and converted into Euler angle sequences. The horizontal yaw and vertical pitch angles are then separated and combined to form a head rotation vector. When dynamically adjusting the slice rendering field of view based on this head rotation vector, the horizontal yaw angle is mapped to a proportional coefficient of the total horizontal span of the three screens to determine the horizontal start and end positions of the slice window. The vertical pitch angle is mapped to the vertical start and end positions of the model. Frustum clipping is performed within the area defined by this rectangular window, rendering only the mesh surfaces within the window. This slice rendering method, which changes in real-time with head posture, ensures that the observer's perspective remains dynamically consistent with the content presented on the triple-screen, greatly enhancing the immersive experience.

[0013] Preferably, during the movement of the slice window, when the horizontal yaw angle increases and the window moves beyond the left screen boundary, the excess portion is cyclically shifted to the far edge of the right screen; when the horizontal yaw angle decreases and the window moves beyond the right screen boundary, the excess portion is cyclically shifted to the far edge of the left screen, and a linear interpolation transition band is introduced at the shift boundary. The cyclic shift mechanism, combined with the interpolation transition band, effectively avoids abrupt changes in content at the edge of the field of view, ensuring a continuous and smooth transition of the image when the observer turns their head in any direction, maintaining a complete sense of spatial enclosure.

[0014] As a further improvement of this invention, after obtaining the head rotation vector, it is input into a Kalman filter. The head rotation angular velocity and angular acceleration are used as state variables, and the original rotation vector obtained through real-time tracking is used as the observation. The filtered head rotation vector is then output to replace the original value in adjusting the slice rendering field of view. The filtering process suppresses high-frequency oscillations in the rendering window caused by minor head jitters, making the field of view adjustment more stable and smooth, and reducing visual fatigue for the observer.

[0015] The technical effects and advantages provided by the present invention in the above technical solution are as follows: Within the transition zone at the angle, the depth misalignment point cloud of the left or right screen model relative to the middle screen model is extracted. Based on this depth misalignment point cloud, a transition mesh sequence with non-uniform stitching weights is generated. The non-uniform stitching weights gradually decrease along the transition zone from the middle screen model towards the left or right screen model; that is, vertices closer to the middle screen retain more positional information of the middle screen model, and vertices closer to the side screen retain more positional information of the side screen model. Viewport stitching is performed on the left and middle screen models, and the right and middle screen models, according to the transition mesh sequence. The boundary vertex coordinates are weighted and averaged according to the non-uniform stitching weights, forming a smooth transition continuous three-fold screen spliced ​​model at the angle transition. This stitching method allows the mesh surface in the transition area to gradually transition from the side screen to the middle screen, avoiding the depth jump caused by hard segmentation. When the observer moves and views within the three-fold screen space, the surface of the architectural model at the transition point presents a natural and continuous visual form. The observer's head rotation vector within the three-fold screen space is tracked in real time, and the slice rendering field angle of the continuous three-fold screen spliced ​​model on the left, middle, and right screens is dynamically adjusted based on the head rotation vector. The horizontal yaw angle in the head rotation vector is mapped to the starting and ending cut columns of the slice window in the horizontal direction of the continuous three-panel spliced ​​model, and the vertical pitch angle is mapped to the starting and ending cut rows in the vertical direction. Only the mesh faces falling within this rectangular window are rendered. When the observer's head turns to the side screen, the rendering area of ​​the side screen shifts accordingly to display the side content of the model that the observer expects to see; when the head returns to center, the rendering area returns to the original field of view. The field of view changes synchronously with the head movement, ensuring that the displayed content of each screen in the three-panel structure always matches the observer's gaze direction, enhancing the sense of spatial immersion and viewing realism. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0017] Figure 1 This is a flowchart of a method for immersive interactive display of three-dimensional architectural models with a three-fold screen. Figure 2 This is a flowchart of the multi-source point cloud fusion and three-fold projection model decomposition process; Figure 3 This is a flowchart of the cloud extraction process for depth misalignment points within the transition zone of the angle. Figure 4 This is a flowchart of non-uniform stitching weighted transition mesh generation and viewport stitching based on deep misaligned point clouds; Figure 5 This is a flowchart of the dynamic adjustment process of the field of view in a three-fold screen slice rendering based on the head rotation vector; Figure 6 This is a flowchart of head rotation vector Kalman filtering and field of view cyclic shift adjustment; Figure 7 This is a statistical diagram of the distribution of face edge lengths in the initial 3D mesh model; Figure 8 This is a schematic diagram of the depth offset distribution and preset depth threshold; Figure 9 This is a comparison curve of the head rotation horizontal yaw angle before and after filtering; Figure 10 This is the head horizontal yaw rate curve after Kalman filtering smoothing. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] See Figure 1This invention provides an immersive interactive display method for a three-dimensional architectural model in a three-fold screen space, comprising: acquiring multi-source surveying data of the target building and constructing an initial three-dimensional mesh model; decomposing the initial three-dimensional mesh model into a left screen model, a middle screen model, and a right screen model according to the folding angle parameters of the three-fold screen space, with the left screen model and the right screen model forming an angled transition zone with the middle screen model respectively; extracting the depth misalignment point cloud of the left screen model or the right screen model relative to the middle screen model within the angled transition zone, the depth misalignment point cloud being used to define the projection overlap area of ​​the left screen model or the right screen model onto the middle screen model; generating a transition mesh sequence with non-uniform stitching weights based on the depth misalignment point clouds, the non-uniform stitching weights gradually decreasing along the angled transition zone from the middle screen model to the left screen model or the right screen model; stitching the left screen model and the right screen model with the middle screen model according to the transition mesh sequence to obtain a continuous three-fold screen splicing model; and tracking the head rotation vector of the observer in the three-fold screen space in real time, and dynamically adjusting the slice rendering field angle of the continuous three-fold screen splicing model on the left screen, middle screen, and right screen according to the head rotation vector. Multi-source mapping data can be acquired through UAV oblique photography and ground-based lidar scanning, and after fusion, an initial 3D mesh model is generated; the folding angle parameter determines the projection cutting surface to achieve model partitioning; the depth misalignment point cloud is extracted based on the depth difference of the boundary vertices to form a transition mesh sequence and the boundary is weighted and stitched, and then the head rotation vector drives the multi-screen rendering area to match the observer's view in real time.

[0020] Example 1: In specific implementation, please refer to Figure 2 The process of acquiring multi-source mapping data of the target building and constructing an initial 3D mesh model includes aligning the spatial coordinates of the image point cloud acquired by UAV oblique photography and the structural point cloud scanned by ground-based lidar. One method for spatial coordinate alignment is to extract 3D scale-invariant feature transformation key points from both the image point cloud and the structural point cloud, and use a fast point feature histogram descriptor to perform feature matching on the two sets of key points to obtain an initial correspondence set. A random sampling consensus algorithm is then used to estimate the rigid transformation matrix from the image point cloud coordinate system to the structural point cloud coordinate system from the initial correspondence set. The rigid transformation matrix includes rotation and translation components. Using the structural point cloud coordinate system as a reference coordinate system, all points in the image point cloud are transformed to the reference coordinate system through the rigid transformation matrix, completing the spatial coordinate alignment and obtaining the aligned fused point cloud.

[0021] Using the aligned and fused point cloud as input, an initial 3D mesh model with normal information is generated using the Poisson surface reconstruction algorithm. During the execution of the Poisson surface reconstruction algorithm, the maximum depth parameter of the octree is set. The maximum depth parameter of the octree is determined according to the point cloud density of the aligned and fused point cloud. The maximum depth parameter of the octree is set to the maximum value that makes the number of points contained in each leaf node fall within a predetermined interval [10,30]. In the initial 3D mesh model output by the Poisson surface reconstruction algorithm, the information storage structure of each vertex includes a 3D coordinate field, a normal vector field, a texture color value field, and a reflection intensity value field. Among them, the 3D coordinate field stores the position of the vertex in the reference coordinate system, the normal vector field stores the normal direction at the vertex calculated by the Poisson surface reconstruction algorithm, the texture color value field stores the color value of the nearest point projected from the image point cloud to the vertex, and the reflection intensity value field stores the reflection intensity value of the nearest point projected from the structure point cloud to the vertex.

[0022] The process of decomposing the initial 3D mesh model into left-screen, middle-screen, and right-screen models based on the folding angle parameters of the three-fold screen space includes: obtaining the first included angle value between the left and middle screens and the second included angle value between the middle and right screens in the three-fold screen space by using structural measurement equipment deployed within the three-fold screen space or reading a pre-calibrated configuration file. Based on the coordinate system of the middle screen's display plane, the normal direction of the middle screen is determined; the normal direction of the middle screen is the vertical direction from the center point of the middle screen's display plane to the observation area. On the horizontal plane, with the center or boundary of the middle screen as the reference origin, the normal direction of the middle screen is rotated to the left according to the first included angle value to obtain the normal direction of the left screen's projection cutting surface; the normal direction of the middle screen is rotated to the right according to the second included angle value to obtain the normal direction of the right screen's projection cutting surface. The position of the left screen's projection cutting surface on the horizontal plane is aligned with the actual boundary plane of the left screen, and the position of the right screen's projection cutting surface on the horizontal plane is aligned with the actual boundary plane of the right screen. Boolean segmentation is performed on the initial 3D mesh model using the left and right screen projection cut surfaces. This segmentation divides the triangular faces in the initial 3D mesh model into three parts: the set of triangular faces completely located to the left of the left screen projection cut surface forms the left screen projection area mesh; the set of triangular faces completely located between the left and right screen projection cut surfaces forms the middle screen projection area mesh; and the set of triangular faces completely located to the right of the right screen projection cut surface forms the right screen projection area mesh. For triangular faces that cross the projection cut surfaces, the Boolean segmentation operation inserts new vertices at the edges where the cut surfaces intersect with them and re-triangulates the faces, dividing them into sub-triangular faces located within their respective projection areas. After segmentation, the left screen projection area mesh serves as the left screen model, the middle screen projection area mesh serves as the middle screen model, and the right screen projection area mesh serves as the right screen model.

[0023] In some embodiments, the positions of the left and right screen projection cut surfaces are fine-tuned so that the distance between them is equal to the projection length of the actual display width of the middle screen onto the normal direction of the middle screen, the angle between the left and right screen projection cut surfaces is equal to a first angle value, and the angle between the right and right screen projection cut surfaces is equal to a second angle value. The left, middle, and right screen models generated by the Boolean segmentation operation maintain connectivity, forming angled transition zones at the boundaries between adjacent left and middle screen models and between adjacent right and middle screen models. These angled transition zones are strip-shaped regions extending a certain width along both sides of the boundary line near the intersection of the two models.

[0024] See Figure 7 The figure shows a histogram of the distribution of all face edge lengths in the initial 3D mesh model. The horizontal axis represents face edge length in units of length, ranging from approximately 1 to 6.5 units. The vertical axis represents the frequency of edges within the corresponding face edge length range. The histogram uses diagonal lines to represent the frequency distribution of face edge lengths, showing an overall approximate normal distribution trend. The face edge lengths with the highest frequency are concentrated in the range of 3.5 to 4.0 units of length. The average edge length is indicated by a black dashed line in the figure. Its location is approximately 3.6 length units, marking the arithmetic mean length of all face edges in the initial 3D mesh model.

[0025] This edge length distribution map corresponds to the mesh edge length characteristics of each leaf node when the Poisson surface reconstruction algorithm generates an initial 3D mesh model from the fused point cloud. This is under the condition that the number of points in each leaf node falls within the predetermined interval [10, 30], reflecting the detailed scale of mesh division within the model. Through this statistical distribution, a preset depth threshold can be dynamically calculated. Its expression is ,in, For the corresponding fold angle in the three-fold screen space, This refers to the average face edge length shown in the figure. This threshold is used to extract the depth offset cloud to determine whether the depth offset of the boundary vertices exceeds the threshold, thereby defining the projection overlap area.

[0026] In summary, the diagram illustrates the statistical characteristics of face edge lengths in the initial 3D mesh model, including the average edge length. The acquisition of these parameters provides the foundation for subsequent dynamic threshold calculations and transition mesh stitching processes, ensuring the fine division of the three-dimensional model of the three-fold curtain space building and the reasonable allocation of stitching weights.

[0027] Example 2: In specific implementation, please refer to Figure 3The process of extracting the depth misalignment point cloud of the left or right screen model relative to the middle screen model within the transition zone includes: obtaining the set of boundary vertices adjacent to the left and middle screen models, and the set of boundary vertices adjacent to the right and middle screen models. The set of boundary vertices adjacent to the left and middle screen models is determined as follows: traversing all vertices of the left screen model, selecting vertices that lie on the left screen projection cut surface during the Boolean segmentation operation, and traversing all vertices of the middle screen model, selecting vertices that lie on the left screen projection cut surface during the Boolean segmentation operation; merging these two sets of vertices and removing duplicate vertices to form the set of boundary vertices adjacent to the left and middle screen models. The set of boundary vertices adjacent to the right and middle screen models is determined in the same way: traversing all vertices of the right screen model, selecting vertices that lie on the right screen projection cut surface during the Boolean segmentation operation, and traversing all vertices of the middle screen model, selecting vertices that lie on the right screen projection cut surface during the Boolean segmentation operation; merging these two sets of vertices and removing duplicate vertices to form the set of boundary vertices adjacent to the right and middle screen models.

[0028] Within the transition zone, each boundary vertex of the left screen model adjacent to the middle screen model is traversed, as are each boundary vertex of the right screen model adjacent to the middle screen model. For each traversed left screen model boundary vertex, the first depth offset from the left screen model boundary vertex to the corresponding projection position on the middle screen model is calculated. The projection position from the left screen model boundary vertex to the corresponding projection position on the middle screen model is determined by the following steps: starting from the left screen model boundary vertex, a ray is emitted along the inward normal direction of the middle screen model display plane, and the intersection point of the ray with the triangular facet of the middle screen model surface is solved; if an intersection point exists, the intersection point closest to the left screen model boundary vertex is taken as the corresponding projection position of the left screen model boundary vertex on the middle screen model; if the ray does not intersect the middle screen model surface, the position of the left screen model boundary vertex to the nearest point on the middle screen model surface is taken as the corresponding projection position. The first depth offset is defined as the Euclidean distance from the spatial coordinates of the left screen model boundary vertex to the spatial coordinates of the corresponding projection position. For each right screen model boundary vertex encountered, following the same method as the left screen model boundary vertex, a ray is emitted from the right screen model boundary vertex along the inner normal direction of the middle screen model display plane to determine the corresponding projection position of the right screen model boundary vertex on the middle screen model, and the Euclidean distance from the right screen model boundary vertex to the corresponding projection position is used as the second depth offset.

[0029] The left screen model boundary vertices when the first depth offset exceeds the preset depth threshold are extracted as left screen depth misalignment point clouds. The right screen model boundary vertices when the second depth offset exceeds the preset depth threshold are extracted as right screen depth misalignment point clouds. The preset depth threshold is dynamically calculated based on the product of the sine of the fold angle between the left and middle screens or the sine of the fold angle between the middle and right screens in the three-fold screen space and the average side length of all face edges in the initial 3D mesh model. The expression is:

[0030] in, This indicates the preset depth threshold, expressed in length units. This indicates the corresponding angle of inversion when extracting the depth misalignment point cloud of the left screen. When extracting the depth misalignment point cloud of the right screen, take the first included angle value between the left and middle screens in the three-fold screen space. The second included angle between the central screen and the right screen in the three-fold screen space is taken. The first and second included angles are determined by the physical structure of the three-fold screen space and range from 90° to 180°. Indicates the angle of the bend. The value of the sine function; This represents the average edge length of all faces in the initial 3D mesh model, expressed in length units.

[0031] Average edge length of all face edges in the initial 3D mesh model The calculation method is as follows: obtain all triangular faces of the initial 3D mesh model, extract the three edges of all triangular faces to form an edge set, and the edge set contains... Edge; for the first edge in the edge set A strip of edge, whose two endpoints are respectively and Calculate the first Length of strip Calculate the arithmetic mean of the lengths of all sides to obtain the average side length of all faces. Preset depth threshold The depth offset point cloud is automatically determined based on the current fold angle and the average side length of the initial 3D mesh model during each calculation, eliminating the need for manual parameter setting. The left-screen depth offset point cloud and the right-screen depth offset point cloud are used to define the projection overlap area from the left-screen model to the middle-screen model and the projection overlap area from the right-screen model to the middle-screen model, respectively.

[0032] See Figure 8The figure shows a histogram of the frequency distribution of depth offsets in Example 2. The horizontal axis represents the depth offset in length units, ranging from 0 to 15, and the vertical axis represents the frequency of the corresponding depth offset. Shaded bars are used to represent the distribution of depth offsets, showing an overall right-skewed, long-tailed distribution trend. The position of the preset depth threshold δ is also marked with a dashed line, with the threshold being approximately 2.0 length units.

[0033] As shown in the figure, the vast majority of depth offsets are concentrated in the range of 0 to 2, with a high frequency. This indicates that the depth offsets of the boundary vertices of the left or right screen model relative to the projection position of the middle screen model are mostly small, meeting the constraints of the preset depth threshold. After the depth offset exceeds the preset threshold δ (approximately 2.0), the frequency decreases significantly, indicating that there are fewer depth misalignment points exceeding the threshold. However, a certain number of vertices still exist, with their depth offset distribution extending over a large range from 2 to 15, reflecting the structural characteristics of significant depth misalignment at the model boundary.

[0034] This figure visually demonstrates the constraint effect of the preset depth threshold δ on the selection of depth misalignment point clouds, supporting the technical solution in Example 2 for extracting left and right screen depth misalignment point clouds based on depth offset and dynamically calculated threshold δ. Detailed statistics on frequency distribution provide data support for generating the transition mesh sequence with non-uniform stitching weights in Example 3, effectively ensuring the accuracy of the stitching process and the smoothness of the model transition zone.

[0035] Example 3: In specific implementation, please refer to Figure 4The process of generating a transition mesh sequence with non-uniform stitching weights based on the depth misalignment point cloud includes: sorting each vertex in the left screen depth misalignment point cloud and sorting each vertex in the right screen depth misalignment point cloud. When sorting each vertex in the left screen depth misalignment point cloud, the vertical distance from each vertex to the boundary line between the left and middle screens is calculated. The boundary line between the left and middle screens is defined as a continuous polyline formed by the set of boundary segments shared by the left and middle screen models in the Boolean partitioning operation. Based on the calculated vertical distance values, all vertices in the left screen depth misalignment point cloud are sorted in ascending order of vertical distance, with the vertex closest to the boundary line at the beginning of the sequence and the vertex farthest from the boundary line at the end. After sorting, the left screen vertex sequence is obtained. When sorting each vertex in the right screen depth misalignment point cloud, the vertical distance from each vertex in the right screen depth misalignment point cloud to the boundary line between the right screen and the middle screen is calculated. The boundary line between the right screen and the middle screen is defined as a continuous polyline formed by the set of boundary line segments shared by the right screen model and the middle screen model in the Boolean partitioning operation. Based on the calculated vertical distance values, all vertices in the right screen depth misalignment point cloud are sorted in ascending order according to the vertical distance from smallest to largest. The vertex closest to the boundary line between the right screen and the middle screen is placed at the beginning of the sequence, and the vertex farthest from the boundary line is placed at the end of the sequence. After sorting, the right screen vertex sequence is obtained.

[0036] Assign a first non-uniform stitching weight sequence to the vertices in the left screen vertex sequence in descending order: Suppose the left screen vertex sequence contains a total of The vertex in the left screen vertex sequence. The vertices are numbered as follows: , The value is ,in, The vertex that is closest to the boundary between the left and middle screens. The vertex that is furthest from the boundary between the left and middle screens; the vertex in the left screen vertex sequence. The first non-uniform stitching weight assigned to each vertex Determined according to the exponential decay function:

[0037] in, Represents the vertex sequence of the left screen. The first non-uniform stitching weights corresponding to each vertex. The range of values ​​is within Between 1 and 1; The base of the natural logarithm is 2.71828. Indicates the attenuation coefficient. Based on the maximum vertical distance from the vertex of the cloud at the left screen depth misalignment point to the boundary line between the left and middle screens. Confirmed, the specific confirmation method is as follows: Maximum vertical distance It is obtained by traversing all vertices in the depth misalignment point cloud of the left screen and calculating the vertical distance of each vertex to the boundary line between the left and middle screens, and then taking the maximum value. Indicates the index of the vertex in the left screen vertex sequence; This represents the total number of vertices in the left screen vertex sequence. The first non-uniform stitching weight sequence decreases from 1 to 0 along the direction perpendicular to the boundary between the left and middle screens, from the middle screen model towards the left screen model, following an exponential decay curve.

[0038] Assign a second non-uniform stitching weight sequence to the vertices in the right-screen vertex sequence in descending order: Suppose the right-screen vertex sequence contains a total of The vertex in the right screen vertex sequence. The vertices are numbered as follows: , The value is ,in, The vertex that is closest to the boundary line between the right and middle screens. The vertex that is furthest from the boundary between the right and middle screens; the vertex in the right screen vertex sequence. The second non-uniform stitching weight is assigned to each vertex. Determined according to the exponential decay function The meaning of each parameter corresponds to that of the left screen side, including the attenuation coefficient. Based on the maximum vertical distance from the vertex of the cloud at the right screen depth misalignment point to the boundary line between the right and middle screens. Sure, The second non-uniform stitching weight sequence decreases from 1 to 0 along the direction perpendicular to the boundary between the right and middle screens, from the middle screen model to the right screen model, following an exponential decay curve.

[0039] Each weight value in the first non-uniform stitching weight sequence is bound to the corresponding vertex in the left screen vertex sequence. The binding operation involves adding a weight field to the data structure of each vertex in the left screen vertex sequence, and then calculating the weight value... Write to the The weight field of each vertex. After binding, all vertices in the left screen vertex sequence carrying the first non-uniform stitching weight are inserted into the boundary between the left and middle screen models according to their spatial positions near the boundary line between the left and middle screens, forming a left transition mesh sequence. Each vertex in the left transition mesh sequence records both its original spatial coordinates in the left screen model and the bound first non-uniform stitching weight value. Each weight value in the second non-uniform stitching weight sequence is bound to the vertex at the corresponding index position in the right screen vertex sequence. All vertices in the bound right screen vertex sequence are then inserted into the boundary between the right and middle screen models, forming a right transition mesh sequence.

[0040] The process of viewport stitching of the left and right screen models with the middle screen model according to the transition mesh sequence includes: reading the first non-uniform stitching weight of each vertex in the left transition mesh sequence. For each vertex in the left transition mesh sequence, based on the read first non-uniform stitching weight, a weighted average is performed on the spatial coordinates of the boundary vertices of the left screen model and the corresponding spatial coordinates of the boundary vertices of the middle screen model to obtain the coordinates of the left stitched vertex. The expression for the weighted average operation is:

[0041] in, The calculated coordinates of the left seam vertex are represented by a three-dimensional vector. This represents the first non-uniform stitching weight bound to the current vertex in the left transition mesh sequence; This represents the original three-dimensional vector of the current vertex's spatial coordinates in the left-side screen model; This represents the three-dimensional spatial coordinates of the current vertex's projected position on the mid-screen model. The calculated left-stitch vertex coordinates replace the original spatial coordinates of the corresponding boundary vertices in the left-screen model, thus updating the left-screen model boundary vertex coordinates. The second non-uniform stitching weight of each vertex in the right transition mesh sequence is read, and the spatial coordinates of the right-screen model boundary vertices and the corresponding mid-screen model boundary vertices are weighted and averaged using the same weighted averaging method to obtain the right-stitch vertex coordinates. These right-stitch vertex coordinates then replace the original spatial coordinates of the corresponding boundary vertices in the right-screen model.

[0042] Laplacian smoothing is performed on the model boundaries after viewport stitching. The scope of Laplacian smoothing is the set of vertices within a predetermined width on both sides of the common boundary between the left and middle screen models, and the set of vertices within a predetermined width on both sides of the common boundary between the right and middle screen models. Laplacian smoothing uses an umbrella operator to calculate the Laplacian coordinates of each vertex. The umbrella operator uses the difference between the average spatial coordinates of all adjacent vertices of the target vertex and the spatial coordinates of the target vertex itself as the Laplacian coordinate vector. The target vertex is moved along the direction of the Laplacian coordinate vector by a preset smoothing factor. After repeating the process a predetermined number of times, the Laplacian smoothing is completed.

[0043] Example 4: In specific implementation, please refer to Figure 5 The process of real-time tracking of an observer's head rotation vector within a triple-screen space includes: capturing depth images of the observer's head using multiple infrared cameras deployed at the top corners of the triple-screen space. The infrared cameras are deployed as follows: one infrared depth sensor is installed at each of the left front top corner, right front top corner, and middle top corner of the triple-screen space. The three infrared depth sensors face the observer's standing area within the triple-screen space, and their field of view covers the entire possible position range of the observer within the triple-screen space. During each tracking iteration, the three infrared depth sensors simultaneously acquire one frame of depth image. Each pixel in the depth image stores the distance value from the corresponding spatial point to the imaging plane of the infrared depth sensor. The three depth images are then timestamped and merged into a set of multi-view depth image frames.

[0044] The process of extracting the quaternion representation of head orientation from depth images involves feeding multi-view depth image frames into a pre-trained head pose estimation convolutional neural network. The core architecture of the pre-trained head pose estimation convolutional neural network consists of a feature extraction backbone network and a pose regression branch network. The feature extraction backbone network adopts a ResNet-18 structure, removing the final global average pooling layer and fully connected layers, retaining the convolutional layer structure from the initial convolutional layer to the fourth residual block. The feature extraction backbone network receives a single-frame depth image as input, with a resolution of 224×224 pixels, single channel, and outputs a 7×7×512 spatial feature map. After the feature extraction backbone network, a pose regression branch network is connected. The pose regression branch network compresses the 7×7×512 spatial feature map into a 512-dimensional feature vector by a global average pooling layer, a fully connected layer maps the 512-dimensional feature vector to a 256-dimensional intermediate feature vector, a modified linear unit activation layer performs a nonlinear transformation on the 256-dimensional intermediate feature vector, and a fully connected layer maps the 256-dimensional intermediate feature vector to a 4-dimensional output vector. The 4-dimensional output vector corresponds to the four components of the head orientation quaternion. When training the pre-trained head pose estimation convolutional neural network, the BIWI deep head pose dataset was used as training data. The BIWI deep head pose dataset contains depth images of multiple people in different head poses and the corresponding ground truth head orientation quaternions. The loss function adopted was the quaternion angle loss, which is defined as the square of the spherical angle distance between the network's predicted quaternion and the ground truth quaternion. The optimizer was a stochastic gradient descent optimizer, with an initial learning rate of 0.01. Every 20 training epochs, the learning rate was decayed to 0.1 times the current learning rate, for a total of 60 training epochs. The batch size was set to 32. Each depth image from the multi-view depth image frames was independently input into the pre-trained head pose estimation convolutional neural network to obtain three candidate quaternion representations. The three candidate quaternion representations were then averaged using quaternion spherical linear interpolation, and the fused quaternion was used as the final head orientation quaternion representation.

[0045] The process of converting the fused head orientation quaternion representation into an Euler angle sequence includes: converting the head orientation quaternion representation... Convert to rotation matrix form, where, Let be the real part of the quaternion. The three imaginary components of the quaternion; extract the horizontal yaw angle from the rotation matrix. and vertical pitch angle Horizontal yaw angle The calculation method is as follows Vertical pitch angle The calculation method is as follows ,in, Represents the arctangent function in the four quadrants. Let represent the arcsine function. The horizontal yaw and vertical pitch angles are combined from the Euler angle sequence to form the head rotation vector, denoted as . .

[0046] The process of dynamically adjusting the slice rendering field of view of the continuous three-fold splicing model on the left, middle, and right screens based on the head rotation vector includes: mapping the horizontal yaw angle in the head rotation vector to a proportional coefficient of the total horizontal span of the left, middle, and right screens. The total horizontal span of the left, middle, and right screens is defined as the maximum scene width in the horizontal direction of the continuous three-fold splicing model after unfolding, expressed as the horizontal axis span of the scene coordinate system. express, The unit is scene length. The mapping method for the scaling factor is:

[0047] in, This represents the normalized positional scaling factor of the slice window in the horizontal direction of the continuous three-fold splicing model. The value range is from 0 to 1; This represents the horizontal yaw angle obtained through real-time tracking in the current frame, in degrees. The minimum physically reachable horizontal yaw angle for an observer within the three-fold screen space is determined through a calibration process. The calibration process involves having the observer face the leftmost edge of the left screen and recording the horizontal yaw angle value at that moment. This represents the maximum physically achievable horizontal yaw angle for an observer within the three-fold screen space. It is determined by having the observer face the rightmost boundary of the right screen and recording the horizontal yaw angle value at that moment.

[0048] Determine the starting and ending cutting columns of the slice window in the horizontal direction of the continuous three-fold screen splicing model based on the scaling factor: Set the horizontal width of the slice window. , The length of the horizontal scene range visible from the observer's current position; starting cut column. The calculation method is as follows Terminate the cutting column The calculation method is as follows ,in, This represents the total horizontal span of the continuous three-fold screen splicing model. The horizontal width of the slice window.

[0049] Map the vertical pitch angle in the head rotation vector to the start and end cut rows in the vertical direction of the continuous three-panel splicing model: Set the total vertical height of the continuous three-panel splicing model to... The vertical height of the slice window is The mapping method for the vertical pitch angle is as follows: ,in, This represents the normalized positional scaling factor of the slice window in the vertical direction of the continuous three-fold splicing model. The value range is from 0 to 1; This represents the vertical pitch angle obtained through real-time tracking in the current frame; The minimum physically achievable vertical pitch angle for an observer within the three-fold screen space is determined by having the observer tilt their head down to the maximum angle and recording the vertical pitch angle value. This represents the maximum physically achievable vertical pitch angle for an observer within the three-fold screen space. It is determined by having the observer tilt their head back to the maximum angle and recording the vertical pitch angle value. (Starting cut line) The calculation method is as follows Terminate cutting line The calculation method is as follows .

[0050] Start cutting column Termination of cutting column Starting cutting row and termination of cutting line The shared rectangular window serves as the viewport for slice rendering, with its four boundaries corresponding to the left boundary in the scene coordinate system. Right boundary Upper boundary Lower boundary Perform view frustum clipping on the visible mesh faces on the left, middle, and right screens respectively: Project each triangular facet in the continuous three-fold splicing model onto the horizontal and vertical planes of the scene coordinate system, and determine whether the projection of the triangular facet overlaps with the rectangular window of the slice rendering field of view; only the mesh faces that fall within the rectangular window are retained, and the retained mesh faces are sent to the rendering pipeline for drawing, while the mesh faces that fall outside the rectangular window are directly discarded in the rendering pipeline.

[0051] See Figure 9 In the graph, the horizontal axis represents time in seconds, and the vertical axis represents the horizontal yaw angle in degrees. The dashed curve marked with an "×" in the legend represents the original horizontal yaw angle obtained in real-time tracking through a head pose estimation convolutional neural network. The solid line curve represents the horizontal yaw angle after filtering by the Kalman filter. As shown in the figure, the original horizontal yaw angle curve exhibits high-frequency fluctuations of a certain amplitude throughout the time period, especially in the intervals of 0 to 10 seconds and 40 to 50 seconds, indicating strong noise interference. After Kalman filtering, the filtered horizontal yaw angle curve is significantly smoother, eliminating low-amplitude random jitter, and its overall trend remains highly consistent with the original curve, demonstrating the filter's effective smoothing and noise suppression capabilities for the head rotation vector.

[0052] The horizontal yaw angle in the figure starts at 0 degrees and gradually rises to a peak of approximately 60 degrees over time (about 20 to 30 seconds), before gradually decreasing back to near 0 degrees. This angle change corresponds to the observer's head rotation within the three-fold screen space, from facing the leftmost edge of the left screen to facing the rightmost edge of the right screen, and then back to the left screen. The smoothness of the filtered curve ensures the stability of the rendering field of view when dynamically adjusting the continuous three-fold screen splicing model slices based on the head rotation vector in Example 4, avoiding jumps and discontinuities in viewpoint switching caused by noise in the original data, and improving the smoothness of the overall immersive interactive experience.

[0053] In summary, this figure visually reflects the technical effect of filtering the head rotation vector obtained by real-time tracking using a Kalman filter in this embodiment. It verifies the effectiveness of the filter in suppressing head rotation angle estimation errors and smoothing rotation curves, providing accurate and reliable input data for the dynamic adjustment of the field of view in subsequent slice rendering.

[0054] Example 5: In specific implementation, please refer to Figure 6 After tracking the observer's head rotation vector in real time within the triple-screen space, the original head rotation vector is input into a Kalman filter to output a filtered head rotation vector. The state equation of the Kalman filter uses the head rotation angular velocity and head rotation angular acceleration as state variables. The time interval is defined. state vector for:

[0055] in, Indicates time The state vector contains four components; Indicates time The horizontal yaw rate, in radians per second; Indicates time The vertical pitch angular velocity, in radians per second; Indicates time The horizontal yaw acceleration, in radians per square second; Indicates time The vertical pitch acceleration, expressed in radians per square second.

[0056] The state equations adopt a constant-velocity rotation model combined with the random walk assumption of angular acceleration. The state transition equations are expressed as follows:

[0057] in, Indicates time The state vector; The state transition matrix, based on kinematic relationships, has the following specific form:

[0058] in, This represents the time interval between two tracking operations, determined by the frame rate of the infrared camera, and is taken as the difference in timestamps between two adjacent depth images. The process noise vector is represented by a vector with a mean of zero and a covariance matrix of... The normal distribution The diagonal elements are set according to the maximum angular acceleration range of the head movement. The noise variance of the horizontal yaw acceleration process is taken as one-tenth of the square of the maximum yaw acceleration amplitude, and the noise variance of the vertical pitch acceleration process is taken as one-tenth of the square of the maximum pitch acceleration amplitude.

[0059] The observation equation of the Kalman filter uses the original head rotation vector obtained from real-time tracking as the observation. The observation vector is defined as follows: for:

[0060] in, Indicates time The original horizontal yaw angle is obtained by using a convolutional neural network for head pose estimation; Indicates time The raw vertical pitch angle is obtained through a convolutional neural network that estimates head pose. The observation equation is expressed as:

[0061] in, The observation matrix represents the mapping of angular velocities in the state vector to the observed angle values ​​at the current moment. Based on the forward Euler integral relation, the specific form of the observation matrix is:

[0062] in, Let the observed noise vector have a mean of zero and a covariance matrix of... The normal distribution The angle estimation error of the head pose estimation convolutional neural network on the BIWI deep head pose dataset was determined by statistical analysis. The variance of the horizontal yaw angle observation noise and the variance of the vertical pitch angle observation noise were respectively taken as the average squared error values ​​of the corresponding angles in the dataset test.

[0063] The recursive process of a Kalman filter includes a prediction step and an update step. In the prediction step, based on the time... Filtered state vector and state transition matrix Calculation time Predicted state vector Simultaneously, the prediction error covariance matrix is ​​updated. In the update step, the Kalman gain matrix is ​​calculated, using the time step... observation vector The predicted state vector is corrected to obtain the time step. Filtered state vector From the filtered state vector The filtered horizontal yaw angle and filtered vertical pitch angle are extracted by multiplying the angle estimate from the previous moment by the filtered angular velocity at the current moment by the time interval to obtain the filtered horizontal yaw angle. and the filtered vertical pitch angle The filtered head rotation vector is used instead of the original head rotation vector during the adjustment of the slice rendering field of view.

[0064] The process of dynamically adjusting the slice rendering field of view of the continuous three-fold screen splicing model on the left, middle, and right screens based on the head rotation vector also includes cyclic shifting. When the observer's horizontal yaw angle increases, causing the slice window to move beyond the left screen boundary, the portion of the slice window that exceeds the left screen boundary is cyclically shifted to the far edge of the right screen. The left screen boundary corresponds to the leftmost coordinate of the continuous three-fold screen splicing model. The left boundary of the slice window is .when When the slice window moves beyond the left screen boundary, the width of the excess portion is... The rectangular area covered by this excess portion is located at the right coordinate of the continuous three-panel splicing model. Starting from point , the cut-off width is The rectangular area is spliced ​​to the left edge of the slice window, causing the entire slice window to shift to the right. The final horizontal range of the slice window is [missing information]. Plus The joint region. When the observer's horizontal yaw angle decreases, causing the slice window to move beyond the right screen boundary, the portion of the slice window exceeding the right screen boundary is cyclically shifted to the far edge of the left screen. The right screen boundary corresponds to the rightmost coordinate of the continuous three-fold screen splicing model. The right boundary of the slice window is .when When, the width of the excess part is The excess portion will be moved from the left coordinate of the continuous three-fold screen splicing model. The width is to be cut off starting from point . The rectangular area is spliced ​​to the right edge of the slice window. During the cyclic shifting process, the horizontal width of the slice window is maintained. constant.

[0065] A linear interpolation transition band is introduced at the shift boundary generated by the cyclic shift. The shift boundary is the seam of the cyclic splicing. Taking the left screen cyclic shift as an example, the seam at the left edge of the slice window is located in the model coordinate system. The mesh patches here come from two non-adjacent regions: region A is... Region B is ,in, The transition band width is set to the horizontal width of the slice window. Five percent. Within the transition zone, two texture sample values ​​at the same pixel location, possibly from region A or region B, are linearly blended. The blending weight changes linearly from the first side to the second side of the transition zone: on the region A side, the texture sample weight for region A is 1, and the texture sample weight for region B is 0; on the region B side, the texture sample weight for region A is 0, and the texture sample weight for region B is 1. After linear interpolation blending, the image transitions smoothly within the transition zone, eliminating visual jumps at the seams. The linear interpolation transition zone processing method during right-screen cyclic shift is symmetrical to that during left-screen cyclic shift, with the shift boundary being... Place.

[0066] See Figure 10 In the graph, the horizontal axis represents time in seconds, ranging from 0 to 50 seconds; the vertical axis represents the horizontal yaw rate in radians per second, ranging from approximately -0.1 to 0.1. The curves represented by dashed lines and triangular markers in the graph represent the original horizontal yaw rate. The curve represented by the solid line is the filtered horizontal yaw rate. .

[0067] As can be seen from the figure, the original horizontal yaw rate curve fluctuates significantly and contains a large amount of high-frequency noise, especially in the time interval of 0 to 15 seconds, where the fluctuations are more intense, with numerical oscillations between approximately ±0.07 radians / second. The filtered horizontal yaw rate curve shows the same overall trend as the original curve, but is significantly smoother, with reduced extreme values, effectively suppressed noise, and a more continuous and stable curve.

[0068] As time gradually increased from 0 seconds to 30 seconds, the filtered angular velocity showed a process of gradually decreasing from positive to negative, and the fluctuation range tended to decrease, indicating that the head rotation speed gradually slowed down from positive rotation and rotated in the opposite direction; after 30 seconds, the filtered angular velocity showed a trend of gradually rising from negative to close to zero.

[0069] This figure illustrates the effect of using a Kalman filter to filter the raw head rotation angular velocity signal obtained from real-time tracking in Example 5. The filtering process, based on the defined state and observation equations, effectively reduces measurement noise and abrupt changes in the angular velocity signal output by the head pose estimation convolutional neural network, improving the smoothness and accuracy of head motion state estimation. The filtered horizontal yaw angular velocity, used as input for subsequent dynamic adjustment of the field of view in the continuous three-fold splicing model slice rendering, enables more stable and natural field of view changes, enhancing the smoothness of the immersive interactive experience and user comfort.

[0070] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for immersive interactive display of three-dimensional architectural models with a three-fold screen, characterized in that, include: Acquire multi-source mapping data of the target building and construct an initial 3D mesh model; Based on the folding angle parameters of the three-fold screen space, the initial three-dimensional mesh model is decomposed into a left screen model, a middle screen model, and a right screen model. The left screen model and the right screen model form an angled transition zone with the middle screen model, respectively. Within the included angle transition zone, extract the depth misalignment point cloud of the left screen model or the right screen model relative to the middle screen model. The depth misalignment point cloud is used to define the projection overlap area of ​​the left screen model or the right screen model onto the middle screen model. A transition mesh sequence with non-uniform stitching weights is generated based on the depth misalignment point cloud, wherein the non-uniform stitching weights gradually decrease along the included angle transition zone from the middle screen model to the left screen model or the right screen model; According to the transition mesh sequence, the left screen model and the right screen model are respectively stitched with the middle screen model to obtain a continuous three-fold screen splicing model; The system tracks the observer's head rotation vector within the three-fold screen space in real time, and dynamically adjusts the slice rendering field of view of the continuous three-fold screen splicing model on the left, middle, and right screens based on the head rotation vector.

2. The immersive interactive display method for a three-dimensional architectural model with a three-fold screen as described in claim 1, characterized in that, The process of acquiring multi-source mapping data of the target building and constructing an initial 3D mesh model includes: The image point cloud acquired by UAV oblique photography is spatially aligned with the structural point cloud scanned by ground lidar. Using the aligned and fused point cloud as input, an initial 3D mesh model with normal information is generated by the Poisson surface reconstruction algorithm. Each vertex in the initial 3D mesh model stores the texture color value of the original image and the reflection intensity value of the structural point cloud.

3. The immersive interactive display method for a three-dimensional architectural model with a three-fold screen as described in claim 1, characterized in that, The process of decomposing the initial 3D mesh model into a left screen model, a middle screen model, and a right screen model based on the folding angle parameters of the three-fold screen space includes: Read the first included angle value between the left screen and the middle screen and the second included angle value between the middle screen and the right screen in the three-fold screen space. Taking the normal direction of the middle screen as the reference, draw the left screen projection cutting surface on the horizontal plane according to the first included angle value and the right screen projection cutting surface on the horizontal plane according to the second included angle value. Use the left screen projection cutting surface and the right screen projection cutting surface to perform a Boolean segmentation operation on the initial three-dimensional mesh model. After segmentation, the mesh part located in the left screen projection area is used as the left screen model, the mesh part located in the middle screen projection area is used as the middle screen model, and the mesh part located in the right screen projection area is used as the right screen model.

4. The immersive interactive display method for a three-dimensional spatial architectural model with a three-fold screen according to claim 1, characterized in that, The process of extracting the depth misalignment point cloud of the left screen model or the right screen model relative to the middle screen model within the included angle transition zone includes: Within the angled transition zone, the set of boundary vertices adjacent to the left screen model and the middle screen model, as well as the set of boundary vertices adjacent to the right screen model and the middle screen model, are traversed respectively. The first depth offset from the boundary vertex of the left screen model to the corresponding projection position of the middle screen model and the second depth offset from the boundary vertex of the right screen model to the corresponding projection position of the middle screen model are calculated. The boundary vertices of the left screen model when the first depth offset exceeds a preset depth threshold are extracted as the left screen depth misalignment point cloud, and the boundary vertices of the right screen model when the second depth offset exceeds the preset depth threshold are extracted as the right screen depth misalignment point cloud.

5. The immersive interactive display method for a three-dimensional architectural model with a three-fold screen as described in claim 4, characterized in that, The preset depth threshold is dynamically calculated based on the product of the sine of the angle between the left and middle screens or the middle and right screens in the three-fold screen space and the average side length of the model surface.

6. The immersive interactive display method for a three-dimensional architectural model with a three-fold screen as described in claim 4, characterized in that, The process of generating a transition mesh sequence with non-uniform stitching weights based on the depth misalignment cloud includes: Each vertex in the left screen depth misalignment point cloud is sorted in ascending order according to its distance to the boundary line between the left and middle screens to obtain the left screen vertex sequence. Each vertex in the right screen depth misalignment point cloud is sorted in ascending order according to its distance to the boundary line between the right and middle screens to obtain the right screen vertex sequence. A first non-uniform stitching weight sequence is assigned to the vertices in the left screen vertex sequence in descending order, and a second non-uniform stitching weight sequence is assigned to the vertices in the right screen vertex sequence in descending order. The first non-uniform stitching weight sequence is bound to the corresponding vertices in the left screen vertex sequence and inserted at the boundary between the left screen model and the middle screen model to form a left transition mesh sequence. The second non-uniform stitching weight sequence is bound to the corresponding vertices in the right screen vertex sequence and inserted at the boundary between the right screen model and the middle screen model to form a right transition mesh sequence.

7. The immersive interactive display method for a three-dimensional architectural model with a three-fold screen as described in claim 6, characterized in that, The non-uniform stitching weight decreases from 1 to 0 along an exponential decay curve in the direction perpendicular to the boundary line.

8. The immersive interactive display method for a three-dimensional architectural model with a three-fold screen as described in claim 6, characterized in that, The process of viewport stitching of the left screen model and the right screen model with the middle screen model according to the transition mesh sequence includes: Read the first non-uniform stitching weight of each vertex in the left transition mesh sequence. Calculate the weighted average of the spatial coordinates of the left screen model boundary vertex and the corresponding middle screen model boundary vertex according to the first non-uniform stitching weight to obtain the left stitching vertex coordinates. Replace the original left screen model boundary vertex coordinates with the left stitching vertex coordinates. Read the second non-uniform stitching weight of each vertex in the right transition mesh sequence. Calculate the weighted average of the spatial coordinates of the right screen model boundary vertex and the corresponding middle screen model boundary vertex according to the second non-uniform stitching weight to obtain the right stitching vertex coordinates. Replace the original right screen model boundary vertex coordinates with the right stitching vertex coordinates. After completing the viewport stitching, perform Laplacian smoothing on the common boundaries of the left and middle screen models and the right and middle screen models.

9. The immersive interactive display method for a three-dimensional architectural model with a three-fold screen as described in claim 1, characterized in that, The process of real-time tracking of the observer's head rotation vector within the triple-screen space includes: Multiple infrared cameras deployed at the top corner of the three-fold screen space capture depth images of the observer's head. Quaternion representations of head orientation are extracted from the depth images. These quaternion representations are then converted into Euler angle sequences. Horizontal yaw and vertical pitch angles are separated from the Euler angle sequences and combined to form a head rotation vector.

10. The immersive interactive display method for a three-dimensional architectural model with a three-fold screen as described in claim 9, characterized in that, The process of dynamically adjusting the slice rendering field of view of the continuous three-fold screen splicing model on the left, middle, and right screens based on the head rotation vector includes: The horizontal yaw angle in the head rotation vector is mapped to a scaling factor of the total horizontal span of the left, middle, and right screens. Based on the scaling factor, the starting and ending cutting columns of the slice window in the horizontal direction of the continuous three-fold splicing model are determined. The vertical pitch angle in the head rotation vector is mapped to the starting and ending cutting rows in the vertical direction of the continuous three-fold splicing model. The rectangular window jointly defined by the starting and ending cutting columns, the starting and ending cutting rows, and the ending cutting rows is used as the slice rendering field of view. The visible mesh patches on the left, middle, and right screens are subjected to frustum clipping, and only the mesh patches falling within the rectangular window are retained for rendering.