Method and system for making dome-screen film based on high-altitude aerial image
By constructing a panoramic space and performing non-rigid geometric deformation processing, pixel-aligned intermediate images are generated, and dynamic stitching lines are generated based on fusion costs. This solves the parallax and stitching line problems in the production of dome films from aerial photography, and achieves high-quality dome film production.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI HENGTIAN YIDA ADVERTISEMENT TRANSMISSION CO LTD
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies for producing dome films using aerial footage suffer from problems such as geometric misalignment due to parallax, lack of semantic understanding in stitching planning, stitching lines appearing in the viewer's field of vision, and an inability to balance sharpness and color seams during the blending process.
By constructing a panoramic space, performing non-rigid geometric deformation processing, generating pixel-aligned intermediate images, and generating dynamic stitching lines based on fusion costs, combined with multi-scale frequency domain adaptive fusion technology, a dome film with high visual consistency is generated.
It effectively corrects geometric misalignment caused by large parallax in high-altitude aerial photography, achieves precise pixel alignment, avoids stitching lines appearing in the core viewing area of the audience, and maintains image sharpness and visual immersion while eliminating stitching seams.
Smart Images

Figure CN121887971A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and digital media technology, and more specifically, to a method and system for producing dome films based on aerial photographs. Background Technology
[0002] Utilizing multi-lens arrays for aerial photography and creating immersive dome films has become a cutting-edge direction in special effects film production. However, when aerial equipment films scenes containing numerous close-up objects, such as cities and mountains, significant parallax occurs between lenses from different perspectives. This parallax causes the same 3D object to appear at different pixel positions in images from different viewpoints. Traditional "simple geometric stitching" methods based on feature point matching and global geometric transformations (such as homography transformation or simple spherical projection) are ineffective in this scenario. Even after initial alignment, obvious geometric misalignment remains in the overlapping areas, manifesting as unacceptable flaws such as torn building edges and "ghosting" of objects. To address residual misalignment, the industry generally adopts a strategy of finding "stitching lines" and blending on both sides. However, existing dome panoramic stitching technologies still face the following problems: First, the stitching line planning lacks an understanding of the semantic content of the image, failing to intelligently avoid important objects. Existing technologies typically plan the stitching path based on low-level features such as color or gradient differences between pixels. This method performs poorly in complex scenes captured by aerial photography, where the stitching line easily passes through rigid objects with clear structures and important semantics, such as buildings and utility poles, causing unnatural cuts to the objects and disrupting visual integrity. Simultaneously, it lacks proactive preference for smooth-textured areas (such as the sky and water surfaces), missing the optimal opportunity to hide seams using these visually less sensitive areas.
[0003] Secondly, there is a lack of optimization for the unique visual characteristics of a dome screen and control for video temporal stability. Traditional panoramic stitching algorithms often treat the image as a flat, two-dimensional picture, ignoring the unique dome structure of a dome screen. This results in stitching lines frequently appearing in the area of the ceiling or directly in front of the viewer, where the viewer's gaze is highest, thus disrupting the immersive experience. Furthermore, when processing consecutive video frames, due to the lack of temporal constraints, even minor camera shake or lighting changes can cause stitching lines to jump randomly between adjacent frames, producing noticeable flickering and dizziness when played on a giant dome screen.
[0004] Third, the "one-size-fits-all" blending process fails to strike a balance between eliminating color seams and protecting structural sharpness. After determining the seam line, a common blending technique involves feathering the edges based on transparency. While this method can mitigate seams in areas with color or brightness differences, it leads to blurring and ghosting of high-frequency details. Especially in areas rich in high-frequency information, such as architectural lines and text edges, simple alpha blending causes edge information from two images to overlap, producing unpleasant double shadows and blurring effects, severely sacrificing image sharpness. This is a fatal flaw for a projection environment where pixel integrity is extremely important and any flaws are magnified and scrutinized on a dome screen. Therefore, this invention proposes a method and system for producing dome films based on aerial photographs to solve the above problems. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a method for producing dome films based on high-altitude aerial photography, comprising: Projecting multi-view aerial video sequences onto a unified spherical coordinate system creates a panoramic space with overlapping areas; Non-rigid geometric deformation processing is applied to the overlapping areas of the panoramic space to generate pixel-aligned intermediate images; The fusion cost of each pixel in the overlapping region of the intermediate image is calculated, and a path search is performed based on the fusion cost of each pixel to generate a dynamic stitching line. The fusion cost is a weighted sum calculated by weighting the texture structure cost, semantic object cost, topological location cost and temporal stability cost. An image mask is generated based on the dynamic stitching line, and the intermediate image is subjected to multi-scale frequency domain adaptive fusion using the image mask to generate a fused panoramic image. The fused panoramic image is then remapped into a projection format suitable for dome screen playback and output.
[0006] Furthermore, constructing a panoramic space with overlapping areas specifically includes: Each frame of the multi-view aerial video sequence is subjected to distortion correction and color consistency correction to generate a corrected image. Feature points of the corrected images are extracted respectively, and feature points of adjacent viewpoints are matched. Based on the matching results, pose parameters of each corrected image relative to the unified spherical coordinate system are calculated. The pixel coordinates of the corrected image are mapped to the unified spherical coordinate system using the pose parameters, and the intersection area of the projections of adjacent viewpoints in the unified spherical coordinate system is marked as the overlapping area to obtain the panoramic space.
[0007] Furthermore, non-rigid geometric deformation processing is applied to the overlapping areas of the panoramic space, specifically including: The corrected images of two adjacent viewpoints within the overlapping region are divided into two-dimensional grids; Calculate the dense optical flow field between two adjacent viewpoint images within the overlapping region, and sample and interpolate each grid vertex of the two-dimensional grid based on the dense optical flow field to calculate the optical flow target displacement vector of each grid vertex; Construct an energy minimization function that includes a data alignment term and a rigidity preservation term, wherein the data alignment term is used to drive the mesh vertices to offset towards the optical flow target displacement vector, and the rigidity preservation term is used to constrain the local deformation amplitude of the mesh; Solving the energy minimization function yields the coordinates of the deformed mesh vertices, and based on these coordinates, the corrected image is remapped to generate the pixel-aligned intermediate image.
[0008] Furthermore, the calculation of the texture structure cost specifically includes: For each pixel in the overlapping region of the intermediate image, calculate the structure tensor of its local neighborhood; Extract two feature values from the structure tensor, and calculate the local anisotropy index of the pixel based on the feature values; For pixels with local anisotropy indices greater than a preset threshold, they are identified as regular geometric edge points and assigned a first texture value. For pixels whose anisotropy index is less than or equal to a preset index threshold, they are identified as disordered texture points and assigned a second texture value. The first texture value is greater than the second texture value.
[0009] Furthermore, the calculation of the semantic object cost specifically includes: The intermediate image is semantically segmented, and each pixel is divided into rigid structure category and non-rigid background category, and corresponding semantic labels are added; The optical flow residual map is calculated based on the pixel difference between adjacent viewpoints in the overlapping region after the non-rigid geometric deformation. For pixels whose semantic label is a non-rigid background category, assign a low semantic value; For pixels with the semantic label of rigid structure, the residual value in the optical flow residual map is further read, and the following judgment is performed: If the residual value is greater than a preset residual threshold, it indicates that the rigid structure is not aligned, and a first high semantic value is assigned. If the residual value is less than a preset residual threshold, it indicates that the rigid structure has been aligned, and a second high semantic value is assigned. Wherein, the first high semantic value is greater than the second high semantic value, and the second high semantic value is greater than the low semantic value.
[0010] Furthermore, the calculation of the topological location cost specifically includes: In the unified spherical coordinate system, the spherical geodesic distance of each pixel in the overlapping area of the intermediate images relative to the key anchor points is calculated, with the dome of the dome and the center of the audience’s frontal view as key anchor points. The spatial position sensitivity coefficient of each pixel is calculated based on the spherical geodesic distance, wherein the smaller the spherical geodesic distance, the higher the spatial position sensitivity coefficient. Visual saliency detection is performed on the intermediate image to obtain the saliency value of each pixel in the overlapping area; The topological location cost of each pixel is obtained by multiplying the spatial location sensitivity coefficient and the significance value of each pixel.
[0011] Furthermore, the calculation of the time-domain stabilization cost specifically includes: The optical flow field is calculated based on the intermediate image between the current frame and the previous frame; Based on the optical flow field, the dynamic stitching line of the intermediate image of the previous frame is mapped to the current frame to obtain the predicted stitching line path of the current frame; Calculate the shortest spatial distance from each pixel within the overlapping region to the predicted path of the suture line; The topological location cost of each pixel is compared with a preset importance threshold, and the overlapping region is divided into high importance region and low importance region based on the comparison result. For pixels located in the high importance region, the shortest spatial distance is weighted by a first weighting coefficient as its temporal stability cost; For pixels located in the low importance region, the shortest spatial distance is weighted by a second weighting coefficient as its temporal stability cost. Wherein, the first weighting coefficient is greater than the second weighting coefficient, and for the first frame image, the temporal stabilization cost is set to a preset constant value.
[0012] Furthermore, the method of path search based on the fusion cost of each pixel includes: The pixels in the overlapping area of the intermediate image are mapped as graph nodes, and connecting edges are established between adjacent pixel nodes to form a mesh graph. The sum of the fusion costs of the pixels at both ends of the connecting edge is used as the weight of the connecting edge; In the grid diagram, a source point and a sink point are set, and the source point and the sink point respectively correspond to the image data of two adjacent viewpoints that constitute the overlapping area. The nodes of the overlapping area relative to the two side boundaries are connected to the source point and the sink point respectively. The maximum flow minimum cut algorithm is used to find the cutting surface that separates the source and sink and minimizes the sum of the weights of the cut edges; The connecting lines corresponding to the cut surfaces are extracted as the dynamic suture lines.
[0013] Furthermore, the multi-scale frequency domain adaptive fusion of the intermediate image specifically includes: Using the dynamic stitching line as the boundary, the image source attribution of each pixel in the overlapping area is determined, and an image mask is generated; Construct the Laplacian pyramid of the intermediate image and the Gaussian pyramid of the image mask; The Gaussian pyramid is divided into low-frequency layer groups corresponding to low-frequency information of the image and high-frequency layer groups corresponding to high-frequency information of the image. Gaussian blurring is applied to the mask layers within the low-frequency layer group to expand the fusion transition region; edge gradients are maintained for the mask layers within the high-frequency layer group to preserve the stitching boundaries. Using the processed masks at each level, the corresponding levels of the Laplace pyramid are weighted and fused to generate the fused pyramid. The fused pyramid is reconstructed by inverse transformation to obtain the fused panoramic image.
[0014] A dome screen film production system based on high-altitude aerial imagery, used to implement the aforementioned method for producing dome screen films based on high-altitude aerial imagery, includes: The panoramic space construction module is used to project multi-view aerial video sequences onto a unified spherical coordinate system to construct a panoramic space with overlapping areas; The pixel alignment optimization module is used to perform non-rigid geometric deformation processing on the overlapping areas of the panoramic space to generate pixel-aligned intermediate images. The stitching generation module is used to calculate the fusion cost of each pixel in the overlapping area of the intermediate image, and perform path search based on the fusion cost of each pixel to generate a dynamic stitching. The fusion cost is a weighted sum calculated by weighting the texture structure cost, semantic object cost, topological location cost and temporal stability cost. The image generation module is used to generate an image mask based on the dynamic stitching line, use the image mask to perform multi-scale frequency domain adaptive fusion of the intermediate image to generate a fused panoramic image, and remap the fused panoramic image into a projection format suitable for dome screen playback for output.
[0015] The technical effects and advantages of this invention are as follows: This invention effectively corrects geometric misalignment caused by large parallax in high-altitude aerial photography by constructing a panoramic space and performing non-rigid geometric deformation processing on overlapping areas, achieving precise pixel alignment. By introducing a multi-dimensional cost function that integrates texture structure, semantic objects, topological position, and temporal stability to generate dynamic stitching lines, it not only achieves intelligent avoidance of rigid semantic objects such as buildings and active utilization of smooth texture areas in the stitching path, but also moves the seams out of the audience's core viewing area (such as the zenith and front) by combining the characteristics of the dome screen's field of view, and effectively suppresses the stitching line jumps and flicker between video frames by using temporal constraints. By employing multi-scale frequency domain adaptive fusion technology, it performs wide-area transition on low-frequency information of the image to eliminate color difference, and maintains gradient on high-frequency information to maintain sharpness. While eliminating stitching seams, it avoids edge blurring and ghosting, thereby generating a dome screen panoramic film with high visual consistency, strong immersion, and clarity, greatly improving the production quality of dome screen films. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of a method for producing dome films based on aerial photography according to the present invention; Figure 2 This is a schematic diagram of a dome screen film production system based on aerial photography, according to the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1 Please see Figure 1 As shown in this embodiment, a method for producing a dome screen film based on aerial photography includes: Multi-view aerial video sequences are projected onto a unified spherical coordinate system to construct a panoramic space with overlapping areas. Since high-altitude aerial lenses typically employ wide-angle or fisheye lenses to cover a larger field of view, and cameras facing different directions are significantly affected by lighting conditions, direct stitching would lead to severe geometric distortion and brightness abrupt changes. This invention employs the following processing steps to avoid this problem: Each frame of the multi-view aerial video sequence is read from the original image. To eliminate geometric distortion caused by wide-angle lenses and solve the problem of uneven lighting during multi-view shooting, distortion correction and color consistency correction are performed on each frame of the multi-view aerial video sequence to generate a corrected image. Specifically, pre-calibrated camera intrinsic parameters and distortion coefficients (radial and tangential distortion parameters) are read, and barrel or pincushion distortion of the original image is eliminated through a reverse mapping algorithm to restore the linear characteristics of the image lines. For the color temperature and brightness differences caused by front lighting or backlighting from different lens orientations, histogram matching or gain compensation algorithms are used. Taking one properly exposed image as a benchmark, the pixel value distribution of the remaining images is adjusted to generate a corrected image that is geometrically realistic and uniform in color and brightness.
[0019] To determine the relative positions of each viewpoint in three-dimensional space, feature points of the corrected images need to be extracted separately, and feature points of adjacent viewpoints need to be matched. Based on the matching results, the pose parameters of each corrected image relative to the unified spherical coordinate system are calculated. Specifically, each corrected image is first scanned, and corner points or texture-rich points are identified as feature points using feature extraction algorithms (such as SIFT or ORB). Then, image pairs of adjacent viewpoints are analyzed to find feature point pairs with the same descriptor, thus completing the feature point matching. Using the geometric constraints between these matched feature point pairs (e.g., the matching points should point in the same direction in the spherical coordinate system), the precise rotation angle and position information of each image relative to the preset unified spherical coordinate system are derived by solving a system of geometric equations (e.g., using a bundle adjustment strategy to optimize the rotation matrix), thereby obtaining the pose parameters and ensuring that all images can be accurately positioned in the same spherical panoramic space.
[0020] To establish a unified stitching benchmark and determine the fusion range, the pixel coordinates of the corrected image need to be mapped to the unified spherical coordinate system using the pose parameters. The intersection of projections from adjacent viewpoints in the unified spherical coordinate system is then marked as the overlapping region, resulting in the panoramic space. Specifically, a virtual unit sphere is established as the panoramic model. Based on the previously calculated rotation matrix and camera intrinsic parameters, each pixel on the 2D image is back-projected onto the 3D spherical coordinate system. During this process, areas on the sphere simultaneously covered by two or more adjacent images are detected and marked as overlapping regions. These overlapping regions are the targets for subsequent non-rigid deformation optimization and stitching line search, thus completing the construction of a panoramic space with a clear topological relationship from independent 2D images.
[0021] The overlapping areas of the panoramic space are subjected to non-rigid geometric deformation processing to generate pixel-aligned intermediate images. The steps are as follows: To achieve localized elastic deformation of the image to compensate for high-altitude parallax, the corrected images of adjacent viewpoints within the overlapping region need to be divided into two-dimensional grids. Unlike traditional methods that apply a single transformation matrix to the entire image, this embodiment discretizes the image (overlapping region) into a regular grid structure (e.g., a 50×50 pixel rectangular grid) composed of numerous vertices and edges. Each grid vertex is treated as an independent control point, capable of moving independently in subsequent steps, thereby causing local stretching or compression of the surrounding image texture. This fine-grained division endows the image with the ability to undergo non-rigid deformation, providing a geometric basis for correcting complex local misalignments.
[0022] To accurately detect pixel misalignment within the overlapping region, it is necessary to calculate the dense optical flow field between adjacent viewpoints within the overlapping region. Based on this dense optical flow field, the vertices of the two-dimensional grid are sampled or interpolated to determine the optical flow target displacement vector for each grid vertex. Specifically, a high-precision optical flow algorithm (such as DeepFlow or Farneback algorithm) is used to calculate the motion vector of each pixel within the overlapping region between two adjacent images, i.e., the dense optical flow field. Since the optical flow field is dense data at the pixel level, while the grid vertices are sparse, it is necessary to sample the optical flow vectors in the neighborhood around each grid vertex and perform bilinear interpolation on the sampling results to calculate the theoretical position that the vertex needs to move to align with the adjacent image. The difference between this position and the current position is the optical flow target displacement vector, which indicates the ideal direction of grid deformation.
[0023] To maintain the naturalness of the image structure while aligning, an energy minimization function needs to be constructed, comprising a data alignment term and a rigidity preservation term. The data alignment term drives the mesh vertices to offset towards the optical flow target displacement vector, while the rigidity preservation term constrains the local deformation amplitude of the mesh. After calculating the new vertex coordinates that minimize the total energy using a sparse linear equation solver, texture remapping is performed on the corrected image based on these mesh vertex coordinates, generating an intermediate image that eliminates parallax misalignment while maintaining visual naturalness and pixel alignment. Energy Minimization Function The function expression is: In the formula, Indicates the transformed first The coordinate vectors of each grid vertex , where is the unknown quantity to be solved; Indicates the transformed first The coordinate vectors of the grid vertices; Indicates the first state before deformation (initial state). The coordinate vectors of the grid vertices; Indicates the first state before deformation (initial state). The coordinate vectors of the grid vertices; Indicates the first The optical flow target displacement vector corresponding to each grid vertex; Indicates the first The confidence weight of each vertex can be set based on the confidence of optical flow calculation or texture richness; the richer the texture, the higher the weight. Represents the set of all grid vertices; Represents the square of the Euclidean distance; This represents the set of all connected edges in the mesh. Indicates two adjacent vertices; This is the balance coefficient, with a value range of [5, 50]. Represents the vertices after deformation With vertex The edge vector between them; Represents the vertex before deformation With vertex The edge vectors between them.
[0024] It should be noted that the above energy minimization function Essentially, it is about an unknown quantity. The quadratic function; to obtain the coordinates of the deformed mesh vertices, it is necessary to differentiate this function and set the derivative to zero, thereby transforming it into a sparse linear system of equations. Solve the problem; the solution obtained This is the optimal vertex coordinate that balances optical flow alignment and structural rigidity.
[0025] The fusion cost of each pixel in the overlapping region of the intermediate image is calculated, and a path search is performed based on the fusion cost of each pixel to generate a dynamic stitching line. The fusion cost is a weighted sum calculated by weighting the texture structure cost, semantic object cost, topological location cost and temporal stability cost, and the sum of the weights is 1.
[0026] To distinguish between complex textured regions and flat regions in an image, it is necessary to calculate the structure tensor of the local neighborhood of each pixel within the overlapping region of the intermediate image. The structure tensor is a second-order matrix that describes the local gradient distribution of an image and can effectively characterize the texture directionality and intensity of the region. For any pixel... First, calculate its horizontal gradient. and vertical gradient (The Sobel operator or Gaussian derivative filter can be used), and then within a local window centered on that pixel (such as a 5×5 or 7×7 neighborhood), construct the following 2×2 structure tensor matrix. Its structural tensor It is a 2×2 symmetric matrix, representing the smoothed result of the outer product of the image gradient vectors within a local window. The specific calculation formula is as follows: In the formula, It is a Gaussian smoothing kernel with a standard deviation of σ, used to perform weighted averaging of gradient information in the local neighborhood, thereby improving robustness to noise.
[0027] To quantify the texture features of a pixel, it is necessary to extract two eigenvalues from the structure tensor and calculate the local anisotropy index of the pixel based on the ratio or difference of the eigenvalues. Specifically, for the aforementioned structure tensor matrix... Eigenvalue decomposition is performed to obtain two non-negative eigenvalues, a1 (representing the maximum rate of change of the local gradient) and a2 (representing the rate of change in the vertical direction), and the local anisotropy index of the pixel is calculated based on these eigenvalues. The calculation formula is: In the formula, It is a very small positive number (such as 10 to the power of negative 6) to prevent the denominator from being zero.
[0028] For pixels with a local anisotropy index higher than a preset threshold (e.g., 0.5), they are identified as regular geometric edge points and assigned a first texture cost (e.g., 1.0). These points typically correspond to prominent structures such as building edges and road lines. If a stitching line passes through such areas, it can easily cause visual breaks, so a higher cost is assigned to avoid them. Conversely, pixels with a local anisotropy index lower than a preset threshold are identified as disordered texture points (or flat points) and assigned a second texture cost (e.g., 0.1). These points typically correspond to the sky, grass, or water surface. Stitching lines passing through such areas are not easily detected, so a lower cost is assigned to attract stitching lines to pass through them. It should be noted that the first texture cost should be greater than the second texture cost, thereby guiding the dynamic stitching algorithm to preferentially select flat or disordered textured areas as stitching paths, hiding stitching traces to the greatest extent.
[0029] To enable the seam line to have content-aware capabilities, semantic segmentation of the overlapping areas in the intermediate images is also required. Using a pre-trained deep learning semantic segmentation model (such as DeepLabV3+ or PSPNet), pixel-level classification of the image content in the overlapping areas is performed, categorizing each pixel into rigid structure and non-rigid background categories, and adding corresponding semantic labels. The rigid structure category includes objects with fixed shapes and obvious geometric features, such as buildings, bridges, vehicles, and fences; the non-rigid background category includes areas with variable shapes or high texture repetition, such as the sky, clouds, water surfaces, and vegetation.
[0030] The optical flow residual map is calculated based on the pixel difference between adjacent viewpoints in the overlapping region after the non-rigid geometric deformation. The optical flow residual map reflects the quality of pixel alignment, and its value is the Euclidean distance between the RGB values of the corresponding pixels after alignment. In the formula, Represents pixels The optical flow residual value at a given pixel quantifies the degree of difference between the two images to be stitched after non-rigid deformation alignment at that pixel location. The smaller the value, the better the alignment effect; the larger the value, the more obvious the misalignment or color difference. Indicates the reference image at pixel points The color value (RGB vector) at the location; when stitching two adjacent images, one of them is usually selected as the reference image, whose geometric position remains unchanged or as the target reference for transformation. Indicates the distortion of the image at the pixel level The color value (RGB vector) at that location. It refers to the color value mapped to the pixel after the non-rigid geometric deformation (based on mesh deformation) process in step 3 from another adjacent image (the source image). Image data of the location. Norm operations on vectors are represented using the L2 norm (Euclidean distance).
[0031] Based on the aforementioned semantic labels and residual information, a hierarchical strategy is adopted to assign semantic values. For pixels with semantic labels belonging to the non-rigid background category, a low semantic value (e.g., 0.1) is assigned. This low semantic value is because areas such as the sky or grass, even with slight misalignments, are not easily detected by the human eye and are suitable as seam hiding areas, thus assigning the lowest cost to encourage the seam line to pass through. For pixels with semantic labels belonging to the rigid structure category, their residual values in the optical flow residual map are further read, and the following judgment is performed: If the residual value is greater than a preset residual threshold (e.g., pixel grayscale difference greater than 30), it indicates that the rigid structure is not aligned; this means that if stitching is done at this location, obvious ghosting or breakage will inevitably occur; therefore, a first high semantic value (e.g., 100.0) is assigned, which is a very large penalty term designed to force the stitching line to absolutely bypass this area.
[0032] If the residual value is less than the preset residual threshold, it indicates that the rigid structure has been aligned. Although the pixel overlap is good at this point, considering the complex light reflection characteristics of rigid objects (such as building facades) and their extremely high requirements for geometric integrity, we still tend to protect their integrity during splicing; therefore, a second high semantic cost value (e.g., 10.0) is assigned. Although this value is much smaller than the cost of the misaligned area, it is still significantly higher than the cost of the background area.
[0033] It should be noted that the above assignment logic satisfies that the first high semantic value is greater than the second high semantic value, and the second high semantic value is greater than the low semantic value, thereby ensuring that the stitching line preferentially selects the background area; when it is impossible to avoid rigid objects, the aligned edges of rigid objects are preferentially selected, and the core area of unaligned rigid objects is avoided.
[0034] In the unified spherical coordinate system, the zenith of the dome and the center of the audience's frontal viewing area are used as key anchor points. The spherical geodesic distance of each pixel within the overlapping area of the intermediate images relative to these key anchor points is calculated. Specifically, two fixed coordinate points are set on the unit sphere model: one is the "zenith anchor point" located at the highest point of the sphere (corresponding to the vertex of the physical dome), and the other is the "center anchor point of the frontal viewing area" located at a specific angle (e.g., 30 degrees elevation) above the equator of the sphere, directly facing the audience seating area. For each pixel within the overlapping area, the radian value of the angle between its corresponding spherical vector and the vectors of the two key anchor points is calculated using the inverse cosine function, and the minimum value is selected as the spherical geodesic distance of that pixel. For example, suppose The zenith anchor point of the dome screen (unit vector) ); Anchor point at the center of the viewer's frontal field of view (unit vector) For any pixel within the overlapping region (Unit vector) Distance between the zenith anchor point and the zenith anchor point The distance between and the center anchor point of the frontal view area They are respectively and The minimum value among them is the pixel. The spherical geodesic distance.
[0035] To convert physical distance into a weighting metric usable by the algorithm, a spatial location sensitivity coefficient for the pixel needs to be calculated based on the spherical geodesic distance. This coefficient quantifies the tolerance of that location to visual defects. The calculation formula is: ;in, This is the normalization coefficient, with a value range of [1.0, 10.0]. The parameter used to control the range of the sensitive area is set to [0.2, 0.8] (unit: radians). This formula indicates that the closer to the zenith or the frontal viewing area (…), the more sensitive the area. (smaller) The higher the value, the more sensitive the area is, and the presence of suture lines is strictly prohibited; as the distance increases, It decays rapidly, and the sensitivity of the edge regions decreases.
[0036] Since simple geometric location is insufficient to cover all visual hotspots (such as a bright bird appearing at the edge of the image), visual saliency detection is also required for the intermediate image to obtain the saliency values of each pixel within the overlapping region. Specifically, an Itti-Koch model or a saliency detection algorithm based on frequency domain residuals is used to generate a saliency map, and the values of each pixel in the saliency map are... That is, a pixel. The saliency value indicates that the higher the value, the more eye-catching the pixel is.
[0037] The spatial location sensitivity coefficient of each pixel is multiplied by its saliency value to obtain the topological location cost of each pixel; this calculation ensures that the pixel is only considered to have a topological location cost if it is located in the edge region of the dome. (low), and not a visually salient target ( When the topological location cost is low, it becomes the ideal path for the suture.
[0038] To establish the pixel correspondence between the current frame and the previous frame, the optical flow field needs to be calculated based on the intermediate image between the current frame and the previous frame. Specifically, two adjacent frames in the time series (i.e., the image at the previous time t-1 and the image at the current time t) are used as input, and a fast dense optical flow algorithm (such as DIS-Flow or Farneback algorithm) is used to calculate the motion vector of each pixel in the overlapping region. This vector describes the specific offset of a pixel from the previous frame to the current frame.
[0039] Based on the optical flow field, the dynamic stitching line of the intermediate image of the previous frame is mapped to the current frame to obtain the predicted stitching line path of the current frame; specifically, the set of coordinates of all pixels on the optimal stitching line determined in the previous frame is obtained. ;for Each coordinate point in Add its corresponding optical flow vector This will give you the expected position of the point in the current frame. Connecting all the expected positions forms the stitching prediction path for the current frame. This path represents the theoretical location where the suture line should be if the object's movement remains continuous.
[0040] To quantify the degree to which a pixel in the current frame deviates from the predicted path, it is also necessary to calculate the shortest spatial distance from each pixel within the overlapping region to the predicted path of the stitching line. Specifically, for any pixel within the overlapping region of the current frame... Calculate its relationship with the predicted path The shortest spatial distance (preferably Euclidean distance) between the nearest points. A larger value for this distance means that if the new stitch line is selected at a pixel... At a certain point, a dramatic spatial shift will occur; conversely, a more stable visual appearance will result.
[0041] Considering that shaking in the center of the dome screen is more likely to cause dizziness than in the edge areas, a differentiated weighting strategy needs to be introduced. This involves comparing the topological position cost of each pixel with a preset importance threshold, and dividing the overlapping area into high-importance and low-importance regions based on the comparison results. Specifically, using the previously calculated topological position cost, an importance threshold is set (preferably 50% to 70% of the maximum topological position cost). Pixels with a topological position cost higher than the importance threshold are marked as "high-importance regions" (i.e., near the zenith or the frontal viewing area), while those with a cost lower than the threshold are marked as "low-importance regions." For pixels located in the high-importance region, the shortest spatial distance is weighted by a first weighting coefficient (e.g., 10.0) as its temporal stability cost; for pixels located in the low-importance region, the shortest spatial distance is weighted by a second weighting coefficient (e.g., 1.0) as its temporal stability cost; wherein the first weighting coefficient is greater than the second weighting coefficient. This design allows the algorithm to impose a strong penalty on the stitching line's movement in the core region of the dome, forcing it to closely follow the path of the previous frame to ensure absolute visual stability; while in the edge region, the stitching line is allowed some degree of freedom to find new paths with better texture alignment.
[0042] It should be noted that for the first frame image, the temporal stabilization cost is set to a preset constant value (e.g., 0), because there is no previous frame and there is no need to consider temporal constraints.
[0043] The pixels within the overlapping region of the intermediate image are mapped to graph nodes, and connecting edges are established between adjacent pixel nodes to form a mesh graph. Specifically, for each discrete coordinate position within the overlapping region... Create a unique graph node As a decision unit, undirected connecting edges are established between adjacent graph nodes in the horizontal and vertical directions according to the spatial adjacency relationship of the image (e.g., using the 4-neighborhood connection rule); these connecting edges constitute potential stitching lines that traverse the network, thereby forming a regular grid topology (n-links, i.e., neighborhood connections).
[0044] It is important to note that although the overlapping region contains aligned image data from two adjacent viewpoints after non-rigid geometric deformation (i.e., first viewpoint data A and second viewpoint data B) at the data level, when constructing the graph model, only a unique set of node and edge topology needs to be established for the spatial coordinate grid of the overlapping region; that is, each graph node represents a spatial location, not a specific image data; at this location, two candidate pixel data can be accessed, and a label is subsequently assigned to the node using a graph cut algorithm (i.e., selecting A or selecting B), thereby determining which viewpoint pixel value is ultimately used for this spatial location.
[0045] To quantify the cost of cutting off adjacent pixel nodes, the sum of the fusion costs of the pixels at both ends of the connecting edge is used as the weight of that edge. This weight reflects the visual cost of cutting off this edge across these two pixels. A larger weight indicates the presence of complex textures, important objects, or key viewpoints, making it unsuitable for cutting; a smaller weight indicates a flat background or a low-sensitivity area, suitable as a stitching boundary. In this way, pixel-level fusion costs can be naturally mapped to cutting costs on the graph structure.
[0046] To represent the binary decision problem of "choosing which side's viewpoint data" in a graphical model, a source and a sink point need to be defined in the mesh graph. The source and sink points correspond to the image data of two adjacent viewpoints constituting the overlapping region, respectively. Nodes on the opposite sides of the overlapping region are connected to the source and sink points, respectively. Specifically, two virtual nodes are introduced: Source point: Represents the distorted image data from the first-person perspective (e.g., the left-hand camera). If a node eventually remains connected to the source node, then that location uses... The pixel value.
[0047] Sum of points: Represents the distorted image data from a second-person perspective (such as the right-hand camera). If a node eventually remains connected to the sink, then that location uses... The pixel value.
[0048] To ensure the topological correctness of the spliced image (i.e., retaining the left image on the left and the right image on the right), hard boundary constraints also need to be applied: The boundary pixel nodes of the overlapping region closest to the first-view side (e.g., all nodes in the leftmost column) are connected to the source node through edges with infinite (or maximum) weights (called t-links, i.e., terminal connections), forcing these nodes to select first-view data. ; Connect the boundary pixel nodes of the overlapping region closest to the second-view side (e.g., all nodes in the rightmost column) to the sink node using edges with infinite weights, thus forcing these nodes to select second-view data. .
[0049] The above processing can be used to construct a standard ST flow network, with the source and sink nodes located at opposite ends of the grid, and the middle nodes connected by ordinary pixel nodes through n-links, forming a complete graph cut model.
[0050] The maximum flow minimum cut algorithm is used to find the cutting plane that separates the source and sink nodes and minimizes the sum of the weights of the cut edges. According to the maximum flow minimum cut theorem, this process is equivalent to finding a cutting plane that divides all graph nodes into two disjoint sets (belonging to...). The set and belonging The algorithm finds a set of n-links that are cut by the cutting plane, and minimizes the sum of the weights of all connecting edges (n-links). This means that the algorithm automatically finds a path that traverses regions with the lowest fusion cost. Finally, this cutting path is plotted on the image plane, resulting in the optimal dynamic stitching line. Pixels on one side of this line are taken from the first-view image, and pixels on the other side are taken from the second-view image, thus achieving seamless stitching with minimal stitching artifacts and optimized visual quality.
[0051] An image mask is generated based on the dynamic stitching line, and the intermediate image is subjected to multi-scale frequency domain adaptive fusion using the image mask to generate a fused panoramic image. The fused panoramic image is then remapped into a projection format suitable for dome screen playback and output.
[0052] Using the dynamic stitching line as the boundary, the image source attribution of each pixel within the overlapping area is determined, and an image mask is generated; specifically, for each pixel within the overlapping area... Based on its attribution label in the graph cut result, a binary mask is generated. If the point belongs to the first-person perspective (source point set), then let Set it to 1; if it belongs to the second perspective (sink set), then let The value is 0; this mask clearly defines the hard boundaries of the stitching.
[0053] To perform targeted processing in different frequency domains, it is necessary to construct the Laplacian pyramid of the intermediate image and the Gaussian pyramid of the image mask; specifically, for the first-view image... Second-view images Construct Laplace's pyramids with N layers (e.g., N=5 or N=7). and The top layer of the pyramid contains low-frequency overview information of the image, while the bottom layer contains high-frequency detail information. A Gaussian pyramid with the corresponding number of layers is constructed from the binary mask. During the construction process, each mask layer is obtained through Gaussian blurring and downsampling, so that high-level masks exhibit smooth grayscale transitions, while low-level masks maintain relatively sharp edges.
[0054] To balance color smoothness and detail clarity, the pyramid needs to be divided into low-frequency layer groups (e.g., the top L layer) corresponding to low-frequency information of the image and high-frequency layer groups (e.g., the bottom 0 to L-1 layers) corresponding to high-frequency information of the image.
[0055] For the mask layers within the low-frequency layer group, a strong Gaussian blur is applied. This smooth mask is then used to perform weighted fusion of the low-frequency layers of the image, allowing the differences in brightness and hue between the two images to transition slowly over a wider area, thus completely eliminating obvious color difference boundaries. For the mask layers within the high-frequency layer group, the edge gradient is preserved, making it close to a binarized state (i.e., preserving the stitching boundary). This sharp mask is then used to fuse the high-frequency layers of the image, ensuring that texture details near the stitching line (such as building lines and tree branches) do not produce ghosting or blurring due to large-scale feathering, thus achieving no reduction in sharpness at the stitching point.
[0056] Using the processed masks at each level, the corresponding levels of the Laplacian pyramid are weighted and fused to generate the fused pyramid. For the first pyramid The layer fusion formula is: .
[0057] The fused pyramid is then reconstructed using an inverse transform to obtain a fused panoramic image. Specifically, starting from the top of the pyramid, each layer is upsampled and detailed information from the next layer is added until the original resolution image is reconstructed. This final image achieves seamless color transitions in low-frequency information and retains clear texture structure in high-frequency information, perfectly meeting the high-definition playback requirements of a dome screen.
[0058] The reconstructed panoramic images are typically in a standard isometric cylindrical projection format, which is convenient for storage and transmission but cannot be directly displayed correctly on the hemispherical screen of a dome theater. Therefore, remapping processing is required to convert the fused panoramic images into a projection format suitable for dome playback (such as fisheye projection format). Specifically, a mapping relationship is established from the target fisheye image plane to the source isometric cylindrical image plane, and each pixel is resampled using a reverse lookup table method. During the remapping process, geometric stretching compensation is performed on the area near the edge of the dome to offset the visual distortion caused by the curvature of the dome, and finally, a high-resolution (such as 4K / 8K) video frame sequence that conforms to the standards of dome projectors is output.
[0059] Example 2 Please see Figure 2 As shown, for parts not described in detail in this embodiment, please refer to the description in Embodiment 1. A method and system for producing dome-screen films based on high-altitude aerial imagery is provided, including: The panoramic space construction module is used to project multi-view aerial video sequences onto a unified spherical coordinate system to construct a panoramic space with overlapping areas; The pixel alignment optimization module is used to perform non-rigid geometric deformation processing on the overlapping areas of the panoramic space to generate pixel-aligned intermediate images. The stitching generation module is used to calculate the fusion cost of each pixel in the overlapping area of the intermediate image, and perform path search based on the fusion cost of each pixel to generate a dynamic stitching. The fusion cost is a weighted sum calculated by weighting the texture structure cost, semantic object cost, topological location cost and temporal stability cost. The image generation module is used to generate an image mask based on the dynamic stitching line, use the image mask to perform multi-scale frequency domain adaptive fusion of the intermediate image to generate a fused panoramic image, and remap the fused panoramic image into a projection format suitable for dome screen playback for output.
[0060] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0061] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0062] In the description of this invention, it should be understood that the terms "first," "second," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0063] In the description of this invention, unless otherwise stated, "a plurality of" means two or more.
[0064] In the description of this invention, "several" means one or more, and "a large number" means two or more.
[0065] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0066] All formulas in this manual are dimensionless and calculated numerically. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.
[0067] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A method for producing dome films based on aerial photographs, characterized in that, include: Projecting multi-view aerial video sequences onto a unified spherical coordinate system creates a panoramic space with overlapping areas; Non-rigid geometric deformation processing is applied to the overlapping areas of the panoramic space to generate pixel-aligned intermediate images; The fusion cost of each pixel in the overlapping region of the intermediate image is calculated, and a path search is performed based on the fusion cost of each pixel to generate a dynamic stitching line. The fusion cost is a weighted sum calculated by weighting the texture structure cost, semantic object cost, topological location cost and temporal stability cost. An image mask is generated based on the dynamic stitching line, and the intermediate image is subjected to multi-scale frequency domain adaptive fusion using the image mask to generate a fused panoramic image. The fused panoramic image is then remapped into a projection format suitable for dome screen playback and output.
2. The method according to claim 1, characterized in that, Constructing panoramic spaces with overlapping areas specifically includes: Each frame of the multi-view aerial video sequence is subjected to distortion correction and color consistency correction to generate a corrected image. Feature points of the corrected images are extracted respectively, and feature points of adjacent viewpoints are matched. Based on the matching results, the pose parameters of each corrected image relative to the unified spherical coordinate system are calculated. The pixel coordinates of the corrected image are mapped to the unified spherical coordinate system using the pose parameters, and the intersection area of the projections of adjacent viewpoints in the unified spherical coordinate system is marked as the overlapping area to obtain the panoramic space.
3. The method according to claim 2, characterized in that, The overlapping areas of the panoramic space are subjected to non-rigid geometric deformation processing, specifically including: The corrected images of two adjacent viewpoints within the overlapping region are divided into two-dimensional grids; Calculate the dense optical flow field between two adjacent viewpoint images within the overlapping region, and sample and interpolate each grid vertex of the two-dimensional grid based on the dense optical flow field to calculate the optical flow target displacement vector of each grid vertex; Construct an energy minimization function that includes a data alignment term and a rigidity preservation term, wherein the data alignment term is used to drive the mesh vertices to offset towards the optical flow target displacement vector, and the rigidity preservation term is used to constrain the local deformation amplitude of the mesh; Solving the energy minimization function yields the coordinates of the deformed mesh vertices, and based on these coordinates, the corrected image is remapped to generate the pixel-aligned intermediate image.
4. The method according to claim 3, characterized in that, The calculation of the texture structure cost specifically includes: For each pixel in the overlapping region of the intermediate image, calculate the structure tensor of its local neighborhood; Extract two feature values from the structure tensor, and calculate the local anisotropy index of the pixel based on the feature values; For pixels with local anisotropy indices greater than a preset threshold, they are identified as regular geometric edge points and assigned a first texture value. For pixels whose anisotropy index is less than or equal to a preset index threshold, they are identified as disordered texture points and assigned a second texture value. The first texture value is greater than the second texture value.
5. The method according to claim 4, characterized in that, The calculation of the semantic object cost specifically includes: The intermediate image is semantically segmented, and each pixel is divided into rigid structure category and non-rigid background category, and corresponding semantic labels are added; The optical flow residual map is calculated based on the pixel difference between adjacent viewpoints in the overlapping region after the non-rigid geometric deformation. For pixels whose semantic label is a non-rigid background category, assign a low semantic value; For pixels with the semantic label of rigid structure, the residual value in the optical flow residual map is further read, and the following judgment is performed: If the residual value is greater than a preset residual threshold, it indicates that the rigid structure is not aligned, and a first high semantic value is assigned. If the residual value is less than a preset residual threshold, it indicates that the rigid structure has been aligned, and a second high semantic value is assigned. Wherein, the first high semantic value is greater than the second high semantic value, and the second high semantic value is greater than the low semantic value.
6. The method according to claim 5, characterized in that, The calculation of the topological location cost specifically includes: In the unified spherical coordinate system, the spherical geodesic distance of each pixel in the overlapping area of the intermediate images relative to the key anchor points is calculated, with the dome of the dome and the center of the audience’s frontal view as key anchor points. The spatial position sensitivity coefficient of each pixel is calculated based on the spherical geodesic distance, wherein the smaller the spherical geodesic distance, the higher the spatial position sensitivity coefficient. Visual saliency detection is performed on the intermediate image to obtain the saliency value of each pixel in the overlapping area; The topological location cost of each pixel is obtained by multiplying the spatial location sensitivity coefficient and the significance value of each pixel.
7. The method according to claim 6, characterized in that, The calculation of the time-domain stability cost specifically includes: The optical flow field is calculated based on the intermediate image between the current frame and the previous frame; Based on the optical flow field, the dynamic stitching line of the intermediate image of the previous frame is mapped to the current frame to obtain the predicted stitching line path of the current frame; Calculate the shortest spatial distance from each pixel within the overlapping region to the predicted path of the suture line; The topological location cost of each pixel is compared with a preset importance threshold, and the overlapping region is divided into high importance region and low importance region based on the comparison result. For pixels located in the high importance region, the shortest spatial distance is weighted by a first weighting coefficient as its temporal stability cost; For pixels located in the low importance region, the shortest spatial distance is weighted by a second weighting coefficient as its temporal stability cost. Wherein, the first weighting coefficient is greater than the second weighting coefficient, and for the first frame image, the temporal stabilization cost is set to a preset constant value.
8. The method according to claim 7, characterized in that, The path search methods based on the fusion cost of each pixel include: The pixels in the overlapping area of the intermediate image are mapped as graph nodes, and connecting edges are established between adjacent pixel nodes to form a mesh graph. The sum of the fusion costs of the pixels at both ends of the connecting edge is used as the weight of the connecting edge; In the grid diagram, a source point and a sink point are set, and the source point and the sink point respectively correspond to the image data of two adjacent viewpoints that constitute the overlapping area. The nodes of the overlapping area relative to the two side boundaries are connected to the source point and the sink point respectively. The maximum flow minimum cut algorithm is used to find the cutting surface that separates the source and sink and minimizes the sum of the weights of the cut edges; The connecting lines corresponding to the cut surfaces are extracted as the dynamic suture lines.
9. The method according to claim 8, characterized in that, The multi-scale frequency domain adaptive fusion of the intermediate image specifically includes: Using the dynamic stitching line as the boundary, the image source attribution of each pixel in the overlapping area is determined, and an image mask is generated; Construct the Laplacian pyramid of the intermediate image and the Gaussian pyramid of the image mask; The Gaussian pyramid is divided into low-frequency layer groups corresponding to low-frequency information of the image and high-frequency layer groups corresponding to high-frequency information of the image. Gaussian blurring is applied to the mask layers within the low-frequency layer group to expand the fusion transition region; edge gradients are maintained for the mask layers within the high-frequency layer group to preserve the stitching boundaries. Using the processed masks at each level, the corresponding levels of the Laplace pyramid are weighted and fused to generate the fused pyramid. The fused pyramid is reconstructed by inverse transformation to obtain the fused panoramic image.
10. A dome screen film production system based on aerial photography, used to implement the dome screen film production method based on aerial photography as described in any one of claims 1 to 9, characterized in that, include: The panoramic space construction module is used to project multi-view aerial video sequences onto a unified spherical coordinate system to construct a panoramic space with overlapping areas; The pixel alignment optimization module is used to perform non-rigid geometric deformation processing on the overlapping areas of the panoramic space to generate pixel-aligned intermediate images. The stitching generation module is used to calculate the fusion cost of each pixel in the overlapping area of the intermediate image, and perform path search based on the fusion cost of each pixel to generate a dynamic stitching. The fusion cost is a weighted sum calculated by weighting the texture structure cost, semantic object cost, topological location cost and temporal stability cost. The image generation module is used to generate an image mask based on the dynamic stitching line, use the image mask to perform multi-scale frequency domain adaptive fusion of the intermediate image to generate a fused panoramic image, and remap the fused panoramic image into a projection format suitable for dome screen playback for output.
Citation Information
Cited By
Combined processing method for power monitoring image denoising and seamless fusion
CN122243803A