X-ray sparse new view synthesis method based on inline prior guidance three-dimensional Gaussian splashing

By constructing an inline prior-guided 3D Gaussian splashing method, and utilizing depth constraints, mask reconstruction, and improved graph Laplacian regularized inline priors, the problem of synthesizing sparse new X-ray views under sparse perspectives was solved, achieving high-quality new view synthesis.

CN121937575APending Publication Date: 2026-04-28ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2026-01-26
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing 3D Gaussian splashing methods suffer from insufficient structural information, inherent suboptimal distribution of rendered images, overfitting of the model to a limited number of viewpoints, and spatial inconsistencies and artifacts caused by uneven density distribution under sparse viewpoint conditions, making it difficult to achieve high-quality synthesis of sparse new X-ray views.

Method used

A three-dimensional Gaussian splashing method based on inline priors is adopted. By constructing depth-constrained inline priors, mask reconstruction inline priors, and improved graph Laplacian regularization inline priors, prior information is directly extracted from the three-dimensional Gaussian representation and volume density field, optimizing the parameters of the three-dimensional Gaussian splashing model and generating high-quality new views.

Benefits of technology

The method significantly improves the stability and reconstruction quality of new view synthesis under sparse view conditions, provides higher accuracy and robustness, solves the problems of inaccurate reconstruction and artifacts under sparse view conditions, and realizes high-quality X-ray new view synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937575A_ABST
    Figure CN121937575A_ABST
Patent Text Reader

Abstract

The invention discloses an X-ray sparse new view synthesis method based on inline prior guidance three-dimensional Gaussian splashing. Compared with an existing method based on three-dimensional Gaussian splashing, the depth constraint inline prior constructed by the method solves the problem of insufficient structure information under a sparse view angle; the constructed mask reconstructs inline prior, and the problems of inherent suboptimum and detail missing of rendered image distribution and over-fitting of three-dimensional gauss to a limited view angle under the sparse view angle condition are solved; on the basis, improved graph Laplacian regularization inline prior is introduced, adaptive weighted constraint is performed on density difference of adjacent Gaussian points, and the problems of space inconsistency and artifacts caused by non-uniform density distribution are solved. Finally, a new solution is provided for X-ray sparse new view synthesis, and the new view synthesis performance and reconstruction quality under the sparse view angle condition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image processing and computer vision technology, and in particular relates to a method for synthesizing new sparse X-ray views based on inline prior-guided three-dimensional Gaussian splashing. Background Technology

[0002] 3D Gaussian Splatting, as an emerging explicit 3D scene representation method, achieves high-quality rasterization rendering and new view synthesis by optimizing parameters such as the position, scale, rotation, and density of each Gaussian point in an explicit 3D Gaussian point cloud. This method has achieved significant success in new view synthesis tasks for natural scenes, demonstrating advantages such as fast rendering speed and high reconstruction quality. However, under extremely sparse view conditions, existing 3D Gaussian Splatting methods (such as X-Gaussian, ...) ... The performance of Gaussian (Gaussian model) differs significantly from that of dense viewpoints. This is mainly attributed to the following issues: First, the structural information is insufficient in sparse viewpoints, and relying solely on pixel-level reconstruction loss is insufficient to adequately constrain the geometric consistency of the 3D structure, leading to a lack of depth information and inaccurate structural reconstruction. Second, due to limited observation data, the 3D Gaussian model tends to perform memorization fitting on the limited viewpoint observation data, resulting in an inherent suboptimal distribution of the rendered image under sparse viewpoint conditions. Problems such as intensity distribution deviation and loss of detail occur in unobserved viewpoints, and the model overfits to the training viewpoint, resulting in poor generalization ability. Third, under sparse viewpoint conditions, the density distribution of 3D Gaussian points is prone to local non-uniformity, with excessive density differences between adjacent Gaussian points, leading to spatial inconsistencies and artifacts and blurring in boundary regions, which seriously affect the reconstruction quality.

[0003] In recent years, spatial consistency constraint methods such as graph Laplacian regularization have been introduced into the 3D Gaussian splashing framework to alleviate the problem of uneven density distribution. The GR-Gaussian method uses graph Laplacian regularization to constrain the density differences between adjacent Gaussian points, which improves spatial consistency to some extent. However, the GR-Gaussian weight function uses fixed parameters and squared distance, which is not adaptable enough to extremely sparse viewpoints and scenarios with uneven local sampling (such as X-ray tasks), resulting in unsatisfactory performance. In addition, existing methods often rely on external large models (such as the DPT depth estimation network) to provide depth priors, which increases computational cost and model complexity, and there may be inconsistencies between the external depth priors and the current 3D Gaussian representation, affecting the optimization effect.

[0004] Existing methods for synthesizing sparse new X-ray views suffer from the following key problems: First, insufficient structural information: it is difficult to obtain accurate depth constraints under sparse perspectives, leading to inaccurate 3D structure reconstruction; Second, inherent suboptimal and overfitting issues in the distribution of rendered images: due to the suboptimal nature of the distribution of rendered images under sparse perspectives, the model's memory-based fitting of limited perspective observation data leads to a decrease in rendering quality under unobserved perspectives; Third, spatial inconsistency and artifact issues: uneven density distribution leads to excessive density differences between adjacent Gaussian points, resulting in visual artifacts.

[0005] Therefore, there is an urgent need for a method for synthesizing sparse new X-ray views based on inline prior-guided 3D Gaussian splashing. This method can effectively utilize the information of the current 3D Gaussian representation itself to construct prior constraints, without relying on external large models. It can achieve high-quality new view synthesis under sparse view conditions and solve problems such as insufficient structural information, inherent suboptimality of rendered image distribution, overfitting, and spatial inconsistency. Summary of the Invention

[0006] The technical problem this invention aims to solve is to overcome the shortcomings of existing technologies for synthesizing sparse new X-ray views, such as insufficient structural information, inherent suboptimal rendering of image distribution under sparse view conditions, overfitting of models to limited observation viewpoints, and spatial inconsistencies and artifacts in the three-dimensional Gaussian density distribution. To this end, this invention proposes a method for synthesizing sparse new X-ray views based on inline prior-guided three-dimensional Gaussian splashing. This method constructs a unified inline prior mechanism within the three-dimensional Gaussian splashing framework, directly extracting prior information from the current three-dimensional Gaussian representation and volume density field. It does not rely on external large-scale pre-trained models or additional deep networks, and can still achieve high-quality synthesis of new X-ray views under conditions of very few projection viewpoints.

[0007] To achieve the above-mentioned objectives, the present invention specifically adopts the following technical solution: In a first aspect, the present invention provides a method for synthesizing new X-ray sparse views based on inline prior-guided three-dimensional Gaussian splashing, comprising the following steps: S1. Obtain initial volume data or point cloud data to construct a 3D Gaussian splash model and set the parameters of each Gaussian kernel; For the input sparse view X-ray real image, obtain the real camera pose set and sparse point cloud from the motion method through structure, and initialize the 3D Gaussian splash model based on the sparse point cloud. S2. Training a 3D Gaussian splash model; During training, the 3D Gaussian splash model is first voxelized under given voxel centers and sizes to obtain a 3D volume density field. A rendering depth map is extracted from this density field via ray casting. The depth difference between adjacent pixels in this rendering depth map is calculated to construct a depth-constrained inline prior. A pseudo-view image is generated through pseudo-camera pose sampling. Simultaneously, a real view image is generated by sampling the real camera pose geometrically closest to the pseudo-camera pose. The real view image is then transformed to the pseudo-camera pose coordinate system to obtain the transformed view image as a pseudo-label image. The pseudo-label image and the pseudo-view image are then combined using a joint mask calculation to construct a mask-reconstructed inline prior. Then… Calculate the Euclidean distance between all Gaussian point pairs, construct a K-nearest neighbor graph for each Gaussian point, and use the square of the density difference between each Gaussian point and its neighbors as the base value. Weight the base value with distance weights that decay exponentially, and average the weighted base value over all Gaussian points and their neighbors to obtain the improved graph Laplacian regularized inline prior. Finally, calculate the pixel reconstruction loss between the rendered image and the real image, and then sum the depth constraint inline prior, mask reconstruction inline prior, improved graph Laplacian regularized inline prior, and pixel reconstruction loss in a weighted sum to form the total loss function. Minimize the total loss function as the optimization objective, and iteratively update the parameters of the 3D Gaussian splash model through backpropagation. S3. In the inference phase, the camera pose of the new perspective is input into the trained 3D Gaussian splash model. Based on the learned 3D Gaussian point cloud representation, the model generates a new view image under the new perspective through a differentiable 3D Gaussian rasterization rendering pipeline, thereby achieving high-quality synthesis of new X-ray views under sparse perspective conditions.

[0008] Based on the above scheme, each step can be implemented in the following preferred manner.

[0009] As a preferred embodiment of the first aspect above, in step S2, the specific implementation of the depth constraint inline prior is as follows: given the voxel center and voxel size, the three-dimensional Gaussian splash model is mapped to a three-dimensional volume density field using the Gaussian voxelization operator; for each voxel's corresponding projection direction, the first voxel position with a density exceeding a preset density threshold is searched in the depth direction of the three-dimensional volume density field, and the depth value of the voxel position is normalized to obtain a rendering depth map; the absolute difference between adjacent pixels is calculated in the horizontal and vertical directions of the rendering depth map, and the average absolute difference is calculated in each direction, and the average values ​​calculated in the two directions are added together and multiplied by a preset scaling factor as the depth smoothing loss; if the training data used when training the three-dimensional Gaussian splash model provides a reference depth map, the Pearson correlation coefficient between the rendering depth map and the reference depth map is calculated, and 1 minus the Pearson correlation coefficient is used as the depth supervision loss, and the depth smoothing loss and the depth supervision loss are weighted and summed as the depth constraint inline prior; if no reference depth map is provided, the depth smoothing loss is directly used as the depth constraint inline prior.

[0010] As a preferred embodiment of the first aspect above, in step S2, the specific implementation of the mask reconstruction inline prior is as follows: A pseudo-camera pose is randomly selected from a preset set of pseudo-camera poses, and a real camera pose geometrically closest to the pseudo-camera pose is selected from a set of real camera poses. The 3D Gaussian splash model is then rendered under both the pseudo-camera pose and the real camera pose, resulting in pseudo-view images and real-view images. Depth maps under the pseudo-camera pose and the real camera pose are extracted through ray projection onto the 3D volume density field, resulting in pseudo-view depth maps and real-view depth maps. Based on the real-view image and the real-view... The depth map is generated by transforming the real view image into the pseudo-camera pose coordinate system using the world view transformation matrix and intrinsic parameter matrix of two camera poses. This generates a pseudo-label image, a projection mask, and a depth-consistent mask. The projection mask and the depth-consistent mask are multiplied to obtain a joint mask. The sum of the absolute differences between corresponding pixels in the pseudo-view image and the pseudo-label image is calculated in the effective region of the joint mask. This sum of absolute differences is divided by the number of effective pixels to obtain the mask reconstruction inline prior between the pseudo-label image and the pseudo-view image. The effective region of the joint mask consists of all effective pixels, and the effective pixels are the pixels with a pixel value of 1 in the joint mask.

[0011] Furthermore, mask reconstruction inline prior The calculation method is as follows: in, This represents the position coordinates of a pixel within the combined mask; Indicates the valid area of ​​the joint mask; Indicates the number of valid pixels; Indicates the pseudo-view image in Pixel value at that location, Indicates that the pseudo-label image is in The pixel value at that location.

[0012] As a preferred embodiment of the first aspect above, in step S2, the specific implementation of the improved graph Laplacian regularized inline prior is as follows: extract the coordinates and density values ​​of each Gaussian point in the three-dimensional Gaussian splash model, calculate the Euclidean distance between each pair of Gaussian points, use the K-nearest neighbor method to identify the k Gaussian points closest to each Gaussian point as neighbor points, and use the Gaussian points as nodes in the K-nearest neighbor graph. Connect the Gaussian points with their corresponding neighbor points to form edges in the K-nearest neighbor graph, thereby constructing a complete K-nearest neighbor graph, where k is a preset number of neighbor points; for each For a Gaussian point, the average Euclidean distance between it and its k neighboring points is taken to obtain the average neighborhood distance corresponding to the Gaussian point, thereby calculating the neighbor weights between the Gaussian point and its neighbors. Then, the square of the density difference between the Gaussian point and its neighbors is calculated and multiplied by the corresponding neighbor weight to obtain the weighted density difference. The weighted density differences of all the calculated values ​​are summed to obtain the weighted loss of the Gaussian point. The weighted losses of all Gaussian points are averaged and normalized, and the averaged and normalized weighted loss is used as the inline prior for the improved graph Laplacian regularization.

[0013] As a preferred embodiment of the first aspect above, the neighbor weight is calculated by an exponential function with a decay rate as the exponent; wherein the decay rate is the negative of the relative neighborhood distance of the Gaussian point, the relative neighborhood distance is the ratio of the Euclidean distance between the Gaussian point and its neighboring points to the corrected average neighborhood distance of the Gaussian point, and the corrected average neighborhood distance of the Gaussian point is the sum of its average neighborhood distance and a constant to prevent division by zero.

[0014] Furthermore, the first The nth Gaussian point and its corresponding nth Neighbor weights among neighboring points Calculate using the following formula: in, For the first The nth Gaussian point and its corresponding nth Euclidean distance between neighboring points; For the first The average neighborhood distance of a Gaussian point; To prevent division by zero of constants.

[0015] Furthermore, improve the inline prior of the graph Laplacian regularization. The calculation method is as follows: in, This represents the total number of Gaussian points. Index for Gaussian points; This represents the index of a neighboring point of a Gaussian point; Indicates by the first The set of indexes consisting of the neighbor indices of each Gaussian point; and They represent the first The Gaussian point and the first Density value of neighboring points.

[0016] As a preferred embodiment of the first aspect mentioned above, the weight coefficient of the depth-constrained inline prior in the total loss function is set to 0.05, the weight coefficient of the mask reconstruction inline prior is set to 0.02, and the weight coefficient of the improved graph Laplacian regularization inline prior is set to... .

[0017] Secondly, the present invention provides a novel X-ray sparse view synthesis system based on inline prior-guided three-dimensional Gaussian splashing, comprising: The data acquisition module is used to acquire initial volume data or point cloud data to construct a 3D Gaussian splash model and set the parameters of each Gaussian kernel. For the input sparse view X-ray real image, the module obtains the real camera pose set and sparse point cloud from the motion method through structure, and initializes the 3D Gaussian splash model based on the sparse point cloud. The model training module is used to train a 3D Gaussian splash model. During training, the 3D Gaussian splash model is first voxelized under given voxel centers and dimensions to obtain a 3D volume density field. A rendering depth map is extracted from this volume density field via ray casting. The depth difference between adjacent pixels in this rendering depth map is calculated to construct a depth-constrained inline prior. A pseudo-view image is generated through pseudo-camera pose sampling. Simultaneously, a real view image is generated by sampling the pose of the real camera that is geometrically closest to the pseudo-camera pose. The real view image is then transformed to the pseudo-camera pose coordinate system, obtaining the transformed view image as a pseudo-label image. The pseudo-label image and the pseudo-view image are then combined using a joint mask calculation to construct a mask reconstruction inline prior. The process involves several steps: First, calculating the Euclidean distance between all Gaussian point pairs and constructing a K-nearest neighbor graph for each Gaussian point. The square of the density difference between each Gaussian point and its neighbors is used as the base value. This base value is then weighted by distance weights that decay exponentially. The weighted base value is then averaged across all Gaussian points and their neighbors to obtain the improved graph Laplacian regularized inline prior. Finally, the pixel reconstruction loss between the rendered image and the real image is calculated. The total loss function is then formed by weighting and summing the depth constraint inline prior, the mask reconstruction inline prior, the improved graph Laplacian regularized inline prior, and the pixel reconstruction loss. Minimizing this total loss function is the optimization objective. The parameters of the 3D Gaussian splash model are iteratively updated through backpropagation. The view synthesis module is used in the inference stage to input the camera pose of the new viewpoint into the trained 3D Gaussian splash model. Based on the learned 3D Gaussian point cloud representation, the model generates a new view image under the new viewpoint through a differentiable 3D Gaussian rasterization rendering pipeline, thereby achieving high-quality synthesis of new X-ray views under sparse viewpoint conditions.

[0018] Thirdly, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can realize the X-ray sparse new view synthesis method based on inline prior-guided three-dimensional Gaussian splashing as described in any of the solutions of the first aspect above.

[0019] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for synthesizing new X-ray sparse views based on inline prior-guided three-dimensional Gaussian splashing as described in any of the solutions of the first aspect above.

[0020] Fifthly, the present invention provides a computer electronic device, which includes a memory and a processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the X-ray sparse new view synthesis method based on inline prior-guided three-dimensional Gaussian splashing as described in any of the above-described first aspects.

[0021] Compared with the prior art, the present invention has the following advantages: This invention introduces multiple inline priors, including depth-constrained inline priors, mask reconstruction inline priors, and improved graph Laplacian regularization inline priors, into the task of synthesizing sparse new X-ray views. Within a 3D Gaussian splashing framework, a unified inline prior-guided method is formed, aiming to address issues such as insufficient structural information, inherent suboptimality of rendered image distribution, overfitting of models to limited observation perspectives, and spatial inconsistencies and artifacts caused by uneven density distribution under sparse viewpoints. Compared to existing 3D Gaussian splashing methods, this invention significantly improves the stability and reconstruction quality of new view synthesis in X-ray scenes with extremely sparse viewpoints and locally uneven sampling, providing higher accuracy and robustness for sparse new X-ray view synthesis tasks. In this invention, each inline prior is directly extracted from the current 3D Gaussian representation and volume density field, without relying on external large-scale pre-trained models or additional deep networks, achieving high-quality X-ray new view synthesis under sparse input view conditions. This provides a novel and effective solution for X-ray new view synthesis applications under sparse viewpoint conditions in medical imaging and related fields. Attached Figure Description

[0022] Figure 1This is an overall framework diagram of the method of the present invention; Figure 2 This is a schematic diagram illustrating the calculation process of the depth-constrained inline prior in this invention; Figure 3 This is a schematic diagram illustrating the calculation process of inline prior for mask reconstruction in this invention; Figure 4 This is a schematic diagram illustrating the calculation process of the improved graph Laplace regularization inline prior in this invention; Figure 5 This is a flowchart illustrating the training and inference process in an embodiment of the present invention; Figure 6 This is a schematic diagram of a real image in an embodiment of the present invention; Figure 7 This is a schematic diagram of the synthesized image in an embodiment of the present invention; Figure 8 This is a system block diagram of the present invention; Figure 9 This is a schematic diagram of the components of a computer electronic device according to the present invention. Detailed Implementation

[0023] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.

[0024] In the description of this invention, it should be understood that the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include at least one of those features.

[0025] like Figure 1 As shown, in a preferred embodiment of the present invention, the above-mentioned method for synthesizing new X-ray sparse views based on inline prior-guided three-dimensional Gaussian splashing includes the following steps S1 to S3. The specific implementation process of each step will be described in detail below.

[0026] S1. Obtain initial volume data or point cloud data to construct a 3D Gaussian splash model and set the parameters of each Gaussian kernel; for the input sparse view X-ray real image, obtain the real camera pose set and sparse point cloud through the Structure from Motion (SfM) method, and initialize the 3D Gaussian splash model based on the sparse point cloud.

[0027] It should be noted that in step S1, the parameters of the Gaussian kernel include position, scale, rotation, and density. The Structure from Motion (SfM) method is existing technology; similar methods can also be used to obtain the real camera pose and sparse point cloud. During the initialization of the 3D Gaussian splash model, the position parameter of each Gaussian point is determined by the point cloud coordinates, the scale parameter is initialized to a small random value, the rotation parameter is initialized to a unit quaternion, and the density parameter is initialized to a preset initial value.

[0028] S2. Training a 3D Gaussian splash model; During training, the 3D Gaussian splash model is first voxelized under given voxel centers and sizes to obtain a 3D volume density field. A rendering depth map is extracted from this density field via ray casting. The depth difference between adjacent pixels in this rendering depth map is calculated to construct a depth-constrained inline prior. A pseudo-view image is generated through pseudo-camera pose sampling. Simultaneously, a real view image is generated by sampling the real camera pose geometrically closest to the pseudo-camera pose. The real view image is then transformed to the pseudo-camera pose coordinate system to obtain the transformed view image as a pseudo-label image. The pseudo-label image and the pseudo-view image are then combined using a joint mask calculation to construct a mask-reconstructed inline prior. Then… The Euclidean distance between all Gaussian point pairs is calculated, and a K-nearest neighbor graph is constructed for each Gaussian point. The square of the density difference between each Gaussian point and its neighbors is used as the base value. The base value is weighted by distance weights that decay exponentially, and the weighted base value is averaged over all Gaussian points and their neighbors to obtain the improved graph Laplacian regularized inline prior. Finally, the pixel reconstruction loss between the rendered image and the real image is calculated. The total loss function is composed of the weighted sum of the depth constraint inline prior, the mask reconstruction inline prior, the improved graph Laplacian regularized inline prior, and the pixel reconstruction loss. The parameters of the 3D Gaussian splash model are iteratively updated through backpropagation with the goal of minimizing the total loss function.

[0029] It should be noted that in step S2, the specific implementation of the depth constraint inline prior is as follows: given the voxel center and voxel size, the 3D Gaussian splash model is mapped to a 3D volume density field using the Gaussian voxelization operator; for the projection direction corresponding to each voxel, the first voxel position with a density exceeding a preset density threshold is searched in the depth direction of the 3D volume density field, and the depth value of the voxel position is normalized to obtain a rendering depth map; the absolute difference between adjacent pixels is calculated in the horizontal and vertical directions of the rendering depth map, and the average absolute difference is calculated in each direction, and the average values ​​calculated in the two directions are added together and multiplied by a preset scaling factor as the depth smoothing loss; if the training data used when training the 3D Gaussian splash model provides a reference depth map, the Pearson correlation coefficient between the rendering depth map and the reference depth map is calculated, and 1 minus the Pearson correlation coefficient is used as the depth supervision loss, and the depth smoothing loss and the depth supervision loss are weighted and summed as the depth constraint inline prior; if no reference depth map is provided, the depth smoothing loss is directly used as the depth constraint inline prior.

[0030] In this embodiment, as Figure 2 As shown, the Gaussian voxelization operator is used to map the 3D Gaussian splash model to a 3D volume density field with a given voxel center and voxel size. In the generated 3D volume density field, a rendering depth map is extracted through sampling, accumulation, and depth output steps, and the depth constraint inline prior is calculated and fed back to the 3D Gaussian splash model. Specifically, the voxelization process in this embodiment is implemented through a query function. For a given voxel center position, number of voxels, and voxel size, the density contribution of all Gaussian points to each voxel position is calculated to form a 3D volume density field.

[0031] For the projection direction of each voxel, a layer-by-layer search is performed along the depth dimension in the depth direction of the 3D volume density field to find the first voxel position whose density exceeds a preset density threshold. The depth value at this position (i.e., the depth dimension index divided by the total number of depth dimensions) is normalized to the range [0,1] to obtain the rendered depth map. Based on this rendered depth map, the depth smoothing loss is calculated to enhance the continuity and stability of the 3D structure without relying on an external depth network. Specifically, for the rendered depth map... Its height and width are respectively and The absolute difference between adjacent pixels in the horizontal direction represents the depth difference between adjacent columns, while the absolute difference between adjacent pixels in the vertical direction represents the depth difference between adjacent rows. This embodiment further averages the absolute differences in the horizontal and vertical directions, and then adds the two averages together with a scaling factor. Multiplication, as a deep smoothing loss The details are as follows: in, It is the average of the absolute differences in the horizontal direction; It is the average of the absolute differences in the vertical direction; Indicates the rendering depth map at the th Line number The depth value at the column; Indicates the rendering depth map at the th Line number The depth value at the column; Indicates the rendering depth map at the th Line number The depth value at the column.

[0032] Furthermore, when the training data includes a reference depth map (ground truth depth), a depth-supervised loss is calculated between the rendered depth map and the reference depth map. Specifically, the rendered depth map and the reference depth map are flattened into one-dimensional vectors, and the Pearson correlation coefficient between them is calculated. The difference between 1 and the correlation coefficient is used as the depth-supervised loss to constrain the consistency between the rendered depth map and the reference depth map. In this case, the depth-supervised loss and the depth smoothing loss together constitute the depth-constrained inline prior. Both are multiplied by their corresponding weight coefficients and added to the total loss to address the problem of insufficient structural information in sparse viewpoints. It is important to note that the depth-supervised loss is only calculated when the training data provides a reference depth map. If the training data does not contain a reference depth map, only the depth smoothing loss is used for self-supervised constraints.

[0033] In this embodiment, the voxel center position is obtained by random sampling of the bounding box and voxel size; the number of voxels is set to 64×64×64; the voxel size is adaptively determined according to the scene bounding box; the density threshold is set to 0.01; and the above scaling factor is set to 0.1.

[0034] It should be noted that in step S2, the specific implementation of the mask reconstruction inline prior is as follows: A pseudo-camera pose is randomly selected from a preset set of pseudo-camera poses, and a real camera pose geometrically closest to the pseudo-camera pose is selected from a set of real camera poses. The 3D Gaussian splash model is then rendered under both the pseudo-camera pose and the real camera pose, resulting in pseudo-view images and real-view images. Depth maps under the pseudo-camera pose and the real camera pose are extracted through ray projection onto the 3D volume density field, resulting in pseudo-view depth maps and real-view depth maps. Based on the real-view image and the real-view depth map... Using the world view transformation matrix and intrinsic parameter matrix of two camera poses, the real view image is transformed into the pseudo camera pose coordinate system, generating a pseudo label image, a projection mask, and a depth-consistent mask. The projection mask and the depth-consistent mask are multiplied to obtain a joint mask. The sum of the absolute differences between corresponding pixels of the pseudo view image and the pseudo label image is calculated in the effective region of the joint mask. The sum of the absolute differences is divided by the number of effective pixels to obtain the mask reconstruction inline prior between the pseudo label image and the pseudo view image. The effective region of the joint mask consists of all effective pixels, and the effective pixels are the pixels with a pixel value of 1 in the joint mask.

[0035] In this embodiment, based on the depth-constrained inline prior, to address the inherent suboptimal nature of the rendered image distribution under sparse viewpoint conditions and the overfitting of the 3D Gaussian to a finite viewpoint, this invention constructs a mask reconstruction inline prior. For example... Figure 3 As shown, specifically, pseudo-camera poses are sampled around the real camera pose, and the real camera pose with the closest geometric distance to the pseudo-camera pose is selected. A 3D Gaussian splash model is rendered from both the pseudo-camera pose and the real camera pose, obtaining view images and their corresponding depth maps in both poses. The aforementioned process has already voxelized the 3D Gaussian splash model to obtain a 3D volume density field. Based on this, depth maps corresponding to the pseudo-camera pose and the real camera pose are extracted from this volume density field through ray casting. The ray casting process involves emitting rays from the camera center along the projection direction of each voxel, searching layer by layer along the depth dimension in the volume density field. When the first position with a density exceeding a preset threshold is found, the normalized depth value at that position is recorded.

[0036] Furthermore, if the extracted depth map size does not match the corresponding image size, bilinear interpolation is used to adjust the depth map to match its corresponding image size: bilinear interpolation treats the depth map as a single-channel image and adjusts the depth map to the target size, ensuring that the depth map and the image size are consistent.

[0037] After obtaining the pseudo-view image, real view image, pseudo-view depth map, and real view depth map, the real view image is transformed into the coordinate system of the pseudo-camera pose using the real view image and its depth map, as well as the world view transformation matrix and intrinsic parameter matrix of the two camera poses, to generate the transformed view image (i.e., pseudo-label image), projection mask, and depth-consistent mask.

[0038] The coordinate transformation process is achieved through the following steps: For each pixel coordinate in the pseudo-camera pose, the depth information is used to project it back into 3D space, and then the transformation matrix of the real camera pose is used to project it onto the image plane of the real camera pose. Bilinear interpolation is then used to obtain the corresponding pixel value, generating the transformed view image. Specifically, in this embodiment, the pixel coordinates in the pseudo-camera pose are first converted to homogeneous coordinates. The inverse of the intrinsic parameter matrix is ​​used to back-project the pixel coordinates onto a 3D point in the pseudo-camera coordinate system. Then, the extrinsic parameter matrix of the pseudo-camera pose is used to transform the 3D point to the world coordinate system. Finally, the extrinsic parameter matrix of the real camera pose is used to transform the world coordinate point to the real camera coordinate system, and the intrinsic parameter matrix is ​​used to project it onto the image plane to obtain the corresponding pixel coordinates. The intrinsic parameter matrix is ​​calculated using the camera's field of view (FoV) and image size, while the focal length can be calculated based on the image width, height, and field of view.

[0039] The projection mask is used to identify the effective region during the deformation process, i.e., whether the projected pixel coordinates are within the image boundaries (coordinate values ​​are within the image width and height range). The depth-consistent mask is used to identify regions with consistent depth information. By comparing the difference between the expected depth value calculated during the deformation process and the actual depth value, regions are considered to have consistent depth when the difference is less than 10% of the expected depth. Furthermore, this invention multiplies the projection mask and the depth-consistent mask to obtain a joint mask, and uses this mask to calculate and reconstruct the inline prior. This is used to constrain the intensity distribution and detail reconstruction of rendered images under sparse camera poses. in, This represents the position coordinates of a pixel within the combined mask; Indicates the valid area of ​​the joint mask; Indicates the number of valid pixels; Indicates the pseudo-view image in Pixel value at that location, Indicates that the pseudo-label image is in The pixel value at that location.

[0040] It is important to note that the inline prior for mask reconstruction described above is calculated only within the effective area of ​​the joint mask. Pixels with a mask value of 0 are not included in the loss calculation, thus ensuring that supervised learning is performed only in geometrically and depth-consistent regions.

[0041] In this embodiment, the pseudo camera pose is generated by adding random noise around the real camera pose, with the noise standard deviation set to 0.05. The number of pseudo camera poses is set to 50, and the above-mentioned mask reconstruction inline prior is introduced after the number of training iterations reaches 2000.

[0042] It should be noted that in step S2, the specific implementation of the improved graph Laplacian regularized inline prior is as follows: extract the coordinates and density values ​​of each Gaussian point in the three-dimensional Gaussian splash model, calculate the Euclidean distance between each pair of Gaussian points, and use the k-nearest neighbor method. The K-Nearest Neighbors (KNN) algorithm takes the k nearest Gaussian points as neighbors for each Gaussian point and treats each Gaussian point as a node in the K-Nearest Neighbors graph. Edges are formed by connecting each Gaussian point to its corresponding neighbors, thus constructing a complete K-Nearest Neighbors graph. Here, k is a preset number of neighboring points. For each Gaussian point, the average Euclidean distance between it and its k neighbors is taken to obtain the average neighborhood distance, thereby calculating the neighbor weights between the Gaussian point and its neighbors. Then, the square of the density difference between the Gaussian point and its neighbors is calculated and multiplied by the corresponding neighbor weight to obtain the weighted density difference. The sum of all calculated weighted density differences yields the weighted loss of the Gaussian point. The weighted losses of all Gaussian points are then averaged and normalized, and the averaged and normalized weighted loss is used as the improved graph Laplacian regularization inline prior.

[0043] Furthermore, the neighbor weights are calculated using an exponential function with a decay rate as the exponent; where the decay rate is the negative of the relative neighborhood distance of the Gaussian point, the relative neighborhood distance is the ratio of the Euclidean distance between the Gaussian point and its neighboring points to the corrected average neighborhood distance of the Gaussian point, and the corrected average neighborhood distance of the Gaussian point is the sum of its average neighborhood distance and a constant to prevent division by zero.

[0044] In this embodiment, to address the spatial inconsistency and artifact problems caused by uneven density distribution, the present invention introduces an improved graph Laplacian regularized inline prior. For example... Figure 4 As shown, starting from the Gaussian point set of the 3D Gaussian splash model, a K-nearest neighbor graph is constructed through KNN adjacency relationship calculation, edge weight (neighbor weight) calculation, and neighborhood density difference calculation. Then, smoothing terms (neighbor proximity), weight control (far point weakening), and improved graph Laplacian regularization are performed to inline prior calculations. Finally, a spatially smooth Gaussian field is generated and fed back to the model to solve the spatial inconsistency and artifact problems caused by uneven density distribution.

[0045] In this embodiment, before calculating the improved graph Laplacian regularized inline prior, the neighbor weights need to be calculated first. The closer the Euclidean distance between two points, the greater the neighbor weight, and the neighbor weight decreases exponentially as the Euclidean distance increases relative to the average neighborhood distance. Specifically, the... The nth Gaussian point and its corresponding nth Neighbor weights among neighboring points Calculate using the following formula: in, For the first The nth Gaussian point and its corresponding nth Euclidean distance between neighboring points; For the first The average neighborhood distance of a Gaussian point; To prevent division by zero, constants are usually taken as minimum values.

[0046] Therefore, the improved Graph Laplace regularization inline prior is obtained. The calculation method is as follows: in, This represents the total number of Gaussian points. Index for Gaussian points; This represents the index of a neighboring point of a Gaussian point; Indicates by the first The set of indexes consisting of the neighbor indices of each Gaussian point; and They represent the first The Gaussian point and the first Density values ​​(opacity) of neighboring points.

[0047] The weights of the squared distance and fixed-scale parameters used in GRGaussian Furthermore, the weighted losses of all Gaussian points and their neighbors are directly summed in the loss function to design a graph Laplacian regularization loss. Compared with the graph Laplacian regularization loss used in GRGaussian, the improvement of this invention is mainly reflected in the design of the aforementioned neighbor weights. to replace weights Furthermore, the weighted loss of all Gaussian points and their neighbors in the loss function is calculated by dividing by... The average normalization method is used to enable the weights to automatically adapt to the local density scale of different regions, and the loss value does not change with the number of Gaussian points. It shows better stability and effectiveness in X-ray tasks with extremely sparse viewpoints and uneven local sampling.

[0048] Among them, the weights in GRGaussian The calculation method is as follows: It should be noted that in step S2, the pixel reconstruction loss is the basic loss of 3DGS. The ground truth image involved in this loss is the sparse viewpoint X-ray real image provided in the training data, and the rendered image is the image generated by the 3D Gaussian splash model from the specified camera viewpoint according to the current parameters. In this embodiment, the above-mentioned pixel reconstruction loss... It is composed of a weighted combination of L1 loss and structural similarity (SSIM) loss between the real image and the rendered image, where L1 loss is... weight and structural similarity loss weight The sum is 1, specifically represented as: In this embodiment, Set to 0.8, Set it to 0.2.

[0049] Therefore, the total loss function adopted in this invention It is expressed as follows: in, , , , All are weighting coefficients. In this embodiment, the weighting coefficients of each item in the total loss function are set as follows: the weighting coefficient of the depth-constrained inline prior is set to 0.05, the weighting coefficient of the mask reconstruction inline prior is set to 0.02, and the weighting coefficient of the improved graph Laplacian regularization inline prior is set to... Furthermore, all the aforementioned weight coefficients can be adjusted within a range not exceeding one order of magnitude to adapt to the needs of new view synthesis under different datasets and different sparse perspectives. Depth-constrained inline priors, mask reconstruction inline priors, and improved graph Laplacian regularization inline priors together constitute the inline prior guidance framework. All inline priors are constructed using the geometric and density information of the current 3D Gaussian set itself, without relying on external large-scale pre-trained models or external deep networks.

[0050] Additionally, it should be noted that in each rendering iteration of the training process, a given Drop ratio is used as a parameter, which can be adjusted within the range [0,1]. For each Gaussian point in the 3D Gaussian splash model, a random number uniformly distributed within the range [0,1] is generated. If the random number is greater than the Drop ratio, the mask value is 1; otherwise, the mask value is 0, so that it does not contribute to this rendering, thus generating a Bernoulli random mask. This mask is multiplied by the density parameter to achieve random deactivation of some Gaussian points during rendering. By randomly deactivating some Gaussian points, the memory-based fitting of the 3D Gaussian model to the limited viewpoint observation data is suppressed, thereby alleviating the overfitting problem. This process is performed independently in each iteration to ensure that randomness is introduced during training and to prevent the model from overfitting the limited observation viewpoint. In this embodiment, the Drop ratio is set to 0.02, that is, approximately 2% of Gaussian points are randomly deactivated in each rendering.

[0051] S3. In the inference phase, the camera pose of the new perspective is input into the trained 3D Gaussian splash model. Based on the learned 3D Gaussian point cloud representation, the model generates a new view image under the new perspective through a differentiable 3D Gaussian rasterization rendering pipeline, thereby achieving high-quality synthesis of new X-ray views under sparse perspective conditions.

[0052] It should also be noted that in step S3, there is no need to input a new X-ray image with a sparse viewpoint. Real-time rendering of any new viewpoint can be achieved solely by relying on the 3D Gaussian point cloud representation information learned during the training phase, which includes attributes such as position, covariance, color, and opacity.

[0053] The present invention will now demonstrate the application effect of the X-ray sparse new view synthesis method based on inline prior guided 3D Gaussian splashing described in S1~S3 of the above embodiments on a specific dataset, so as to facilitate understanding of the essence of the present invention.

[0054] Example The overall process of the X-ray sparse new view synthesis method based on inline prior-guided 3D Gaussian splashing used in this embodiment can be divided into three stages: data preprocessing, model training, and image prediction, as detailed below. Figure 5 As shown.

[0055] 1. Data Preprocessing Stage Step 1: Preprocess the raw X-ray images with sparse perspectives. Obtain X-ray datasets containing five different organs (chest, foot, head, abdomen, and pancreas), with each organ containing X-ray images from multiple perspectives. Perform preprocessing operations such as size normalization and grayscale value normalization on the raw X-ray images to ensure that the image data format is uniform and facilitates subsequent processing.

[0056] Step 2: Obtain camera pose and initial point cloud. A NAF format data file is used, containing X-ray images, CT scanner geometric configuration parameters, and projection angle information. The camera pose is calculated using geometric transformations based on the CT scanner's geometric configuration parameters and projection angles to obtain the camera's rotation matrix and translation vector. The initial sparse point cloud is loaded from a pre-generated initial point cloud file, which contains the point cloud's 3D coordinates and density information. In this embodiment, X-ray images from three sparse viewpoints are used as training input for each organ.

[0057] 2. Model Training Phase Step 1: Construct the training dataset and process it in batches. The preprocessed X-ray images and corresponding camera pose data are used to construct the training dataset, which is then divided into batches according to a fixed batch size, totaling several parts. One batch.

[0058] Step 2, select the batch index in sequence. A batch of training samples is used to train a 3D Gaussian splash model guided by inline priors. During training, for each training sample, pixel reconstruction loss, depth constraint inline prior, mask reconstruction inline prior, and improved graph Laplacian regularization inline prior are calculated sequentially. Based on the total loss of all training samples in the batch, the network parameters of the entire model are adjusted using the backpropagation algorithm, updating the position, scale, rotation, and density parameters of the 3D Gaussian set. The training process continues until all batches in the training dataset have participated in the training, and the model converges after a specified number of iterations, indicating that training is complete.

[0059] 3. Image Prediction The camera poses corresponding to the X-ray images in the test set are used as input and directly fed into the trained 3D Gaussian splash model. The model uses the trained 3D Gaussian set and a 3D Gaussian splash rendering algorithm to generate X-ray images from the given test camera poses, thus achieving new view synthesis.

[0060] In this embodiment, the test results are as follows: Figure 6 and Figure 7 As shown, this invention introduces multiple inline priors, such as depth constraints, mask reconstruction, and improved graph Laplacian regularization, to form a unified inline prior-guided method within a 3D Gaussian splashing framework. The resulting synthesized new view is highly consistent with real X-ray images in terms of structural details, bone density distribution, and visual quality. It accurately reproduces the shape, size, and relative position of bones, with clearly visible joint spaces and no obvious artifacts, demonstrating that this invention can generate realistic and structurally accurate new X-ray views.

[0061] Compared with existing methods, this invention significantly improves the stability and reconstruction quality of new view synthesis under extremely sparse view conditions, providing a new and effective solution for the application of X-ray new view synthesis under sparse view conditions in medical imaging and related fields, and providing strong technical support for possible future applications such as medical image diagnosis assistance and surgical planning visualization.

[0062] It should also be noted that the X-ray sparse new view synthesis method based on inline prior-guided 3D Gaussian splashing in the above embodiments can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a X-ray sparse new view synthesis system based on inline prior-guided 3D Gaussian splashing, corresponding to the X-ray sparse new view synthesis method based on inline prior-guided 3D Gaussian splashing provided in the above embodiments, such as... Figure 8 As shown, it includes: The data acquisition module is used to acquire initial volume data or point cloud data to construct a 3D Gaussian splash model and set the parameters of each Gaussian kernel. For the input sparse view X-ray real image, the module obtains the real camera pose set and sparse point cloud from the motion method through structure, and initializes the 3D Gaussian splash model based on the sparse point cloud. The model training module is used to train a 3D Gaussian splash model. During training, the 3D Gaussian splash model is first voxelized under given voxel centers and dimensions to obtain a 3D volume density field. A rendering depth map is extracted from this volume density field via ray casting. The depth difference between adjacent pixels in this rendering depth map is calculated to construct a depth-constrained inline prior. A pseudo-view image is generated through pseudo-camera pose sampling. Simultaneously, a real view image is generated by sampling the pose of the real camera that is geometrically closest to the pseudo-camera pose. The real view image is then transformed to the pseudo-camera pose coordinate system, obtaining the transformed view image as a pseudo-label image. The pseudo-label image and the pseudo-view image are then combined using a joint mask calculation to construct a mask reconstruction inline prior. The process involves several steps: First, calculating the Euclidean distance between all Gaussian point pairs and constructing a K-nearest neighbor graph for each Gaussian point. The square of the density difference between each Gaussian point and its neighbors is used as the base value. This base value is then weighted by distance weights that decay exponentially. The weighted base value is then averaged across all Gaussian points and their neighbors to obtain the improved graph Laplacian regularized inline prior. Finally, the pixel reconstruction loss between the rendered image and the real image is calculated. The total loss function is then formed by weighting and summing the depth constraint inline prior, the mask reconstruction inline prior, the improved graph Laplacian regularized inline prior, and the pixel reconstruction loss. Minimizing this total loss function is the optimization objective. The parameters of the 3D Gaussian splash model are iteratively updated through backpropagation. The view synthesis module is used in the inference stage to input the camera pose of the new viewpoint into the trained 3D Gaussian splash model. Based on the learned 3D Gaussian point cloud representation, the model generates a new view image under the new viewpoint through a differentiable 3D Gaussian rasterization rendering pipeline, thereby achieving high-quality synthesis of new X-ray views under sparse viewpoint conditions.

[0063] It is understood that the X-ray sparse new view synthesis method based on inline prior-guided three-dimensional Gaussian splashing described in S1~S3 above can essentially be implemented by a computer program. Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer program product corresponding to the X-ray sparse new view synthesis method based on inline prior-guided three-dimensional Gaussian splashing provided in the above embodiments. This product includes a computer program / instructions that, when executed by a processor, can implement the X-ray sparse new view synthesis method based on inline prior-guided three-dimensional Gaussian splashing as described in the above embodiments.

[0064] Similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the X-ray sparse new view synthesis method based on inline prior-guided three-dimensional Gaussian splashing provided in the above embodiments, such as... Figure 9 As shown, it includes a memory and a processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the X-ray sparse new view synthesis method based on inline prior-guided three-dimensional Gaussian splashing in the above embodiments.

[0065] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0066] Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer-readable storage medium corresponding to the X-ray sparse new view synthesis method based on inline prior guided three-dimensional Gaussian splashing provided in the above embodiments. The storage medium stores a computer program, which, when executed by a processor, can realize the X-ray sparse new view synthesis method based on inline prior guided three-dimensional Gaussian splashing in the above embodiments.

[0067] Specifically, in the computer-readable storage medium of the above three embodiments, the stored computer program is executed by a processor, which can perform the aforementioned steps S1 to S3.

[0068] It is understood that the aforementioned storage media may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage media may also be various media capable of storing program code, such as USB flash drives, external hard drives, magnetic disks, or optical discs.

[0069] It is understood that the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0070] It should also be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. In the embodiments provided in this application, the division of steps or modules in the system and method is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.

[0071] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A method for synthesizing new X-ray sparse views based on inline prior-guided 3D Gaussian splashing, characterized in that, Includes the following steps: S1. Obtain initial volume data or point cloud data to construct a 3D Gaussian splash model and set the parameters of each Gaussian kernel; For the input sparse viewpoint X-ray real image, obtain the real camera pose set and sparse point cloud from the motion method through structure, and initialize the 3D Gaussian splash model based on the sparse point cloud. S2. Training the 3D Gaussian splash model; During training, the 3D Gaussian splash model is first voxelized under given voxel centers and dimensions to obtain a 3D volume density field. A rendering depth map is then extracted from this density field via ray casting. The depth difference between adjacent pixels in this rendering depth map is calculated to construct a depth-constrained inline prior. A pseudo-view image is generated through pseudo-camera pose sampling. Simultaneously, a real view image is generated by sampling the pose of the real camera that is geometrically closest to the pseudo-camera pose. The real view image is then transformed to the pseudo-camera pose. In the pose coordinate system, the transformed view image is obtained as a pseudo-label image. The pseudo-label image and the pseudo-view image are used to construct a mask reconstruction inline prior by calculating a joint mask. Then, the Euclidean distance between all Gaussian point pairs is calculated, and a K-nearest neighbor graph is constructed for each Gaussian point. The square of the density difference between each Gaussian point and its neighbors is used as the base value. The base value is weighted by the distance weights that decay exponentially between the two. The weighted base value is averaged over all Gaussian points and their neighbors to obtain the improved graph Laplacian regularized inline prior. Finally, the pixel reconstruction loss between the rendered image and the real image is calculated. The total loss function is formed by weighted summation of the depth constraint inline prior, the mask reconstruction inline prior, the improved graph Laplacian regularization inline prior, and the pixel reconstruction loss. The parameters of the 3D Gaussian splash model are iteratively updated by backpropagation with the goal of minimizing the total loss function. S3. In the inference phase, the camera pose of the new perspective is input into the trained 3D Gaussian splash model. Based on the learned 3D Gaussian point cloud representation, the model generates a new view image under the new perspective through a differentiable 3D Gaussian rasterization rendering pipeline, thereby achieving high-quality synthesis of new X-ray views under sparse perspective conditions.

2. The method for synthesizing new X-ray sparse views based on inline prior-guided 3D Gaussian splashing as described in claim 1, characterized in that, In step S2, the specific implementation of the depth constraint inline prior is as follows: given the voxel center and voxel size, the three-dimensional Gaussian splash model is mapped to a three-dimensional volume density field using the Gaussian voxelization operator; for each voxel's corresponding projection direction, the first voxel position with a density exceeding a preset density threshold is searched in the depth direction of the three-dimensional volume density field, and the depth value of the voxel position is normalized to obtain a rendering depth map; the absolute difference between adjacent pixels is calculated in the horizontal and vertical directions of the rendering depth map, and the average absolute difference is calculated in each direction, and the average values ​​calculated in the two directions are added together and multiplied by a preset scaling factor as the depth smoothing loss; If the training data used to train the 3D Gaussian splash model provides a reference depth map, then the Pearson correlation coefficient between the rendered depth map and the reference depth map is calculated. The Pearson correlation coefficient is subtracted from 1 as the depth supervision loss, and the depth smoothing loss and the depth supervision loss are weighted and summed as the depth constraint inline prior. If no reference depth map is provided, then the depth smoothing loss is directly used as the depth constraint inline prior.

3. The method for synthesizing new X-ray sparse views based on inline prior-guided 3D Gaussian splashing as described in claim 2, characterized in that, In step S2, the specific implementation of the mask reconstruction inline prior is as follows: randomly select a pseudo camera pose from the preset set of pseudo camera poses, and select a real camera pose that is geometrically closest to the pseudo camera pose from the set of real camera poses, so that the three-dimensional Gaussian splash model is rendered under the pseudo camera pose and the real camera pose respectively, and the pseudo view image and the real view image are obtained accordingly. By ray projection onto a three-dimensional volume density field, depth maps under pseudo-camera pose and real camera pose are extracted respectively, resulting in pseudo-view depth map and real view depth map. Based on the real view image and the real view depth map, the real view image is transformed into the pseudo camera pose coordinate system using the world view transformation matrix and intrinsic parameter matrix of two camera poses. This generates a pseudo label image, a projection mask, and a depth-consistent mask. The projection mask and the depth-consistent mask are multiplied to obtain a joint mask. The sum of the absolute differences between corresponding pixels in the pseudo view image and the pseudo label image is calculated in the effective region of the joint mask. This sum of absolute differences is divided by the number of effective pixels to obtain the mask reconstruction inline prior between the pseudo label image and the pseudo view image. The effective region of the joint mask consists of all effective pixels, and the effective pixels are the pixels with a pixel value of 1 in the joint mask.

4. The method for synthesizing new X-ray sparse views based on inline prior-guided 3D Gaussian splashing as described in claim 1, characterized in that, In step S2, the specific implementation of the improved graph Laplacian regularized inline prior is as follows: extract the coordinates and density values ​​of each Gaussian point in the 3D Gaussian splash model, calculate the Euclidean distance between each pair of Gaussian points, use the K-nearest neighbor method to identify the k Gaussian points closest to each Gaussian point as neighbors, and use the Gaussian points as nodes in the K-nearest neighbor graph. Connect the Gaussian point with its corresponding neighbors to form edges in the K-nearest neighbor graph, thereby constructing a complete K-nearest neighbor graph, where k is the preset number of neighbor points; for each Gaussian point, its... The mean Euclidean distances to the k neighboring points are taken to obtain the average neighborhood distance of the Gaussian point, thereby calculating the neighbor weights between the Gaussian point and its neighbors. Then, the square of the density difference between the Gaussian point and its neighbors is calculated and multiplied by the corresponding neighbor weight to obtain the weighted density difference. All the calculated weighted density differences are summed to obtain the weighted loss of the Gaussian point. The weighted losses of all Gaussian points are averaged and normalized, and the averaged and normalized weighted loss is used as the inline prior for the improved graph Laplacian regularization.

5. The method for synthesizing new X-ray sparse views based on inline prior-guided 3D Gaussian splashing as described in claim 4, characterized in that, The neighbor weights are calculated using an exponential function with a decay rate as the exponent; where the decay rate is the negative of the relative neighborhood distance of the Gaussian point, the relative neighborhood distance is the ratio of the Euclidean distance between the Gaussian point and its neighboring points to the corrected average neighborhood distance of the Gaussian point, and the corrected average neighborhood distance of the Gaussian point is the sum of its average neighborhood distance and a constant to prevent division by zero.

6. The method for synthesizing new X-ray sparse views based on inline prior-guided 3D Gaussian splashing as described in claim 1, characterized in that, In the total loss function, the weight coefficients for the depth-constrained inline prior are set to 0.05, the weight coefficients for the mask reconstruction inline prior are set to 0.02, and the weight coefficients for the improved graph Laplacian regularization inline prior are set to... .

7. A novel X-ray sparse view synthesis system based on inline prior-guided 3D Gaussian splashing, characterized in that, include: The data acquisition module is used to acquire initial volume data or point cloud data to construct a 3D Gaussian splash model and set the parameters of each Gaussian kernel. For the input sparse view X-ray real image, the module obtains the real camera pose set and sparse point cloud from the motion method through structure, and initializes the 3D Gaussian splash model based on the sparse point cloud. The model training module is used to train a 3D Gaussian splash model. During training, the 3D Gaussian splash model is first voxelized at a given voxel center and size to obtain a 3D volume density field. A rendering depth map is then extracted from this volume density field via ray casting. Depth constraint inline priors are constructed by calculating the depth differences between adjacent pixels in this rendering depth map. A pseudo-view image is generated through pseudo-camera pose sampling, and a real view image is generated by sampling the pose of the real camera that is geometrically closest to the pseudo-camera pose. The real view image is then transformed to... In the pseudo-camera pose coordinate system, the transformed view image is obtained as the pseudo-label image. The pseudo-label image and the pseudo-view image are used to construct the mask reconstruction inline prior by calculating the joint mask. Then, the Euclidean distance between all Gaussian point pairs is calculated, and a K-nearest neighbor graph is constructed for each Gaussian point. The square of the density difference between each Gaussian point and its neighbors is used as the base value. The base value is weighted by the distance weights that decay exponentially between the two. The weighted base value is averaged over all Gaussian points and their neighbors to obtain the improved graph Laplacian regularized inline prior. Finally, the pixel reconstruction loss between the rendered image and the real image is calculated. The total loss function is formed by weighted summation of the depth constraint inline prior, the mask reconstruction inline prior, the improved graph Laplacian regularization inline prior, and the pixel reconstruction loss. The parameters of the 3D Gaussian splash model are iteratively updated by backpropagation with the goal of minimizing the total loss function. The view synthesis module is used in the inference stage to input the camera pose of the new viewpoint into the trained 3D Gaussian splash model. Based on the learned 3D Gaussian point cloud representation, the model generates a new view image under the new viewpoint through a differentiable 3D Gaussian rasterization rendering pipeline, thereby achieving high-quality synthesis of new X-ray views under sparse viewpoint conditions.

8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it can realize the method for synthesizing new X-ray sparse views based on inline prior-guided three-dimensional Gaussian splashing as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method for synthesizing new X-ray sparse views based on inline prior-guided three-dimensional Gaussian splashing as described in any one of claims 1 to 6.

10. A computer electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the method for synthesizing new X-ray sparse views based on inline prior-guided three-dimensional Gaussian splashing as described in any one of claims 1 to 6.