A method for disparity control and editing of stereoscopic video based on neural radiance fields

By adopting a stereoscopic video conversion method based on neural radiation fields, the problems of inaccurate parallax calculation and visual fatigue are solved, and high-quality, visually comfortable and artistically effective stereoscopic video generation is achieved, improving the generation efficiency and artistic expression of stereoscopic videos.

CN115482323BActive Publication Date: 2026-08-04SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI UNIV
Filing Date
2022-08-10
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing stereoscopic video conversion methods suffer from inaccurate parallax calculation, severe visual fatigue, and insufficient local parallax editing capabilities, making it difficult to generate high-quality, visually comfortable, and artistically appealing stereoscopic videos.

Method used

A neural radiation field-based approach is adopted to construct a 4D spatiotemporal dynamic neural radiation field, adaptively control parallax, and realize local parallax editing during stereo rendering. High-quality view reconstruction and local object separation are achieved by using multilayer perceptron and visual saliency analysis.

Benefits of technology

It achieves the generation of high-quality stereoscopic videos with precise parallax control and high visual comfort. It can also perform local parallax editing, which enhances the artistic effect and viewing experience of stereoscopic videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115482323B_ABST
    Figure CN115482323B_ABST
Patent Text Reader

Abstract

This invention relates to a method for parallax control and editing in stereoscopic video based on neural radiation fields. First, a neural radiation field with bidirectional temporal flow is introduced to generate a dynamic video field with a new perspective. Second, the ideal parallax is adaptively and accurately calculated based on viewing conditions and video scene characteristics, generating stereoscopic videos with significant stereoscopic effects and visual comfort. Finally, the parallax of individual objects is re-edited during the stereoscopic rendering process based on the neural radiation field. Compared with existing technologies, this method achieves the highest overall performance in image reconstruction quality metrics. Simultaneously, this framework achieves a lower visual fatigue index and a stereoscopic effect including both positive and negative parallax. Experiments collected feedback from 10 non-professional viewers and 10 professional filmmakers regarding the results of this method, demonstrating the value of our framework in optimizing visual experience and artistic expression in 3D film production from three perspectives: stereoscopic effect, comfort, and local parallax editing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of stereoscopic video conversion, and in particular to a method for controlling and editing the parallax of stereoscopic video based on neural radiation fields. Background Technology

[0002] 3D stereoscopic display technology has been developing for a very long time. With the popularization and development of stereoscopic display devices, the demand for stereoscopic content continues to rise, and stereoscopic conversion is one of the important means of generating 3D content.

[0003] Existing stereoscopic conversion methods can be categorized into three types: manual conversion, traditional conversion, and learning-based conversion. Manual conversion offers the greatest controllability, but it requires significant time and manpower, and demands a certain level of experience from the artist performing the operation. Traditional stereoscopic conversion requires extracting depth from image cues or manually drawing depth maps, which suffers from inaccurate depth map estimation. Deep learning-based methods are currently a research hotspot, specifically divided into two-stage and single-stage methods. The two-stage method requires two stages: monocular depth estimation and missing pixel completion. The training steps are cumbersome, and the conversion relies on depth maps estimated from monocular videos, making precise parallax control difficult. Furthermore, the image completion stage cannot effectively utilize 3D information, resulting in synthesized textures that do not conform to real physical laws. When constructing stereoscopic effects, differences between left and right views can cause 3D dizziness and other visual fatigue. Single-stage automatic stereoscopic conversion is the most ideal solution, but its lack of controllability and the difficulty in acquiring large amounts of 3D video data limit further research development.

[0004] Parallax control faces three main challenges: First, inaccurate parallax calculation. To control parallax to match different monitor sizes and viewing distances and to cope with the diversity of 3D display devices, adaptive parallax control urgently needs to be addressed, requiring the stereoscopic conversion process to have strong and precise parallax control capabilities. Second, visual fatigue. Converted films are often criticized for their lack of comfort. To reduce visual fatigue, comfort zone constraints need to be added to parallax control. Third, local parallax editing. To achieve artistic effects, conversion methods should be able to edit local parallax, such as generating visually comfortable rather than physically accurate stereoscopic videos to create a strong stereoscopic impact from prominent objects. Methods that directly process depth maps to achieve depth control often produce object deformation and holes. One type of method maintains the structure of the selected object by adding multiple constraints during the distortion process, but this type of method relies on the depth map. Summary of the Invention

[0005] To address the above problems, this invention proposes a method for controlling and editing parallax in stereoscopic video based on neural radiation fields. This method fully utilizes the editable properties of neural radiation fields to achieve high-quality viewpoint synthesis with precise parallax control. Furthermore, it can adaptively generate stereoscopic videos with high visual comfort based on viewing conditions and scene attributes. Finally, this method also enables the re-editing of local parallax, resulting in superior local parallax artistic effects.

[0006] This method treats the input video as a left view. First, it constructs a dynamic neural radiation field using multiple consecutive frames of the input video. Simultaneously, it calculates adaptive parallax based on viewing conditions, thereby controlling the virtual camera within the neural radiation field to render a specified new perspective. Volumetric rendering is then used to obtain the right view rendering result. To adapt to the needs of video parallax editing, this method employs layered stereoscopic rendering, achieving the separation of specified objects within the neural radiation field, and thus enabling parallax control of local objects in the video.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A method for controlling and editing the parallax of stereoscopic video based on neural radiation fields, the operation steps of which are as follows:

[0009] Step 1: Construct a 4D spatiotemporal dynamic neural radiation field using the sequence frames of the input view, and then synthesize the static and dynamic objects in the scene after processing them separately;

[0010] Step 2: In the 4D neural radiation field constructed in Step 1, parallax is adaptively and precisely controlled to generate a new view through stereo rendering;

[0011] Step 3: In the 4D neural radiation field constructed in Step 1, separate local objects and modify the parallax of the objects. After stereo rendering, realize the new view after local parallax editing.

[0012] Furthermore, the specific operation steps of step 1 are as follows:

[0013] 1-1: For each frame of the video, the 3D position coordinates of the static part of the motion-masking region and the viewing direction are taken as input. After prediction by two cascaded multilayer perceptrons (MLPs), a field state with continuous color and density in three-dimensional space is obtained as the implicit representation of the scene. The continuous field state is divided into multiple equally spaced layers. Under a given camera model, simulated rays from each pixel position are calculated, and ray tracing is simulated in this field state. The intersection with the layered radiation field is obtained by numerical integration to obtain the pixel color value and cumulative transparency in the rendered image, thereby obtaining the reconstructed color value of the current frame k. A novel view from this viewpoint is then rendered.

[0014] 1-2: When training on a static scene, the results are continuously optimized by minimizing the difference between the reconstructed color value and the real value of each pixel in the current frame, thus obtaining the final weights of the MLP;

[0015] 1-3: When representing dynamic scenes, a bidirectional scene flow is used to calculate the offset of each pixel position at the time before and after the current time, which is used to represent the degree of scene distortion in the temporal domain. Thus, the model is trained using three perspectives, including the current frame and the images at the time before and after it. At the same time, the model is also used to predict the mixed weights at the current time, which is used to represent the combined weight allocation of dynamic and static scenes at a certain pixel position.

[0016] Furthermore, the specific operation steps of step 2 are as follows:

[0017] 2-1: First, identify the objects whose parallax needs to be adjusted based on visual saliency, and then add a mask to them;

[0018] 2-2: Then, layer rendering is performed with a specified depth interval in the neural radiation field, and the depth value is integrated layer by layer while calculating the cumulative color and transparency. When the depth in the salient mask continues to accumulate and tends to stabilize, the depth of the rendering layer is considered to be the depth of the optimal zero parallax plane.

[0019] 2-3: After obtaining the depth value of the zero parallax plane, the relationship function between the optical axis translation degree and the camera baseline a can be further determined based on the off-axis parallel model;

[0020] 2-4: Based on the Shibata comfort zone theory, and according to the current viewing conditions such as viewing distance, interpupillary distance and screen resolution, by controlling the maximum positive and negative parallax within a reasonable range, the specific camera baseline and lens translation amount are further determined to achieve an adaptive parallax control model.

[0021] Furthermore, the specific steps of step 3 are as follows:

[0022] 3-1: Analyze the scene content, add masks to local objects based on the viewer's visual attention points, and determine the 2D bounding box of the target in the image based on the mask;

[0023] 3-2: Accumulate and integrate the depth values ​​in the layered neural radiation field. When the depth values ​​in the mask area start to accumulate and the area stabilizes and reaches the threshold, the depth of the layer is considered to be the near and far depth of the 3D bounding box of the local editing object. Thus, the 3D bounding box of the target object is determined, and the local object is extracted in the neural radiation field.

[0024] 3-3: During the volumetric rendering stage, when performing ray tracing on the entire scene, the light inside the bounding box is distorted in the horizontal direction. This allows for parallax control of local objects and the background after culling local objects, resulting in the effect of magnifying or reducing the horizontal parallax of local objects, thereby enhancing the three-dimensionality of prominent objects in the scene.

[0025] Compared with the prior art, the present invention has the following significant features and advantages:

[0026] 1. This invention proposes a novel automatic stereoscopic conversion method, which for the first time applies neural radiation field technology to complete stereoscopic conversion, and can generate high-quality stereoscopic video with controllability;

[0027] 2. This invention proposes an adaptive parallax control method, which can accurately control the magnitude of parallax and the ratio of positive and negative parallax in the video according to viewing conditions and scene attributes, thereby generating stereoscopic videos with high visual comfort.

[0028] 3. This invention proposes a method for local parallax editing based on neural radiation fields, which can separate local targets in neural radiation fields and achieve artistic effects of local parallax magnification or reduction in stereoscopic rendering. Attached Figure Description

[0029] Figure 1 This is a flowchart illustrating the overall framework of the present invention.

[0030] Figure 2 It is a 4D spatiotemporal dynamic neural radiation field network architecture.

[0031] Figure 3 This is a schematic diagram of the parallax control method for off-axis parallel models.

[0032] Figure 4 This is a flowchart of local parallax editing.

[0033] Figure 5 This invention is compared with three advanced stereo conversion methods and a real value conversion method to demonstrate its effectiveness.

[0034] Figure 6 Angle disparity maps for different zero disparity plane depth values.

[0035] Figure 7 Visual fatigue indices for multiple scenarios under different zero-parallax plane models.

[0036] Figure 8 A red-cyan stereoscopic effect image with adaptive optimal parallax under different viewing conditions.

[0037] Figure 9 These are before-and-after images of local parallax editing. Detailed Implementation

[0038] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0039] like Figure 1 As shown, a method for controlling and editing the parallax of stereoscopic video based on neural radiation fields is described, with the following steps:

[0040] Step 1: Using the sequence frames of the input view, construct a 4D spatiotemporal dynamic neural radiation field, process and synthesize the static and dynamic objects in the scene separately, such as... Figure 2 As shown;

[0041] 1-1: For each frame of the video, the 3D position coordinates γ(x,y,z) of the static part of the motion-masked region and the viewing direction υ(φ,θ) are used as input. After prediction by two cascaded multilayer perceptrons (MLPs), a field state with continuous color and density in 3D space is obtained as the implicit representation of the scene. Then, the continuous field state is divided into M equally spaced layers. Under a given camera model, simulated rays from each pixel position (i,j) are calculated, and ray tracing is simulated in this field state. The intersection with the layered radiation field is numerically integrated to obtain the pixel color value and cumulative transparency T in the rendered image. p This allows us to obtain the reconstructed color value of the current frame k. Render the novel view from that viewpoint;

[0042]

[0043]

[0044] in, Let represent a static multilayer perceptron defined by a set of parameters Θ; γ represents the 3D position coordinates; v represents the viewing direction; c represents the RGB color value; σ represents the volume density; for each frame of the video, k∈(0,...,N-1), δ represents the reconstructed RGBA color. p Indicates the distance between adjacent samples;

[0045] 1-2: When training on a static scene, the results are continuously optimized by minimizing the difference between the reconstructed color value and the real value of each pixel in the current frame, thus obtaining the final weights of the MLP;

[0046] 1-3: When representing dynamic scenes, a bidirectional scene flow is used to calculate the offset of each pixel position at the current time step and the time steps before and after the current time step. This offset represents the degree of scene distortion in the temporal domain, thus enabling the model to be trained using three perspectives, including the current frame and images at the time steps before and after it. Simultaneously, the model predicts the mixed weights at the current time step, which represent the combined weight allocation of dynamic and static scenes at a given pixel position.

[0047] Step 2: In the 4D neural radiation field constructed in Step 1, parallax is adaptively and precisely controlled to generate a new view through stereoscopic rendering, such as... Figure 3 As shown;

[0048] 2-1: Calculation of the Optimal Zero Parallax Plane Based on Visual Saliency: In this invention, the optimal zero parallax plane depth is defined as the boundary value between the foreground and background depths. This depth can distinguish between foreground and background, thus creating a negative parallax effect with the foreground appearing out of the screen, and a positive parallax effect with the background appearing extended. First, the objects whose parallax needs to be adjusted are determined based on visual saliency, and a mask is added to them.

[0049] 2-2: Subsequently, layered rendering is performed within the neural radiation field at specified depth intervals, and the depth value is integrated layer by layer while calculating the cumulative color and transparency. When the depth within the salient mask continuously accumulates and tends to stabilize, the depth of that rendering layer is considered to be the depth of the optimal zero parallax plane.

[0050] 2-3: After obtaining the depth value of the zero parallax plane, the relationship function between the optical axis translation degree Δl and the camera baseline a can be further determined as shown in (3). Based on the off-axis parallel model, according to the obtained zero parallax plane, the relationship expression between the camera baseline a and the lens translation amount Δl can be obtained, where Z is Figure 3 The maximum depth of the scene shown, Z zero Depth of the zero parallax plane;

[0051] (ZZ zero )a-2Z zero Δl=0 (3)

[0052] 2-4: Based on the Shibata comfort zone theory shown in (4), and according to the viewing conditions such as current viewing distance, interpupillary distance, and screen resolution, by controlling the maximum positive and negative parallax within a reasonable range, the specific camera baseline a and lens translation Δl are further determined as shown in (5), thus realizing an adaptive parallax control model, where Z near for Figure 3 The minimum depth of the scene shown. This step achieves optimal stereoscopic effect while providing greater visual comfort;

[0053]

[0054]

[0055] Where, d near d represents the near-plane depth of the comfort zone. far Indicates the far-plane depth of the comfort zone; d irepresents the distance from the eyes to the screen; e represents the interpupillary distance; the Shibata comfort zone theory defines the relationship between viewing distance and convergence distance using a function, and obtains estimates of the near and far boundaries of the comfort zone through numerous experiments, using the parameter T. near ,m near ,T far ,m far M represents the proportional scaling of pixels during image display;

[0056] Step 3: In the 4D neural radiation field constructed in Step 1, separate local objects and modify their parallax. After stereo rendering, realize the new view after local parallax editing, such as... Figure 4 As shown.

[0057] 3-1: Analyze the scene content, add masks to local objects based on the viewer's visual attention points, and determine the 2D bounding box of the target in the image based on the mask;

[0058] 3-2: Accumulate and integrate the depth values ​​in the layered neural radiation field. When the depth values ​​in the mask area start to accumulate and the area stabilizes and reaches the threshold, the depth of the layer is considered to be the near and far depth of the 3D bounding box of the local editing object. This can determine the 3D bounding box of the target object and realize the extraction of local objects in the neural radiation field.

[0059] 3-3: During the volumetric rendering stage, when performing ray tracing on the entire scene, the light inside the bounding box is distorted in the horizontal direction. This allows for parallax control of local objects and the background after culling local objects, resulting in the effect of magnifying or reducing the horizontal parallax of local objects, thereby enhancing the three-dimensionality of prominent objects in the scene.

[0060] Ten stereoscopic videos were downloaded from the internet, including side-by-side 3D movies and side-by-side 3D videos manually converted by professional editors. A 2-3 second video clip was selected from each movie or video, and its left-view video was used as the object scene for stereoscopic conversion. An algorithm was used to generate its right-view video. The final effect can be viewed in a VR environment using an HTC Vive, or processed into a red-green stereoscopic effect demonstration. In this embodiment, a 23-inch LED desktop monitor was used as the display, and the viewing distance was set to 1 meter. The red-green stereoscopic images generated by this invention shown in the accompanying drawings are all based on these viewing conditions. In this embodiment, COLMAP was used to predict camera pose beforehand, and Python programming language PyTorch 1.6.0 and CUDA 10.0 were used to test the code. All experiments were conducted on a machine equipped with an Intel(R) Xeon(R) E5-2620 CPU, a 2.10GHz processor, 64GB of RAM, and an Nvidia Titan Xp GPU. The preferred embodiments of this invention are described in detail below from two aspects: adaptive parallax control and local parallax editing.

[0061] Example 1

[0062] The adaptive disparity control method based on neural radiation fields includes the following steps:

[0063] Step 1: Input the decoded 2D video to... Figure 2 The neural radiation field network structure shown utilizes a time-dependent dynamic neural radiation field to implicitly represent the scene of the input video.

[0064] Step 2: Add constraints on interpupillary distance and viewing distance, construct a virtual camera model in the implicit field state, and realize adaptive parallax calculation and control based on the constraints of parallax comfort zone and stereoscopic perception.

[0065] Step 3: After stereoscopic rendering and compositing, a high-quality novel view is obtained, and the output is a left-right format 3D sequence frame or a red-cyan format 3D sequence frame for viewing under VR glasses or red-green glasses.

[0066] The results of Example 1 will be further analyzed in conjunction with the accompanying drawings:

[0067] 1. Figure 5This demonstrates the advancements in image quality and detail achieved by this invention compared to existing state-of-the-art methods. Previous methods exhibited significant drawbacks, such as pixel jitter and distortion, blurring and artifacts resulting from using only foreground pixels to fill in background holes, and incorrect color value prediction for hole areas. In contrast, this invention leverages NeRF's powerful depth reconstruction capabilities while fully utilizing continuous frame information to more accurately reconstruct occluded scenes. Among these comparative methods, the generated results are closer to the true values.

[0068] 2. Figure 6 , Figure 7 This demonstrates the advantages of the invention's adaptive parallax in terms of visual comfort. Figure 6 It is an angular disparity map reconstructed from a disparity map, where Z0 is the angular disparity map for adaptively calculating the depth of the zero disparity plane in this invention. This method uses the image visual attention region represented by the saliency map as weights, calculates the visual fatigue index using angular disparity values, and displays the results. Figure 7 As can be seen from the data, the disparity control model constructed using the zero disparity plane obtained by our method has a lower visual fatigue index compared to the zero disparity plane in the vicinity of the depth, which confirms that our method can achieve better visual comfort.

[0069] 3. Figure 8 The presentation showcases red-cyan stereoscopic images with varying parallaxes generated adaptively for four film clips under different interpupillary distances (IPDs) and viewing distances (Dv). The actual parallax values ​​increase with both IPD and Dv, demonstrating the impact of viewing distance and IPD on the parallax results. Furthermore, points A, B, and C exhibit different parallax types (positive parallax, zero parallax, and negative parallax), proving that our workflow effectively utilizes the zero parallax plane to segment the foreground and background while simultaneously reflecting visually comfortable positive and negative parallax.

[0070] Example 2

[0071] The local parallax re-editing method based on neural radiation fields is a supplementary operation to Example 1, which specifically includes the following operation steps:

[0072] Steps one and two: Same as in Example 1;

[0073] Step 3: The visual attention region of the image is obtained using a salient region analysis algorithm as a mask, and the 2D bounding box of the target in the image is determined based on the mask; the invention further determines the 3D bounding box of the salient object by accumulating the depth value in the layered neural radiation field, thereby realizing the extraction of local objects in the neural radiation field.

[0074] Step 4: Add a local parallax offset. During the stereoscopic rendering stage, the algorithm, while performing implicit field-state simulation ray tracing, warps the light rays inside the 3D bounding box along the horizontal direction by an offset value, thereby achieving the effect of magnifying or reducing the horizontal parallax of local objects. The warped ray tracing path is then used for stereoscopic rendering, outputting a left-right format 3D frame sequence or a red-cyan format 3D frame sequence for viewing in VR glasses or red-green glasses.

[0075] The results of Example 2 will be further analyzed in conjunction with the accompanying drawings:

[0076] Figure 9 This demonstrates the effect of parallax magnification of prominent areas in a film clip, following the artist's guidance. Compared to post-production software, this invention saves significant costs associated with manual image matting and layering. Furthermore, the foreground obtained through depth measurement is of relatively high quality, providing a basis for judgment and assistance for post-production staff in performing ROTO operations.

[0077] The above embodiments present a method for controlling and editing parallax in stereoscopic video based on neural radiation fields. First, a neural radiation field with bidirectional temporal flow is introduced to generate a dynamic video field with a new perspective. Second, it can adaptively and accurately calculate the ideal parallax based on viewing conditions and video scene characteristics, generating stereoscopic videos with significant stereoscopic effects and visual comfort. Finally, the parallax of individual objects is re-edited during the stereoscopic rendering process based on neural radiation fields. Compared with existing technologies, this method achieves the highest overall performance in image reconstruction quality metrics. Simultaneously, this framework achieves a lower visual fatigue index and a stereoscopic effect including both positive and negative parallax. Experiments in the above embodiments collected feedback from 10 non-professional viewers and 10 professional filmmakers regarding the results of this method. From the perspectives of stereoscopic effect, comfort, and local parallax editing effect, the results demonstrate the value of our framework in optimizing visual experience and artistic expression in 3D film production.

[0078] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made according to the purpose of the invention. Any changes, modifications, substitutions, combinations or simplifications made based on the spirit and principle of the technical solution of the present invention shall be equivalent substitutions. As long as they meet the purpose of the invention and do not deviate from the technical principle and inventive concept of the present invention, they shall fall within the protection scope of the present invention.

Claims

1. A method for neural-radiance-field based stereo video disparity control and editing, characterized in that, The operation steps are as follows: Step 1: Construct a 4D spatiotemporal dynamic neural radiation field using the sequence frames of the input view, and then synthesize the static and dynamic objects in the scene after processing them separately; Step 2: In the 4D neural radiation field constructed in Step 1, parallax is adaptively and precisely controlled to generate a new view through stereo rendering; Step 3: In the 4D neural radiation field constructed in Step 1, separate local objects and modify the parallax of the objects. After stereo rendering, realize the new view after local parallax editing. The specific steps for step 2 are as follows: 2-1: First, identify the objects whose parallax needs to be adjusted based on visual saliency, and then add a mask to them; 2-2: Then, layer rendering is performed with a specified depth interval in the neural radiation field, and the depth value is integrated layer by layer while calculating the cumulative color and transparency. When the depth in the salient mask continues to accumulate and tends to stabilize, the depth of the rendering layer is considered to be the depth of the optimal zero parallax plane. 2-3: After obtaining the depth value of the zero parallax plane, the relationship function between the optical axis translation degree and the camera baseline a can be further determined based on the off-axis parallel model; 2-4: Based on the Shibata comfort zone theory, and according to the current viewing distance, interpupillary distance and screen resolution, by controlling the maximum positive and negative parallax within a reasonable range, the specific camera baseline and lens translation amount are further determined to achieve an adaptive parallax control model.

2. The method of controlling and editing the parallax of a neural radiance field based stereoscopic video according to claim 1, wherein, The specific steps for step 1 are as follows: 1-1: For each frame of the video, the 3D position coordinates of the static part of the motion region and the viewing direction are taken as input. After two cascaded multilayer perceptrons (MLPs) for prediction, a field state with continuous color and density in three-dimensional space is obtained as the implicit representation of the scene. The continuous field is divided into multiple equally spaced layers. Under a given camera model, the simulated light rays from each pixel position are calculated, and ray tracing is simulated in this field. The pixel color value and cumulative transparency in the rendering image are obtained by numerical integration at the intersection with the layered radiation field, thereby obtaining the reconstructed color value of the current frame k and rendering a novel view from this viewpoint. 1-2: When training on a static scene, the results are continuously optimized by minimizing the difference between the reconstructed color value and the real value of each pixel in the current frame, thus obtaining the final weights of the MLP; 1-3: When representing dynamic scenes, a bidirectional scene flow is used to calculate the offset of each pixel position at the time before and after the current time, which is used to represent the degree of scene distortion in the temporal domain. Thus, the model is trained using three perspectives, including the current frame and the images at the time before and after it. At the same time, the model is also used to predict the mixed weights at the current time, which is used to represent the combined weight allocation of dynamic and static scenes at a certain pixel position.

3. The method for controlling and editing stereoscopic video parallax based on neural radiation fields according to claim 1, characterized in that, The specific steps for step 3 are as follows: 3-1: Analyze the scene content, add masks to local objects based on the viewer's visual attention points, and determine the 2D bounding box of the target in the image based on the mask; 3-2: Accumulate and integrate the depth values ​​in the layered neural radiation field. When the depth values ​​in the mask area start to accumulate and the area stabilizes and reaches the threshold, the depth of the layer is considered to be the near and far depth of the 3D bounding box of the local editing object. Thus, the 3D bounding box of the target object is determined, and the local object is extracted in the neural radiation field. 3-3: During the volumetric rendering stage, when performing ray tracing on the entire scene, the light inside the bounding box is distorted in the horizontal direction. This allows for parallax control of local objects and the background after culling local objects, resulting in the effect of magnifying or reducing the horizontal parallax of local objects, thereby enhancing the three-dimensionality of prominent objects in the scene.