Multi-source video collaborative imaging algorithm and device
Through the multi-source video collaborative imaging algorithm, including motion trajectory acquisition, preprocessing, alignment and feature extraction, combined with geometric, radiation and dynamic field modeling, the problems of poor imaging coherence and detail loss in multi-source video processing are solved, and high-quality image fusion is achieved.
Patent Information
- Application Number
- CN202510660542.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing multi-source video processing technology cannot fully explore the information value of multiple video sources, resulting in poor picture coherence and serious loss of details in imaging, especially in intelligent driving and on-board live video broadcast applications, which may lead to wrong judgments or picture jitter.
A multi-source video collaborative imaging algorithm is proposed, including obtaining the motion trajectory and preprocessing of different cameras, spatial and timing alignment, feature extraction and construction of geometric, radiation, and dynamic field modeling subnetworks, through mutation detection and color completion, a complete radiation field is generated, and finally image fusion is performed.
It achieves image fusion with good definition and coherence, and has strong anti-shake capability, solves the problem of object position deviation caused by wide-angle lenses, and improves the imaging quality of intelligent driving and on-board camera live broadcast.
Smart Images

Figure CN120182328A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-source video collaborative imaging algorithm and device, belonging to the technical field of image processing. Background Art
[0002] In many video application scenarios, such as large event live broadcasts, intelligent driving, and complex industrial monitoring, there are generally multiple video source inputs, and it is necessary to fuse multiple video sources to obtain a comprehensive video.
[0003] Traditional multi-video processing methods are relatively simple. Usually, only multiple video frames are simply spliced or switched for display, and the information value of multiple video sources cannot be fully exploited. This processing method may lead to problems such as poor picture coherence and serious loss of details in the final imaging, making it difficult to meet scenarios with high requirements for imaging quality and information integrity.
[0004] Especially in actual use, due to different device models, positions, angles, and installation methods of different video source shooting devices, the obtained videos will have various abnormalities, including jitter, partial occlusion or loss of the field of view, partial or complete area overexposure, color deviation, etc.
[0005] In addition, in order to obtain a wider field of view, existing shooting devices, especially some vehicle-mounted lenses and some outdoor live broadcast lenses, use wide-angle lenses, resulting in phenomena such as distortion, which causes serious deviations in the positions of objects in the fused video. Such deviations, if applied to intelligent driving, may lead to incorrect judgments during intelligent driving and even cause serious consequences; if applied to vehicle-mounted camera live broadcasts, it will cause picture jitter and distortion, affecting the viewing effect.
[0006] Therefore, it is necessary to conduct more in-depth research on the existing multi-source video collaborative imaging algorithm to solve the above problems. Summary of the Invention
[0007] In order to overcome the above problems, in-depth research has been carried out, and a multi-source video collaborative imaging algorithm is proposed, including the following steps: S1. Obtain the motion trajectories of different cameras, and perform preprocessing on all video sources respectively. The preprocessing includes dynamic stabilization and distortion correction; S2. Based on the preprocessed videos, combined with the motion trajectories of their respective cameras, perform spatial alignment and temporal alignment between different source videos; S3. Perform feature extraction on the spatially and temporally aligned different videos respectively to obtain a feature pyramid; S4. Construct a geometric modeling sub-network, and based on the feature pyramid of any source video, obtain the density field of this source video to represent the geometric structure of the scene; Construct a radiation modeling sub-network. Based on the density field of any source video, obtain the radiation field of the source video, which is used to characterize the surface optical properties of objects in the scene; Construct a dynamic field modeling sub-network. Based on the radiation field of any source video, obtain the dynamic field of the source video, which is used to characterize the dynamic changes in the scene; S5. Perform mutation detection based on the density field of any source video to obtain the occluded area of the source video. Based on other source videos, perform color completion on the occluded area to obtain a completed radiation field; Use the completed radiation field to replace the occluded area in the radiation field to obtain a scene field; S6. Fuse the scene fields of all source videos to synthesize an image frame.
[0008] In a preferred embodiment, in S2, during the time sequence alignment process, by generating virtual frames, the frame rates of all video sources are made the same, including the following steps: For any source video, S221. Obtain the optical flow field of the video; S222. Based on the motion trajectory of its camera, predict the relative motion between adjacent frames of the video and obtain the global motion optical flow field prediction; S223. Align the optical flow field with the global motion optical flow field prediction to generate a residual optical flow; S224. Generate virtual frames based on the residual optical flow and supplement them to the actual frames so that different source videos can achieve frame alignment.
[0009] In a preferred embodiment, in S4, the geometric modeling sub-network is a multi-layer MLP structure, which takes the feature pyramid and the corresponding time stamps as inputs and outputs the density fields at different times.
[0010] In a preferred embodiment, the radiation modeling sub-network is a multi-layer MLP structure, and its inputs are the density field and the viewing direction of the camera lens.
[0011] In a preferred embodiment, the dynamic field modeling sub-network includes a gated recurrent unit and a fully connected network, and its inputs are the density field and the radiation field.
[0012] In a preferred embodiment, in S4, a distortion field sub-network is also constructed to perform distortion correction on the radiation field. The distortion field sub-network is a fully connected network with a multi-layer MLP structure. Its inputs are the normalized pixel coordinates and the camera focal length parameters, and the output is the offset of the pixel coordinates. The pixel coordinate offset is used to correct the pixel coordinates in the radiation modeling sub-network.
[0013] In a preferred embodiment, in S5, the completion includes the following sub-steps: S51. Extract feature blocks from the feature pyramids corresponding to other source videos according to the position of the occlusion area of the current video source; S52. Project the extracted feature blocks onto the perspective coordinate system of the current video source to obtain projected feature blocks; S53. Obtain the similarity between the occlusion area and the projected feature blocks corresponding to different source videos; S54. Take the similarity as the weight and fuse multiple projected feature blocks to obtain a completed feature pyramid; S55. Based on the completed feature pyramid, obtain a completed radiation field through a geometric modeling sub-network and a radiometric modeling sub-network.
[0014] In a preferred embodiment, in S6, different scene fields in the fusion process have different weights, and the weights are obtained based on the perspective of the camera that captured the source video.
[0015] The present invention also provides an electronic device, including: At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the above algorithms.
[0016] The present invention also provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any one of the above algorithms.
[0017] The beneficial effects of the present invention include: (1) The finally obtained image has high clarity and good picture coherence; (2) It has strong anti-shake ability and solves the problem of picture distortion caused by field of view occlusion; (3) Solves the problem of object position deviation caused by wide-angle lenses. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic flowchart of a multi-source video collaborative imaging algorithm according to a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The present invention will be further described in detail below with reference to the drawings and embodiments. Through these descriptions, the features and advantages of the present invention will become more clearly defined.
[0020] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein need not be construed as superior to or better than other embodiments. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0021] A multi-source video collaborative imaging algorithm provided according to the present invention, as Figure 1 shown, includes the following steps: S1. Obtain the motion trajectories of different cameras, and perform preprocessing on all video sources respectively. The preprocessing includes dynamic stabilization and distortion correction; S2. Based on the preprocessed videos, in combination with the motion trajectories of their respective cameras, perform spatial alignment and temporal alignment between different source videos; S3. Perform feature extraction on the spatially and temporally aligned different videos respectively to obtain a feature pyramid; S4. Construct a geometric modeling sub-network, and based on the feature pyramid of any source video, obtain the density field of this source video for characterizing the geometric structure of the scene; Construct a radiation modeling sub-network, and based on the density field of any source video, obtain the radiation field of this source video for characterizing the surface optical properties of the objects in the scene; Construct a dynamic field modeling sub-network, and based on the radiation field of any source video, obtain the dynamic field of this source video for characterizing the dynamic changes of the scene; S5. Perform mutation detection based on the density field of any source video to obtain the occlusion area of this source video, and perform color completion on this occlusion area based on other source videos to obtain a completed radiation field; Use the completed radiation field to replace the occlusion area in the radiation field to obtain a scene field; S6. Fuse the scene fields of all source videos to synthesize an image frame.
[0022] In S1, the motion trajectory of the camera is the camera 6DOF parameters and timestamps, which are obtained based on the IMU data of the camera. The 6DOF parameters are commonly used parameters in computer vision, and the specific obtaining method thereof is not limited in the present invention, and those skilled in the art can freely perform it according to experience.
[0023] The dynamic stabilization is used to eliminate non-active motion jitter. The optical flow method is used to calculate the local motion of adjacent frames, which is fused with the IMU data, and the current frame is reversely distorted through a motion compensation matrix to eliminate high-frequency jitter.
[0024] The distortion correction is used to reduce the influence of lens distortion, and inverse distortion is applied based on the camera calibration parameters (focal length, distortion coefficient) for pixel remapping.
[0025] Preferably, in S1, the preprocessing further includes exposure equalization processing. By constructing a luminance histogram pyramid, the dynamic ranges of each camera are aligned, and then the overexposed / underexposed regions are adjusted through an adaptive Gamma curve.
[0026] In S2, the different source videos are spatially aligned by constructing the geometric relationship between different cameras.
[0027] Preferably, the spatial alignment is performed through the following steps: S211. Perform ORB feature extraction on each preprocessed video respectively to obtain feature points; S212. Match the feature points of different source videos to obtain multiple pairs of matching points; S213. Perform linear triangulation on the pairs of matching points of each source video to generate the 3D point cloud of this source video.
[0028] In S211, ORB feature extraction is a commonly used algorithm in SLAM, and its specific process is not elaborated in this invention.
[0029] In S212, a distance metric algorithm is used to obtain the similarity between feature points, and the matching result is obtained based on the similarity. For example, the similarity between feature points is represented by the Hamming distance or the Euclidean distance, so as to achieve matching.
[0030] In S213, linear triangulation is a commonly used algorithm in SLAM, and it is not elaborated in this invention.
[0031] Preferably, there is also step S214. Optimize the 3D point cloud through bundle adjustment. Bundle adjustment (BA) is an algorithm that optimizes multi-view geometric parameters by minimizing the reprojection error, and is widely used in photogrammetry and computer vision. Its specific process is not elaborated in this invention.
[0032] The temporal alignment refers to aligning different source videos. There will be a time deviation between the original frame sequences captured by different cameras, that is, the frame rates of shooting are different, and this deviation will cause blurring of the subsequent synthesized images.
[0033] Furthermore, by generating virtual frames, the frame rates of all video sources are made the same.
[0034] Preferably, the temporal alignment includes the following steps: For any source video, S221. Obtain the optical flow field of this video; S222. Based on the motion trajectory of its camera, predict the relative motion between adjacent frames of the video and obtain the global motion optical flow field prediction; S223. Align the optical flow field with the predicted global motion optical flow field to generate a residual optical flow; S224. Generate virtual frames based on the residual optical flow and supplement them to the actual frames so that different source videos can achieve frame alignment.
[0035] In S221, the optical flow field is obtained based on the Horn-Schunck algorithm.
[0036] In S222, the global motion prediction is expressed as:
[0037] Among them, represents t the position of the camera at time which is obtained from the camera's motion trajectory, t represents the motion prediction at time t for the
[0038] +1 moment, that is, the relative motion between adjacent frames of the video. The predicted global motion optical flow field
[0039] Among them, is the projection function, is the pixel point coordinate, is the depth.
[0040] In S223, the residual optical flow is the difference between the optical flow field and the predicted global motion optical flow field, characterizing the real motion of local dynamic objects.
[0041] In S224, the virtual synchronization frame is expressed as:
[0042] Among them, (x, y) is the pixel point coordinate, is the virtual synchronization frame corresponding to the virtual time τ, is the set of times of the virtual frame, is the time integral offset of the residual perfusion, is the weight coefficient, is the real frame at time t.
[0043] Among them, the weight coefficient is set as:
[0044] Among them, is the time in the virtual frame, represents the residual optical flow, is the weight decay rate that can be set.
[0045] In S3, the feature extraction is implemented using the ConvNeXt network, which is a supervised convolutional neural network widely used in different visual processes. Its structure is not elaborated in this invention.
[0046] Preferably, the feature pyramid includes feature tensors at at least 4 scales.
[0047] In S4, the geometric modeling sub-network is a multi-layer MLP (Multi-Layer Perceptron) structure. It takes the feature pyramid and the corresponding timestamps as inputs and outputs the density field at different times.
[0048] By adding the input of the time dimension, the geometric field has time-varying characteristics, adapting to the subsequent fusion process.
[0049] In a preferred embodiment, the input of the geometric modeling sub-network further includes the depth of the image, which more accurately represents the geometric structure of the scene.
[0050] Preferably, the geometric modeling sub-network is a 5-layer MLP structure. The first four layers use the LeakyReLU activation function, and the last layer is the Softplus activation function.
[0051] The radiance field is the base color of the image frame pixels, and the optical properties of the objects in the image are characterized by the base color.
[0052] Preferably, the input of the radiance modeling sub-network further includes the viewing direction of the camera lens. More preferably, the viewing direction is input in the form of spherical harmonic encoding. For example, the viewing direction is expanded into a 9-dimensional vector by a 3rd-order spherical harmonic function.
[0053] According to the present invention, the viewing direction is concatenated after the density field and used as the input of the radiance modeling sub-network together.
[0054] By adding the viewing direction, dynamic features are introduced in the prediction process, improving the accuracy of color prediction for different source videos.
[0055] According to the present invention, the radiance modeling sub-network is a multi-layer MLP structure.
[0056] Preferably, the radiance modeling sub-network includes a 3-layer MLP structure. The first two layers use the ReLU activation function, and the last layer uses the Sigmoid function.
[0057] The dynamic field is the rate of change of color over time, reflecting the motion blur effect.
[0058] Preferably, the input of the dynamic field modeling sub-network is the density field and the radiance field.
[0059] Preferably, the dynamic field modeling sub-network includes a gated recurrent unit (GRU) and a fully connected network. After splicing the density field and the radiation field, the result is input into the gated recurrent unit, and the output of the gated recurrent unit passes through a multi-layer fully connected network to output the dynamic field.
[0060] In the present invention, using the density field as one of the inputs, which acts as a carrier signal, enables the dynamic field modeling sub-network to selectively absorb the characteristic information of the radiation field and improve the accuracy.
[0061] In the present invention, by setting up a geometric modeling sub-network, a radiation modeling sub-network, and a dynamic field modeling sub-network, different physical effects are separated, the parameter interpretability is enhanced, the accuracy of the model is improved, and the blurring effect caused by motion is greatly reduced.
[0062] Preferably, in S4, a distortion field sub-network is further constructed to perform distortion correction on the radiation field.
[0063] The distortion field sub-network is a fully connected network with a multi-layer MLP structure, preferably a 4-layer MLP structure. The ReLU activation function is used between the multi-layer MLP structures. Its input is the normalized pixel coordinates and camera focal length parameters, and the output is the offset of the pixel coordinates.
[0064] Preferably, the output layer of the distortion field sub-network uses the Tanh function to constrain the offset range and prevent overcorrection.
[0065] Furthermore, the pixel coordinates in the radiation modeling sub-network are corrected using the pixel coordinate offset, thereby participating in the prediction of the radiation field to achieve the effect of secondary distortion correction.
[0066] According to the present invention, the distortion field sub-network splices the coordinates and focal length parameters of each pixel, predicts the coordinate offset at this position through the MLP layer, and then dynamically compensates for the distortion.
[0067] Furthermore, the loss of the distortion field sub-network is set as the pixel difference between the synthesized image and the original uncorrected frame, enabling the distortion field sub-network to learn the deformation of the real distortion.
[0068] According to the present invention, by splicing the coordinates and focal length parameters of each pixel, the consistency of the optical characteristics of different cameras is maintained through the input of the focal length parameters, and the compensation strategy for different cameras is automatically adapted.
[0069] Furthermore, in the training stage of the present invention, it is necessary to first freeze the distortion field sub-network, optimize the geometric modeling sub-network, the radiation modeling sub-network, and the dynamic field modeling sub-network preferentially. After the optimization is completed, the distortion field sub-network is unfrozen, and then all the networks are trained to obtain the final network parameters.
[0070] Traditional distortion correction is carried out through camera calibration parameters, such as the preprocessing process in the present invention. However, it is found that this algorithm does not perform well in the dynamic shooting process of the lens outdoors. It is analyzed that the distortion correction in the preprocessing process can only eliminate the main geometric distortions of the lens (such as barrel distortion, pincushion distortion, etc.), but has no effect on dynamic distortions, such as temperature drift, water vapor refraction, hot air distortion, etc. And water vapor refraction and hot air distortion are common phenomena during the driving of an automobile. In the present invention, by setting a distortion field sub-network to perform secondary correction on the distortion, the distortion can be compensated more accurately, thereby greatly improving the phenomena of jitter and boundary blurring in the synthesized picture area.
[0071] In S5, by calculating the depth change rate in the density field, when it is higher than the threshold, it is considered that a mutation occurs, and the mutation area is the occlusion area.
[0072] In the present invention, the specific value of the threshold is not limited, and those skilled in the art can freely set it according to the actual situation.
[0073] According to the present invention, the completed radiation field is a color distribution field in a three-dimensional space. Preferably, the completion includes the following sub-steps: S51. Extract feature blocks from the feature pyramids corresponding to other source videos according to the position of the occlusion area of the current video source; S52. Project the extracted feature blocks into the perspective coordinate system of the current video source to obtain projected feature blocks; S53. Obtain the similarity between the occlusion area and the projected feature blocks corresponding to different source videos; S54. Take the similarity as the weight and fuse multiple projected feature blocks to obtain a completed feature pyramid; S55. Based on the completed feature pyramid, obtain the completed radiation field through the geometric modeling sub-network and the radiation modeling sub-network.
[0074] Preferably, it further includes S56. Smooth the boundary of the completed radiation field to make its connection with the non-occluded area of the current video source smooth.
[0075] In S51, since there are multiple source videos, the extracted feature blocks are also multiple.
[0076] Preferably, before extracting feature blocks from the feature pyramids of other source videos, the source videos are also detected to confirm whether they are also occluded. If they are occluded, the feature blocks of these videos are not extracted.
[0077] In S52, since the positions and angles of the cameras for shooting different source videos are different, according to the movement trajectory of the camera, the feature blocks are projected into the current perspective coordinate system to solve the spatial misalignment caused by the perspective difference.
[0078] In S53, the similarity is preferably calculated using cosine similarity. Specifically, the cosine similarity between the features at each position in the occluded region and the corresponding position in the feature projection feature block is calculated, and the set of similarities at different positions is used as the weight matrix of the video source.
[0079] Preferably, the weight matrix further includes the dynamic field of other video sources at this position, that is, the product of the similarity and the dynamic field is used as the weight matrix of the video source. The introduction of the dynamic field suppresses the time-varying gradient in the filled region and avoids non-physical motion in the repaired region.
[0080] In S54, the projection feature blocks of different source videos are fused according to the weight matrix to obtain a filled feature pyramid.
[0081] In S56, the dynamic field corresponding to the boundary position of the occluded region of the current video source is used as the confidence to smooth the boundary of the filled radiation field.
[0082] The filling algorithm in the present invention combines the colors, temporal coherence, and motion trends of different source videos, making the accuracy of the supplemented image higher, significantly improving the PSNR for filling dynamic regions, especially having a better filling effect in scenarios such as water vapor occlusion, reflective occlusion, and smoke occlusion.
[0083] In S6, preferably, different scene fields in the fusion process have different weights, and the weights are obtained based on the perspective of the camera that captured the source video.
[0084] Preferably, the weight is expressed as:
[0085] where represents the weight of the th video source at the point (x, y), is a hyperparameter; represents the shooting direction of the th video source and the fusion rendering direction angle.
[0086] The setting of the hyperparameter controls the intensity of the perspective angle attenuation, determines the speed at which the weight decreases with the perspective difference, ensures a smooth transition between adjacent perspectives, and eliminates seams.
[0087] Obviously, the weights of different source videos obtained above need to be normalized before use.
[0088] Furthermore, the rendering is performed based on the ray casting (NeRF) algorithm, which is not elaborated in the present invention.
[0089] The various embodiments of the algorithms described above in the present invention can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0090] It should be understood that the various forms of the flow shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the disclosure of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and this is not limited herein.
Claims
1. A multi-source video collaborative imaging algorithm, characterized in that, It includes the following steps: S1. Obtain the motion trajectories of different cameras, and preprocess all video sources respectively. The preprocessing includes dynamic stabilization and distortion correction; S2. Based on the preprocessed videos, combined with the motion trajectories of their respective cameras, perform spatial alignment and temporal alignment between different source videos; S3. Extract features from the spatially and temporally aligned different videos respectively to obtain a feature pyramid; S4. Construct a geometric modeling sub-network. Based on the feature pyramid of any source video, obtain the density field of this source video, which is used to characterize the geometric structure of the scene; Construct a radiation modeling sub-network. Based on the density field of any source video, obtain the radiation field of this source video, which is used to characterize the surface optical properties of the objects in the scene; Construct a dynamic field modeling sub-network. Based on the radiation field of any source video, obtain the dynamic field of this source video, which is used to characterize the dynamic changes of the scene; S5. Perform mutation detection based on the density field of any source video to obtain the occluded area of this source video, and perform color completion on the occluded area based on other source videos to obtain a completed radiation field; Use the completed radiation field to replace the occluded area in the radiation field to obtain a scene field; S6. Fuse the scene fields of all source videos to synthesize an image frame.
2. The multi-source video collaborative imaging algorithm according to claim 1, characterized in that, In S2, during the temporal alignment process, by generating virtual frames, the frame rates of all video sources are made the same, including the following steps: For any source video, S221. Obtain the optical flow field of this video; S222. Based on the motion trajectory of its camera, predict the relative motion between adjacent frames of the video and obtain the global motion optical flow field prediction; S223. Align the optical flow field with the global motion optical flow field prediction to generate a residual optical flow; S224. Generate virtual frames based on the residual optical flow and supplement them to the actual frames so that different source videos can achieve frame alignment.
3. The multi-source video collaborative imaging algorithm according to claim 1, characterized in that, In S4, the geometric modeling sub-network is a multi-layer MLP structure, which takes the feature pyramid and the corresponding timestamps as inputs and outputs the density fields at different times.
4. The multi-source video collaborative imaging algorithm according to claim 1, characterized in that, The radiation modeling sub-network is a multi-layer MLP structure, and its input is the density field and the viewing direction of the camera lens.
5. The multi-source video collaborative imaging algorithm according to claim 1, characterized in that, The dynamic field modeling sub-network includes a gated recurrent unit and a fully connected network, and its input is the density field and the radiation field.
6. The multi-source video collaborative imaging algorithm according to claim 1, characterized in that, In S4, a distortion field sub-network is also constructed to perform distortion correction on the radiation field. The distortion field sub-network is a fully connected network with a multi-layer MLP structure. Its input is the normalized pixel coordinates and the camera focal length parameters, and the output is the offset of the pixel coordinates. The pixel coordinate offset is used to correct the pixel coordinates in the radiation modeling sub-network.
7. The multi-source video collaborative imaging algorithm according to claim 1, characterized in that, In S5, the completion includes the following sub-steps: S51. Extract feature blocks from the corresponding feature pyramids of other source videos according to the position of the occluded area of the current video source; S52. Project the extracted feature blocks into the view coordinate system of the current video source to obtain projected feature blocks; S53. Obtain the similarity between the occluded area and the projected feature blocks corresponding to different source videos; S54. Fuse multiple projected feature blocks with the similarity as the weight to obtain a completed feature pyramid; S55. Based on the completed feature pyramid, obtain the completed radiation field through the geometric modeling sub-network and the radiation modeling sub-network.
8. The multi-source video collaborative imaging algorithm according to claim 1, characterized in that, In S6, different scenarios in the fusion process have different weights, and the weights are obtained based on the perspective of the camera that captured the source video.
9. An electronic device, characterized in that, Comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the algorithm according to any one of claims 1-8.
10. A computer-readable storage medium storing computer instructions, characterized in that, Wherein, the computer instructions are used to cause the computer to execute the algorithm according to any one of claims 1-8.
Citation Information
Patent Citations
Human body free view angle synthesis method based on generalizable neural radiation field
CN115953476A
Face image synthesis optimization method based on neural radiation field
CN117422829A
Step-by-step depth completion method based on neural radiation field and terminal
CN118839745A
Scene space three-dimensional model dynamic modeling method based on multi-modal data
CN119339008A
Space-time representation of dynamic scenes
US11748940B1