Dynamic scene reconstruction method and electronic equipment

By inserting transition timestamp images in 3D Gaussian Splatting technology and using super-resolution models, the problems of motion blur and texture loss in dynamic scenes are solved, achieving a smoother three-dimensional reconstruction effect of dynamic scenes.

CN120431246APending Publication Date: 2025-08-05UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510421739.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

When handling dynamic scenes, the existing 3D Gaussian Splatting technology is prone to blur and texture loss in moving areas and complex backgrounds, resulting in a decline in visual quality, especially when objects move faster.

Method used

By acquiring image sets from multiple perspectives, analyzing image changes in adjacent timestamps, inserting images with transition timestamps to generate updated image sets, and using super-resolution models to improve image resolution, and finally building a four-dimensional Gaussian point cloud to improve the three-dimensional reconstruction effect of dynamic scenes.

Benefits of technology

It improves the three-dimensional reconstruction effect of dynamic scenes, avoids blurred motion of objects and texture loss, and improves visual quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431246A_ABST
    Figure CN120431246A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic scene reconstruction method and electronic equipment, and the method comprises the steps: obtaining a first image set, obtained by shooting a target dynamic scene at a plurality of first timestamps of the target dynamic scene under a first visual angle, for any first visual angle; for the first image set corresponding to any first visual angle, according to any two adjacent first images of the first timestamps in the first image set, determining whether time frame insertion processing is carried out on the any two adjacent first images of the first timestamps; under the condition that it is determined that time frame insertion processing is carried out on the first images of any two adjacent first timestamps, for any first view angle, a second image corresponding to a second timestamp under the first view angle is inserted between the first images of any two adjacent first timestamps in a first image set corresponding to the first view angle, obtaining an updated first image set corresponding to the first view angle; and determining a four-dimensional Gaussian point cloud corresponding to the target dynamic scene according to the plurality of updated first image sets corresponding to the plurality of first visual angles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional modeling, and more specifically, to a dynamic scene reconstruction method and electronic equipment. Background Art

[0002] With the development of computer vision and computer graphics, three-dimensional reconstruction of dynamic scenes and synthesis of new perspectives have become one of the core technologies in fields such as augmented reality (AR), virtual reality (VR) and movie special effects.

[0003] Currently, 3D Gaussian Splatting (3DGS) is an efficient method for reconstructing three-dimensional scenes. 3DGS uses explicit 3D Gaussian primitives to represent three-dimensional scenes and accelerates the rendering process of three-dimensional scenes through differentiable rasterization, achieving efficient real-time performance. However, 3DGS is mainly suitable for static scenes. When extended to dynamic scenes (i.e., 4D Gaussian Splatting, or 4DGS for short), it is prone to blurring and texture loss when processing moving areas and complex backgrounds, resulting in a decrease in visual quality. In particular, when objects move rapidly, the degree of motion blur is severely deepened, which seriously affects the 3D reconstruction effect of dynamic scenes. Summary of the Invention

[0004] One purpose of the present invention is to provide a new technical solution for dynamic scene reconstruction to solve the technical problem that when constructing dynamic scenes based on 3DGS technology, moving areas and complex backgrounds are prone to blurring and texture loss, which in turn leads to a decline in visual quality, thereby improving the three-dimensional reconstruction effect of dynamic scenes.

[0005] According to a first aspect of the present invention, a dynamic scene reconstruction method is provided, comprising:

[0006] For any first perspective among the multiple first perspectives, obtaining a first image set obtained by photographing the target dynamic scene at multiple first time stamps under the first perspective, to obtain multiple first image sets corresponding to the multiple first perspectives;

[0007] For a first image set corresponding to any first perspective, determining whether to perform time interpolation processing on any two first images with adjacent first timestamps in the first image set;

[0008] When it is determined that temporal interpolation processing is performed on any two first images with adjacent first timestamps, for each first perspective among the multiple first perspectives, a second image corresponding to a second timestamp at the first perspective is inserted between any two first images with adjacent first timestamps in the first image set corresponding to the first perspective, to obtain an updated first image set corresponding to the first perspective; wherein the second timestamp is a transition timestamp between the any two adjacent first timestamps;

[0009] A four-dimensional Gaussian point cloud corresponding to the target dynamic scene is determined according to the multiple updated first image sets corresponding to the multiple first perspectives.

[0010] Optionally, inserting the second image corresponding to the second timestamp at the first perspective between any two adjacent first images with first timestamps in the first image set corresponding to the first perspective to obtain an updated first image set corresponding to the first perspective includes:

[0011] Determine a Gaussian point cloud corresponding to the first timestamp based on the multiple first images corresponding to the same first timestamp in the multiple first image sets;

[0012] determining a motion trajectory of the Gaussian point cloud according to the plurality of Gaussian point clouds corresponding to the plurality of first timestamps;

[0013] determining a Gaussian point cloud corresponding to the second timestamp according to the second timestamp and the motion trajectory of the Gaussian point cloud;

[0014] Determining, based on the Gaussian point cloud corresponding to the second timestamp and the first viewing angle, a two-dimensional image of the Gaussian point cloud of the second timestamp at the first viewing angle as a third image corresponding to the second timestamp at the first viewing angle;

[0015] Inputting the third image corresponding to the second timestamp at the first viewing angle into a super-resolution model to obtain a second image corresponding to the second timestamp at the first viewing angle; wherein the second image resolution of the second image is greater than the third image resolution of the third image;

[0016] According to the second timestamp corresponding to the second image, the second image is inserted as the first image into the first image set corresponding to the first perspective to obtain an updated first image set corresponding to the first perspective.

[0017] Optionally, the determining whether to perform time interpolation processing on any two first images with adjacent first timestamps in the first image set includes:

[0018] Calculating the motion vector of each pixel in the first images of any two adjacent first time stamps by an optical flow method;

[0019] When the number of pixel points of the target motion vector in the first images of any two adjacent first timestamps is greater than or equal to a number threshold, it is determined to perform time interpolation processing on the first images of any two adjacent first timestamps; wherein the target motion vector is a motion vector greater than or equal to a motion vector threshold.

[0020] Optionally, before determining whether to perform time interpolation processing on any two first images with adjacent first timestamps in the first image set, the method further includes:

[0021] Determining a second perspective corresponding to any two adjacent first perspectives from the plurality of first perspectives, thereby obtaining a plurality of second perspectives; wherein the second perspective is a transition perspective between the first perspectives adjacent to the any two adjacent perspectives;

[0022] For any second perspective, determining a fourth image corresponding to the second perspective at the first timestamp based on the Gaussian point cloud corresponding to the second perspective and any first timestamp;

[0023] Inputting the fourth image of the second perspective at the first timestamp into a super-resolution model to obtain a fifth image of the second perspective at the first timestamp; wherein a fifth image resolution corresponding to the fifth image is greater than a fourth image resolution corresponding to the fourth image;

[0024] A first image set corresponding to the first perspective is determined according to the fifth image of the second perspective at each first time stamp in the plurality of first time stamps.

[0025] Optionally, each of the multiple first perspectives is determined by a camera pose, and determining a second perspective corresponding to any two adjacent first perspectives among the multiple first perspectives according to the first perspectives includes:

[0026] Determining, based on a camera pose corresponding to each of the plurality of first perspectives, a plurality of adjacent first perspectives among the plurality of first perspectives;

[0027] For any two adjacent first perspectives among the multiple adjacent first perspectives, determine the second perspective corresponding to the any two adjacent first perspectives based on the camera postures corresponding to the any two adjacent first perspectives, and obtain multiple second perspectives corresponding to the multiple adjacent first perspectives.

[0028] Optionally, the camera pose includes a rotation matrix and a translation vector, and determining the second perspective corresponding to any two adjacent first perspectives according to the camera pose corresponding to the any two adjacent first perspectives includes:

[0029] Inserting a transition rotation matrix into the rotation matrices corresponding to the first perspectives adjacent to any two perspectives by spherical interpolation;

[0030] Inserting a transition translation vector into the translation vectors corresponding to the first perspectives adjacent to any two perspectives by linear interpolation;

[0031] A second perspective corresponding to the first perspectives adjacent to any two perspectives is determined according to the transition rotation matrix and the transition translation vector.

[0032] Optionally, the super-resolution model is determined by the following steps:

[0033] Acquire a training image set; wherein the training image set includes a high-resolution image set and a low-resolution image set, the high-resolution image set includes multiple first images corresponding to the same initial first timestamp in the multiple first image sets, and the low-resolution image set is an image set rendered under the multiple first perspectives based on the Gaussian point cloud corresponding to the initial first timestamp;

[0034] The super-resolution model is trained using the training image set to obtain a trained super-resolution model.

[0035] Optionally, training the super-resolution model using the training image set to obtain a trained super-resolution model includes:

[0036] Inputting the low-resolution image set into a super-resolution model to obtain a predicted high-resolution image set;

[0037] constructing a loss function based on the predicted high-resolution image set and the high-resolution image set;

[0038] According to the loss function, the model parameters of the super-resolution model are adjusted to obtain the trained super-resolution model.

[0039] Optionally, determining the four-dimensional Gaussian point cloud corresponding to the target dynamic scene according to the multiple updated first image sets corresponding to the multiple first perspectives includes:

[0040] Determining an updated Gaussian point cloud corresponding to a third timestamp based on a plurality of first images corresponding to a same third timestamp in a plurality of updated first image sets corresponding to the plurality of first perspectives; wherein the third timestamp includes the first timestamp and the second timestamp;

[0041] For any first perspective among the multiple first perspectives, determining a reference image set corresponding to the first perspective based on the updated Gaussian point clouds of the multiple third timestamps and the first perspective;

[0042] optimizing, based on any two reference images with adjacent third time stamps in the reference image set corresponding to the first perspective, two first images corresponding to any two adjacent third time stamps in the updated first image set corresponding to the first perspective, to obtain an optimized first image set corresponding to the first perspective;

[0043] According to the multiple optimized first images belonging to the same third timestamp in the multiple optimized first image sets corresponding to the multiple first perspectives, the optimized Gaussian point cloud corresponding to the third timestamp is determined to obtain a four-dimensional Gaussian point cloud corresponding to the target dynamic scene.

[0044] According to a second aspect of the present invention, an electronic device is further provided, comprising a memory and a processor, wherein the memory is used to store executable instructions; the processor is used to operate under the control of the instructions to execute the method as described in the first aspect of the present invention.

[0045] One beneficial effect of the present invention is that, by analyzing the first images of any two adjacent first timestamps in any first image set, it is determined whether it is necessary to perform time interpolation processing on the first images of the two adjacent first timestamps. If it is determined that time interpolation processing is required, a second image is inserted between the first images corresponding to any two adjacent first timestamps in the first image set corresponding to each first perspective. In this way, when an object moves faster in a target dynamic scene, the second image can be inserted into the two first images corresponding to the faster movement of the object in each first image set corresponding to the first perspective, so as to obtain an updated first image set corresponding to the first perspective. Since the second image is inserted into the two first images of the faster movement of the object in the updated first image set, the movement of the object presented in the three-dimensional reconstruction of the target dynamic scene can be made smoother, avoiding the problems of blur and texture loss that are easy to occur when processing moving areas and complex backgrounds, and improving the visual quality of the three-dimensional reconstruction of the target dynamic scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0047] Figure 1 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention;

[0048] Figure 2 is a schematic flow chart of a dynamic scene reconstruction method according to an embodiment of the present invention;

[0049] Figure 3 FIG. 1 is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0050] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention.

[0051] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.

[0052] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0053] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0054] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0055] <Hardware Configuration>

[0056] Figure 1 is a schematic structural diagram of an electronic device 100 according to an embodiment of the present invention.

[0057] like Figure 1 As shown, the electronic device 100 can be any electronic device, such as a PC, a notebook computer, a server, etc.

[0058] In this embodiment, referring to Figure 1 As shown, the electronic device 100 may include a processor 1100, a memory 1200, an interface device 1300, a communication device 1400, a display device 1500, an input device 1600, a speaker 1700, a microphone 1800, and the like.

[0059] The processor 1100 may be a mobile processor. The memory 1200 may include, for example, ROM (read-only memory), RAM (random access memory), and non-volatile memory such as a hard disk. The interface device 1300 may include, for example, a USB interface, a headphone jack, and the like. The communication device 1400 may be capable of wired or wireless communication. The communication device 1400 may include a short-range communication device, such as any device that performs short-range wireless communication based on a short-range wireless communication protocol such as Hilink protocol, WiFi (IEEE 802.11 protocol), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, LiFi, etc. The communication device 1400 may also include a long-range communication device, such as any device that performs WLAN, GPRS, 2G / 3G / 4G / 5G long-range communication. The display device 1500 may be, for example, an LCD display, a touch screen display, etc. The display device 1500 is used to display the captured first image. The input device 1600 may include, for example, a touch screen, a keyboard, etc. The user may input / output voice information through the speaker 1700 and the microphone 2800.

[0060] In this embodiment, the memory 1200 of the electronic device 100 is used to store instructions for controlling the processor 1100 to perform at least the dynamic scene reconstruction method according to any embodiment of the present invention. A skilled person can design instructions based on the disclosed solution. How instructions control the processor's operation is well known in the art and will not be described in detail here.

[0061] Despite Figure 1 , multiple devices of the electronic device 100 are shown; however, the present invention may only relate to some of the devices, for example, the electronic device 100 only relates to the memory 1200 , the processor 1100 and the display device 1500 .

[0062] <Method Example>

[0063] Figure 2 2 is a flow chart of a dynamic scene reconstruction method according to an embodiment of the present invention, which can be implemented by the electronic device 2000.

[0064] according to Figure 2 As shown, the dynamic scene reconstruction method of this embodiment may include the following steps S2100 to S2600:

[0065] Step S2100: For any first perspective among multiple first perspectives, obtain a first image set obtained by photographing the target dynamic scene at multiple first time stamps under the first perspective, and obtain multiple first image sets corresponding to the multiple first perspectives.

[0066] In this embodiment, a dynamic scene may refer to a scene in which objects, background, or lighting conditions in the scene change over time.

[0067] The target dynamic scene may be a dynamic scene that requires three-dimensional modeling. The target dynamic scene may be, for example, a scene of a chef cutting a steak, a vehicle driving on the road, etc. Those skilled in the art should understand that the specific type of the target dynamic scene is not limited here.

[0068] The multiple first perspectives may be, for example, 5 first perspectives, and the 5 first perspectives are different.

[0069] The first perspective is determined by the camera pose, that is, each first perspective corresponds to a camera pose, and multiple first perspectives correspond to multiple camera poses.

[0070] Those skilled in the art should understand that the multiple first perspectives may also be other numbers of first perspectives, which is not limited here.

[0071] The first image set includes a plurality of first images, and the first images are two-dimensional RGB images, that is, the first images include RGB values of each pixel.

[0072] The multiple first timestamps are arranged in chronological order, for example, 1s, 2s, 3s, 4s, etc.

[0073] For example, if the multiple first perspectives are five first perspectives and the multiple first timestamps are fifteen first timestamps, then for any one of the five first perspectives, two-dimensional RGB images (i.e., first images) of the target dynamic scene at the first perspective at the fifteen first timestamps are captured to obtain a first image set corresponding to the first perspective. Thus, five first image sets corresponding to the five perspectives are obtained, and each first image set includes the fifteen first images at the fifteen first timestamps arranged in chronological order.

[0074] The inventors found that, under normal circumstances, the number of first perspectives in step S2100 is small. If multiple first image sets corresponding to multiple first perspectives are directly used as rendering image sets, and the four-dimensional Gaussian point cloud of the target dynamic scene (i.e., the three-dimensional Gaussian point cloud that changes over time) is rendered through the rendering image sets, due to the insufficient number of first perspectives, the rendering results are prone to discontinuity, which affects the reconstruction effect of the target dynamic scene.

[0075] In this regard, the inventors expanded multiple second perspectives based on the multiple first perspectives, and then used the second perspective as a first perspective to obtain multiple first image sets corresponding to the new first perspectives, thereby improving the perspective richness of the rendered image set. This can solve the problem in related technologies that rendering results are prone to discontinuity due to insufficient number of first perspectives, and improve the ability to restore details of dynamic scenes.

[0076] Based on this, in some embodiments, after step S2100 and before executing step S2200, the method further includes: steps S3100 to S3400.

[0077] Step S3100 , determining a second perspective corresponding to any two adjacent first perspectives among the multiple first perspectives, to obtain multiple second perspectives.

[0078] In this embodiment, the two adjacent first perspectives are two first perspectives with the shortest perspective distance in spatial position among the multiple first perspectives.

[0079] For the convenience of description, two adjacent first perspectives may be referred to as a first perspective pair. Then, there are multiple first perspective pairs in the multiple first perspectives.

[0080] In some embodiments, each of the plurality of first perspectives is determined by a camera pose.

[0081] In this embodiment, one first perspective corresponds to one camera pose, and multiple first perspectives correspond to multiple camera poses.

[0082] The camera pose can be the pitch angle and rotation angle of the camera. The camera pose can also be the rotation matrix and translation vector of the camera, which is not limited here.

[0083] In these embodiments, step S3100 determines a second perspective corresponding to any two adjacent first perspectives among the multiple first perspectives based on the first perspectives of the multiple first perspectives, including: step S3100.1 and step S3100.2.

[0084] Step S3100.1: Determine adjacent first perspectives among the multiple first perspectives based on a camera pose corresponding to each first perspective among the multiple first perspectives.

[0085] In this embodiment, in an embodiment where the first perspective is determined by a camera posture, two adjacent first perspectives may also refer to two first perspectives that are closest in spatial position among a plurality of camera postures corresponding to the plurality of first perspectives.

[0086] Exemplarily, the multiple first perspectives are five first perspectives, namely, perspective A, perspective B, perspective C, perspective D, and perspective E, each of which corresponds to a camera pose. For perspective A among the five first perspectives, the first perspective closest to perspective A is determined based on the camera pose of perspective A and the camera poses of the other first perspectives (i.e., perspective B, perspective C, perspective D, and perspective E). For example, the first perspective closest to perspective A is perspective B. Perspective A and perspective B are then two adjacent first perspectives, i.e., perspective A and perspective B form a first perspective pair. Using the same method, it can be obtained that perspective B and perspective C form a first perspective pair, perspective C and perspective D form a first perspective pair, and perspective D and perspective E form a first perspective pair. Perspective A and perspective B, perspective B and perspective C, perspective C and perspective D, and perspective D and perspective E are multiple adjacent first perspectives.

[0087] Step S3100.2, for any two adjacent first perspectives among the multiple adjacent first perspectives, determine the second perspective corresponding to any two adjacent first perspectives based on the camera postures corresponding to the any two adjacent first perspectives, and obtain multiple second perspectives corresponding to the multiple adjacent first perspectives.

[0088] In this embodiment, for any two adjacent first perspectives among multiple adjacent first perspectives, that is, any first perspective pair among multiple first perspective pairs, a second perspective is determined through perspective interpolation, thereby obtaining multiple second perspectives corresponding to the multiple first perspective pairs. The second perspective is a transitional perspective between the two adjacent first perspectives.

[0089] Continuing with the above example, perspectives A and B form a first perspective pair, perspectives B and C form a first perspective pair, perspectives C and D form a first perspective pair, and perspectives D and E form a first perspective pair, for a total of four first perspective pairs. Then, for the first perspective pair consisting of perspectives A and B, the transition perspective a between perspectives A and B is determined. For the first perspective pair consisting of perspectives B and C, the transition perspective b between perspectives B and C is determined. For the first perspective pair consisting of perspectives C and D, the transition perspective c between perspectives C and D is determined. For the first perspective pair consisting of perspectives D and E, the transition perspective d between perspectives D and E is determined. This results in four transition perspectives corresponding to the four first perspective pairs, i.e., four second perspectives. These four second perspectives are transition perspective a, transition perspective b, transition perspective c, and transition perspective d.

[0090] In some embodiments, the camera pose includes a rotation matrix and a translation vector.

[0091] In this embodiment, the rotation matrix can be used to represent the rotation of the camera coordinate system relative to the world coordinate system, and the translation vector can be used to represent the translation of the camera coordinate system relative to the world coordinate system.

[0092] In these embodiments, step S3100.2 determines the second perspective corresponding to any two adjacent first perspectives based on the camera pose corresponding to the first perspectives of any two adjacent perspectives, including steps S3100.21 to S3100.23.

[0093] Step S3100.21: insert a transition rotation matrix into the rotation matrices corresponding to the first perspectives adjacent to any two perspectives through spherical interpolation.

[0094] For example, for a first perspective pair consisting of perspective A and perspective B, according to the rotation matrix R of perspective A a and the rotation matrix R of the view B b , determine a transition rotation matrix R by the following formula (1) k :

[0095] R k =Slerp(R a ,R b ,α) (1)

[0096] The above formula (1) is the calculation formula for spherical interpolation.

[0097] Those skilled in the art should know that spherical interpolation calculates the optimal interpolation path based on the center of the object, so that the generated perspective data conforms to the constraints of the spherical coordinate system, thereby optimizing the camera trajectory and reducing the geometric distortion that may be caused by linear interpolation.

[0098] Step S3100.22: insert a transition translation vector into the translation vectors corresponding to the first perspectives adjacent to any two perspectives through linear interpolation.

[0099] For example, for a first perspective pair consisting of perspective A and perspective B, according to the translation vector T of perspective A a and the translation vector T of the viewing angle B b , a transition translation vector T is determined by the following formula (2): k :

[0100] T k =(1-α)T a +αT b (2)

[0101] The above formula (2) is the calculation formula for linear interpolation.

[0102] Step S3100.23: Determine the second perspective corresponding to the first perspectives adjacent to any two perspectives based on the transition rotation matrix and the transition translation vector.

[0103] Continuing with the above example, the transition rotation matrix R calculated according to formula (1) is k The transition translation vector T calculated by formula (2) k , get a second perspective.

[0104] Step S3200: For any second perspective, determine a fourth image corresponding to the second perspective at the first timestamp based on the Gaussian point cloud corresponding to the second perspective and any first timestamp.

[0105] In this embodiment, the Gaussian point cloud corresponding to any first timestamp can be a three-dimensional Gaussian representation of the target dynamic scene at that first timestamp. The Gaussian point cloud corresponding to any first timestamp can be obtained based on the multiple first images in the multiple first image set acquired in step S2100 that belong to that first timestamp. If the multiple first timestamps in step S2100 are 15 first timestamps, then the Gaussian point cloud corresponding to each of the 15 first timestamps can be obtained, thus obtaining 15 Gaussian point clouds corresponding to the 15 first timestamps.

[0106] Those skilled in the art should understand that the method of determining the Gaussian point cloud corresponding to the first timestamp based on multiple first images belonging to the same first timestamp in multiple first image sets is well known in the art and will not be elaborated here.

[0107] Among them, a three-dimensional Gaussian point cloud corresponding to a first timestamp can be composed of multiple Gaussian basis elements, denoted as G = {g1, g2, ..., g N}, where N represents the number of Gaussian basis elements in the static scene of the target dynamic scene at the first timestamp. Moreover, for any Gaussian basis element in the Gaussian point cloud, it can be represented as a three-dimensional Gaussian distribution with a mean vector μ and a covariance matrix Σ:

[0108]

[0109] Among them, μ∈R3 represents the position of the Gaussian basis element in the Gaussian point cloud, and Σ∈R3×3 is the anisotropic covariance matrix.

[0110] In this embodiment, for any second perspective, based on the second perspective and the Gaussian point cloud corresponding to any first timestamp, a two-dimensional image of the Gaussian point cloud corresponding to the first timestamp at the second perspective is obtained as the fourth image corresponding to the second perspective at the first timestamp.

[0111] Continuing with the above example, five first perspectives are interpolated to determine four second perspectives, namely transition perspective a, transition perspective b, transition perspective c, and transition perspective d. The plurality of first timestamps is 15 first timestamps. At this point, for transition perspective a, based on transition perspective a and the Gaussian point cloud corresponding to any of the 15 first timestamps, a two-dimensional image of the Gaussian point cloud corresponding to that first timestamp is determined at transition perspective a, and this image is used as the fourth image of transition perspective a at that first timestamp, resulting in 15 fourth images of transition perspective a at the 15 first timestamps. Similarly, 15 fourth images of transition perspective b at the 15 first timestamps, 15 fourth images of transition perspective c at the 15 first timestamps, and 15 fourth images of transition perspective d at the 15 first timestamps can be obtained.

[0112] Since the fourth image is obtained by rendering the Gaussian point cloud corresponding to any first timestamp, the resolution of the fourth image is relatively low. Therefore, super-resolution processing needs to be performed on the fourth image to improve the resolution of the fourth image.

[0113] Step S3300: Input the fourth image of the second perspective at the first timestamp into a super-resolution model to obtain a fifth image of the second perspective at the first timestamp.

[0114] In this embodiment, the fifth image resolution corresponding to the fifth image is greater than the fourth image resolution corresponding to the fourth image, wherein the fifth image resolution is the image resolution of the fifth image, and the fourth image resolution is the image resolution of the fourth image.

[0115] Continuing with the above example, 15 fourth images at 15 first time stamps for transition view a, 15 fourth images at 15 first time stamps for transition view b, 15 fourth images at 15 first time stamps for transition view c, and 15 fourth images at 15 first time stamps for transition view d are input into the super-resolution model, resulting in 15 fifth images corresponding to transition view a, 15 fifth images corresponding to transition view b, 15 fifth images corresponding to transition view c, and 15 fifth images corresponding to transition view d. The fifth image resolution of the fifth image is greater than the fourth image resolution of the fourth image.

[0116] Step S3400: Determine a first image set corresponding to a first perspective according to the fifth image of the second perspective at each first timestamp in the plurality of first timestamps.

[0117] In this embodiment, the second perspective is used as a new first perspective, and the multiple fifth images of the second perspective at multiple first timestamps are used as the multiple first images of the new first perspective at multiple first timestamps to obtain the first image set corresponding to the new first perspective.

[0118] Continuing with the above example, after obtaining the 15 fifth images of transition perspective a at the 15 first time stamps, transition perspective a is used as a new first perspective. The 15 fifth images of transition perspective a at the 15 first time stamps are then used as the 15 first images of the new first perspective at the 15 first time stamps, obtaining the first image set corresponding to the new first perspective. Similarly, transition perspective b, transition perspective c, and transition perspective d can be used as the first image set corresponding to the new first perspective. Adding these to the five first image sets of the original five first perspectives, we obtain nine first image sets corresponding to nine first perspectives.

[0119] Step S2200 : For a first image set corresponding to any first perspective, determine whether to perform time interpolation processing on any two first images with adjacent first timestamps in the first image set.

[0120] In this embodiment, the change information between any two first images with adjacent first timestamps in any first image set can be identified to determine whether to perform temporal interpolation processing on the first images with adjacent first timestamps. The change information includes optical flow information, illumination change information, etc.

[0121] Those skilled in the art should understand that the first perspective in step S2200 may be an original first perspective or a second perspective generated by perspective interpolation, which is not limited here.

[0122] For example, the 15 first timestamps are 0s, 1s, 2s, 3s, 4s, 5s, 6s, 7s, 8s, 9s, 10s, 11s, 12s, 13s, and 14s. For a first image set corresponding to any first perspective, based on the first images of any two adjacent first timestamps in the first image set, determine whether to perform time interpolation processing on the first images of any two adjacent first timestamps. If it is finally determined that time interpolation processing is required for: the first images corresponding to 0s and 1s, the first images corresponding to 5s and 6s, the first images corresponding to 8s and 9s, the first images corresponding to 10s and 11s, and the first images corresponding to 13s and 14s.

[0123] In some embodiments, step S2200 determines whether to perform time interpolation processing on any two first images with adjacent first timestamps in the first image set, including: step S2200.1 and step S2200.2.

[0124] Step S2200.1: Calculate the motion vector of each pixel in the first image of any two adjacent first time stamps by using an optical flow method.

[0125] Those skilled in the art should understand that the method of calculating the motion direction of each pixel in two images by using the optical flow method is well known in the art and will not be elaborated on here.

[0126] Step S2200.2: When the number of pixels of the target motion vector in the first images of any two adjacent first time stamps is greater than or equal to a number threshold, determine to perform temporal interpolation processing on the first images of any two adjacent first time stamps.

[0127] In this embodiment, the quantity threshold may be a critical value of the number of pixels of the target motion vector corresponding to the start of the temporal interpolation process, wherein the target motion vector is a motion vector greater than or equal to the motion vector threshold.

[0128] If the number of pixels in the target motion vector in the first images of any two adjacent first time stamps is greater than or equal to a threshold, it indicates that the first images of the two adjacent first time stamps have changed significantly, that is, the object in the two adjacent first time stamps is moving rapidly. In this case, to avoid the problem of increased motion blur caused by the rapid object movement in the two adjacent first time stamps, which would seriously affect the 3D reconstruction effect of the dynamic scene, the first images of the two adjacent first time stamps are temporally interpolated.

[0129] Step S2300, when it is determined that time interpolation processing is performed on the first images of any two adjacent first timestamps, for each first perspective of the multiple first perspectives, the second image corresponding to the second timestamp under the first perspective is inserted between the first images of any two adjacent first timestamps in the first image set corresponding to the first perspective to obtain an updated first image set corresponding to the first perspective.

[0130] In this embodiment, the second timestamp is a transition timestamp between any two adjacent first timestamps.

[0131] After the second image is inserted into the first image set, the second image in the updated first image set is the first image, so that subsequent processing can be performed based on the updated first image set.

[0132] Continuing with the above example, if it is determined that time interpolation processing is required for the following: the first image corresponding to 0s and 1s, the first image corresponding to 5s and 6s, the first image corresponding to 8s and 9s, the first image corresponding to 10s and 11s, and the first image corresponding to 13s and 14s, then for any first perspective, the second image corresponding to 0.5s under the first perspective is inserted between the first images corresponding to 0s and 1s, the second image corresponding to 5.5s under the first perspective is inserted between the first images corresponding to 5s and 6s, the second image corresponding to 8.5s under the first perspective is inserted between the first images corresponding to 8s and 9s, the second image corresponding to 10.5s under the first perspective is inserted between the first images corresponding to 10s and 11s, and the second image corresponding to 13.5s under the first perspective is inserted between the first images corresponding to 13s and 14s, thereby obtaining five second images corresponding to the first perspective. These five second images are then inserted as new first images into the first image set corresponding to the first perspective, thereby obtaining an updated first image set corresponding to the first perspective. Similarly, an updated first image set corresponding to each of the multiple first perspectives can be obtained, wherein 0.5s, 5.5s, 8.5s, 10.5s, and 13.5s are the second timestamps.

[0133] Since the second timestamp is a transition timestamp between two adjacent first timestamps, in order to obtain the second image corresponding to the second timestamp at any first perspective, it is necessary to first obtain the Gaussian point cloud corresponding to the second timestamp. Then, based on the Gaussian point cloud corresponding to the second timestamp, rendering, super-resolution, and other operations at the first perspective are performed to obtain the second image corresponding to the second timestamp at any first perspective. Therefore, the following describes in detail the specific process of determining the second image corresponding to the second timestamp at any first perspective through steps S2300.1 to S2300.5:

[0134] In some embodiments, step S2300 inserts the second image corresponding to the second timestamp under the first perspective between any two first images with adjacent first timestamps in the first image set corresponding to the first perspective to obtain an updated first image set corresponding to the first perspective, including: steps S2300.1 to S2300.6.

[0135] Step S2300.1: Determine a Gaussian point cloud corresponding to the first timestamp based on the multiple first images corresponding to the same first timestamp in the multiple first image sets.

[0136] In one example, the multiple first image sets are five first image sets corresponding to the original five first viewpoints, and the multiple first timestamps are 15. Therefore, based on the five first images corresponding to any first timestamp in the five first image sets, the Gaussian point cloud corresponding to that first timestamp can be determined, resulting in 15 Gaussian point clouds corresponding to the 15 first timestamps. Due to the large number of first viewpoints in this example, the generated 15 Gaussian point clouds have a lower level of detail.

[0137] In another example, the multiple first image sets are nine first image sets corresponding to the original five first perspectives and four new first perspectives generated by perspective interpolation. Thus, based on the nine first images corresponding to any first timestamp in these nine first image sets, the Gaussian point cloud corresponding to that first timestamp can be determined, resulting in 15 Gaussian point clouds corresponding to the 15 first timestamps. Because this example includes a large number of first perspectives, the resulting 15 Gaussian point clouds have a high level of detail.

[0138] Step S2300.2: Determine the motion trajectory of the Gaussian point cloud according to the multiple Gaussian point clouds corresponding to the multiple first timestamps.

[0139] In this embodiment, the motion trajectory of the Gaussian point cloud is used to characterize the position change of each Gaussian basis element in the Gaussian point cloud over time.

[0140] In one example, the motion trajectory of a Gaussian point cloud is a set of displacement-time equations. Each displacement-time equation describes the motion trajectory of a Gaussian basis element. The displacement-time equation is used to represent the position change of a Gaussian basis element over time.

[0141] The motion trajectory of each Gaussian basis element in the Gaussian point cloud is determined by tracking the position of the Gaussian basis element in the plurality of first time stamps.

[0142] Those skilled in the art should understand that the method of determining the motion trajectory of each Gaussian basis element in the Gaussian point cloud by tracking the position of the Gaussian basis element in multiple first time stamps is relatively well known in the art and will not be elaborated here.

[0143] Step S2300.3: Determine the Gaussian point cloud corresponding to the second timestamp based on the second timestamp and the motion trajectory of the Gaussian point cloud.

[0144] In this embodiment, since the motion trajectory of the Gaussian point cloud is used to characterize the position change of each Gaussian basis element in the Gaussian point cloud over time, and the position change of a Gaussian basis element over time can be expressed by a displacement-time relationship equation, the second timestamp can be substituted into the displacement-time relationship equation corresponding to each Gaussian basis element to determine the position of each Gaussian basis element, and thus the Gaussian point cloud corresponding to the second timestamp can be obtained based on the positions of multiple Gaussian basis elements corresponding to the second timestamp.

[0145] For example, if the second timestamp is 0.5s, the Gaussian point cloud corresponding to 0.5s can be obtained.

[0146] Step S2300.4: Determine, based on the Gaussian point cloud corresponding to the second timestamp and the first perspective, a two-dimensional image of the Gaussian point cloud of the second timestamp at the first perspective as a third image corresponding to the second timestamp at the first perspective.

[0147] Continuing with the above example, after obtaining the Gaussian point cloud corresponding to 0.5 seconds, the Gaussian point cloud corresponding to 0.5 seconds is rendered at any of the nine first-view angles, and a 2D image of the Gaussian point cloud at that first angle is obtained as the third image corresponding to 0.5 seconds at that first angle. In other words, each third image corresponds to a second timestamp and a first angle.

[0148] Since the third image is obtained by rendering the Gaussian point cloud of the second timestamp from the first perspective, its resolution is relatively low. In order to improve the effect of dynamic scene reconstruction, the third image needs to be super-resolution processed, that is, execute step S2300.5.

[0149] Step S2300.5: Input the third image corresponding to the second timestamp under the first perspective into the super-resolution model to obtain the second image corresponding to the second timestamp under the first perspective.

[0150] In this embodiment, the second image resolution of the second image is greater than the third image resolution of the third image, wherein the second image resolution is the image resolution of the second image, and the third image resolution is the image resolution of the third image.

[0151] Continuing with the above example, after obtaining the third image corresponding to any first viewing angle at 0.5 s, super-resolution processing is performed on these third images to obtain the second image corresponding to the first viewing angle at 0.5 s.

[0152] Step S2300.6: insert the second image as the first image into the first image set corresponding to the first perspective according to the second timestamp corresponding to the second image, to obtain an updated first image set corresponding to the first perspective.

[0153] For example, the second timestamps are 0.5s, 5.5s, 8.5s, 10.5s, and 13.5s. Thus, by executing steps S2300.1 to S2300.5, nine second images corresponding to nine first view angles at 0.5s, nine second images corresponding to nine first view angles at 5.5s, nine second images corresponding to nine first view angles at 8.5s, nine second images corresponding to nine first view angles at 10.5s, and nine second images corresponding to nine first view angles at 13.5s are obtained. For any second image among the nine second images corresponding to nine first view angles at 0.5s, based on the first view angle corresponding to the second image and the second timestamp of 0.5, the second image is inserted between the first images corresponding to 0s and 1s of the first image set corresponding to the first view angle. For any second image among the nine second images corresponding to the nine first view angles at 5.5s, based on the first view angle corresponding to the second image and the second timestamp of 5.5, the second image is inserted between the first images corresponding to 5s and 6s of the first image set corresponding to the first view angle. For any second image among the nine second images corresponding to the nine first view angles at 8.5s, based on the first view angle corresponding to the second image and the second timestamp of 8.5, the second image is inserted between the first images corresponding to 8s and 9s of the first image set corresponding to the first view angle. For any second image among the nine second images corresponding to the nine first view angles at 10.5s, based on the first view angle corresponding to the second image and the second timestamp of 10.5, the second image is inserted between the first images corresponding to 10s and 11s of the first image set corresponding to the first view angle. For any second image among the nine second images corresponding to the nine first view angles at 13.5s, based on the first view angle corresponding to the second image and the second timestamp of 13.5s, the second image is inserted between the first images corresponding to 13s and 14s of the first image set corresponding to the first view angle. After performing all the above insertions, an updated first image set corresponding to each of the nine first perspectives is obtained. For ease of description, the first timestamp and the second timestamp can be collectively referred to as the third timestamp. Then, the updated first image set includes 20 first images corresponding to the 20 third timestamps.

[0154] Step S2400 : determining a four-dimensional Gaussian point cloud corresponding to the target dynamic scene based on the multiple updated first image sets corresponding to the multiple first perspectives.

[0155] In this embodiment, the four-dimensional Gaussian point cloud corresponding to the target dynamic scene represents the change of the Gaussian point cloud corresponding to the target dynamic scene over time.

[0156] Continuing with the above example, after obtaining the updated first image set corresponding to each of the 9 first perspectives, the Gaussian point cloud corresponding to the target dynamic scene is constructed according to the updated first image set corresponding to each of the 9 first perspectives, that is, the four-dimensional Gaussian point cloud corresponding to the target dynamic scene is constructed.

[0157] In order to further improve the temporal stability and reconstruction coherence of the interpolated perspective in dynamic scenes, pixel optimization can be performed on the updated first image set corresponding to the first perspective based on the reference image set corresponding to the first perspective to obtain the optimized first image set, and a four-dimensional Gaussian point cloud can be obtained based on multiple optimized first image sets corresponding to multiple first perspectives.

[0158] Based on this, in some embodiments, step S2400 determines the four-dimensional Gaussian point cloud corresponding to the target dynamic scene based on the multiple updated first image sets corresponding to the multiple first perspectives, including: steps S2400.1 to S2400.4.

[0159] Step S2400.1: Determine an updated Gaussian point cloud corresponding to the third timestamp based on the multiple first images corresponding to the same third timestamp in the multiple updated first image sets corresponding to the multiple first perspectives.

[0160] In this embodiment, each updated first image set includes a plurality of first images corresponding to a plurality of third timestamps, wherein the third timestamp includes a first timestamp and a second timestamp.

[0161] Exemplarily, after perspective interpolation, four new first perspectives are inserted, forming nine first perspectives together with the original five first perspectives. After temporal interpolation processing is performed on the first image set corresponding to any first perspective, an updated first image set corresponding to that first perspective is obtained. The updated first image set includes 20 first images corresponding to 20 third timestamps. Based on the nine first images in the nine updated first image sets corresponding to the nine first perspectives that belong to the same third timestamp, an updated Gaussian point cloud corresponding to the third timestamp is determined, resulting in 20 updated Gaussian point clouds corresponding to the 20 third timestamps.

[0162] Step S2400.2: For any first perspective among the multiple first perspectives, determine a reference image set corresponding to the first perspective based on the updated Gaussian point clouds of multiple third timestamps and the first perspective.

[0163] In this embodiment, a reference image set corresponding to a first perspective includes multiple reference images of the updated Gaussian point cloud at the first perspective at multiple third timestamps. Each third timestamp corresponds to one reference image. A reference image is a two-dimensional image of the updated Gaussian point cloud at the first perspective at the third timestamp.

[0164] Continuing with the above example, after obtaining 20 updated Gaussian point clouds corresponding to 20 third time stamps, for each of the nine first view angles, these 20 updated Gaussian point clouds corresponding to the third time stamps are projected onto that first view angle to obtain a reference image set corresponding to that first view angle. This reference image set includes the 20 reference images corresponding to the 20 third time stamps. Thus, for each of the nine first view angles, nine reference image sets are obtained.

[0165] Step S2400.3, based on the reference images of any two adjacent third timestamps in the reference image set corresponding to the first perspective, optimize the two first images corresponding to any two adjacent third timestamps in the updated first image set corresponding to the first perspective to obtain the optimized first image set corresponding to the first perspective.

[0166] In this embodiment, for any first perspective, the optical flow field is calculated between any two reference images with adjacent third timestamps in the reference image set corresponding to the first perspective. Based on the optical flow field between the two reference images with adjacent third timestamps, pixel optimization is performed on the two first images corresponding to the two adjacent third timestamps in the updated first image set corresponding to the first perspective, thereby obtaining an optimized first image set corresponding to the first perspective. For multiple first perspectives, multiple optimized first image sets are corresponding.

[0167] Exemplarily, the multiple third timestamps are 20 third timestamps, where the 20 third timestamps include 0s and 0.5s, and the multiple first perspectives include perspective A. Then, the reference image set corresponding to perspective A includes 20 reference images corresponding to the 20 third timestamps, and the updated first image set corresponding to perspective A includes 20 first images corresponding to the 20 third timestamps. In this case, an optical flow field is calculated between the two reference images corresponding to 0s and 0.5s in the reference image set corresponding to perspective A, and then pixel optimization is performed on the two first images corresponding to 0s and 0.5s in the updated first image set corresponding to perspective A based on the optical flow field.

[0168] Step S2400.4, based on the multiple optimized first images belonging to the same third timestamp in the multiple optimized first image sets corresponding to the multiple first perspectives, determine the optimized Gaussian point cloud corresponding to the third timestamp, and obtain the four-dimensional Gaussian point cloud corresponding to the target dynamic scene.

[0169] In this embodiment, the four-dimensional Gaussian point cloud includes multiple optimized Gaussian point clouds corresponding to multiple third timestamps, and one third timestamp corresponds to one optimized Gaussian point cloud.

[0170] Exemplarily, the multiple first perspectives are 9 first perspectives, corresponding to 9 optimized first image sets, and the multiple third timestamps are 20 third timestamps. Based on the 9 optimized first images in these 9 optimized first image sets that belong to the same third timestamp, an optimized Gaussian point cloud corresponding to the third timestamp is obtained. For the 20 third timestamps, 20 optimized Gaussian point clouds corresponding to the third timestamps can be obtained. These 20 optimized Gaussian point clouds corresponding to the third timestamps constitute a four-dimensional Gaussian point cloud of the target dynamic scene.

[0171] According to an embodiment of the present application, through the above-mentioned steps S2400.1 to S2400.4, by collecting multiple optimized first images corresponding to multiple first perspectives and belonging to the same third timestamp, the optimized Gaussian point cloud corresponding to the third timestamp is determined, and the four-dimensional Gaussian point cloud corresponding to the target dynamic scene is obtained. This can optimize the visual consistency between different perspectives, so that the four-dimensional Gaussian point cloud of the target dynamic scene changes more smoothly in the time series, avoiding jitter or discontinuity problems caused by the interpolation and reconstruction process, which not only improves the rendering clarity, but also enhances the stability of the target dynamic scene during dynamic changes.

[0172] The inventors discovered that in related art, when performing super-resolution processing, pixel constraints and pre-trained 2D super-resolution models can be introduced through optimization in a high-resolution space to enhance the representation capabilities of Gaussian primitives. However, this method is primarily targeted at static scenes, and suffers from problems such as loss of detail when reconstructing high-resolution dynamic scenes. To address this, the inventors proposed a technical solution for fine-tuning the super-resolution model to mitigate the occurrence of these issues. The details are as follows:

[0173] In some embodiments, the super-resolution model in step S3300 or step S2300.5 is determined by the following steps S4100 and S4200:

[0174] Step S4100: Obtain a training image set.

[0175] In this embodiment, the training image set includes a high-resolution image set and a low-resolution image set. The high-resolution image set includes multiple first images corresponding to the same initial first timestamp in multiple first image sets. The initial first timestamp is the earliest timestamp among the multiple first timestamps. The low-resolution image set is an image set rendered from multiple first perspectives based on the Gaussian point cloud corresponding to the initial first timestamp. The high-resolution image set can be referred to as a target high-resolution image set, i.e., an image set obtained by super-resolution processing the low-resolution image set using a super-resolution model. The target high-resolution image set can be the target output of the super-resolution model, used to evaluate the model's performance and optimize model training.

[0176] In one example, there are five first perspectives (i.e., the first perspectives do not include new first perspectives generated after perspective interpolation), which correspond to five first image sets. The number of first timestamps is 15 first timestamps, and the initial first timestamp is the earliest timestamp among the 15 first timestamps. At this time, the five first images corresponding to the initial first timestamp in the five first image sets are used as five high-resolution images to obtain a high-resolution image set. Based on the five first images corresponding to the initial first timestamp in the five first image sets, a Gaussian point cloud corresponding to the initial first timestamp is obtained. The Gaussian point cloud corresponding to the initial first timestamp is then rendered from five first perspectives to obtain five two-dimensional rendered images. These five two-dimensional rendered images are used as five low-resolution images to obtain a low-resolution image set.

[0177] In another example, the number of multiple first perspectives is 9 (i.e., the multiple first perspectives include a new first perspective generated after perspective interpolation), which corresponds to 9 first image sets. The number of first timestamps is 15 first timestamps, and the initial first timestamp is the earliest timestamp among the 15 first timestamps. At this time, the 9 first images corresponding to the initial first timestamp in the 9 first image sets are used as 9 high-resolution images to obtain a high-resolution image set. Based on the 9 first images corresponding to the initial first timestamp in the 9 first image sets, a Gaussian point cloud corresponding to the initial first timestamp is obtained. The Gaussian point cloud corresponding to the initial first timestamp is then rendered from 9 first perspectives to obtain 9 two-dimensional rendered images, which are used as 9 low-resolution images to obtain a low-resolution image set.

[0178] Step S4200: training a super-resolution model using the training image set to obtain a trained super-resolution model.

[0179] In this step, the super-resolution model is trained using conventional model training methods, which will not be elaborated here.

[0180] In some embodiments, step S4200 trains the super-resolution model using the training image set to obtain the trained super-resolution model, including steps S4200.1 to S4200.3.

[0181] Step S4200.1: Input the low-resolution image set into a super-resolution model to obtain a predicted high-resolution image set.

[0182] In this embodiment, the super-resolution model performs super-resolution processing on each low-resolution image in the input low-resolution image set to obtain a predicted high-resolution image corresponding to the low-resolution image.

[0183] In an example where a low-resolution image set includes 5 low-resolution images corresponding to a first perspective, and a high-resolution image set includes 5 high-resolution images corresponding to the first perspective, super-resolution processing is performed on each low-resolution image in the low-resolution image set to obtain 5 expected high-resolution images corresponding to the 5 low-resolution images, which can also be called 5 expected high-resolution images corresponding to the 5 first perspectives.

[0184] Step S4200.2: construct a loss function based on the predicted high-resolution image set and the high-resolution image set.

[0185] In this embodiment, the expected high-resolution images in the expected high-resolution image set are images predicted by the super-resolution model, and the high-resolution images in the high-resolution image set are real high-resolution images, that is, target high-resolution images.

[0186] The loss function is used to characterize the difference between the expected high-resolution image and the target high-resolution image.

[0187] Continuing with the above example, after obtaining the five expected high-resolution images corresponding to the five first perspectives, for each first perspective, the loss function corresponding to the first perspective is obtained based on the expected high-resolution image of the first perspective and the high-resolution image corresponding to the first perspective. For the five first perspectives, five loss functions corresponding to the five first perspectives can be obtained.

[0188] In some embodiments, the loss function includes at least one of: L1 loss, L2 loss, generator loss, perceptual loss, brightness loss, and style loss.

[0189] In this embodiment, the L1 loss and the L2 loss are used to characterize the pixel value difference between the expected high-resolution image and the target high-resolution image.

[0190] The generator loss is a loss function for the generator in a generative adversarial network (GAN). It is used to characterize the difference between the expected high-resolution image generated by the generator and the target high-resolution image, as well as the generator's ability to deceive the discriminator. The generator loss can enhance the texture sharpness of the expected high-resolution image and optimize the local details of the expected high-resolution image, making the generated expected high-resolution image more visually realistic.

[0191] Perceptual loss is used to characterize the difference between the predicted high-resolution image and the target high-resolution image in feature space. It can improve the detail quality of the predicted high-resolution image, making the generated predicted high-resolution image closer to the target high-resolution image in terms of global structure, reducing pixel-level errors, and optimizing potential artifacts and distortion.

[0192] The brightness loss is used to characterize the brightness difference between the predicted high-resolution image and the target high-resolution image. This loss can address the brightness bias that can occur when training a super-resolution model based solely on L1 and L2 losses, making the predicted high-resolution image more visually natural and avoiding overbrightness or darkness.

[0193] Style loss is used to characterize the style difference between the predicted high-resolution image and the target high-resolution image. Style loss can improve the overall visual quality of the predicted high-resolution image, making the predicted high-resolution image closer in style to the target high-resolution image.

[0194] Step S4200.3: Adjust the model parameters of the super-resolution model according to the loss function to obtain the trained super-resolution model.

[0195] By including a high-resolution image set and a low-resolution image set as the training image set, the high-resolution image set includes multiple first images corresponding to the same initial first timestamp in multiple first image sets, and the low-resolution image set is an image set rendered under multiple first perspectives based on the Gaussian point cloud corresponding to the initial first timestamp, the richness of the training samples of the super-resolution model can be increased. Moreover, by adding the super-resolution model trained with such training image sets, after super-resolution processing is performed on the two-dimensional low-resolution image rendered by the Gaussian point cloud, the problem of detail loss can be reduced, and the reconstruction effect of dynamic scenes can be improved.

[0196] By determining whether to perform time interpolation processing on the first images of any two adjacent first timestamps in any first image set, and inserting a second image between the first images corresponding to any two adjacent first timestamps in the first image set corresponding to each first perspective when it is determined that time interpolation processing is to be performed on any two adjacent first timestamps, it is possible to achieve that when an object in the target dynamic scene moves faster, the second image is inserted into the two first images corresponding to the faster movement of the object in the first image set corresponding to each first perspective, so as to obtain an updated first image set corresponding to the first perspective. Moreover, since the second image is inserted into the two first images with faster movement of the object in the updated first image set, the movement of the object presented in the three-dimensional reconstruction of the target dynamic scene can be made smoother, thereby avoiding the problems of blur and texture loss that are easy to occur when processing moving areas and complex backgrounds, and improving the visual quality of the three-dimensional reconstruction of the target dynamic scene.

[0197] <Storage Medium Embodiment>

[0198] An embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in any of the above method embodiments are implemented.

[0199] <Device Example>

[0200] Figure 3 FIG. 3 is a structural block diagram of an electronic device 3000 according to an embodiment of the present invention.

[0201] In this embodiment, if Figure 3 As shown, the electronic device 3000 includes a memory 3001 and a processor 3002, wherein the memory 3001 is used to store executable instructions, and the processor 3002 is used to operate according to the control of the instructions to execute the method described in any of the above embodiments.

[0202] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0203] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0204] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0205] The computer program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. The computer readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), is personalized by utilizing the state information of the computer readable program instructions, and the electronic circuit can execute the computer readable program instructions, thereby realizing various aspects of the present invention.

[0206] Various aspects of the present invention are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0207] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0208] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0209] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of an instruction, and the module, program segment or part of the instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are all equivalent.

[0210] While various embodiments of the present invention have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A dynamic scene reconstruction method, characterized in that: The method comprises: For any first perspective among the multiple first perspectives, obtaining a first image set obtained by photographing the target dynamic scene at multiple first time stamps under the first perspective, to obtain multiple first image sets corresponding to the multiple first perspectives; For a first image set corresponding to any first perspective, determining whether to perform time interpolation processing on any two first images with adjacent first timestamps in the first image set; When it is determined that temporal interpolation processing is performed on any two first images with adjacent first timestamps, for each first perspective among the multiple first perspectives, a second image corresponding to a second timestamp at the first perspective is inserted between any two first images with adjacent first timestamps in the first image set corresponding to the first perspective, to obtain an updated first image set corresponding to the first perspective; wherein the second timestamp is a transition timestamp between the any two adjacent first timestamps; A four-dimensional Gaussian point cloud corresponding to the target dynamic scene is determined according to the multiple updated first image sets corresponding to the multiple first perspectives.

2. The method according to claim 1, characterized in that Inserting the second image corresponding to the second timestamp at the first perspective between any two first images with adjacent first timestamps in the first image set corresponding to the first perspective to obtain an updated first image set corresponding to the first perspective includes: Determine a Gaussian point cloud corresponding to the first timestamp based on the multiple first images corresponding to the same first timestamp in the multiple first image sets; determining a motion trajectory of the Gaussian point cloud according to the plurality of Gaussian point clouds corresponding to the plurality of first timestamps; determining a Gaussian point cloud corresponding to the second timestamp according to the second timestamp and the motion trajectory of the Gaussian point cloud; Determining, based on the Gaussian point cloud corresponding to the second timestamp and the first viewing angle, a two-dimensional image of the Gaussian point cloud of the second timestamp at the first viewing angle as a third image corresponding to the second timestamp at the first viewing angle; Inputting the third image corresponding to the second timestamp at the first viewing angle into a super-resolution model to obtain a second image corresponding to the second timestamp at the first viewing angle; wherein the second image resolution of the second image is greater than the third image resolution of the third image; According to the second timestamp corresponding to the second image, the second image is inserted as the first image into the first image set corresponding to the first perspective to obtain an updated first image set corresponding to the first perspective.

3. The method according to claim 1, characterized in that The determining, based on any two first images with adjacent first timestamps in the first image set, whether to perform time interpolation processing on the any two first images with adjacent first timestamps includes: Calculating the motion vector of each pixel in the first images of any two adjacent first time stamps by an optical flow method; When the number of pixel points of the target motion vector in the first images of any two adjacent first timestamps is greater than or equal to a number threshold, it is determined to perform time interpolation processing on the first images of any two adjacent first timestamps; wherein the target motion vector is a motion vector greater than or equal to a motion vector threshold.

4. The method according to claim 1, wherein Before determining whether to perform time interpolation processing on any two first images with adjacent first timestamps in the first image set, the method further includes: Determining a second perspective corresponding to any two adjacent first perspectives from the plurality of first perspectives, thereby obtaining a plurality of second perspectives; wherein the second perspective is a transition perspective between the first perspectives adjacent to the any two adjacent perspectives; For any second perspective, determining a fourth image corresponding to the second perspective at the first timestamp based on the Gaussian point cloud corresponding to the second perspective and any first timestamp; Inputting the fourth image of the second perspective at the first timestamp into a super-resolution model to obtain a fifth image of the second perspective at the first timestamp; wherein a fifth image resolution corresponding to the fifth image is greater than a fourth image resolution corresponding to the fourth image; A first image set corresponding to the first perspective is determined according to the fifth image of the second perspective at each first time stamp in the plurality of first time stamps.

5. The method according to claim 1, wherein Each of the plurality of first perspectives is determined by a camera posture, and determining a second perspective corresponding to any two adjacent first perspectives among the plurality of first perspectives according to the first perspectives includes: Determining, based on a camera pose corresponding to each of the plurality of first perspectives, a plurality of adjacent first perspectives among the plurality of first perspectives; For any two adjacent first perspectives among the multiple adjacent first perspectives, determine the second perspective corresponding to the any two adjacent first perspectives based on the camera postures corresponding to the any two adjacent first perspectives, and obtain multiple second perspectives corresponding to the multiple adjacent first perspectives.

6. The method according to claim 5, characterized in that The camera pose includes a rotation matrix and a translation vector, and determining the second perspective corresponding to any two adjacent first perspectives according to the camera pose corresponding to the first perspective of any two adjacent perspectives includes: Inserting a transition rotation matrix into the rotation matrices corresponding to the first perspectives adjacent to any two perspectives by spherical interpolation; Inserting a transition translation vector into the translation vectors corresponding to the first perspectives adjacent to any two perspectives by linear interpolation; A second perspective corresponding to the first perspectives adjacent to any two perspectives is determined according to the transition rotation matrix and the transition translation vector.

7. The method according to claim 2 or 4, characterized in that The super-resolution model is determined by the following steps: Acquire a training image set; wherein the training image set includes a high-resolution image set and a low-resolution image set, the high-resolution image set includes multiple first images corresponding to the same initial first timestamp in the multiple first image sets, and the low-resolution image set is an image set rendered under the multiple first perspectives based on the Gaussian point cloud corresponding to the initial first timestamp; The super-resolution model is trained using the training image set to obtain a trained super-resolution model.

8. The method according to claim 7, characterized in that The step of training the super-resolution model using the training image set to obtain the trained super-resolution model includes: Inputting the low-resolution image set into a super-resolution model to obtain a predicted high-resolution image set; constructing a loss function based on the predicted high-resolution image set and the high-resolution image set; According to the loss function, the model parameters of the super-resolution model are adjusted to obtain the trained super-resolution model.

9. The method according to claim 1, characterized in that The determining, based on the multiple updated first image sets corresponding to the multiple first perspectives, a four-dimensional Gaussian point cloud corresponding to the target dynamic scene includes: Determining an updated Gaussian point cloud corresponding to a third timestamp based on a plurality of first images corresponding to a same third timestamp in a plurality of updated first image sets corresponding to the plurality of first perspectives; wherein the third timestamp includes the first timestamp and the second timestamp; For any first perspective among the multiple first perspectives, determining a reference image set corresponding to the first perspective based on the updated Gaussian point clouds of the multiple third timestamps and the first perspective; optimizing, based on any two reference images with adjacent third time stamps in the reference image set corresponding to the first perspective, two first images corresponding to any two adjacent third time stamps in the updated first image set corresponding to the first perspective, to obtain an optimized first image set corresponding to the first perspective; According to the multiple optimized first images belonging to the same third timestamp in the multiple optimized first image sets corresponding to the multiple first perspectives, the optimized Gaussian point cloud corresponding to the third timestamp is determined to obtain a four-dimensional Gaussian point cloud corresponding to the target dynamic scene.

10. An electronic device comprising a memory and a processor, wherein the memory is configured to store executable instructions; and the processor is configured to operate under the control of the instructions to execute the method according to any one of claims 1 to 9.