Three-dimensional object interpolation method based on low-rank semantic sampling

By rendering three-dimensional objects as multi-view images and utilizing a multi-view diffusion model with a low-rank adapter, the problem of insufficient utilization of semantic information in three-dimensional interpolation generation is solved, and a semantically reasonable and geometrically smooth interpolation effect is achieved.

CN120689480APending Publication Date: 2025-09-23SHANGHAI TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510759223.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-23

Smart Images

  • Figure CN120689480A_ABST
    Figure CN120689480A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional object interpolation method based on low-rank semantic sampling. The method comprises the following steps: acquiring two three-dimensional objects to be interpolated and corresponding text descriptions thereof; rendering each three-dimensional object into a multi-view image; inverting the multi-view image to obtain potential noise, and interpolating the potential noise according to a preset interpolation coefficient; calculating an embedded vector described by the text, and interpolating the embedded vector according to the interpolation coefficient; generating an interpolated multi-view image by using a multi-view diffusion model integrated with a low-rank adapter according to the interpolated potential noise, the embedded vector and the multi-view image; and obtaining a three-dimensional object sequence between the two three-dimensional objects according to the interpolated multi-view image. An interpolation multi-view image with reasonable semantics and smooth geometric and texture transition is directly generated through a multi-view diffusion model integrated with a low-rank adapter, and a three-dimensional object sequence obtained based on the interpolation multi-view image is natural, more real, reasonable and smooth in transition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional object interpolation, and in particular to a three-dimensional object interpolation method based on low-rank semantic sampling. Background Art

[0002] With the rapid development of deep learning technology, diffusion-based image generation models have achieved remarkable success in the field of two-dimensional image processing, leveraging their exceptional generative capabilities. These models are capable of generating highly realistic and detailed images, representing a significant advancement in image generation technology. Furthermore, the continuous emergence of large-scale datasets of three-dimensional objects has provided a solid data foundation for research in the field of 3D generation. Researchers can leverage this rich data to conduct deeper exploration and experimentation, further advancing 3D generation technology.

[0003] Against this backdrop, the field of 3D generation has seen significant breakthroughs, with 3D interpolation generation becoming a hot research topic. The core goal of this task is to generate an interpolation result that represents a transition between two 3D objects and their corresponding textual descriptions. This transition must not only exhibit smooth geometric changes but also maintain semantic coherence and plausibility.

[0004] In practical applications, 3D interpolation has a wide range of uses. For example, in film and television production, this technology can be used to generate transition effects between characters and scenes with varying shapes, enhancing the visual impact of the images. In game development, it can achieve natural transitions between different models, enriching the visual experience. In product design, designers can use 3D interpolation technology to quickly explore different product designs and find the optimal design solution.

[0005] Currently, for 3D interpolation generation tasks, most methods still rely on traditional graphics interpolation techniques. These methods are mainly based on geometric transformations and mathematical calculations, and achieve interpolation between objects by operating on geometric elements such as vertices, edges, and faces of 3D objects. However, this traditional method has obvious limitations in utilizing semantic information. In practical applications, 3D objects are not only composed of geometric shapes, but also contain rich semantic information, such as the category, function, and attributes of the object. When faced with large differences in 3D interpolation requirements, traditional methods cannot effectively utilize this semantic information, which easily leads to inconsistent generation results at the semantic level. For example, when interpolating two objects of different categories, the generated intermediate result may be semantically confusing, neither resembling the previous object nor the next object, and cannot meet the needs of practical applications. Summary of the Invention

[0006] In view of the above-mentioned defects of the prior art, the present invention provides a three-dimensional object interpolation method based on low-rank semantic sampling to solve the technical problems of unnatural geometric and texture transitions and loss of semantic and identity features in the existing three-dimensional object interpolation technology.

[0007] To achieve the above-mentioned purpose and other related purposes, the present invention provides a three-dimensional object interpolation method based on low-rank semantic sampling, comprising: obtaining two three-dimensional objects to be interpolated and their corresponding text descriptions, the three-dimensional objects including a first three-dimensional object and a second three-dimensional object; rendering each of the three-dimensional objects as a multi-view image; inverting the multi-view image to obtain potential noise, and interpolating the potential noise according to a preset interpolation coefficient; calculating an embedding vector of the text description, and interpolating the embedding vector according to the interpolation coefficient; generating an interpolated multi-view image based on the interpolated potential noise and embedding vector, and the multi-view image using a multi-view diffusion model integrated with a low-rank adapter; and obtaining a three-dimensional object sequence between the two three-dimensional objects based on the interpolated multi-view image.

[0008] In one embodiment of the present invention, each of the three-dimensional objects is rendered as a multi-view image, including: normalizing the three-dimensional object; rendering the normalized three-dimensional object according to preset parameters to obtain images at different azimuths and different horizontal angles; and obtaining a set of multi-view images based on four images at the same horizontal angle and in a preset order of azimuths.

[0009] In one embodiment of the present invention, the multi-view images are inverted to obtain potential noise, and the potential noise is interpolated according to preset interpolation coefficients, including: inverting the multi-view images of the first three-dimensional object to obtain first potential noise; inverting the multi-view images of the second three-dimensional object to obtain second potential noise; and interpolating the first potential noise and the second potential noise by spherical linear interpolation method according to the interpolation coefficients.

[0010] In one embodiment of the present invention, an embedding vector of the text description is calculated, and the embedding vector is interpolated according to the interpolation coefficient, including: calculating a first embedding vector of the text description of the first three-dimensional object; calculating a second embedding vector of the text description of the second three-dimensional object; and interpolating the first embedding vector and the second embedding vector by linear interpolation according to the interpolation coefficient.

[0011] In one embodiment of the present invention, the low-rank adapter includes a first low-rank adapter and a second low-rank adapter, and the low-rank adapter is used to fine-tune the residual component corresponding to the attention layer parameters of the multi-view diffusion model.

[0012] In one embodiment of the present invention, an interpolated multi-view image is generated based on the interpolated potential noise and embedding vector, and the multi-view image, using a multi-view diffusion model integrated with a low-rank adapter, including: using multiple groups of multi-view images of the first three-dimensional object to train the first low-rank adapter, and obtaining first low-rank parameters after the training is completed; using multiple groups of multi-view images of the second three-dimensional object to train the second low-rank adapter, and obtaining second low-rank parameters after the training is completed; interpolating the first low-rank parameters and the second low-rank parameters by linear interpolation according to the interpolation coefficients; and generating the interpolated multi-view image based on the interpolated potential noise, embedding vector, and low-rank parameters using the multi-view diffusion model.

[0013] In one embodiment of the present invention, the first low-rank adapter is trained using multi-view images of the first three-dimensional object, including: in the step of training the first low-rank adapter using the multi-view images of the first three-dimensional object, in each training iteration, a random number is generated and it is determined whether the random number is less than a preset trade-off coefficient: if so, the first low-rank adapter is trained using the mean of the initial latent distribution output by the variational autoencoder; if not, the first low-rank adapter is trained using the result sampled from the initial latent distribution output by the variational autoencoder.

[0014] In one embodiment of the present invention, a sequence of three-dimensional objects between two of the three-dimensional objects is obtained based on the interpolated multi-view images, including: obtaining a perceived distance between adjacent frames based on a set of the interpolated multi-view images; optimizing the interpolation coefficients based on the perceived distance; regenerating multiple sets of optimized interpolated multi-view images based on the optimized interpolation coefficients; and obtaining a sequence of three-dimensional objects between the two of the three-dimensional objects based on the multiple sets of optimized interpolated multi-view images.

[0015] In one embodiment of the present invention, optimizing the interpolation coefficient according to the perceived distance includes: obtaining a total perceived distance according to the perceived distances of all adjacent frames; obtaining a gradient of the relative perceived distance with respect to the interpolation coefficient according to the perceived distance of each adjacent frame and the total perceived distance; obtaining a relative perceived distance corresponding to each interpolation coefficient according to the gradient of the relative perceived distance with respect to the interpolation coefficient; and obtaining an optimized interpolation coefficient according to the relative perceived distances corresponding to all the interpolation coefficients.

[0016] In one embodiment of the present invention, a three-dimensional object sequence between two of the three-dimensional objects is obtained based on the interpolated multi-view images, including: using a four-dimensional Gaussian sputtering reconstruction interpolation process based on multiple groups of the interpolated multi-view images to obtain a spatiotemporally continuous decomposable three-dimensional Gaussian representation; and decomposing the spatiotemporally continuous decomposable three-dimensional Gaussian representation into a three-dimensional object sequence under the interpolation coefficients.

[0017] Beneficial effects of the present invention: The present invention proposes a three-dimensional object interpolation method based on low-rank semantic sampling. The method inverts the rendered multi-view images to obtain potential noise and calculates interpolation. At the same time, the corresponding interpolation is also calculated for the text description. Based on these interpolations, a multi-view diffusion model integrated with a low-rank adapter is used to directly generate semantically reasonable, geometrically and texture-smooth interpolated multi-view images. The three-dimensional object sequence obtained based on this is naturally more realistic and reasonable, and the transition is smooth. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 A flowchart of a three-dimensional object interpolation method provided by one embodiment of the present invention;

[0020] Figure 2 A schematic diagram of a three-dimensional object interpolation method provided by an embodiment of the present invention;

[0021] Figure 3 A detailed flow chart of step S200 provided in one embodiment of the present invention;

[0022] Figure 4 A detailed flow chart of step S300 provided in one embodiment of the present invention;

[0023] Figure 5 A detailed flowchart of step S400 provided in one embodiment of the present invention;

[0024] Figure 6 A detailed flow chart of step S500 provided in one embodiment of the present invention;

[0025] Figure 7 A detailed flow chart of interpolation coefficient optimization in step S600 provided in one embodiment of the present invention

[0026] Figure 8 A detailed flowchart of step S620 provided in one embodiment of the present invention;

[0027] Figure 9 This is a flowchart of generating a three-dimensional object sequence in step S600 according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The following describes the embodiments of the present invention through specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. It should be noted that the following embodiments and the features in the embodiments can be combined with each other unless they conflict. In addition to the specific methods, equipment, and materials used in the embodiments, based on the understanding of the prior art by those skilled in the art and the description of the present invention, any methods, equipment, and materials of the prior art that are similar or equivalent to the methods, equipment, and materials in the embodiments of the present invention can also be used to implement the present invention.

[0029] It should be understood that the terms used in the examples of the present invention are for describing specific embodiments rather than for limiting the scope of protection of the present invention. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those generally understood by those skilled in the art.

[0030] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In some of the embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0031] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functions and operations that may be implemented by the methods and computer program products of various embodiments disclosed in the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0032] See Figure 1 , Figure 1 An embodiment of the present invention provides a three-dimensional object interpolation method based on low-rank semantic sampling, including steps S100 to S600.

[0033] Step S100: Get two 3D objects to be interpolated and their corresponding text descriptions. One of the two 3D objects is used as the source object, which is recorded as the first 3D object. Figure 2 The corresponding text description of Object A is, for example, "a man's head"; the other is the target object, recorded as the second three-dimensional object, corresponding to Figure 2 For example, the corresponding text description of Object B is “Sculpture of a man's head made of black stone”. In the present invention, it is necessary to perform interpolation from the source object to the target object to obtain the intermediate transition state of the two three-dimensional objects.

[0034] Step S200: Render each three-dimensional object into a multi-view image.

[0035] See Figure 3 In a specific embodiment of the present invention, step S200 includes steps S201 to S203.

[0036] Step S201: Normalize the 3D objects. Normalization unifies the spatial coordinates and scales of two 3D objects into a standard reference frame, ensuring comparability of different objects under the same camera parameters. It also ensures consistency in perspective rendering and alignment in subsequent feature spaces, eliminating the impact of differences in the original object sizes on low-rank adapter training.

[0037] Step S202: Render the normalized three-dimensional object according to preset parameters to obtain images at different azimuth angles and horizontal angles. The preset parameters may be, for example, rendering with a radius of 2.5 and a field of view of 49.1°. During rendering, the azimuth angle φ may, for example, vary from 0° to 360° in increments of 10°, and the horizontal angle θ may, for example, be selected from the set {-15°, 0°, 15°}, so that almost all details of the 3D object can be captured. At each azimuth angle and each horizontal angle, an image of a three-dimensional object may be obtained. Taking the above parameter selection as an example, 36×3=108 images may be obtained after rendering of the first three-dimensional object. Similarly, 108 images may also be obtained for the second three-dimensional object.

[0038] Step S203: Obtain a set of multi-view images based on four images at the same horizontal angle and in a preset order of azimuth angles. The preset order of azimuth angles may be, for example, φ-90°, φ, φ+90°, and φ+180°. In this way, the above 108 images can be used to obtain 27 sets of multi-view images, each of which consists of four images. For example, four images with azimuth angles of -90°, 0°, 90°, and 180° and a horizontal angle of 0° (based on which camera parameters, i.e., the camera extrinsic parameter matrix, can be obtained) constitute a set of multi-view images; four images with azimuth angles of -70°, 20°, 110°, and 200° and a horizontal angle of 15° can also constitute a set of multi-view images.

[0039] In the subsequent training process of the low-rank adapter, these 27 sets of multi-view images will be used. During the latent noise inversion and interpolated multi-view image generation, if only one set of multi-view images is used, the azimuth and horizontal angles of the finally generated interpolated multi-view image will correspond to the azimuth and horizontal angles of the set of multi-view images used; if 27 sets of multi-view images are used, the finally generated interpolated multi-view image will also correspond to multiple azimuths and horizontal angles. In the following embodiment, two generations are mentioned. The first time, only one set of multi-view images is randomly selected (generally, four images with azimuths of -90°, 0°, 90°, 180° and a horizontal angle of 0° are directly selected). The camera parameters of the finally generated interpolated multi-view image will correspond to it, that is, the azimuths of the generated interpolated multi-view image are -90°, 0°, 90°, 180° and the horizontal angle is 0°. Then, the interpolation coefficients are optimized based on the multiple sets of interpolated multi-view images under the azimuth and horizontal angles. The second time, all 27 sets of multi-view images will be selected for noise inversion and optimized interpolation multi-view image generation, so that the final generated three-dimensional object sequence will be more accurate.

[0040] Step S300: Invert the multi-view image to obtain potential noise, and interpolate the potential noise according to a preset interpolation coefficient.

[0041] See Figure 4 In a specific embodiment of the present invention, step S300 includes steps S301 to S303.

[0042] Step S301: Invert the multi-view image of the first three-dimensional object to obtain the first potential noise. During the first inversion, take the image of the first three-dimensional object with azimuth angles of -90°, 0°, 90°, 180° and horizontal angle of 0° as an example, and its multi-view image is recorded as MV0. During the inversion, the denoising diffusion implicit model (corresponding to Figure 2 The DDIM in

[15] is used as a sampler, which is more stable than DDPM during inversion and interpolation. Specifically, the inverse noise z of MV0 can be obtained by inverting the denoising diffusion implicit model in the multi-view diffusion model integrated with the low-rank adapter. T0 During the second inversion, all 27 sets of multi-view images of the first three-dimensional object need to be inverted.

[0043] Step S302: Invert the multi-view image of the second 3D object to obtain the second potential noise. During the first inversion, the multi-view image of the second 3D object also selects images with azimuth angles of -90°, 0°, 90°, 180° and horizontal angle of 0°. The multi-view image is recorded as MV1. After inversion, the reverse noise z of MV1 can be obtained. T1 During the second inversion, all 27 sets of multi-view images of the second three-dimensional object need to be inverted.

[0044] Step S303: Interpolate the first potential noise and the second potential noise using spherical linear interpolation based on the interpolation coefficient. The interpolation coefficient is obtained based on the interpolation order. The interpolation coefficient is in the range [0, 1], including the two endpoints 0 and 1. For example, if the interpolation order is 11, the interpolation coefficient α is {0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1}.

[0045] The formula for spherical linear interpolation is as follows:

[0046]

[0047] Where φ in the first formula is calculated using the second formula, which is different from the previous azimuth. For z T0 The transpose of is the vector inner product of the reverse noise, and |||| is the L2 norm. After the calculation of the above formula, the interpolated potential noise z of the first potential noise and the second potential noise can be obtained Tα .

[0048] Step S400: Calculate the embedding vector of the text description and interpolate the embedding vector according to the interpolation coefficient.

[0049] See Figure 5 In a specific embodiment of the present invention, step S400 includes: S401, calculating a first embedding vector c0 of the text description of the first three-dimensional object; S402, calculating a second embedding vector c1 of the text description of the second three-dimensional object; S403, interpolating the first embedding vector and the second embedding vector by linear interpolation according to the interpolation coefficient. The interpolated text embedding vector c α =(1-α)c0+αc1.

[0050] Step S500 : generating an interpolated multi-view image based on the interpolated potential noise and the embedding vector, and the multi-view image, using a multi-view diffusion model integrated with a low-rank adapter.

[0051] In a specific embodiment of the present invention, a low-rank adapter includes a first low-rank adapter and a second low-rank adapter, and the low-rank adapter is used to fine-tune the residual components corresponding to the attention layer parameters of the multi-view diffusion model. The low-rank adapter is an efficient fine-tuning method that can replenish visual semantics and identity information using lower computing resources. In the present invention, the low-rank adapter does not fine-tune all parameters of the multi-view diffusion model, but only adjusts the residual components corresponding to the attention layer parameters of the multi-view diffusion model. These residual parameters can be decomposed into a combination of low-rank matrices.

[0052] See Figure 6In a specific embodiment of the present invention, step S500 includes steps S501 to S504.

[0053] Step S501: train a first low-rank adapter using multiple sets of multi-view images of a first three-dimensional object, and obtain first low-rank parameters after the training is completed.

[0054] Step S502: train a second low-rank adapter using multiple sets of multi-view images of the second three-dimensional object, and obtain second low-rank parameters after the training is completed.

[0055] In the above two steps, multiple sets of multi-view images are used to train the low-rank adapter. After training, the first low-rank parameter Δθ0 and the second low-rank parameter Δθ1 can be obtained. The loss function formula during training is as follows:

[0056]

[0057] Where ∈ is random noise, usually standard Gaussian noise ∈~N(0,1); t is the time step of the diffusion process; z 0i =E(MV i ) is E(MV obtained using variational autoencoder i ), the variational autoencoder corresponds to Figure 2 In the present invention, the multi-view image needs to be processed by a variational autoencoder to obtain a potential distribution before it can be used for training the low-rank adapter; c i Condition information (such as text prompts or camera parameters), including the text prompt P i Encoded text embedding; is the noisy latent embedding at diffusion step t, α t is the noise scheduling coefficient (decreasing with t); ∈ θ is a denoising model with parameter θ; Δθ i is the trainable parameter of the low-rank adapter, i can be 0 or 1, corresponding to the first low-rank parameter Δθ0 and the second low-rank parameter Δθ1, respectively. The loss function is achieved by minimizing the prediction noise ∈ θ The mean square error with the true noise ∈, optimize the parameters Δθ of the low-rank adapter i .

[0058] In a specific embodiment of the present invention, in the step of training a first low-rank adapter using multiple groups of multi-view images of a first three-dimensional object, in each training iteration, a random number is generated and it is determined whether the random number is less than a preset trade-off coefficient: if so, the first low-rank adapter is trained using the mean of the initial latent distribution output by the variational autoencoder; if not, the first low-rank adapter is trained using the result sampled from the initial latent distribution output by the variational autoencoder.

[0059] From z 0i Sampling latent codes from a distribution to optimize the neighboring latent codes of the target image’s latent code is effective in preserving semantic and identity features. However, this approach often fails when processing multi-view images. For example, in the case of 3D face objects, semantic and identity features are mainly concentrated in the front view, while the back view lacks semantic features and clear identity representation. In previous methods, this problem causes the back view image to be optimized to have obvious semantic and identity features similar to the front view image. By using the 0i In-distribution sampling as the encoding feature of multi-view images can enhance the consistency of adjacent latent codes in terms of semantics and identity features. In addition, using z 0i The mean of the distribution can strengthen the interpolation result and ensure that it is aligned with the correct back view. To achieve this goal, a trade-off coefficient β is introduced to maintain the correct back view image and smooth interpolation effect.

[0060] In a specific embodiment of the present invention, β can be set to any value between 0.2 and 0.4, for example, it can be set to 0.3. In this case, in each optimization step, there is a 30% probability of using z 0i The mean of the distribution as a result of the variational autoencoder has a 70% probability of using the value from z 0i The result obtained by sampling from the distribution is used as the output of the variational autoencoder.

[0061] Step S503: interpolate the first low-rank parameter and the second low-rank parameter using linear interpolation according to the interpolation coefficient. For any interpolation coefficient α, interpolation can be performed as follows:

[0062] Δθ α =(1-α)Δθ0+αΔθ1;

[0063] Since the two low-rank parameters effectively capture the semantic and identity information, the interpolated low-rank parameters can successfully maintain a balanced transition in semantics and identity.

[0064] Step S504: Generate an interpolated multi-view image using the multi-view diffusion model based on the interpolated potential noise, embedding vector, and low-rank parameter. Tα , embedding vector c α , low-rank parameter Δθ α , then the interpolated multi-view images can be generated based on these interpolation results.

[0065] The multi-view diffusion model can generate consistent multi-view images. It combines 2D and 3D data to achieve the generalization of the 2D diffusion model and the consistency of 3D rendering. Its core idea is to implicitly learn a general 3D prior knowledge that is independent of the 3D representation and can be applied to 3D generation through fractional distillation sampling (SDS), enhancing the consistency and stability of existing 2D enhancement methods.

[0066] Given a set of noisy images x t ∈R F×H×W×C (corresponding to the potential noise z after interpolation Tα ), as the conditional text prompt y (corresponding to the interpolated embedding vector c α ) and a set of external camera parameters c∈R F×16 (camera extrinsic matrix), the multi-view diffusion model is trained to generate a set of images x0∈R of the same scene from F different perspectives F×H×W×C The training loss function of the multi-view diffusion model is defined as: Where θ is the parameter of the multi-view diffusion model; X is the real multi-view image; X mv is the multi-view image generated by the model; t is the diffusion time step, which is used to control the noise scheduling intensity.

[0067] The multi-view diffusion model is a latent diffusion model that encodes the input image into a latent code and adds noise. In the latent space, the denoising U-Net∈ θ It consists of a series of basic blocks, each of which contains a self-attention module, a cross-attention module, and a residual block. The attention mechanism in U-Net can be expressed as:

[0068]

[0069] Where Q represents the query feature derived from the spatial feature, K and V correspond to the key and value features obtained from the spatial feature or text embedding, respectively, and their respective projection matrices.

[0070] For each set of multi-view images, multiple sets of interpolated multi-view images can be generated through the multi-view diffusion model. Each set of interpolated multi-view images consists of four pictures with the same camera parameters (for example, azimuth angles of -90°, 0°, 90°, 180° and horizontal angle of 0°), which can be recorded as MV α , specifically including: MV0, MV 1 / n 、MV 2 / n ,…,MV 1-1 / n , MV1, where n is the number of interpolation times minus 1. Taking the above interpolation coefficients as an example, the interpolated multi-view images are MV0, MV 0.1 、MV 0.2 ,…,MV0.9 , MV1. That is, for a set of multi-view images, 11 sets of interpolated multi-view images can be obtained.

[0071] Step S600: Obtain a three-dimensional object sequence between two three-dimensional objects based on the interpolated multi-view images.

[0072] See Figure 7 In a specific embodiment of the present invention, a multi-view diffusion model integrated with a low-rank adapter effectively maintains a smooth transition between semantics and appearance by uniformly interpolating in the low-rank parameter and latent noise spaces. To further ensure that the interpolated three-dimensional object is also smooth in both semantics and appearance, this embodiment optimizes the interpolation coefficients based on the results of the initial interpolation, and regenerates the interpolation result based on the optimized interpolation coefficients. This results in an interpolation result with improved smoothness. Specifically, step S600 includes steps S610 to S640.

[0073] Step S610: Obtain the perceived distance of adjacent frames based on a set of interpolated multi-view images. 1 / n 、MV 2 / n ,…,MV 1-1 / n For example, a large-scale multi-view Gaussian model (LGM) can be used as a feature extractor for three-dimensional objects to obtain a multi-view image MV i and MV i+1 / n The perceived distance D(MV i ,MV i+1 / n ), where i∈{0,1 / n,2 / n,…,1-1 / n}.

[0074] Step S620: Optimize the interpolation coefficient according to the perceived distance.

[0075] See Figure 8 In a specific embodiment of the present invention, step S620 includes steps S621 to S624.

[0076] Step S621: Obtain the total perception distance based on the perception distances of all adjacent frames. This can be expressed as:

[0077]

[0078] Step S622: Obtain the gradient of the relative perception distance with respect to the interpolation coefficient based on the perception distance of each adjacent frame and the total perception distance. This can be expressed as:

[0079]

[0080] Step S623: Obtain the relative perceptual distance corresponding to each interpolation coefficient based on the gradient of the relative perceptual distance with respect to the interpolation coefficient. For any given interpolation coefficient α, the relative perceptual distance to the first frame (i.e., MV0) can be estimated by integrating ΔD(x):

[0081]

[0082] Step S624: According to the relative perception distances corresponding to all interpolation coefficients, the optimized interpolation coefficients are obtained. Using D0 and its inverse function D0′, the re-scheduled interpolation coefficient α can be derived. i , defined as {α i =D0′(y)|y=0,1 / n,…,1}.

[0083] Step S630: Regenerate multiple sets of optimized interpolated multi-view images based on the optimized interpolation coefficients. Specifically, the aforementioned steps are repeated using the optimized interpolation coefficients to recalculate the interpolated potential noise, embedding vector, and low-rank parameters. Multiple sets of optimized interpolated multi-view images are regenerated according to step S504. "Optimized" is added to the name to distinguish them from the initially generated interpolated multi-view images. Specifically, for one set of multi-view images, 11 sets of optimized interpolated multi-view images are generated; for 27 sets of multi-view images, a total of 27*11 sets of optimized interpolated multi-view images are generated.

[0084] Step S640: Interpolate the multi-view images based on the multiple sets of optimized images to obtain a sequence of three-dimensional objects between the two three-dimensional objects. Resampling enhances the interpolation result, shifting the smoothness from low-rank parameters and potential noise to semantic and appearance smoothness.

[0085] See Figure 9 In a specific embodiment of the present invention, step S600 can obtain a 3D object sequence based on multiple sets of interpolated multi-view images, or it can obtain a 3D object sequence based on multiple sets of optimized interpolated multi-view images. The latter method is used as an example below and specifically includes: S601, reconstructing the interpolation process using a four-dimensional Gaussian sputtering technique based on the multiple sets of optimized interpolated multi-view images to obtain a spatiotemporally continuous, decomposable 3D Gaussian representation; S602, decomposing the spatiotemporally continuous, decomposable 3D Gaussian representation into a 3D object sequence at the interpolation coefficients.

[0086] In other words, if the interpolation coefficient optimization step is not introduced, then 27*11 sets of interpolated multi-view images are directly generated, and processing in step S601 can be performed based on the multiple sets of interpolated multi-view images. If the interpolation coefficient optimization step is introduced, the interpolated multi-view images generated for the first time have only 11 sets (i.e., only the potential noise of one multi-view image is calculated). Based on this optimized interpolation coefficient, the optimized interpolated multi-view images are generated for the second time, with the number of sets being 27*11, and processing in step S601 is performed based on the 27*11 sets of optimized interpolated multi-view images.

[0087] Four-dimensional Gaussian sputtering models dynamic scenes using a set of four-dimensional Gaussian basis points, each defined by its mean vector and covariance matrix. These are decomposed into a conditional three-dimensional Gaussian and a marginal one-dimensional Gaussian to capture the spatial and temporal aspects of the scene. After resampling, the optimized interpolation coefficients are non-uniform. During the reconstruction phase, since the appearance of multi-view images is uniform, the non-uniformly sampled interpolation coefficients are matched to the uniformly sampled time t. Four-dimensional Gaussian sputtering is used to reconstruct the interpolation process, resulting in a 3D object sequence with uniform interpolation coefficients during decomposition.

[0088] It should be noted that the step division of the various methods above is only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they contain the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this patent.

[0089] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.

Claims

1. A three-dimensional object interpolation method based on low-rank semantic sampling, characterized in that: include: Obtaining two three-dimensional objects to be interpolated and their corresponding text descriptions, the three-dimensional objects including a first three-dimensional object and a second three-dimensional object; rendering each of the three-dimensional objects into a multi-view image; Inverting the multi-view image to obtain potential noise, and interpolating the potential noise according to a preset interpolation coefficient; Calculating an embedding vector of the text description, and interpolating the embedding vector according to the interpolation coefficient; generating an interpolated multi-view image using a multi-view diffusion model integrated with a low-rank adapter according to the interpolated latent noise and the embedding vector and the multi-view image; A three-dimensional object sequence between the two three-dimensional objects is obtained according to the interpolated multi-view images.

2. The three-dimensional object interpolation method based on low-rank semantic sampling according to claim 1, characterized in that: Rendering each of the three-dimensional objects into a multi-view image, comprising: performing normalization processing on the three-dimensional object; Rendering the normalized three-dimensional object according to preset parameters to obtain images at different azimuth angles and horizontal angles; A set of multi-view images is obtained based on four images at the same horizontal angle and at preset sequential azimuth angles.

3. The three-dimensional object interpolation method based on low-rank semantic sampling according to claim 2, characterized in that: Inverting the multi-view image to obtain potential noise, and interpolating the potential noise according to a preset interpolation coefficient, comprising: inverting the multi-view image of the first three-dimensional object to obtain a first potential noise; inverting the multi-view image of the second three-dimensional object to obtain second potential noise; The first potential noise and the second potential noise are interpolated using a spherical linear interpolation method according to the interpolation coefficient.

4. The three-dimensional object interpolation method based on low-rank semantic sampling according to claim 1, characterized in that: Calculating an embedding vector of the text description and interpolating the embedding vector according to the interpolation coefficient includes: Calculating a first embedding vector of the text description of the first three-dimensional object; calculating a second embedding vector of the text description of the second three-dimensional object; The first embedding vector and the second embedding vector are interpolated by linear interpolation according to the interpolation coefficient.

5. The three-dimensional object interpolation method based on low-rank semantic sampling according to claim 2, characterized in that: The low-rank adapter includes a first low-rank adapter and a second low-rank adapter, and the low-rank adapter is used to fine-tune the residual components corresponding to the attention layer parameters of the multi-view diffusion model.

6. The three-dimensional object interpolation method based on low-rank semantic sampling according to claim 5, characterized in that: The method generates an interpolated multi-view image according to the interpolated potential noise and the embedding vector and the multi-view image using a multi-view diffusion model integrated with a low-rank adapter, comprising: Training the first low-rank adapter using the multiple sets of multi-view images of the first three-dimensional object, and obtaining first low-rank parameters after the training is completed; Training the second low-rank adapter using the multiple sets of multi-view images of the second three-dimensional object, and obtaining second low-rank parameters after the training is completed; interpolating the first low-rank parameter and the second low-rank parameter by linear interpolation according to the interpolation coefficient; The interpolated multi-view image is generated using the multi-view diffusion model according to the interpolated potential noise, the embedding vector, and the low-rank parameter.

7. The three-dimensional object interpolation method based on low-rank semantic sampling according to claim 6, characterized in that: In the step of training the first low-rank adapter using the multiple sets of multi-view images of the first three-dimensional object, generating a random number and determining whether the random number is less than a preset trade-off coefficient in each training iteration: If yes, use the mean of the initial latent distribution output by the variational autoencoder to train the first low-rank adapter; If not, the first low-rank adapter is trained using results sampled from the initial latent distribution output by the variational autoencoder.

8. The three-dimensional object interpolation method based on low-rank semantic sampling according to claim 1, characterized in that: Obtaining a three-dimensional object sequence between two of the three-dimensional objects according to the interpolated multi-view images, comprising: Obtaining a perceptual distance between adjacent frames based on a set of said interpolated multi-view images; Optimizing the interpolation coefficient according to the perception distance; regenerating multiple sets of optimized interpolation multi-view images according to the optimized interpolation coefficients; A three-dimensional object sequence between two three-dimensional objects is obtained according to the multiple sets of optimized interpolated multi-view images.

9. The three-dimensional object interpolation method based on low-rank semantic sampling according to claim 8, characterized in that: Optimizing the interpolation coefficient according to the perception distance includes: According to the perception distances of all adjacent frames, the total perception distance is obtained; Obtaining a gradient of the relative perceptual distance with respect to the interpolation coefficient according to the perceptual distance of each adjacent frame and the total perceptual distance; Obtaining a relative perception distance corresponding to each interpolation coefficient according to a gradient of the relative perception distance with respect to the interpolation coefficient; According to the relative perception distances corresponding to all the interpolation coefficients, optimized interpolation coefficients are obtained.

10. The three-dimensional object interpolation method based on low-rank semantic sampling according to claim 9, characterized in that: Obtaining a three-dimensional object sequence between two of the three-dimensional objects according to the interpolated multi-view images, comprising: Reconstructing the interpolation process using four-dimensional Gaussian sputtering based on the plurality of sets of optimized interpolated multi-view images to obtain a spatiotemporally continuous decomposable three-dimensional Gaussian representation; The spatiotemporally continuous decomposable three-dimensional Gaussian representation is decomposed into a three-dimensional object sequence under the interpolation coefficients.