Image processing method, apparatus, device, storage medium and program product

CN122657451APending Publication Date: 2026-08-28BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510220647.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0002]相关技术中,使用单张图像生成多视图图像的方案大都采用Zero123、One-2-3-45等三维重建方法,这些方法虽然能够生成单张图像对应的多视图图像,但是各个视图图像的图像内容都不太一致,存在多视图图像质量较低的问题

Benefits of technology

[0053] The image processing method provided in this application constructs a first cost body using an initial image set corresponding to the image to be processed. Then, it optimizes the first cost body to obtain a second cost body, and performs consistency correction on initial view images from different perspectives in the initial image set based on the second cost body. That is, by optimizing the cost body, inaccurate initial view images are corrected, making the corrected new view images more consistent. This improves the image quality of each view image and ensures the effectiveness of the 3D reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122657451A_ABST
    Figure CN122657451A_ABST
Patent Text Reader

Abstract

The application provides an image processing method and device, a computer readable storage medium and a computer program product. The method comprises the following steps: constructing a first cost volume by using an initial image set corresponding to a to-be-processed image; the initial image set comprises initial view images of at least two different perspectives corresponding to the to-be-processed image; optimizing the first cost volume to obtain a second cost volume; and performing consistency correction on the initial view images in the initial image set by using the second cost volume to obtain a target image set corresponding to the to-be-processed image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to image processing technology, and more particularly to an image processing method, apparatus, computer-readable storage medium, and computer program product. Background Technology

[0002] In related technologies, most schemes for generating multi-view images from a single image employ 3D reconstruction methods such as Zero123 and One-2-3-45. Although these methods can generate multi-view images corresponding to a single image, the image content of each view image is not very consistent, resulting in low quality of the multi-view images. Summary of the Invention

[0003] This application provides an image processing method, apparatus, computer-readable storage medium, and computer program product that can improve the quality of multi-view images.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] This application provides an image processing method, the method comprising:

[0006] A first cost body is constructed using an initial image set corresponding to the image to be processed; the initial image set includes initial view images from at least two different perspectives corresponding to the image to be processed;

[0007] The first cost body is constructed using the initial image set;

[0008] The first cost body is optimized to obtain the second cost body;

[0009] The second cost body is used to perform consistency correction on the initial view images in the initial image set to obtain the target image set corresponding to the image to be processed.

[0010] In the above scheme, the first cost body is optimized to obtain the second cost body, including:

[0011] The symbolic distance function is modified using the first cost body to obtain the modified symbolic distance function;

[0012] The modified symbolic distance function is regularized to obtain the regularized symbolic distance function;

[0013] The second cost body is determined based on the regularized symbolic distance function.

[0014] In the above scheme, the consistency correction of the initial view images in the initial image set is performed using the second cost body, including:

[0015] The radiation function is corrected using the second cost body to obtain the corrected radiation function;

[0016] The modified radiation function is used to perform consistency correction on the initial view images in the initial image set.

[0017] In the above scheme, after obtaining the regularized symbolic distance function, the method further includes:

[0018] Obtain the geometric deviation factor;

[0019] The regularized symbolic distance function is optimized using the geometric deviation factor to obtain the optimized symbolic distance function;

[0020] Accordingly, determining the second cost body based on the regularized symbolic distance function includes:

[0021] The second cost body is determined using the optimized symbolic distance function.

[0022] In the above scheme, constructing the first cost body using an initial image set corresponding to the image to be processed includes:

[0023] Feature extraction is performed on the initial view images in the initial image set to obtain the feature map corresponding to the initial view image;

[0024] The feature maps corresponding to the initial view image are aggregated to obtain the first cost body.

[0025] In the above scheme, before constructing the first cost body using the initial image set corresponding to the image to be processed, the method further includes:

[0026] The image to be processed is processed using at least one preset diffusion model to obtain an initial image set corresponding to the image to be processed.

[0027] This application embodiment also provides an image processing apparatus, the image processing apparatus comprising:

[0028] A construction module is used to construct a first cost body using an initial image set corresponding to the image to be processed; the initial image set includes initial view images from at least two different perspectives corresponding to the image to be processed;

[0029] An optimization module is used to optimize the first cost body to obtain a second cost body;

[0030] The correction module is used to correct the initial view images in the initial image set using the second cost body to obtain a target image set corresponding to the image to be processed.

[0031] In the above scheme, the optimization module is further used for:

[0032] The symbolic distance function is modified using the first cost body to obtain the modified symbolic distance function;

[0033] The modified symbolic distance function is regularized to obtain the regularized symbolic distance function;

[0034] The second cost body is determined based on the regularized symbolic distance function.

[0035] In the above scheme, the correction module is further configured to:

[0036] The radiation function is corrected using the second cost body to obtain the corrected radiation function;

[0037] The modified radiation function is used to perform consistency correction on the initial view images in the initial image set.

[0038] In the above scheme, after obtaining the regularized symbolic distance function, the optimization module is further used to:

[0039] Obtain the geometric deviation factor;

[0040] The regularized symbolic distance function is optimized using the geometric deviation factor to obtain the optimized symbolic distance function;

[0041] The optimized symbolic distance function is used to determine the second cost body.

[0042] In the above scheme, the construction module is further used for:

[0043] Feature extraction is performed on the initial view images in the initial image set to obtain the feature map corresponding to the initial view image;

[0044] The feature maps corresponding to the initial view image are aggregated to obtain the first cost body.

[0045] In the above scheme, before constructing the first cost body using the initial image set corresponding to the image to be processed, the construction module is further configured to:

[0046] The image to be processed is processed using at least one preset diffusion model to obtain an initial image set corresponding to the image to be processed.

[0047] This application embodiment also provides an image processing device, the image processing device comprising:

[0048] Memory is used to store executable instructions for a computer;

[0049] A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the image processing method as described in any of the preceding claims.

[0050] This application also provides a computer-readable storage medium storing computer-executable instructions or a computer program, which, when executed by a processor, implement the image processing method as described in any of the preceding claims.

[0051] This application also provides a computer program product, including computer-executable instructions or a computer program, which, when executed by a processor, can implement the image processing method as described in any of the preceding claims.

[0052] The embodiments of this application have the following beneficial effects:

[0053] The image processing method provided in this application constructs a first cost body using an initial image set corresponding to the image to be processed. Then, it optimizes the first cost body to obtain a second cost body, and performs consistency correction on initial view images from different perspectives in the initial image set based on the second cost body. That is, by optimizing the cost body, inaccurate initial view images are corrected, making the corrected new view images more consistent. This improves the image quality of each view image and ensures the effectiveness of the 3D reconstruction. Attached Figure Description

[0054] Figure 1 A schematic flowchart of an image processing method provided in an embodiment of this application;

[0055] Figure 2 This is a schematic diagram illustrating the result of an image processing method provided in an embodiment of this application;

[0056] Figure 3 This is a schematic diagram of the structure of the image processing apparatus provided in the embodiments of this application;

[0057] Figure 4 This is a schematic diagram of the structure of the image processing device provided in the embodiments of this application. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0059] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0060] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0061] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0062] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0063] In related technologies, most schemes for generating multi-view images from a single image employ 3D reconstruction methods such as Zero123 and One-2-3-45. Although these methods can generate multi-view images corresponding to a single image, the image content of each view image is not very consistent, resulting in low quality of the multi-view images.

[0064] Based on the above technical problems, embodiments of this application provide an image processing method, apparatus, computer-readable storage medium, and computer program product.

[0065] Figure 1 This is a schematic flowchart of an image processing method provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps:

[0066] S101. Construct the first cost body using the initial image set corresponding to the image to be processed.

[0067] The image processing method provided in this application embodiment is applicable to scenarios that require the generation of three-dimensional views. Here, a three-dimensional view is a projection of a three-dimensional model observed from different perspectives in three-dimensional space, and it may include multiple view images. Exemplarily, the execution subject of the image processing method is an image processing module.

[0068] In this embodiment of the application, before constructing the first cost body using the initial image set corresponding to the image to be processed, the image to be processed can be obtained first; for example, the image to be processed includes a target object, and the image to be processed refers to the source image used to generate a three-dimensional view of the target object, which can be a two-dimensional red (R), green (G), and blue (B) image.

[0069] For example, the target object included in the image to be processed can be one object or a combination of multiple objects; wherein, the object can be static, such as: buildings, natural landscapes, etc.; or it can be dynamic, such as: human body, vehicle, etc.

[0070] Here, there are no specific limitations on how the image to be processed is acquired. For example, the image to be processed can be captured by a camera, obtained directly from a database, or generated based on text prompts.

[0071] In this embodiment of the application, after obtaining the image to be processed, an initial image set corresponding to the image to be processed can be determined first. The initial image set may include initial view images from at least two different perspectives corresponding to the image to be processed; here, the initial view images from at least two different perspectives refer to view images of the target object in the image to be processed from at least two different perspectives; for example, different perspectives may include front view, side view, top view, etc.

[0072] For example, determining the initial image set corresponding to the image to be processed may include: processing the image to be processed using at least one preset diffusion model to obtain the initial image set corresponding to the image to be processed.

[0073] Here, the preset diffusion model, also known as the multi-view diffusion model, is used to generate view images from different perspectives. The type of preset diffusion model is not specifically limited here. For example, the preset diffusion model can be the Zero123 model, the MVDream model, the Zero123++ model, etc.

[0074] Understandably, using at least one preset diffusion model to process the image to be processed can yield initial view images from different perspectives corresponding to the image to be processed. However, there is a consistency problem among these initial view images. That is, the shape or structure of the target object in different initial view images is inconsistent. Therefore, it is necessary to perform consistency correction on the initial view images in the initial image set to make the corrected view images more consistent, thereby ensuring the accuracy of the three-dimensional reconstruction of the image. The following is an exemplary description of this process.

[0075] In this embodiment of the application, after obtaining the initial image set corresponding to the image to be processed, the initial image set can be used to construct the first cost body; here, the first cost body is the initial three-dimensional cost body.

[0076] For example, in order to efficiently represent 3D priors, after obtaining the initial image set... Next, a bounded volume space for the 3D region of interest defined by the initial image set I is first defined, and this bounded volume space is defined in the camera coordinate system of the image, having a fixed local volume resolution. In order to efficiently use the initial image set I as a priori for the preset diffusion model, the first cost volume is constructed using the initial image set I.

[0077] In some embodiments, constructing a first cost body using an initial image set corresponding to the image to be processed includes: extracting features from the initial view images in the initial image set to obtain feature maps corresponding to the initial view images; and aggregating the feature maps corresponding to the initial view images to obtain the first cost body.

[0078] For example, a feature extraction network can be used to extract features from the initial view image in the initial image set to obtain the feature map corresponding to the initial view image; here, the feature extraction network can be a two-dimensional convolutional neural network, or other feature extraction networks, without specific limitations.

[0079] Wherein, a two-dimensional convolutional neural network can be represented as f 2D The feature map corresponding to the initial view image can be represented as: N is the number of feature maps, i.e., the number of initial view images; these extracted feature maps can be used for subsequent cost volume construction.

[0080] In this embodiment of the application, the feature map corresponding to the initial view image is obtained. Then, the feature map corresponding to the initial view image can be processed. Aggregation is performed to obtain the first cost body.

[0081] For example, the first cost body V can convert the feature map Aggregating these priors into a three-dimensional space provides a valuable reference for geometric reasoning in the prior refinement stage. (Using image I) i The camera parameters can project each voxel in the first cost volume V into the image space to obtain a pixel-level index [u i ,v i Then, feature maps are obtained through interpolation. Furthermore, this can be achieved through a three-dimensional CNN, i.e., f 3DThe variance of the projected features across all camera views is processed and used as a feature for each voxel in the first cost volume V. To represent the first cost volume V more efficiently, f can also be... 3D It is represented as a sparse 3D CNN, with empty voxels that are not visible from any view.

[0082] S102. Optimize the first cost body to obtain the second cost body.

[0083] In this embodiment of the application, after constructing the first cost body using the initial image set, the first cost body can be optimized to obtain the second cost body.

[0084] In some embodiments, optimizing the first cost body to obtain the second cost body may include: modifying the symbolic distance function using the first cost body to obtain a modified symbolic distance function; regularizing the modified symbolic distance function to obtain a regularized symbolic distance function; and determining the second cost body based on the regularized symbolic distance function.

[0085] For example, a sign distance function (SDF) can be initialized first, denoted as f. sdf (ε(r(t))), then the symbolic distance function is corrected using the first cost body to obtain f sdf (ε(r(t)),V(r(t))), as the symbolic distance function f sdf The simplified form of (ε(r(t)),V(r(t))), f sdf (r(t)) is almost everywhere differentiable, and its gradient Satisfying the Eikonal equation This means that SDF can be trained using Eikonal regularization.

[0086] For example, the Eikonal loss can be used as a regularizer to make the implicit field behave like a symbolic distance field. However, training with a fixed Eikonal regularization can lead to suboptimal convergence, meaning that the SDF has converged based on the Eikonal regularization, but the rendering can still be further optimized.

[0087] For example, to address the problem of thin structure reconstruction failure, the applicant found that even if the SDF does not contain the corresponding structure, it should still be ensured that these structures can be rendered from the radiation field so that gradients can be subsequently backpropagated from the radiation field to the SDF. Based on the overall loss function L... total =L rgb +L sdf We propose weighting the Eikonal regularization to adaptively optimize the SDF, where the loss function L... sdf As shown in formula (1):

[0088]

[0089] For the i-th ray r i Set the ray weight λ r (r i )for:

[0090]

[0091] Where, d r (r i () is about light r i Radiation distance α is a positive hyperparameter less than 1, for example, 1.10. -6 In implementation, the approximate normal is... It is f sdf The derivative of (r(·)), λ E It is usually set to 0.1. To avoid extreme values ​​in different scenarios, d r (r) is restricted to a range [c min ,c max ]Inside.

[0092] For example, considering the impact of simple changes in Equation (1) on the results, a similar relaxation term is applied to each sampling point of the selected ray during training for Eikonal regularization. Ideally, the number of points contributing to the space behind the observed geometric surface is far less than the number of blank spaces in neural rendering. While this is a good phenomenon for efficient rendering, it also leads to a problem where the integral rendering points, which can be used as the estimated depth points for each ray, are biased by the zero crossover of the rays. The parameterized geometric model does not guarantee an ideal SDF distribution, even though the parameters of the geometric model are explicitly initialized to produce a spherical SDF, the radiation-based supervision does not impose explicit regularization on the underlying SDF field. The consistency problem between the radiation field and the SDF makes it difficult to optimize the internal space of the SDF, especially within small structures with small negative spaces.

[0093] Considering that D-NeuS defines a weighted rendering point r(t) r That is, the rendering point t is obtained by discretizing and integrating the volume between the nearest point tn and the farthest point tf. r Calculate o+t r ·v. Where o is the center of the camera and v is the view direction. The radiation of the reference ray r is expressed in formula (3) and the transmittance is expressed as T. s (t):

[0094]

[0095] Formula (3) will be explained further later, and will not be repeated here; where, the rendering point t r The weighted summation of ω(t) is obtained as shown in formula (4):

[0096]

[0097] Wherein, the weight ω(t) j This can be expressed by formula (5):

[0098]

[0099] Furthermore, a geometric deviation factor λ is introduced. g Its definition is shown in formula (6). The geometric deviation factor is used to optimize the regularized symbolic distance function to obtain the optimized symbolic distance function.

[0100]

[0101] Where r(t) s The zero-crossing point is approximated by evaluating the SDF value on the ray r(·). Therefore, in this case, Eikonal regularization is only fully enforced if there is no geometric deviation between the zero-crossing surface of the SDF and the rendered point in the radiation field. This is achieved through λ. g (r) Weighted SDF loss will also backpropagate gradients to T. s The factor s used in (t) adjusts the projection from the SDF to the radiation field. The final loss function L is obtained. sdf Defined as:

[0102]

[0103] Based on loss function L sdf The symbolic distance function is optimized to obtain an optimized symbolic distance function; then, the optimized symbolic distance function is used to determine the second cost body; wherein, the optimized symbolic distance function includes the second cost body.

[0104] S103. Use the second cost body to perform consistency correction on the initial view images in the initial image set to obtain the target image set corresponding to the image to be processed.

[0105] In this embodiment of the application, after obtaining the second cost body according to the above steps, the second cost body can be used to correct the initial view image in the initial image set to obtain the target image set corresponding to the image to be processed.

[0106] In some embodiments, correcting the initial view images in the initial image set using the second cost body may include: correcting the radiation function using the second cost body to obtain a corrected radiation function; and using the corrected radiation function to perform consistency correction on the initial view images in the initial image set.

[0107] For example, referring to formula (3) above, the radiation function can be expressed as f rad (r(t),v), by using the first cost body to correct the symbolic distance function, we can obtain the corrected radiation function frad(r(t),v,V(r(t))), which is explained below.

[0108] For example, firstly, n points are sampled along the camera ray r, which is represented as {r(t)} i )=o+t r The RGB values ​​of each pixel in the image are generated using the formula ·v|i=1,...,n}; where o is the center of the camera, t i The sampling interval is along the ray r, and v is the view direction. The volume density σ(r(t)) and radiation function f are accumulated from the sampled points. rad (r(t),v) can be used to calculate the color of light. As shown in formula (3) above.

[0109] The transmittance T(t) is derived from the volume density σ(r(t)). It should be noted that the optimized sign distance function is initialized with the volume density. T(t) represents the distance from the nearest point t along the ray r. n To the farthest point t f The cumulative transmittance can be expressed as formula (8):

[0110]

[0111] Where T(t) is a monotonically decreasing function, with an initial value of T(t). n The product T(t)·σ(r(t)) is 1. The product T(t)·σ(r(t)) is used as the weight ω(t) of radiation in volume rendering.

[0112] Since the rendering process is differentiable, the model can learn the radiation field c from multiple view images and minimize the rendered pixels using a loss function. The color difference between the ground truth pixel C(r) corresponding to i∈{1,...,m} does not require 3D supervision, and the loss function L rgb This can be expressed as formula (9):

[0113]

[0114] Here, m represents the batch size during training. In practice, to improve convergence speed, positional hashing encoding ε(r(t)) is used to obtain the color. Furthermore, to improve the robustness of multi-view supervision, the pointwise features presented in the second cost volume V(r(t)) are also used as conditional information. Finally, the modified radiation function can be expressed as f rad (ε(r(t)),v,V(r(t))).

[0115] For example, in order to extract a 3D mesh from a region of interest in neural rendering, it is reasonable to project from the radiation field of the symbolic distance function. Here, a function Φ can be found to transform the symbolic distance function so as to adapt the symbolic distance function to calculate the density-related term T(t)·σ(r(t)) in Equation (3). The corresponding solution is to set Φ(r(t)) as the transmittance T(t).

[0116] For example, the derivative of transmittance T(t) is a negative weighting function, as shown in Equation (10):

[0117]

[0118] According to the above representation, the SDF surface lies on the maximum radiation weight. The maximum value is calculated by setting the derivative of the weighting function to zero, as shown in Equation (11):

[0119]

[0120] To satisfy the criteria in formulas (10) and (11), the transmittance T(t) can be defined as a normalized sigmoid function [1+exp(s·f(r(t)))]. -1 It includes a trainable parameter s. The scalar s reveals the strength of the correlation between the SDF and the radiation field, which is typically increased gradually during training. In the initial stages of training, a smaller s can reduce the correlation between the SDF and the radiation field, allowing the radiation parameter model to be optimized without relying on a correct SDF.

[0121] Furthermore, after obtaining the corrected radiation function f rad After (r(t),v), the modified radiation function can be used to perform consistency correction on the initial view images in the initial image set.

[0122] For example, the initial view images in the initial image set can be treated as a whole for consistency correction. Here, the specific implementation of consistency correction is not limited. For example, a correction function can be determined based on the corrected radiation function, and the correction function can be multiplied with the initial view images in the initial image set to obtain the corrected initial view image.

[0123] The image processing method provided in this application constructs a first cost body using an initial image set corresponding to the image to be processed. Then, it optimizes the first cost body to obtain a second cost body, and corrects the initial view images from different perspectives in the initial image set based on the second cost body. That is, by optimizing the cost body, inaccurate initial view images are corrected, making the corrected new view images more consistent. This improves the image quality of each view image and ensures the effectiveness of the 3D reconstruction.

[0124] Below, in conjunction with Figure 2 The image processing process in the embodiments of this application will be further illustrated with examples, such as... Figure 2 As shown, the image to be processed includes a mining truck, that is, the target object in the image to be processed is a mining truck. If the image to be processed is processed according to the One-2-3-45 model, the four view images shown in the upper right position can be obtained. If the image processing flow provided in this application embodiment is used to further correct the four view images shown in the upper right position, the four view images shown in the lower right position can be obtained. It can be seen that compared with the four view images before correction, the four view images after correction are significantly more consistent.

[0125] Based on the foregoing embodiments, this application also provides an image processing apparatus. Figure 3 This is a schematic diagram of the structure of the image processing apparatus provided in the embodiments of this application, such as... Figure 3 As shown, the image processing device 3 may include:

[0126] The construction module 303 is used to construct a first cost body using an initial image set corresponding to the image to be processed; the initial image set includes initial view images from at least two different perspectives corresponding to the image to be processed;

[0127] Optimization module 304 is used to optimize the first cost body to obtain the second cost body;

[0128] The correction module 305 is used to correct the initial view images in the initial image set using the second cost body to obtain the target image set corresponding to the image to be processed.

[0129] In some embodiments, the optimization module 304 is further configured to:

[0130] The symbolic distance function is corrected using the first cost body to obtain the corrected symbolic distance function;

[0131] The modified symbolic distance function is regularized to obtain the regularized symbolic distance function;

[0132] The second cost body is determined based on the regularized symbolic distance function.

[0133] In some embodiments, the correction module 305 is further configured to:

[0134] The radiation function is corrected using a second cost body to obtain the corrected radiation function;

[0135] The initial view images in the initial image set are corrected for consistency using the modified radiation function.

[0136] In some embodiments, after obtaining the regularized symbolic distance function, the optimization module 304 is further configured to:

[0137] Obtain the geometric deviation factor;

[0138] The regularized symbolic distance function is optimized using a geometric deviation factor to obtain the optimized symbolic distance function.

[0139] The second cost body is determined using the optimized symbolic distance function.

[0140] In some embodiments, the construction module 303 is further configured to:

[0141] Feature extraction is performed on the initial view images in the initial image set to obtain the feature maps corresponding to the initial view images;

[0142] The feature maps corresponding to the initial view image are aggregated to obtain the first cost body.

[0143] In some embodiments, before constructing the first cost body using an initial image set corresponding to the image to be processed, the construction module 303 is further configured to:

[0144] The image to be processed is processed using at least one preset diffusion model to obtain an initial image set corresponding to the image to be processed.

[0145] In practical applications, the acquisition module 301, determination module 302, and execution module 303 can all be implemented by a processor located in the image processing device. The processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), controller, microcontroller, and microprocessor.

[0146] Based on the foregoing embodiments, this application also provides an image processing device. Figure 4 This is a schematic diagram of the structure of the image processing device provided in the embodiments of this application, such as... Figure 4 As shown, the image processing device 4 may include a processor 401 and a memory 402; wherein,

[0147] Memory 402 is used to store computer-executable instructions;

[0148] The processor 401, when executing computer-executable instructions or computer programs stored in the memory 402, implements the image processing method as described above.

[0149] Based on the foregoing embodiments, this application also provides a computer-readable storage medium storing computer-executable instructions or a computer program, which, when executed by a processor, implement the image processing method as described in any of the preceding embodiments.

[0150] Based on the foregoing embodiments, this application also provides a computer program product, including computer-executable instructions or a computer program, which, when executed by a processor, can implement the image processing method as described in any of the preceding embodiments.

[0151] In some embodiments, the computer-readable storage medium may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it may be a device that includes one or any combination of the above-mentioned memories.

[0152] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0153] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

Claims

1. An image processing method, characterized in that, The method includes: A first cost body is constructed using an initial image set corresponding to the image to be processed; the initial image set includes initial view images from at least two different perspectives corresponding to the image to be processed; The first cost body is optimized to obtain the second cost body; The second cost body is used to perform consistency correction on the initial view images in the initial image set to obtain the target image set corresponding to the image to be processed.

2. The method according to claim 1, characterized in that, The first cost body is optimized to obtain a second cost body, which includes: The symbolic distance function is modified using the first cost body to obtain the modified symbolic distance function; The modified symbolic distance function is regularized to obtain the regularized symbolic distance function; The second cost body is determined based on the regularized symbolic distance function.

3. The method according to claim 1 or 2, characterized in that, The initial view images in the initial image set are subjected to consistency correction using the second cost body, including: The radiation function is corrected using the second cost body to obtain the corrected radiation function; The modified radiation function is used to perform consistency correction on the initial view images in the initial image set.

4. The method according to claim 2, characterized in that, After obtaining the regularized symbolic distance function, the method further includes: Obtain the geometric deviation factor; The regularized symbolic distance function is optimized using the geometric deviation factor to obtain the optimized symbolic distance function; Accordingly, determining the second cost body based on the regularized symbolic distance function includes: The second cost body is determined using the optimized symbolic distance function.

5. The method according to any one of claims 1 to 4, characterized in that, The construction of the first cost body using the initial image set corresponding to the image to be processed includes: Feature extraction is performed on the initial view images in the initial image set to obtain the feature map corresponding to the initial view image; The feature maps corresponding to the initial view image are aggregated to obtain the first cost body.

6. The method according to any one of claims 1 to 5, characterized in that, Before constructing the first cost body using the initial image set corresponding to the image to be processed, the method further includes: The image to be processed is processed using at least one preset diffusion model to obtain an initial image set corresponding to the image to be processed.

7. An image processing apparatus, characterized in that, The image processing device includes: A construction module is used to construct a first cost body using an initial image set corresponding to the image to be processed; the initial image set includes initial view images from at least two different perspectives corresponding to the image to be processed; An optimization module is used to optimize the first cost body to obtain a second cost body; The correction module is used to correct the initial view images in the initial image set using the second cost body to obtain a target image set corresponding to the image to be processed.

8. An image processing device, characterized in that, The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the image processing method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the image processing method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the image processing method according to any one of claims 1 to 6.