Image processing method and device, equipment, storage medium and program product

Through feature mapping and pixel ray sampling technology, the target object image synthesized by new perspectives is generated, which solves the problem of inefficiency in the existing technology and achieves more efficient and high-quality new perspective synthesis.

CN120032028APending Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311564164.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing new perspective synthesis technology is relatively inefficient, mainly due to the need for three-dimensional reconstruction.

Method used

By acquiring the reference object image and its viewing angle parameters, feature mapping is performed to obtain a three-dimensional feature plane with an axis orthogonal axis, point features are obtained using pixel rays to sample and obtain point features, and a target object image with a viewing angle conforming to the desired viewing angle parameters are generated.

Benefits of technology

The efficiency of new perspective synthesis is improved, and the image is characterized in the feature space can be achieved with higher representation speed and resolution, and the spatial perception ability is enhanced through pixel ray sampling, improving image generation quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032028A_ABST
    Figure CN120032028A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, equipment, a storage medium and a program product, relates to the technical field of artificial intelligence, can be used for a new view angle synthesis task, and comprises the steps: obtaining a reference object image, and a current view angle parameter and an expected view angle parameter of the reference object image; according to the expected view angle parameters, performing feature mapping on the reference object image to obtain a three-dimensional feature plane with orthogonal axes; acquiring a pixel ray of a pixel point in the reference object image according to the current visual angle parameter; according to the pixel rays, point features corresponding to the pixel points are obtained through sampling in the three-dimensional feature plane; and according to the point features corresponding to the pixel points, generating a target object image of which the viewing angle accords with expected viewing angle parameters. Compared with the prior art, the new view angle synthesis can be carried out in a more efficient and high-quality manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and specifically to an image processing method, an image processing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] The new perspective synthesis task is to render the object image of the desired perspective given the original perspective image of an object. The new perspective synthesis technology is widely used in various pan-entertainment applications.

[0003] At present, new perspective synthesis is mainly realized based on 3D reconstruction and physical rendering. For objects that need to be synthesized from a new perspective, such as people, animals, plants, etc., the object is first 3D reconstructed and related maps are applied. Then, by changing the viewing angle conditions, the physical rendering method is controlled to obtain the object image at a specified viewing angle. However, due to the need for 3D reconstruction, this new perspective synthesis method is less efficient. Summary of the invention

[0004] The embodiments of the present application provide an image processing method, an image processing device, an electronic device, a computer-readable storage medium, and a computer product, which can improve the efficiency of new perspective synthesis.

[0005] In a first aspect, the image processing method provided by the present application comprises:

[0006] Acquire a reference object image, and a current viewing angle parameter and an expected viewing angle parameter of the reference object image;

[0007] According to the desired viewing angle parameters, feature mapping is performed on the reference object image to obtain a three-dimensional feature plane orthogonal to the axis;

[0008] According to the current viewing angle parameter, a pixel ray of a pixel point in the reference object image is obtained;

[0009] According to the pixel ray, the point features corresponding to the pixel points are sampled in the three-dimensional feature plane;

[0010] According to the point features corresponding to the pixel points, an image of the target object whose viewing angle meets the expected viewing angle parameters is generated.

[0011] In a second aspect, the image processing device provided by the present application includes:

[0012] A reference acquisition module, used to acquire a reference object image, and current viewing angle parameters and expected viewing angle parameters of the reference object image;

[0013] A feature mapping module, used for performing feature mapping on the reference object image according to the expected viewing angle parameters to obtain a three-dimensional feature plane with orthogonal axes;

[0014] A ray acquisition module, used to acquire pixel rays of pixel points in the reference object image according to the current viewing angle parameter;

[0015] A feature sampling module is used to sample point features corresponding to pixel points in a three-dimensional feature plane according to pixel rays;

[0016] The image rendering module is used to generate a target object image whose viewing angle meets the expected viewing angle parameters according to the point features corresponding to the pixel points.

[0017] Optionally, in one embodiment, the feature sampling module is used to determine multiple sampling points corresponding to pixel points on the pixel ray, and obtain the projection features of each sampling point on the three-dimensional feature plane; and fuse the projection features of each sampling point to obtain the point features corresponding to the pixel point based on the fused features of each sampling point.

[0018] Optionally, in one embodiment, the feature sampling module is used to determine whether the parameter difference between the current viewing parameter and the expected viewing parameter reaches a difference threshold; and if the parameter difference between the current viewing parameter and the expected viewing parameter does not reach the difference threshold, the projection feature of each sampling point on the three-dimensional feature plane is obtained.

[0019] Optionally, in one embodiment, the feature sampling module is also used to convert each sampling point from the camera coordinate system to the object coordinate system corresponding to the object in the reference object image according to the current viewing angle parameter if the parameter difference between the current viewing angle parameter and the expected viewing angle parameter does not reach the difference threshold, so as to obtain the converted sampling point; and obtain the projection feature of each converted sampling point on the three-dimensional feature plane; and fuse the projection feature of each converted sampling point to obtain the point feature corresponding to the pixel point according to the fused feature of each converted sampling point.

[0020] Optionally, in one embodiment, the image rendering module is used to predict the corresponding color value and density value based on each point feature corresponding to the pixel point; and generate a target object image whose viewing angle meets the expected viewing angle parameters based on the color value and density value corresponding to each point feature.

[0021] Optionally, in one embodiment, the image rendering module is used to generate an initial object image whose viewing angle meets the expected viewing angle parameters according to the color value and density value corresponding to each point feature, and the resolution of the initial object image is smaller than the resolution of the reference object image; and perform super-resolution processing on the initial object image, and use the obtained super-resolution image as the target object image.

[0022] Optionally, in one embodiment, the image processing device is executed through an image generation model, and the image processing device also includes a model training module, which is used to obtain a sample object image and a sample current viewing angle parameter of the sample object image; and obtain a random latent code, and generate a corresponding sample initial image and a sample super-resolution image through the image generation model according to the random latent code and the sample current viewing angle parameter; and upsample the sample initial image to obtain an upsampled initial image with the same resolution as the sample super-resolution image; and splice the upsampled initial image and the sample super-resolution image to obtain a first image to be judged; and blur the sample object image to obtain a blurred sample object image; and splice the blurred sample object image and the sample object image to obtain a second image to be judged; and use the sample current viewing angle parameter as a discrimination reference of the discriminator network, and update the network parameters of the image generation model with the constraint that the discrimination result of the discriminator network for the first image to be judged is false and the discrimination result of the second image to be judged is true until a preset stop condition is met.

[0023] Optionally, in one embodiment, the feature sampling module is used to determine a first preset number of sampling points equidistantly on the pixel ray; and determine a second preset number of sampling points on the pixel ray according to depth information of the pixel points.

[0024] Optionally, in one embodiment, the feature mapping module is used to perform inverse mapping on the reference object image to obtain an implicit code of the reference object image; and perform feature mapping on the implicit code according to the desired viewing angle parameters to obtain a multi-channel feature map; and reconstruct the multi-channel feature map into an axis-orthogonal three-dimensional feature plane, and the number of channels of the feature map corresponding to each feature plane is the same.

[0025] Optionally, in one embodiment, the feature mapping module is used to perform feature mapping on the latent code according to the expected viewing angle parameters to obtain a temporary multi-channel feature map; and reconstruct the temporary multi-channel feature map into an axis-orthogonal temporary three-dimensional feature plane; and sample the temporary point features corresponding to the pixels in the temporary three-dimensional feature plane according to the pixel rays; and generate a temporary object image according to the temporary point features corresponding to the pixels; and update the latent code according to the difference between the temporary object image and the reference object image at the same object position to obtain an updated latent code; and perform feature mapping on the updated latent code according to the expected viewing angle parameters to obtain a multi-channel feature map.

[0026] Optionally, in one embodiment, the reference acquisition module is used to acquire an object image to be restored of the object, wherein the object region in the object image to be restored is blocked; and acquire a historical object image corresponding to the object image to be restored as the reference object image, wherein the object region in the historical object image is not blocked; and acquire a viewing angle parameter of the object image to be restored as the expected viewing angle parameter;

[0027] The image processing device also includes an image restoration module, which is used to restore the image of the object to be restored according to the target object image to obtain a restored image of the object.

[0028] In a third aspect, the electronic device provided in the present application includes a memory and a processor, the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the steps in the image processing method provided in the present application.

[0029] In a fourth aspect, the computer-readable storage medium provided in the present application stores a computer program, which is suitable for being executed by a processor to implement the steps in the image processing method provided in the present application.

[0030] In a fifth aspect, the computer program product provided in the present application includes a computer program, which is suitable for being run by a processor to implement the steps in the image processing method provided in the present application.

[0031] The image processing scheme provided by the present application obtains a reference object image, as well as the current viewing angle parameters and the expected viewing angle parameters of the reference object image, and then performs feature mapping on the reference object image according to the expected viewing angle parameters to obtain an axis-orthogonal three-dimensional feature plane. In addition, according to the current viewing angle parameters, the pixel rays of the pixel points in the reference object image are obtained, and the point features corresponding to the pixel points sampled by the pixel rays in the three-dimensional feature plane are used. Finally, according to the point features corresponding to the pixel points, a target object image whose viewing angle meets the expected viewing angle parameters is generated. In this way, by using the three-dimensional feature plane to represent the image in the feature space, a higher representation speed and resolution can be obtained, thereby realizing detailed content under the same capacity, and the use of pixel rays to sample the point features of the pixel points can increase the spatial perception ability, thereby improving the image generation quality. Compared with the related art, the image processing scheme provided by the present application can synthesize new viewing angles more efficiently and with higher quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0033] Figure 1a is a schematic diagram of a scene of an image processing system provided in an embodiment of the present application;

[0034] Figure 1b It is a schematic diagram showing the process of the image processing method provided in the embodiment of the present application;

[0035] Figure 1c is an example diagram of a three-dimensional feature plane obtained by performing feature mapping on a reference object image in an embodiment of the present application;

[0036] Figure 1d is a structural schematic diagram of an image generation model provided in an embodiment of the present application;

[0037] Figure 1e is another structural schematic diagram of the image generation model provided in the embodiments of the present application;

[0038] Figure 1f is another example diagram of obtaining a three-dimensional feature plane by performing feature mapping on a reference object image in an embodiment of the present application;

[0039] Figure 1g is an example diagram of generating a target object image in a side view from a reference object image in a front view in an embodiment of the present application;

[0040] Figure 2 is another flowchart of the image processing method provided by an embodiment of the present application;

[0041] Figure 3 is a structural schematic diagram of an image processing device provided in an embodiment of the present application;

[0042] Figure 4 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] It should be noted that the principles of the present application are illustrated by implementing them in an appropriate computing environment. The following description is based on the illustrated specific embodiments of the present application and should not be considered as limiting other specific embodiments of the present application that are not described in detail herein.

[0044] In the following description of the present application, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0045] In the following description of the present application, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0047] It should be noted that artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0048] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, and mechatronics. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes machine learning (ML) technology, among which deep learning (DL) is a new research direction in machine learning. It is introduced into machine learning to make it closer to the original goal, namely artificial intelligence. At present, deep learning is mainly used in machine vision, natural language processing and other fields. Deep learning is the inherent laws and representation levels of learning sample data. The information obtained in these learning processes is of great help to the interpretation of data such as text, images and sound. Using deep learning technology and corresponding training sets, network models that implement different functions can be trained. For example, taking the generative model as an example, based on different types of training sets, we can train generative models that can generate different types of content, such as generative models that can generate images, generative models that can generate text, and generative models that can generate speech.

[0049] Computer Vision (CV) is a science that studies how to make machines "see". To put it more specifically, it refers to using cameras and computers to replace human eyes to identify and measure targets, and further perform image processing to make computer processing into images that are more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.

[0050] The present application mainly relates to the field of computer vision technology of artificial intelligence technology, and provides an image processing method, an image processing device, an electronic device, a computer-readable storage medium, and a computer program product. The image processing method can be executed by the image processing device, or by an electronic device integrated with the image processing device.

[0051] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described below are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0052] Please refer to Figure 1aThe present application also provides an image processing system, which includes an electronic device 100, which is used to execute the image processing method provided by the present application. The electronic device 100 can be any device equipped with a processor and has processing capabilities, such as a mobile device with a processor such as a smart phone, a tablet computer, a PDA, a laptop computer, a virtual reality device, an augmented reality device or a mixed reality device, or a fixed device with a processor such as a desktop computer, a television, a server, an industrial device, etc., wherein the reference object image of the original perspective is obtained, and the current perspective parameters describing the original perspective, and the expected perspective parameters describing the new perspective to be generated are obtained, and then according to the expected perspective parameters, the reference object image is feature mapped to obtain a three-dimensional feature plane with axes orthogonal to three feature planes, and then according to the current perspective parameters, the pixel rays of the pixel points in the reference object image are obtained, and according to the pixel rays, the point features corresponding to the pixel points are sampled in the three-dimensional feature plane, and finally, according to the point features corresponding to the pixel points, the target object image whose perspective meets the expected perspective parameters is generated.

[0053] In addition, if Figure 1a As shown, the image processing system may also include a memory 200 for storing relevant data in the image processing process, such as original data such as the acquired reference object image, current viewing angle parameters and expected viewing angle parameters, intermediate data such as three-dimensional feature planes, pixel rays, point features, etc. in the processing process, and result data such as the final target object image.

[0054] It should be noted that the image processing system described above is merely an example, which is intended to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided in the embodiment of the present application. A person of ordinary skill in the art can appreciate that with the evolution of the image processing system and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.

[0055] It should be noted that the serial numbers of the following embodiments are not intended to limit the preferred order of the embodiments.

[0056] Please refer to Figure 1b , Figure 1b FIG. 1 is a flow chart of the image processing method provided in this embodiment. Figure 1b As shown, the process of the image processing method can be as follows:

[0057] In 110 , a reference object image, and current viewing angle parameters and expected viewing angle parameters of the reference object image are acquired.

[0058] In an embodiment of the present application, in response to a demand for generating a new perspective image, an object image of an object for which a new perspective image needs to be generated is obtained as a reference for generating the new perspective image, and is correspondingly recorded as a reference object image. The object may be any object, including but not limited to people, animals, plants, and buildings, etc. The reference object image may be obtained by photographing the object at a certain perspective through a device with image capturing capability such as a camera or a camera, for example, photographing the object at a frontal perspective through a camera to obtain a frontal perspective image of the object. In addition, the expected perspective parameter is used to indicate the perspective of the new perspective image to be generated, for example, the obtained expected perspective parameter indicates that the perspective of the new perspective image to be generated is the side perspective of the object.

[0059] The current viewing angle parameters may be the current camera parameters of the reference object image, and the expected viewing angle parameters may be the expected camera parameters of the new viewing angle image to be generated. The camera parameters include the internal and external parameters of the camera; the camera internal parameters may be parameters related to the camera's own characteristics, such as the camera's focal length, pixels and other data; the camera external parameters may be parameters of the camera in the world coordinate system, such as the camera's position (coordinates) and the camera's rotation matrix and other data.

[0060] In this embodiment, the current camera parameters can be obtained by detecting the reference object image using a pose detection algorithm in the related technology. There is no specific restriction on which pose detection algorithm to use, and it can be selected by technical personnel in this field according to actual needs. For example, the pose detection algorithm in DECA can be used to perform pose detection on the reference object image, and the current camera parameters of the reference object image can be obtained accordingly.

[0061] It should be noted that this embodiment does not impose any specific restrictions on the triggering method for the generation requirement of the new viewing angle image, for example, it may be triggered in response to an external input, or it may be automatically triggered when a preset trigger condition is met, etc. In addition, this embodiment does not impose any specific restrictions on the method for obtaining the desired viewing angle parameter, for example, it may be obtained through external input, or it may be obtained by self-recognition, etc.

[0062] In 120 , feature mapping is performed on the reference object image according to the desired viewing angle parameter to obtain a three-dimensional feature plane with orthogonal axes.

[0063] As above, please refer to Figure 1c After obtaining the reference object image and the expected viewing angle parameters, the reference object image is feature mapped according to the expected viewing angle parameters to obtain feature maps of multiple channels, and then the feature maps of the multiple channels are divided into three equal parts, and each feature map is reconstructed into a feature plane, and the three-dimensional feature planes with axes orthogonal to each other composed of the XY feature plane, the XZ feature plane and the ZX feature plane are obtained accordingly.

[0064] It should be noted that in this embodiment, a pre-trained image generation model is configured to take an object image at an original viewing angle and a desired viewing angle parameter as input, and take an object image at a viewing angle indicated by the desired viewing angle parameter as output. Figure 1d The image generation model may include a feature mapping network and an image rendering network, wherein the feature mapping network is configured to map the input object image of the original perspective to the feature space according to the desired perspective parameters to obtain its hidden layer representation, and the image rendering network adopts the architecture of the neural radiation field network and is configured to map the hidden layer representation of the object image back to the image space to obtain the object image of the new perspective.

[0065] Optionally, in one embodiment, feature mapping is performed on the reference object image according to the desired viewing angle parameter to obtain an axis-orthogonal three-dimensional feature plane, including:

[0066] Performing inverse mapping on the reference object image to obtain a latent code of the reference object image;

[0067] Performing feature mapping on the latent code according to the expected viewing angle parameter to obtain a multi-channel feature map;

[0068] The multi-channel feature map is reconstructed into a three-dimensional feature plane with orthogonal axes, and the number of channels of the feature map corresponding to each feature plane is the same.

[0069] Please refer to Figure 1e , the feature mapping network may include a mapping subnetwork and a generating subnetwork, wherein the generating subnetwork may adopt the generator network in the generative adversarial network, such as the generator network of StyleGAN as the generating subnetwork, and the input data of the corresponding generating subnetwork needs to be the latent code of the latent space corresponding to the generating subnetwork, and the generating subnetwork is configured to generate the corresponding feature map according to the input latent code; the mapping subnetwork is configured to map the input original latent code and the expected viewing angle parameter to a new latent code that conforms to the aforementioned latent code form as the input of the generating subnetwork. There is no specific limitation on the architecture of the mapping subnetwork here, and it can be configured by those skilled in the art according to actual needs.

[0070] In addition, in this embodiment, a generative adversarial network inversion model is pre-trained, and the generative adversarial network inversion model is configured to inversely map the image from the image space back to the latent space corresponding to the generative subnetwork. The architecture and training method of the generative adversarial network inversion model are not specifically limited here, and can be configured by those skilled in the art according to actual needs.

[0071] Accordingly, in this embodiment, please refer to Figure 1fFirst, the reference object image is inversely mapped through the generative adversarial network inversion model, and the inverse mapping is performed back to the latent space corresponding to the generative subnetwork to obtain the latent code of the reference object image; then, the latent code of the reference object image and the expected viewing angle parameter (represented in the form of a vector) are mapped to a new latent code of the latent space corresponding to the generative subnetwork through the mapping subnetwork; finally, the new latent code is input into the generative subnetwork, and a multi-channel feature map (denoted as a multi-channel feature map) is generated through the generative subnetwork, and the multi-channel feature map is reconstructed into an axis-orthogonal three-dimensional feature plane, where the number of channels of the feature map corresponding to each dimensional feature plane is the same.

[0072] For example, taking the generator network of StyleGAN adopted in the generative subnetwork as an example, the reference object image is inversely mapped through the generative adversarial network inversion model, and it is inversely mapped back to the latent space corresponding to the generative subnetwork to obtain the latent code of the reference object image (a 512-dimensional vector); then, the latent code of the reference object image and the expected viewing angle parameter (a 25-dimensional vector) are mapped to a new latent code (a 512-dimensional vector) of the latent space corresponding to the generative subnetwork through the mapping subnetwork; finally, the new latent code is input into the generative subnetwork, and a multi-channel feature map (for example, 96 channels) is generated through the generative subnetwork, and the multi-channel feature map is reconstructed into an axis-orthogonal three-dimensional feature plane, where each dimensional feature plane corresponds to a feature map of 32 channels.

[0073] In 130 , pixel rays of pixel points in the reference object image are obtained according to the current viewing angle parameter.

[0074] As described above, after feature mapping is performed on the reference object image to obtain the axis-orthogonal three-dimensional feature plane, the corresponding pixel ray is obtained for the pixel point in the reference object image according to the current viewing angle parameter. Among them, the pixel ray of a pixel point can be generally understood as the ray emitted from the pixel point to the camera.

[0075] Exemplarily, the imaging plane of the reference object image may be determined first according to the current viewing angle parameter, and for a pixel point in the reference object image, a ray perpendicular to the imaging plane is determined based on the pixel point as the pixel ray of the pixel point.

[0076] In 140 , point features corresponding to the pixel points are sampled in the three-dimensional feature plane according to the pixel rays.

[0077] In this embodiment, for a pixel point in the reference object image, feature sampling is performed in a three-dimensional feature plane according to a pixel ray corresponding to the pixel point, and a point feature corresponding to the pixel point is obtained by corresponding sampling.

[0078] Optionally, in one embodiment, sampling in a three-dimensional feature plane according to a pixel ray to obtain a point feature corresponding to a pixel point includes:

[0079] Determine multiple sampling points corresponding to the pixel points on the pixel ray, and obtain the projection feature of each sampling point on the three-dimensional feature plane;

[0080] The projection features of each sampling point are fused, and the point features corresponding to the pixel points are obtained according to the fused features of each sampling point.

[0081] In this embodiment, for a pixel point in the reference object image, multiple sampling points corresponding to the pixel point can be determined on the pixel ray corresponding to the pixel point according to the pre-configured sampling point selection rule. The configuration of the sampling point selection rule is not limited here, and can be configured by those skilled in the art according to actual needs. For example, the sampling point selection rule can be configured to select a preset number of sampling points equidistantly on the pixel ray, and the sampling point selection rule can also be configured to randomly select a preset number of sampling points on the pixel ray, etc.

[0082] As described above, for a pixel point in the reference object image, after determining multiple sampling points corresponding to the pixel point from its corresponding pixel ray, for a sampling point, it is projected onto each dimensional feature plane of the three-dimensional feature plane to obtain the projection features of the sampling point on each dimensional feature plane. 0 ,y 0 , z 0 ), project the sampling point to the XY feature plane sampling (x 0 ,y 0 ) point features to obtain the projection feature F xy , project the sampling point to the YZ feature plane sampling ( y0 , z 0 ) point features to obtain the projection feature F yz , project the sampling point to the ZX feature plane sampling (z 0 , x 0 ) point to obtain feature F zx .

[0083] Finally, for each sampling point, the projection features of the sampling point on each dimensional feature plane are fused to obtain the fused features of the sampling point as the point features of the pixel point. In this way, for a pixel point in the reference object image, the corresponding number of point features is obtained according to the determined number of sampling points.

[0084] Among them, the process of fusing the projection features of the sampling points can be expressed as:

[0085] F=F xy +F yz +F zx;

[0086] Among them, F represents the fusion feature of a sampling point, that is, the point feature of the corresponding pixel point, F xy Indicates the projection feature of the sampling point on the XY feature plane, F yz Indicates the projection feature of the sampling point on the YZ feature plane, F zx Represents the projection feature of the sampling point on the XY feature plane.

[0087] In order to ensure the image quality of the generated new viewing angle image, this embodiment adopts different point feature acquisition methods according to the difference between the expected viewing angle parameter and the current viewing angle parameter. Before acquiring the projection feature of each sampling point on the three-dimensional feature plane, the following method is further included:

[0088] Determine whether a parameter difference between a current viewing angle parameter and an expected viewing angle parameter reaches a difference threshold;

[0089] If the parameter difference between the current viewing angle parameter and the expected viewing angle parameter does not reach the difference threshold, the projection feature of each sampling point on the three-dimensional feature plane is obtained.

[0090] It should be noted that the current viewing angle parameter represents the current viewing angle of the object in the reference object image, and the expected viewing angle parameter represents the expected viewing angle of the object in the expected new viewing angle image to be generated. Accordingly, the parameter difference between the current viewing angle parameter and the expected viewing angle parameter represents the viewing angle difference between the current viewing angle and the expected viewing angle. The greater the parameter difference, the greater the viewing angle difference. In addition, the sampling space of the sampling points determined based on the pixel ray in the above embodiments is in the camera coordinate system. The present application finds that when the viewing angle difference is small, directly sampling the point features in the camera coordinate system can obtain better stability and ensure the image quality of the generated new viewing angle image. However, when the viewing angle difference is large, directly sampling the point features in the camera coordinate system has poor stability and cannot ensure the image quality of the generated new viewing angle image. Therefore, the present embodiment adopts different point feature acquisition methods according to the difference between the expected viewing angle parameter and the current viewing angle parameter.

[0091] First, the parameter difference between the current viewing angle parameter and the expected viewing angle parameter is obtained, and it is determined whether the parameter difference between the current viewing angle parameter and the expected viewing angle parameter reaches a difference threshold. The difference threshold is used as a dividing line for whether point feature sampling is stable in the camera coordinate system, and can be determined by those skilled in the art according to actual needs, and is not specifically limited here.

[0092] As above, if it is determined that the parameter difference between the current viewing angle parameter and the expected viewing angle parameter does not reach the difference threshold, it means that the sampling of point features in the camera coordinate system is still stable, and the point features of the pixel points obtained by sampling according to the point feature sampling method recorded in the above embodiment will not be repeated here.

[0093] Optionally, in one embodiment, after determining whether the parameter difference between the current viewing angle parameter and the expected viewing angle parameter reaches a difference threshold, the method further includes:

[0094] If the parameter difference between the current viewing angle parameter and the expected viewing angle parameter does not reach the difference threshold, then, according to the current viewing angle parameter, each sampling point is converted from the camera coordinate system to the object coordinate system corresponding to the object in the reference object image to obtain the converted sampling point;

[0095] Obtain the projection features of each transformed sampling point on the three-dimensional feature plane;

[0096] The projection features of each transformed sampling point are fused, and the point features corresponding to the pixel points are obtained according to the fused features of each transformed sampling point.

[0097] As described above, if it is determined that the parameter difference between the current viewing angle parameter and the expected viewing angle parameter reaches the difference threshold, it means that the sampling of point features in the camera coordinate system will become unstable, and this embodiment will no longer perform point feature sampling in the camera coordinate system. To solve the problem of unstable point feature sampling, this embodiment converts the sampling space from the camera coordinate system to the object coordinate system of the object in the reference object image.

[0098] It should be noted that the object coordinate system is associated with a specific object, and each object has its own specific coordinate system. The object coordinate systems of different objects are independent of each other and can be the same or different, without any connection between them. At the same time, the object coordinate system is bound to the object. Binding means that when the object moves or rotates, the object coordinate system undergoes the same translation or rotation, and the object coordinate system and the object move synchronously and are bound to each other.

[0099] In this embodiment, an object coordinate system of the object in the reference object image is also established. According to the current viewing angle parameter, each sampling point is converted from the camera coordinate system to the object coordinate system corresponding to the object in the reference object image to obtain the converted sampling point, which can be expressed as:

[0100] P object =R T s -1 (P cam -t);

[0101] Among them, P object represents the converted sampling point, R represents the rotation matrix in the current viewing angle parameter, T represents the transpose, s represents the scaling parameter in the current viewing angle parameter, P cam Represents the sampling point in the camera coordinate system, and t represents the translation vector in the current viewing angle parameter.

[0102] As mentioned above, after each sampling point is converted from the camera coordinate system to the object coordinate system corresponding to the object in the reference object image and the corresponding converted sampling point is obtained, for each converted sampling point, the projection feature of the converted sampling point on the three-dimensional feature plane is obtained, and the projection feature of each converted sampling point is fused, and the point feature corresponding to the pixel point is obtained according to the fused feature of each converted sampling point.

[0103] The point feature sampling in the object coordinate system may be implemented accordingly with reference to the point feature sampling method in the camera coordinate system in the above embodiment, which will not be described in detail here.

[0104] Optionally, in one embodiment, an optional sampling point selection rule is provided to determine multiple sampling points corresponding to pixel points on the pixel ray, including:

[0105] Determine a first preset number of sampling points equidistantly on the pixel ray;

[0106] According to the depth information of the pixel point, a second preset number of sampling points is determined on the pixel ray.

[0107] In this embodiment, multiple sampling points corresponding to the pixel points are determined on the pixel ray according to two sampling modes at the same time.

[0108] Among them, according to the pre-configured sampling interval, a first preset number of sampling points are equidistantly determined on the pixel ray. There is no specific limitation on the values ​​of the sampling interval and the first preset number, which can be configured by those skilled in the art according to actual needs.

[0109] In addition, for the pixel points in the reference object image, a second preset number of sampling points are determined near the depth thereof based on the depth information thereof. For example, the second preset number of sampling points can be determined within a preset distance range from the depth of the pixel point on the pixel ray. There is no specific limitation on the value of the preset distance range here, and the value can be determined by those skilled in the art according to actual needs.

[0110] Thus, for a pixel in the reference object image, a first preset number plus a second preset number of sampling points are determined. This embodiment can further improve the image quality of the generated new perspective image by selecting additional sampling points near the corresponding depth of the pixel as a supplement.

[0111] In 150, a target object image whose viewing angle meets the expected viewing angle parameters is generated according to the point features corresponding to the pixel points.

[0112] In this embodiment, for a pixel point in the reference object image, the color value of the pixel point is predicted based on the point feature corresponding to the pixel point. Thus, a target object image whose viewing angle meets the expected viewing angle parameters can be generated based on the predicted color value of the pixel point in the reference object image.

[0113] For example, please refer to Figure 1g , the reference object image is a facial image of a person from a frontal perspective. Assuming that the perspective indicated by the expected perspective parameter is a side perspective of the person, the target object image generated by the image processing method provided by the present application is a facial image of the person from a side perspective.

[0114] Optionally, in one embodiment, generating a target object image whose viewing angle meets the expected viewing angle parameter according to the point features corresponding to the pixel points includes:

[0115] According to the features of each point corresponding to the pixel, predict the corresponding color value and density value;

[0116] According to the color value and density value corresponding to each feature point, an image of the target object whose viewing angle meets the expected viewing angle parameters is generated.

[0117] In this embodiment, according to the point features of each sampling point corresponding to the pixel point, the color value and density value of the sampling point are predicted through the image rendering network; for a pixel point, according to the expected viewing angle parameters, the color value and density value of all the corresponding sampling points are integrated through the image rendering network to obtain the color value of the pixel point, and a new multi-channel feature map is obtained; then the feature maps of the first three channels of the new multi-channel feature map are merged into a new object image, that is, a target object image whose viewing angle meets the expected viewing angle parameters. For example, assuming that each dimensional feature plane of the generated three-dimensional feature plane corresponds to a feature map of 32 channels, and a feature map of 32 channels is obtained through the image rendering network, the feature maps of the first three channels are merged into the target object image in this embodiment.

[0118] Optionally, in one embodiment, to improve the efficiency of generating a new viewing angle image, the target object image is generated in two stages, wherein the target object image whose viewing angle meets the expected viewing angle parameters is generated according to the color value and density value corresponding to each point feature, including:

[0119] Generate an initial object image whose viewing angle meets the expected viewing angle parameter according to the color value and density value corresponding to each feature point, wherein the resolution of the initial object image is smaller than the resolution of the reference object image;

[0120] The initial object image is subjected to super-resolution processing, and the obtained super-resolution image is used as the target object image.

[0121] In this embodiment, the image rendering network includes a rendering sub-network and a supramolecular network, wherein the rendering sub-network is configured to generate a new perspective object image having a resolution less than the resolution of the original input object image, and the supramolecular network is configured to perform super-resolution processing on the new perspective object image generated by the rendering sub-network, and output a new perspective object image having a resolution greater than or equal to the original input object image. The rendering sub-network may adopt the architecture of a neural radiation field network, and the supramolecular network may adopt an architecture composed of multiple residual blocks, each of which is composed of a convolutional layer, an upsampling layer, and a ReLU function layer.

[0122] Correspondingly, when generating the target object image, first, according to the color value and density value corresponding to each feature point, a new perspective object image with a resolution smaller than the resolution of the benchmark object image is generated through the rendering subnetwork, which is recorded as the initial object image. Then, the initial object image generated by the rendering subnetwork is super-resolution processed through the supramolecular network to obtain a super-resolution image with a resolution greater than or equal to the resolution of the benchmark object image, and the super-resolution image is used as the target object image.

[0123] This embodiment can improve the rendering efficiency of the image by reducing the resolution of the object image actually rendered, thereby achieving the purpose of improving the overall processing efficiency. For example, assuming that the resolution of the original input object image is configured to be 512*512, the resolution of the object image generated by the image rendering subnetwork can be configured to be 128*128, and the output resolution of the super-resolution processing of the supramolecular network can be configured to be 512*512.

[0124] Optionally, in one embodiment, the original latent code of the reference object image is not directly input into the image generation model for processing, but is first optimized so that it can accurately represent the reference object image in the latent space corresponding to the generation subnetwork, and then input into the image generation model for processing to ensure the image quality of the generated new view image. According to the expected view parameters, the latent code is feature mapped to obtain a multi-channel feature map, including:

[0125] According to the expected viewing angle parameters, feature mapping is performed on the latent code to obtain a temporary multi-channel feature map;

[0126] Reconstruct the temporary multi-channel feature map into a temporary three-dimensional feature plane orthogonal to the axis;

[0127] According to the pixel ray, a temporary point feature corresponding to the pixel point is sampled in the temporary three-dimensional feature plane;

[0128] Generate a temporary object image according to the temporary point features corresponding to the pixel points;

[0129] Update the latent code according to the difference between the temporary object image and the reference object image at the same object position to obtain an updated latent code;

[0130] According to the expected viewing angle parameters, feature mapping is performed on the updated latent code to obtain a multi-channel feature map.

[0131] In this embodiment, the original latent code of the reference object image is input into the image generation model, and the latent code is updated according to the processed data of the image generation model.

[0132] Correspondingly, after the reference object image is inversely mapped by the generative adversarial network inversion model and inversely mapped back to the latent space corresponding to the generative subnetwork to obtain the latent code of the reference object image, the latent code of the reference object image and the expected viewing angle parameter (represented in the form of a vector) are mapped to a new latent code of the latent space corresponding to the generative subnetwork by the mapping subnetwork, which is recorded as a temporary latent code; then, the temporary latent code is input into the generative subnetwork, and a multi-channel feature map is generated by the generative subnetwork, which is recorded as a temporary multi-channel feature map, and the temporary multi-channel feature map is correspondingly reconstructed into an axis-orthogonal three-dimensional feature plane, which is recorded as a temporary three-dimensional feature plane; then, according to the pixel ray, the temporary point feature corresponding to the pixel point is sampled in the temporary three-dimensional feature plane; and according to the temporary point feature corresponding to the pixel point, the corresponding object image is generated by the image rendering network, which is recorded as a temporary object image; finally, according to the difference between the temporary object image and the reference object image at the same object position, the rendering loss is determined, and the hidden code is updated with the reduction of the rendering loss as a constraint until the second preset stop condition is met to obtain the updated hidden code. There is no specific restriction on the setting of the second preset stop condition here, and it can be set by technical personnel in this field according to actual needs. For example, the second preset stop condition can be set to the number of updates of the implicit code of the reference object image reaching a second preset number, or the second preset stop condition can be set to the rendering loss reaching a second loss threshold.

[0133] Among them, the rendering loss can be expressed as:

[0134]

[0135] L RGB represents the rendering loss, P represents the pixel of the temporary object image, and P in represents the pixel points belonging to the object content area, C p represents the color value of a position in the object content area in the reference object image, C p 'Indicates the color value of the position in the temporary object image.

[0136] As mentioned above, after completing the update of the latent code of the reference object image and obtaining the corresponding updated latent code, feature mapping is performed on the updated latent code according to the expected viewing angle parameters to obtain a multi-channel feature map for generating the target object image. Please refer to the relevant description in the above embodiments for details, which will not be repeated here.

[0137] This embodiment updates the implicit code of the reference object image in the above manner, which can constrain the geometry and texture of the object.

[0138] Optionally, in one embodiment, an optional training scheme for an image generation model is further provided, and before obtaining the reference object image, and the current viewing angle parameter and the expected viewing angle parameter of the reference object image, the method further includes:

[0139] Obtaining a sample object image and a sample current viewing angle parameter of the sample object image;

[0140] Obtain a random latent code, and generate the corresponding sample initial image and sample super-resolution image through an image generation model according to the random latent code and the current viewing angle parameters of the sample;

[0141] Upsampling the sample initial image to obtain an upsampled initial image with the same resolution as the sample super-resolution image;

[0142] Splicing the upsampled initial image and the sample super-resolution image to obtain a first image to be judged;

[0143] Performing blur processing on the sample object image to obtain a blurred sample object image;

[0144] splicing the blurred sample object image and the sample object image to obtain a second image to be judged;

[0145] The current viewing angle parameters of the sample are used as the discrimination reference of the discriminator network. The network parameters of the image generation model are updated with the constraint that the discrimination result of the discriminator network for the first image to be discriminated is false and the discrimination result of the second image to be discriminated is true until the preset stop condition is met.

[0146] In this embodiment, for different types of objects, sample object images of that type are used to train the image generation model. For example, for people, sample person images are used to train the image generation model. For a certain type of animal, sample images of that type of animal are used to train the image generation model, and so on.

[0147] Exemplarily, taking human objects as an example, face images can be obtained from the FFHQ and / or CelebA datasets as sample object images. For each sample object image, the pose detection algorithm in DECA is used to detect the sample object image, and the camera parameters obtained by the detection are recorded as the sample current viewing angle parameters of the sample object image.

[0148] In addition, a multi-dimensional random vector that conforms to the latent space corresponding to the image generation network is generated as a random hidden code, and according to the random hidden code and the current viewing angle parameter of the sample, a corresponding low-resolution image and super-resolution image are generated through the image generation model, and the generated low-resolution image is recorded as the sample initial image, and the generated super-resolution image is recorded as the sample super-resolution image. Please refer to the relevant description of the above embodiment for the image generation process, which will not be repeated here.

[0149] As above, after the corresponding sample initial image and sample super-resolution image are generated by the image generation model, the generated sample initial image is further upsampled to obtain an upsampled initial image with the same resolution as the sample super-resolution image, and the upsampled initial image and the sample super-resolution image are spliced ​​on the channel, and the obtained spliced ​​image is recorded as the first image to be judged.

[0150] In addition, the sample object image is blurred to obtain a blurred sample object image. For example, the sample object image is first downsampled by 4 times to obtain a downsampled image, and then the downsampled image is upsampled by 4 times to achieve blurring of the sample object image to obtain a blurred sample object image.

[0151] As described above, after the sample object image is blurred to obtain the blurred sample object image, the blurred sample object image and the sample object image are further spliced ​​on the channel, and the obtained spliced ​​image is recorded as the second image to be determined.

[0152] Finally, the sample current viewing angle parameter of the sample object image is used as the discriminant reference of the discriminator network, and the discriminator network's discriminant result for the first image to be discriminated is false and the discriminant result for the second image to be discriminated is true is constrained to update the network parameters of the image generation model until the preset stop condition is met. Among them, the discriminator network can adopt the discriminator network architecture of StyleGAN, and the training loss can adopt the loss of Pi-GAN (Periodic Implicit GenerativeAdversarial Networks, periodic implicitly generated adversarial network), which will not be elaborated here.

[0153] It should be noted that this embodiment does not impose any specific restrictions on the setting of the above preset stop conditions, which can be set by those skilled in the art according to actual needs. For example, the preset stop condition can be set to the number of updates of the network parameters of the image generation model reaching a first preset number of times, or the preset stop condition can be set to the training loss reaching a first loss threshold.

[0154] Optionally, in one embodiment, acquiring a reference object image includes:

[0155] Acquire an image of an object to be repaired, wherein an object region in the image of the object to be repaired is blocked;

[0156] Acquire a historical object image corresponding to the object image to be restored as a reference object image, wherein the object area in the historical object image is not blocked;

[0157] Obtain the desired viewing angle parameters of the reference object image, including:

[0158] Obtaining a viewing angle parameter of the image of the object to be repaired as an expected viewing angle parameter;

[0159] After generating a target object image whose viewing angle meets the expected viewing angle parameter according to the point features corresponding to the pixel points, the method further includes:

[0160] According to the target object image, the image of the object to be repaired is repaired to obtain a repaired image of the object.

[0161] This embodiment provides an optional application for generating a new-viewing angle image, and the generated new-viewing angle image is used for restoring the content of an image.

[0162] Among them, an object image to be repaired of the object is obtained, and the object area of ​​the object in the object image to be repaired (that is, the image content of the object in the object image to be repaired) is blocked, and a historical object image corresponding to the object image to be repaired is obtained, and the object area of ​​the object in the historical object image is not blocked. For example, taking the object as a talk show performer as an example, for the performance video of the talk show performer captured, there are parts of the image where the talk show performer is blocked. Accordingly, the blocked image of the talk show performer can be obtained from the performance video as the object image to be repaired, and the historical object image of the talk show performer that was not blocked before the object image to be repaired is obtained. In addition, the camera parameters of the historical object image detected by the posture detection algorithm are used as the current viewing angle parameters, and the camera parameters of the object image to be repaired are detected by the posture detection algorithm as the expected viewing angle parameters.

[0163] Based on the historical object image (reference object image), current viewing angle parameters and expected viewing angle parameters obtained above, the image processing method provided in the above embodiment of the present application is used to generate a target object image whose viewing angle meets the expected viewing angle parameters. Afterwards, the object image to be repaired is repaired according to the generated target object image, for example, the image content of the object area of ​​the object image to be repaired is directly replaced with the image content of the object area of ​​the generated target object image, so as to repair its image content and obtain a repaired image in which the object area is not blocked.

[0164] As can be seen from the above, the image processing solution provided by the present application obtains the reference object image, as well as the current viewing angle parameters and the expected viewing angle parameters of the reference object image, and then performs feature mapping on the reference object image according to the expected viewing angle parameters to obtain an axis-orthogonal three-dimensional feature plane. In addition, according to the current viewing angle parameters, the pixel rays of the pixel points in the reference object image are obtained, and the point features corresponding to the pixel points sampled by the pixel rays in the three-dimensional feature plane are used. Finally, according to the point features corresponding to the pixel points, a target object image whose viewing angle meets the expected viewing angle parameters is generated. In this way, by using the three-dimensional feature plane to represent the image in the feature space, a higher representation speed and resolution can be obtained, thereby realizing detailed content under the same capacity, and the use of pixel rays to sample the point features of the pixel points can increase the spatial perception ability, thereby improving the image generation quality. Compared with the related art, the image processing solution provided by the present application can synthesize new viewing angles more efficiently and with higher quality.

[0165] Please refer to Figure 2 The image processing method provided by the present application is described below by taking an electronic device as the execution subject and a human face as the object. Figure 2 As shown, the process of the image processing method can also be as follows:

[0166] In 210 , the electronic device obtains a reference facial image, and a current viewing angle parameter and an expected viewing angle parameter of the reference facial image.

[0167] In an embodiment of the present application, the electronic device obtains a facial image of a person for which a new perspective image needs to be generated in response to a demand for generating a new perspective image, and uses it as a reference for generating the new perspective image, and accordingly records it as a reference facial image. The reference facial image can be obtained by photographing the face of a person at a certain perspective through a device with image capturing capability such as a camera or a camera. For example, the electronic device photographs the face of a person at a frontal perspective through a camera to obtain a frontal perspective image of the face. In addition, the expected perspective parameter is used to indicate the perspective of the new perspective image that is expected to be generated. For example, the obtained expected perspective parameter indicates that the perspective of the new perspective image that is expected to be generated is a side perspective of the face.

[0168] In 220, the electronic device performs inverse mapping on the reference facial image by using a generative adversarial network inversion model to obtain a latent code of the reference facial image.

[0169] In this embodiment, an image generation model and a generative adversarial network inversion model are pre-trained. The image generation model is configured to take a facial image of the original perspective (in the form of hidden coding) and a desired perspective parameter as input, and a facial image of the perspective indicated by the desired perspective parameter as output. The image generation model may include a feature mapping network and an image rendering network, wherein the feature mapping network is configured to map the input facial image of the original perspective to a feature space according to the desired perspective parameter to obtain its hidden layer representation, and the image rendering network adopts the architecture of a neural radiation field network and is configured to map the hidden layer representation of the facial image of the original perspective back to the image space to obtain a facial image of a new perspective. Among them, the feature mapping network may include a mapping subnetwork and a generation subnetwork. The generating subnetwork can use the generator network in the generative adversarial network, such as the generator network of StyleGAN as the generating subnetwork, and the input data of the corresponding generating subnetwork needs to be the latent code of the latent space corresponding to the generating subnetwork, and the generating subnetwork is configured to generate the corresponding feature map according to the input latent code; the mapping subnetwork is configured to map the input original latent code and the expected viewing angle parameter to a new latent code that conforms to the aforementioned latent code form as the input of the generating subnetwork. There is no specific limitation on the architecture of the mapping subnetwork here, and it can be configured by those skilled in the art according to actual needs.

[0170] The generative adversarial network inversion model is configured to inversely map the image from the image space back to the latent space corresponding to the generative subnetwork to obtain the corresponding latent code. The architecture and training method of the generative adversarial network inversion model are not specifically limited here, and can be configured by those skilled in the art according to actual needs.

[0171] Correspondingly, in this embodiment, the electronic device performs inverse mapping on the reference facial image through the generative adversarial network inversion model, and inverse maps it back to the latent space corresponding to the generative subnetwork to obtain the latent code of the reference facial image.

[0172] In 230, the electronic device performs feature mapping on the latent code according to the desired viewing angle parameter to obtain a multi-channel feature map.

[0173] In 240, the electronic device reconstructs the multi-channel feature map into three-dimensional feature planes with orthogonal axes, and the number of channels of the feature map corresponding to each feature plane is the same.

[0174] As mentioned above, after the electronic device obtains the latent code of the reference facial image by inverse mapping, it further maps the latent code of the reference facial image and the expected viewing angle parameter (represented in the form of a vector) into a new latent code of the latent space corresponding to the generation subnetwork through the mapping subnetwork; finally, the new latent code is input into the generation subnetwork, and a multi-channel feature map (denoted as a multi-channel feature map) is generated through the generation subnetwork, and the multi-channel feature map is reconstructed into an axis-orthogonal three-dimensional feature plane, where the number of channels of the feature map corresponding to each dimensional feature plane is the same.

[0175] For example, taking the generator network of StyleGAN adopted in the generative subnetwork as an example, the reference facial image is inversely mapped through the generative adversarial network inversion model, and it is inversely mapped back to the latent space corresponding to the generative subnetwork to obtain the latent code of the reference facial image (a 512-dimensional vector); then, the latent code of the reference facial image and the expected viewing angle parameter (a 25-dimensional vector) are mapped to a new latent code (a 512-dimensional vector) of the latent space corresponding to the generative subnetwork through the mapping subnetwork; finally, the new latent code is input into the generative subnetwork, and a multi-channel feature map (for example, 96 channels) is generated through the generative subnetwork, and the multi-channel feature map is reconstructed into an axis-orthogonal three-dimensional feature plane, where each dimensional feature plane corresponds to a feature map of 32 channels.

[0176] In 250, the electronic device obtains pixel rays of pixel points in the reference facial image according to the current viewing angle parameter.

[0177] As described above, after feature mapping is performed on the reference facial image to obtain the axis-orthogonal three-dimensional feature plane, the corresponding pixel ray is obtained for the pixel point in the reference facial image according to the current viewing angle parameter. The pixel ray of a pixel point can be generally understood as the ray emitted from the pixel point to the camera.

[0178] Exemplarily, the imaging plane of the reference facial image may be determined first according to the current viewing angle parameter, and for a pixel point in the reference facial image, a ray perpendicular to the imaging plane is determined based on the pixel point as the pixel ray of the pixel point.

[0179] In 260 , the electronic device determines a plurality of sampling points corresponding to the pixel points on the pixel ray, and transforms the plurality of sampling points from the camera coordinate system to the object coordinate system corresponding to the face in the reference facial image to obtain a plurality of transformed sampling points.

[0180] In this embodiment, for a pixel point in the reference facial image, multiple sampling points corresponding to the pixel point can be determined on the pixel ray corresponding to the pixel point according to the pre-configured sampling point selection rule. For example, for a pixel point, the electronic device determines a first preset number of sampling points equidistantly on the pixel ray corresponding to the pixel point; and determines a second preset number of sampling points on the pixel ray according to the depth information of the pixel point.

[0181] In this embodiment, an object coordinate system of the face in the reference facial image is also established. The electronic device converts each sampling point from the camera coordinate system to the object coordinate system corresponding to the face in the reference facial image according to the current viewing angle parameter, and obtains the converted sampling point, which can be expressed as:

[0182] P object =R T s -1 (P cam -t);

[0183] Among them, P object represents the converted sampling point, R represents the rotation matrix in the current viewing angle parameter, T represents the transpose, s represents the scaling parameter in the current viewing angle parameter, P cam Represents the sampling point in the camera coordinate system, and t represents the translation vector in the current viewing angle parameter.

[0184] In 270 , the electronic device samples the plurality of converted sampling points in the three-dimensional feature plane to obtain point features corresponding to the pixel points.

[0185] As described above, for a pixel point in the reference facial image, after determining the corresponding multiple sampling points from the corresponding pixel ray and converting them from the camera coordinate system to the object coordinate system, for a converted sampling point, project it to each dimensional feature plane of the three-dimensional feature plane to obtain the projection features of the converted sampling point on each dimensional feature plane. For example, for a converted sampling point (x 0 ,y 0 , z 0 ), project the transformed sampling point to the XY feature plane sampling (x 0 ,y 0 ) point features to obtain the projection feature F xy , project the transformed sampling point to the YZ feature plane sampling (y 0 , z 0 ) point features to obtain the projection feature F yz , project the transformed sampling point to the ZX feature plane sampling (z 0 , x 0 ) point to obtain feature F zx .

[0186] Finally, for each converted sampling point, the projection features of the converted sampling point on each dimensional feature plane are fused to obtain the fused features of the converted sampling point as the point features of the pixel point. In this way, for a pixel point in the reference facial image, the corresponding number of point features are obtained according to the determined number of sampling points.

[0187] Among them, the process of fusing the projection features of the transformed sampling points can be expressed as:

[0188] F=F xy +F yz +F zx ;

[0189] Among them, F represents the fusion feature of a converted sampling point, that is, the point feature of the corresponding pixel point, F xy Indicates the projection feature of the transformed sampling point on the XY feature plane, F yz Represents the projection feature of the transformed sampling point on the YZ feature plane, F zx Represents the projection feature of the transformed sampling point on the XY feature plane.

[0190] In 280, the electronic device generates a target facial image whose viewing angle meets the expected viewing angle parameters according to the point features corresponding to the pixel points.

[0191] In this embodiment, the image rendering network includes a rendering sub-network and a supramolecular network, wherein the rendering sub-network is configured to generate a new-perspective facial image having a resolution less than the resolution of the original input facial image, and the supramolecular network is configured to perform super-resolution processing on the new-perspective facial image generated by the rendering sub-network, and output a new-perspective facial image having a resolution greater than or equal to the original input facial image. The rendering sub-network may adopt the architecture of a neural radiation field network, and the supramolecular network may adopt an architecture composed of multiple residual blocks, each of which is composed of a convolutional layer, an upsampling layer, and a ReLU function layer.

[0192] Accordingly, when generating a target facial image, the electronic device first predicts the corresponding color value and density value through the rendering sub-network according to the point features corresponding to the pixel points, and then generates a new perspective facial image with a resolution less than the resolution of the reference facial image through the rendering sub-network according to the color value and density value corresponding to each point feature, which is recorded as the initial facial image, and then the initial facial image generated by the rendering sub-network is super-resolution processed through the supramolecular network to obtain a super-resolution image with a resolution greater than or equal to the resolution of the reference facial image, and the super-resolution image is used as the target facial image. For example, assuming that the resolution of the original input facial image is configured to be 512*512, the resolution of the facial image generated by the image rendering sub-network can be configured to be 128*128, and the output resolution of the super-resolution processing of the supramolecular network can be configured to be 512*512.

[0193] In order to better implement the above image processing method, the present application embodiment also provides a corresponding image processing device, wherein the meanings of the terms are the same as those in the above image processing method, and the specific implementation details can be found in the description of the above method embodiment.

[0194] Please refer to Figure 3 , Figure 3 The structure diagram of the image processing device provided in the embodiment of the present application is as follows. The image processing device may include a reference acquisition module 310, a feature mapping module 320, a ray acquisition module 330 and a feature sampling module 340, wherein:

[0195] A reference acquisition module 310, for acquiring a reference object image, and current viewing angle parameters and expected viewing angle parameters of the reference object image;

[0196] A feature mapping module 320 is used to perform feature mapping on the reference object image according to the desired viewing angle parameters to obtain a three-dimensional feature plane with orthogonal axes;

[0197] A ray acquisition module 330, used to acquire pixel rays of pixel points in the reference object image according to the current viewing angle parameter;

[0198] A feature sampling module 340 is used to sample point features corresponding to pixel points in a three-dimensional feature plane according to pixel rays;

[0199] The image rendering module 350 is used to generate a target object image whose viewing angle meets the expected viewing angle parameters according to the point features corresponding to the pixel points.

[0200] Optionally, in one embodiment, the feature sampling module 340 is used to determine multiple sampling points corresponding to pixel points on the pixel ray, and obtain the projection features of each sampling point on the three-dimensional feature plane; and fuse the projection features of each sampling point to obtain the point features corresponding to the pixel point according to the fused features of each sampling point.

[0201] Optionally, in one embodiment, the feature sampling module 340 is used to determine whether the parameter difference between the current viewing parameter and the expected viewing parameter reaches a difference threshold; and if the parameter difference between the current viewing parameter and the expected viewing parameter does not reach the difference threshold, the projection feature of each sampling point on the three-dimensional feature plane is obtained.

[0202] Optionally, in one embodiment, the feature sampling module 340 is also used to convert each sampling point from the camera coordinate system to the object coordinate system corresponding to the object in the reference object image according to the current viewing angle parameter if the parameter difference between the current viewing angle parameter and the expected viewing angle parameter does not reach the difference threshold, so as to obtain the converted sampling point; and obtain the projection feature of each converted sampling point on the three-dimensional feature plane; and fuse the projection feature of each converted sampling point to obtain the point feature corresponding to the pixel point according to the fused feature of each converted sampling point.

[0203] Optionally, in one embodiment, the image rendering module 350 is used to predict the corresponding color value and density value based on each point feature corresponding to the pixel point; and generate a target object image whose perspective meets the expected perspective parameters based on the color value and density value corresponding to each point feature.

[0204] Optionally, in one embodiment, the image rendering module 350 is used to generate an initial object image whose perspective meets the expected perspective parameters based on the color value and density value corresponding to each point feature, and the resolution of the initial object image is smaller than the resolution of the reference object image; and to perform super-resolution processing on the initial object image and use the obtained super-resolution image as the target object image.

[0205] Optionally, in one embodiment, the image processing device is executed through an image generation model, and the image processing device also includes a model training module, which is used to obtain a sample object image and a sample current viewing angle parameter of the sample object image; and obtain a random latent code, and generate a corresponding sample initial image and a sample super-resolution image through the image generation model according to the random latent code and the sample current viewing angle parameter; and upsample the sample initial image to obtain an upsampled initial image with the same resolution as the sample super-resolution image; and splice the upsampled initial image and the sample super-resolution image to obtain a first image to be judged; and blur the sample object image to obtain a blurred sample object image; and splice the blurred sample object image and the sample object image to obtain a second image to be judged; and use the sample current viewing angle parameter as a discrimination reference of the discriminator network, and update the network parameters of the image generation model with the constraint that the discrimination result of the discriminator network for the first image to be judged is false and the discrimination result of the second image to be judged is true until a preset stop condition is met.

[0206] Optionally, in one embodiment, the feature sampling module 340 is used to determine a first preset number of sampling points equidistantly on the pixel ray; and determine a second preset number of sampling points on the pixel ray according to depth information of the pixel point.

[0207] Optionally, in one embodiment, the feature mapping module 320 is used to perform inverse mapping on the reference object image to obtain an implicit code of the reference object image; and perform feature mapping on the implicit code according to the desired viewing angle parameters to obtain a multi-channel feature map; and reconstruct the multi-channel feature map into an axis-orthogonal three-dimensional feature plane, and the number of channels of the feature map corresponding to each feature plane is the same.

[0208] Optionally, in one embodiment, the feature mapping module 320 is used to perform feature mapping on the latent code according to the expected viewing angle parameters to obtain a temporary multi-channel feature map; and reconstruct the temporary multi-channel feature map into an axis-orthogonal temporary three-dimensional feature plane; and sample the temporary point features corresponding to the pixels in the temporary three-dimensional feature plane according to the pixel rays; and generate a temporary object image according to the temporary point features corresponding to the pixels; and update the latent code according to the difference between the temporary object image and the reference object image at the same object position to obtain an updated latent code; and perform feature mapping on the updated latent code according to the expected viewing angle parameters to obtain a multi-channel feature map.

[0209] Optionally, in one embodiment, the reference acquisition module 310 is used to acquire an object image to be restored of the object, wherein the object region in the object image to be restored is blocked; and acquire a historical object image corresponding to the object image to be restored as the reference object image, wherein the object region in the historical object image is not blocked; and acquire a viewing angle parameter of the object image to be restored as the expected viewing angle parameter;

[0210] The image processing device also includes an image restoration module, which is used to restore the image of the object to be restored according to the target object image to obtain a restored image of the object.

[0211] The specific implementation of each of the above modules can be found in the previous embodiments and will not be described in detail here.

[0212] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the processor is used to execute the steps of the image processing method provided in the above embodiment by calling a computer program stored in the memory.

[0213] Please refer to Figure 4 , Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0214] The electronic device may include components such as a processor 101 with one or more processing cores, a memory 102 with one or more computer-readable storage media, a power supply 103, and an input unit 104. Those skilled in the art will appreciate that Figure 4 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0215] The processor 101 is the control center of the electronic device, which uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 102, and calling data stored in the memory 102. Optionally, the processor 101 may include one or more processing cores; optionally, the processor 101 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 101.

[0216] The memory 102 can be used to store software programs and modules. The processor 101 executes various functional applications and data processing by running the software programs and modules stored in the memory 102. The memory 102 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 102 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 102 may also include a memory controller to provide the processor 101 with access to the memory 102.

[0217] The electronic device also includes a power supply 103 for supplying power to each component. Optionally, the power supply 103 can be logically connected to the processor 101 through a power management system, so as to manage charging, discharging, and power consumption through the power management system. The power supply 103 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0218] The electronic device may further include an input unit 104, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0219] Although not shown, the electronic device may also include a display unit, an image acquisition component, etc., which will not be described in detail here. Specifically in this embodiment, the processor 101 loads the executable code corresponding to one or more computer programs into the memory 102, and the processor 101 executes the steps in the image processing method provided in this application, such as:

[0220] Acquire a reference object image, and a current viewing angle parameter and an expected viewing angle parameter of the reference object image;

[0221] According to the desired viewing angle parameters, feature mapping is performed on the reference object image to obtain a three-dimensional feature plane orthogonal to the axis;

[0222] According to the current viewing angle parameter, a pixel ray of a pixel point in the reference object image is obtained;

[0223] According to the pixel ray, the point features corresponding to the pixel points are sampled in the three-dimensional feature plane;

[0224] According to the point features corresponding to the pixel points, an image of the target object whose viewing angle meets the expected viewing angle parameters is generated.

[0225] It should be noted that the electronic device provided in the embodiment of the present application and the image processing method in the above embodiment belong to the same concept, and its specific implementation process is detailed in the above related embodiments and will not be repeated here.

[0226] The present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program stored therein is executed on a processor of an electronic device provided in an embodiment of the present application, the processor of the electronic device implements the steps in the image processing method provided in the present application. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0227] The present application also provides a computer program product, which includes a computer program. When the computer program is executed on a processor of an electronic device provided in an embodiment of the present application, the processor of the electronic device implements the steps in the image processing method provided in the present application.

[0228] The image processing method, image processing device, electronic device, computer-readable storage medium and computer program product provided by the present application are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

[0229] It should be noted that when the above embodiments of the present application are applied to specific products or technologies, the relevant data of the user is involved, and the user's permission or consent is required, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

Claims

1. An image processing method, It is characterized in that include: Acquire a reference object image, and a current viewing angle parameter and an expected viewing angle parameter of the reference object image; Performing feature mapping on the reference object image according to the desired viewing angle parameter to obtain an axis-orthogonal three-dimensional feature plane; According to the current viewing angle parameter, obtaining pixel rays of pixel points in the reference object image; According to the pixel ray, sampling is performed in the three-dimensional feature plane to obtain a point feature corresponding to the pixel point; A target object image whose viewing angle meets the expected viewing angle parameter is generated according to the point features corresponding to the pixel points.

2. The image processing method according to claim 1, It is characterized in that The sampling, according to the pixel ray, in the three-dimensional feature plane to obtain a point feature corresponding to the pixel point includes: Determine a plurality of sampling points corresponding to the pixel point on the pixel ray, and obtain a projection feature of each sampling point on the three-dimensional feature plane; The projection features of each sampling point are fused, and the point features corresponding to the pixel point are obtained according to the fused features of each sampling point.

3. The image processing method according to claim 2, It is characterized in that Before obtaining the projection feature of each sampling point on the three-dimensional feature plane, the method further includes: Determining whether a parameter difference between the current viewing angle parameter and the expected viewing angle parameter reaches a difference threshold; If the parameter difference between the current viewing angle parameter and the expected viewing angle parameter does not reach the difference threshold, the projection feature of each sampling point on the three-dimensional feature plane is obtained.

4. The image processing method according to claim 2, It is characterized in that After determining whether the parameter difference between the current viewing angle parameter and the expected viewing angle parameter reaches a difference threshold, the method further includes: If the parameter difference between the current viewing angle parameter and the expected viewing angle parameter does not reach the difference threshold, then, according to the current viewing angle parameter, each sampling point is converted from the camera coordinate system to the object coordinate system corresponding to the object in the reference object image to obtain a converted sampling point; Obtaining the projection feature of each converted sampling point on the three-dimensional feature plane; The projection features of each transformed sampling point are fused, and the point features corresponding to the pixel point are obtained according to the fused features of each transformed sampling point.

5. The image processing method according to claim 2, It is characterized in that Generating a target object image whose viewing angle meets the expected viewing angle parameter according to the point features corresponding to the pixel points comprises: According to each point feature corresponding to the pixel point, predict the corresponding color value and density value; According to the color value and density value corresponding to each feature point, a target object image whose viewing angle meets the expected viewing angle parameters is generated.

6. The image processing method according to claim 5, It is characterized in that Generating a target object image whose viewing angle meets the expected viewing angle parameter according to the color value and density value corresponding to each point feature includes: Generate an initial object image whose viewing angle meets the expected viewing angle parameter according to the color value and density value corresponding to each point feature, wherein the resolution of the initial object image is smaller than the resolution of the reference object image; The initial object image is subjected to super-resolution processing, and the obtained super-resolution image is used as the target object image.

7. The image processing method according to claim 6, It is characterized in that The image processing method is performed by an image generation model, and before acquiring the reference object image, and the current viewing angle parameter and the expected viewing angle parameter of the reference object image, it also includes: Acquire a sample object image and a sample current viewing angle parameter of the sample object image; Obtaining a random latent code, and generating a corresponding sample initial image and a sample super-resolution image through the image generation model according to the random latent code and the current viewing angle parameter of the sample; Upsampling the sample initial image to obtain an upsampled initial image with the same resolution as the sample super-resolution image; splicing the upsampled initial image and the sample super-resolution image to obtain a first image to be determined; Performing blur processing on the sample object image to obtain a blurred sample object image; splicing the blurred sample object image and the sample object image to obtain a second image to be determined; The current viewing angle parameter of the sample is used as a discrimination reference of the discriminator network, and the network parameters of the image generation model are updated with the constraint that the discrimination result of the discriminator network for the first image to be discriminated is false and the discrimination result of the second image to be discriminated is true, until a preset stop condition is met.

8. The image processing method according to claim 2, It is characterized in that The determining of a plurality of sampling points corresponding to the pixel point on the pixel ray comprises: Determining a first preset number of sampling points equidistantly on the pixel ray; A second preset number of sampling points is determined on the pixel ray according to the depth information of the pixel point.

9. The image processing method according to claim 1, It is characterized in that The step of performing feature mapping on the reference object image according to the expected viewing angle parameter to obtain an axis-orthogonal three-dimensional feature plane comprises: Performing inverse mapping on the reference object image to obtain a latent code of the reference object image; Performing feature mapping on the latent code according to the expected viewing angle parameter to obtain a multi-channel feature map; The multi-channel feature map is reconstructed into a three-dimensional feature plane with orthogonal axes, and the number of channels of the feature map corresponding to each feature plane is the same.

10. The image processing method according to claim 9, It is characterized in that The step of performing feature mapping on the latent code according to the expected viewing angle parameter to obtain a multi-channel feature map includes: Performing feature mapping on the latent code according to the expected viewing angle parameter to obtain a temporary multi-channel feature map; Reconstructing the temporary multi-channel feature map into a temporary three-dimensional feature plane with axes orthogonal to each other; According to the pixel ray, sampling is performed in the temporary three-dimensional feature plane to obtain a temporary point feature corresponding to the pixel point; generating a temporary object image according to the temporary point features corresponding to the pixel points; updating the latent code according to the difference between the temporary object image and the reference object image at the same object position to obtain an updated latent code; According to the expected viewing angle parameter, feature mapping is performed on the updated latent code to obtain a multi-channel feature map.

11. The image processing method according to any one of claims 1 to 10, It is characterized in that The obtaining of the reference object image comprises: Acquire an image of an object to be repaired, wherein an object region in the image of the object to be repaired is blocked; Acquire a historical object image corresponding to the object image to be restored as a reference object image, wherein the object area in the historical object image is not blocked; Obtain the desired viewing angle parameters of the reference object image, including: Acquiring a viewing angle parameter of the image of the object to be repaired as an expected viewing angle parameter; After generating a target object image whose viewing angle meets the expected viewing angle parameter according to the point features corresponding to the pixel points, the method further includes: The image of the object to be repaired is repaired according to the target object image to obtain a repaired image of the object.

12. An image processing device, It is characterized in that include: A reference acquisition module, used to acquire a reference object image, and current viewing angle parameters and expected viewing angle parameters of the reference object image; A feature mapping module, used for performing feature mapping on the reference object image according to the expected viewing angle parameter to obtain a three-dimensional feature plane with axes orthogonal to each other; A ray acquisition module, used for acquiring pixel rays of pixel points in the reference object image according to the current viewing angle parameter; A feature sampling module, used for sampling in the three-dimensional feature plane according to the pixel ray to obtain a point feature corresponding to the pixel point; The image rendering module is used to generate a target object image whose viewing angle meets the expected viewing angle parameters according to the point features corresponding to the pixel points.

13. An electronic device, It is characterized in that The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program in the memory to implement the steps in the image processing method according to any one of claims 1 to 11.

14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being executed by a processor to implement the steps in the image processing method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the steps of the image processing method according to any one of claims 1 to 11 are implemented.