Depth image generation method and device, electronic equipment and storage medium

By using reference images with high ambient brightness and generation adversarial networks in the iterative training of the target generation model, the accuracy of image generation depth images is solved under dark light conditions, and more efficient depth feature extraction and image generation are achieved.

CN120259394APending Publication Date: 2025-07-04ORIENTAL POWER HOLDINGS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410008701.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When the prior art images are generated by images captured under dark light conditions, it is difficult to accurately capture object details, especially the depth differences caused by the strong light and reflective areas, which cannot accurately generate depth images.

Method used

By training the target generation model, reference images with a higher ambient brightness than the sample image in each iterative training, combined with the feature differences between the samples and reference images in multiple feature dimensions, parameter adjustments are made to the initial generation model, and the generative adversarial network and discriminator are used to align features and depth images to reduce pixel-level depth information annotation.

Benefits of technology

It improves the accuracy of image generation depth features under different ambient brightness, reduces the time to label samples, and improves model training efficiency and the accuracy of depth images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259394A_ABST
    Figure CN120259394A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a depth image generation method and device, electronic equipment and a storage medium, and is used for improving the accuracy of a generated depth image. The method comprises the steps of inputting a target image into a trained target generation model for depth information extraction, and obtaining a target depth feature; obtaining a depth image of the target image based on the target depth feature; and in each round of iterative training of the target generation model, performing parameter adjustment on the initial generation model by using the feature difference of the selected sample image and the reference image under multiple feature dimensions and combining the difference between the depth images of the sample image and the reference image. The environment brightness of the shot reference image is higher than the environment brightness of the shot sample image, the extracted features and the generated depth image are more accurate, the difference between the features and the depth image is aligned in the training process, the depth information extraction capability of the model can be improved, and then the more accurate depth image is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In a computer vision system, three-dimensional scene information provides more possibilities for various computer vision applications such as image segmentation and object detection. Correspondingly, as a common expression of three-dimensional scene information, depth images have also been widely used. The pixel value of each pixel point in a depth image can be used to represent the distance of a certain point in the scene from the image collector. Through the depth image, the geometric shape of the object surface in the scene can be reflected.

[0003] Currently, for images taken under well-lit conditions, corresponding depth images can already be generated relatively accurately through neural network models. However, for images taken under low-light conditions (hereinafter referred to as night images), due to poor lighting conditions, it is difficult to accurately capture the details of the objects in the images, and the accuracy of the generated depth images is poor.

[0004] In related technologies, in order to improve the generation accuracy of depth images corresponding to night images, mainly some image enhancement means (for example, adjusting image brightness and saturation) are first used to enhance the night images, and then the enhanced images are input into the neural network model to obtain the output depth images.

[0005] However, image enhancement can only help restore the details of the regions with poor brightness difference in night images. There may also be some regions with strong light and reflection in night images. The distance between these regions with strong light and reflection and the image collector is actually not much different from the distance between the surrounding regions and the image collector, but there will be an obvious brightness difference in the night images, which will then lead to a large depth difference in the depth images. However, the above-mentioned image enhancement means cannot eliminate the regions with strong light and reflection in night images, nor can they eliminate the brightness difference with the surrounding regions, so they cannot accurately generate the depth images corresponding to night images. Summary of the Invention

[0006] Embodiments of the present application provide a method, device, electronic device and storage medium for generating depth images to improve the accuracy of the generated depth images.

[0007] A method for generating a depth image provided by an embodiment of the present application includes:

[0008] Input the obtained target image into a trained target generation model for depth information extraction to obtain target depth features;

[0009] Based on the target depth features, obtain the depth image of the target image;

[0010] Among them, in each round of iterative training of the target generation model, the feature differences of the selected sample image and the reference image in multiple feature dimensions are used, combined with the differences between the depth images of the sample image and the reference image respectively, to adjust the parameters of the initial generation model; the environmental brightness when shooting the reference image is higher than the environmental brightness when shooting the sample image.

[0011] A depth image generation device provided by an embodiment of the present application includes:

[0012] An input unit, configured to input an acquired target image into a trained target generation model for depth information extraction to obtain target depth features;

[0013] A processing unit, configured to obtain a depth image of the target image based on the target depth features;

[0014] Among them, in each round of iterative training of the target generation model, the feature differences of the selected sample image and the reference image in multiple feature dimensions are used, combined with the differences between the depth images of the sample image and the reference image respectively, to adjust the parameters of the initial generation model; the environmental brightness when shooting the reference image is higher than the environmental brightness when shooting the sample image.

[0015] Optionally, the device includes a training unit, configured to, in each round of iterative training, adjust the parameters of the initial generation model through the following steps:

[0016] Based on the differences between multiple sample extraction features of the sample image and multiple reference extraction features of the reference image, construct a first loss function, where the multiple sample extraction features are obtained by performing multiple rounds of feature extraction on the sample image;

[0017] Based on the differences between the depth images of the sample image and the reference image respectively, construct a second loss function;

[0018] Obtain a target loss function based at least on the first loss function and the second loss function;

[0019] Based on the target loss function, adjust the parameters of the initial generation model.

[0020] Optionally, the training unit is specifically configured to:

[0021] For each of the multiple sample extraction features, perform the following operations respectively: input a sample extraction feature and a corresponding reference extraction feature into a discriminator to obtain a corresponding feature discrimination result; where the size of the one sample extraction feature is the same as that of the corresponding reference extraction feature;

[0022] Construct the first loss function based on the discrimination results of each feature.

[0023] Optionally, the training unit is further configured to:

[0024] Perform feature fusion on the feature extracted from the sample and the preset template feature to obtain a sample fusion feature, where the preset template feature is used to represent: the reference distribution of the depth information of each pixel point included in the sample image, and the depth information is used to represent: the distance between the object to which the corresponding pixel point belongs in the sample image and the image collector that captures the sample image;

[0025] Perform feature fusion on the reference extraction feature and the preset template feature to obtain a reference fusion feature;

[0026] Then, inputting the feature extracted from one sample and the corresponding reference extraction feature into the discriminator includes:

[0027] Input the sample fusion feature and the reference fusion feature into the discriminator.

[0028] Optionally, the training unit is further configured to:

[0029] Adjust the parameters of the discriminator based on the discrimination results of each feature.

[0030] Optionally, the training unit is specifically configured to:

[0031] Construct a third loss function based on the difference between multiple sample reconstruction features of the sample image and multiple reference reconstruction features of the reference image, where the multiple sample reconstruction features are obtained by performing multiple rounds of feature reconstruction on the sample extraction features obtained by the last round of feature extraction;

[0032] Obtain the target loss function based on at least the first loss function, the second loss function, and the third loss function.

[0033] Optionally, the input of the first round of feature extraction is the sample image, and the input of each subsequent round of feature extraction is the output of the previous round of feature extraction;

[0034] The input of the first round of feature reconstruction is the sample extraction feature obtained by the last round of feature extraction, and the input of each subsequent round of feature reconstruction is the output of the previous round of feature reconstruction.

[0035] Optionally, the training unit is specifically configured to:

[0036] Crop the sample depth images of the respective reconstructed features of the multiple samples to obtain respective target region images, where the sample depth images are obtained by convolving the corresponding sample reconstructed features, and each target region image includes: pixel points on the diagonal of the corresponding sample depth image;

[0037] After sorting the target region images according to the sizes of the corresponding sample reconstructed features, construct a fourth loss function based on the pixel differences between every two adjacent target region images;

[0038] Obtain the target loss function based on the first loss function, the second loss function, the third loss function, and the fourth loss function.

[0039] Optionally, the training unit is specifically configured to:

[0040] For every two adjacent target region images, perform the following operations respectively:

[0041] Upsample the target region image with the smaller size among the two target region images to obtain a target sampled image, and the target sampled image has the same size as the target region image with the larger size among the two target region images;

[0042] Based on the difference in pixel values between the same-position pixel points included in the target sampled image and the target region image with the larger size, obtain a pixel loss function;

[0043] Construct the fourth loss function based on the obtained pixel loss functions.

[0044] Optionally, the processing unit is specifically configured to obtain the depth image of the sample image in the following manner:

[0045] Extract depth information from the sample image to obtain sample depth features;

[0046] Perform convolution based on the sample depth features to obtain the depth image of the sample image.

[0047] An electronic device provided by an embodiment of the present application includes a processor and a memory, where the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of any of the above depth image generation methods.

[0048] An embodiment of the present application provides a computer-readable storage medium, which includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of any of the above depth image generation methods.

[0049] An embodiment of the present application provides a computer program product, the computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any of the above-mentioned depth image generation methods.

[0050] The beneficial effects of the present application are as follows:

[0051] An embodiment of the present application provides a depth image generation method, device, electronic device and storage medium. With the help of a trained target generation model, the present application extracts depth information from a target image to obtain target depth features, and then obtains a depth image of the target image based on the target depth features. Regarding the target generation model used, its main function is to extract the depth features of the image. In order to improve the accuracy of the features extracted by the model, during the process of training the target generation model, the depth features of the sample image are first extracted by the initial generation model. At this time, the extraction ability of the model is limited, and it may also be affected by factors such as lighting conditions, strong light, and reflection, resulting in the inability to accurately capture the depth features in the sample image.

[0052] Based on this, a reference image is introduced in the present application. Since the environmental brightness when shooting the reference image is higher than the environmental brightness when shooting the sample image, as described in the background art, for images taken under good lighting conditions, features can already be accurately extracted through a neural network model and corresponding depth images can be generated. Therefore, in each round of iterative training of the target generation model, the feature differences between the selected sample image and the reference image in multiple feature dimensions are used, combined with the differences between the depth images of the sample image and the reference image respectively, to adjust the parameters of the initial generation model, providing more dimensions of comparison basis for the parameter adjustment process and helping to improve the training effect of the model.

[0053] In this way, through the dual constraints in the feature domain and the image domain, the features of the sample image extracted by the generation model are made to be close to the features of the reference image. Furthermore, regardless of the lighting conditions of the image input to the generation model, the extracted depth features can always be close to the actual feature distribution, enabling the target generation model to extract accurate depth features for images under different environmental brightnesses. Further, based on the depth features output by the target generation model, a more accurate depth image can be obtained.

[0054] In addition, during the process of adjusting the model parameters, since the differences between the sample image and the reference image in the feature dimension and the depth image dimension are mainly considered, there is no need to perform pixel-level depth information annotation on the sample image, which reduces the time for annotating samples. Moreover, the features of the reference image and the depth image can also be directly obtained using the neural network model in related technologies, reducing the difficulty of constructing training samples, improving the construction efficiency of training samples, thereby reducing the time for the entire model training and improving the model training efficiency.

[0055] Other features and advantages of the present application will be described in the following specification. Moreover, some of them will become apparent from the specification or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings

[0056] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0057] Figure 1A It is a schematic diagram of the workflow of the generative adversarial network;

[0058] Figure 1B It is a schematic diagram of a night image and the corresponding depth image;

[0059] Figure 1C It is an optional schematic diagram of an application scenario in an embodiment of the present application;

[0060] Figure 2 It is an implementation flowchart of a depth image generation method in an embodiment of the present application;

[0061] Figure 3 It is a schematic diagram of the structure of a target generation model in an embodiment of the present application;

[0062] Figure 4 It is a schematic diagram of a method for constructing a second loss function in an embodiment of the present application;

[0063] Figure 5A It is a schematic diagram of a target region image in an embodiment of the present application;

[0064] Figure 5B It is another schematic diagram of a target region image in an embodiment of the present application;

[0065] Figure 6 It is a schematic diagram of the construction process of the fourth loss function in an embodiment of the present application;

[0066] Figure 7A It is a schematic diagram of the distribution of objects in an image in an embodiment of the present application;

[0067] Figure 7B It is a schematic diagram of a divided area in an embodiment of the present application;

[0068] Figure 7C It is a schematic diagram of preset template features in an embodiment of the present application;

[0069] Figure 8 It is a schematic diagram of a feature discrimination method in an embodiment of the present application;

[0070] Figure 9 It is a schematic diagram of the execution process of one round of iterative training in an embodiment of the present application;

[0071] Figure 10 It is a schematic diagram of the structure of an image processing device in an embodiment of the present application;

[0072] Figure 11 It is a schematic diagram of a hardware composition structure of an electronic device applying an embodiment of the present application;

[0073] Figure 12 It is a schematic diagram of a hardware composition structure of another electronic device applying an embodiment of the present application. Detailed implementation manners

[0074] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the technical solutions of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts shall fall within the scope of protection of the technical solutions of the present application.

[0075] Terms such as "first" and "second" in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here.

[0076] Some concepts involved in the embodiments of the present application are introduced below.

[0077] Depth information: It includes the distance between the pixel points in the image and the object to which they belong and the image acquisition device that captures the image. After obtaining the depth information of the image, it is possible to combine the internal parameters of the image acquisition device to obtain data such as three-dimensional information or point cloud corresponding to the image, which can be applied to many fields such as games, rendering, virtual reality (VR), augmented reality (AR), and artificial intelligence generated content (AIGC). In the above fields, a large amount of three-dimensional assets or three-dimensional models are usually required for scene design, interaction, etc. The method in the embodiments of this application can assist in generating three-dimensional assets under low-light conditions based on real scenes, including scenes, characters, objects, etc., effectively supporting the research and development and advancement of the above fields.

[0078] Generative Adversarial Network (GAN): It is a generative model that learns through the mutual game of two neural networks, consisting of a generator G and a discriminator D. Through training with a large amount of sample data, the generation ability of the generator and the discrimination ability of the discriminator are gradually improved in the confrontation. The generator randomly samples from the latent space as input, and its output result needs to imitate the real samples in the training set as much as possible. The input of the discriminator is either real samples or the output of the generator, and its purpose is to distinguish the output of the generator from the real samples as much as possible. The generator and the discriminator confront and learn from each other, and the ultimate goal is to make the discriminator unable to judge whether the output result of the generator is real.

[0079] As Figure 1A shown, it is a schematic diagram of the working process of the generative adversarial network. On the one hand, real samples are selected from the dataset. On the other hand, the generator generates fake samples based on the input feature Z, and the real samples and fake samples are input into the discriminator for discrimination. The parameters of the generator and the discriminator are updated according to the discrimination results of the discriminator.

[0080] In the training process of the target generation model in the embodiments of this application, the idea of the generative adversarial network is used. The initial generation model is regarded as the generator. Through continuous learning, taking sample feature extraction as an example, the sample extraction feature and the reference extraction feature are input into the discriminator, so that the discriminator cannot distinguish the sample extraction feature from the reference extraction feature, ensuring that the features extracted by the initial generation model are close to the reference extraction features.

[0081] Encoder: When the input is a picture, it is used to encode and map the picture to a high-dimensional latent space. In the embodiments of this application, the target generation model can include an encoder for feature extraction to obtain sample extraction features.

[0082] Decoder: During the process of processing images, it is used to decode and reconstruct the features in the high-dimensional latent space to obtain the corresponding reconstructed features. In the embodiments of the present application, the target generation model may include a decoder for feature reconstruction to obtain sample reconstructed features.

[0083] AIGC: Also known as generative AI, it is the combination of artificial intelligence and content creation, enabling machines to automatically generate high-quality and high-efficiency content, and has wide applications in fields such as intelligent hardware and big data analysis.

[0084] Artificial Intelligence (AI): It uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0085] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0086] Machine Learning (ML): It is an interdisciplinary subject in multiple fields, involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The target generation model in the embodiments of the present application is precisely trained based on the idea of machine learning.

[0087] The depth image generation method in the embodiments of the present application can be applied to the field of intelligent transportation. For example, the image of the road traffic scene captured can be input into the target generation model to obtain the depth information within the scene, and then a three-dimensional scene can be reconstructed based on the depth information to assist in realizing driverless driving.

[0088] The design concept of the embodiments of the present application will be briefly introduced as follows:

[0089] In a computer vision system, three-dimensional scene information provides more possibilities for various computer vision applications such as image segmentation, object detection, and object tracking. Correspondingly, depth images, as a common expression of three-dimensional scene information, have also been widely used. The pixel value of each pixel point in a depth image can be used to represent the distance between a point in the scene and the image collector. Through the depth image, the geometric shape of the surface of the object in the scene can be reflected.

[0090] Currently, for images taken under well-lit conditions, corresponding depth images can already be generated relatively accurately through a neural network model. However, for images taken under low-light conditions (hereinafter referred to as night images), due to poor lighting conditions, it is difficult to accurately capture the details of the objects in the images, and the accuracy of the generated depth images is poor.

[0091] In related technologies, in order to improve the generation accuracy of depth images corresponding to night images, mainly some image enhancement means (for example, adjusting image brightness and saturation) are first used to enhance the night images, and then the enhanced images are input into the neural network model to obtain the output depth images.

[0092] However, image enhancement can only help restore the details of the regions with poor brightness difference in night images. There may also be some strong light and reflective regions in night images. The distances between these strong light and reflective regions and the image collector are actually not much different from the distances between the surrounding regions and the image collector, but there will be obvious brightness differences in the night images, which will in turn result in large depth differences in the depth images. However, the above-mentioned image enhancement means cannot eliminate the strong light and reflective regions in night images, nor can they eliminate the brightness differences with the surrounding regions, so they cannot accurately generate depth images corresponding to night images.

[0093] For example, referring to Figure 1B , it is a schematic diagram of a night image and the corresponding depth image. The first row is the night image, and the second row is the depth image. Due to the requirements of the patent document format, both the night image and the depth image are shown in the form of grayscale images. In fact, both the night image and the depth image are color images. Some strong light and reflective regions are framed in the night image. It can be seen that in the corresponding depth image, the colors of these regions are significantly different from those of the surrounding regions, that is, there are errors in the depth image.

[0094] In view of this, the embodiments of the present application provide a depth image generation method, apparatus, electronic device, and storage medium. With the help of a trained target generation model, the present application extracts depth information from a target image to obtain target depth features, and then obtains a depth image of the target image based on the target depth features. Regarding the target generation model used, its main function is to extract the depth features of an image. In order to improve the accuracy of the features extracted by the model, during the process of training the target generation model, the depth features of a sample image are first extracted by an initial generation model. At this time, the extraction ability of the model is limited, and it may also be affected by factors such as lighting conditions, strong light, and reflection, resulting in the inability to accurately capture the depth features in the sample image.

[0095] Based on this, a reference image is introduced in the present application. Since the environmental brightness when shooting the reference image is higher than that when shooting the sample image, as described in the background art, for images taken under good lighting conditions, features can already be accurately extracted and corresponding depth images can be generated through a neural network model. Therefore, in each round of iterative training of the target generation model, the feature differences between the selected sample image and the reference image in multiple feature dimensions are used, combined with the differences between the respective depth images of the sample image and the reference image, to adjust the parameters of the initial generation model, providing more-dimensional comparison bases for the parameter adjustment process and helping to improve the training effect of the model.

[0096] In this way, through the dual constraints in the feature domain and the image domain, the features of the sample image extracted by the generation model are made to be close to the features of the reference image. Furthermore, regardless of the lighting conditions of the image input to the generation model, the extracted depth features can always be close to the actual feature distribution, enabling the target generation model to extract accurate depth features for images under different environmental brightnesses. Further, a more accurate depth image can be obtained based on the depth features output by the target generation model.

[0097] In addition, since during the process of adjusting the model parameters, the differences between the sample image and the reference image in the feature dimension and the depth image dimension are mainly considered, there is no need to perform pixel-level depth information annotation on the sample image, which reduces the time for annotating samples. Moreover, the features and depth images of the reference image can also be directly obtained using neural network models in related technologies, reducing the difficulty of constructing training samples, improving the construction efficiency of training samples, and thus reducing the overall time for model training and improving the model training efficiency.

[0098] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0099] As shown Figure 1C in the figure, it is a schematic diagram of the application scenario of the embodiment of the present application. The application scenario diagram includes two terminal devices 110 and a server 120.

[0100] In the embodiment of the present application, the terminal device includes but is not limited to devices such as mobile phones, tablet computers, laptop computers, desktop computers, e - book readers, intelligent voice interaction devices, intelligent home appliances, vehicle terminals, etc.; a client related to depth image generation can be installed on the terminal device, and the client can be software (such as a browser, 3D reconstruction software, etc.), or a web page, a small program, etc. The server is the background server corresponding to the software, web page, small program, etc., or a server dedicated to depth image generation, and the present application does not make specific limitations. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0101] It should be noted that the depth image generation method in the embodiment of the present application can be executed by an electronic device, and the electronic device can be a server or a terminal device, that is, the method can be executed independently by the server or the terminal device, or jointly executed by the server and the terminal device. For example, when jointly executed by the server and the terminal device, the server implements the training process of the target generation model, deploys the target generation model to the terminal device, and the terminal device inputs the acquired target image into the target generation model for depth information extraction to obtain the target depth feature, and based on the target depth feature, obtains the depth image of the target image.

[0102] In an alternative embodiment, the terminal device and the server can communicate through a communication network.

[0103] In an alternative embodiment, the communication network is a wired network or a wireless network.

[0104] It should be noted that Figure 1C the illustration shown is only an example, and actually the number of terminal devices and servers is not limited, and no specific limitation is made in the embodiment of the present application.

[0105] In the embodiments of the present application, when there are multiple servers, the multiple servers can form a blockchain, and the servers are nodes on the blockchain; for example, in the depth image generation method disclosed in the embodiments of the present application, the reference image, reference extraction features, reference reconstruction features, and depth image of the reference image involved therein can all be stored on the blockchain, etc.

[0106] In addition, the embodiments of the present application can be applied to various scenarios, including not only the depth image generation scenario, but also, without limitation, scenarios such as games, cloud technologies, artificial intelligence, intelligent transportation, and assisted driving. For example, when the embodiments of the present application are applied to the game scenario, the depth image of the target image can be generated based on the method in the present application, and then a three-dimensional scene in the game can be constructed based on the depth image.

[0107] Next, in combination with the above-described application scenarios, the depth image generation method provided by the exemplary embodiment of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.

[0108] Refer to Figure 2 As shown, it is a flowchart of the implementation of a depth image generation method provided by the embodiments of the present application. Taking the server as the execution subject as an example, the specific implementation process of the method includes the following steps S21 - S22:

[0109] S21: The server inputs the obtained target image into the trained target generation model for depth information extraction to obtain target depth features;

[0110] Specifically, the target generation model performs depth information extraction on the target image. The specific operations of depth information extraction can include operations such as feature extraction, feature reconstruction, and feature fusion, which are not specifically limited herein. In the embodiments of the present application, the target generation model has an encoder - decoder structure. Among them, the encoder is used to perform feature extraction on the target image to obtain intermediate depth features, and the decoder performs feature reconstruction on the target depth features to obtain target depth features.

[0111] For example, as Figure 3 shown, it is a schematic structural diagram of a target generation model in the embodiments of the present application. For image 1, the encoder performs 3 rounds of feature extraction on image 1. The input of the first round of feature extraction is image 1, and the input of each subsequent round of feature extraction is the output of the previous round of feature extraction. The sizes of the features output by the 3 rounds of feature extraction decrease from large to small, and the output of the third round of feature extraction is intermediate depth feature 1. The decoder performs 3 rounds of feature reconstruction on intermediate depth feature 1. The input of the first round of feature reconstruction is intermediate depth feature 1, and the input of each subsequent round of feature reconstruction is the output of the previous round of feature reconstruction. The sizes of the features output by the 3 rounds of feature reconstruction increase from small to large.

[0112] S22: The server obtains the depth image of the target image based on the target depth feature.

[0113] Specifically, by performing convolutional processing on the target depth feature, the corresponding depth image can be obtained.

[0114] In each round of iterative training of the target generation model, the feature differences between the selected sample image and the reference image in multiple feature dimensions are used, combined with the differences between the depth images of the sample image and the reference image respectively, to adjust the parameters of the initial generation model.

[0115] Specifically, in each round of iterative training, a sample image and a reference image are selected from the training sample set, and the ambient brightness of the reference image during shooting is higher than that of the sample image during shooting to ensure the reference effect of the reference image in terms of features and depth images. The better the lighting conditions during shooting the reference image, the more accurate the extracted features and depth images will be. Therefore, in the embodiments of the present application, the ambient brightness of the reference image is not specifically limited and can be set according to requirements.

[0116] In addition, the reference image can also be limited by setting the shooting time. For example, according to common sense, the lighting is usually the most sufficient from 12:00 to 14:00 every day, so the shooting time of the reference image can be set from 12:00 to 14:00, and specifically, it can also be flexibly adjusted according to the differences between different regions and different seasons.

[0117] Optionally, the acquisition method of the depth image of the sample image is as follows: the initial generation model is used to extract the depth information of the sample image to obtain the sample depth feature; the sample depth feature is convolved to obtain the depth image of the sample image.

[0118] In each round of iterative training, the initial generation model is mainly used to extract the depth information of the sample image to obtain the features of the sample image in different dimensions, and the depth image of the sample image is obtained based on the features. Therefore, from the dual constraints of the feature dimension and the depth image dimension, the feature extraction ability of the model can be guaranteed, and thus the accuracy of the depth image can be improved.

[0119] In the embodiments of the present application, with the help of the trained target generation model, the depth information of the target image is extracted to obtain the target depth feature, and then the depth image of the target image is obtained based on the target depth feature. As for the used target generation model, its main function is to extract the depth features of the image. In order to improve the accuracy of the features extracted by the model, during the process of training the target generation model, first, the depth features of the sample image are extracted by the initial generation model. At this time, the extraction ability of the model is limited, and it may also be affected by factors such as lighting conditions, strong light, and reflection, resulting in the inability to accurately capture the depth features in the sample image.

[0120] Based on this, a reference image is introduced in this application. Since the environmental brightness when taking the reference image is higher than that when taking the sample image, as described in the background art, for images taken under well-lit conditions, features can already be accurately extracted through a neural network model and corresponding depth images can be generated. Therefore, in each round of iterative training of the target generation model, the feature differences between the selected sample image and the reference image in multiple feature dimensions are used, combined with the differences between the respective depth images of the sample image and the reference image, to adjust the parameters of the initial generation model, providing more dimensions of comparison basis for the parameter adjustment process and helping to improve the training effect of the model.

[0121] In this way, through the dual constraints in the feature domain and the image domain, the features of the sample image extracted by the generation model are close to the features of the reference image. Furthermore, regardless of the lighting conditions of the image input to the generation model, the extracted depth features can always be close to the actual feature distribution, enabling the target generation model to extract accurate depth features for images under different environmental brightnesses. Further, based on the depth features output by the target generation model, a more accurate depth image can be obtained.

[0122] In addition, since during the process of model parameter adjustment, the differences between the sample image and the reference image in the feature dimension and the depth image dimension are mainly considered, there is no need to perform pixel-level depth information annotation on the sample image, which reduces the time for annotating samples. Also, the features and depth images of the reference image can be directly obtained using the neural network model in related technologies, reducing the difficulty of constructing training samples, improving the construction efficiency of training samples, and thus reducing the overall model training time and improving the model training efficiency.

[0123] The following specifically introduces the training process of the target generation model in the embodiments of this application. Optionally, in each round of iterative training, the parameters of the initial generation model are adjusted through the following steps S31 - S34:

[0124] S31: Based on the differences between the multiple sample extraction features of the sample image and the multiple reference extraction features of the reference image, construct a first loss function;

[0125] S32: Based on the differences between the respective depth images of the sample image and the reference image, construct a second loss function;

[0126] S33: Obtain a target loss function based on at least the first loss function and the second loss function;

[0127] S34: Based on the target loss function, adjust the parameters of the initial generation model.

[0128] Specifically, multiple sample extraction features are obtained by performing multiple rounds of feature extraction on the sample image. The input for the first round of feature extraction is the sample image, and the input for each subsequent round of feature extraction is the output of the previous round of feature extraction. In the embodiments of the present application, mainly taking the feature extraction of the sample image by an encoder as an example for illustration, in fact, it can also be a recurrent neural network, a convolutional neural network, a residual network, etc., which are not specifically limited herein. Correspondingly, the reference extraction features are also obtained by performing feature extraction on the reference image.

[0129] In step S31, the sample extraction features and the reference extraction features can be calculated for the loss one by one, and then the calculated losses are weighted and summed to obtain the first loss function, where the corresponding sample extraction features and reference extraction features have the same size.

[0130] For example, the size of sample extraction feature 1 is 256×256, the size of sample extraction feature 2 is 128×128, the size of sample extraction feature 3 is 64×64, the size of reference extraction feature 1 is 256×256, the size of reference extraction feature 2 is 128×128, and the size of reference extraction feature 3 is 64×64.

[0131] The loss 1 is calculated from sample extraction feature 1 and reference extraction feature 1, the loss 2 is calculated from sample extraction feature 2 and reference extraction feature 2, and the loss 3 is calculated from sample extraction feature 3 and reference extraction feature 3. The losses 1, 2, and 3 are weighted and summed to obtain the first loss function. The weights of each loss can be set according to requirements. In the embodiments of the present application, an example is given where the weights of each loss are all 1.

[0132] In step S31, after the sample extraction features are fused, the loss can also be calculated with the features obtained by fusing with the reference extraction features to obtain the first loss function.

[0133] Still taking sample extraction features 1-3 as an example, the sample extraction features 1-3 are fused to obtain the sample fusion feature, the reference extraction features 1-3 are fused to obtain the reference fusion feature, and the difference between the sample fusion feature and the reference fusion feature is calculated to obtain the first loss function.

[0134] It should be noted that the methods of obtaining the first loss function listed above are only for illustrative purposes, and the difference between the sample extraction features and the reference extraction features can also be calculated based on other methods, which are not specifically limited herein.

[0135] In step S32, since the pixel values of the pixels in the depth image can represent depth information, the second loss function can be obtained based on the pixel value differences between the pixels at the same positions in the depth image of the sample image and the depth image of the reference image.

[0136] Taking Figure 4 as an example, it is a schematic diagram of a construction method of a second loss function in an embodiment of the present application. The pixel points at the same position are: pixel 1 and pixel 5, pixel 2 and pixel 6, pixel 3 and pixel 7, pixel 4 and pixel 8. The difference in pixel values between pixel 1 and pixel 5 is a, the difference in pixel values between pixel 2 and pixel 6 is b, the difference in pixel values between pixel 3 and pixel 7 is c, and the difference in pixel values between pixel 4 and pixel 8 is d. The second loss function is obtained by weighted summation of a, b, c, and d.

[0137] The second loss function can also be obtained based on the difference between the pixel value distribution of the depth image of the sample image and the pixel value distribution of the depth image of the reference image, which is not specifically limited here.

[0138] After obtaining the first loss function and the second loss function, the target loss function can be obtained only based on the first loss function and the second loss function. For example, the target loss function is obtained by weighted summation of the first loss function and the second loss function; the target loss function can also be obtained based on the first loss function, the second loss function, and the regularization term, or the target loss function is obtained based on the first loss function, the second loss function, and the loss function in other dimensions.

[0139] Optionally, the third loss function is obtained in the following manner:

[0140] Based on the difference between multiple sample reconstruction features of the sample image and multiple reference reconstruction features of the reference image, a third loss function is constructed.

[0141] Among them, the multiple sample reconstruction features are obtained by performing multiple rounds of feature reconstruction on the sample extraction features obtained by the last round of feature extraction. The input of the first round of feature reconstruction is the sample extraction features obtained by the last round of feature extraction, and the input of each subsequent round of feature reconstruction is the output of the previous round of feature reconstruction. In the embodiment of the present application, it is mainly illustrated by taking the decoder to perform feature reconstruction on the sample extraction features obtained by the last round of feature extraction as an example. In fact, it can also be a recurrent neural network, a convolutional neural network, a residual network, etc., which is not specifically limited here. Correspondingly, the reference reconstruction features are also obtained by performing feature extraction through feature reconstruction on the reference extraction features obtained from the last round of features.

[0142] When calculating the third loss function, the sample reconstruction features and the reference reconstruction features can be calculated for loss one by one, and then the calculated losses are weighted and summed to obtain the third loss function, where the corresponding sample reconstruction features and reference reconstruction features have the same size.

[0143] For example, the size of sample reconstruction feature 1 is 256×256, the size of sample reconstruction feature 2 is 128×128, the size of sample reconstruction feature 3 is 64×64, the size of reference reconstruction feature 1 is 256×256, the size of reference reconstruction feature 2 is 128×128, and the size of reference reconstruction feature 3 is 64×64.

[0144] The loss 1 is calculated from sample reconstruction feature 1 and reference reconstruction feature 1, and the loss 2 is calculated from sample reconstruction feature 2 and reference reconstruction feature 2. The weighted sum of loss 1 and loss 2 is used to obtain the third loss function. The weights of each loss can be set according to requirements. In the embodiments of the present application, it is exemplified that the weights of each loss are all 1.

[0145] When calculating the third loss function, the sample reconstruction features can also be feature-fused first, and then the loss is calculated from the features obtained by fusing with the reference reconstruction features to obtain the third loss function.

[0146] Still taking sample reconstruction features 1-2 as an example, the sample reconstruction features 1-2 are feature-fused to obtain sample reconstruction fused features, and the reference reconstruction features 1-2 are feature-fused to obtain reference reconstruction fused features. The difference between the sample reconstruction fused features and the reference reconstruction fused features is calculated to obtain the third loss function.

[0147] It should be noted that the methods for obtaining the third loss function listed above are only for illustration, and the difference between the sample reconstruction features and the reference reconstruction features can also be calculated based on other methods, which are not specifically limited herein.

[0148] Furthermore, in step S33, the target loss function is obtained based on at least the first loss function, the second loss function, and the third loss function.

[0149] The target loss function can be obtained only based on the first loss function, the second loss function, and the third loss function. For example, the weighted sum of the first loss function, the second loss function, and the third loss function is used to obtain the target loss function; the target loss function can also be obtained based on the first loss function, the second loss function, and the third loss function, as well as a regularization term, or the target loss function can be obtained based on the first loss function, the second loss function, and the third loss function, as well as loss functions in other dimensions.

[0150] Furthermore, the fourth loss function can be constructed in the following way:

[0151] The sample depth images of multiple sample reconstruction features are respectively cropped to obtain respective target region images. After sorting the target region images according to the sizes of the corresponding sample reconstruction features, the fourth loss function is constructed based on the pixel differences between every two adjacent target region images.

[0152] Among them, the sample depth image is obtained by convolving the corresponding sample reconstruction feature. For each sample reconstruction feature, a sample depth image of a corresponding size can be obtained through convolution. For each sample depth image, a small rectangular region, that is, the target region image, can be cropped along the diagonal. Each target region image contains the pixel points on the diagonal of the corresponding sample depth image.

[0153] As Figure 5A shown, it is a schematic diagram of a target region image in an embodiment of the present application. For the sample depth images 1-3, the target region images 1-3 are respectively cropped along the diagonals L1-L4. The positions of the target region images 1-3 corresponding to the sample depth images are the same.

[0154] Figure 5A In the following, taking the example of cropping one target region image from each sample depth image for illustration, actually, multiple target region images can also be cropped from one sample depth image. As Figure 5B shown, it is another schematic diagram of a target region image in an embodiment of the present application. Four target region images are cropped from each sample depth image.

[0155] The target region images at the same positions of different sample depth images actually correspond to the objects in the same region of the sample image. Therefore, there should be consistency among these target region images. The sample depth images can be arranged in descending order or ascending order according to the sizes of the corresponding sample reconstruction features. After the arrangement, based on the pixel differences between adjacent target region images, a fourth loss function is constructed.

[0156] For example, the arranged order is: target region image 1 → target region image 2 → target region image 3 → target region image 4. Calculate the pixel difference 1 between target region image 1 and target region image 2, the pixel difference 2 between target region image 2 and target region image 3, and the pixel difference 3 between target region image 3 and target region image 4. By performing weighted summation on the pixel differences 1, 2, and 3, the fourth loss function is obtained.

[0157] Furthermore, after obtaining the fourth loss function, the target loss function can be obtained based on the first loss function, the second loss function, the third loss function, and the fourth loss function, or the target loss function can also be obtained based on the first loss function, the second loss function, and the fourth loss function.

[0158] Due to the size differences between the sample depth images, in order to obtain the pixel differences between adjacent target regions, the sample depth images also need to be processed. Optionally, for two adjacent target region images, the following operations are performed:

[0159] Upsample the target region image with the smaller size among the two target region images to obtain a target sampled image, where the size of the target sampled image is the same as that of the target region image with the larger size among the two target region images; based on the difference in pixel values between the target sampled image and the pixel points at the same positions included in the target region image with the larger size, obtain a pixel loss function.

[0160] For example, if the sizes of the two target region images are 256×256 and 128×128 respectively, then upsample the 128×128 target region image to 256×256, then calculate the difference in pixel values between the pixel points at the same positions respectively, and sum up the calculated differences in pixel values to obtain a pixel loss function. Finally, based on the obtained pixel loss functions, construct a fourth loss function.

[0161] By applying multi-scale consistency constraints, obtain better depth information, thereby achieving a better depth information estimation effect.

[0162] As Figure 6 shown, it is a schematic diagram of the construction process of the fourth loss function in the embodiment of this application. The size of target region image 1 is 64×64, the size of target region image 2 is 128×128, and the size of target region image 3 is 256×256. Upsample target region image 1 to 128×128, and calculate pixel loss function 1 with target region image 2. Upsample target region image 2 to 256×256, and calculate pixel loss function 2 with target region image 3. Sum up pixel loss function 1 and pixel loss function 2 to obtain the fourth loss function.

[0163] Taking the reconstruction feature dn = [dn1,..., dnn] of n samples as an example, each sample reconstruction feature dni can obtain a corresponding-sized sample depth image Si through convolution. For Si and Si-1, along Figure 5A from l1 to l4 and l2 to l3 in, small rectangular regions can be cropped out. There is consistency in the corresponding rectangular regions in Si and Si-1, that is:

[0164] S_crop_i-1 = Crop(Si-1)

[0165] S_crop_i = Crop(Si)

[0166] Loss_crop = ||up_sample(S_crop_i-1) - S_crop_i||1

[0167] where Crop represents the cropping operation along Figure 5A from L1 to L4 and L2 to L3 in, up_sample represents the upsampling operation, and Loss_crop represents the pixel loss function.

[0168] Next, the specific construction processes of the first loss function, the second loss function, and the third loss function in the embodiments of the present application will be introduced.

[0169] First, the construction method of the first loss function is as follows:

[0170] For multiple samples, extract features and perform the following operations respectively: input a sample-extracted feature and the corresponding reference-extracted feature into a discriminator to obtain the corresponding feature discrimination result; wherein, the size of a sample-extracted feature is the same as that of the corresponding reference-extracted feature; based on each feature discrimination result, construct the first loss function.

[0171] Specifically, during the training process, the training objective is that the sample-extracted feature is close to the reference-extracted feature. Therefore, in the embodiments of the present application, the discriminator is used to distinguish between the sample-extracted feature and the reference-extracted feature with the same size, and it is expected that through iterative training, the discriminator cannot distinguish between the sample-extracted feature and the reference-extracted feature.

[0172] Based on the above method, through the discrimination of the sample-extracted feature and the reference-extracted feature by the discriminator, the features extracted by the initial generation model are getting closer and closer to the reference-extracted feature, realizing the alignment at the feature end, enabling the model to focus on the high-frequency part of the extracted features, improving the feature extraction ability of the initial generation model, and further improving the accuracy of the depth image.

[0173] Optionally, the sample-extracted feature and the reference-extracted feature can also be processed before being input into the discriminator:

[0174] Fuse a sample-extracted feature with a preset template feature to obtain a sample-fused feature; fuse the reference-extracted feature with the preset template feature to obtain a reference-fused feature; input the sample-fused feature and the reference-fused feature into the discriminator.

[0175] As described above, the depth information of a pixel point is used to represent the distance between the corresponding pixel point and the image acquisition device of the sample image. Then, through the preset template feature, the reference distribution of the depth information of each pixel point included in the sample image can be represented to constrain the extracted sample-extracted feature to be more in line with the visual perception law.

[0176] According to the distribution law of the acquired image, it can be simplified to obtain Figure 7A , that is, usually the upper part of the image is the sky, the lower part is the ground, and the left and right sides are buildings, etc. Therefore, the image area can be divided into 4 parts, as shown in Figure 7B the regions 1, 2, 3, and 4 included in

[0177] Based on this, set the preset template feature, such as Figure 7CAs shown in the figure, it is a schematic diagram of the preset template features in the embodiment of the present application. In the preset template features, the template values of regions 1-3 are set to represent a gradual change from 1 to 0. This is because in regions 1, 2, and 3, the depth information of the pixel points generally gradually increases from the outside to the center region of the image (that is, the distance between the objects in regions 1-3 and the image collector gradually becomes farther). In region 4, the depth information of the pixel points generally gradually decreases from the outside to the center region of the image. Therefore, in the preset template features, the template value of region 4 is set to gradually increase from 0.2 to 0.4. The preset template features set in the above manner can be more in line with the visual perception in reality, that is, in regions 1-3, the distance between the object located in the center of the region and the image collector is farther, and in region 4, the distance between the object located in the center of the region and the image collector is closer.

[0178] Through the prior template design that conforms to visual perception, while making the extracted features close to the reference extraction features, it is more in line with the visual perception law, further improving the accuracy of the extracted features, and then improving the accuracy of the depth image. It can effectively handle situations such as strong light and jumps in night scenes, achieving a better effect.

[0179] There are many ways of feature fusion. For example, feature stitching, merging (concat), superposition (add), etc. In the embodiment of the present application, the way of feature fusion is concat, that is, to increase the number of channels of the sample extraction features and the reference extraction features. Taking the sample extraction features as 64×256×256 and the preset template features as 1×256×256 as an example, the sample fusion features after feature fusion are 65×256×256.

[0180] As Figure 8 shown in the figure, it is a schematic diagram of a feature discrimination method in the embodiment of the present application. After performing concat operations on the sample extraction features and the reference extraction features with the preset template features respectively, sample fusion features and reference fusion features are obtained. The sample fusion features and the reference fusion features are input into a discriminator to obtain a feature discrimination result.

[0181] Based on the idea of the generative adversarial network, both the generative network and the adversarial network need to continuously learn. Therefore, it is also necessary to adjust the parameters of the discriminator based on each feature discrimination result to enhance the discrimination ability of the discriminator.

[0182] The construction process of the first loss function (which can also be called GAN loss) will be illustrated by an example. For the sample extraction feature ed and the reference extraction feature en, for each feature edi and eni of the same size among them, a concat operation is respectively performed with the prior template Mp (preset template feature) to obtain new features edmi (sample fusion feature) and enmi (reference fusion feature). Taking edi and eni as examples, through the concate operation, edi, eni and the prior template Mp are combined into features edmi and enmi, that is:

[0183] edmi = concat(edi, Mp)

[0184] enmi = concat(eni, Mp)

[0185] Then, edmi and enmi are input into the discriminator, making the discriminator unable to distinguish edmi from enmi, thereby ensuring that the information of eni and edi is close and guaranteeing the consistency of information.

[0186] By feature alignment based on the prior template, the feature extraction ability of the model is improved.

[0187] The construction method of the third loss function is as follows:

[0188] For the reconstructed features of multiple samples, the following operations are respectively performed: input a sample reconstructed feature and the corresponding reference reconstructed feature into the discriminator to obtain the corresponding reconstructed feature discrimination result; among them, a sample reconstructed feature and the corresponding reference reconstructed feature have the same size; based on each reconstructed feature discrimination result, the third loss function is constructed.

[0189] Specifically, during the training process, the training objective is that the sample reconstructed feature is close to the reference reconstructed feature. Therefore, in the embodiments of the present application, the discriminator is used to distinguish the sample reconstructed feature and the reference reconstructed feature of the same size, and it is expected that through iterative training, the discriminator cannot distinguish the sample reconstructed feature and the reference reconstructed feature.

[0190] Based on the above method, through the discrimination of the sample reconstructed feature and the reference reconstructed feature by the discriminator, the features reconstructed by the initial generation model are getting closer and closer to the reference reconstructed feature, realizing the alignment at the feature end, making the model focus on the high-frequency part of the reconstructed features, improving the feature reconstruction ability of the initial generation model, and further improving the accuracy of the depth image.

[0191] Optionally, the sample reconstructed feature and the reference reconstructed feature can also be processed and then input into the discriminator:

[0192] Fuse a sample reconstruction feature with a preset template feature to obtain a sample fusion feature; fuse a reference reconstruction feature with a preset template feature to obtain a reference fusion feature; input the sample fusion feature and the reference fusion feature into a discriminator.

[0193] Specifically, the way of feature fusion can be feature splicing, concat, add, etc., which is not specifically limited here. In the embodiments of this application, the way of feature fusion is concat, that is, to increase the number of channels of the sample reconstruction feature and the reference reconstruction feature. Taking the sample reconstruction feature as 128×256×256 and the preset template feature as 1×256×256 as an example, the sample fusion feature after feature fusion is 129×256×256.

[0194] Based on the idea of the generative adversarial network, both the generative network and the adversarial network need to continuously learn. Therefore, it is also necessary to adjust the parameters of the discriminator based on the discrimination results of each reconstruction feature to enhance the discrimination ability of the discriminator.

[0195] Through the design of a prior template that conforms to visual perception, while making the reconstructed feature close to the reference reconstructed feature, it is more in line with the laws of visual perception, further improving the accuracy of the reconstructed feature, and then improving the accuracy of the depth image. It can effectively handle strong light, jumps, etc. in night scenes and achieve better results.

[0196] The construction method of the second loss function is as follows:

[0197] Input the depth image of the sample image (sample depth image) and the depth image of the reference image (reference depth image) into the discriminator to obtain the corresponding image discrimination results; construct the second loss function based on the image discrimination results.

[0198] Specifically, during the training process, the training objective is for the sample depth image to be close to the reference depth image. Therefore, in the embodiments of this application, the discriminator is used to distinguish the sample depth image and the reference depth image, and it is expected that through iterative training, the discriminator cannot distinguish the sample depth image and the reference depth image.

[0199] Based on the above method, through the discrimination of the sample depth image and the reference depth image by the discriminator, the alignment at the image end is realized, enabling the model to focus on the high-frequency part of the image, improving the depth information extraction ability of the initial generation model, and then improving the accuracy of the depth image.

[0200] Optionally, the sample depth image and the reference depth image can also be processed before being input into the discriminator:

[0201] Fuse the features of the sample depth image and the preset template features to obtain a sample fused image; fuse the features of the reference depth image and the preset template features to obtain a reference fused image; input the sample fused image and the reference fused image into a discriminator.

[0202] Specifically, the way of feature fusion can be feature splicing, concat, add, etc., which is not specifically limited here. In the embodiments of this application, the way of feature fusion is concat, that is, to increase the number of channels of the sample depth image and the reference depth image. Taking the sample depth image as 1×256×256 and the preset template features as 1×256×256 as an example, the sample fused image after feature fusion is 2×256×256.

[0203] Based on the idea of the generative adversarial network, both the generative network and the adversarial network need to continuously learn. Therefore, it is also necessary to adjust the parameters of the discriminator based on the discrimination results of each image to enhance the discrimination ability of the discriminator.

[0204] Through the design of a prior template that conforms to visual perception, while making the sample depth image close to the reference depth image, it more conforms to the laws of visual perception, further improving the accuracy of the depth image, and can effectively handle strong light, jumps, etc. in night scenes, achieving a better effect.

[0205] It should be noted that the discriminators used to distinguish sample extraction features and reference extraction features, sample reconstruction features and reference reconstruction features, and sample depth images and reference depth images are three different discriminators, which are respectively used to achieve discrimination in different dimensions.

[0206] In the embodiments of this application, the reference extraction features, reference reconstruction features, and reference depth images of the reference image can be directly obtained by using a trained reference generation model to extract features and reconstruct features of the reference image, or a reference generation model can be obtained by pre-training the basic generation model, and self-supervised loss or supervised loss is used during the training process. Among them, the basic generation model has the same structure as the initial generation model, which is an encoder-decoder structure.

[0207] Input the reference image into the reference generation model. The encoder extracts features from the reference image to obtain ed = [ed1,..., edn], a total of n features, with the size from large to small. The decoder reconstructs features from edn to obtain dd = [dd1,..., ddn], a total of n features, with the size from small to large. At the same time, the last feature ddn of the decoder can obtain the final depth map d_day) (reference depth image) through convolution. The reference image can also be called a daytime image, and the corresponding reference generation model is called a daytime depth information estimation network.

[0208] Such as Figure 9As shown in the figure, it is a schematic diagram of the execution process of one round of iterative training in an embodiment of the present application. The sample image is input into the initial generation model, and through the encoder, multi-round feature extraction is performed to obtain Feature 1, Feature 2, and Feature 3. Through the decoder, multi-round feature reconstruction is performed on Feature 3 to obtain Feature 4, Feature 5, and Feature 6. Feature 6 is convolved to obtain the sample depth image. The discriminator 1 is used to distinguish Feature 1 from Feature 11, Feature 2 from Feature 21, and Feature 3 from Feature 31 respectively, to obtain Discrimination Result 1, Discrimination Result 2, and Discrimination Result 3. Loss Function 1 is constructed using Discrimination Results 1-3. The discriminator 2 is used to distinguish Feature 4 from Feature 41, Feature 5 from Feature 51, and Feature 6 from Feature 61 respectively, to obtain Discrimination Result 4, Discrimination Result 5, and Discrimination Result 6. Loss Function 2 is constructed using Discrimination Results 4-6. The discriminator 3 is used to distinguish the sample depth image from the reference depth image to obtain Discrimination Result 6. Loss Function 3 is constructed based on Discrimination Result 6. The sum of Loss Function 1, Loss Function 2, and Loss Function 3 is obtained as the target loss function, which is used to adjust the parameters of the initial generation model.

[0209] Based on the same inventive concept, an embodiment of the present application also provides a depth image generation device. As Figure 10 shown, it is a schematic structural diagram of the depth image generation device 1000, which may include:

[0210] An input unit 1001, configured to input the acquired target image into the trained target generation model for depth information extraction to obtain target depth features;

[0211] A processing unit 1002, configured to obtain the depth image of the target image based on the target depth features;

[0212] Wherein, in each round of iterative training of the target generation model, the feature differences of the selected sample image and reference image in multiple feature dimensions are used, combined with the differences between the depth images of the sample image and the reference image respectively, to adjust the parameters of the initial generation model; the environmental brightness of the reference image taken is higher than the environmental brightness of the sample image taken.

[0213] In an embodiment of the present application, with the help of the trained target generation model, depth information of the target image is extracted to obtain target depth features, and then the depth image of the target image is obtained based on the target depth features. Regarding the target generation model used, its main function is to extract the depth features of the image. In order to improve the accuracy of the features extracted by the model, during the process of training the target generation model, first, the depth features of the sample image are extracted by the initial generation model. At this time, the extraction ability of the model is limited, and it may also be affected by factors such as lighting conditions, strong light, and reflection, resulting in the inability to accurately capture the depth features in the sample image.

[0214] Based on this, a reference image is introduced in this application. Since the environmental brightness when capturing the reference image is higher than that when capturing the sample image, as described in the background art, for images captured under well-lit conditions, features can already be accurately extracted through a neural network model and corresponding depth images can be generated. Therefore, in each round of iterative training of the target generation model, the feature differences between the selected sample image and the reference image in multiple feature dimensions are used, combined with the differences between the depth images of the sample image and the reference image respectively, to adjust the parameters of the initial generation model, providing more dimensions of comparison basis for the parameter adjustment process and helping to improve the training effect of the model.

[0215] In this way, through the dual constraints in the feature domain and the image domain, the features of the sample image extracted by the generation model are made to be close to the features of the reference image. Furthermore, regardless of the lighting conditions of the image input to the generation model, the extracted depth features can always be close to the actual feature distribution, enabling the target generation model to extract accurate depth features for images under different environmental brightnesses. Further, based on the depth features output by the target generation model, a more accurate depth image can be obtained.

[0216] In addition, since in the process of model parameter adjustment, it is mainly to consider the differences between the sample image and the reference image in the feature dimension and the depth image dimension, there is no need to perform pixel-level depth information annotation on the sample image, which reduces the time for annotating samples. Also, the features and depth images of the reference image can be directly obtained using the neural network model in related technologies, reducing the difficulty of constructing training samples, improving the construction efficiency of training samples, and thus reducing the overall model training time and improving the model training efficiency.

[0217] Optionally, the device includes a training unit 1003, which is used to adjust the parameters of the initial generation model through the following steps in each round of iterative training:

[0218] Based on the differences between multiple sample extraction features of the sample image and multiple reference extraction features of the reference image, construct a first loss function, where the multiple sample extraction features are obtained by performing multiple rounds of feature extraction on the sample image;

[0219] Based on the differences between the depth images of the sample image and the reference image respectively, construct a second loss function;

[0220] Obtain a target loss function based on at least the first loss function and the second loss function;

[0221] Based on the target loss function, adjust the parameters of the initial generation model.

[0222] Optionally, the training unit 1003 is specifically used for:

[0223] Extract features for multiple samples and perform the following operations respectively: Input a sample-extracted feature and its corresponding reference-extracted feature into a discriminator to obtain corresponding feature discrimination results; among them, the size of a sample-extracted feature is the same as that of its corresponding reference-extracted feature.

[0224] Construct a first loss function based on each feature discrimination result.

[0225] Optionally, the training unit 1003 is further configured to:

[0226] Perform feature fusion on a sample-extracted feature and a preset template feature to obtain a sample fusion feature. The preset template feature is used to represent: the reference distribution of the respective depth information of each pixel point included in the sample image, and the depth information is used to represent: the distance between the object to which the corresponding pixel point belongs in the sample image and the image acquisition device that captures the sample image;

[0227] Perform feature fusion on the reference-extracted feature and the preset template feature to obtain a reference fusion feature;

[0228] Then inputting a sample-extracted feature and its corresponding reference-extracted feature into the discriminator includes:

[0229] Input the sample fusion feature and the reference fusion feature into the discriminator.

[0230] Optionally, the training unit 1003 is further configured to:

[0231] Adjust the parameters of the discriminator based on each feature discrimination result.

[0232] Optionally, the training unit 1003 is specifically configured to:

[0233] Construct a third loss function based on the difference between multiple sample reconstruction features of the sample image and multiple reference reconstruction features of the reference image. The multiple sample reconstruction features are obtained by performing multiple rounds of feature reconstruction on the sample-extracted features obtained from the last round of feature extraction;

[0234] Obtain a target loss function based on at least the first loss function, the second loss function, and the third loss function.

[0235] Optionally, the input of the first round of feature extraction is the sample image, and the input of each subsequent round of feature extraction is the output of the previous round of feature extraction;

[0236] The input of the first round of feature reconstruction is the sample-extracted feature obtained from the last round of feature extraction, and the input of each subsequent round of feature reconstruction is the output of the previous round of feature reconstruction.

[0237] Optionally, the training unit 1003 is specifically configured to:

[0238] Crop the sample depth images of the respective reconstructed features of multiple samples to obtain respective target region images, where the sample depth images are obtained by convolving the corresponding sample reconstructed features, and each target region image includes: pixel points on the diagonal of the corresponding sample depth image;

[0239] After sorting the target region images according to the sizes of the corresponding sample reconstructed features, construct a fourth loss function based on the pixel differences between every two adjacent target region images;

[0240] Based on the first loss function, the second loss function, the third loss function, and the fourth loss function, obtain the target loss function.

[0241] Optionally, the training unit 1003 is specifically configured to:

[0242] For every two adjacent target region images, respectively perform the following operations:

[0243] Upsample the target region image with the smaller size among the two target region images to obtain a target sampled image, and the target sampled image has the same size as the target region image with the larger size among the two target region images;

[0244] Based on the difference in pixel values between the same-position pixel points included in the target sampled image and the target region image with the larger size, obtain a pixel loss function;

[0245] Based on the obtained pixel loss functions, construct a fourth loss function.

[0246] Optionally, the processing unit 1002 is specifically configured to obtain the depth image of the sample image in the following manner:

[0247] Extract depth information from the sample image to obtain sample depth features;

[0248] Based on the sample depth features, perform convolution to obtain the depth image of the sample image.

[0249] For the convenience of description, the above parts are divided into respective modules (or units) according to functions and described separately. Of course, when implementing this application, the functions of the respective modules (or units) can be implemented in the same or multiple software or hardware.

[0250] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0251] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0252] Based on the same inventive concept as the above method embodiments, an electronic device is also provided in the embodiments of the present application. In one embodiment, the electronic device can be a server, such as Figure 1C the server shown. In this embodiment, the structure of the electronic device can be as Figure 11 shown, including a memory 1101, a communication module 1103, and one or more processors 1102.

[0253] The memory 1101 is used to store the computer program executed by the processor 1102. The memory 1101 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, programs required to run the instant messaging function, etc.; the data storage area can store various instant messaging information and operation instruction sets, etc.

[0254] The memory 1101 can be a volatile memory, such as a random-access memory (RAM); the memory 1101 can also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1101 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1101 can be a combination of the above memories.

[0255] The processor 1102 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1102 is used to implement the above-mentioned depth image generation method when calling the computer program stored in the memory 1101.

[0256] The communication module 1103 is used to communicate with the terminal device and other servers.

[0257] In the embodiments of the present application, the specific connection medium between the above-mentioned memory 1101, communication module 1103, and processor 1102 is not limited. In the embodiments of the present application Figure 11 it is described that the memory 1101 and the processor 1102 are connected through the bus 1104, and the bus 1104 is described in thick lines in Figure 11 The connection manners between other components are only for illustrative purposes and are not limited thereto. The bus 1104 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of description, Figure 11 it is only described by a thick line in

[0258] The memory 1101 stores a computer storage medium, and the computer storage medium stores computer-executable instructions for implementing the depth image generation method of the embodiments of the present application. The processor 1102 is used to execute the above-mentioned depth image generation method, as Figure 2 shown.

[0259] In another embodiment, the electronic device may also be other electronic devices, such as Figure 1C the terminal device shown. In this embodiment, the structure of the electronic device may be as Figure 12 shown, including components such as a communication component 1210, a memory 1220, a display unit 1230, a camera 1240, a sensor 1250, an audio circuit 1260, a Bluetooth module 1270, and a processor 1280.

[0260] The communication component 1210 is used to communicate with the server. In some embodiments, it may include a Wireless Fidelity (WiFi) module. The WiFi module belongs to short-range wireless transmission technology, and the electronic device can help users send and receive information through the WiFi module.

[0261] The memory 1220 can be used to store software programs and data. The processor 1280 executes various functions and data processing of the terminal device by running the software programs or data stored in the memory 1220. The memory 1220 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. The memory 1220 stores an operating system that enables the terminal device to operate. In this application, the memory 1220 can store the operating system and various application programs, and can also store a computer program for executing the depth image generation method according to the embodiments of this application.

[0262] The display unit 1230 can also be used to display information input by the user or information provided to the user, as well as the graphical user interface (GUI) of various menus of the terminal device. Specifically, the display unit 1230 may include a display screen 1232 disposed on the front of the terminal device. Among them, the display screen 1232 can be configured in the form of a liquid crystal display, a light-emitting diode, etc. The display unit 1230 can be used to display the depth image generation user interface and the like in the embodiments of this application.

[0263] The display unit 1230 can also be used to receive input digital or character information, and generate signal inputs related to the user settings and function control of the terminal device. Specifically, the display unit 1230 may include a touch screen 1231 disposed on the front of the terminal device, which can collect touch operations of the user on or near it, such as clicking buttons, dragging scroll boxes, etc.

[0264] Among them, the touch screen 1231 can cover the display screen 1232, or the touch screen 1231 and the display screen 1232 can be integrated to implement the input and output functions of the terminal device. After integration, it can be simply called a touch display screen. In this application, the display unit 1230 can display application programs and corresponding operation steps.

[0265] The camera 1240 can be used to capture static images, and the user can post comments on the images captured by the camera 1240 through an application. The camera 1240 can be one or more. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 1280 to be converted into a digital image signal.

[0266] The terminal device may further include at least one sensor 1250, such as an acceleration sensor 1251, a distance sensor 1252, a fingerprint sensor 1253, and a temperature sensor 1254. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.

[0267] The audio circuit 1260, the speaker 1261, and the microphone 1262 can provide an audio interface between the user and the terminal device. The audio circuit 1260 can transmit the electrical signal converted from the received audio data to the speaker 1261, and the speaker 1261 converts it into a sound signal for output. The terminal device may also be configured with volume buttons for adjusting the volume of the sound signal. On the other hand, the microphone 1262 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1260, converted into audio data, and then the audio data is output to the communication component 1210 to be sent to, for example, another terminal device, or the audio data is output to the memory 1220 for further processing.

[0268] The Bluetooth module 1270 is used to interact with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 1270 to perform data interaction.

[0269] The processor 1280 is the control center of the terminal device, connecting various parts of the entire terminal using various interfaces and lines. By running or executing the software programs stored in the memory 1220 and calling the data stored in the memory 1220, it executes various functions of the terminal device and processes data. In some embodiments, the processor 1280 may include one or more processing units; the processor 1280 may also integrate an application processor and a baseband processor, where the application processor mainly processes the operating system, the user interface, and application programs, etc., and the baseband processor mainly processes wireless communication. It can be understood that the above baseband processor may not be integrated into the processor 1280. In this application, the processor 1280 can run the operating system, application programs, user interface display, and touch response, as well as the depth image generation method of the embodiments of this application. In addition, the processor 1280 is coupled to the display unit 1230.

[0270] In some possible implementation manners, various aspects of the depth image generation method provided in this application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the depth image generation method according to various exemplary embodiments of this application described above in this specification. For example, the electronic device can execute the steps as Figure 2 shown in.

[0271] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0272] The program product of the embodiments of the present application may adopt a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on an electronic device. However, the program product of the present application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with a command execution system, apparatus, or device.

[0273] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.

[0274] The computer program contained on the readable medium may be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0275] The computer program for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The computer program can be executed entirely on the user's electronic device, partially on the user's device, executed as a stand-alone software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In cases involving a remote electronic device, the remote electronic device can be connected to the user's electronic device through any type of network including a local area network (LAN) or a wide area network (WAN), or, it can be connected to an external electronic device (e.g., by connecting through the Internet using an Internet service provider).

[0276] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described units can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0277] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0278] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable computer programs.

[0279] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.

[0280] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.

[0281] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.

[0282] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0283] Obviously, those skilled in the art can make various changes and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A method for generating a depth image, characterized in that, The method includes: Inputting the obtained target image into a trained target generation model for depth information extraction to obtain target depth features; Obtaining a depth image of the target image based on the target depth features; Wherein, in each round of iterative training of the target generation model, the feature differences between the selected sample image and the reference image in multiple feature dimensions are used, combined with the differences between the depth images of the sample image and the reference image respectively, to adjust the parameters of the initial generation model; the environmental brightness when shooting the reference image is higher than the environmental brightness when shooting the sample image.

2. The method according to claim 1, wherein In each round of iterative training, the parameters of the initial generation model are adjusted through the following steps: Based on the differences between multiple sample extraction features of the sample image and multiple reference extraction features of the reference image, constructing a first loss function, where the multiple sample extraction features are obtained by performing multiple rounds of feature extraction on the sample image; Based on the differences between the depth images of the sample image and the reference image respectively, constructing a second loss function; Obtaining a target loss function based on at least the first loss function and the second loss function; Adjusting the parameters of the initial generation model based on the target loss function.

3. The method according to claim 2, wherein The constructing of the first loss function based on the differences between multiple sample extraction features of the sample image and multiple reference extraction features of the reference image includes: For the multiple sample extraction features, the following operations are respectively performed: inputting a sample extraction feature and a corresponding reference extraction feature into a discriminator to obtain corresponding feature discrimination results; wherein, the size of the one sample extraction feature is the same as that of the corresponding reference extraction feature; Constructing the first loss function based on each feature discrimination result.

4. The method according to claim 3, characterized in that, Before inputting a sample extraction feature and a corresponding reference extraction feature into the discriminator, it further includes: Performing feature fusion on the one sample extraction feature and a preset template feature to obtain a sample fusion feature, where the preset template feature is used to represent: the reference distribution of the depth information of each pixel point included in the sample image, and the depth information is used to represent: the distance between the object to which the corresponding pixel point belongs in the sample image and the image collector for shooting the sample image; Performing feature fusion on the reference extraction feature and the preset template feature to obtain a reference fusion feature; Then the inputting of a sample extraction feature and a corresponding reference extraction feature into the discriminator includes: Inputting the sample fusion feature and the reference fusion feature into the discriminator.

5. The method according to claim 3, characterized in that, The method further includes: Adjusting the parameters of the discriminator based on each feature discrimination result.

6. The method according to claim 2, characterized in that The obtaining of the target loss function based on at least the first loss function and the second loss function includes: Based on the differences between multiple sample reconstruction features of the sample image and multiple reference reconstruction features of the reference image, constructing a third loss function, where the multiple sample reconstruction features are obtained by performing multiple rounds of feature reconstruction on the sample extraction features obtained from the last round of feature extraction; Obtain the target loss function based at least on the first loss function, the second loss function, and the third loss function.

7. The method according to claim 6, wherein The input of the first-round feature extraction is the sample image, and the input of each subsequent round of feature extraction is the output of the previous round of feature extraction; The input of the first-round feature reconstruction is the sample extraction features obtained from the last-round feature extraction, and the input of each subsequent round of feature reconstruction is the output of the previous round of feature reconstruction.

8. The method according to claim 6, wherein The obtaining the target loss function based at least on the first loss function, the second loss function, and the third loss function includes: Crop the respective sample depth images of the multiple sample reconstruction features to obtain respective target region images, where the sample depth images are obtained by convolving the corresponding sample reconstruction features, and each target region image includes: the pixel points on the diagonal of the corresponding sample depth image; After sorting the target region images according to the size of the corresponding sample reconstruction features, construct a fourth loss function based on the pixel differences between every two adjacent target region images; Obtain the target loss function based on the first loss function, the second loss function, the third loss function, and the fourth loss function.

9. The method according to claim 8, wherein The constructing the fourth loss function based on the pixel differences between every two adjacent target region images includes: For every two adjacent target region images, perform the following operations respectively: Upsample the target region image with the smaller size among the two target region images to obtain a target sampling image, where the target sampling image has the same size as the target region image with the larger size among the two target region images; Obtain a pixel loss function based on the difference in pixel values between the same-position pixel points included in the target sampling image and the target region image with the larger size; Construct the fourth loss function based on the obtained pixel loss functions.

10. The method according to any one of claims 1 to 9, characterized in that, Obtain the depth image of the sample image in the following manner: Extract depth information from the sample image to obtain sample depth features; Perform convolution based on the sample depth features to obtain the depth image of the sample image.

11. A depth image generation device, characterized in that, It includes: An input unit, configured to input the acquired target image into a trained target generation model for depth information extraction to obtain target depth features; A processing unit, configured to obtain the depth image of the target image based on the target depth features; Wherein, in each round of iterative training of the target generation model, the feature differences of the selected sample image and the reference image in multiple feature dimensions are used, combined with the differences between the respective depth images of the sample image and the reference image, to adjust the parameters of the initial generation model; the environmental brightness when shooting the reference image is higher than the environmental brightness when shooting the sample image.

12. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, It includes a computer program which, when running on an electronic device, is used to cause the electronic device to execute the steps of any one of claims 1 to 10.

14. A computer program product, characterized in that, It includes a computer program which is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any one of claims 1 to 10.