Training method of light main direction model and electronic device
Patent Information
- Application Number
- CN202410874326.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-01
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2044-07-01
AI Technical Summary
[0005]本申请实施例提供一种光照主方向模型的训练方法及电子设备,用于解决训练数据集的获取难度大的技术问题
Smart Images

Figure CN121330413B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a training method and electronic device for a principal illumination direction model. Background Technology
[0002] With the development of image processing technology, how to process images based on illumination parameters has become a key focus in the field. Objects in an image will exhibit different lighting and shadow effects under different illumination parameters. Therefore, processing images based on different illumination parameters will yield different image effects. The principal direction of illumination is one of the illumination parameters that has a significant impact on image effects. Generally, a principal direction of illumination estimation model is used to predict and process the principal direction of illumination in an image.
[0003] However, current illumination principal direction estimation models often require a large amount of labeled training data in advance, and model training has high requirements for the image quality in the dataset.
[0004] Therefore, how to reduce the difficulty of obtaining training datasets has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides a training method and electronic device for a principal illumination direction model, which solves the technical problem of difficulty in obtaining training datasets.
[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0007] Firstly, a training method for a principal illumination direction model is provided, the method comprising:
[0008] Positive sample image pairs (including the first and second images) are obtained by relighting data augmentation of an original image, or negative sample image pairs (including the third and fourth images) are obtained by relighting data augmentation of two different original images; the first encoder is trained unsupervised based on the positive or negative sample image pairs to obtain the weight parameters of the first encoder; the weight parameters of the first encoder are transferred to the illumination principal direction estimation model to be trained, and the illumination principal direction estimation model to be trained is trained in a supervised manner based on the labeled fifth image.
[0009] In this way, during the unsupervised pre-training stage, the dataset can be expanded by adding a re-illumination data enhancement strategy, increasing the diversity and complexity of the training dataset. This helps the model learn more about the changes and invariants in the data, thereby improving the model's performance in different scenarios. The encoder's weight parameters are then iteratively pre-trained in unsupervised mode, and the encoder weight parameters obtained during the unsupervised pre-training stage are loaded into the supervised training stage. Thus, the supervised training stage can complete the training of the illumination principal direction estimation model with a small amount of dataset, achieving the technical effect of reducing the difficulty of obtaining datasets.
[0010] In one possible implementation of the first aspect, obtaining a pair of positive sample images (including the first and second images) through relighting data enhancement of a single original image, or obtaining a pair of negative sample images (including the third and fourth images) through relighting data enhancement of two different original images, may include:
[0011] The first original image is relit based on the first and second principal illumination directions to obtain positive sample image pairs (including the first and second images); wherein the first principal illumination direction is different from the second principal illumination direction; and / or, the second original image is relit based on the third principal illumination direction, and the third original image is relit based on the fourth principal illumination direction to obtain negative sample image pairs (including the third and fourth images); wherein the third principal illumination direction is different from the fourth principal illumination direction; and the second original image is different from the third original image.
[0012] In this way, the dataset can be expanded by adding re-lighting data augmentation strategies, increasing the diversity and complexity of the training dataset, helping the model learn more about the changes and invariants in the data, thereby improving the model's performance in different scenarios.
[0013] In another possible implementation of the first aspect, the first principal direction of illumination includes a first azimuth angle and a first elevation angle; the step of obtaining the second principal direction of illumination includes: obtaining the second azimuth angle by rotating the first azimuth angle to the left or right at regular intervals in the horizontal direction, i.e., changing the magnitude of the azimuth angle; obtaining the second elevation angle by rotating the first elevation angle upward or downward at regular intervals in the vertical direction, i.e., changing the magnitude of the elevation angle; determining the second principal direction of illumination based on the second azimuth angle and the second elevation angle; and / or,
[0014] The third principal direction of illumination includes the third azimuth angle and the third elevation angle; the steps to obtain the fourth principal direction of illumination include: rotating the third azimuth angle to the left or right at certain intervals in the horizontal direction, i.e. changing the size of the azimuth angle, to obtain the fourth azimuth angle; rotating the third elevation angle upward or downward at certain intervals in the vertical direction, i.e. changing the size of the elevation angle, to obtain the fourth elevation angle; and then determining the second principal direction of illumination based on the fourth azimuth angle and the fourth elevation angle.
[0015] In this way, by employing a relighting data augmentation strategy, multiple different principal lighting directions can be selected to relight the same original image, or different principal lighting directions can be selected to relight different original images, thereby generating multiple views under different principal lighting directions and expanding the dataset. This results in multiple sample image pairs that can be used as training image pairs, thus reducing the difficulty of obtaining training image pairs.
[0016] In another possible implementation of the first aspect, the first original image is relit according to the first principal lighting direction and the second principal lighting direction to obtain the first image and the second image as a positive sample image pair, including: determining first rendering information according to the first principal lighting direction; determining second rendering information according to the second principal lighting direction; rendering the first original image according to the first rendering information and the second rendering information respectively to obtain the first image and the second image as a positive sample image pair; and / or,
[0017] The process involves relighting the second original image according to the third principal lighting direction and relighting the third original image according to the fourth principal lighting direction to obtain a third image and a fourth image as a negative sample image pair. This includes: determining third rendering information according to the third principal lighting direction; determining fourth rendering information according to the fourth principal lighting direction; rendering the second original image according to the third rendering information and the third original image according to the fourth rendering information to obtain a third image and a fourth image as a negative sample image pair.
[0018] In this way, the relighting data enhancement strategy can rearrange the light and shadow in the image, increasing the light and shadow details of the image. After rendering, it can generate a more realistic, three-dimensional image with higher resolution and clarity and obvious light and dark contrast, thereby achieving a significant improvement in image quality and reducing the difficulty of obtaining datasets during model training.
[0019] In another possible implementation of the first aspect, the first encoder is trained unsupervised based on the first sample set to obtain the weight parameters of the first encoder, including:
[0020] One image from the sample image pair in the first sample set is input into the first encoder to obtain the first feature; the other image from the sample image pair is input into another encoder to obtain the second feature; the sample image pair can be a positive sample image pair or a negative sample image pair; the first loss value is obtained based on the difference between the first feature and the second feature; the weight parameters of the first encoder are iteratively adjusted in the direction of reducing the first loss value until the iteration ends, and the weight parameters of the first encoder are obtained.
[0021] In this way, by minimizing the loss value, the similarity between positive sample image pairs can be minimized, while the similarity between negative sample image pairs can be maximized. This brings positive sample image pairs closer together while pushing negative sample image pairs further apart, resulting in positive sample image pairs being closer in the high-dimensional feature space and negative sample image pairs being farther apart. This can be used to extract features related to the estimation of the principal direction of illumination from the input training images, such as the shape, texture, and color of objects in the image under illumination. These features can help the principal direction of illumination estimation model estimate illumination information more accurately.
[0022] In another possible implementation of the first aspect, determining a first loss value based on the difference between the first feature and the second feature includes: performing a similarity analysis on the first feature and the second feature to obtain the first loss value.
[0023] In this way, the contrast loss value is obtained by analyzing the similarity between the first feature and the second feature; the weight parameters are updated during the unsupervised training phase to minimize the loss value until the model training stopping condition is met, and the weight parameters of the first encoder are obtained after the unsupervised pre-training iteration is completed.
[0024] In another possible implementation of the first aspect, a similarity analysis is performed on the first feature and the second feature to obtain a first loss value, including:
[0025] Determine the first mapping feature obtained from the first feature mapping; the dimension of the first mapping feature is lower than that of the first feature;
[0026] Determine the second mapping feature obtained from the second feature mapping; the dimension of the second mapping feature is lower than that of the second feature;
[0027] Add the first mapping feature to the reference feature queue; wherein the reference feature queue includes multiple reference features;
[0028] A similarity analysis is performed on multiple reference features and second projected features to obtain the first loss value.
[0029] This allows the reference features in the queue to be updated slowly and periodically, enabling the illumination principal direction estimation model to better learn the similarity between samples of the same type and distinguish samples of different categories, thus helping the model learn effective feature representations and improving model performance.
[0030] In another possible implementation of the first aspect, adding the first mapping feature to the reference feature queue includes:
[0031] The first mapping feature is added to the end of the reference feature queue; if the length of the reference queue exceeds a first preset threshold, the oldest historical mapping feature is removed from the front of the reference feature queue to update the reference features in the reference queue.
[0032] This allows the reference features in the queue to be updated slowly and periodically, enabling the illumination principal direction estimation model to better learn the similarity between samples of the same type and distinguish samples of different categories, thus helping the model learn effective feature representations and improving model performance.
[0033] In another possible implementation of the first aspect, the label includes the true principal direction of illumination; the principal direction of illumination estimation model to be trained includes a third encoder to be trained and a classification head to be trained.
[0034] The weight parameters of the first encoder are transferred to the illumination principal direction estimation model to be trained, and the illumination principal direction estimation model to be trained is trained in a supervised manner based on the third image, including:
[0035] The weight parameters of the first encoder are assigned to the second encoder of the transfer learning model;
[0036] In each round of supervised iterative training, the third image is input into the second encoder to extract features from the third image and obtain the third feature.
[0037] The third feature is input into the third encoder to be trained for feature extraction, resulting in the fourth feature;
[0038] The main direction of illumination is predicted using the classification head to be trained based on the fourth feature;
[0039] The second loss value is calculated based on the difference between the predicted principal direction of illumination and the actual principal direction of illumination;
[0040] Adjust the weight parameters of the third encoder and the classifier head to be trained in the direction of reducing the second loss value until the iteration ends, and obtain the trained illumination principal direction estimation model.
[0041] Thus, the training method provided in this application allows users to use the weight parameters of a pre-trained feature encoder as the starting point for a new task. Users do not need to train the model from scratch, but can leverage the knowledge and skills of feature encoders already trained on other large datasets to quickly train a principal direction of illumination estimation model, reducing the difficulty of obtaining training sets.
[0042] In a second aspect, an electronic device is provided, the electronic device including at least a memory and one or more processors; the memory is used to store computer instructions, which, when executed by one or more processors, cause the electronic device to perform the method as described in any of the first aspects above.
[0043] Thirdly, a computer-readable storage medium is provided that stores computer instructions or programs that, when executed on a computer, cause the method described in any of the first aspects above to be performed.
[0044] Fourthly, a computer program product is provided, the computer program product including computer instructions; when some or all of the computer instructions are run on a computer, the method as described in any of the first aspects above is performed. Attached Figure Description
[0045] Figure 1 This diagram illustrates the training process of the principal direction estimation model for illumination in related technologies. Figure 1 ;
[0046] Figure 2 This diagram illustrates the training process of the principal direction estimation model for illumination in related technologies. Figure 2 ;
[0047] Figure 3 A schematic diagram of the training method for the principal direction estimation model of illumination provided in the embodiments of this application. Figure 1 ;
[0048] Figure 4 This is a schematic diagram of the main direction of illumination provided in an embodiment of this application;
[0049] Figure 5 This is a schematic diagram of the main direction of illumination provided in an embodiment of this application;
[0050] Figure 6 Illustration of the relighting data enhancement strategy provided in the embodiments of this application Figure 1 ;
[0051] Figure 7 Illustration of the relighting data enhancement strategy provided in the embodiments of this application Figure 2 ;
[0052] Figure 8Illustration of the relighting data enhancement strategy provided in the embodiments of this application Figure 3 ;
[0053] Figure 9 Illustration of the relighting data enhancement strategy provided in the embodiments of this application Figure 4 ;
[0054] Figure 10 A schematic diagram of the training method for the principal direction estimation model of illumination provided in the embodiments of this application. Figure 2 ;
[0055] Figure 11 A schematic diagram of the training method for the principal direction estimation model of illumination provided in the embodiments of this application. Figure 3 ;
[0056] Figure 12 A schematic diagram of the training method for the principal direction estimation model of illumination provided in the embodiments of this application. Figure 4 ;
[0057] Figure 13 This is a schematic diagram comparing the performance of the embodiments of this application with that of the traditional principal direction estimation model;
[0058] Figure 14 This is a schematic diagram illustrating the application of the illumination principal direction estimation model provided in the embodiments of this application;
[0059] Figure 15 A hardware system diagram of a terminal device applicable to this application;
[0060] Figure 16 A software structure block diagram of a terminal device provided in an embodiment of this application. Detailed Implementation
[0061] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.
[0062] Before introducing specific embodiments, for ease of understanding, some concepts related to the embodiments of this application are explained by way of example for reference.
[0063] Unsupervised learning is a major training method in machine learning and deep learning. In unsupervised training, the training dataset only contains input data and has no corresponding output labels. The goal of the model is to discover the structure and patterns in the data.
[0064] Contrastive learning is an unsupervised learning method that learns data representations by comparing similar and dissimilar samples. The core idea of contrastive learning is to bring similar samples (positive sample pairs) closer together while pushing dissimilar samples (negative sample image pairs) further apart. Manual labeling is typically not required in contrastive learning, making it suitable for unsupervised learning scenarios. It should be understood that the solution in this application is applied to image processing scenarios. Therefore, in the solution of this application, a positive sample pair refers to a positive sample image pair, i.e., an image pair formed by data augmentation of the same original image, and a negative sample image pair refers to a negative sample image pair, i.e., an image pair formed by different original images.
[0065] Transfer learning is a machine learning method that allows users to use the weight parameters of a pre-trained feature encoder as a starting point for a new task. Users do not need to train the model from scratch; instead, they can leverage the knowledge and skills of feature encoders that have already been trained on other large datasets.
[0066] Supervised learning is a primary training method in machine learning and deep learning. In supervised training, the training dataset contains input data and corresponding output labels. The goal of the model is to learn a mapping relationship from input to output. During training, the model predicts the output through forward propagation, then measures the difference between the predicted output and the true label using a loss value, and updates the model weight parameters through backpropagation to reduce this difference.
[0067] Principal illumination direction estimation refers to determining the direction of the primary light source for a given scene or image. The primary light source is the light source that has a major / decisive impact on the light intensity of the scene or image. For example, in an outdoor scene, the intensity of the sun directly affects the overall brightness of the environment; therefore, the sun is the primary light source. By estimating the principal illumination direction, we can better understand and simulate the impact of lighting on the appearance of objects, thereby improving rendering quality, optimizing image processing algorithms, or enhancing the navigation capabilities of robots in complex environments. The principal illumination direction is typically defined by two angles: azimuth and elevation.
[0068] Azimuth angle (Az): Also known as horizontal longitude, it is the horizontal angle between a point on the north-pointing line and the light source in a clockwise direction. It describes the lateral position of the light source relative to north. When adjusting the azimuth angle, the parabolic surface moves left and right on the horizontal plane, and the rotation range of the azimuth angle on the horizontal plane is 0°–360°.
[0069] Elevation angle: This refers to the angle at which a light source is above the horizon, describing the vertical position of the light source relative to the horizon. When the light source is the sun, the elevation angle is also called the solar elevation angle, which is the angle between the sunlight at a certain location and the cross-section of the Earth's surface connecting that location to the Earth's center. The solar elevation angle is 0° at sunrise and sunset, 90° at noon, and can be -90° when below the horizon.
[0070] Principal direction estimation of illumination has wide applications in many fields, especially in scenarios that require accurate illumination models to enhance visual effects or simulate natural lighting conditions. For example, in Geographic Information Systems (GIS), azimuth and elevation angles are two important parameters used in principal direction estimation to describe the sun's position in the sky. By using azimuth and elevation angles, the sun's position in the sky can be determined, allowing for the calculation of illumination and shadow effects, and the simulation of the impact of illumination on terrain.
[0071] In related technologies, supervised training methods are used to train models that can predict / estimate the main direction of illumination, and the main direction of illumination is then predicted based on the model.
[0072] Next, we will introduce two schemes for training illumination principal direction estimation models using supervised training methods provided by relevant technologies.
[0073] Option 1: Figure 1 This diagram illustrates the training process of the principal direction estimation model for illumination in related technologies. Figure 1 .like Figure 1As shown, in the supervised training phase, based on labeled training images (labels including the true principal direction of illumination), the encoding and decoding structures perform normal estimation to generate a normal map, albedo estimation to generate an albedo map, and sky mask prediction to obtain the sky mask prediction result. The sky mask prediction involves extracting the sky region from the training image, distinguishing the sky region from other objects (such as buildings, ground, etc.) in the training image, and predicting the illumination information contained in the sky region. Illumination information is obtained by using an encoder and a fully connected layer (FC layer) structure or an encoding and decoding structure based on the normal map, albedo map, and sky mask prediction result. Illumination information can include the position of the sun, the color and brightness of the sky, etc. Based on the illumination information, spherical harmonics (SH) is converted into an environment map, that is, the spherical harmonic function coefficients are converted into an environment texture map. The spherical harmonic function is used to approximate the illumination distribution in the environment, and the environment map is an image reflecting the characteristics of the ambient illumination. The predicted principal direction of illumination is further calculated based on the environmental map, and this predicted principal direction includes both elevation and azimuth angles. Then, based on the difference between the predicted principal direction of illumination and the actual principal direction of illumination in the labels, the loss value is determined. The model's weight parameters are adjusted in the direction that reduces the loss value, thus completing supervised training.
[0074] In Scheme 1, a large training dataset, typically exceeding 500,000 images, needs to be prepared in advance during the supervised training phase, and the image quality of the dataset is required to be high. Therefore, Scheme 1 suffers from the disadvantage of difficulty in acquiring the dataset.
[0075] Option 2: Figure 2 This diagram illustrates the training process of the principal direction estimation model for illumination in related technologies. Figure 2 .like Figure 2 As shown, the illumination principal direction estimation model includes convolutional layers and fully connected layers. The model estimates the principal direction of illumination based on labeled training images and outputs the predicted azimuth and elevation angles. Here, the labels refer to the true values of the principal direction of illumination for each training image, i.e., the true azimuth and true elevation angles. The model then performs self-optimization and adjusts its weight parameters based on the predicted azimuth and elevation angles and the labels.
[0076] In Scheme 2, the illumination principal direction estimation model is directly trained using labeled training images. To improve model performance and prevent overfitting, a large labeled dataset, typically exceeding 500,000 images, is usually required in advance, and the image quality of the dataset must be high. Therefore, Scheme 2 also suffers from the disadvantage of difficulty in acquiring the dataset.
[0077] In summary, when electronic devices estimate the main direction of illumination based on Scheme 1 or Scheme 2, a large dataset of high-quality images needs to be prepared in advance, which makes it difficult to obtain the dataset.
[0078] To address the aforementioned issues, this application provides a training method for a principal direction of illumination estimation model, comprising two stages: unsupervised pre-training and supervised training. In the unsupervised pre-training stage, the dataset is expanded by adding a re-illumination data augmentation strategy, increasing the diversity and complexity of the training dataset. This helps the model learn more about data variations and invariants, thereby improving the model's performance in different scenarios. The encoder weight parameters are iteratively pre-trained in unsupervised mode, and the encoder weight parameters obtained in the unsupervised pre-training stage are loaded into the supervised training stage. Thus, the supervised training stage can complete the training of the principal direction of illumination estimation model with a small dataset, achieving the technical effect of reducing the difficulty of acquiring datasets.
[0079] Figure 3 A schematic diagram of the training method for the principal direction estimation model of illumination provided in the embodiments of this application. Figure 1 .like Figure 3 As shown, the training method for the illumination principal direction estimation model provided in this application includes an unsupervised training stage and a supervised training stage, wherein the unsupervised pre-training stage is based on contrastive learning.
[0080] Please see Figure 3 The training method in the unsupervised training stage is as follows: the dataset is augmented by a re-illumination data enhancement strategy to obtain training image pairs, i.e., the first sample set, which includes the first image and the second image; the first encoder is trained unsupervised based on the first sample set to obtain the weight parameters of the first encoder.
[0081] Specifically, in the unsupervised training phase, the second feature of the second image can be extracted by the fourth encoder, and the first feature of the first image can be extracted by the first encoder. The contrast loss value is obtained by analyzing the similarity between the first feature and the second feature. The weight parameters are updated in the unsupervised training phase to minimize the loss value until the model training stopping condition is met. After the unsupervised pre-training iteration is completed, the weight parameters of the first encoder are obtained.
[0082] It should be noted that the relighting data augmentation strategy can generate multiple views under different dominant lighting directions by using relighting techniques, thereby expanding the dataset. Specifically, the relighting data augmentation strategy refers to relighting one or more original images by changing the dominant lighting direction and then rendering them to obtain multiple rendered training images.
[0083] When the first and second images are image pairs obtained by relighting and rendering the same original image after changing the main lighting direction, the training image pair is a positive sample image pair. When the first and second images are image pairs obtained by relighting and rendering the same original image after changing the main lighting direction, the training image pair is a negative sample image pair.
[0084] Furthermore, minimizing the loss value can be achieved by minimizing the similarity between positive sample image pairs and maximizing the similarity between negative sample image pairs. That is, bringing positive sample image pairs closer together while pushing negative sample image pairs further apart, resulting in positive sample image pairs being closer in the high-dimensional feature space and negative sample image pairs being farther apart. The model training can stop when the maximum number of iterations is reached or the loss function converges. The loss function converges when the loss value is less than or equal to a first preset value, and / or when the change in the loss value is less than or equal to a second preset value. Both the first and second preset values can be set in advance as needed. The maximum number of iterations can be 100, 1000, etc.
[0085] Please continue reading. Figure 3 The supervised training phase training method is as follows: obtain a second sample set, which includes a fifth image with a label; transfer the weight parameters of the first encoder to the illumination principal direction estimation model to be trained, and train the illumination principal direction estimation model to be trained in a supervised manner based on the second sample set.
[0086] Specifically, in the supervised training phase, the weight parameters of the first encoder from the unsupervised training phase can be obtained and assigned to the second encoder in the transfer learning model during the supervised training phase. The second encoder of the transfer learning model is used to extract features from the labeled fifth image to obtain the third feature of the fifth image. The predicted principal direction of illumination is obtained by analyzing the third feature. The loss value is calculated based on the predicted principal direction of illumination and the true principal direction of illumination in the label. The weight parameters are then updated during the supervised training phase to minimize the loss value until the model training stopping condition is met.
[0087] It should be understood that in the embodiments of this application, the transfer learning model in the supervised training phase is usually used as a feature extractor. The second encoder of the transfer learning model is responsible for extracting general features of the labeled input data (features of the labeled fifth image in this embodiment of the application, where the label is the true main illumination direction of the fifth image, including azimuth and elevation angles).
[0088] It is understood that the first encoder, the fourth encoder, and the second encoder of the transfer learning model can be used to extract features related to the estimation of the principal direction of illumination from the input training images, such as the shape, texture, and color of objects in the images under illumination. These features can help the principal direction of illumination estimation model estimate illumination information more accurately. The first encoder, the fourth encoder, and the second encoder in the transfer learning model can be a CNN structure, a recurrent neural network (RNN) structure, or a Transformer (TF) structure; this application does not limit this.
[0089] In this embodiment, since the weight parameters of the first encoder need to be assigned to the second encoder in the transfer learning model after the unsupervised pre-training iterations, the first encoder and the second encoder in the transfer learning model adopt the same structure, i.e., they have the same number of layers, the same type of layers, and the same number of parameters. For example, they both adopt the same CNN structure. During the supervised training phase, some parameters of the transfer learning model are frozen to retain the feature extraction capability of the unsupervised pre-training phase, and the remaining parameters can be fine-tuned according to the specific task of the supervised training phase.
[0090] As discussed above, the re-illumination data enhancement strategy is one of the core inventions of this case. Therefore, the following will... Figures 4 to 5 In the next section, the principle of relighting technology will be explained in more detail by combining the concept of the main direction of illumination.
[0091] Figure 4 This is a schematic diagram of the main illumination direction provided in an embodiment of this application. Figure 4 As shown, the x-axis, y-axis, and z-axis represent the horizontal, vertical, and axial axes, respectively, with the origin at o. Together, they form a three-dimensional rectangular coordinate system (e.g., the world coordinate system). The principal direction of illumination includes the azimuth angle θ and the elevation angle φ. The azimuth angle determines the direction of the light source on the horizontal plane xoy. The azimuth angle is calculated clockwise and ranges from 0° to 360°. The elevation angle represents the angle between the projection of the light source onto the horizontal plane xoy and the light source in the vertical direction (i.e., the perpendicular direction), and ranges from -90° to 90°.
[0092] Figure 5 This is a schematic diagram of the main direction of illumination provided in an embodiment of this application. Figure 4 The principal direction of illumination shown is projected onto a plane; please refer to [link / reference]. Figure 5The x-axis is the horizontal axis, and the y-axis is the vertical axis. The azimuth angle θ is mapped to the x-coordinate, while the elevation angle φ is mapped to the y-coordinate. To simplify the problem, we can assume that this specific projection method results in a rectangle with an aspect ratio of 2:1. This means that changes in the elevation angle φ determine the height of the rectangle, while changes in the azimuth angle θ determine its width. The horizontal axis ranges from twice the height of the rectangle. In this way, the complex 3D lighting problem can be simplified into a 2D problem, making it easier to analyze and process.
[0093] Please continue reading. Figure 4 and Figure 5 Relighting is a technique that involves rotating the image horizontally (range 0° to 360°) at regular intervals in either a first or second direction, thus changing the azimuth angle. The first and second directions are completely opposite in the horizontal direction, such as left or right. Similarly, the technique involves rotating vertically (range -90° to 90°) at regular intervals in either a third or fourth direction, thus changing the elevation angle. These third and fourth directions are completely opposite in the vertical direction, such as upward or downward. These rotations alter the dominant lighting direction of the original image, effectively relighting it. The horizontal direction refers to the direction parallel to the xoy plane, while the vertical direction, also known as the vertical axis, is perpendicular to the xoy plane.
[0094] Please continue reading. Figure 4 and Figure 5 For example, relighting can also be achieved by rotating the original image 10° to the right every 10° in the horizontal direction (range 0°–360°) and upwards every 10° in the vertical direction (range -30°–30°), thus changing the principal direction of illumination and relighting the original image to obtain multiple images processed using relighting techniques. Taking the principal direction of illumination including the origin as an example, the principal direction of illumination in this example is represented as... Figure 5 As shown by the line segment ab passing through the origin, the two endpoints of the line segment are (-30°, -30°) which is (-3 / 18π, -3 / 18π) and (30°, 30°) which is (3 / 18π, 3 / 18π).
[0095] For example, relighting can be achieved by rotating the image horizontally (range 0° to 360°) every 40° and vertically (range -40° to 40°) every 20°, thereby changing the main direction of illumination of the original image and relighting it to obtain multiple images processed by relighting techniques. Figure 6 Illustration of the relighting data enhancement strategy provided in the embodiments of this application Figure 1 .like Figure 6 As shown, Figure 6 (a) in the image is the first original image. Figure 6 (b1) is the image obtained by relighting the first original image according to the main lighting direction (θ1, φ1), as shown by the arrow. Figure 6 (b2) is the image obtained by relighting and rendering the first original image. The main lighting direction (θ2, φ2) of relighting (b2) is shown by the arrow. Figure 6 (c) in the image is the second original image. Figure 6 In the image, (d1) is the image obtained by relighting and rendering the second original image. The main lighting direction (θ3, φ3) for relighting (d1) is shown by the arrow. Figure 6 In the image, (d2) is the image obtained by relighting and rendering the second original image. The main lighting direction (θ4, φ4) for relighting (d2) is shown by the arrow.
[0096] Please continue reading. Figure 6 ,exist Figure 6 When (b1) and (b2) are used as training image pairs, since (b1) and (b2) are obtained by relighting and rendering the same original image after changing the main lighting direction, (b1) and (b2) are positive sample image pairs. Similarly, in Figure 6 When (d1) and (d2) are used as training image pairs, they are also positive sample image pairs. Figure 6 When (b1) and (d1) are used as training image pairs, since (b1) and (d1) are obtained by relighting and rendering different original images after changing the main lighting direction, (b1) and (d1) are negative sample image pairs. Similarly, Figure 6 When (b2) and (d2), (b1) and (d2), and (b2) and (d1) are used as training image pairs, they are also negative sample image pairs.
[0097] The above Figure 6 In this context, it's understandable that while the image content of (b1) and (b2) appears the same, their actual display effects are completely different due to the different angles of the main direction of the relighting. Similarly, (d1) and (d2) are also different from each other.
[0098] In this way, by employing the relighting data augmentation strategy, multiple different dominant lighting directions can be selected to relight the same original image, generating multiple views under different dominant lighting directions, thus expanding the dataset. This yields multiple sample image pairs that can be used as training images, thereby reducing the difficulty of obtaining training image pairs. Furthermore, because the relighting data augmentation strategy rearranges the light and shadow in the image, it increases the detail of the light and shadow, resulting in higher resolution, clarity, and more realistic, three-dimensional images with clear contrast after rendering. This significantly improves image quality and reduces the difficulty of obtaining datasets during model training.
[0099] In this embodiment, the relighting technique can render the original image after changing the principal direction of illumination based on the following rendering equation to obtain the rendered image:
[0100] I = A⊙LB(N)
[0101] Where I represents rendering information, i.e., light intensity, also known as lighting rendering information; A represents albedo, describing the contribution of ambient light; L represents spherical harmonic coefficients, describing the radiance of the light source; N represents the normal vector; B(N) represents the spherical harmonic basis functions calculated based on the normal vector, describing how light interacts with the object surface. ⊙ represents element-wise multiplication between corresponding positions of the matrices.
[0102] Understandably, rendering equations are commonly used in computer graphics for global illumination rendering, where the principal direction of illumination is related to the spherical harmonic coefficient L. The rendering equation can be interpreted as follows: the final illumination intensity I is the result of ambient light A and light source L adjusted by the object surface B(N). The rendering equation considers both direct and indirect lighting (through reflections from other objects), thus enabling the generation of more realistic images.
[0103] Figure 7 Illustration of the relighting data enhancement strategy provided in the embodiments of this application Figure 2 Please see. Figure 7 , Figure 7 In the image (a), the original image is unrendered. Since the original image contains no lighting information, there is almost no obvious contrast between light and dark areas. Meanwhile... Figure 7 Image (b) is the image after horizontal relighting and rendering. After horizontal relighting, Figure 7 The dashed boxes in (b) exemplarily mark several highlighted areas, from Figure 7 As can be clearly seen in (b), the sky area within the dashed box and the non-sky area outside the dashed box produce a contrast in light and dark in the horizontal direction. Figure 7Image (c) shows the image after both horizontal and vertical relighting and rendering. After horizontal and vertical relighting, Figure 7 The dashed boxes in (c) exemplarily mark several highlighted areas, such as... Figure 7 As shown in (c), the sky area marked by the dashed box and the non-sky area outside the dashed box have a more pronounced contrast in light and dark in the picture.
[0104] Thus, when the original image lacks lighting information due to weather or other reasons, the image lacks clear contrast between light and dark, resulting in poor imaging quality. The relighting technique provided in this application embodiment can relight objects horizontally and / or vertically, rearranging light and shadow to enhance image detail and clarity. After rendering, it can generate more realistic, three-dimensional images with clear contrast between light and dark, significantly improving image quality. Furthermore, the relighting data augmentation strategy can effectively improve the imaging quality of training images, reducing the image quality requirements of the dataset used for model training, thereby reducing the difficulty of acquiring the dataset.
[0105] Next, we will discuss the combination of appendix. Figure 8 and attached Figure 9 The data enhancement strategy for relighting will be further introduced.
[0106] Relighting data augmentation strategies can also combine relighting techniques with data augmentation techniques. This involves relighting the original image by changing the main direction of illumination, and then performing data augmentation (such as rotation, cropping, and color enhancement). By simulating different lighting conditions and scenes using a limited number of original images, and combining this with data augmentation techniques, the model can learn the characteristics of objects under different viewpoints and changes. Relighting data augmentation strategies improve the model's generalization and adaptability to real-world scenes, enhance its robustness, and save on data costs.
[0107] Figure 8 Illustration of the relighting data enhancement strategy provided in the embodiments of this application Figure 3 One possible implementation is as follows: Figure 8 As shown, another data enhancement strategy is to first change the main illumination direction of the original image to relight the original image, and then perform data enhancement on the relit image. Figure 8 (a) in the image is the original image. Figure 8 In the image, (b1), (b2), and (b3) are the images obtained after relighting the original image by changing the principal direction of illumination. The principal direction of illumination for relighting the original image is shown by the arrow. Figure 8 In the image, (c1), (c2), and (c3) are the images obtained by rotating the re-illuminated images (b1), (b2), and (b3) 90° to the right during data augmentation. Figure 8 In the image, (d1), (d2), and (d3) are the images obtained by rotating the re-illuminated images (b1), (b2), and (b3) 90° to the left during data augmentation. Figure 8 In the image, (e1), (e2), and (e3) are images obtained by cropping the images (b1), (b2), and (b3) after relighting, respectively, during data augmentation.
[0108] and Figure 6 Similarly, the above Figure 8 In the equation (c1), (c2), and (c3), all are different; (d1), (d2), and (d3) are different; and (e1), (e2), and (e3) are also different.
[0109] Figure 9 Illustration of the relighting data enhancement strategy provided in the embodiments of this application Figure 4 Another possible implementation is, for example... Figure 9 As shown, the relighting data enhancement strategy can be to first perform data enhancement on the original image, and then relight the enhanced image by changing the main direction of illumination. Figure 9 (a) in the image is the original image. Figure 9 In the image, (b1), (b2), and (b3) are the images obtained after data augmentation by rotating 90° to the right, rotating 90° to the left, and cropping, respectively. Figure 9 In the image, (c1), (c2), and (c3) are the images obtained after relighting the image (b1) after it has been rotated 90° to the right. The main lighting direction for relighting (c1), (c2), and (c3) is shown by the arrow. Figure 9 In the image, (d1), (d2), and (d3) are the images obtained after relighting the image (b2) after it has been rotated 90° to the left. The main lighting direction for relighting (d1), (d2), and (d3) is shown by the arrow. Figure 9 In the image, (e1), (e2), and (e3) are the images obtained after relighting the cropped image (b3). The main lighting directions for relighting (e1), (e2), and (e3) are shown by the arrows.
[0110] and Figure 6 Similarly, the above Figure 9 In the given equation, (c1), (c2), and (c3) are all different; (d1), (d2), and (d3) are all different; and (e1), (e2), and (e3) are all different.
[0111] Please see Figure 8 and Figure 9When rotating, multiple different rotation directions and angles can be selected; when cropping, different proportions can be cropped; and when transforming colors, different colors can be transformed. Furthermore, data augmentation techniques can include changing image size, flipping, Gaussian blur, affine transformation, or grayscale transformation. These data augmentation techniques can be combined individually with relighting techniques, or they can be combined in combination with relighting techniques to generate more training images.
[0112] Thus, by combining relighting techniques and data augmentation strategies, we can either select multiple different principal illumination directions to relight the same original image to expand the dataset, or combine data augmentation to further augment the image. Therefore, with a limited training dataset, we can effectively increase the quantity and diversity of the training dataset, thereby reducing the difficulty of acquiring the dataset. Furthermore, the increased diversity of the training dataset enhances the diversity of the encoder's output features, effectively improving the model's generalization ability and performance, and enhancing the accuracy of the principal illumination direction estimation model.
[0113] It should be understood that the above Figure 3 The unsupervised training architecture is illustrated by taking the first and second images as a pair of positive sample images. It should be noted that the sample image pairs in the first sample set of the unsupervised training architecture can also include the third and fourth images, i.e., negative sample image pairs.
[0114] Figure 10 A schematic diagram of the training method for the principal direction estimation model of illumination provided in the embodiments of this application. Figure 2 .like Figure 10 As shown in the embodiments of this application, the illumination principal direction estimation model includes an unsupervised training phase and a supervised training phase.
[0115] Please see Figure 10 The training method for the unsupervised training phase includes at least the following steps:
[0116] (1) Re-lighting;
[0117] Specifically, the dataset is augmented using a re-lighting data enhancement strategy to obtain training image pairs, which include a first image and a second image, or a third image and a fourth image. Figure 10 The training method for the unsupervised training phase is illustrated using a training image pair including the first image and the second image as an example.
[0118] (2) Extract the second feature q of the second image through the fourth encoder, and map the second feature q to a low-dimensional representation space through the second multilayer perceptron projection head (MLP projection head) to obtain the second projected feature q'; extract the first feature p of the first image through the first encoder, and map the first feature p to a low-dimensional representation space through the first MLP projection head to obtain the first projected feature p'; obtain the contrast loss value by analyzing the similarity between the second projected feature q' and the first projected feature p'; update the weight parameters in the unsupervised training stage to minimize the loss value until the model training stopping condition is met, and assign the weight parameters of the first encoder after the unsupervised pre-training iteration is completed to the second encoder of the transfer learning model in the supervised training stage.
[0119] Please continue reading. Figure 10 The supervised training phase involves the following steps: First, the weight parameters of the first encoder from the unsupervised training phase are obtained and loaded into the second encoder of the transfer learning model. Second, the second encoder of the transfer learning model extracts features from the labeled fifth image to obtain a third feature o. Third, the third encoder processes this third feature o into a third classification feature o'. A fully connected layer classification head (FC) maps the third classification feature o' to the final category label, generating the predicted principal illumination direction. The loss value is calculated based on the predicted principal illumination direction and the true principal illumination direction in the label. Weight parameters are then updated during the supervised training phase to minimize the loss value until the model training stopping condition is met.
[0120] During the supervised training phase, the second encoder of the transfer learning model can extract general features from the labeled fifth image, obtaining the general features of the fifth image, namely the third feature o. The third encoder can then process the third feature o into a third classification feature o' adapted for classification, which can be used in the subsequent FC classification head for classification tasks.
[0121] It's important to note that the FC (Fully Connected) classifier head is located at the end of the third encoder. The FC classifier head can include one or more fully connected layers, and an output layer that uses either a softmax activation function (for multi-class classification) or a sigmoid activation function (for binary classification) to generate a probability distribution. By adjusting the number of layers and neurons in the FC classifier head, users can control the model's complexity and learning capacity. The FC layer can output the predicted principal direction of illumination.
[0122] The hidden high-dimensional feature representation of an image is transformed into a form used in contrastive learning tasks through an MLP projection head, enhancing the discriminative power of the features through nonlinear transformations. For example, the second feature q is transformed into the second projected feature q'.
[0123] Thus, the MLP projection head implemented through a multilayer perceptron, and the training method for the illumination principal direction estimation model provided in this embodiment, can learn high-dimensional feature representations of the data, thereby improving the model's ability to fit complex data. By mapping high-dimensional features to a low-dimensional space, the MLP projection head helps reduce overfitting, allowing the model to perform better on data outside the training set. By outputting the illumination principal direction through an FC classification head, the training method for the illumination principal direction estimation model provided in this embodiment can also effectively reduce the dimensionality of the feature space, thereby reducing model complexity, which helps avoid overfitting and improve the model's generalization ability. Therefore, the training method for the illumination principal direction estimation model provided in this embodiment can improve the model performance and estimation accuracy of the illumination principal direction estimation model.
[0124] Figure 11 A schematic diagram of the training method for the principal direction estimation model of illumination provided in the embodiments of this application. Figure 3 .like Figure 11 As shown in the embodiments of this application, the illumination principal direction estimation model includes an unsupervised training phase and a supervised training phase.
[0125] Please see Figure 11 The training method in the unsupervised training phase is as follows: The dataset is augmented using a re-illumination data enhancement strategy to obtain training image pairs, including a first image and a second image. A second feature q is extracted from the second image using a momentum encoder. This second feature q is mapped to a low-dimensional representation space using a second MLP projection head to obtain a second projected feature q', which is then added to the reference feature queue. A first feature p is extracted from the first image using a first encoder. This first feature p is mapped to a low-dimensional representation space using a first MLP projection head to obtain a first projected feature p'. The similarity between the first projected feature p' and each reference feature in the reference feature queue is calculated, and the contrastive loss is obtained. The weight parameters in the unsupervised training phase are updated to minimize the loss value until the model training stopping condition is met. Finally, the weight parameters of the first encoder after the unsupervised pre-training iterations are assigned to the second encoder of the transfer learning model in the supervised training phase.
[0126] The training methods for the supervised training phase can be referred to Figure 10 The explanation will not be repeated here.
[0127] Understandably, a momentum encoder is a deep neural network, which can employ a CNN structure. A CNN consists of multiple convolutional layers, activation functions, and pooling layers to extract features from a first image. As the layers deepen, CNNs can capture increasingly abstract features. Unlike traditional first-level encoders, the weight parameters of a momentum encoder are not updated directly through backpropagation, but rather slowly through a momentum update strategy. This ensures that the encoder's weight parameter updates are not entirely dependent on the current input, but also consider the output from the previous time step, allowing the unsupervised pre-training architecture to maintain a more stable feature representation. Through momentum updates, the momentum encoder can gradually learn the knowledge from the previous time step, making the extracted features more consistent across different time points. This helps maintain the consistency of the features extracted by the encoder, which in turn helps the model see more negative examples during training, thus improving the model's generalization ability. This generalization ability allows the model to better learn the intrinsic structure of the data, thereby performing better on unseen data.
[0128] It's important to note that the reference feature queue stores features extracted from training images up to the current time step; these features are called reference features. In the reference feature queue, reference features are typically stored as a dictionary, and multiple reference features are concatenated together using tensor concatenation. The purpose of the reference feature queue is to provide a reference set, enabling the network to perform contrastive learning over a larger dataset. During training, the first projected feature p' is added to the end of the reference feature queue. If the length of the reference queue exceeds a preset maximum value, the oldest element is removed from the front of the queue, allowing the reference features in the queue to be updated slowly and periodically.
[0129] Specifically, commonly used similarity analysis methods include cosine similarity analysis, which calculates the cosine angle between the first projected feature p' and each reference feature in the reference feature queue. Reference features in the reference feature queue that are similar to the first projected feature p' are considered as positive samples k. + Reference features that are not similar to the first projected feature p' are considered as negative samples k. ﹣ .
[0130] In the above Figure 11 In the process illustrated, the objective of the contrastive loss is to minimize the similarity between positive sample image pairs (i.e., minimize the difference between positive sample image pairs) while maximizing the similarity between negative sample image pairs (i.e., maximize the difference between positive sample image pairs). The positive samples k in the reference queue... + and negative sample k ﹣The first projected feature p' can be used to calculate the contrastive loss to guide the model's learning. The illumination principal direction estimation model can better learn the similarity between samples of the same type and distinguish samples of different categories, helping the model learn effective feature representations, thereby improving the model's performance.
[0131] The contrast loss function can be cosine similarity contrast loss, Euclidean distance contrast loss, Mahalanobis distance contrast loss, or InfoNCE (Information Noise Contrast Estimation) loss, etc. For example, when the contrast loss function is InfoNCE loss, its calculation formula is:
[0132] L(q',k)=-log(exp(s im(q',k + )) / Σexp(s im(q',k ﹣ )))
[0133] Where, s im(q',k + ) represents the first projected feature q' and the positive sample k + The similarity between them, Σexp(s im(q',k) ﹣ )) represents all negative samples k ﹣ The sum of similarities.
[0134] In this way, the unsupervised pre-training stage can combine relighting data augmentation strategies with existing data augmentation techniques to obtain views of the same original image under multiple different principal illumination directions. Using these views as training data pairs not only reduces the number of training data points required and simplifies data acquisition, but also enhances the diversity of encoder output features. Furthermore, the unsupervised pre-training stage learns feature representations that distinguish different categories of samples through contrastive learning. By optimizing the loss value, the principal illumination direction estimation model can continuously update the weight parameters in the unsupervised training framework, thereby improving its performance.
[0135] The supervised training framework provided in this application can reduce the training time and the number of data entries required for the illumination principal direction estimation model, while also improving the performance of the illumination principal direction estimation model.
[0136] In the embodiments of this application, Figure 12 A schematic diagram of the training method for the principal direction estimation model of illumination provided in the embodiments of this application. Figure 4 .like Figure 12 As shown, in a specific example, the training method for the supervised training phase is as follows: Figure 10The supervised training phase can include: a transfer learning model, a third encoder, and a fully connected (FC) classifier head. The third encoder consists of four convolutional layers: conv7-64, conv5-127, conv3-256, and conv3-256. conv7-64 indicates that the convolutional layer has seven kernels, each with 64 channels. This means the CNN will learn seven different features, each consisting of 64 channels, such as object shape, texture, and color. The other three convolutional layers are similar and will not be elaborated further. The FC classifier head consists of four fully connected layers: FC-64, FC-32, FC-16, and FC-2. FC-64 indicates that the fully connected layer has 64 neurons. This means the CNN will learn to map the features extracted by the convolutional layers to a 64-dimensional space, and the weights between the 64 neurons will be trained to capture the complex relationships between the input features. The other three fully connected layers are similar and will not be elaborated further.
[0137] Thus, the supervised model provided in this application embodiment can learn a large number of different features through the encoder, with each convolutional layer learning features at different levels of the input data. Shallower convolutional layers may learn lower-level features (such as edges and textures), while deeper convolutional layers may learn higher-level features (such as object parts and scene structures). By stacking multiple convolutional layers, the encoder can perform non-linear transformations on the training images, thereby extracting more abstract and complex features. Furthermore, the model can capture spatial information over a wider range, thus better understanding the entire image or scene and improving the estimation accuracy of the principal direction of illumination. The fully connected (FC) classification head can map the features extracted by the encoder onto the probability distribution of multiple categories. By stacking multiple fully connected layers, the classification head can progressively learn the mapping from low-level features to high-level category information. In addition, fully connected layers allow the model to fuse features from different convolutional layers and channels, thereby comprehensively considering various information to make classification decisions.
[0138] Specifically, the FC (Fully Continuous) classifier head is responsible for mapping the third classification feature o' generated by the third encoder to the final class label, i.e., classifying the fifth image based on the third classification feature o'. The FC classifier head maps the input to the predicted scores of the class label through an activation function (such as softmax). These predicted scores are used to calculate the loss value and update the weight parameters of the supervised training architecture through the backpropagation algorithm.
[0139] During supervised training, the fifth image is used to generate the predicted principal direction of illumination through forward propagation. This supervised training method learns the mapping relationship between the input fifth image and the principal direction of illumination, thus achieving the task of predicting the principal direction of illumination without the need for manual feature design and selection. The fifth image, through forward propagation, generates the predicted principal direction of illumination information, enabling the task of predicting the principal direction of illumination.
[0140] from Figures 3 to 12 As can be seen, the illumination principal direction estimation model provided in this application, during the supervised training phase, loads the target encoder weight parameters into the second encoder of the transfer learning model. Since the target encoder weight parameters in the unsupervised pre-training phase already contain general features learned from a large amount of data, the model can converge to better performance faster during supervised training, reducing training time. The target encoder weight parameters help the model learn more representative features, which can better generalize to new data, thereby improving the performance of the illumination principal direction estimation model on new tasks. Because the target encoder weight parameters provide rich feature representations, the model is more likely to avoid overfitting when facing limited labeled data, thus achieving better performance on the test set. Reusing the target encoder weight parameters in the supervised training phase also saves on this part of the computational resources. In addition, transfer learning enables the model to transfer knowledge between different domains or tasks. Even if the unsupervised pre-training is performed on a completely different dataset, the target encoder weight parameters may still be helpful for new tasks.
[0141] In summary, the training method for the illumination principal direction estimation model provided in this application includes two stages: unsupervised pre-training and supervised training. During the unsupervised pre-training stage, a re-illumination data augmentation strategy can be used to expand the training dataset and improve image quality, thereby reducing the difficulty of obtaining the training dataset. It also increases the diversity of encoder output features, enabling the model to adapt to different data distributions and enhancing its generalization ability to new data and its robustness to noise. Furthermore, after the unsupervised pre-training stage iterations are completed, the target encoder weight parameters can be obtained, which can be used to accelerate the convergence speed of the supervised training stage, reducing training time while improving the generalization ability of the illumination principal direction estimation model.
[0142] The illumination principal direction estimation model provided in this application removes the transfer learning model during use, retaining only the third encoder and the FC classification head. The classification features of the input image are extracted based on the trained third encoder, and the FC classification head outputs the predicted illumination principal direction based on the classification features.
[0143] Table 1 is a comparison table of the effects of the traditional illumination principal direction estimation model obtained by the training method provided in Scheme 2 and the illumination principal direction estimation model obtained by the training method provided in the embodiments of this application. Figure 13 This is a schematic diagram comparing the performance of the embodiments of this application with that of the traditional principal direction estimation model. Figure 13 (a) in the diagram is a schematic diagram of the elevation angle CDF (<20°) in the traditional principal direction of illumination estimation model. Figure 13 (b) in the diagram is a schematic diagram of the azimuth CDF (<30°) in the traditional principal direction of illumination estimation model. Figure 13 (c) is a schematic diagram of the elevation angle CDF (<20°) in the principal direction of illumination estimation model of this application embodiment. Figure 13 (d) is a schematic diagram of the azimuth angle CDF (<30°) in the illumination principal direction estimation model of this application embodiment.
[0144] Table 1
[0145]
[0146] Please refer to Table 1. The average elevation angle error is used to characterize the estimation accuracy of the model. As shown in Table 1, compared with the traditional main illumination direction estimation model, the average elevation angle error and the average azimuth angle error of the main illumination direction estimation model in this application embodiment are reduced to a certain extent. The main illumination direction estimation model in this application embodiment has higher estimation accuracy and better model performance.
[0147] Please refer to Table 1 and Figure 13 Elevation CDF (<20°) represents the proportion of images with an elevation angle not exceeding 20° among all images, and azimuth CDF (<30°) represents the proportion of images with an azimuth angle not exceeding 30° among all images. Both elevation and azimuth CDF are directly proportional to generalization ability; the larger the elevation and azimuth CDF, the better the generalization ability of the model. (From Table 1 and...) Figure 13 As can be seen, compared with the traditional illumination principal direction estimation model, the illumination principal direction estimation model of this application embodiment has improved the elevation angle CDF (<20°) and azimuth angle CDF (<30°) to a certain extent, and the illumination principal direction estimation model of this application embodiment has better generalization.
[0148] The cumulative distribution function (CDF) is a method in probability theory for describing the distribution of random variables. The CDF indicates the proportion of variables that satisfy certain constraints. Generalization is an important concept in transfer learning; it helps ensure that pre-trained models can capture similar feature representations across different tasks, thereby achieving better generalization ability and cross-domain applications.
[0149] Thus, the illumination principal direction estimation model of this application embodiment has a significantly improved generalization ability because both the elevation angle CDF (<20°) and azimuth angle CDF (<30°) are greater than those of the traditional illumination principal direction estimation model. This helps to capture similar feature representations across different tasks, thereby achieving better performance on different tasks.
[0150] The principal direction of illumination estimation model provided in this application can be applied not only to the aforementioned GIS field, but also to many other fields such as games. For example, when creating 3D models and animations, the estimation of the principal direction of illumination determines how light interacts with objects, playing a crucial role in the contrast and shadow effects of the image. In lightmapping and global illumination rendering techniques (such as ray tracing), accurate estimation of the principal direction of illumination can significantly improve rendering quality. Furthermore, in virtual reality (VR) and augmented reality (AR) applications, accurate simulation of real-world lighting conditions is required to create immersive experiences.
[0151] Figure 14 This diagram illustrates the application of the principal direction of illumination estimation model provided in this application. The principal direction of illumination estimation model trained using the method of this application can be applied to the AR (Augmented Reality) field. In AR, by fusing digital information from the virtual world with the real world, users can perceive and interact with the virtual world in the real world. Figure 14 As shown, in real outdoor scenarios, including Figure 14 (a) shows a real first cylinder 1401, which is vertically positioned on a horizontal surface. The shadow of the first cylinder 1401 under sunlight is a second cylinder 1402. A virtual third cylinder 1403 needs to be inserted directly in front of the first cylinder 1401 (i.e., on the same meridian as the first cylinder 1401, closer to the observer's direction) using AR technology. The theoretical shadow of the third cylinder 1403 is simulated as a fourth cylinder 1404 using the principal direction of illumination estimation model provided in this application (including azimuth and elevation angles).
[0152] Please continue reading. Figure 14 In (a), in a real outdoor scene, when sunlight encounters the first cylinder 1401, the first cylinder 1401 will block the sunlight, forming a shadow of the second cylinder 1402. Since the first cylinder 1401 is vertically set on the horizontal ground, the first angle A between the second cylinder 1402 and the due north direction is the actual solar azimuth angle of the first cylinder 1401.
[0153] Please continue reading. Figure 14In (b), the third cylinder 1403 is a virtual object, and the fourth cylinder 1404 is the theoretical shadow calculated based on the position and main direction of illumination of the third cylinder 1403. Therefore, the second angle B between the fourth cylinder 1404 and the due north direction is the theoretical solar azimuth angle of the third cylinder.
[0154] The difference between the actual solar azimuth and the theoretical solar azimuth is called the azimuth error, which is the difference between the first included angle A and the second included angle B, and is also the third included angle C.
[0155] Thus, by Figure 14 It can be seen that the theoretical solar azimuth angle calculated based on the principal direction of illumination (including azimuth and elevation angle) obtained by the principal direction of illumination estimation model provided in this application is small in difference from the actual solar azimuth angle. That is to say, the third included angle A is small. Therefore, the principal direction of illumination estimation model provided in this application has a small error, high estimation accuracy of the principal direction of illumination, and good model performance.
[0156] In addition to its applications in the AR field, the illumination principal direction estimation model provided in this application can also be applied to the field of terminal devices. Terminal devices can perform model training, which can be used to implement the training method for the illumination principal direction estimation model described in the above method embodiments, and can also be used to apply the illumination principal direction estimation model described in the above method embodiments to perform post-processing lighting on images in the terminal device.
[0157] The terminal device performing unsupervised pre-training and the terminal device performing supervised training can be the same terminal device or different terminal devices; this application embodiment does not impose specific limitations in this regard. As an example and not a limitation, the terminal device can be, but is not limited to, tablet computers, desktop computers, laptop computers, handheld computers, laptops, in-vehicle devices, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), mobile phones, etc.; this application embodiment does not impose limitations in this regard.
[0158] The following will combine Figures 15 to 16 This application describes in detail the hardware and software systems of the terminal devices to which it applies.
[0159] It should be understood that the hardware and software systems in the embodiments of this application can execute the various methods described in the foregoing embodiments of this application. That is, the specific working processes of the various products described below can be referred to the corresponding processes in the foregoing method embodiments.
[0160] Figure 15This is a hardware system for a terminal device applicable to this application.
[0161] Terminal device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, display screen 170, camera 193, etc.
[0162] It should be noted that, Figure 15 The structure shown does not constitute a specific limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may include a... Figure 15 The components shown may include more or fewer components, or the terminal device 100 may include... Figure 15 The components shown may be a combination of certain components, or the terminal device 100 may include... Figure 15 Sub-components of some of the components shown. Figure 15 The components shown can be implemented in hardware, software, or a combination of both.
[0163] Processor 110 may include one or more processing units. For example, processor 110 may include at least one of the following processing units: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and neural network processing unit (NPU). These different processing units may be independent devices or integrated devices. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.
[0164] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0165] Figure 15 The connection relationships between the modules shown are merely illustrative and do not constitute a limitation on the connection relationships between the modules of the terminal device 100. Optionally, the modules of the terminal device 100 may also adopt a combination of various connection methods described in the above embodiments.
[0166] The charging management module 140 is used to receive power from the charger.
[0167] The wireless communication function of terminal device 100 can be implemented through devices such as antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.
[0168] The mobile communication module 150 can provide a wireless communication solution for use on the terminal device 100.
[0169] Similar to the mobile communication module 150, the wireless communication module 160 can also provide a wireless communication solution for use on the terminal device 100.
[0170] In some embodiments, the antenna 1 of the terminal device 100 is coupled to the mobile communication module 150, and the antenna 2 of the terminal device 100 is coupled to the wireless communication module 160, so that the terminal device 100 can communicate with the network and other terminal devices through wireless communication technology.
[0171] The ISP (Integrated Photo Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can perform algorithmic optimization on image noise, brightness, and color. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0172] For example, the ISP can process the image based on the principal illumination direction output by the principal illumination direction estimation model provided in this application embodiment. As the principal illumination direction processed by the ISP changes, the lighting and shadow effects of the image change accordingly. The image can be an image stored in the terminal device 100 or an image captured by the camera 193.
[0173] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into a standard red-green-blue (RGB), YUV, or other image signal format. In some embodiments, the terminal device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0174] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals. For example, when terminal device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0175] The hardware system of terminal device 100 has been described in detail above. The software system of terminal device 100 is described below. The software system can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment takes a layered architecture as an example to exemplarily describe the software system of terminal device 100.
[0176] Figure 16 A software structure block diagram of a terminal device provided in an embodiment of this application.
[0177] A layered architecture divides the system into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the system is divided into five layers, from top to bottom: application layer, application framework layer, hardware abstraction layer, driver layer, and hardware layer.
[0178] The application layer may include a series of application packages. In the embodiments of this application, the application packages may include camera applications, video applications, AR applications, etc.
[0179] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes some predefined functions. In this embodiment, the application framework layer may include a camera access interface, which may include camera management and camera devices. The camera access interface is used to provide APIs and programming frameworks for camera applications.
[0180] The hardware abstraction layer is an interface layer located between the application framework layer and the driver layer, providing a virtual hardware platform for the operating system. In this embodiment, the hardware abstraction layer may include a camera hardware abstraction layer and a camera algorithm library.
[0181] The camera hardware abstraction layer can provide virtual hardware for the camera device, detecting the actual total illuminance and / or actual dynamic range of the test environment based on the image within the camera's field of view. The camera algorithm library may include the runtime code and data required for the terminal device involved in this application's embodiments to perform shooting, such as data needed to detect the actual total illuminance and / or actual dynamic range of the test environment.
[0182] The driver layer is the layer between hardware and software. It includes drivers for various hardware components, such as camera drivers, digital signal processor drivers, sensor drivers, and image processor drivers.
[0183] The camera device driver is used to drive the camera sensor to acquire images and to drive the image signal processor to preprocess the images. The digital signal processor driver is used to drive the digital signal processor to process images. The image processor driver is used to drive the graphics processor to process images. For example, the image processor driver can obtain the corresponding principal direction of illumination based on the principal direction of illumination estimation model provided in the embodiments of this application, and then perform image rendering and other processing based on the principal direction of illumination.
[0184] The hardware layer may include image processors, digital signal processors, sensors, and image signal processors, etc.
[0185] The following example, using a photographing scene as a case study, illustrates the workflow of the software and hardware systems of the terminal device 100 in processing images according to the main direction of illumination.
[0186] For example, the camera 193 is controlled by the camera driver to take pictures and obtain the original image to be processed. After processing by calling the illumination principal direction estimation model provided in this application (i.e., the illumination principal direction estimation model trained based on the method of this application), the relevant data of the illumination principal direction is output. The original image is then lit and rendered according to the illumination principal direction to simulate the lighting and shadow effects of objects in the original image under the illumination principal direction. The processed image is then stored or sent to other processes for use.
[0187] This application also provides an electronic device, which includes at least a memory and one or more processors; the memory is used to store computer instructions, and when one or more processors execute the computer instructions, the electronic device performs the functions or steps described in the above method embodiments.
[0188] This application also provides a computer-readable storage medium including computer instructions that, when executed on the electronic device, cause the electronic device to perform the various functions or steps described in the method embodiments.
[0189] This application also provides a computer program product that, when run on an electronic device, causes the electronic device to perform the functions or steps described in the above method embodiments.
[0190] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0191] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0192] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0193] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0194] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A training method for a principal illumination direction model, characterized in that, include: Based on at least one original image, a first sample set is generated, comprising sample image pairs; the sample image pairs include positive sample image pairs and / or negative sample image pairs; the positive sample image pairs include a first image and a second image obtained by relighting data enhancement of one of the original images, and the negative sample image pairs include a third image and a fourth image obtained by relighting data enhancement of two different original images; the relighting data enhancement refers to: relighting one or more original images by changing the principal direction of illumination and rendering to obtain multiple rendered training images, the principal direction of illumination including azimuth and elevation angles, the rendering based on illumination rendering information to obtain the rendered images, the illumination rendering information referencing direct illumination and indirect illumination, the indirect illumination including reflected light from other objects; The first encoder is trained unsupervised based on the first sample set to obtain the weight parameters of the first encoder. Obtain a second sample set; the second sample set includes a fifth image carrying a label; The weight parameters of the first encoder are transferred to the illumination principal direction estimation model to be trained, and the illumination principal direction estimation model to be trained is trained in a supervised manner based on the second sample set.
2. The method according to claim 1, characterized in that, The step of generating a first sample set comprising sample image pairs based on at least one original image includes: The first original image is relit based on the first illumination principal direction and the second illumination principal direction respectively, and the first image and the second image are used as positive sample image pairs; wherein, the first illumination principal direction is different from the second illumination principal direction; And / or, The second original image is relit based on the third principal illumination direction, and the third original image is relit based on the fourth principal illumination direction to obtain the third image and the fourth image as a negative sample image pair; wherein the third principal illumination direction is different from the fourth principal illumination direction; and the second original image is different from the third original image.
3. The method according to claim 2, characterized in that, The first principal direction of illumination includes a first azimuth angle and a first elevation angle; The steps for obtaining the second principal direction of illumination include: rotating the first azimuth angle along the horizontal direction by a first preset angle to obtain the second azimuth angle; rotating the first elevation angle along the vertical direction by a second preset angle to obtain the second elevation angle; and determining the second principal direction of illumination based on the second azimuth angle and the second elevation angle. And / or, The third principal direction of illumination includes a third azimuth angle and a third elevation angle; the step of obtaining the fourth principal direction of illumination includes: rotating the third azimuth angle along the horizontal direction by a first preset angle to obtain a fourth azimuth angle; rotating the third elevation angle along the vertical direction by a second preset angle to obtain a fourth elevation angle; and determining the fourth principal direction of illumination based on the fourth azimuth angle and the fourth elevation angle.
4. The method according to claim 2, characterized in that, The step of relighting the first original image based on the first principal lighting direction and the second principal lighting direction to obtain the first image and the second image as a positive sample image pair includes: determining first rendering information based on the first principal lighting direction; determining second rendering information based on the second principal lighting direction; and rendering the first original image based on the first rendering information and the second rendering information to obtain the first image and the second image as a positive sample image pair. And / or, The step of relighting the second original image based on the third principal lighting direction and relighting the third original image based on the fourth principal lighting direction to obtain the third image and the fourth image as a negative sample image pair includes: determining third rendering information based on the third principal lighting direction; determining fourth rendering information based on the fourth principal lighting direction; rendering the second original image based on the third rendering information and rendering the third original image based on the fourth rendering information to obtain the third image and the fourth image as a negative sample image pair.
5. The method according to any one of claims 1 to 4, characterized in that, The step of unsupervised training of the first encoder based on the first sample set to obtain the weight parameters of the first encoder includes: The sample image pairs in the first sample set are input to the first encoder to obtain a first feature of one image and a second feature of the other image in the sample image pair; the sample image pair is either the positive sample image pair or the negative sample image pair. A first loss value is determined based on the difference between the first feature and the second feature; The weight parameters of the first encoder are iteratively adjusted in the direction of reducing the first loss value until the iteration ends, thus obtaining the weight parameters of the first encoder.
6. The method according to claim 5, characterized in that, Determining the first loss value based on the difference between the first feature and the second feature includes: A similarity analysis is performed on the first feature and the second feature to obtain the first loss value.
7. The method according to claim 6, characterized in that, The step of performing similarity analysis on the first feature and the second feature to obtain the first loss value includes: Determine a first mapping feature obtained from the first feature mapping; the dimension of the first mapping feature is lower than that of the first feature; Determine a second mapping feature obtained from the second feature mapping; the dimension of the second mapping feature is lower than that of the second feature; The first mapping feature is added to the reference feature queue; wherein, the reference feature queue includes multiple historical mapping features; the historical mapping features refer to the mapping features that were added to the reference feature queue before the first mapping feature was added to the reference feature queue. A similarity analysis is performed on multiple historical mapping features and the second mapping feature to obtain the first loss value.
8. The method according to claim 7, characterized in that, Adding the first mapped feature to the reference feature queue includes: Add the first mapping feature to the end of the reference feature queue; If the length of the reference feature queue exceeds a first preset threshold, the earliest historical mapping feature is removed from the front of the reference feature queue to update the historical mapping features in the reference feature queue.
9. The method according to any one of claims 1 to 4 or 6 to 8, characterized in that, The label includes the true principal direction of illumination; the principal direction of illumination estimation model to be trained includes a third encoder to be trained and a classification head to be trained. The step of transferring the weight parameters of the first encoder to the illumination principal direction estimation model to be trained, and training the illumination principal direction estimation model to be trained in a supervised manner based on the second sample set, includes: The weight parameters of the first encoder are assigned to the second encoder of the transfer learning model; In each round of supervised iterative training, the fifth image is input to the second encoder to extract features from the fifth image to obtain the third feature; The third feature is input into the third encoder to be trained for feature extraction to obtain the fourth feature; The classification head to be trained is used to predict the main direction of illumination based on the fourth feature; The second loss value is calculated based on the difference between the predicted principal direction of illumination and the actual principal direction of illumination; The weight parameters of the third encoder and the classifier to be trained are adjusted in the direction of reducing the second loss value until the iteration ends, and the trained illumination principal direction estimation model is obtained.
10. An electronic device, characterized in that, The electronic device includes at least a memory and one or more processors; the memory stores computer instructions that, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 9.
12. A computer program product, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Illumination estimation method and device
CN111063017A
Text recognition system training method in self-supervised contrast learning natural scene
CN114973226A