Stereoscopic scene automatic light control method and system thereof
By constructing a panoramic image of a three-dimensional scene and using artificial intelligence algorithms to extract feature information, the problem of unnatural light effects in game scenes is solved, intelligent light adjustment is achieved, and realistic and adaptable are enhanced.
Patent Information
- Application Number
- CN202410950088.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-07-16
AI Technical Summary
Traditional light control methods have stiff and unnatural light and shadow effects in game scenes, and lack flexibility, making it difficult to adapt to the lighting needs of different scenes.
By constructing a panoramic image of a three-dimensional scene, and using an artificial intelligence-based image processing and analysis algorithm, feature information of the foreground and background are extracted, and the foreground-background semantic joint interaction network is used to adjust the light intensity, color and irradiation direction.
It realizes intelligent adjustment of light, making the scene more in line with the actual lighting conditions, enhances the sense of reality, and adapts to the shooting or rendering needs of different scenes.
Smart Images

Figure CN118918236B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent light control, and more specifically, to a method and system for automatically adjusting light in a three-dimensional scene. Background Art
[0002] With the development of technology, people have higher requirements for the visual realism of game scenes. As the most commonly used realism rendering algorithm at present, ray tracing plays a huge role in fields such as three-dimensional animation, virtual reality, and digital twin.
[0003] However, in traditional light control methods, light sources and shadows are usually pre-calculated and rendered, which may limit the interactivity of the game world and result in the light and shadow effects being possibly rigid and unnatural. In addition, some traditional light control technologies are only applicable to specific types of game scenes and have poor effects in other scenes, lacking flexibility.
[0004] Therefore, a scheme for automatically adjusting light in a three-dimensional scene is desired. Summary of the Invention
[0005] This application aims at the deficiencies in the prior art and provides a method and system for automatically adjusting light in a three-dimensional scene. By constructing a panoramic image of the three-dimensional scene and using artificial intelligence-based image processing and analysis algorithms to perform feature analysis and extraction of the foreground features and background features of the panoramic image of the three-dimensional scene, the light intensity, light color, and light irradiation direction are adaptively adjusted based on the semantic joint interaction features between the foreground and background of the three-dimensional scene. In this way, the light can be intelligently adjusted to make the scene where the characters are located more in line with the actual lighting conditions, enhancing the sense of realism, and the adaptive adjustment of light can meet different shooting or rendering requirements, whether it is an indoor scene or an outdoor environment.
[0006] According to one aspect of this application, a method for automatically adjusting light in a three-dimensional scene is provided, which includes:
[0007] Construct a panoramic image of the three-dimensional scene;
[0008] Extract the foreground part and the background part of the panoramic image to obtain the foreground part of the three-dimensional scene and the background part of the three-dimensional scene;
[0009] Perform foreground feature extraction and background feature extraction on the foreground part of the three-dimensional scene and the background part of the three-dimensional scene respectively to obtain a foreground feature map of the foreground part of the three-dimensional scene and a background feature map of the background part of the three-dimensional scene;
[0010] Input the foreground feature map of the three-dimensional scene and the background feature map of the three-dimensional scene into the feature multi-scale perception enhancement module respectively to obtain the enhanced foreground feature map of the three-dimensional scene and the enhanced background feature map of the three-dimensional scene;
[0011] Input the enhanced foreground feature map of the three-dimensional scene and the enhanced background feature map of the three-dimensional scene into the foreground-background semantic joint interaction network to obtain the three-dimensional scene foreground-background semantic joint interaction feature map as the three-dimensional scene foreground-background semantic joint interaction feature;
[0012] Based on the three-dimensional scene foreground-background semantic joint interaction feature, obtain the light parameter decoding result, where the light parameter decoding result includes light intensity, light color, and light irradiation direction.
[0013] According to another aspect of the present application, a three-dimensional scene automatic light control system is provided, which includes:
[0014] A panoramic image construction module for constructing a panoramic image of the three-dimensional scene;
[0015] A panoramic image segmentation module for extracting the foreground part and the background part of the panoramic image to obtain the foreground part of the three-dimensional scene and the background part of the three-dimensional scene;
[0016] A panoramic image segmentation partial feature extraction module for respectively performing foreground feature extraction and background feature extraction on the foreground part of the three-dimensional scene and the background part of the three-dimensional scene to obtain the foreground feature map of the three-dimensional scene and the background feature map of the three-dimensional scene;
[0017] A panoramic image segmentation partial feature enhancement module for respectively inputting the foreground feature map of the three-dimensional scene and the background feature map of the three-dimensional scene into the feature multi-scale perception enhancement module to obtain the enhanced foreground feature map of the three-dimensional scene and the enhanced background feature map of the three-dimensional scene;
[0018] A panoramic image segmentation partial semantic joint interaction module for inputting the enhanced foreground feature map of the three-dimensional scene and the enhanced background feature map of the three-dimensional scene into the foreground-background semantic joint interaction network to obtain the three-dimensional scene foreground-background semantic joint interaction feature map as the three-dimensional scene foreground-background semantic joint interaction feature;
[0019] A light parameter decoding result generation module for obtaining the light parameter decoding result based on the three-dimensional scene foreground-background semantic joint interaction feature, where the light parameter decoding result includes light intensity, light color, and light irradiation direction.
[0020] Due to the adoption of the above technical solutions, the present application has significant technical effects:
[0021] The automatic light control method and system for a three-dimensional scene provided by this application construct a panoramic image of the three-dimensional scene and use image processing and analysis algorithms based on artificial intelligence to perform feature analysis and extraction of the foreground features and background features of the panoramic image of the three-dimensional scene, so as to adaptively adjust the light intensity, light color, and light irradiation direction based on the semantic joint interaction features between the foreground and background of the three-dimensional scene. In this way, the light can be intelligently adjusted to make the scene where the person is located more in line with the actual lighting conditions, enhancing the sense of reality, and the adaptive adjustment of light can meet different shooting or rendering requirements, whether it is an indoor scene or an outdoor environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0023] Figure 1 It is a flowchart of the automatic light control method for a three-dimensional scene according to an embodiment of the present application.
[0024] Figure 2 It is a schematic structural diagram of the automatic light control method for a three-dimensional scene according to an embodiment of the present application.
[0025] Figure 3 It is a flowchart of inputting the enhanced feature map of the foreground part of the three-dimensional scene and the enhanced feature map of the background part of the three-dimensional scene into the foreground-background semantic joint interaction network to obtain the foreground-background semantic joint interaction feature map of the three-dimensional scene according to an embodiment of the present application.
[0026] Figure 4 It is a flowchart of training the foreground feature extractor based on the dilated convolutional neural network model, the background feature extractor based on the dilated convolutional neural network model, the feature multi-scale perception enhancement module, the foreground-background semantic joint interaction network, and the light parameter configurator based on the decoder according to an embodiment of the present application.
[0027] Figure 5 It is a system block diagram of the automatic light control system for a three-dimensional scene according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0029] It should be understood that the various steps recited in the method embodiments of the present disclosure may be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0030] In the description of the embodiments of the present disclosure, the terms "including" and its like should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0031] It should be noted that the modifications referring to "one" and "plural" in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless explicitly stated otherwise in the context, it should be understood as "one or more".
[0032] With the progress of technology, people have put forward higher standards for the visual realism of game scenes. Ray tracing technology, as a currently widely used rendering algorithm, has brought significant impacts to application fields such as 3D animation, virtual reality, and digital twins. However, traditional ray control methods have limitations. The light sources and shadow effects are often pre-set, which may affect the interactivity of the game environment and make the light and shadow effects appear less smooth and realistic. In addition, some traditional ray control technologies may only be applicable to specific types of game scenes and may perform poorly in other scenes, showing a lack of adaptability.
[0033] Therefore, in view of the above technical problems, the technical concept of the present application is to construct a panoramic image of a three-dimensional scene and use artificial intelligence-based image processing and analysis algorithms to perform feature analysis and extraction of the foreground features and background features of the panoramic image of the three-dimensional scene, so as to adaptively adjust the light intensity, light color, and light irradiation direction based on the semantic joint interaction features between the foreground and background of the three-dimensional scene. In this way, the light can be intelligently adjusted to make the scene where the character is located more in line with the actual lighting conditions, enhancing the sense of reality, and the adaptive adjustment of light can adapt to different shooting or rendering requirements, whether it is an indoor scene or an outdoor environment.
[0034] Figure 1The flowchart of the automatic light control method for a three-dimensional scene according to an embodiment of the present application. As Figure 1 shown, the automatic light control method for a three-dimensional scene according to an embodiment of the present application includes: S110, constructing a panoramic image of the three-dimensional scene; S120, extracting the foreground part and the background part of the panoramic image to obtain the foreground part of the three-dimensional scene and the background part of the three-dimensional scene; S130, respectively performing foreground feature extraction and background feature extraction on the foreground part of the three-dimensional scene and the background part of the three-dimensional scene to obtain a foreground feature map of the foreground part of the three-dimensional scene and a background feature map of the background part of the three-dimensional scene; S140, respectively inputting the foreground feature map of the foreground part of the three-dimensional scene and the background feature map of the background part of the three-dimensional scene into a feature multi-scale perception enhancement module to obtain an enhanced foreground feature map of the foreground part of the three-dimensional scene and an enhanced background feature map of the background part of the three-dimensional scene; S150, inputting the enhanced foreground feature map of the foreground part of the three-dimensional scene and the enhanced background feature map of the background part of the three-dimensional scene into a foreground-background semantic joint interaction network to obtain a foreground-background semantic joint interaction feature map of the three-dimensional scene as the foreground-background semantic joint interaction feature; S160, inputting the enhanced foreground feature map of the foreground part of the three-dimensional scene and the enhanced background feature map of the background part of the three-dimensional scene into a foreground-background semantic joint interaction network to obtain a foreground-background semantic joint interaction feature map of the three-dimensional scene as the foreground-background semantic joint interaction feature.
[0035] In step S110, a panoramic image of the three-dimensional scene is constructed. It should be understood that the panoramic image provides 360-degree omnidirectional image information of the three-dimensional scene, which can completely record every detail in the three-dimensional scene, including lighting conditions, object layout, spatial relationships, etc., providing rich data support for subsequent light control. Based on this, in the technical solution of the present application, it is necessary to first construct a panoramic image of the three-dimensional scene. Specifically, in a specific embodiment of the present application, an implementable way to construct a panoramic image of the three-dimensional scene can be to obtain multi-angle scene images around the center of the scene that needs to be light-controlled, and then use image stitching and stereo calibration techniques to form a complete panoramic image.
[0036] In step S120, the foreground part and the background part of the panoramic image are extracted to obtain the foreground part of the stereoscopic scene and the background part of the stereoscopic scene. Accordingly, considering that the panoramic image of the stereoscopic scene contains the perspective information of the entire stereoscopic scene within a 360-degree range, and the foreground part and the background part of the panoramic image usually have different lighting requirements. Specifically, the foreground part in the panoramic image refers to the area part that is closer to the observer in the scene and is usually located at the center of the image, while the background part refers to the area part that is farther from the observer in the scene and is usually located at the edge of the image or at a relatively long distance. That is, stronger light may be required for the foreground to highlight the details of the object, while soft light may be required for the background to create an atmosphere. Based on this, in order to make the foreground of the stereoscopic scene and the background of the stereoscopic scene more coordinated and unified in terms of light, color, and illumination direction, in the technical solution of this application, the foreground part and the background part of the panoramic image are extracted to obtain the foreground part of the stereoscopic scene and the background part of the stereoscopic scene. In particular, in a specific embodiment of this application, the foreground part and the background part of the panoramic image can be extracted by image segmentation technology, so as to obtain the foreground part of the stereoscopic scene and the background part of the stereoscopic scene.
[0037] In step S130, foreground feature extraction and background feature extraction are respectively performed on the foreground part and the background part of the stereoscopic scene to obtain a foreground feature map of the stereoscopic scene foreground part and a background feature map of the stereoscopic scene background part. Specifically, in the embodiments of the present application, performing foreground feature extraction and background feature extraction on the foreground part and the background part of the stereoscopic scene respectively to obtain a foreground feature map of the stereoscopic scene foreground part and a background feature map of the stereoscopic scene background part includes: inputting the foreground part of the stereoscopic scene into a foreground feature extractor based on an atrous convolutional neural network model to obtain the foreground feature map of the stereoscopic scene foreground part; inputting the background part of the stereoscopic scene into a background feature extractor based on an atrous convolutional neural network model to obtain the background feature map of the stereoscopic scene background part. Correspondingly, considering that the foreground part of the stereoscopic scene contains important implicit feature information about the stereoscopic scene foreground, and the atrous convolutional neural network model can expand the receptive field without increasing the number of parameters by introducing the atrous convolution operation, capturing deeper feature information, which is particularly important for processing the foreground part of the stereoscopic scene. Therefore, in the technical solution of the present application, the foreground part of the stereoscopic scene is input into a foreground feature extractor based on an atrous convolutional neural network model to extract and capture the deeper foreground feature information hidden in the foreground part of the stereoscopic scene, thereby obtaining the foreground feature map of the stereoscopic scene foreground part. Similarly, considering that the background part of the stereoscopic scene also reflects the implicit feature information of the background part, based on this, in order to more accurately extract and mine the deep implicit background feature information of the background part of the stereoscopic scene, in the technical solution of the present application, the background part of the stereoscopic scene is input into a background feature extractor based on an atrous convolutional neural network model to obtain the background feature map of the stereoscopic scene background part.
[0038] In step S140, the foreground feature map of the stereoscopic scene foreground part and the background feature map of the stereoscopic scene background part are respectively input into a feature multi-scale perception enhancement module to obtain an enhanced foreground feature map of the stereoscopic scene foreground part and an enhanced background feature map of the stereoscopic scene background part. It should be understood that considering that both the foreground feature map of the stereoscopic scene foreground part and the background feature map of the stereoscopic scene background part express local detail feature information and overall global feature information about the stereoscopic scene foreground and background. Therefore, in order to more finely understand and analyze the detail and overall feature information expressed by different local regions in the foreground feature map of the stereoscopic scene foreground part and the background feature map of the stereoscopic scene background part to obtain a comprehensive and detailed feature representation, in the technical solution of the present application, the foreground feature map of the stereoscopic scene foreground part and the background feature map of the stereoscopic scene background part are respectively input into a feature multi-scale perception enhancement module to obtain an enhanced foreground feature map of the stereoscopic scene foreground part and an enhanced background feature map of the stereoscopic scene background part.
[0039] It is worth mentioning that the feature multi-scale perception enhancement module performs feature processing on the input feature map through three different branches to capture and extract local fine-grained information and overall global information in the input feature map respectively, so as to perform feature enhancement processing on the input feature map. Specifically, in the first branch, local fine-grained feature information about the foreground part is extracted by performing channel compression, global average pooling, and non-linear activation processing on the foreground part feature map of the three-dimensional scene. In the second branch, point convolution processing is performed on the foreground part feature map of the three-dimensional scene to retain features useful for specific tasks, making the foreground part features more compact and efficient. In the third branch, dilated convolution encoding is performed on the foreground part feature map of the three-dimensional scene to expand the global perception information of the foreground part features. Then, the features obtained from the first branch and the third branch are respectively fused and added to the features of the second branch, and then through the expansion of the receptive field, the important details and overall structural feature information in the foreground part feature map of the three-dimensional scene are enhanced, so as to obtain an enhanced foreground part feature map of the three-dimensional scene containing rich semantic information. In particular, the processing method of the background part feature map of the three-dimensional scene is consistent with the processing method of the foreground part feature map of the three-dimensional scene.
[0040] Specifically, in the embodiments of the present application, inputting the foreground part feature map of the three-dimensional scene and the background part feature map of the three-dimensional scene into the feature multi-scale perception enhancement module respectively to obtain an enhanced foreground part feature map of the three-dimensional scene and an enhanced background part feature map of the three-dimensional scene includes: in the first branch, performing first-channel feature extraction on the foreground part feature map of the three-dimensional scene to obtain a first-channel local activation feature vector of the foreground part of the three-dimensional scene; in the second branch, performing point convolution processing on the foreground part feature map of the three-dimensional scene to obtain a second-channel compressed foreground part feature map of the three-dimensional scene; in the third branch, performing receptive field expansion on the foreground part feature map of the three-dimensional scene to obtain a receptive field-expanded global activation feature matrix of the foreground part of the three-dimensional scene; multiplying the receptive field-expanded global activation feature matrix of the foreground part of the three-dimensional scene with the corresponding feature matrices along the channel dimension of the second-channel compressed foreground part feature map of the three-dimensional scene by position to obtain a second-channel compressed global activation feature map of the foreground part of the three-dimensional scene; multiplying the first-channel local activation feature vector of the foreground part of the three-dimensional scene with the feature matrices along the channel dimension of the second-channel compressed foreground part feature map of the three-dimensional scene by position to obtain a second-channel compressed local activation feature map of the foreground part of the three-dimensional scene; adding the second-channel compressed global activation feature map of the foreground part of the three-dimensional scene and the second-channel compressed local activation feature map of the foreground part of the three-dimensional scene by position to obtain a second-channel compressed multi-scale fusion activation feature map of the foreground part of the three-dimensional scene; performing dilated convolution encoding on the second-channel compressed multi-scale fusion activation feature map of the foreground part of the three-dimensional scene to obtain the enhanced foreground part feature map of the three-dimensional scene.
[0041] More specifically, in the embodiments of the present application, in the first branch, first-channel feature extraction is performed on the foreground part feature map of the stereo scene to obtain a first-channel local activation feature vector of the foreground part of the stereo scene, including: performing point convolution processing on the foreground part feature map of the stereo scene to obtain a first-channel compressed feature map of the foreground part of the stereo scene; performing global average pooling on each feature matrix along the channel dimension in the first-channel compressed feature map of the foreground part of the stereo scene to obtain a first-channel compressed feature vector of the foreground part of the stereo scene; and performing non-linear activation on the first-channel compressed feature vector of the foreground part of the stereo scene to obtain the first-channel local activation feature vector of the foreground part of the stereo scene.
[0042] More specifically, in the embodiments of the present application, in the third branch, receptive field expansion is performed on the foreground part feature map of the stereo scene to obtain a globally activated feature matrix of the receptive field expansion of the foreground part of the stereo scene, including: performing dilated convolution encoding on the foreground part feature map of the stereo scene to obtain a feature map of the receptive field expansion of the foreground part of the stereo scene; performing point convolution processing on the feature map of the receptive field expansion of the foreground part of the stereo scene to obtain a globally feature matrix of the receptive field expansion of the foreground part of the stereo scene; and performing non-linear activation on the globally feature matrix of the receptive field expansion of the foreground part of the stereo scene to obtain the globally activated feature matrix of the receptive field expansion of the foreground part of the stereo scene.
[0043] In the embodiments of the present application, specifically, the foreground part feature map of the stereo scene and the background part feature map of the stereo scene are respectively input into the feature multi-scale perception enhancement module to obtain an enhanced foreground part feature map of the stereo scene and an enhanced background part feature map of the stereo scene, including: inputting the foreground part feature map of the stereo scene into the feature multi-scale perception enhancement module and processing it according to the following enhancement formula to obtain the enhanced foreground part feature map of the stereo scene; where the enhancement formula is:
[0044]
[0045] Where F represents the foreground part feature map of the stereo scene, C 1 (·) represents performing point convolution processing on the feature map, Avg(·) is performing global average pooling processing on each feature matrix along the channel dimension of the feature map, δ(·) is non-linear activation processing, C 3,2 and C 3,3 are respectively 3×3 dilated convolution operations with dilation rates of 3 and 2, is addition by position, is multiplication by position, F cRepresents the feature map of the foreground part of the enhanced three-dimensional scene. In particular, the encoding method of the feature map of the background part of the three-dimensional scene is consistent with the encoding method of the feature map of the foreground part of the three-dimensional scene.
[0046] In step S150, the feature map of the foreground part of the enhanced three-dimensional scene and the feature map of the background part of the enhanced three-dimensional scene are input into the foreground-background semantic joint interaction network to obtain the foreground-background semantic joint interaction feature map of the three-dimensional scene as the foreground-background semantic joint interaction feature. It should be understood that in order to promote the feature information exchange and interaction between the feature map of the foreground part of the enhanced three-dimensional scene and the feature map of the background part of the enhanced three-dimensional scene, so that the model can better understand the semantic interaction relationship between the features of the foreground part of the three-dimensional scene and the features of the background part of the three-dimensional scene, and thus for more accurate light adjustment in the subsequent process. In the technical solution of this application, the feature map of the foreground part of the enhanced three-dimensional scene and the feature map of the background part of the enhanced three-dimensional scene are input into the foreground-background semantic joint interaction network to obtain the foreground-background semantic joint interaction feature map of the three-dimensional scene. That is to say, the foreground-background semantic joint interaction network can effectively integrate the feature map of the foreground part of the enhanced three-dimensional scene and the feature map of the background part of the enhanced three-dimensional scene, and strengthen the complex semantic association and interaction between the foreground and background of the three-dimensional scene, so as to more accurately capture the overall semantic information of the three-dimensional scene, improve the model's understanding ability of the entire three-dimensional scene, and make the light control and visual effect optimization more accurate and effective.
[0047] Specifically, Figure 3 It is a flowchart of inputting the feature map of the foreground part of the enhanced three-dimensional scene and the feature map of the background part of the enhanced three-dimensional scene into the foreground-background semantic joint interaction network to obtain the foreground-background semantic joint interaction feature map of the three-dimensional scene in the automatic light control method for three-dimensional scenes according to the embodiments of this application. As Figure 3As shown, inputting the enhanced stereoscopic scene foreground partial feature map and the enhanced stereoscopic scene background partial feature map into the foreground-background semantic joint interaction network to obtain the stereoscopic scene foreground-background semantic joint interaction feature map includes: S151, performing shape transformation on the enhanced stereoscopic scene foreground partial feature map and the enhanced stereoscopic scene background partial feature map to obtain the enhanced stereoscopic scene foreground partial feature matrix and the enhanced stereoscopic scene background partial feature matrix; S152, calculating the matrix multiplication between the enhanced stereoscopic scene foreground partial feature matrix and the enhanced stereoscopic scene background partial feature matrix to obtain the enhanced stereoscopic scene foreground-background correlation feature matrix; S153, inputting the enhanced stereoscopic scene foreground-background correlation feature matrix into the Softmax function to obtain the enhanced stereoscopic scene foreground-background dependence relationship matrix, and then performing linear interpolation on the enhanced stereoscopic scene foreground-background dependence relationship matrix to obtain the enhanced stereoscopic scene foreground-background dependence relationship interaction matrix; S154, performing shape transformation on the feature matrix obtained by multiplying the enhanced stereoscopic scene foreground-background dependence relationship interaction matrix and the enhanced stereoscopic scene foreground partial feature matrix to obtain the stereoscopic scene foreground-background semantic joint interaction feature map.
[0048] In step S160, based on the stereoscopic scene foreground-background semantic joint interaction feature, obtain the light parameter decoding result, and the light parameter decoding result includes light intensity, light color, and light irradiation direction. Specifically, in the embodiment of the present application, obtaining the light parameter decoding result based on the stereoscopic scene foreground-background semantic joint interaction feature includes: inputting the stereoscopic scene foreground-background semantic joint interaction feature map into the light parameter configurator based on the decoder to obtain the light parameter decoding result. That is, perform decoding processing on the stereoscopic scene foreground-background semantic joint interaction feature obtained by performing semantic joint interaction using the enhanced stereoscopic scene foreground partial feature map and the enhanced stereoscopic scene background partial feature map, so as to adaptively adjust to obtain the light intensity, light color, and light irradiation direction. In this way, it is possible to intelligently adjust the light to make the scene where the person is located more in line with the actual lighting conditions, enhance the sense of reality, and the adaptive adjustment of the light can adapt to different shooting or rendering requirements, whether it is an indoor scene or an outdoor environment.
[0049] It is worth mentioning that those of ordinary skill in the art should be aware that before applying the deep neural network model for inference, the deep neural network model needs to be trained first so that the deep neural network can implement specific function capabilities.
[0050] Specifically, in the embodiments of the present application, a training step is further included: for training the foreground feature extractor based on the dilated convolutional neural network model, the background feature extractor based on the dilated convolutional neural network model, the feature multi-scale perception enhancement module, the foreground-background semantic joint interaction network, and the ray parameter configurator based on the decoder.
[0051] Figure 4 It is a flowchart of the training step in the automatic ray regulation method for a three-dimensional scene according to the embodiments of the present application. As Figure 4 shown, the training step includes: S210, obtaining training data, where the training data includes training panoramic images for constructing a three-dimensional scene and real ray parameter decoding results, and the real ray parameter decoding results are the true values of ray intensity, ray color, and ray illumination direction; S220, extracting the foreground part and the background part of the training panoramic image to obtain the foreground part of the training three-dimensional scene and the background part of the training three-dimensional scene; S230, inputting the foreground part of the training three-dimensional scene into the foreground feature extractor based on the dilated convolutional neural network model to obtain a feature map of the foreground part of the training three-dimensional scene; S240, inputting the background part of the training three-dimensional scene into the background feature extractor based on the dilated convolutional neural network model to obtain a feature map of the background part of the training three-dimensional scene; S250, respectively inputting the feature map of the foreground part of the training three-dimensional scene and the feature map of the background part of the training three-dimensional scene into the feature multi-scale perception enhancement module to obtain a feature map of the enhanced foreground part of the training three-dimensional scene and a feature map of the enhanced background part of the training three-dimensional scene; S260, inputting the feature map of the enhanced foreground part of the training three-dimensional scene and the feature map of the enhanced background part of the training three-dimensional scene into the foreground-background semantic joint interaction network to obtain a foreground-background semantic joint interaction feature map of the training three-dimensional scene; S270, inputting the foreground-background semantic joint interaction feature map of the training three-dimensional scene into the ray parameter configurator based on the decoder to obtain a training ray parameter decoding result; S280, calculating the cross-entropy loss function value between the training ray parameter decoding result and the real ray parameter decoding result to obtain a decoding loss function value; S290, training the foreground feature extractor based on the dilated convolutional neural network model, the background feature extractor based on the dilated convolutional neural network model, the feature multi-scale perception enhancement module, the foreground-background semantic joint interaction network, and the ray parameter configurator based on the decoder based on the decoding loss function value through backpropagation of gradient descent.
[0052] It should be understood that here, the foreground part feature map of the training enhanced stereo scene and the background part feature map of the training enhanced stereo scene respectively represent the multi-scale perception enhanced image semantic features of the image semantic space distribution of the foreground part and the background part of the training panoramic image. In this way, after inputting the foreground part feature map of the training enhanced stereo scene and the background part feature map of the training enhanced stereo scene into the foreground-background semantic joint interaction network, the foreground-background semantic joint interaction feature map of the training stereo scene will also have unevenness in feature aggregation decoding regression due to the difference in the foreground-background semantic interaction weight distribution, and the probability density distribution regression convergence logic of the intra-class features and inter-class features after the clustering operation is inconsistent.
[0053] Based on this, in this preferred embodiment, when training the foreground feature extractor based on the dilated convolutional neural network model, the background feature extractor based on the dilated convolutional neural network model, the feature multi-scale perception enhancement module, the foreground-background semantic joint interaction network, and the ray parameter configurator based on the decoder according to the classification loss function value, wherein, in each iteration of the training, the foreground-background semantic joint interaction feature map of the training stereo scene is iteratively optimized.
[0054] Among them, the iterative optimization process includes: performing a clustering operation on the foreground-background semantic joint interaction feature map of the training stereo scene, for example, performing a clustering operation based on the distance between feature values, and determining the number of feature values within the clustering set obtained through the clustering operation, that is, the number of intra-class features, and subtracting the number of intra-class features from the total number of feature values of the foreground-background semantic joint interaction feature map of the training stereo scene to obtain the number of inter-class features; dividing the total number of feature values by the number of intra-class features and the number of inter-class features respectively to obtain the class importance value and the inter-class ratio value, and calculating the reciprocal of the class importance value to obtain the class constraint value; calculating the power function of each feature value of the foreground-background semantic joint interaction feature map of the training stereo scene with the class importance value as the exponent, adding it to the exponential value with the natural function as the base and the class constraint value as the exponent, and then multiplying by the inter-class ratio value to obtain the foreground-background semantic joint interaction modulation map of the training stereo scene; performing a dot multiplication of the foreground-background semantic joint interaction feature map of the training stereo scene with the class importance value to obtain the foreground-background semantic joint interaction ontology map of the training stereo scene; calculating the weighted sum of the foreground-background semantic joint interaction modulation map of the training stereo scene and the foreground-background semantic joint interaction ontology map of the training stereo scene with a weight hyperparameter to obtain the optimized foreground-background semantic joint interaction feature map of the training stereo scene.
[0055] Specifically, in this preferred embodiment, the training stereo scene foreground-background semantic joint interaction feature map is optimized to obtain an optimized training stereo scene foreground-background semantic joint interaction feature map, and the process is represented by the following formula:
[0056]
[0057] Where F is the training stereo scene foreground-background semantic joint interaction feature map, n is the total number of eigenvalue of the training stereo scene foreground-background semantic joint interaction feature map, k is the number of intra-class features of the training stereo scene foreground-background semantic joint interaction feature map, ⊙ is element-wise multiplication, β is the weight hyperparameter, is element-wise addition, represents performing an exponential operation on the eigenvalue at each position in the feature map, and F′ is the optimized training stereo scene foreground-background semantic joint interaction feature map.
[0058] Therefore, while performing a clustering operation on the training stereo scene foreground-background semantic joint interaction feature map, by combining the clustering importance measure of the eigenvalue of the training stereo scene foreground-background semantic joint interaction feature map with the class constraint measure of the clustering operation, further responsive modulation is performed with the out-of-class ratio factor, and based on the clustering scaling of the feature ontology representation of the training stereo scene foreground-background semantic joint interaction feature map, the feature purification simplicity and effectiveness of the clustering operation for the entire feature set of the training stereo scene foreground-background semantic joint interaction feature map are established, and the probability density distribution regression convergence logic inconsistency caused by the clustering operation is suppressed, so as to improve the decoding regression iteration effect of the feature set input of the training stereo scene foreground-background semantic joint interaction feature map based on the decoder's ray parameter configurator, that is, to improve the speed of decoding training and the accuracy of the decoding result. In this way, the light can be intelligently adjusted to make the scene more in line with the actual lighting conditions, enhancing the sense of reality, and the adaptive adjustment of light can adapt to different shooting or rendering requirements, whether it is an indoor scene or an outdoor environment.
[0059] In summary, the stereo scene automatic light control method based on the embodiments of the present application is elucidated. It constructs a panoramic image of the stereo scene and uses artificial intelligence-based image processing and analysis algorithms to perform feature analysis and extraction of the foreground features and background features of the panoramic image of the stereo scene, so as to adaptively adjust the light intensity, light color, and light irradiation direction based on the semantic joint interaction features between the foreground and background of the stereo scene. In this way, the light can be intelligently adjusted to make the scene where the person is located more in line with the actual lighting conditions, enhancing the sense of reality, and the adaptive adjustment of light can adapt to different shooting or rendering requirements, whether it is an indoor scene or an outdoor environment.
[0060] Figure 5 This is a system block diagram of a stereoscopic scene automatic light control system according to an embodiment of the present application. As Figure 5 shown, the stereoscopic scene automatic light control system 100 according to an embodiment of the present application includes: a panoramic image construction module 110 for constructing a panoramic image of the stereoscopic scene; a panoramic image segmentation module 120 for extracting the foreground part and the background part of the panoramic image to obtain a stereoscopic scene foreground part and a stereoscopic scene background part; a panoramic image segmentation part feature extraction module 130 for respectively performing foreground feature extraction and background feature extraction on the stereoscopic scene foreground part and the stereoscopic scene background part to obtain a stereoscopic scene foreground part feature map and a stereoscopic scene background part feature map; a panoramic image segmentation part feature enhancement module 140 for respectively inputting the stereoscopic scene foreground part feature map and the stereoscopic scene background part feature map into a feature multi-scale perception enhancement module to obtain an enhanced stereoscopic scene foreground part feature map and an enhanced stereoscopic scene background part feature map; a panoramic image segmentation part semantic joint interaction module 150 for inputting the enhanced stereoscopic scene foreground part feature map and the enhanced stereoscopic scene background part feature map into a foreground-background semantic joint interaction network to obtain a stereoscopic scene foreground-background semantic joint interaction feature map as a stereoscopic scene foreground-background semantic joint interaction feature; a light parameter decoding result generation module 160 for obtaining a light parameter decoding result based on the stereoscopic scene foreground-background semantic joint interaction feature, where the light parameter decoding result includes light intensity, light color, and light irradiation direction.
[0061] Here, those skilled in the art can understand that the specific functions and operations of each unit and module in the above stereoscopic scene automatic light control system 100 have been described in detail above with reference to Figures 1 to 4 the description of the stereoscopic scene automatic light control method, and therefore, the repeated description thereof will be omitted.
[0062] In summary, the stereoscopic scene automatic light control system 100 based on the embodiment of the present application is clarified. It constructs a panoramic image of the stereoscopic scene and uses artificial intelligence-based image processing and analysis algorithms to perform feature analysis and extraction of the foreground features and background features of the panoramic image of the stereoscopic scene, so as to adaptively adjust the light intensity, light color, and light irradiation direction based on the semantic joint interaction features between the foreground and background of the stereoscopic scene. In this way, it can intelligently adjust the light to make the scene where the person is located more in line with the actual lighting conditions, enhance the sense of reality, and the adaptive adjustment of the light can meet different shooting or rendering requirements, whether it is an indoor scene or an outdoor environment.
[0063] As described above, the three-dimensional scene automatic light control system 100 according to the embodiments of the present application can be implemented in various wireless terminals, such as a server for three-dimensional scene automatic light control, etc. In one example, the three-dimensional scene automatic light control system 100 according to the embodiments of the present application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the three-dimensional scene automatic light control system 100 can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the three-dimensional scene automatic light control system 100 can also be one of the many hardware modules of the wireless terminal.
[0064] Alternatively, in another example, the three-dimensional scene automatic light control system 100 and the wireless terminal can also be separate devices, and the three-dimensional scene automatic light control system 100 can be connected to the wireless terminal through a wired and / or wireless network, and transmit and interact information in accordance with a predefined data format.
[0065] The above are only examples of the principles of the present disclosure, and those skilled in the art can make various modifications without departing from the scope of the present disclosure. The above embodiments are presented for illustrative purposes rather than limitations. The present disclosure can also take many forms other than those explicitly described herein.
Claims
1. A method for automatic light control of a stereoscopic scene, characterized in that: include: Constructing a panoramic image of a stereoscopic scene; Extracting a foreground portion and a background portion of the panoramic image to obtain a stereoscopic scene foreground portion and a stereoscopic scene background portion; Performing foreground feature extraction and background feature extraction on the foreground part of the stereoscopic scene and the background part of the stereoscopic scene respectively to obtain a feature map of the foreground part of the stereoscopic scene and a feature map of the background part of the stereoscopic scene; Inputting the stereoscopic scene foreground feature map and the stereoscopic scene background feature map into a feature multi-scale perception enhancement module to obtain an enhanced stereoscopic scene foreground feature map and an enhanced stereoscopic scene background feature map; Inputting the enhanced stereoscopic scene foreground part feature map and the enhanced stereoscopic scene background part feature map into a foreground-background semantic joint interaction network to obtain a stereoscopic scene foreground-background semantic joint interaction feature map as a stereoscopic scene foreground-background semantic joint interaction feature; Based on the foreground-background semantic joint interaction feature of the stereoscopic scene, a light parameter decoding result is obtained, wherein the light parameter decoding result includes light intensity, light color and light irradiation direction; The method of inputting the stereoscopic scene foreground feature map and the stereoscopic scene background feature map into a feature multi-scale perception enhancement module to obtain an enhanced stereoscopic scene foreground feature map and an enhanced stereoscopic scene background feature map comprises: In the first branch, first channel feature extraction is performed on the stereoscopic scene foreground part feature map to obtain a first channel stereoscopic scene foreground part local activation feature vector; In the second branch, point convolution processing is performed on the stereoscopic scene foreground part feature map to obtain a second channel compressed stereoscopic scene foreground part feature map; In the third branch, the receptive field of the foreground part feature map of the stereoscopic scene is expanded to obtain a receptive field expansion global activation feature matrix of the foreground part of the stereoscopic scene; Multiplying the receptive field expansion global activation feature matrix of the stereoscopic scene foreground part by position with each corresponding feature matrix of the second channel compressed stereoscopic scene foreground part feature map along the channel dimension to obtain the second channel compressed stereoscopic scene foreground part global activation feature map; Multiplying the local activation feature vector of the foreground part of the stereoscopic scene of the first channel by each feature matrix along the channel dimension of the feature map of the foreground part of the compressed stereoscopic scene of the second channel by position to obtain the local activation feature map of the foreground part of the compressed stereoscopic scene of the second channel; Adding the second channel compressed stereo scene foreground part global activation feature map and the second channel compressed stereo scene foreground part local activation feature map according to position to obtain a second channel compressed stereo scene foreground part multi-scale fused activation feature map; Performing dilated convolution coding on the multi-scale fused activation feature map of the foreground part of the second channel compressed stereo scene to obtain the enhanced stereo scene foreground part feature map; Wherein, in the first branch, performing first channel feature extraction on the stereoscopic scene foreground part feature map to obtain a first channel stereoscopic scene foreground part local activation feature vector includes: Performing point convolution processing on the stereoscopic scene foreground part feature map to obtain a first channel stereoscopic scene foreground part compressed feature map; Performing global mean pooling on each feature matrix along the channel dimension in the first channel stereoscopic scene foreground part compression feature map to obtain a first channel stereoscopic scene foreground part compression feature vector; Performing nonlinear activation on the first channel stereoscopic scene foreground part compression feature vector to obtain the first channel stereoscopic scene foreground part local activation feature vector; Among them, in the third branch, the receptive field expansion is performed on the feature map of the foreground part of the stereoscopic scene to obtain a receptive field expansion global activation feature matrix of the foreground part of the stereoscopic scene, including: Performing dilated convolution coding on the feature map of the foreground part of the stereoscopic scene to obtain a feature map of the receptive field expansion of the foreground part of the stereoscopic scene; Performing point convolution processing on the receptive field expansion feature map of the foreground part of the stereoscopic scene to obtain a global feature matrix of the receptive field expansion of the foreground part of the stereoscopic scene; Performing nonlinear activation on the receptive field expansion global feature matrix of the foreground part of the stereoscopic scene to obtain the receptive field expansion global activation feature matrix of the foreground part of the stereoscopic scene; The method of inputting the enhanced stereoscopic scene foreground part feature map and the enhanced stereoscopic scene background part feature map into a foreground-background semantic joint interaction network to obtain a stereoscopic scene foreground-background semantic joint interaction feature map comprises: Performing shape transformation on the enhanced stereoscopic scene foreground part feature map and the enhanced stereoscopic scene background part feature map to obtain an enhanced stereoscopic scene foreground part feature matrix and an enhanced stereoscopic scene background part feature matrix; Calculating matrix multiplication between the enhanced stereoscopic scene foreground part feature matrix and the enhanced stereoscopic scene background part feature matrix to obtain an enhanced stereoscopic scene foreground-background correlation feature matrix; After inputting the enhanced stereo scene foreground and background correlation feature matrix into a Softmax function to obtain an enhanced stereo scene foreground and background dependency matrix, linear interpolation is performed on the enhanced stereo scene foreground and background dependency matrix to obtain an enhanced stereo scene foreground and background dependency interaction matrix; The feature matrix obtained by matrix multiplication of the enhanced stereoscopic scene foreground-background dependency interaction matrix and the enhanced stereoscopic scene foreground part feature matrix is transformed in shape to obtain the stereoscopic scene foreground-background semantic joint interaction feature map.
2. The method for automatic light control of a stereoscopic scene according to claim 1, characterized in that: The foreground feature extraction and the background feature extraction are respectively performed on the foreground part of the stereoscopic scene and the background part of the stereoscopic scene to obtain a feature map of the foreground part of the stereoscopic scene and a feature map of the background part of the stereoscopic scene, including: Inputting the foreground part of the stereoscopic scene into a foreground feature extractor based on a dilated convolutional neural network model to obtain a feature map of the foreground part of the stereoscopic scene; The background part of the stereoscopic scene is input into a background feature extractor based on a hole convolutional neural network model to obtain a feature map of the background part of the stereoscopic scene.
3. The method for automatic light control of a stereoscopic scene according to claim 2, characterized in that: Based on the stereoscopic scene foreground-background semantic joint interaction feature, a light parameter decoding result is obtained, including: inputting the stereoscopic scene foreground-background semantic joint interaction feature map into a decoder-based light parameter configurator to obtain the light parameter decoding result.
4. The method for automatic light control of a stereoscopic scene according to claim 3, characterized in that: It also includes a training step: for training the foreground feature extractor based on the hole convolutional neural network model, the background feature extractor based on the hole convolutional neural network model, the feature multi-scale perception enhancement module, the foreground-background semantic joint interaction network and the decoder-based light parameter configurator.
5. The method for automatic light control of a stereoscopic scene according to claim 4, characterized in that: The training step comprises: Acquire training data, where the training data includes a training panoramic image for constructing a stereoscopic scene, and a real light parameter decoding result, where the real light parameter decoding result is a real value of light intensity, light color, and light irradiation direction; Extracting the foreground portion and the background portion of the training panoramic image to obtain a training stereoscopic scene foreground portion and a training stereoscopic scene background portion; Inputting the foreground part of the training stereoscopic scene into the foreground feature extractor based on the hole convolutional neural network model to obtain a feature map of the foreground part of the training stereoscopic scene; Inputting the training stereoscopic scene background part into the background feature extractor based on the hole convolutional neural network model to obtain a training stereoscopic scene background part feature map; Inputting the training stereo scene foreground part feature map and the training stereo scene background part feature map into the feature multi-scale perception enhancement module to obtain the training enhanced stereo scene foreground part feature map and the training enhanced stereo scene background part feature map; Inputting the training enhanced stereo scene foreground part feature map and the training enhanced stereo scene background part feature map into the foreground-background semantic joint interaction network to obtain a training stereo scene foreground-background semantic joint interaction feature map; Inputting the training stereoscopic scene foreground-background semantic joint interaction feature map into the decoder-based light parameter configurator to obtain a training light parameter decoding result; Calculating a cross entropy loss function value between the training light parameter decoding result and the real light parameter decoding result to obtain a decoding loss function value; The foreground feature extractor based on the hole convolutional neural network model, the background feature extractor based on the hole convolutional neural network model, the feature multi-scale perception enhancement module, the foreground-background semantic joint interaction network and the decoder-based light parameter configurator are trained based on the decoding loss function value and through back propagation of gradient descent.
6. A stereoscopic scene automatic light control system, used to execute the stereoscopic scene automatic light control method according to claim 1, characterized in that: include: A panoramic image construction module, used to construct a panoramic image of a stereoscopic scene; A panoramic image segmentation module, used for extracting a foreground portion and a background portion of the panoramic image to obtain a stereoscopic scene foreground portion and a stereoscopic scene background portion; A panoramic image segmentation part feature extraction module, used for performing foreground feature extraction and background feature extraction on the stereoscopic scene foreground part and the stereoscopic scene background part respectively to obtain a stereoscopic scene foreground part feature map and a stereoscopic scene background part feature map; A panoramic image segmentation part feature enhancement module, used to input the stereoscopic scene foreground part feature map and the stereoscopic scene background part feature map into a feature multi-scale perception enhancement module to obtain an enhanced stereoscopic scene foreground part feature map and an enhanced stereoscopic scene background part feature map; A panoramic image segmentation part semantic joint interaction module, used for inputting the enhanced stereo scene foreground part feature map and the enhanced stereo scene background part feature map into a foreground-background semantic joint interaction network to obtain a stereo scene foreground-background semantic joint interaction feature map as a stereo scene foreground-background semantic joint interaction feature; The light parameter decoding result generating module is used to obtain the light parameter decoding result based on the foreground-background semantic joint interaction feature of the stereoscopic scene, and the light parameter decoding result includes light intensity, light color and light irradiation direction.
Citation Information
Patent Citations
Image semantic segmentation method and system based on semantic propagation and foreground and background perception
CN114494699A