A dark-light light field image enhancement system and method for opera scenes
By using a lighting prior and implicit neural representation processing module, combined with a spatial-angle Transformer network, the problems of uneven lighting and feature redundancy in low-light light field image enhancement are solved, achieving high-quality enhancement of Huangmei Opera light field images and improving the viewing experience in low-light environments.
Patent Information
- Application Number
- CN202510451722.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Existing low-light image enhancement methods struggle to effectively handle uneven illumination and feature redundancy in low-light environments, resulting in poor image quality for Huangmei Opera light fields and negatively impacting the viewing experience.
By employing an illumination prior processing module and an implicit neural representation processing module, combined with a spatial-angle Transformer network, the system adaptively adjusts brightness and restores illumination information through illumination prior processing and implicit neural representation techniques, thereby reducing redundant features and improving image clarity and detail fidelity.
It significantly improves image quality and detail preservation in low-light environments, restores natural lighting effects, and enhances the visibility and viewing experience of Huangmei Opera light field images. It is applicable to fields such as low-light image enhancement, medical image processing, video surveillance, and computer vision.
Smart Images

Figure CN120374477B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, specifically relating to a low-light field image enhancement system and method for opera scenes. Background Technology
[0002] Light field imaging technology can capture four-dimensional information of a scene, namely the spatial and angular information of light, which is crucial for achieving realistic 3D visual effects. For a long time, low-light light field image enhancement has been an important research direction in the field of image processing, aiming to improve the quality of light field images in low-light environments. However, existing traditional low-light light field image enhancement methods have certain limitations. For example, some methods rely too heavily on a single module for image feature extraction, which bottlenecks performance improvement and makes it difficult to effectively cope with complex low-light environments. Furthermore, traditional methods face significant challenges in improvement, as their performance is highly dependent on the level of each module, resulting in limited optimization space. Therefore, there is an urgent need for a more efficient enhancement method that can not only significantly improve image quality but also enhance the visual effect of the image, thereby optimizing the user experience.
[0003] However, most existing deep learning-based low-light image enhancement models only address the spatial dimension, removing redundancy between features while neglecting feature redundancy in dimensions such as angle, channel, and epiplanet. This limitation makes it difficult to further improve the compression performance of low-light images. This problem is particularly prominent in the digital presentation of Huangmei Opera. As a traditional Chinese opera art, Huangmei Opera emphasizes delicate performances and rich stage lighting variations. However, in low-light environments, the lighting effects on the opera stage are often insufficient, making it difficult to clearly present key details such as actors' facial expressions, costume textures, and stage scenery, affecting the audience's viewing experience. In particular, the unique costumes of Huangmei Opera, such as the water sleeves, long gowns, and headdresses, are prone to losing detail in low-light scenes, affecting the authentic reproduction of opera culture. Summary of the Invention
[0004] To address the problems existing in the prior art, this invention provides a low-light light field image enhancement system and method for opera scenes. By effectively improving the visibility and detail retention of light field images under low-light conditions, it solves the problems of poor image quality and high noise in Huangmei Opera light field images under low-light conditions, making traditional opera more vivid and realistic in the process of digital presentation.
[0005] A low-light light field image enhancement system for opera scenes includes:
[0006] The illumination prior processing module is used to perform illumination prior processing on the acquired dark-light Huangmei Opera light field image to obtain a brightened image;
[0007] An implicit neural representation processing module is used to perform implicit neural representation processing on the brightened image;
[0008] The first feature extraction module is used to extract the first global feature and the first local feature from the brightened image after implicit neural representation processing;
[0009] The enhanced feature acquisition module is used to stitch together the extracted first global feature and the first local feature and combine them with the dark light Huangmei Opera light field image to obtain enhanced features;
[0010] The enhanced image acquisition module is used to extract and stitch together the second global feature and the second local feature of the enhanced features to obtain the dark light field enhanced image of Huangmei Opera, thus completing the dark light field image enhancement for opera scenes.
[0011] Preferably, the illumination prior processing module includes:
[0012] The image segmentation unit is used to segment the sub-aperture image of the light field of Huangmei Opera in dark light to obtain block images of preset dimensions;
[0013] The mean extraction unit is used to perform illumination prior processing on the segmented image, extract the maximum, minimum and average values of each segmented image in the channel dimension, and obtain statistical information;
[0014] The channel stitching unit is used to perform channel-level stitching of the statistical information with the sub-aperture image of the dark light field of Huangmei Opera to obtain the brightened image.
[0015] Preferably, the implicit neural representation processing module includes:
[0016] The position encoding unit is used to perform position encoding on the spatial coordinates of the brightened image based on the Fourier position encoding method;
[0017] The feature extraction unit is used to extract local neighborhood features from the spatial coordinates after position encoding and the brightened image to obtain the spatial coordinate information of the brightened image.
[0018] The illumination enhancement unit is used to perform deep learning processing on the spatial coordinate information and the local neighborhood features using a multilayer perceptron to obtain a brightened image after implicit neural representation processing.
[0019] Preferably, the first feature extraction module includes:
[0020] The first global feature extraction unit is used to extract feature information of different angles and spatial positions in the brightened image after implicit neural representation processing based on space-angle transformation, and obtain the first global feature.
[0021] The first local feature extraction unit is used to perform a local convolution operation on the brightened image after implicit neural representation processing to extract the disparity information features of the brightened image after implicit neural representation processing and obtain the first global feature.
[0022] Preferably, the enhanced feature acquisition module includes:
[0023] The first feature splicing unit is used to splice the first local feature with the first global feature to obtain the first spliced feature;
[0024] An attention processing unit is used to adjust the feature channel weights of the first spliced feature based on the SE attention module to obtain the first spliced feature after channel weight adjustment.
[0025] The first convolutional unit is used to perform 2D convolution on the first spliced feature after channel weight adjustment;
[0026] The feature combining unit is used to combine the first stitched feature after 2D convolution with the dark-light Huangmei Opera light field image to obtain the enhanced feature.
[0027] Preferably, the enhanced image acquisition module includes:
[0028] The second convolutional unit is used to perform 2D convolution on the enhanced features;
[0029] The second feature extraction unit is used to extract the second global feature and the second local feature of the enhanced features after 2D convolution based on the low-light enhancement module;
[0030] The second feature splicing unit is used to splice the second global feature and the second local feature to obtain the second spliced feature, and to perform 2D convolution on the second spliced feature;
[0031] An enhanced image acquisition unit is used to fuse the second stitched feature after 2D convolution with the dark-light Huangmei Opera light field image to obtain the enhanced image.
[0032] This invention also provides a method for enhancing low-light light field images for opera scenes, the system comprising:
[0033] Illumination prior processing was performed on the acquired dark-light Huangmei Opera light field image to obtain a brightened image;
[0034] The brightened image is subjected to implicit neural representation processing;
[0035] First global features and first local features are extracted from the brightened image after implicit neural representation processing;
[0036] The extracted first global features and first local features are spliced together and combined with the dark-light Huangmei Opera light field image to obtain enhanced features;
[0037] The second global feature and the second local feature of the enhancement features are extracted and stitched together to obtain the dark light field enhancement image of Huangmei Opera, thus completing the dark light field image enhancement for opera scenes.
[0038] Preferably, the method for obtaining the brightened image includes:
[0039] Segment the sub-aperture image of the light field of Huangmei Opera in dark light to obtain block images of a preset dimension;
[0040] The image blocks are processed with illumination priors to extract the maximum, minimum, and average values of each image block in the channel dimension, thereby obtaining statistical information.
[0041] The statistical information is then stitched together with the sub-aperture image of the dark-light Huangmei Opera light field at the channel level to obtain the brightened image.
[0042] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention introduces illumination priors, providing more precise illumination adjustment for image enhancement. Through illumination priors, the model can automatically infer and recover illumination information during the enhancement process, solving the dilemma that traditional methods cannot effectively handle uneven illumination in low-light environments. Illumination priors can adaptively adjust brightness and contrast during image enhancement, resulting in a more natural and realistic illumination effect in the enhanced image, avoiding problems such as over-enhancement or brightness distortion. Simultaneously, it introduces implicit neural representation (INR) technology, which can represent complex geometric structures and details in an image in a continuous manner, avoiding the spatial limitations of traditional representation methods, thus providing more accurate feature representation, especially effectively recovering imperceptible details in low-light environments. Secondly, implicit representation reduces redundant information through low-dimensional mapping, compressing unnecessary noise and features, improving image clarity and processing efficiency. Furthermore, this technology has strong robustness, capable of handling problems such as uneven illumination, effectively restoring image brightness and contrast. Because implicit neural representation can handle light field image features from different viewpoints, it ensures geometric consistency and detail fidelity of the image, resulting in more precise enhancement effects. Meanwhile, implicit neural representations have strong scalability and generalization ability, and can be applied to image enhancement tasks in Huangmei Opera light field images and other low-light environments, thus promoting the further development of image enhancement technology.
[0043] By combining a spatial angle Transformer network and illumination priors, this invention can effectively recover illumination information under low-light conditions while preserving spatial details of the image, thereby achieving a more refined and realistic image enhancement effect. Overall, this invention not only solves the problem of light field image enhancement in low-light environments but also significantly improves image details and quality, showing broad application prospects, especially in low-light image enhancement, computer vision, and film and television production. Furthermore, by improving the quality, enhancing details and colors of low-light Huangmei Opera light field images, and restoring natural visual effects, this invention helps Huangmei Opera achieve a clearer, more vivid, and layered performance in low-light environments, enhancing the audience's viewing experience and the effect of cultural inheritance. Attached Figure Description
[0044] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is the overall process of the low-light light field image enhancement system for opera scenes according to an embodiment of the present invention;
[0046] Figure 2 This is an overall framework diagram of a low-light field image enhancement system for opera scenes according to an embodiment of the present invention.
[0047] Figure 3 This is a schematic diagram of the implicit neural representation processing module (INR) according to an embodiment of the present invention;
[0048] Figure 4 This is a schematic diagram of the Angular Transformer module structure according to an embodiment of the present invention;
[0049] Figure 5 This is a schematic diagram of the Spatial Transformer module structure according to an embodiment of the present invention;
[0050] Figure 6 This is a schematic diagram of the Local Detail Feature Extraction (LDFEM) module structure according to an embodiment of the present invention;
[0051] Figure 7 This is a schematic diagram of the low-light enhancement module (GLFEM) according to an embodiment of the present invention. Detailed Implementation
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0054] Example 1
[0055] like Figure 1 , Figure 2 As shown, a low-light field image enhancement system for opera scenes includes: an illumination prior processing module, an implicit neural representation processing module, a first feature extraction module, an enhancement feature acquisition module, and an enhancement image acquisition module.
[0056] The illumination prior processing module is used to perform illumination prior processing on the acquired dark-light Huangmei Opera light field image (x) to obtain a brightened image. The original light field image can be represented as L(u, v, h, w, c), where (u, v) represents the angular dimension (sub-aperture view) of the light field image, (h, w) is the spatial dimension, and c is the number of color channels (RGB channels). In order to enable the image to be processed by the network and reduce the computational cost during model training and inference, the image data is first preprocessed to crop or adjust the dark-light light field image to the input size required by the network. The input dimension of the processed dark-light light field image (x') is (5, 5, 3, 32, 32);
[0057] A further implementation method includes a light prior processing module comprising:
[0058] The image segmentation unit is used to segment the sub-aperture image of the dark light field of Huangmei Opera to obtain a patch image of a preset dimension. In this embodiment, the dimension of each patch is (batch_size, 5, 5, 3, 32, 32), and illumination prior processing is performed on each small patch.
[0059] The mean extraction unit performs illumination prior processing on the segmented image, extracting the maximum, minimum, and average values of each pixel in each segmented image across the channel dimensions to obtain statistical information. The batch value is 1, and the first and second dimensions (5) represent the angular resolution, reflecting the image's brightness, contrast, and detail. The maximum value represents the highlight areas, the minimum value represents the shadow areas, and the average value reflects the overall brightness. Based on this statistical information, illumination enhancement processing such as brightness enhancement, contrast adjustment, and non-linear mapping is performed to improve the image's dark detail and overall sharpness.
[0060] The channel stitching unit is used to perform channel-level stitching (cat) with the sub-aperture image of the dark-light Huangmei Opera light field to obtain a brightened image. After processing, the resulting feature matrix has dimensions of (batch_size, 5, 5, 5, 32, 32), where the third dimension of 5 represents the two channels corresponding to the maximum and average values, as well as the channels of the original image, ensuring that the multidimensional information of the image is fully preserved during processing.
[0061] The implicit neural representation processing module is used to perform implicit neural representation processing on the brightened image.
[0062] A further implementation method is, such as Figure 3 As shown, the implicit neural representation (INR) processing module includes:
[0063] The positional encoding unit is used to encode the spatial coordinates of the brightened image based on the Fourier positional encoding method, so that the coordinates of each pixel have a unique representation in high-dimensional space. In this embodiment, the input dimension of the fed feature (x') is (1, 5, 160, 160). The implicit neural representation first performs positional encoding on the input image, which uses Fourier positional encoding, that is, performing sine and cosine transformations on the input coordinates and mapping them using a series of different frequencies to obtain a higher-dimensional representation:
[0064] γ(x)=[sin(2 l πx),cos(2 l πx)],l=0,1,…,L-1,
[0065] Specifically, the input coordinates are subjected to sine and cosine transformations and mapped using a series of frequencies 2l (where l = 0, 1, ..., L-1) to obtain a higher-dimensional representation. The advantage of this positional encoding is that it enhances the model's ability to represent high-frequency information, effectively alleviating the problem of excessive smoothing of input coordinates in MLP structures (i.e., difficulty in learning high-frequency details), thereby improving performance in interpolation and super-resolution tasks. Furthermore, Fourier positional encoding enhances the model's coordinate awareness, making it easier to learn complex function mappings and improving the model's generalization and fidelity.
[0066] The feature extraction unit extracts local neighborhood features from the position-encoded spatial coordinates and the brightened image to obtain the spatial coordinate information of the brightened image. It creates all coordinate points (i.e., the spatial positions of pixels) in the image grid. These spatial coordinates are further processed along with image features to capture spatial relationships. In the low-light enhancement task, this method uses local neighborhood features to enable the model to focus on local contrast changes and combines this with globally normalized coordinate information to ensure spatial consistency of the enhancement effect. Fourier position encoding improves the ability to capture high-frequency details, helping to recover texture and edge information in low-light areas and reducing blur and noise interference. Furthermore, stitching together relative coordinate information gives the model stronger spatial awareness, enabling it to adapt to different brightness distributions and achieve more refined enhancement effects.
[0067] The illumination enhancement unit employs a multilayer perceptron (MLP) to perform deep learning processing on spatial coordinate information and local neighborhood features, obtaining a brightened image after implicit neural representation processing. The MLP directly fits the coordinate-to-color mapping, without relying on a fixed convolutional kernel structure, giving it stronger generalization ability under different lighting conditions. This provides clearer and more natural enhancement results for visual tasks in low-light environments, ultimately outputting an image processed by implicit neural representations. Specifically, the MLP model consists of multiple fully connected layers, each using the ReLU activation function to increase non-linearity. Its calculation process is as follows: s = f θ (γ(x)). (x is the input spatial coordinate, γ(x) is the coordinate representation after Fourier encoding, f θ This is an MLP model, where s represents the enhanced image features output. The role of the MLP is to process the input features using deep learning, ultimately outputting 5 channel values for each location. This process effectively combines prior lighting information with the original low-light image, thereby improving image clarity and detail. Finally, the feature matrix after implicit neural representation has dimensions of (batch_size, 5, 160, 160).
[0068] The first feature extraction module is used to extract the first global and first local features from the brightened image after implicit neural representation processing. Specifically, after implicit neural representation processing, the obtained features are reshaped to (batch_size*u*v, 5, 32, 32) and then subjected to a 2D convolution operation to increase the number of image channels. The feature matrix dimension after the convolution operation is (batch_size*u*v, 64, 32, 32), and then the feature matrix dimension is reshaped to (batch_size, u, v, 64, 32, 32). Subsequently, global and local features are extracted from the processed features.
[0069] In this embodiment, a low-light enhancement module (GLFEM) is used for global and local feature extraction. For example... Figure 7 As shown, the low-light enhancement module consists of a Spatial Transformer, an Angular Transformer, an SE attention mechanism, and a Local Detail Feature Extraction (LDFEM) module in parallel.
[0070] A further implementation method wherein the first feature extraction module includes:
[0071] The first global feature extraction unit is used to extract feature information from different angles and spatial locations in the brightened image after implicit neural representation processing based on spatial-angular transformation, thereby obtaining the first global feature. The formula is as follows:
[0072] F SA =T SA (F INR )
[0073] F INR These are the highlighted image features after implicit neural representation processing;
[0074] T SA Represents a space-angle transformation operation;
[0075] F SA It is the extracted global light field feature.
[0076] Specifically, spatial-angle transformation extracts features from different angles and spatial locations in a light field image by transforming its angle and spatial dimensions, thereby capturing the image's global structural information. This process is implemented through convolutional operations (such as Conv_ang and Conv_spa) and other related modules, fusing the spatial and angular information of the light field image to obtain a richer global feature representation. In this embodiment, in the low-light enhancement task, SpatialTransform effectively improves the enhancement effect of the light field image through the collaborative work of local detail feature extraction, self-attention mechanism, and feedforward network. Figure 5 As shown, SpaTrans is similar to traditional Transformers. Each Spatial Transformer contains two LayerNorms, one Multi-Head Self-Attention (MHSA) mechanism, one Feed-Forward Network (FFN), and two residual connections. First, the image is flattened to extract local region features, which are then mapped using an MLP to capture local spatial structure information. Subsequently, Multi-Head Self-Attention (MHSA) is used to calculate global correlations. The calculation formula for the MHSA mechanism is as follows:
[0077]
[0078] Simultaneously, using `atten_mask` to limit the attention range allows for more precise information exchange between distant pixels, thereby improving contrast in dark areas. LayerNorm normalization and residual connections enhance training stability and improve the model's generalization ability. Finally, Conv3D is used to reconstruct the light field data, ensuring complete preservation of angular information. This method not only effectively suppresses noise but also enhances image details under complex lighting conditions, resulting in a clearer and more natural enhanced light field image. Figure 4As shown, the Angular Transformer is similar to the traditional Transformer. Each Angular Transformer contains two LayerNorms, one Multi-Head Self-Attention (MHSA) mechanism, one Feed-Forward Network (FFN), and two residual connections. However, unlike the traditional Transformer, the Angular Transformer focuses on the relationship between pixels at the same location across multiple angles. Therefore, it needs to transform the feature matrix into feature data with dimensions (32*32, 64, 5, 5) using the rearrange function, where (u, v) represents the angular dimension (sub-aperture view) of the light field image (5, 5). After the Angular Transformer performs angular feature enhancement extraction, the feature dimension is still (32*32, 64, 5, 5), and it is transformed again into feature data with dimensions (batch_size, 64, 32, 32) using the rearrange function. Figure 4 As shown, Angular Transform primarily utilizes a multi-head self-attention (MHSA) mechanism for feature extraction in light field images, focusing on the angular dimension. This module first converts the input light field image (buffer) into a flat "angle token" representation for easier processing and learning. Then, through layer normalization and MHSA, the model captures long-range dependencies between different angles, enhancing its understanding of angular features. Next, a feedforward neural network further extracts richer features. Finally, the model transforms the processed features back into the original spatial structure (SAI) form. The advantage of this method is its efficient integration of information from different angles and its significant improvement in feature expressiveness through self-attention, especially when processing complex light field images. It better captures the correlations between angles, thereby improving the model's performance in image processing tasks.
[0079] like Figure 6 As shown, the first local feature extraction unit is used to perform a local convolution operation on the brightened image after implicit neural representation processing to extract the disparity information features of the brightened image after implicit neural representation processing, thereby obtaining the first local feature. The local detail feature extraction module (LDFEM) performs a local convolution operation on the brightened image after implicit neural representation processing, as shown in the following formula:
[0080] F LDFEM =T LDFEM (FINR )
[0081] F INR These are the highlighted image features after implicit neural representation processing;
[0082] T LDFEM This represents the operation of extracting local detail features;
[0083] F LDFEM It is the extracted global light field feature.
[0084] In this embodiment, the LDFEM module focuses on local feature extraction and enhancement of the image, implementing various feature extraction operations, primarily applied to the extraction of spatial, angular, and EPI (parallax image) information from light field images. First, the model extracts features of different dimensions (spatial, angular, and EPI) through multiple convolutional layers (such as Conv_spa, Conv_ang, Conv_epi_h, and Conv_epi_v). Then, multiple enhancement modules (such as epi_boost, sa_boost, and Conv_mixray) further enhance the representational power of these features. By using multi-scale convolution and attention mechanisms (such as SEAttention), the model can adaptively select and enhance important features, effectively capturing the complex relationships between spatial, angular, and EPI information. Furthermore, through residual connections and multi-layer fusion, the model can integrate various feature information, thereby improving the final feature representation capability. This multi-dimensional, multi-scale, and enhanced feature extraction method effectively improves the performance of light field image processing tasks, especially in complex tasks such as low-light enhancement, providing more detailed and richer feature information, thus improving task performance.
[0085] The enhanced feature acquisition module is used to concatenate the extracted first global features and first local features, and combine them with the dark-light Huangmei Opera light field image to obtain enhanced features. The formula is as follows:
[0086] F final =Cat(F INR ,F LDFEM )
[0087] A further implementation method includes an enhanced feature acquisition module comprising:
[0088] The first feature concatenation unit is used to concatenate the first local feature with the first global feature to obtain the first concatenated feature. Specifically, the dimension of the processed feature matrix is (batch_size, 5, 5, 128, 32, 32). The concatenated feature is first reshaped into (batch_size*u*v, 128, 32, 32), and then it will be further processed by a Squeeze-and-Excitation (SE) attention module.
[0089] The attention processing unit is used to adjust the feature channel weights of the first stitched features based on the SE attention module, allowing the model to automatically learn and focus on more important features, thereby improving the model's sensitivity to key information. This helps the model to more accurately focus on important regions in the image when enhancing low-light images, improving the enhancement effect and obtaining the first stitched features after channel weight adjustment.
[0090] The first convolutional unit is used to perform 2D convolution on the first spliced feature after channel weight adjustment;
[0091] The feature combination unit combines the first concatenated features from the 2D convolution with the light field image of the dark-light Huangmei Opera to obtain enhanced features. Specifically, these features processed by the SE module are reduced by a 2D convolution to decrease the number of image channels. The resulting feature matrix has dimensions of (batch_size*u*v, 3, 32, 32), further controlling the computational complexity of the model while preserving effective feature information. It is then reshaped to (batch_size, 5, 5, 3, 32, 32). Finally, the processed features are compared with the original image to calculate the difference, yielding the enhanced feature F1 score. This enhanced feature F1 score reflects the difference between the model output and the denoised image, providing important feedback signals for subsequent optimization steps and helping the model further improve the image enhancement effect.
[0092] The enhanced image acquisition module is used to extract and stitch together the second global feature and the second local feature of the enhancement features to obtain the dark light field enhanced image of Huangmei Opera, thus completing the dark light field image enhancement for opera scenes.
[0093] A further embodiment of the embodiment includes: the enhanced image acquisition module comprising:
[0094] The second convolutional unit performs 2D convolution on the enhanced features. The purpose of this convolution is to transform the enhanced features F1 into a richer feature representation, allowing for better processing of image details in subsequent stages. Specifically, the obtained enhanced features F1 are fed into a 2D convolution to increase their channel count. The resulting feature matrix is (batch_size*u*v, 64, 32, 32), which is then reshaped to (batch_size, u, v, 64, 32, 32) and fed into the low-light enhancement module for further extraction of global and local features.
[0095] The second feature extraction unit is used to extract the second global and second local features of the enhanced features after 2D convolution based on the low-light enhancement module. Specifically, the processed feature matrix has dimensions (batch_size*u*v, 64, 32, 32), which is then reshaped to (batch_size, 5, 5, 64, 32, 32) and fed into the low-light enhancement module (GLFEM) for further extraction of global and local features. Here, the purpose of the 2D convolution operation is to transform the enhanced features F1 into a richer feature representation, facilitating better processing of image details in subsequent stages.
[0096] The second feature concatenation unit concatenates the second global feature and the second local feature to obtain the second concatenated feature, and then performs a 2D convolution on the second concatenated feature. Specifically, the convolutionally processed global feature is concatenated with the local feature to fuse more spatial and detailed information for more accurate image enhancement. The concatenated feature is then reshaped to (batch_size*u*v, 128, 32, 32) and fed into another 2D convolution to reduce the number of image channels.
[0097] The enhanced image acquisition unit is used to fuse the second concatenated features after 2D convolution with the dark-light Huangmei Opera light field image to obtain an enhanced image. Specifically, the processed feature matrix (batch_size*u*v, 64, 32, 32) controls the computational cost, making the final feature map more concise and efficient in its representation. This operation helps extract the most useful key information from the image and avoids interference from redundant features. Subsequently, this processed feature is used in the calculation along with the original image to generate the model's output. In this way, the model can fuse the original image and the enhanced features to generate an optimized enhanced image that represents the transformation process from a dark-light light field image to a light field image under normal lighting. Finally, a normal light field image (batch_size, 5, 3, 32, 32) is generated.
[0098] In summary, this invention innovatively introduces illumination priors to optimize illumination information inference in low-light environments, improving the model's robustness and detail recovery capabilities under complex lighting conditions. Addressing the unique stage lighting style of Huangmei Opera, such as the soft warm-toned lighting, the complex textures of opera costumes, and the actors' delicate facial expressions, this invention proposes an image enhancement framework based on implicit neural representations. Leveraging the powerful expressive capabilities of neural networks, it achieves high-quality end-to-end image enhancement. Furthermore, by combining a spatial angle Transformer network, it accurately processes the spatial and angular information of the light field image, ensuring that the geometric consistency and detail richness of the light field image are maintained during the enhancement process.
[0099] This invention can be widely applied in fields such as low-light image enhancement, medical image processing, video surveillance, and computer vision, particularly in the digital preservation and dissemination of traditional opera, such as optimizing the VR immersive experience of Huangmei Opera. By enhancing stage lighting effects, key visual elements in Huangmei Opera performances, such as water sleeves, headdresses, and costume patterns, become clearer, showcasing the unique artistic charm of the opera. Simultaneously, the enhanced light field imagery better restores the layering of the stage setting, allowing the opera performance to maintain its original lighting atmosphere in a digital environment. Even in low-light conditions, viewers can experience the realistic atmosphere of the stage and the actors' nuanced performances, thereby improving the quality of the digital presentation of Huangmei Opera and promoting the modern dissemination of traditional opera culture.
[0100] Example 2
[0101] This invention also provides a method for enhancing low-light light field images for opera scenes, and an application system comprising:
[0102] Illumination prior processing was performed on the acquired dark-light Huangmei Opera light field image to obtain a brightened image;
[0103] Implicit neural representation processing is applied to the brightened image;
[0104] First global features and first local features are extracted from the brightened image after implicit neural representation processing;
[0105] The extracted first global feature and first local feature are concatenated and combined with the dark light Huangmei Opera light field image to obtain enhanced features;
[0106] The second global feature and the second local feature of the enhancement feature are extracted and stitched together to obtain the dark light field enhancement image of Huangmei Opera, thus completing the dark light field image enhancement for opera scenes.
[0107] A further embodiment of the method for obtaining a brightened image includes:
[0108] Segment the sub-aperture image of the light field of Huangmei Opera in dark light to obtain block images of a preset dimension;
[0109] Illumination prior processing is performed on the segmented images to extract the maximum, minimum, and average values of each segmented image in the channel dimension, thereby obtaining statistical information;
[0110] The statistical information is combined with the sub-aperture image of the dark-light Huangmei Opera light field at the channel level to obtain the brightened image.
[0111] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A low-light light field image enhancement system for opera scenes, characterized in that, include: The illumination prior processing module is used to perform illumination prior processing on the acquired dark-light Huangmei Opera light field image to obtain a brightened image; An implicit neural representation processing module is used to perform implicit neural representation processing on the brightened image; The first feature extraction module is used to extract the first global feature and the first local feature from the brightened image after implicit neural representation processing; The enhanced feature acquisition module is used to stitch together the extracted first global feature and the first local feature and combine them with the dark light Huangmei Opera light field image to obtain enhanced features; The enhanced image acquisition module is used to extract and stitch together the second global feature and the second local feature of the enhancement features to obtain the dark light field enhancement image of Huangmei Opera, thus completing the dark light field image enhancement for opera scenes; The implicit neural representation processing module includes: The position encoding unit is used to perform position encoding on the spatial coordinates of the brightened image based on the Fourier position encoding method; The feature extraction unit is used to extract local neighborhood features from the spatial coordinates after position encoding and the brightened image to obtain the spatial coordinate information of the brightened image. The illumination enhancement unit is used to perform deep learning processing on the spatial coordinate information and the local neighborhood features using a multilayer perceptron to obtain a brightened image after implicit neural representation processing. The first feature extraction module includes: The first global feature extraction unit is used to extract feature information of different angles and spatial positions in the brightened image after implicit neural representation processing based on space-angle transformation, and obtain the first global feature. The first local feature extraction unit is used to perform a local convolution operation on the brightened image after implicit neural representation processing, extract the disparity information features of the brightened image after implicit neural representation processing, and obtain the first local feature. The enhanced feature acquisition module includes: The first feature splicing unit is used to splice the first local feature with the first global feature to obtain the first spliced feature; An attention processing unit is used to adjust the feature channel weights of the first spliced feature based on the SE attention module to obtain the first spliced feature after channel weight adjustment. The first convolutional unit is used to perform 2D convolution on the first spliced feature after channel weight adjustment; The feature combining unit is used to combine the first stitched feature after 2D convolution with the dark-light Huangmei Opera light field image to obtain the enhanced feature.
2. The system according to claim 1, characterized in that, The illumination prior processing module includes: The image segmentation unit is used to segment the sub-aperture image of the light field of Huangmei Opera in dark light to obtain block images of preset dimensions; The mean extraction unit is used to perform illumination prior processing on the segmented image, extract the maximum, minimum and average values of each segmented image in the channel dimension, and obtain statistical information; The channel stitching unit is used to perform channel-level stitching of the statistical information with the sub-aperture image of the dark light field of Huangmei Opera to obtain the brightened image.
3. The system according to claim 1, characterized in that, The enhanced image acquisition module includes: The second convolutional unit is used to perform 2D convolution on the enhanced features; The second feature extraction unit is used to extract the second global feature and the second local feature of the enhanced features after 2D convolution based on the low-light enhancement module; The second feature splicing unit is used to splice the second global feature and the second local feature to obtain the second spliced feature, and to perform 2D convolution on the second spliced feature; An enhanced image acquisition unit is used to fuse the second stitched feature after 2D convolution with the dark-light Huangmei Opera light field image to obtain the enhanced image.
4. A method for enhancing low-light light field images for opera scenes, using the system described in any one of claims 1-3, characterized in that, include: Illumination prior processing was performed on the acquired dark-light Huangmei Opera light field image to obtain a brightened image; The brightened image is subjected to implicit neural representation processing; First global features and first local features are extracted from the brightened image after implicit neural representation processing; The extracted first global features and first local features are spliced together and combined with the dark-light Huangmei Opera light field image to obtain enhanced features; The second global feature and the second local feature of the enhancement features are extracted and stitched together to obtain the dark light field enhancement image of Huangmei Opera, thus completing the dark light field image enhancement for opera scenes.
5. The method according to claim 4, characterized in that, The method for obtaining the brightened image includes: Segment the sub-aperture image of the light field of Huangmei Opera in dark light to obtain block images of a preset dimension; The image blocks are processed with illumination priors to extract the maximum, minimum, and average values of each image block in the channel dimension, thereby obtaining statistical information. The statistical information is then stitched together with the sub-aperture image of the dark-light Huangmei Opera light field at the channel level to obtain the brightened image.
Citation Information
Patent Citations
Dark light image processing method based on multi-stage joint enhancement mechanism
CN117391987A
Low-light image enhancement method and system based on global-local illumination perception
CN118195947A
Mine image three-dimensional reconstruction method based on neural radiation field
CN118279490A