Dark light field image enhancement system and method for opera scene
Through light prior processing and implicit neural representation processing, combined with the space-angle Transformer network, the problem of uneven light illumination in dark light field image enhancement is solved, and the high-quality enhancement of images in Huangmei Opera scenes is achieved, the visibility and detail retention of images are improved, and the viewing experience of the audience is improved.
Patent Information
- Application Number
- CN202510451722.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing dark light field image enhancement methods are difficult to effectively improve image quality in low-light environments, especially in Huangmei Opera performances. Uneven light makes it difficult to clearly present the actor's expressions, clothing textures and stage set details, affecting the viewing experience.
The lighting prior processing module and the implicit neural characterization processing module are adopted, combined with the space-angle Transformer network, and through the lighting prior processing and implicit neural characterization processing, global and local features are extracted and spliced, lighting information is restored and image details are enhanced.
Significantly improve image quality and detail retention under low light conditions, enhance image visibility and authenticity, enhance the digital presentation effect of Huangmei Opera, and improve the viewing experience of the audience.
Smart Images

Figure CN120374477A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a low-light light field image enhancement system and method for opera scenes. Background Art
[0002] Light field imaging technology can capture the four-dimensional information of a scene, namely the spatial and angular information of light, which is crucial for achieving a realistic three-dimensional visual effect. For a long time, low-light light field image enhancement has been an important research direction in the field of image processing, aiming to improve the quality of light field images in low-light environments. However, existing traditional low-light light field image enhancement methods have certain limitations. For example, some methods rely too much on a single module for image feature extraction, resulting in a bottleneck in performance improvement and difficulty in effectively dealing with complex low-light environments. In addition, traditional methods face great challenges in the improvement process, and their performance highly depends on the level of each module, leading to limited optimization space. Therefore, there is an urgent need for a more efficient enhancement method that can not only significantly improve image quality but also enhance the visual effect of images, thereby optimizing the user experience.
[0003] However, most existing deep learning-based low-light light field image enhancement models only start from the spatial dimension to remove redundant features between features, ignoring the feature redundancy between dimensions such as angles, channels, and epipolar planes. This limitation makes it difficult to further improve the compression performance of low-light light field images. In the digital presentation process of Huangmei Opera, this problem is particularly prominent. As a traditional Chinese opera art, Huangmei Opera pays attention to delicate performances and rich stage lighting changes. However, in low-light environments, the lighting effects on the opera stage are often insufficient, making it difficult to clearly present key details such as the facial expressions of actors, the texture of costumes, and the stage scenery, affecting the viewing experience of the audience. Especially the unique costumes of Huangmei Opera, such as water sleeves, long gowns, and headdresses, are prone to losing details in low-light scenes, affecting the true reproduction of opera culture. Summary of the Invention
[0004] To solve the problems existing in the prior art, the present invention provides a low-light light field image enhancement system and method for opera scenes, which can effectively improve the visibility and detail retention of light field images under low-light conditions, thereby solving the problems of poor quality and high noise of Huangmei Opera light field images in low-light environments, and making traditional operas more vivid and realistic in the digital presentation process.
[0005] A low-light light field image enhancement system for opera scenes includes:
[0006] A lighting prior processing module for performing lighting prior processing on the captured low-light Huangmei Opera light field image to obtain a brightened image;
[0007] An implicit neural representation processing module for performing implicit neural representation processing on the brightened image;
[0008] A first feature extraction module for extracting a first global feature and a first local feature from the brightened image after implicit neural representation processing;
[0009] An enhanced feature acquisition module for splicing the extracted first global feature and the first local feature and combining with the low-light Huangmei Opera light field image to obtain an enhanced feature;
[0010] An enhanced image acquisition module for extracting and splicing a second global feature and a second local feature of the enhanced feature to obtain a low-light light field enhanced image of Huangmei Opera, completing the enhancement of the low-light light field image for the opera scene.
[0011] Preferably, the illumination prior processing module includes:
[0012] An image segmentation unit for segmenting the sub-aperture image of the low-light Huangmei Opera light field to obtain a block image with a preset dimension;
[0013] An average value extraction unit for performing illumination prior processing on the block image, extracting the maximum value, minimum value, and average value of each block image in the channel dimension to obtain statistical information;
[0014] A channel splicing unit for performing channel-level splicing of the statistical information and the sub-aperture image of the low-light Huangmei Opera light field to obtain the brightened image.
[0015] Preferably, the implicit neural representation processing module includes:
[0016] A position encoding unit for performing position encoding on the spatial coordinates of the brightened image based on the Fourier position encoding method;
[0017] A feature extraction unit for performing local neighborhood feature extraction on the spatially encoded coordinates and the brightened image to obtain spatial coordinate information of the brightened image;
[0018] An illumination enhancement unit for performing deep learning processing on the spatial coordinate information and the local neighborhood features using a multi-layer perceptron to obtain the brightened image after implicit neural representation processing.
[0019] Preferably, the first feature extraction module includes:
[0020] A first global feature extraction unit for extracting feature information at different angles and spatial positions in the brightened image after implicit neural representation processing based on spatial-angle transformation to obtain the first global feature;
[0021] The first local feature extraction unit is used to perform local convolution operations on the brightened image after implicit neural representation processing, extract the parallax information features of the brightened image after implicit neural representation processing, and obtain the first global feature.
[0022] Preferably, the enhanced feature acquisition module includes:
[0023] The first feature splicing unit is used to splice the first local feature and the first global feature to obtain a first spliced feature;
[0024] The attention processing unit is used to adjust the feature channel weights of the first spliced feature based on the SE attention module to obtain the first spliced feature with adjusted channel weights;
[0025] The first convolution unit is used to perform 2D convolution on the first spliced feature with adjusted channel weights;
[0026] The feature combination unit is used to combine the first spliced feature after 2D convolution with the low-light Huangmei Opera light field image to obtain the enhanced feature.
[0027] Preferably, the enhanced image acquisition module includes:
[0028] The second convolution unit is used to perform 2D convolution on the enhanced feature;
[0029] The second feature extraction unit is used to extract the second global feature and the second local feature of the enhanced feature after 2D convolution based on the low-light enhancement module;
[0030] The second feature splicing unit is used to splice the second global feature and the second local feature to obtain a second spliced feature, and perform 2D convolution on the second spliced feature;
[0031] The enhanced image acquisition unit is used to perform fusion calculation on the second spliced feature after 2D convolution and the low-light Huangmei Opera light field image to obtain the enhanced image.
[0032] The present invention also provides a low-light light field image enhancement method for opera scenes, applying the system, including:
[0033] Perform illumination prior processing on the collected low-light Huangmei Opera light field image to obtain a brightened image;
[0034] Perform implicit neural representation processing on the brightened image;
[0035] Extract the first global feature and the first local feature from the brightened image after implicit neural representation processing;
[0036] Concatenate the extracted first global feature and the first local feature, and combine them with the low-light Huangmei Opera light field image to obtain an enhanced feature;
[0037] Extract and concatenate the second global feature and the second local feature of the enhanced feature to obtain a low-light light field enhanced image of Huangmei Opera, completing the low-light light field image enhancement for the opera scene.
[0038] Preferably, the method for obtaining the brightened image includes:
[0039] Segment the sub-aperture image of the low-light Huangmei Opera light field to obtain a block image with a preset dimension;
[0040] Perform illumination prior processing on the block image, and extract the maximum value, minimum value, and average value of each block image in the channel dimension to obtain statistical information;
[0041] Perform channel-level concatenation of the statistical information and the sub-aperture image of the low-light Huangmei Opera light field to obtain the brightened image.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention introduces illumination prior to provide more precise illumination adjustment for image enhancement. Through the illumination prior, the model can automatically infer and restore the illumination information during the enhancement process, solving the dilemma that traditional methods cannot effectively handle uneven illumination in low-light environments. The illumination prior can adaptively adjust the brightness and contrast during the image enhancement process, making the enhanced image have a more natural and realistic illumination effect, avoiding problems such as over-enhancement or brightness distortion. At the same time, the implicit neural representation (INR) technology is introduced. This method can represent the complex geometric structures and detailed information in the image in a continuous manner, avoiding the spatial limitations of traditional representation methods, thereby providing more precise feature expression, especially effectively restoring imperceptible details in low-light environments. Secondly, the implicit representation reduces redundant information through low-dimensional mapping, compresses unnecessary noise and features, and improves the clarity and processing efficiency of the image. In addition, this technology has strong robustness, can handle problems such as uneven illumination, and effectively restores the brightness and contrast of the image. Since the implicit neural representation can handle the light field image features under different perspectives, it ensures the geometric consistency and detail fidelity of the image, and the enhancement effect is more precise. At the same time, the implicit neural representation has strong scalability and generalization ability, and can be applied to the Huangmei Opera light field image and other image enhancement tasks in low-light environments, promoting the further development of image enhancement technology.
[0043] By combining a spatial angle Transformer network and a lighting prior, the present invention can effectively restore the lighting information under low-light conditions while retaining the spatial details of the image, thereby achieving a more delicate and realistic image enhancement effect. Generally speaking, the present invention not only solves the problem of light field image enhancement in low-light environments, but also significantly improves the image details and quality, and has broad application prospects especially in the fields of low-light image enhancement, computer vision, and film and television production. At the same time, the present invention restores the natural visual effect by improving the quality of low-light Huangmei Opera light field images, enhancing details and colors, helping Huangmei Opera achieve a clearer, more vivid and hierarchical performance in low-light environments, and enhancing the viewing experience of the audience and the effect of cultural inheritance. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] To more clearly illustrate the technical solutions of the present invention, the following briefly introduces the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0045] Figure 1 is the overall process of the low-light light field image enhancement system for the opera scene in the embodiment of the present invention;
[0046] Figure 2 is the overall framework diagram of the low-light light field image enhancement system for the opera scene in the embodiment of the present invention;
[0047] Figure 3 is the structural schematic diagram of the implicit neural representation processing module (INR) in the embodiment of the present invention;
[0048] Figure 4 is the structural schematic diagram of the Angular Transformer module in the embodiment of the present invention;
[0049] Figure 5 is the structural schematic diagram of the Spatial Transformer module in the embodiment of the present invention;
[0050] Figure 6 is the structural schematic diagram of the local detail feature extraction (LDFEM) module in the embodiment of the present invention;
[0051] Figure 7 is the structural schematic diagram of the low-light enhancement module (GLFEM) in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0054] Embodiment 1
[0055] As Figure 1 、 Figure 2 shown, a low-light light field image enhancement system for the opera scene includes: a light prior processing module, an implicit neural representation processing module, a first feature extraction module, an enhanced feature acquisition module, and an enhanced image acquisition module.
[0056] The light prior processing module is used to perform light prior processing on the collected low-light Huangmei Opera light field image (x) to obtain a brightened image. The original light field image can be expressed as L(u, v, h, w, c), where (u, v) represents the angular dimension (sub-aperture view) of the light field image, (h, w) is the spatial dimension, and c is the number of color channels (RGB channels). In order to enable the image to be processed by the network and reduce the computational complexity during model training and inference, the image data is first preprocessed to crop or adjust the low-light light field image to the input size required by the network. The input dimension of the processed low-light light field image (x') is (5, 5, 3, 32, 32);
[0057] A further implementation manner is that the light prior processing module includes:
[0058] An image segmentation unit for segmenting the sub-aperture images of the low-light Huangmei Opera light field to obtain block images (patches) of a preset dimension; in this embodiment, the dimension of each patch is (batch_size, 5, 5, 3, 32, 32), and light prior processing is performed on each small block.
[0059] The mean extraction unit is used to perform illumination prior processing on the segmented images, extract the maximum value, minimum value, and average value of each pixel point of each segmented image in the channel dimension, and obtain statistical information. Here, the value of batch is 1, and the 5 in the first and second dimensions is the angular resolution, which reflects the brightness, contrast, and details of the image. The maximum value represents the highlight area of the image, the minimum value represents the shadow part, and the average value reflects the overall brightness. Based on these statistical information, illumination enhancement processing such as brightness enhancement, contrast adjustment, and non-linear mapping is performed to improve the dark part details and overall clarity of the image.
[0060] The channel splicing unit is used to perform channel-level splicing (cat) on the statistical information and the sub-aperture image of the low-light Huangmei Opera light field to obtain a brightened image. After processing, the dimension of the obtained feature matrix is (batch_size, 5, 5, 5, 32, 32), where the 5 in the third dimension represents the two channels corresponding to the maximum value and the average value, as well as the channels of the original image, ensuring that the multi-dimensional information of the image is fully retained during the processing.
[0061] The implicit neural representation processing module is used to perform implicit neural representation processing on the brightened image.
[0062] A further implementation manner is that, as Figure 3 shown, the implicit neural representation (INR) processing module includes:
[0063] The position encoding unit is used to perform position encoding on the spatial coordinates of the brightened image based on the Fourier position encoding method, so that the coordinates of each pixel have a unique representation in the high-dimensional space. In this embodiment, the input dimension of the fed feature (x’) is (1, 5, 160, 160). The implicit neural representation first performs position encoding on the input image. Here, the Fourier Positional Encoding method is used, that is, sine (sin) and cosine (cos) transformations are performed on the input coordinates, and a series of different frequencies are used for mapping to obtain a higher-dimensional representation:
[0064] γ(x) = [sin(2 l πx), cos(2 l πx)], l = 0, 1, …, L - 1,
[0065] Specifically, the input coordinates are subjected to sine (sin) and cosine (cos) transformations and mapped using a series of frequencies 2l (where l = 0, 1, ..., L - 1) to obtain a higher-dimensional representation. The advantage of this positional encoding is that it can enhance the model's ability to express high-frequency information, effectively alleviate the problem of over-smoothing of the input coordinates by the MLP structure (i.e., difficulty in learning high-frequency details), and thus improve the performance of interpolation and super-resolution tasks. In addition, Fourier positional encoding can enhance the model's coordinate perception ability, making it easier to learn complex function mapping relationships and improving the generalization and fidelity of the model.
[0066] A feature extraction unit is used to perform local neighborhood feature extraction on the spatially encoded coordinates and the brightened image to obtain the spatial coordinate information of the brightened image, create all coordinate points in the image grid (i.e., the spatial positions of pixels), and these spatial coordinates will be further processed together with the features of the image to capture spatial relationships. In the low-light enhancement task, this method enables the model to focus on local contrast changes through local neighborhood features and combines global normalized coordinate information to ensure the spatial consistency of the enhancement effect. Fourier positional encoding improves the ability to capture high-frequency details, helps to restore the texture and edge information in low-light areas, and reduces blur and noise interference. In addition, concatenating relative coordinate information enables the model to have a stronger spatial perception ability and adapt to different brightness distributions, thus achieving a more refined enhancement effect.
[0067] A lighting enhancement unit is used to perform deep learning processing on the spatial coordinate information and local neighborhood features using a multi-layer perceptron to obtain the brightened image after implicit neural representation processing. The MLP directly fits the mapping from coordinates to colors and does not rely on a fixed convolutional kernel structure, making it have stronger generalization ability under different lighting conditions and providing clearer and more natural enhancement results for visual tasks in low-light environments. Finally, the image after implicit neural representation processing is output. Specifically, the MLP model consists of multiple fully connected layers, and the ReLU activation function is used in each layer to increase non-linearity. The calculation process is as follows: s = f θ (γ(x)). (x is the input spatial coordinate, γ(x) is the coordinate representation after Fourier encoding, f θ is the MLP model, and s is the enhanced image feature output). The role of the MLP is to perform deep learning processing on the input features and finally output 5-channel values for each position. This process can effectively combine the lighting prior information and the original low-light image, thereby improving the clarity and detail performance of the image. Finally, the dimension of the feature matrix after implicit neural representation is (batch_size, 5, 160, 160).
[0068] The first feature extraction module is used to extract the first global feature and the first local feature from the brightened image processed by the implicit neural representation. Specifically, after being processed by the implicit neural representation, the obtained features will be reshaped into (batch_size * u * v, 5, 32, 32), and then through a 2D convolution operation to increase the number of channels of the image. The dimension of the feature matrix after the convolution operation is (batch_size * u * v, 64, 32, 32), and then the dimension of the feature matrix is reshaped into (batch_size, u, v, 64, 32, 32). Subsequently, the processed features are used to extract global and local features.
[0069] In this embodiment, the global and local features are extracted by using the low-light enhancement module (GLFEM). As Figure 7 shown, the low-light enhancement module is composed of a Spatial Transformer, an Angular Transformer, an SE attention mechanism, and a local detail feature extraction (LDFEM) module in parallel.
[0070] A further implementation manner is that the first feature extraction module includes:
[0071] The first global feature extraction unit is used to extract the feature information of different angles and spatial positions in the brightened image processed by the implicit neural representation based on the Spatial-Angular Transform to obtain the first global feature. The formula is as follows:
[0072] F SA =T SA (F INR )
[0073] F INR is the feature of the brightened image processed by the implicit neural representation;
[0074] T SA represents the Spatial-Angular Transform operation;
[0075] F SA is the extracted global light field feature.
[0076] Specifically, the spatial-angle transformation captures the global structural information of the image by converting the angular and spatial dimensions of the light field image and extracting the features at different angles and spatial positions in the image. This process is implemented through convolutional operations (such as Conv_ang and Conv_spa) and other related modules, which can fuse the spatial and angular information of the light field image to obtain a more abundant global feature representation. In this embodiment, in the low-light enhancement task, SpatialTransform works in cooperation with local detail feature extraction, self-attention mechanism, and feed-forward network to effectively improve the enhancement effect of the light field image. As Figure 5 shown, SpaTrans is similar to the traditional Transformer. Each Spatial Transformer contains 2 LayerNorms, 1 multi-head self-attention mechanism (MHSA), 1 feed-forward neural network (FFN), and 2 residual connections. First, the picture is flattened to extract local region features and mapped through MLP to capture the spatial local structural information. Subsequently, multi-head self-attention (MHSA) is used to calculate the global correlation. For the MHSA mechanism, its calculation formula is as follows:
[0077]
[0078] At the same time, atten_mask is used to limit the attention range, making the information interaction between distant pixels more accurate, thereby enhancing the contrast of the low-light area. Through LayerNorm normalization and residual connection, the training stability is improved and the generalization ability of the model is enhanced. Finally, the light field data is restored through Conv3D to ensure that the angular information is completely retained. This method can not only effectively suppress noise but also enhance the image details in complex lighting environments, making the enhanced light field image clearer and more natural. As Figure 4As shown, the Angular Transformer is also similar to the traditional Transformer. Each Angular Transformer contains two LayerNorms, one multi-head self-attention mechanism (MHSA), one feed-forward network (FFN), and two residual connections. However, different from the traditional Transformer, the Angular Transformer focuses on the relationship between pixel points at the same position among multiple angles. Therefore, it is necessary to deform the feature matrix through the rearrange function into feature data with dimensions of (32*32, 64, 5, 5), where (u, v) being (5, 5) represents the angular dimension (sub-aperture view) of the light field image. After the angular feature enhancement extraction by the Angular Transformer, the feature dimension is still (32*32, 64, 5, 5), and then it is deformed through the rearrange function again into feature data with dimensions of (batch_size, 64, 32, 32). As Figure 4 As shown, the Angular Transform mainly extracts features through the multi-head self-attention (MHSA) mechanism for the angular dimension in the light field image. This module first converts the input light field image (buffer) into a flat "angular token" representation for easy processing and learning. Then, through layer normalization and the multi-head self-attention mechanism, the model can capture long-range dependencies between different angles and enhance the understanding of angular features. Next, the feed-forward neural network is used to further extract richer features. Finally, the model converts the processed features back to the original spatial structure (SAI) form. The advantage of this method is that it can efficiently integrate information from different angles and significantly improve the feature expression ability through the self-attention mechanism. Especially when dealing with complex light field images, it can better capture the correlation between angles, thereby improving the performance of the model in image processing tasks.
[0079] As Figure 6 As shown, the first local feature extraction unit is used to perform local convolution operations on the brightened image after implicit neural representation processing to extract the disparity information features of the brightened image after implicit neural representation processing and obtain the first local feature. The local convolution operation on the brightened image after implicit neural representation processing is performed through the local detail feature extraction module LDFEM module (Local Detail Feature ExtractionModule), and the formula is as follows:
[0080] F LDFEM = T LDFEM (FINR )
[0081] F INR is the brightened image feature after implicit neural representation processing;
[0082] T LDFEM represents the local detail feature extraction operation;
[0083] F LDFEM is the extracted global light field feature.
[0084] In this embodiment, the LDFEM module focuses on the local feature extraction and enhancement of images, implements a variety of feature extraction operations, and is mainly applied to the extraction of spatial, angular, and EPI (disparity image) information of light field images. First, the model extracts features of different dimensions (spatial, angular, EPI) through multiple convolutional layers (such as Conv_spa, Conv_ang, Conv_epi_h, Conv_epi_v) respectively, and then further improves the representation ability of the features through multiple enhancement modules (such as epi_boost, sa_boost, Conv_mixray). By using multi-scale convolution and attention mechanisms (such as SEAttention), the model can adaptively select and enhance important features, effectively capturing the complex relationships between space, angle, and EPI. In addition, through residual connections and multi-layer fusion, the model can integrate various feature information, thereby improving the final feature expression ability. This multi-dimensional, multi-scale, and enhanced feature extraction method can effectively improve the performance of light field image processing tasks, especially in complex tasks such as low-light enhancement, and can provide more detailed and rich feature information, thus improving the performance of the tasks.
[0085] An enhanced feature acquisition module is used to splice the extracted first global feature and the first local feature and combine them with the low-light Huangmei opera light field image to obtain enhanced features. The formula is as follows:
[0086] F final = Cat(F INR ,F LDFEM )
[0087] A further implementation manner is that the enhanced feature acquisition module includes:
[0088] The first feature concatenation unit is used to concatenate (concat) the first local feature and the first global feature to obtain the first concatenated feature. Specifically, the dimension of the processed feature matrix is (batch_size, 5, 5, 128, 32, 32). The concatenated feature is first reshaped into (batch_size * u * v, 128, 32, 32), and then it will be further processed by a Squeeze-and-Excitation (SE) attention module.
[0089] The attention processing unit is used to adjust the feature channel weights of the first concatenated feature based on the SE attention module, allowing the model to automatically learn and focus on more important features, thereby improving the sensitivity of the model to key information. When helping the model enhance low-light images, it can more accurately focus on important regions in the image, improve the enhancement effect, and obtain the first concatenated feature with adjusted channel weights.
[0090] The first convolution unit is used to perform 2D convolution on the first concatenated feature with adjusted channel weights.
[0091] The feature combination unit is used to combine the first concatenated feature after 2D convolution with the low-light Huangmei Opera light field image to obtain the enhanced feature. Specifically, these features processed by the SE module are reduced in the number of image channels through a 2D convolution. The dimension of the processed feature matrix is (batch_size * u * v, 3, 32, 32), which further controls the computational complexity of the model while retaining effective feature information. Then it is reshaped into (batch_size, 5, 5, 3, 32, 32). Finally, the processed feature is subjected to a difference calculation with the original image to obtain the enhanced feature F1. This enhanced feature F1 reflects the gap between the model output and the denoised image, and it provides an important feedback signal for subsequent optimization steps to help the model further improve the image enhancement effect.
[0092] The enhanced image acquisition module is used to extract and concatenate the second global feature and the second local feature of the enhanced feature to obtain the enhanced low-light light field image of Huangmei Opera, completing the enhancement of the low-light light field image for the opera scene.
[0093] A further implementation manner is that the enhanced image acquisition module includes:
[0094] The second convolutional unit is used to perform 2D convolution on the enhanced features. At this time, the purpose of the convolution operation is to convert the enhanced feature F1 into a richer feature representation, so as to better process the details in the image in the subsequent stage. Specifically, the obtained enhanced feature F1 is sent into the 2D convolution to increase its number of channels. The processed feature matrix is (batch_size*u*v, 64, 32, 32), and then it is reshaped into (batch_size, u, v, 64, 32, 32) and sent into the low-light enhancement module to further extract global features and local features.
[0095] The second feature extraction unit is used to extract the second global feature and the second local feature of the enhanced feature after 2D convolution based on the low-light enhancement module. Specifically, the dimension of the processed feature matrix is (batch_size*u*v, 64, 32, 32), and then it is reshaped into (batch_size, 5, 5, 64, 32, 32) and sent into the low-light enhancement module (GLFEM) to further extract global features and local features. At this time, the purpose of the 2D convolution operation is to convert the enhanced feature F1 into a richer feature representation, so as to better process the details in the image in the subsequent stage.
[0096] The second feature splicing unit is used to splice the second global feature and the second local feature to obtain the second spliced feature, and perform 2D convolution on the second spliced feature. Specifically, the globally and locally features after convolution processing are spliced together, so as to fuse more spatial information and detail information for more accurate image enhancement. The spliced feature is reshaped into (batch_size*u*v, 128, 32, 32) and sent into a 2D convolution again to reduce the number of channels of the image.
[0097] The enhanced image acquisition unit is used to fuse and calculate the second spliced feature after 2D convolution with the low-light Huangmei opera light field image to obtain the enhanced image. Specifically, the processed feature matrix is (batch_size*u*v, 64, 32, 32) to control the computational amount, so that the final feature map is more concise and has efficient expression ability. This operation helps to extract the most useful key information in the image and avoid the interference of redundant features. Subsequently, this processed feature participates in the calculation together with the original image, and finally generates the output of the model. In this way, the model can fuse the original image and the enhanced features to generate an optimized enhanced image, which represents the conversion process from the low-light light field image to the light field image under normal illumination. Finally, a normal light field image (batch_size, 5, 3, 32, 32) is generated.
[0098] In summary, the present invention innovatively introduces lighting priors to optimize the inference of lighting information in dark environments and improve the robustness and detail recovery capabilities of the model under complex lighting conditions. In view of the unique stage lighting style of Huangmei Opera, such as soft warm-toned lighting, complex textures of opera costumes, and delicate facial expressions of actors, the present invention proposes an image enhancement framework based on implicit neural representation, which uses the powerful expressive power of neural networks to achieve end-to-end high-quality image enhancement. In addition, combined with the spatial angle Transformer network, the spatial and angle information of the light field image is accurately processed to ensure that the geometric consistency and detail richness of the light field image are maintained during the enhancement process.
[0099] The present invention can be widely used in the fields of low-light image enhancement, medical image processing, video surveillance and computer vision, especially in the digital protection and dissemination of traditional operas, such as the optimization of VR immersive experience of Huangmei Opera. By enhancing the stage lighting effect, key visual elements such as sleeves, headdresses, and clothing patterns in Huangmei Opera performances are made clearer, showing the unique artistic charm of the opera. At the same time, the enhanced light field image can better restore the layering of the stage setting, so that the opera performance can maintain the original light and shadow atmosphere in a digital environment. Even if the audience watches in a low-light environment, they can feel the real atmosphere of the stage and the delicate performances of the actors, thereby improving the digital presentation quality of Huangmei Opera and promoting the modern dissemination of traditional opera culture.
[0100] Embodiment 2
[0101] The present invention also provides a dark light field image enhancement method for opera scenes and an application system, comprising:
[0102] Perform illumination prior processing on the collected dark-light Huangmei Opera light field images to obtain brightened images;
[0103] Implicit neural representation processing of brightened images;
[0104] Extracting a first global feature and a first local feature from the brightened image processed by the implicit neural representation;
[0105] The extracted first global feature and the first local feature are spliced and combined with the dark-light Huangmei Opera light field image to obtain an enhanced feature;
[0106] The second global feature and the second local feature of the enhanced feature are extracted and spliced to obtain the dark-light light field enhanced image of Huangmei Opera, thus completing the dark-light light field image enhancement for opera scenes.
[0107] A further embodiment is that the method for obtaining a brightened image comprises:
[0108] Segment the sub-aperture image of the dark light Huangmei Opera light field to obtain a block image of a preset dimension;
[0109] Perform illumination prior processing on the segmented image, extract the maximum value, minimum value, and average value of each segmented image in the channel dimension, and obtain statistical information;
[0110] Perform channel-level splicing on the statistical information and the sub-aperture image of the dim Huangmei Opera light field to obtain a brightened image.
[0111] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. An enhanced system for low-light light field images in an opera scene, characterized in that, Including: A light prior processing module, configured to perform light prior processing on the collected low-light Huangmei Opera light field image to obtain a brightened image; An implicit neural representation processing module, configured to perform implicit neural representation processing on the brightened image; A first feature extraction module, configured to extract a first global feature and a first local feature from the brightened image after implicit neural representation processing; An enhanced feature acquisition module, configured to splice the extracted first global feature and the first local feature and combine them with the low-light Huangmei Opera light field image to obtain an enhanced feature; An enhanced image acquisition module, configured to extract and splice a second global feature and a second local feature of the enhanced feature to obtain a low-light light field enhanced image of Huangmei Opera, completing the enhancement of the low-light light field image for the opera scene.
2. The system according to claim 1, characterized in that, The light prior processing module includes: An image segmentation unit, configured to segment the sub-aperture image of the low-light Huangmei Opera light field to obtain a block image of a preset dimension; A mean extraction unit, configured to perform light prior processing on the block image, extract the maximum value, minimum value, and average value of each block image in the channel dimension to obtain statistical information; A channel splicing unit, configured to perform channel-level splicing of the statistical information and the sub-aperture image of the low-light Huangmei Opera light field to obtain the brightened image.
3. The system according to claim 1, characterized in that, The implicit neural representation processing module includes: A position encoding unit, configured to perform position encoding on the spatial coordinates of the brightened image based on the Fourier position encoding method; A feature extraction unit, configured to perform local neighborhood feature extraction on the spatially encoded coordinates and the brightened image to obtain spatial coordinate information of the brightened image; A light enhancement unit, configured to perform deep learning processing on the spatial coordinate information and the local neighborhood features using a multi-layer perceptron to obtain the brightened image after implicit neural representation processing.
4. The system according to claim 1, wherein The first feature extraction module includes: A first global feature extraction unit, configured to extract feature information at different angles and spatial positions in the brightened image after implicit neural representation processing based on spatial-angle transformation to obtain the first global feature; A first local feature extraction unit, configured to perform local convolution operations on the brightened image after implicit neural representation processing to extract the disparity information features of the brightened image after implicit neural representation processing to obtain the first global feature.
5. The system according to claim 1, wherein The enhanced feature acquisition module includes: A first feature splicing unit, configured to splice the first local feature and the first global feature to obtain a first spliced feature; An attention processing unit, configured to adjust the feature channel weights of the first spliced feature based on the SE attention module to obtain the first spliced feature after channel weight adjustment; A first convolution unit, configured to perform 2D convolution on the first spliced feature after channel weight adjustment; A feature combination unit, configured to combine the first spliced feature after 2D convolution with the low-light Huangmei Opera light field image to obtain the enhanced feature.
6. The system according to claim 1, wherein The enhanced image acquisition module includes: A second convolution unit, configured to perform 2D convolution on the enhanced feature; A second feature extraction unit, configured to extract a second global feature and a second local feature of the enhanced feature after 2D convolution based on a low-light enhancement module; A second feature splicing unit, configured to splice the second global feature and the second local feature to obtain a second spliced feature, and perform 2D convolution on the second spliced feature; An enhanced image acquisition unit, configured to perform fusion calculation on the second spliced feature after 2D convolution and the low-light Huangmei Opera light field image to obtain the enhanced image.
7. A method for enhancing low-light light field images in an opera scene, applying the system according to any one of claims 1-6, characterized in that It includes: Performing illumination prior processing on the collected low-light Huangmei Opera light field image to obtain a brightened image; Performing implicit neural representation processing on the brightened image; Extracting a first global feature and a first local feature from the brightened image after implicit neural representation processing; Splicing the extracted first global feature and the first local feature and combining them with the low-light Huangmei Opera light field image to obtain an enhanced feature; Extracting and splicing a second global feature and a second local feature of the enhanced feature to obtain a low-light light field enhanced image of Huangmei Opera, completing the enhancement of the low-light light field image for the opera scene.
8. The method according to claim 7, wherein The method for obtaining the brightened image includes: Segmenting the sub-aperture image of the low-light Huangmei Opera light field to obtain a block image of a preset dimension; Performing illumination prior processing on the block image, and extracting the maximum value, minimum value, and average value of each block image in the channel dimension to obtain statistical information; Performing channel-level splicing on the statistical information and the sub-aperture image of the low-light Huangmei Opera light field to obtain the brightened image.
Citation Information
Patent Citations
Low-light image enhancement method for extracting and fusing local and global features
CN114972134A
Multi-view three-dimensional reconstruction method based on implicit neural representation
CN115761178A
Image segmentation method and system based on step feature fusion and attention mechanism
CN116309622A
Camouflage target image segmentation method and system based on multilevel feature fusion
CN116703950A
Image deblurring system and method thereof
CN117314770A