Image processing method and device, equipment and storage medium
By obtaining preview images and multimodal sensor data, combining regional segmentation models and brightness difference indexes, dynamically determining exposure parameters and performing layered image fusion, the processing deficiencies of HDR technology in complex lighting and dynamic scenes are resolved, thereby improving image quality and efficiency.
Patent Information
- Application Number
- CN202510893758.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-03
AI Technical Summary
Existing high dynamic range imaging (HDR) technology has shortcomings in processing effect and efficiency, especially in complex lighting and dynamic scenes, which can easily lead to loss of details in overexposed/underexposed areas and ghosting.
By obtaining the preview image and multimodal sensor data of the current scene, using the regional segmentation model and brightness difference index, dynamically determining the exposure parameters, and performing layered image fusion, dynamic partition exposure and fusion are achieved.
The HDR processing effect and efficiency have been optimized, the dynamic range, energy efficiency and adaptability of application scenarios have been improved, ghosting and overexposure/underexposure phenomena have been reduced, and image quality and user experience have been improved.
Smart Images

Figure CN120751270A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an image processing method and apparatus, device and storage medium. Background Art
[0002] High Dynamic Range Imaging (HDR) preserves more levels of detail by expanding the dynamic range of the brightest and darkest parts of an image.
[0003] Common HDR solutions need to be improved in terms of HDR processing effects and efficiency. Summary of the Invention
[0004] The embodiments of the present application provide an image processing method and apparatus, a device, and a storage medium, which can take into account both image processing efficiency and image processing effects.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] In a first aspect, an embodiment of the present application provides an image processing method, the method comprising:
[0007] Acquire a first preview image of the current scene and / or multimodal sensor data of the current scene;
[0008] determining one or more exposure parameters for the current scene based on the first preview image and / or the multimodal sensor data;
[0009] acquiring one or more first images of the current scene based on one or more exposure parameters of the current scene;
[0010] A second image of the current scene is determined by performing layered image fusion based on one or more first images of the current scene.
[0011] In a second aspect, an embodiment of the present application provides an image processing device, the image processing device comprising:
[0012] an acquiring unit, configured to acquire a first preview image of a current scene and / or multimodal sensor data of the current scene;
[0013] a determining unit, configured to determine one or more exposure parameters of a current scene based on the first preview image and / or the multimodal sensor data;
[0014] The acquiring unit is further configured to acquire one or more first images of the current scene based on one or more exposure parameters of the current scene;
[0015] The determining unit is further configured to perform layered image fusion based on one or more first images of the current scene to determine a second image of the current scene.
[0016] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory storing instructions executable by the processor. When the instructions are executed by the processor, the method of the first aspect is implemented.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the method of the first aspect described above is implemented.
[0018] The embodiments of the present application provide an image processing method and apparatus, a device and a storage medium, which obtain a first preview image of the current scene and / or multimodal sensor data of the current scene; determine one or more exposure parameters of the current scene based on the first preview image and / or multimodal sensor data; obtain one or more first images of the current scene based on the one or more exposure parameters of the current scene; perform layered image fusion based on the one or more first images of the current scene to determine the second image of the current scene. That is to say, in the embodiments of the present application, one or more exposure parameters can be dynamically determined for the current scene based on the first preview image of the current scene and / or multimodal sensor data of the current scene, and after image acquisition is performed using one or more exposure parameters, layered image fusion is performed on the acquired images. In summary, the present application can realize dynamic partitioned exposure and dynamic partitioned fusion through preview images and / or multimodal sensor data, thereby optimizing the processing effect and processing efficiency of HDR, and achieving significant improvements in multiple dimensions such as dynamic range, energy efficiency ratio, and application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic diagram of the implementation flow of the image processing method proposed in the embodiment of the present application;
[0020] Figure 2 A schematic diagram of the implementation flow of the image processing method proposed in the embodiment of the present application;
[0021] Figure 3 A schematic diagram of the implementation flow of the image processing method proposed in the embodiment of the present application;
[0022] Figure 4 A schematic diagram of the implementation process of dynamic partition fusion proposed in an embodiment of the present application;
[0023] Figure 5 This is a schematic diagram of the implementation process of heterogeneous multi-camera dynamic collaboration proposed in an embodiment of the present application;
[0024] Figure 6 A schematic diagram of an application scenario of the image processing method proposed in an embodiment of the present application;
[0025] Figure 7 This is a schematic diagram before the application of the image processing method proposed in the embodiment of the present application;
[0026] Figure 8 This is a schematic diagram after the application of the image processing method proposed in an embodiment of the present application;
[0027] Figure 9 A schematic diagram of the structure of the image processing device proposed in an embodiment of the present application;
[0028] Figure 10 This is a schematic diagram of the structure of the electronic device proposed in the embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. It should be understood that the specific embodiments described herein are only used to explain the different applications and are not intended to limit the application. It should also be noted that for ease of description, the drawings only show the parts that differ from the related applications.
[0030] High Dynamic Range (HDR) preserves more detail by expanding the dynamic range between the brightest and darkest parts of an image. Traditional Standard Dynamic Range (SDR) imaging, due to its limited brightness range, can easily cause blown highlights or loss of detail in shadows. HDR, however, significantly enhances the realism of images by leveraging high peak brightness, deep blacks, and a wide color gamut.
[0031] One HDR implementation involves a single-camera, multi-exposure fusion solution. In this solution, a single camera captures multiple frames using a fixed sequence of exposure parameters (e.g., low, medium, and high). Image alignment algorithms (e.g., SIFT feature matching) are used to eliminate the effects of jitter, and weighted fusion (e.g., luminance gradient-based fusion) is used to generate an HDR image.
[0032] However, the fixed exposure parameters of single-camera multi-exposure fusion solutions cannot adapt to complex lighting scenarios (such as high-light ratio or low-light scenarios, or complex lighting strategies such as nighttime bars, KTVs, and streets), resulting in loss of detail in overexposed or underexposed areas. Furthermore, multi-frame capture takes a long time, making ghosting more common in dynamic scenes.
[0033] Another way to implement HDR is to use a dual-camera setup, where the primary camera is responsible for regular shooting, while the secondary camera collects auxiliary data (such as long-exposure frames) and adjusts the primary camera parameters based on the secondary camera data.
[0034] However, the dual-camera division of labor is fixed (e.g., primary camera for photos, secondary camera for light metering), and the division of labor cannot be adjusted dynamically. The secondary camera data is not used for scene content analysis, and parameter selection relies on simple statistics (e.g., average brightness).
[0035] Another HDR implementation involves an exposure parameter selection scheme based on image statistics, where the brightness histogram or average brightness value of a single frame is calculated and the exposure parameters for the next frame are selected based on a preset rule (such as a histogram distribution threshold).
[0036] However, exposure parameter selection schemes based on image statistics only rely on brightness statistics and ignore scene content information such as color and texture, resulting in parameter selection bias and failure to cover dynamic scenes (such as rapid lighting changes or moving objects).
[0037] In other words, common HDR solutions need to be improved in terms of HDR processing effects and efficiency.
[0038] In order to solve the above problems, the embodiments of the present application provide an image processing method and apparatus, a device and a storage medium to obtain a first preview image of the current scene and / or multimodal sensor data of the current scene; determine one or more exposure parameters of the current scene based on the first preview image and / or multimodal sensor data; obtain one or more first images of the current scene based on the one or more exposure parameters of the current scene; perform layered image fusion based on the one or more first images of the current scene to determine the second image of the current scene. That is to say, in the embodiments of the present application, one or more exposure parameters can be dynamically determined for the current scene based on the first preview image of the current scene and / or multimodal sensor data of the current scene, and after image acquisition is performed using one or more exposure parameters, layered image fusion is performed on the acquired images. In summary, the present application can realize dynamic partitioned exposure and dynamic partitioned fusion through preview images and / or multimodal sensor data, thereby optimizing the processing effect and processing efficiency of HDR, and achieving significant improvements in multiple dimensions such as dynamic range, energy efficiency, and application scenarios.
[0039] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.
[0040] An embodiment of the present application provides an image processing method, which can be applied to an image processing device or electronic device, and can also be applied to any terminal including an image processing device or electronic device.
[0041] It is understood that the image processing method proposed in the embodiment of the present application may include an HDR solution in a single-camera or multi-camera scenario, and specifically may include an HDR solution based on scene perception and dynamic partitioning.
[0042] Below, the image processing method proposed in the embodiment of the present application is exemplarily described by taking an image processing device as an example.
[0043] Furthermore, in the embodiments of the present application, Figure 1 This is a schematic diagram of the image processing method implementation process proposed in the embodiment of the present application, as shown in FIG. Figure 1 As shown, the image processing method may include the following steps:
[0044] Step 101: Acquire a first preview image of a current scene and / or multimodal sensor data of the current scene.
[0045] In an embodiment of the present application, the image processing device may obtain a first preview image of the current scene, and / or the image processing device may also obtain multimodal sensor data of the current scene.
[0046] In an embodiment of the present application, the first preview image of the current scene may be an image of the current scene captured by one or more cameras and having a lower resolution, wherein the first preview image can be used to quickly perceive the scene content of the current scene.
[0047] In an embodiment of the present application, the multimodal sensor data of the current scene may be different types of sensor data collected by one or more sensors for the current scene, wherein the multimodal sensor data can be used to accurately evaluate scene information of the current scene.
[0048] In an embodiment of the present application, when obtaining a first preview image of the current scene, the first preview image of the current scene can be obtained through a configured first camera, and the first preview image can be downsampled to obtain the first preview image.
[0049] In some embodiments, if the image processing device is configured with a camera, such as a first camera, the first preview image may be a downsampled preview image obtained by the first camera for the current scene. The image processing device may first capture an initial preview image of the current scene using the first camera, and then further downsample the initial preview image to ultimately obtain the first preview image.
[0050] That is, in an embodiment of the present application, for a single-camera application scenario, the first preview image may include downsampled main camera data.
[0051] In an embodiment of the present application, when obtaining the first preview image of the current scene, the first preview image of the current scene can be obtained through the configured second camera.
[0052] In some embodiments, if the image processing device is configured with at least two cameras, such as a first camera and a second camera, the first preview image may be a preview image with a lower resolution obtained by a secondary camera (such as the second camera) for the current scene.
[0053] That is to say, in an embodiment of the present application, for a multi-camera application scenario, the first preview image may include low-resolution secondary camera data.
[0054] For example, in one implementation scenario, the image processing device is configured with a plurality of different cameras, wherein the number of cameras configured with the image processing device may be greater than or equal to 2. That is, in this application, the image processing device is a multi-camera device.
[0055] In some embodiments, the multiple different cameras are generally designed with different functions and features to meet the shooting requirements in different scenarios. This application does not specifically limit the performance and types of the multiple different cameras.
[0056] Exemplarily, in some embodiments, the multiple different cameras may include, but are not limited to, at least two of: a main camera, a wide-angle camera, a telephoto camera, a macro camera, a depth-of-field camera, and the like.
[0057] The main camera is the most basic camera, typically boasting higher resolution and better image quality. It handles most everyday photography tasks, such as landscapes and portraits. The main camera's sensor size, aperture, and pixel count are all relatively high to ensure clear, detailed photos in all lighting conditions.
[0058] The wide-angle camera offers a wider viewing angle than the main camera, allowing you to capture more of the scene, making it ideal for capturing scenes like landscapes and architecture. The wide-angle camera delivers a stronger visual impact, showcasing a wider field of view. It also excels in capturing images in confined spaces, giving them a greater sense of space and three-dimensionality.
[0059] Telephoto cameras are typically used for optical zoom, bringing distant objects closer for close-up shots. They're ideal for capturing distant landscapes and close-up portraits. Their shallow depth of field allows them to highlight the subject and blur the background, creating a more professional look. Some high-end phones also feature telephoto cameras with high-magnification optical zoom, enabling longer shooting distances while maintaining clarity.
[0060] Macro cameras support close-range focusing, allowing you to capture the wonderful details of the microscopic world, such as flowers and insects. Macro cameras typically have high magnification and good focusing performance, capable of presenting a delicate and clear microscopic world.
[0061] The depth-of-field camera is primarily used to enhance the blur effect in photos, making the subject stand out more clearly and the background blurrier. Through algorithmic processing, the depth-of-field camera can achieve a more natural and gentle blur effect, enhancing the artistic quality of photos.
[0062] Of course, in addition to the common camera types mentioned above, multiple different camera types can also include specialized cameras such as Time of Flight (TOF) lenses and cine lenses. ToF lenses are primarily used for 3D perception and depth measurement, enabling features like augmented reality and facial recognition. Cine lenses typically have higher resolutions and excellent color reproduction, making them suitable for capturing high-quality video.
[0063] In an embodiment of the present application, when acquiring multimodal sensor data of a current scene, the multimodal sensor data may be acquired through a configured multimodal sensor.
[0064] In an embodiment of the present application, the multimodal sensor includes at least one or more of the following: a first camera; a second camera; a TOF sensor, an ambient light sensor, and a Danxia original color lens.
[0065] That is to say, in an embodiment of the present application, multimodal sensor data may include information obtained from one or more devices including a main camera (such as a first camera), a secondary camera (such as a second camera), a depth sensor (TOF), an ambient light sensor, and a Danxia original color lens.
[0066] For example, in some embodiments, the first camera and the second camera each represent two independent image acquisition devices, typically with different optical parameters (such as focal length, field of view, and amount of light entering), and can be used to capture image information of the same scene from different perspectives or under different exposure conditions. For example, the first camera can be the primary camera and the second camera can be the secondary camera. The first camera can be responsible for primary imaging, while the second camera can assist with low-resolution scene perception or depth measurement.
[0067] For example, in some embodiments, a time-of-flight (TOF) sensor calculates distance based on the time difference between the emission and reflection signals of a laser or infrared light source and is commonly used to obtain depth information of a scene. This depth information helps distinguish foreground from background, supports region division in dynamic partitioning algorithms, and improves the accuracy of HDR image fusion.
[0068] For example, in some embodiments, the ambient light sensor is used to detect the light intensity of the surrounding environment to help determine whether the current scene is in strong light, weak light, or other specific lighting conditions, thereby optimizing exposure parameter selection and image enhancement strategy.
[0069] For example, in some embodiments, the Danxia Original Color lens is a specialized lens with multispectral acquisition capabilities, capable of capturing richer color information and improving the color reproduction accuracy of images. This lens is particularly suitable for applications requiring high-fidelity color rendering, such as portrait photography and artistic creation.
[0070] It is understood that in the embodiments of this application, one or more sensors, including the first camera, the second camera, the Time of Flight sensor, the ambient light sensor, and the Danxia Original Color lens, work together to form a complete multimodal data acquisition system, providing high-quality input data for subsequent scene recognition, dynamic exposure prediction, and image fusion. Compared to traditional single-sensor solutions, multimodal sensors can significantly improve the quality and adaptability of HDR images, especially in complex lighting and dynamic scenes, showing greater robustness and flexibility.
[0071] Step 102: Determine one or more exposure parameters of the current scene based on the first preview image and / or multimodal sensor data.
[0072] In an embodiment of the present application, after acquiring a first preview image of the current scene and / or multimodal sensor data of the current scene, the image processing device can further determine one or more exposure parameters of the current scene based on the first preview image and / or the multimodal sensor data.
[0073] That is to say, in an embodiment of the present application, one or more exposure parameters determined by dynamic derivation can be obtained based on the first preview image of the current scene, or based on the multimodal sensor data of the current scene, or based on the first preview image of the current scene and the multimodal sensor data of the current scene. This application does not make specific limitations.
[0074] In the embodiments of the present application, Figure 2 This is a schematic diagram of the image processing method implementation process proposed in the embodiment of the present application, as shown in FIG. Figure 2 As shown, the image processing method proposed in the embodiment of the present application may include the following steps:
[0075] Step 105 : Based on the first preview image and / or the multimodal sensor data, determine one or more image regions and region parameters of the one or more image regions corresponding to the first preview image by using a region segmentation model.
[0076] In an embodiment of the present application, after obtaining a first preview image of the current scene and / or multimodal sensor data of the current scene, the image processing device can further perform region segmentation on the first preview image based on the first preview image and / or the multimodal sensor data to obtain one or more image regions, and further obtain region parameters of one or more image regions corresponding to the one or more image regions.
[0077] Exemplarily, in some embodiments, the region parameters may include at least one or more of the following: a region mask of one or more image regions, a region type of one or more image regions, and a region boundary of one or more image regions.
[0078] For example, in some embodiments, the region parameter may be a quantitative indicator describing characteristics of an image region, such as region area, shape characteristics, edge distribution, etc. The region parameter may help identify different types of image content, such as faces, text, sky, buildings, etc.
[0079] Exemplarily, in some embodiments, the region segmentation model is a deep learning model for dividing an image into multiple semantic regions. The region segmentation model can adopt a lightweight structure (such as MobileNetV3+DeepLabV3) to ensure real-time performance and computational efficiency.
[0080] Exemplarily, in some embodiments, the model input of the region segmentation model may be a first preview image and / or multimodal sensor data, and the model output of the region segmentation model may include different semantic regions in the image (such as face, sky, text) and their corresponding region mask maps.
[0081] That is, in the embodiment of the present application, the region segmentation model can be used to identify key content in the first preview image and perform differentiated processing based on its characteristics.
[0082] For example, in some embodiments, when a user turns on the camera and enters HDR photography mode, a region segmentation model automatically activates to analyze the current image in real time. If a face is detected in the image, the model marks it and generates a corresponding region mask for subsequent processing modules. This mechanism not only improves the intelligence of image processing but also enhances the quality of HDR imaging, making the resulting composite image more natural and realistic.
[0083] For example, in some embodiments, a region mask for one or more image regions may include a binary image output by a region segmentation model, which identifies the location and shape of each semantic region in the image. Each pixel in the mask indicates whether the location belongs to a specific region. For example, in a mask for a face region, all pixels belonging to the face are marked as 1, while pixels in other regions are marked as 0. The mask plays an important role in subsequent processing and can be used to guide operations such as exposure parameter prediction and fusion strategy selection.
[0084] For example, in some embodiments, the region type of one or more image regions is the classification result of different regions in the image, such as faces, sky, text, trees, etc. The classification of region types depends on the training data and algorithm design of the region segmentation model. Each region type has different visual characteristics and processing requirements. For example, the face region generally requires higher color reproduction and detail preservation, while the sky region may place greater emphasis on a balanced brightness distribution.
[0085] For example, in some embodiments, the region boundaries of one or more image regions may include contour lines describing the shape and size of each image region. Region boundary information is crucial for subsequent operations such as motion compensation and fusion weight calculation. For example, in dynamic scenes, the motion trajectory of an object can be determined by analyzing changes in region boundaries, and the fusion strategy can be adjusted accordingly to reduce ghosting and blurring.
[0086] In the embodiments of this application, there is a close connection between the region mask, region type, and region boundary at both the logical and physical levels. The mask provides information about the spatial distribution of the region, the region type determines how the region is processed, and the region boundary affects the transition effect between regions. These three parameters together constitute the semantic expression system of the image, laying the foundation for the subsequent HDR imaging process.
[0087] Thus, the present application embodiment introduces the concept of region segmentation model and region parameters to achieve refined image processing. This allows accurate identification of key regions in the image and allocation of optimal exposure parameters and fusion strategies based on their characteristics, effectively improving the quality of HDR images and, in turn, enhancing the user's shooting experience under complex lighting conditions.
[0088] In the embodiments of the present application, Figure 3 This is a schematic diagram of the image processing method implementation process proposed in the embodiment of the present application, as shown in FIG. Figure 3 As shown, the image processing method proposed in the embodiment of the present application may include the following steps:
[0089] Step 106 : Determine a brightness difference index of one or more image regions corresponding to the first preview image based on the first preview image and / or the multimodal sensor data.
[0090] In an embodiment of the present application, after determining one or more image regions corresponding to the first preview image through a region segmentation model based on the first preview image and / or multimodal sensor data, the brightness difference index of one or more image regions corresponding to the first preview image can be further determined based on the first preview image and / or multimodal sensor data.
[0091] In an embodiment of the present application, the luminance difference index (RLDI) is used to quantify the difference in luminance distribution between different image regions, reflecting whether a high dynamic range lighting condition exists in the scene.
[0092] In some embodiments, the brightness difference index can be obtained by calculating parameters such as the brightness mean difference, variance contrast, and gradient discontinuity between regions.
[0093] Exemplarily, in some embodiments, based on the first preview image and / or the multimodal sensor data, brightness difference data of one or more image areas corresponding to the first preview image is determined; wherein the brightness difference data includes at least one or more of the following: brightness mean difference; variance contrast; gradient discontinuity parameter; based on the brightness difference data of the one or more image areas, a brightness difference index of the one or more image areas is determined.
[0094] That is, in an embodiment of the present application, calculating the regionalized brightness difference index (RLDI) of the first preview image may include, but is not limited to, calculating one or more of the inter-region brightness mean difference (ΔL_mean), the variance contrast (C_var), and the gradient discontinuity (i.e., the gradient discontinuity parameter G_dis).
[0095] For example, in some embodiments, in a backlit portrait scene, the brightness difference between the background and the subject's face is large, and the brightness difference index will be significantly increased. The brightness difference index can be used to determine whether to enter HDR mode and provide a basis for subsequent exposure parameter generation.
[0096] Step 107 : Determine whether to determine one or more exposure parameters based on the brightness difference index of the one or more image regions and a first index threshold.
[0097] In an embodiment of the present application, after determining the brightness difference index of one or more image areas corresponding to the first preview image based on the first preview image and / or multimodal sensor data, it is possible to further determine one or more exposure parameters based on the brightness difference index of the one or more image areas and the first index threshold, that is, determine whether to perform dynamic derivation of exposure parameters, or, in other words, determine whether to perform subsequent HDR processing.
[0098] In some embodiments, the first index threshold is a preset value used to measure whether the brightness difference index meets the criteria for triggering HDR processing. For example, when the brightness difference index is greater than or equal to the first index threshold, it indicates that the current scene has a large brightness variation, and it is necessary to generate multiple sets of exposure parameters to ensure image quality. Otherwise, when the brightness difference index is less than the first index threshold, it indicates that the current scene brightness variation is relatively gradual, and it is not necessary to dynamically promote multiple sets of exposure parameters to ensure image quality.
[0099] For example, in some embodiments, in night scene shooting, if the brightness difference index is greater than a set threshold (first index threshold), three sets of parameters, long exposure, medium exposure, and short exposure, can be generated to capture dark, mid-tone, and bright details, respectively, thereby achieving more comprehensive dynamic range coverage.
[0100] In the embodiment of the present application, by comparing the brightness difference index of the image region with a preset threshold (the first index threshold), it is possible to accurately determine whether the current scene requires an HDR processing strategy. This avoids wasting computing resources when HDR is not required, thereby improving imaging efficiency, and further enhancing the user's photography experience and device battery life.
[0101] In an embodiment of the present application, when determining the exposure strategy of the current scene based on the regionalization information of one or more image areas, the scene type corresponding to the current scene can be determined based on the regional parameters of one or more image areas, and / or the brightness difference index of one or more image areas, and / or the regionalization information of one or more image areas; and the exposure strategy of the current scene is determined based on the scene type.
[0102] In the embodiments of the present application, the luminance difference index (RLDI) is a metric used to measure the differences in luminance distribution within and between image regions. Its calculation formula can be comprehensively evaluated based on factors such as luminance mean difference, variance contrast, and gradient discontinuity. A high RLDI value indicates significant illumination variation or complex textures in the area, possibly representing a scene with a high light ratio or large dynamic range. The luminance difference index can be used to identify areas requiring special attention and assign independent exposure strategies to them, thereby improving overall image quality.
[0103] In the embodiments of this application, by combining regional parameters with the luminance difference index (RLDI), key and non-key areas in a scene can be more accurately determined and classified accordingly. The role of regional parameters is to provide semantic information support, making the subsequent exposure strategy generation more intelligent and precise.
[0104] In an embodiment of the present application, regionalization information may include a set of information obtained after dividing an image into multiple sub-regions with specific semantics or functions. Regionalization information generally includes the boundaries of each region, category labels (such as face, text, sky), brightness distribution characteristics, etc. By analyzing the regionalization information, the understanding of the scene can be further refined, and the exposure strategy can be adjusted according to the content of each region. For example, in a backlit portrait scene, the face area can be identified through regionalization information and its exposure quality can be prioritized.
[0105] For example, in some embodiments, the regionalization information may be normalized information entropy (NIE).
[0106] In some embodiments, the pixel value distribution probability of one or more image regions can be determined based on the first preview image and / or the multimodal sensor data; and regionalization information of one or more image regions can be determined based on the pixel value distribution probability of one or more image regions.
[0107] For example, in some embodiments, assuming that the pixel value distribution probability of an image region is the region pixel value distribution probability p_i of the image region, and the total number of pixels of the first preview image is N, the normalized information entropy (NIE) of the image region can be determined by the following formula:
[0108] NIE=-Σ(p_i*log2(p_i)) / log2(N) (1)
[0109] That is to say, in an embodiment of the present application, based on one or more of the regional parameters of one or more image areas, the brightness difference index of one or more image areas, and the regionalization information of one or more image areas, the current scene can be distinguished from different dimensions to determine the scene type corresponding to the current scene.
[0110] In embodiments of the present application, the scene type may include the results of classifying the current shooting scene based on image content and environmental characteristics. Common scene types include night scenes, backlit portraits, indoor mixed light, moving objects, and document reprints. Different scene types correspond to different imaging requirements and optimization goals. For example, night scenes require enhanced dark details, backlit portraits require balancing the brightness of the subject and background, and document reprints require improving text clarity. By identifying the scene type, predefined optimization strategies can be invoked, enabling more intelligent exposure control.
[0111] In an embodiment of the present application, the exposure strategy may include a set of exposure parameter selection rules and fusion algorithm configurations formulated for specific scene types. The exposure strategy typically includes multiple sets of candidate exposure parameters (such as different shutter speeds, ISO values, gain coefficients), regionalized exposure weight allocation, layered fusion methods (such as static background multi-frame fusion, dynamic area single-frame enhancement), etc. A reasonable exposure strategy can effectively improve the quality of HDR images, avoid overexposure or underexposure, and maintain good compatibility and stability under different hardware conditions.
[0112] In summary, in the embodiments of the present application, scene types are identified based on the regional parameters, brightness difference index, and regionalization information of the image region, and a highly adaptable exposure strategy is generated accordingly. This ensures that high-quality HDR images can be obtained under different lighting conditions and scene contents, thereby improving the visual effect and practicality of the image, and further meeting the diverse shooting needs of users.
[0113] In an embodiment of the present application, when determining one or more exposure parameters of the current scene based on the first preview image and / or multimodal sensor data, regionalization information of one or more image areas corresponding to the first preview image can be determined based on the first preview image and / or multimodal sensor data; based on the regionalization information of the one or more image areas, the exposure strategy of the current scene can be determined; based on the exposure strategy of the current scene, one or more exposure parameters of the current scene can be determined.
[0114] In embodiments of the present application, the exposure strategy can be used to determine the method for calculating exposure parameters and / or the number of exposure parameters for the current scene. Specifically, the exposure strategy can be used to determine the calculation method for determining exposure parameters and / or the number of exposure parameters for the current scene.
[0115] In an embodiment of the present application, regionalization information of one or more image regions corresponding to the first preview image may be determined based on the first preview image and / or multimodal sensor data.
[0116] For example, in some embodiments, regionalization information may include semantic segmentation of an image into multiple regions with clear semantic meaning, and extraction of visual features and attributes for each region. These regions may include face regions, sky regions, text regions, dark regions, highlight regions, etc. Regionalization information includes not only the boundary information of each region, but may also include the brightness distribution, texture complexity, color characteristics, etc. of the region.
[0117] In some embodiments, by introducing regionalized information, a more detailed understanding of image content can be achieved, allowing independent exposure processing strategies to be developed for different regions. For example, in a backlit portrait scene, the brightness difference between the face area and the background area is large, and traditional global exposure strategies are difficult to meet the needs of both simultaneously. By using regionalized information, the lighting conditions of these two areas can be evaluated separately, thereby generating a more reasonable exposure parameter combination. In this way, overexposure or underexposure problems caused by a single exposure parameter can be avoided, thereby improving the overall dynamic range performance of the image and enhancing the ability to retain details in key areas, such as skin color reproduction and text clarity.
[0118] In an embodiment of the present application, an exposure strategy for the current scene may be determined based on regionalized information of one or more image regions; wherein the exposure strategy is used to determine a calculation method for exposure parameters and / or the number of exposure parameters for the current scene.
[0119] For example, in some embodiments, the exposure strategy may include a set of rules and methods for selecting and generating exposure parameters based on scene content and region segmentation. This determines whether to employ single-frame optimization, multi-frame fusion, or whether different exposure times or intensification values should be used for different regions. For example, in scenes with high light contrast, the exposure strategy may trigger multi-frame HDR synthesis; while in scenes with low dynamic range, only single-frame capture may be used to conserve resources.
[0120] For example, in some embodiments, the exposure strategy can also control the number of exposure parameters, for example, generating three different sets of exposure parameters for complex scenes, while using only one set for simple scenes. This dynamic decision-making mechanism enables a balance between image quality and performance.
[0121] In some embodiments, the generation logic of exposure parameters can be flexibly adjusted according to different scenarios, thereby improving imaging quality and reducing unnecessary computing overhead, thereby optimizing system energy efficiency and ensuring efficient operation on mobile devices.
[0122] In an embodiment of the present application, one or more exposure parameters of the current scene may be determined based on the exposure strategy of the current scene.
[0123] For example, in some embodiments, exposure parameters refer to the settings actually applied to the camera when capturing images, including but not limited to exposure time (shutter speed), ISO sensitivity value, aperture size, etc. These parameters directly affect the brightness, noise level, and dynamic range of the image.
[0124] It can be understood that in the embodiments of the present application, due to the introduction of regionalized information and dynamic exposure strategy, the exposure parameters are no longer fixed, but are dynamically generated for different regions and scenes.
[0125] For example, in some embodiments, in a night portrait scene, a longer exposure time and lower ISO may be selected for the subject area to reduce noise and preserve more details; while a shorter exposure time and higher ISO may be used for the background area to prevent overexposure and capture more environmental details. This differentiated exposure strategy helps obtain higher quality images in complex lighting conditions.
[0126] In some embodiments, by customizing optimal exposure parameters for different areas, image quality can be significantly improved, especially under extreme lighting conditions, thereby enhancing the user's photography experience and image satisfaction.
[0127] In an embodiment of the present application, the dynamically generated one or more exposure parameters of the current scene and the one or more image areas corresponding to the first preview image may correspond one-to-one or one-to-many, which is not limited in the present application.
[0128] That is, in the embodiment of the present application, the number of image regions obtained after dividing the first preview image and the number of exposure parameters finally dynamically generated may be the same or different.
[0129] Step 103: Acquire one or more first images of the current scene based on one or more exposure parameters of the current scene.
[0130] In an embodiment of the present application, after determining one or more exposure parameters of the current scene based on the first preview image and / or multimodal sensor data, one or more first images of the current scene can be further acquired based on the one or more exposure parameters of the current scene.
[0131] In an embodiment of the present application, one or more first images can be understood as an image sequence obtained by the image processing device according to the generated exposure parameters, that is, an image sequence obtained by the image processing device performing image capture for the current scene according to the generated one or more exposure parameters.
[0132] In an embodiment of the present application, when obtaining one or more first images of the current scene based on one or more exposure parameters of the current scene, one or more first images of the current scene can be obtained based on one or more exposure parameters through the configured first camera.
[0133] In some embodiments, if the image processing device is configured with a camera, such as a first camera, the first image may be a full-resolution image captured by the first camera for the current scene.
[0134] That is to say, in the embodiment of the present application, for a single-camera application scenario, the first image may include full-resolution image data captured by the main camera.
[0135] In an embodiment of the present application, the image processing device may further acquire a second preview image of the current scene through a configured first camera, and determine a third image of the current scene based on the second preview image.
[0136] In some embodiments, if the image processing device is configured with a camera, such as a first camera, the third image may be a downsampled preview image obtained by the first camera for the current scene. The image processing device may first capture a second preview image of the current scene using the first camera, and then further downsample the second preview image to ultimately obtain the third image.
[0137] That is to say, in the embodiment of the present application, for a single-camera application scenario, the third image may include downsampled main camera data.
[0138] In an embodiment of the present application, the image processing device may further acquire a third image of the current scene through a configured second camera.
[0139] In some embodiments, if the image processing device is configured with at least two cameras, such as a first camera and a second camera, the third image may be a preview image with a lower resolution obtained by a secondary camera (such as the second camera) for the current scene.
[0140] That is to say, in the embodiment of the present application, for a multi-camera application scenario, the third image may include low-resolution secondary camera data.
[0141] In an embodiment of the present application, the third image may be a low-resolution preview image used to assist in subsequent HDR image processing. As an auxiliary image for HDR image processing, the third image may provide auxiliary information such as motion vectors and depth data to optimize subsequent fusion effects.
[0142] For example, in some embodiments, the third image may include an input image used to extract scene motion features, typically a low-resolution preview frame or an auxiliary image captured by the secondary camera. The primary function of this image is to provide dynamic information in the scene, such as object movement and camera shake, rather than to influence the final image quality. For example, in a multi-camera system, the third image may be quickly captured at a lower resolution by the secondary camera to facilitate real-time analysis of scene motion. In a single-camera system, the third image may be obtained by downsampling the preview data from the primary camera.
[0143] Step 104: Perform layered image fusion based on one or more first images of the current scene to determine a second image of the current scene.
[0144] In an embodiment of the present application, after acquiring one or more first images of the current scene based on one or more exposure parameters of the current scene, layered image fusion can be further performed based on the one or more first images of the current scene to determine a second image of the current scene.
[0145] In an embodiment of the present application, when performing layered image fusion based on one or more first images of the current scene to determine the second image of the current scene, the first type of image and the second type of image of the current scene can be determined based on the one or more first images of the current scene; then image fusion is performed based on the first type of image and the second type of image to determine the second image of the current scene.
[0146] In an embodiment of the present application, different types of images corresponding to the current scene, such as first type images and second type images, may be determined based on one or more first images of the current scene.
[0147] In some embodiments, by classifying the first image into different types, a fusion algorithm can be formulated more specifically, thereby improving the overall imaging quality.
[0148] For example, in some embodiments, the number of types obtained after division is not limited to the first type of image and the second type of image, and there may be more types, which is not specifically limited in this application.
[0149] For example, in some embodiments, the first type of image may be a static image, and the second type of image may be a dynamic image.
[0150] For example, in some embodiments, the first type of image may be a text image, and the second type of image may be a background image.
[0151] For example, in some embodiments, the first type of image may be a foreground image, and the second type of image may be a background image.
[0152] In some embodiments, the second type of image includes multiple dynamic type images. When determining the second type of image of the current scene based on one or more first images of the current scene, the motion information corresponding to the current scene is determined based on the third image of the current scene; the one or more first images are aligned based on the motion information to obtain one or more aligned images; and the second type of image is determined based on the one or more aligned images.
[0153] For example, in some embodiments, the motion information may include data on the motion of an object or camera in the scene obtained by analyzing the third image, typically including parameters such as motion vector, speed, direction, etc.
[0154] It can be understood that in the embodiments of the present application, by utilizing the third image to extract motion information, the accuracy of multi-frame image fusion can be effectively improved, especially in dynamic scenes, the generation of ghosting and blurring artifacts can be reduced, thereby improving the quality of the second type of image finally output.
[0155] In some embodiments, one or more first images are aligned based on the motion information to obtain one or more aligned images. The alignment process may include registering the multiple first images in the spatial domain based on the known motion information so that they are consistent in geometric position, thereby facilitating subsequent fusion processing. The core of the alignment process is to eliminate image offsets caused by camera shake or object motion, and ensure that the pixels in the same area of images under different exposures can correspond correctly. For example, when the motion detection results indicate that an object has shifted 5 pixels between two frames, the alignment process will adjust the displacement of the second frame image relative to the first frame image accordingly, so that the pixel distribution of the two in the target area tends to be consistent.
[0156] For example, in some embodiments, the alignment process is typically implemented in combination with an optical flow method, a feature point matching algorithm (such as SIFT, SURF), or a deep learning-based image registration model.
[0157] For example, in some embodiments, in actual applications, in order to improve efficiency, sparse processing technology can be used to perform high-precision alignment only on key areas (such as faces, text, etc. output by semantic segmentation), while low-complexity methods are used for global areas.
[0158] It can be understood that in the embodiments of the present application, through alignment processing, the alignment accuracy of multiple frames of images in the spatial dimension can be significantly improved, and image misalignment caused by motion or jitter can be avoided, thereby improving the clarity and naturalness of the final HDR composite image.
[0159] In summary, in the embodiments of the present application, by extracting motion information based on the third image, aligning the first image accordingly, and then generating the second type of image, the fusion quality of multiple frames in dynamic scenes can be effectively improved, thereby reducing the generation of ghosting and blurring artifacts, and thus achieving high-quality HDR image synthesis.
[0160] In an embodiment of the present application, when performing image fusion based on a first type of image and a second type of image to determine a second image of the current scene, fusion can be performed based on the first type of image to determine a first fused image; and based on the first fused image and the second type of image, the second image of the current scene can be determined.
[0161] In some embodiments, different types of images, such as first type images and second type images, may be processed accordingly using different processing strategies.
[0162] For example, in some embodiments, different processing strategies may be employed for the first and second types of images during the image fusion process. For example, a multi-frame fusion approach may be employed for the first type of image to obtain a first fused image, while a single-frame strategy may be employed for the second type of image to directly fuse the image with the first fused image to ultimately obtain a second image of the current scene.
[0163] In an embodiment of the present application, the first type of image includes multiple static images. When fusion is performed based on the first type of image to determine the first fused image, the fusion weight corresponding to the first type of image is determined based on the image information of the first type of image; the first type of image is fused based on the fusion weight to determine the first fused image.
[0164] Exemplarily, in some embodiments, taking static fusion as an example, assuming that the first type of image is multiple static images, a multi-frame weighted fusion strategy can be adopted to further determine the fusion weights, and use the fusion weights to perform weighted fusion on the multiple static images to finally obtain the first fused image.
[0165] For example, in some embodiments, image information may include comprehensive data on visual features such as brightness, color, and texture contained in the image. Image information can be used to determine the clarity, contrast, and detail richness of the image, thereby helping to determine the weight that should be assigned to each image during the fusion process. By dynamically calculating fusion weights based on image information, the present application can intelligently assign the importance of each image in the fusion process based on its actual quality, thereby avoiding the problems of local detail loss or over-smoothing that may be caused by fixed weights.
[0166] In the embodiments of the present application, by introducing a fusion weight mechanism, the present application can reduce artifacts and distortion while ensuring image details, especially in scenes under complex lighting conditions, and can significantly improve the realism and visual effects of the image.
[0167] In an embodiment of the present application, when determining the second image of the current scene based on the first fused image and the second type of image, the second type of image is preset processed based on the first fused image to obtain a processed image; and the second image of the current scene is determined based on the first fused image and the processed image.
[0168] In some embodiments, the preset processing may include at least one or more of the following: local contrast enhancement processing; deblurring processing; brightening processing; and trusted multi-view classification processing.
[0169] For example, in some embodiments, local contrast enhancement, deblurring, brightening, TMC and other algorithms may be applied to the dynamic region to align the image quality of the static region.
[0170] That is, in the embodiment of the present application, by performing preset processing such as local contrast enhancement and deblurring on the dynamic area, its clarity and detail expression can be significantly improved. Subsequently, the processed dynamic image is fused with the static background image to finally obtain a high-quality HDR image.
[0171] In summary, the image processing method proposed in the embodiments of this application, through scene semantic perception, dynamic zone exposure, multi-camera heterogeneous collaboration and other processing, combined with intelligent fusion algorithms, energy efficiency optimization strategies, and exception handling strategies, systematically breaks through the performance and effect bottlenecks of traditional HDR technology, achieving the following core advantages:
[0172] 1. Full scenario coverage: various user scenarios can be covered through targeted exposure strategies;
[0173] 2. Refined processing: Through scene recognition, a zoned exposure strategy is applied to different subject areas, ensuring that no subject is missed or sacrificed, ensuring that both the foreground and background look good.
[0174] 3. Improved hardware efficiency: Improved multi-camera collaborative utilization;
[0175] 4. Motion blur reduction: Precise exposure strategy and partition fusion algorithm can effectively reduce the blur of moving objects;
[0176] 5. Robustness in extreme environments: Increase HDR algorithm fault tolerance to ensure 100% image quality.
[0177] This systematically solves the inherent problems of existing HDR imaging technology and achieves improvements in dynamic range, energy efficiency, user experience and other dimensions.
[0178] For example, in some embodiments, the implementation process of dynamic partition fusion is as follows: Figure 4As shown, scene analysis (such as RLDI calculation) is performed based on the preview frames captured by the main camera or the auxiliary camera, and further based on the calculation result of RLDI and a predetermined threshold, it is determined whether to trigger the precise partitioning exposure process. If the calculation result of RLDI is greater than the threshold, the precise partitioning exposure process is triggered, otherwise the ordinary multi-frame process is executed. When executing the precise partitioning exposure process, semantic segmentation and dynamic partitioning are performed respectively, and then the exposure parameters are dynamically determined, that is, the optimal exposure strategy is determined according to different partitions, and a photo-taking frame request is sent to the hardware layer, and then multi-exposure shooting can be performed through multiple sensors, and the images obtained under different exposure parameters are layered and fused. Among them, for the layered fusion process, the segmented areas can be determined separately to determine different types of areas, such as motion areas and static areas. For the motion areas, single-frame enhancement and motion compensation processing can be performed, and for the static areas, multi-frame weighted fusion processing can be performed to finally synthesize an HDR image.
[0179] For example, in some embodiments, the implementation process of heterogeneous multi-camera dynamic collaboration is as follows: Figure 5 As shown, a heterogeneous camera capability matrix can be designed to dynamically allocate tasks based on camera hardware characteristics (wide-angle lens light intake, telephoto lens resolution, time-of-flight depth accuracy, and multispectral sensor). The wide-angle lens (secondary camera, monochrome 8M) performs global scene analysis and low-exposure frame capture; the telephoto lens (main camera, RGB 64M) performs super-resolution imaging of key areas (such as faces and text); and the time-of-flight sensor performs depth-assisted exposure parameter calculation and motion compensation. Data collected by the camera hardware passes through the ISP processor and depth perception module before being input into the ASD scene detection and segmentation exposure algorithm. Simultaneously, the RLDI algorithm has been upgraded to incorporate multispectral RLDI. In addition to visible light brightness, it incorporates infrared / ultraviolet spectral difference analysis to improve segmentation accuracy in complex lighting conditions (such as haze and backlighting). Near-infrared band data is used to enhance the signal-to-noise ratio (SNR) in dark areas, while visible light band data is used to preserve color accuracy. Cross-spectral motion compensation: In dynamic scenes, the near-infrared band can penetrate some haze / smog, providing more stable motion estimation data. After fusion, the output is a partition mask to guide the selection of multi-exposure parameters. Ultimately, the subsequent multi-camera fusion HDR algorithm is executed based on the determined multi-exposure parameters. For the ASD scene detection partition exposure algorithm and the multi-camera fusion HDR algorithm, homomorphic computing power can be called, reducing the terminal computing power requirements.
[0180] An embodiment of the present application proposes an image processing method, which obtains a first preview image of the current scene and / or multimodal sensor data of the current scene; determines one or more exposure parameters of the current scene based on the first preview image and / or multimodal sensor data; obtains one or more first images of the current scene based on the one or more exposure parameters of the current scene; and performs layered image fusion based on the one or more first images of the current scene to determine the second image of the current scene. That is to say, in an embodiment of the present application, one or more exposure parameters can be dynamically determined for the current scene based on the first preview image of the current scene and / or multimodal sensor data of the current scene, and after image acquisition is performed using one or more exposure parameters, layered image fusion is performed on the acquired image. In summary, the present application can realize dynamic partitioned exposure and dynamic partitioned fusion through preview images and / or multimodal sensor data, thereby optimizing the processing effect and processing efficiency of HDR, and achieving significant improvements in multiple dimensions such as dynamic range, energy efficiency ratio, and application scenarios.
[0181] Based on the above embodiments, another embodiment of the present application proposes an image processing method and designs a multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning. This method introduces a lightweight semantic segmentation model and regional information analysis, combined with multimodal sensor data (such as preview images, depth information, ambient light sensors, TOF sensors, Danxia original color lenses, etc.), to achieve real-time perception and intelligent partitioning of the current scene, thereby generating differentiated exposure strategies for different areas.
[0182] High dynamic range (HDR) is a technology that expands the brightness range of an image by fusing multiple frames. It is widely used in photography, video processing, and computer vision. HDR technology aims to preserve more detail in complex lighting conditions, improving the visual quality and realism of images.
[0183] In related technologies, HDR image generation typically relies on methods such as single-camera multi-exposure fusion, dual-camera division of labor, or parameter selection based on image statistics. The single-camera approach captures multiple frames of images using different exposure parameters and aligns and fuses them, but the fixed exposure strategy is difficult to adapt to dynamic scenes. The dual-camera approach divides the work of collecting data between the main and auxiliary cameras, but lacks a collaborative optimization mechanism. While image statistics-based methods can achieve a certain degree of parameter adjustment, they often ignore the characteristics of scene content, resulting in limited image quality.
[0184] These solutions commonly suffer from the following issues: First, the exposure parameter selection mechanism is rigid and cannot be flexibly adjusted based on scene content; second, they are highly dependent on hardware architecture, limiting the technology's universal applicability; third, they are poorly adaptable to dynamic scenes, prone to ghosting or blurring; fourth, they consume large amounts of computing resources, impacting real-time performance and power consumption control; and fifth, they lack targeted optimization of key local areas, resulting in the loss of important information. These issues severely restrict the effectiveness of HDR technology in complex environments.
[0185] This demonstrates that HDR image processing technologies struggle to balance image detail and overall quality under complex lighting conditions due to rigid exposure parameter selection mechanisms, strong hardware dependency, poor adaptability to dynamic scenes, high computational resource consumption, and coarse-grained scene recognition optimization. For example, in scenes like backlit portraits and night scenes of moving objects, traditional methods often suffer from localized overexposure / underexposure, ghosting, blurring, high power consumption, and loss of detail in key areas.
[0186] To address the above-mentioned issues, the image processing method proposed in the embodiment of the present application dynamically allocates tasks and reduces dependence on specific hardware through a heterogeneous collaborative architecture of the main and secondary cameras, supporting flexible adaptation of single-camera or multi-camera systems. On this basis, a layered image fusion strategy is adopted to implement differentiated fusion algorithms for static backgrounds and dynamic objects respectively to eliminate artifacts, improve clarity, and significantly reduce motion blur and color deviation. Ultimately, the present invention can effectively reduce power consumption, improve imaging stability, and be compatible with a variety of device configurations while ensuring image quality, thereby achieving a more natural and realistic HDR imaging effect in complex lighting scenarios.
[0187] In some embodiments, a first preview image of the current scene and / or multimodal sensor data of the current scene can be obtained; one or more exposure parameters of the current scene can be determined based on the first preview image and / or multimodal sensor data; one or more first images of the current scene can be obtained based on the one or more exposure parameters of the current scene; and layered image fusion can be performed based on the one or more first images of the current scene to determine a second image of the current scene.
[0188] Thus, in the embodiments of the present application, by acquiring a first preview image and / or multimodal sensor data of the current scene and dynamically determining one or more exposure parameters based on this information, the exposure strategy can be intelligently adjusted according to the scene content, avoiding overexposure or underexposure problems caused by fixed parameters. Furthermore, layered image fusion is performed based on the multiple acquired first images, and different fusion methods can be used for static and dynamic areas, respectively, to improve overall imaging quality and reduce ghosting. Compared with the existing method that relies on global unified processing, this significantly improves the image detail retention rate and the naturalness of the image quality.
[0189] In some embodiments, when obtaining a first preview image of the current scene, the first preview image of the current scene can be obtained through a configured first camera, and the first preview image can be downsampled to obtain the first preview image; and / or, the first preview image of the current scene can be obtained through a configured second camera.
[0190] Exemplarily, in some embodiments, a first preview image of the current scene is acquired through a configured first camera, and the first preview image is downsampled to obtain the first preview image.
[0191] For example, in some embodiments, the first camera may include a main camera for capturing images of the current scene. This camera typically has high resolution and imaging quality, and can provide clear image data as a basis for subsequent processing. In this application, the first camera is responsible for not only capturing the original image but also downsampling it to generate a low-resolution version suitable for fast processing and real-time display. The purpose of downsampling is to reduce the image data volume and computational complexity, thereby speeding up processing and reducing resource consumption.
[0192] Downsampling can include the process of converting a high-resolution image into a low-resolution image. This process is achieved by reducing the number of pixels in the image or merging adjacent pixels. Common downsampling methods include nearest neighbor interpolation, bilinear interpolation, etc. In this application, downsampling can ensure that the image meets the algorithm's requirements for computational efficiency while maintaining key visual information. For example, in a single-camera project, the main camera preview data will be downsampled and fed into the algorithm for processing, thereby avoiding the high latency problem caused by the full-resolution image.
[0193] In this way, the response speed and processing power can be significantly improved while ensuring image quality, thereby improving the user experience, especially on resource-constrained mobile devices.
[0194] Exemplarily, in some embodiments, a first preview image of the current scene is acquired through a configured second camera.
[0195] For example, in some embodiments, the second camera may include an auxiliary camera, which typically has a lower resolution than the first camera but offers more flexible task division capabilities. In this application, the primary task of the second camera is to assist the first camera in scene perception and parameter prediction. For example, in a multi-camera project, the second camera can capture low-resolution auxiliary frames for tasks such as motion detection, semantic segmentation, and color correction.
[0196] The second camera works closely with the first. The first camera is responsible for high-quality image acquisition, while the second camera focuses on low-power, low-latency scene perception tasks. The advantages of this heterogeneous multi-camera collaborative architecture are: on the one hand, it fully utilizes the hardware characteristics of different cameras to improve overall imaging efficiency; on the other hand, it effectively reduces the burden on a single camera and extends device battery life.
[0197] In practical applications, the second camera can calculate motion vectors using optical flow, helping to identify dynamic objects in the scene and adjust exposure strategies accordingly. Furthermore, the second camera can also assist with color correction, particularly when using Danxia True Color lenses, helping to improve image color reproduction accuracy.
[0198] By introducing a second camera, finer scene control and higher image processing efficiency can be achieved, thereby providing better HDR imaging effects under complex lighting conditions.
[0199] In summary, in the embodiment of the present application, the first preview image is obtained and down-sampled by the configured first camera, and the second preview image is obtained by the configured second camera. This can reduce the computational complexity of image processing, improve response speed and processing power, and thus achieve efficient HDR imaging processing on resource-constrained devices, thereby improving the user's shooting experience and image quality.
[0200] In some embodiments, when acquiring multimodal sensor data of the current scene, the multimodal sensor data can be acquired through a configured multimodal sensor; wherein the multimodal sensor may include a combination of multiple types of sensors for collecting data of different physical properties, the purpose of which is to acquire information about the current scene from multiple dimensions, thereby improving the understanding of the scene content and processing accuracy.
[0201] Exemplarily, in some embodiments, the multimodal sensor may include, but is not limited to, a first camera, a second camera, a time-of-flight (TOF) sensor, an ambient light sensor, and a Danxia original color lens.
[0202] In other words, in the embodiments of this application, by integrating multiple types of sensors such as TOF and ambient light sensors to obtain rich multimodal data, it helps to more comprehensively perceive scene information and provide a more reliable basis for subsequent exposure parameter prediction and image fusion. This multimodal data fusion approach can significantly improve the ability to adapt to complex scenes compared to single image input.
[0203] For example, in some embodiments, such as when shooting at night, the ambient light sensor can detect low light conditions and prompt the user to activate long exposure mode. Meanwhile, the TOF sensor can assist in identifying moving objects, preventing blur or artifacts during image fusion. In backlit portraits, the Danxia True Color lens ensures the faithful reproduction of skin tones, enhancing the overall visual effect.
[0204] In summary, the embodiments of this application utilize multimodal sensors, combined with different types of data sources, to more comprehensively perceive and understand the current scene, thereby generating more natural, clear, and realistic HDR images. This effectively addresses the challenges of complex lighting and dynamic changes, thereby improving image quality and meeting user demands for high-quality images.
[0205] In some embodiments, based on the first preview image and / or multimodal sensor data, one or more image regions corresponding to the first preview image and region parameters of one or more image regions are determined through a region segmentation model; wherein the region parameters include at least one or more of the following: region masks of one or more image regions, region types of one or more image regions, and region boundaries of one or more image regions.
[0206] For example, in some embodiments, the region segmentation model can adopt a lightweight structure (such as MobileNetV3+DeepLabV3), and its input can be a downscale version of the main camera preview image, or a low-resolution image captured by the secondary camera, and can also include multimodal sensor data. The output results include different semantic regions in the image (such as faces, sky, text) and their corresponding regional masks. Through the region segmentation model, the key content in the image can be identified and differentiated according to its characteristics.
[0207] For example, in some embodiments, the region parameters include at least one or more of the following: a region mask of one or more image regions, a region type of one or more image regions, and a region boundary of one or more image regions. The region mask may include a binary image output by a region segmentation model, used to identify the location and shape of each semantic region in the image. The region type is a classification result of different regions in the image, such as faces, sky, text, trees, etc. The region boundary may include the outline of each image region, which describes the shape and size of the region.
[0208] For example, in some embodiments, there is a close connection between region masks, region types, and region boundaries at both a logical and physical level. The mask provides information about the spatial distribution of regions, the region type determines how the region is processed, and the region boundaries influence the transition between regions. Together, these three parameters constitute the semantic representation of the image, laying the foundation for the subsequent HDR imaging process.
[0209] In other words, in the embodiments of this application, a region segmentation model is used to perform semantic segmentation on the image, extracting the mask, type, and boundary information of each region. This helps to formulate differentiated exposure parameters and fusion strategies based on the attributes of different regions. Compared with the traditional simple brightness range division, this method can more accurately identify key areas such as faces and text, thereby ensuring their image quality.
[0210] In some embodiments, a brightness difference index of one or more image areas corresponding to the first preview image is determined based on the first preview image and / or multimodal sensor data; and based on the brightness difference index of the one or more image areas and a first index threshold, it is determined whether to determine one or more exposure parameters.
[0211] For example, in some embodiments, the brightness difference index is used to quantify the brightness distribution difference between different image regions, reflecting whether there is a high dynamic range lighting condition in the scene.
[0212] For example, in some embodiments, the first index threshold is a preset value used to measure whether the brightness difference index meets the criteria for triggering HDR processing. For example, when the brightness difference index exceeds the first index threshold, it indicates that the current scene has a large brightness variation and it is necessary to generate multiple sets of exposure parameters to ensure image quality.
[0213] For example, in some embodiments, by comparing the brightness difference index of the image region with a first index threshold, it is possible to accurately determine whether the current scene requires an HDR processing strategy. This avoids wasting computing resources when HDR is not required, thereby improving imaging efficiency, and further enhancing the user's photography experience and device battery life.
[0214] In other words, in the embodiments of this application, by quantitatively analyzing the brightness difference index of the image area and combining it with a set threshold to determine whether to generate multiple exposure parameters, HDR mode can be automatically triggered in high-light-ratio scenes. This adaptive mechanism, based on the actual scene complexity, effectively avoids redundant operations, saves resources, and improves efficiency compared to traditional fixed exposure combination methods.
[0215] In summary, the embodiments of this application implement an intelligent recognition and response mechanism for the dynamic range of image scenes by introducing a brightness difference index and a first index threshold. In actual implementation, image regions are analyzed for brightness, extracting key brightness features that serve as the basis for subsequent judgments. These features are then compared with preset thresholds to determine whether to initiate the HDR processing process. This seamless integration of the entire process ensures both accurate judgments and improved operational efficiency.
[0216] In some embodiments, when determining one or more exposure parameters of the current scene based on the first preview image and / or multimodal sensor data, regionalization information of one or more image areas corresponding to the first preview image can be determined based on the first preview image and / or multimodal sensor data; based on the regionalization information of the one or more image areas, the exposure strategy of the current scene is determined; wherein the exposure strategy is used to determine the calculation method of the exposure parameters and / or the number of exposure parameters of the current scene; based on the exposure strategy of the current scene, one or more exposure parameters of the current scene are determined.
[0217] For example, in some embodiments, exposure parameters include, but are not limited to, exposure time (shutter speed), ISO sensitivity, aperture size, etc. These parameters directly affect the brightness, noise level, and dynamic range of the image. Because this solution introduces regionalized information and a dynamic exposure strategy, the exposure parameters are no longer fixed but are dynamically generated for different regions and scenes.
[0218] For example, in some embodiments, in a night portrait scene, a longer exposure time and lower ISO may be selected for the subject area to reduce noise and preserve more details; while a shorter exposure time and higher ISO may be used for the background area to prevent overexposure and capture more environmental details. This differentiated exposure strategy helps obtain higher quality images in complex lighting conditions.
[0219] In this way, by customizing the optimal exposure parameters for different areas, image quality can be significantly improved, especially in extreme lighting conditions, thereby enhancing the user's photography experience and image satisfaction.
[0220] In summary, in the embodiments of the present application, by introducing regionalized information, it is possible to identify the lighting characteristics of different semantic areas in the image, and formulate differentiated exposure parameters accordingly, thereby effectively solving the problem that traditional global exposure strategies cannot take into account all areas under complex lighting.
[0221] It can be understood that in the embodiments of this application, first, the image is divided into regions and regional information is extracted to provide a basis for subsequent exposure strategy formulation. Second, the complexity of the current scene is determined based on the regional information to determine whether multi-frame fusion or regional differentiated exposure is required. Finally, based on the formulated exposure strategy, specific exposure parameters are generated for different regions. The entire process achieves closed-loop control from image understanding to parameter generation, improving intelligence and adaptability.
[0222] Thus, in the embodiments of this application, by identifying different regions in an image and analyzing their regionalized information, such as brightness differences and texture complexity, a more refined exposure strategy can be developed. This strategy dynamically determines whether to generate multiple exposure parameters and how to calculate them based on regional characteristics. Compared to traditional methods based on global brightness statistics, it can more accurately match the lighting conditions of local areas, thereby improving the imaging quality of key areas.
[0223] In some embodiments, when determining the exposure strategy of the current scene based on the regionalization information of one or more image areas, the scene type corresponding to the current scene is determined based on the regional parameters of one or more image areas, and / or the brightness difference index of one or more image areas, and / or the regionalization information of one or more image areas; and the exposure strategy of the current scene is determined based on the scene type.
[0224] For example, in some embodiments, the scene type may include the result of classifying the current shooting scene according to the image content and environmental characteristics. Common scene types include night scenes, backlit portraits, indoor mixed light, moving objects, document copying, etc. Different scene types correspond to different imaging requirements and optimization goals.
[0225] For example, in some embodiments, the exposure strategy may include a set of exposure parameter selection rules and fusion algorithm configurations developed for specific scene types. The exposure strategy typically includes multiple sets of candidate exposure parameters (such as different shutter speeds, ISO values, and gain coefficients), regionalized exposure weight allocation, and layered fusion methods (such as multi-frame fusion of static backgrounds and single-frame enhancement of dynamic areas). A reasonable exposure strategy can effectively improve the quality of HDR images, avoid overexposure or underexposure, and maintain good compatibility and stability under different hardware conditions.
[0226] In the embodiments of the present application, scene types are identified based on regional parameters, brightness difference index, and regionalization information of the image region, and an adaptive exposure strategy is generated accordingly. This ensures that high-quality HDR images can be obtained under different lighting conditions and scene contents, thereby improving the visual effect and practicality of the image and meeting the diverse shooting needs of users.
[0227] In summary, in the embodiments of this application, by comprehensively analyzing regional parameters, brightness difference index, and regionalization information, the specific type of the current scene (e.g., backlit portraits, night scenes of buildings, etc.) can be identified and the corresponding optimization strategy can be invoked accordingly. This approach makes the selection of exposure strategies more targeted and can maintain good imaging results under various complex lighting conditions.
[0228] In some embodiments, when acquiring one or more first images of the current scene based on one or more exposure parameters of the current scene, the one or more first images of the current scene can be acquired based on the one or more exposure parameters through the configured first camera.
[0229] For example, in some embodiments, the first camera may include a main camera module for capturing images, which typically has high resolution and imaging quality, can output full-size image data (e.g., 16MP or higher), and has complete image signal processing capabilities. The first camera is responsible for actual image acquisition based on multiple sets of dynamically generated exposure parameters to ensure that high-quality original image data can be obtained under different lighting conditions. For example, when shooting night scenes, the first camera may use a long exposure time to capture background details, while when shooting moving objects, a short exposure time may be used to reduce blur.
[0230] In some embodiments of the present application, by using the configured first camera to capture multiple sets of images under different exposure parameters, image quality and dynamic range coverage can be improved. This ensures that each area is recorded under optimal exposure conditions, thereby avoiding local overexposure or underexposure, and thus improving the clarity and detail of the final composite image.
[0231] In summary, in the embodiments of the present application, by using the main camera to capture multiple first images according to dynamically generated exposure parameters, a wider brightness range can be covered, providing high-quality basic material for subsequent layered image fusion. Compared to shooting with fixed exposure parameters, this method can better adapt to imaging requirements under different lighting conditions.
[0232] In an embodiment of the present application, a second preview image of the current scene may be acquired through the configured first camera, and a third image of the current scene may be determined based on the second preview image.
[0233] In an embodiment of the present application, a third image of the current scene can also be acquired through a configured second camera.
[0234] For example, in some embodiments, the first camera can also provide preview images for scene perception and parameter prediction. This camera typically has high sensor performance and image processing capabilities, such as support for multi-frame synthesis and dynamic range expansion. In practical applications, the first camera can be the main camera of a mobile phone, the main lens of an in-vehicle camera, or the core imaging unit of industrial inspection equipment.
[0235] For example, in some embodiments, the second preview image may include a low-resolution image captured by the first camera for rapid scene recognition and preliminary processing. The preview image is typically output at a lower resolution (e.g., 1080p or 720p) to enable real-time scene perception without consuming excessive computing resources. During the HDR imaging process, the second preview image is used for tasks such as semantic segmentation and regionalized brightness difference index (RLDI) calculation, thereby providing a basis for subsequent exposure parameter prediction and image fusion.
[0236] Exemplarily, in some embodiments, the second camera may include a secondary camera to assist the first camera in completing scene perception and image processing tasks. In the present application, the second camera generally has a lower resolution (such as 2MP to 4MP), but has fast response capabilities and low power consumption characteristics. It can be used to shoot auxiliary images, provide motion vector estimation, update semantic segmentation results, and correct color deviations of the main camera. For example, in night portrait shooting, the second camera can provide low-resolution auxiliary frames for optical flow method to detect micro-motion of the character, thereby helping the main camera to perform motion compensation and partition fusion.
[0237] For example, in some embodiments, the third image can be an optimized image formed after preliminary processing, or a supplementary image generated by the second camera. The third image may contain information about moving objects, high dynamic range data for a local area, or an enhanced image of specific features. The third image works in conjunction with the image generated by the first camera to contribute to the final HDR image synthesis. For example, in a moving scene, the third image generated by the second camera can be used to distinguish static backgrounds from dynamic objects and guide the main camera in selecting the most appropriate exposure parameters and fusion strategy.
[0238] In summary, in the embodiments of the present application, after the primary camera captures the first image, a second preview image can be acquired to assist in generating the third image, or the secondary camera can directly generate the third image. That is, the first camera captures the second preview image and generates the third image based on its content, or the second camera captures the third image, which serves as supplementary information to enhance robustness and adaptability. This approach helps to timely update image content in dynamic scenes and improves real-time responsiveness.
[0239] In an embodiment of the present application, when performing layered image fusion based on one or more first images of the current scene to determine the second image of the current scene, the first type of image and the second type of image of the current scene are determined based on the one or more first images of the current scene; and image fusion is performed based on the first type of image and the second type of image to determine the second image of the current scene.
[0240] For example, in some embodiments, the classification of first and second type images is based on the semantic characteristics of the image content, lighting conditions, and dynamic / static properties. For example, the first type of image may contain high dynamic range areas (such as a backlit portrait background) and key semantic objects (such as faces and text), while the second type of image may contain low dynamic range areas or non-key objects (such as ordinary buildings). This classification facilitates subsequent differentiation processing, thereby improving the quality of HDR images.
[0241] In some embodiments, by classifying images into different types, more targeted exposure strategies and fusion algorithms can be developed, thereby improving overall image quality. Furthermore, due to the use of a lightweight semantic segmentation model, this method is more efficient in terms of resource consumption and is suitable for use on mobile devices.
[0242] For example, in some embodiments, different processing strategies are employed for the first and second image types during the image fusion process. For example, for high-dynamic regions in the first image type, a multi-frame weighted fusion approach may be employed to preserve more detail; whereas for low-dynamic regions in the second image type, a single-frame enhancement strategy may be employed to reduce computational burden and speed up processing. Ultimately, the two image types are combined through mask blending to generate a high-quality HDR image.
[0243] In other words, in the embodiments of this application, through differentiated image fusion strategies, problems such as ghosting and blurring that occur in traditional methods can be effectively avoided while ensuring overall image quality. In addition, the coordinated operation of the main and auxiliary cameras can further enhance robustness and adaptability, making it suitable for HDR imaging requirements in complex lighting and dynamic scenes.
[0244] For example, in some embodiments, the first image is divided into two categories: static background and dynamic objects, and fusion processing is performed separately. A multi-frame weighted fusion is first performed on the static region, and then combined with the single-frame enhancement results of the dynamic region to form a high-quality HDR image. This approach preserves the rich details of the background while clearly presenting the moving objects, achieving a clear distinction between static and dynamic images.
[0245] In some embodiments, image fusion can be performed based on the first type of image and the second type of image to determine the second image of the current scene, and fusion can be performed based on the first type of image to determine the first fused image. The second image of the current scene can be determined based on the first fused image and the second type of image.
[0246] For example, in some embodiments, the first type of image may include a first set of images with specific exposure parameters acquired during the imaging process, typically used to capture local high dynamic range areas or key semantic areas (such as faces, text, etc.) in a scene. For example, in night portrait photography, the first type of image may be an image acquired with a longer exposure time to preserve background details; while in a backlit text recognition scenario, the first type of image may be an image acquired with a shorter exposure time to ensure clear text edges.
[0247] For example, in some embodiments, the first fused image may include an intermediate result after preliminary fusion processing of the first type of image. The fusion process may be based on weighted averaging, gradient-aware weight mapping, or other algorithms, aiming to improve image quality and provide a basis for subsequent multi-frame fusion. Through this processing, a reference image with high information fidelity can be quickly obtained with low complexity, laying the foundation for the final HDR synthesis. That is, while reducing the consumption of computing resources, the overall quality of the image can be improved, especially retaining more detailed information in key areas (such as faces and text), thereby improving the user's subjective satisfaction with the imaging quality.
[0248] For example, in some embodiments, the second-type image may include images acquired during the imaging process with different exposure parameters, typically used to supplement the brightness range not covered by the first-type image or to enhance certain specific features. For example, in night portrait photography, the second-type image may be an image acquired with a medium or short exposure time to avoid overexposure and preserve details of the foreground subject; while in motion scenes, the second-type image may be a single-frame image acquired with a high shutter speed to eliminate motion blur.
[0249] For example, in some embodiments, the second image may include a final output image that combines the advantages of the first fused image and the second type of image. Through a layered fusion strategy, it preserves the rich details of the static background while ensuring the clarity of dynamic objects. This image can more comprehensively reflect the actual lighting conditions of the scene and meet user requirements for image quality and color reproduction.
[0250] In the embodiments of this application, by splitting the image fusion process into two steps—first generating a first fused image based on the first type of image, and then performing the final fusion with the second type of image—computational complexity can be reduced while simultaneously improving image quality. This effectively balances energy efficiency and imaging quality, thereby reducing redundant computations and extending device battery life, thus meeting the dual needs of mobile terminals for real-time imaging and high-quality output.
[0251] It can be seen that in some embodiments, through the above-mentioned layered fusion processing, the information of multiple frames of images can be effectively integrated, and the optimization processing of static and dynamic separation can be achieved, thereby significantly improving the visual quality and content integrity of the image, especially showing higher robustness and adaptability under complex lighting and dynamic scenes.
[0252] In summary, in this embodiment, a high-quality reference image is generated through the initial fusion of the first type of image, serving as the basis for subsequent fusion. Based on this, the second type of image is introduced, and a layered fusion strategy is used to further optimize the dynamic range and detail of the overall image. This two-stage fusion approach not only reduces the computational load but also improves the imaging quality, enabling efficient operation even on resource-constrained mobile devices, thus balancing performance and user experience.
[0253] In an embodiment of the present application, a fusion weight corresponding to the first type of image may be determined based on image information of the first type of image; and the first type of image may be fused based on the fusion weight to determine a first fused image.
[0254] For example, in some embodiments, image information may include comprehensive data on visual features such as brightness, color, and texture contained in the image. This information can be used to determine the image's clarity, contrast, and level of detail, thereby helping to determine the weight to be assigned to each image during the fusion process. For example, in an image containing both bright and dark areas, this image information can be used to identify which areas are more suitable as the primary reference image, thereby improving the quality of the final fused image.
[0255] For example, in some embodiments, the fusion weight can be a numerical representation of the contribution of each image to the final result during the fusion process of multiple images. This weight can be dynamically adjusted based on factors such as image clarity, noise level, and detail preservation. For example, between an image containing motion blur and a clear image, a higher fusion weight may be assigned to the clear image to ensure that the final image retains more effective information.
[0256] It can be understood that in the embodiments of the present application, by dynamically calculating the fusion weights based on image information, the present application can intelligently allocate the importance of each image in the fusion process according to its actual quality, thereby avoiding the problem of local detail loss or over-smoothing that may be caused by fixed weights.
[0257] For example, in some embodiments, the fusion weight determines the contribution ratio of each input image in the synthesis process. In practical applications, each image can be weighted averaged or otherwise combined according to its weight value based on the calculated fusion weight, thereby generating a higher quality fused image. For example, in HDR image processing, images shot with different exposure parameters will be assigned different fusion weights to ensure that the details of both bright and dark areas are reasonably preserved in the final image, which can reduce artifacts and distortion while ensuring image details, especially in scenes under complex lighting conditions, and can significantly improve the realism and visual effects of the image.
[0258] In some embodiments, when determining the second type of image of the current scene based on one or more first images of the current scene, the motion information corresponding to the current scene can be determined based on the third image of the current scene; one or more first images are aligned based on the motion information to obtain one or more aligned images; and the second type of image is determined based on the one or more aligned images.
[0259] For example, in some embodiments, the primary purpose of the third image is to provide dynamic information about the scene, such as object movement and camera shake, rather than to enhance the final image quality. For example, in a multi-camera system, the third image can be quickly captured at a lower resolution by the secondary camera to facilitate real-time analysis of scene motion. In a single-camera system, the third image may be obtained by downsampling the preview data from the primary camera.
[0260] For example, in some embodiments, motion information can be data about the motion of objects or the camera in the scene obtained by analyzing the third image, typically including parameters such as motion vectors, speed, and direction. This information is important for subsequent image alignment and fusion. For example, motion information can be used to estimate the relative displacement between adjacent frames, thereby eliminating ghosting or blurring caused by motion during multi-frame fusion. Furthermore, motion information can be used to distinguish static backgrounds from dynamic foregrounds, enabling differentiated processing in fusion strategies.
[0261] For example, in some embodiments, by extracting motion information from a third image, the accuracy of multi-frame image fusion can be effectively improved, especially by reducing the generation of ghosting and blurring artifacts in dynamic scenes, thereby improving the quality of the second type of image finally output.
[0262] For example, in some embodiments, the alignment process may include registering multiple first images in the spatial domain based on known motion information so that their geometric positions remain consistent, thereby facilitating subsequent fusion processing. The core of the alignment process is to eliminate image offsets caused by camera shake or object motion, ensuring that pixels in the same area of images with different exposures correctly correspond. For example, if the motion detection results indicate that an object has shifted by 5 pixels between two frames, the alignment process will adjust the displacement of the second frame image relative to the first frame image accordingly, so that the pixel distribution of the two in the target area tends to be consistent.
[0263] For example, in some embodiments, alignment processing can significantly improve the alignment accuracy of multiple frames of images in the spatial dimension, avoid image misalignment caused by motion or jitter, and thus enhance the clarity and naturalness of the final HDR composite image.
[0264] For example, in some embodiments, in actual applications, the selection of the second type of image depends on the specific scenario requirements. For example, in a night portrait scene, the second type of image may be an image composed of a long-exposure background and a short-exposure subject; while in a document copying scene, the second type of image may be an image fused and sharpened from multiple frames.
[0265] In summary, in the embodiments of the present application, by extracting motion information based on the third image, aligning the first image accordingly, and then generating the second type of image, the fusion quality of multiple frames in dynamic scenes can be effectively improved, thereby reducing the generation of ghosting and blurring artifacts, and thus achieving high-quality HDR image synthesis.
[0266] In some embodiments, when determining the second image of the current scene based on the first fused image and the second type of image, the second type of image can be preset processed based on the first fused image to obtain a processed image; based on the first fused image and the processed image, the second image of the current scene is determined; wherein the preset processing includes at least one or more of the following: local contrast enhancement processing; deblurring processing; brightening processing; and trusted multi-view classification processing.
[0267] It is understood that in the embodiments of this application, by performing preset processing such as local contrast enhancement and deblurring on the dynamic area, its clarity and detail expression can be significantly improved. Subsequently, the processed dynamic image is fused with the static background image to ultimately obtain a high-quality HDR image that can better meet the user's image quality requirements.
[0268] The embodiment of the present application proposes an image processing method that can dynamically determine one or more exposure parameters for the current scene based on the first preview image of the current scene and / or the multimodal sensor data of the current scene, and after using one or more exposure parameters for image acquisition, perform layered image fusion on the acquired image. In summary, the present application can realize dynamic partitioned exposure and dynamic partitioned fusion through preview images and / or multimodal sensor data, thereby optimizing the processing effect and processing efficiency of HDR, and achieving significant improvements in multiple dimensions such as dynamic range, energy efficiency, and application scenarios.
[0269] Based on the above embodiments, another embodiment of the present application proposes a multi-camera collaborative HDR imaging method and system based on scene perception and dynamic partitioning, which can improve HDR imaging quality and system efficiency.
[0270] There are many limitations to related HDR imaging technologies, which can be categorized as follows:
[0271] 1. Rigid exposure parameter selection mechanisms: Related technologies often rely on fixed exposure parameter sequences or simple statistical rules based on single-frame imagery, making them unable to dynamically adapt to complex and changing shooting scenarios. This approach can easily lead to partial overexposure or underexposure in extreme high-light-contrast scenes, requiring the capture of more redundant frames and increasing processing time and power consumption.
[0272] 2. Strong dependence on hardware architecture: Dual-camera solutions usually require specific hardware configuration (such as a main camera + a black and white secondary camera), which is costly and has poor compatibility, making it difficult to adapt to single-camera devices or heterogeneous camera combinations.
[0273] 3. Insufficient adaptability to dynamic scenes: Multi-frame HDR technology is prone to ghosting or blurring in dynamic scenes due to object movement or sudden changes in lighting. The related anti-shake algorithm is not optimized for fast-moving objects, resulting in a decrease in HDR image quality.
[0274] 4. Excessive consumption of computing resources: Related technologies require complex analysis of full-resolution images, resulting in high processing delays and high power consumption on mobile devices, and a poor user experience.
[0275] 5. Coarse granularity of scene recognition and optimization: Related technologies only recognize scenes at a global level and lack targeted optimization of local areas (such as faces and text), resulting in the loss of details in important areas.
[0276] For example, in a backlit portrait shooting scene, the related technology uses fixed exposure parameters, which easily leads to underexposure of the face area and overexposure of the background.
[0277] For example, in a night scene of shooting moving objects, when shooting a moving vehicle with related technology, the light trails are prone to have trailing shadows and the outline of the vehicle body is blurred.
[0278] For example, in a document reshoot scenario, related technologies fail to detect text areas, resulting in insufficient text edge contrast and a decreased OCR recognition rate.
[0279] To address the above issues, the multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiments of this application may include the following parts: The following core innovations are proposed to improve HDR imaging quality and system efficiency:
[0280] 1. Dynamic parameter prediction engine: This engine uses a lightweight semantic segmentation model to identify scene content (such as faces, sky, and text) in real time, combined with regionalized image information analysis, to dynamically generate multiple exposure parameter combinations.
[0281] 2. Multi-camera asynchronous collaborative architecture: The primary and secondary cameras utilize an asymmetric division of labor to reduce hardware dependencies and support single-camera or multi-camera projects. For example, the primary camera is responsible for high-quality image acquisition, while the secondary camera performs low-resolution scene perception and parameter pre-calculation. Single-camera projects can use the primary camera for preview downscale, while multi-camera projects can use the low-resolution secondary camera to feed into the scene detection algorithm.
[0282] 3. Single-frame and multi-frame layered fusion strategy: Differentiated fusion strategies are implemented for static backgrounds and dynamic objects to improve image quality and stability. For example, static areas use multi-frame weighted fusion based on the base frame, while dynamic areas use single-frame enhancement and motion compensation corresponding to the optimal exposure scene.
[0283] 4. Preview lightweight real-time refresh results: Introducing sparse processing and dynamic frame skipping strategies to ensure a smooth preview experience. Among them, whether it is a single-camera project using the main camera preview for scene recognition, or a multi-camera project using a low-resolution secondary camera for scene recognition, they all run in an independent pipeline and will not block the camera interface main thread drawing, ensuring a smooth preview experience; when using the main camera preview for scene recognition, the downscale mechanism is used to reduce the algorithm computation load; introducing sparse processing technology, only high-precision calculations are performed on key areas (such as semantic segmentation output), and a low-complexity algorithm is used for the global area; the scene detection algorithm has a built-in dynamic frame skipping strategy. If the difference between the previous and next frames is less than the threshold, the parameter prediction results are reused, that is, the subsequent frame can reuse one or more exposure parameters of the previous frame.
[0284] In the embodiments of the present application, on the one hand, dynamic semantic partitioning and multimodal sensor fusion are used to improve the ability to restore local details and enhance adaptability to extreme scenes. On the other hand, dynamic scheduling of multi-camera heterogeneous resources is used to optimize imaging efficiency and improve hardware compatibility. On the other hand, an adaptive partition fusion algorithm is used to eliminate artifacts and color casts, and improve naturalness and color restoration accuracy. On the other hand, scene-adaptive computing resource allocation is used to optimize energy efficiency and enhance the basic camera experience. On the other hand, a multi-camera collaborative fault-tolerant mechanism is used to improve robustness and ensure the success rate of taking photos in extreme scenes.
[0285] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiment of the present application can support HDR data sharing and joint optimization among multiple devices, and is suitable for scenarios such as panoramic live broadcast and security monitoring. It can integrate AI semantic segmentation and motion partitioning to improve partitioning accuracy and generalization capabilities. It can dynamically select fusion algorithms based on NIE values to improve real-time performance and image quality. It can be expanded to medical imaging, industrial inspection, autonomous driving and other fields. It is possible to develop CMOS sensors that support partitioned exposure to improve HDR performance. It is possible to combine AI prediction models with edge computing to achieve more efficient HDR processing.
[0286] It can be seen that the multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiment of the present application includes dynamic semantic partitioning, multi-dimensional semantic feature fusion, heterogeneous multi-camera collaboration, gradient perception fusion, dynamic computing resource allocation, robustness enhancement, etc., which comprehensively improves HDR imaging quality and system efficiency.
[0287] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiment of the present application can be applied to Figure 6 The image processing scenario shown includes the following parts:
[0288] 1. Start the camera to send a preview request;
[0289] 2. Return to the preview frame;
[0290] 3. The preview frame is displayed in the normal interface and a copy of the preview data is sent to the scene detection algorithm;
[0291] 4. The ASD algorithm returns the current scene perception and partitioning results to the app;
[0292] 5. ASD scene detection algorithm;
[0293] 6. Combine the ASD results and issue the most suitable exposure strategy for the current scene to different sensors and frames respectively;
[0294] 7. Return multiple frames with different exposure strategies;
[0295] 8. Send multiple frames of photos that are most suitable for the current scene exposure to the algorithm for post-processing;
[0296] 9. Algorithm post-processing;
[0297] 10. Use different algorithm processing strategies for static and moving areas, such as multi-frame fusion and single-frame enhancement;
[0298] 11. Determine whether the processing is normal; if the algorithm processing is normal, execute 10; otherwise, execute 11;
[0299] 12. The results of multi-frame algorithm processing are stored in the album;
[0300] 13. With fault-tolerant processing, the Base frame is saved in the album after passing the basic 3A module.
[0301] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiments of the present application may include scene perception and semantic segmentation, dynamic exposure parameter prediction, multi-camera asynchronous collaborative acquisition, and layered image fusion.
[0302] During scene perception and semantic segmentation, the input information includes downscaled primary camera data for single-camera projects and low-resolution secondary camera data for multi-camera projects. The output information includes a region mask that identifies the region type and boundaries.
[0303] In an embodiment of the present application, semantic segmentation can use a lightweight model (such as MobileNetV3+DeepLabV3) (parameter volume <2MB) to segment key areas (such as faces, sky, and text).
[0304] High light ratio detection, calculation of the regionalized brightness difference index (RLDI) of the preview frame, including but not limited to: calculating the mean brightness difference (ΔL_mean), variance contrast (C_var), and gradient discontinuity (G_dis) between regions.
[0305] High RLDI areas reflect complex lighting scenes (such as backlighting and strong contrast textures), triggering multi-exposure fusion and local tone mapping.
[0306] When predicting dynamic exposure parameters, the regional information amount is calculated, and the normalized information entropy (NIE) is calculated for each semantic region, that is, the regional information:
[0307] NIE = -Σ(p_i * log2(p_i)) / log2(N) (1)
[0308] Among them, p_i is the distribution probability of regional pixel values, and N is the total number of pixels.
[0309] High NIE areas (such as text) are marked as critical areas, and their exposure accuracy is prioritized.
[0310] For example, during the exposure parameter generation process, the input is: semantic region mask + NIE distribution + RLDI value, and the output is: three sets of exposure parameters (e.g., T1 = 1 / 100s, T2 = 1 / 30s, T3 = 1 / 10s, etc.).
[0311] Dynamically generate multiple sets of candidate exposure parameters:
[0312] For key areas: use NIE as weight to calculate the optimal exposure time:
[0313] T_opt = K * (1 / NIE) (2)
[0314] For non-critical areas: use binary iterative approximation to ensure coverage of the brightness range.
[0315] During asynchronous multi-camera collaborative acquisition, the main camera's task is to capture full-resolution images (e.g., 16MP) using the exposure parameters generated in step 2, outputting an image sequence {I1, I2, I3, I4, I5, ...}. For example, for a scene depicting a woman walking with a dog through a nighttime street scene, scene perception uses the AI scene perception of the corresponding frame to determine the optimal exposure parameters for each frame, avoiding compromise and ensuring a beautiful view of both people and scenery.
[0316] The first group: the most skin-friendly long exposure frame for girls: I1 frame
[0317] Group 2: Ultra-long exposure frame with the best night sky purity and details: frame I2
[0318] The third group: The most friendly ultra-short exposure frame for capturing dog sports: I3 frame
[0319] Group 4: Short-exposure frame with the best street light detail: Frame I4
[0320] Group 5: The most friendly medium exposure frame to the colors of the trees and flowers along the street: Frame 15
[0321] Secondary camera tasks: Synchronously capture low-resolution auxiliary frames (e.g., 2MP) for motion vector estimation (calculating the displacement of adjacent frames using optical flow) and real-time updating of semantic segmentation results (to cope with scene changes).
[0322] When fusion of layered images is performed, for the fusion of static backgrounds, multi-exposure weighted fusion (multi-frame Raw domain fusion algorithm) is used for static areas (such as the sky and buildings). The weight calculation is:
[0323] W_i = exp(-(L_i - L_target)^2 / σ2 ) (3)
[0324] Where L_target is the ideal brightness and σ is the preset value.
[0325] When fusion is performed on layered images, the motion vector field (MVF) is calculated using the auxiliary frame from the secondary camera to align moving objects across multiple frames. A single-frame enhancement strategy for underexposed frames is employed for moving areas: a single frame (such as frame I3) with exposure closest to the optimal NIE value is selected. Local contrast enhancement, deblurring, brightening, and TMC algorithms are applied to align the image quality of static areas.
[0326] Finally, the static fusion result and the dynamic enhancement result are mixed through a mask to output an HDR image with beautiful people and scenery, pure static scenes, clear moving objects, bright but not blown, dark but not black, and accurate color restoration.
[0327] Example 1: Night portrait scene (high light ratio + static subject), the subject is in a backlit environment (with a strong background light source), the face is underexposed (average brightness 30), and the background is overexposed (average brightness 220). The implementation steps are as follows:
[0328] Scene perception: The secondary camera captures preview frames, and the semantic segmentation model identifies "face" (15% area), "sky" (60%), and "building" (25%).
[0329] Calculate regionalized RLDI:
[0330] RLDI=(220-30) / (125+0.1)≈1.52>threshold 1.2, then HDR mode is triggered;
[0331] Dynamic exposure prediction:
[0332] The NIE of the key area (face) is calculated to be 0.85 (high information content), and the NIE of the non-key area (sky) is calculated to be 0.35.
[0333] Generate exposure parameters:
[0334] Face priority: T1=1 / 20s (ISO 800)
[0335] Global Balance: T2 = 1 / 50s (ISO 400)
[0336] Background suppression: T3 = 1 / 200s (ISO 200)
[0337] Multi-camera collaborative acquisition:
[0338] The main camera shoots 16MP triple frames (T1 / T2 / T3).
[0339] The secondary camera simultaneously captures 8MP auxiliary frames, and the optical flow method is used to detect slight movements of the person (displacement < 2 pixels).
[0340] Layered Fusion:
[0341] Static background (sky / buildings): weighted fusion of T2+T3 frames, weight formula: Wi = exp(-(Li-128)^2 / 50^2)
[0342] Dynamic subjects (faces): Select T1 frames and local CLAHE enhancement (block size 32x32, contrast limit 2.0).
[0343] Motion compensation: Sub-pixel alignment of the face area is performed based on the secondary camera optical flow data.
[0344] Effect verification: This solution can take into account both portraits and backgrounds, compared to the existing solution that can only choose between portraits and backgrounds.
[0345] Example 2: Backlit text scene (high dynamic details), before applying this solution, Figure 7 As shown, the document is placed next to a window, and the text area (brightness 10-50) coexists with the overexposed background outside the window (brightness 250). The implementation steps are as follows:
[0346] Semantic segmentation: Identify “text regions” (NIE=0.92, key regions) and “windows” (NIE=0.15).
[0347] Exposure parameter generation: Text priority: T1 = 1 / 10s (ISO 1600), T2 = 1 / 30s (ISO 800), T3 = 1 / 100s (ISO 400).
[0348] Dynamic fusion: Text area: fuse T1+T2 frames, using gradient-preserving weighting (weight is proportional to local gradient).
[0349] Window area: Use only T3 frames to avoid overexposure.
[0350] Effect verification: After applying this solution, if Figure 8 As shown in the following table:
[0351] Table 1
[0352] index Traditional methods This program Text readability (OCR) Recognition rate 78% Recognition rate 95% (↑22%) Bloom area 15% 3%(↓80%)
[0353] Example 3: Motion scene (dynamic object + complex background), running person (displacement speed 5 pixels / frame) and static forest background, uneven lighting, implementation steps are as follows:
[0354] Motion detection: The auxiliary frame optical flow method is used to detect the motion vector of the person (5 pixels / frame).
[0355] Layered processing:
[0356] Background: Multi-frame fusion (T1=1 / 50s, T2=1 / 100s).
[0357] Motion figures: Select a single frame (T2) and apply the mobile device platform's ISP to deblur, brighten, and align multiple frames using TMC.
[0358] Motion compensation: Correct motion blur in the human area based on optical flow data.
[0359] Effect comparison: This solution is superior to the original solution in terms of ghost area ratio and motion blur PSNR indicators.
[0360] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiment of the present application can realize dynamic scene perception and semantic partitioning, and improve the ability to restore local details. Among them, through multimodal sensor fusion (preview image, TOF depth information, ambient light sensor, Danxia original color lens can be freely matched and selected), real-time perception of scene type (such as backlit portrait, night scene architecture, indoor mixed light). A deep learning segmentation network is used to dynamically divide the image area (such as highlight area, dark area, face area, text area, sky area), and assign independent exposure strategies to different areas.
[0361] Based on this, precise local exposure control can be achieved, for example, preserving skin detail in facial areas (improving color reproduction and signal-to-noise ratio) and enhancing edge sharpness in text areas (improving OCR recognition rate). It can also achieve adaptability to extreme scenes. For example, in ultra-high dynamic range scenes (such as a backlit window at noon), dark noise is reduced and highlight overflow suppression is improved.
[0362] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiment of the present application can realize dynamic scheduling of multi-camera heterogeneous resources and optimize imaging efficiency. Among them, the acquisition tasks are dynamically allocated based on the hardware characteristics of the camera (such as the amount of light entering the wide-angle lens, the resolution of the telephoto lens, the TOF depth accuracy, and the color of the Danxia original color lens). The wide-angle lens is responsible for global scene analysis, the telephoto lens is responsible for super-resolution acquisition of key areas (such as faces), TOF assists in the calculation of partition exposure parameters, and the Danxia original color lens focuses on collecting color-related parameters for color restoration.
[0363] This maximizes resource utilization. For example, by leveraging the advantages of multiple sensors in traditional single-camera solutions, multi-camera collaboration can achieve more precise, differentiated scene exposure strategy control and algorithm processing. Hardware compatibility can also be enhanced, such as supporting heterogeneous systems with 4 to 6 cameras (such as multi-camera modules in mobile phones) and compatibility with sensors from different manufacturers.
[0364] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiment of the present application can realize an adaptive partition fusion algorithm to eliminate artifacts and color casts. Among them, through gradient perception weight mapping, combined with the previous scene segmentation and perception results, multi-frame fusion is adopted for static areas, single-frame optimization is adopted for moving areas, and the fusion weight is dynamically adjusted at the partition boundary to avoid the "ghosting" problem caused by traditional algorithms. Cross-camera color consistency calibration is introduced, and the white balance and color gamut deviation are corrected by linking the Danxia original color lens with the multi-camera ISP (image signal processor).
[0365] This improves the naturalness of fusion, such as improving the User Subjective Score (MOS) and reducing artifacts. It also enables accurate color reproduction, such as reducing Delta E color differences, which is closer to what the human eye sees.
[0366] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiment of the present application can realize scene-adaptive computing resource allocation and optimize the basic camera experience. Among them, the fusion mode is dynamically selected according to the complexity of the scene. Low complexity mode (such as uniform illumination): single-camera multi-frame synthesis optimizes power consumption; high complexity mode (such as night scene + moving object): multi-camera multi-frame + partition AI enhancement, giving priority to image quality. Utilize heterogeneous computing architecture (CPU + GPU + NPU) for parallel processing, and place the corresponding operations on the most suitable hardware core. For different sensors, various matching architectures are compatible according to the requirements of the project. For single-camera projects: the main camera preview data is downscaled and fed into the algorithm for processing; for dual-camera projects: the secondary camera small-size image is fed into the algorithm for processing; TOF and Danxia original color lens: the combination can make segmentation and color restoration more accurate. The parallel processing of the preview algorithm optimizes the preview smoothness without blocking the main preview display of the camera.
[0367] This allows for optimized energy efficiency. For example, while maintaining the same image quality, power consumption can be reduced compared to existing solutions by utilizing partitioned computing, downscale computing, and heterogeneous computing. Compatibility can also be improved, with the framework supporting a variety of camera combinations, from single to multiple cameras, allowing for customization based on project requirements.
[0368] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiments of this application can implement a multi-camera collaborative fault-tolerance mechanism and improve robustness. Specifically, a fault-tolerance mechanism for algorithm processing failure is designed: if the scene segmentation perception result is not available in the exposure strategy, a lucky exposure strategy is adopted. If anomalies occur in the post-processing multi-frame fusion and single-frame optimization algorithms, the necessary ISP algorithm is re-performed based on the reference frame to prevent image loss.
[0369] Based on this, enhanced imaging stability can be achieved, for example, the success rate of handheld shooting can be increased to 100%.
[0370] In summary, the multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiments of this application systematically breaks through the performance and effect bottlenecks of traditional HDR technology through scene semantic perception, dynamic partitioning exposure, multi-camera heterogeneous collaboration, and other processing, combined with intelligent fusion algorithms, energy efficiency optimization strategies, and exception handling strategies, achieving the following core advantages:
[0371] 1. Full scenario coverage: various user scenarios can be covered through targeted exposure strategies;
[0372] 2. Refined processing: Through scene recognition, a zoned exposure strategy is applied to different subject areas, ensuring that no subject is missed or sacrificed, ensuring that both the foreground and background look good.
[0373] 3. Improved hardware efficiency: Improved multi-camera collaborative utilization;
[0374] 4. Motion blur reduction: Precise exposure strategy and partition fusion algorithm can effectively reduce the blur of moving objects;
[0375] 5. Robustness in extreme environments: Increase HDR algorithm fault tolerance to ensure 100% image quality.
[0376] This systematically solves the inherent problems of existing HDR imaging technology and achieves improvements in dynamic range, energy efficiency, user experience and other dimensions.
[0377] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiment of this application can be widely used in smart phones, vehicle-mounted imaging, industrial inspection and other fields, and has counting barriers and market competitiveness.
[0378] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiments of this application includes the following extended applications of the multi-camera collaborative architecture:
[0379] Multi-camera module collaboration: supports the collaborative work of three or more cameras (such as wide-angle + telephoto + ToF), and optimizes the partitioning strategy based on depth information.
[0380] Cross-device collaboration: HDR data sharing and joint optimization between multiple devices such as mobile phones, drones, and surveillance cameras can be achieved through wireless communication (such as Wi-Fi P2P).
[0381] Technical implementation:
[0382] A distributed computing framework is introduced to allocate exposure testing, partition analysis, and fusion computing tasks to different devices.
[0383] Application scenarios: Panoramic HDR live broadcast of large-scale events, multi-view security monitoring.
[0384] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiments of this application includes the following enhanced and generalized extended applications of the dynamic partitioning algorithm:
[0385] Partition dimension expansion:
[0386] Semantic Partitioning: Integrates AI semantic segmentation models (such as U-Net) to identify semantic areas such as the sky, faces, and text, and optimizes exposure parameters in a targeted manner (for example, prioritizing facial areas to ensure skin tone restoration).
[0387] Motion Zoning: Combines optical flow to detect moving object areas and dynamically adjusts exposure timing to avoid motion blur.
[0388] RLDI algorithm upgrade:
[0389] Introducing multispectral RLDI: In addition to visible light brightness, it integrates infrared / ultraviolet spectral difference analysis to improve the partitioning accuracy under complex lighting conditions (such as haze and backlight).
[0390] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiments of this application includes the following extended applications:
[0391] Dynamic parameter adjustment of NIE driver:
[0392] Dynamically select the fusion algorithm based on the NIE value:
[0393] Low NIE area (simple information): fast weighted average method is used;
[0394] High NIE areas (complex textures): Enable multi-scale pyramid fusion or generative adversarial network (GAN) to restore details.
[0395] Real-time optimization:
[0396] Hardware acceleration: ISP chips or NPUs are used to implement parallel computing of RLDI / NIE to meet real-time HDR processing requirements.
[0397] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiments of this application has horizontal extensions of application scenarios including:
[0398] Medical imaging:
[0399] Endoscopy / microscopy scenarios: Multi-camera collaborative HDR is used to enhance the texture visibility of low-light tissues, and RLDI is combined with lesion areas (such as blood vessel rupture points).
[0400] Industrial testing:
[0401] Highly reflective metal surface defect detection: Multi-exposure fusion is used to eliminate reflective interference and improve defect recognition accuracy.
[0402] Autonomous driving:
[0403] On-board multi-camera HDR system: Optimizes imaging quality in scenes such as tunnel entrances and exits, and nighttime bright road signs through dynamic zoning.
[0404] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiments of this application includes the following collaborative innovations in hardware architecture:
[0405] Customized sensor design:
[0406] Develop a CMOS sensor that supports Region-based Exposure (RBE), allowing exposure parameters to be set independently for different regions within a single frame.
[0407] Multi-camera array design:
[0408] Ring multi-camera module: 6-8 cameras synchronously capture images with different exposures, combined with parallax correction to achieve 360° panoramic HDR.
[0409] The multi-camera collaborative HDR imaging method based on scene perception and dynamic partitioning proposed in the embodiments of this application integrates with emerging technologies including:
[0410] AI Joint Optimization:
[0411] Training RLDI / NIE prediction models: Use deep learning to predict scene dynamic range and information distribution, and plan exposure parameters in advance.
[0412] Generative HDR enhancement: HDR reconstruction of low dynamic range (LDR) images is performed using a diffusion model, and rationality is generated by combining multi-camera data constraints.
[0413] Edge computing and cloud collaboration:
[0414] Basic fusion is completed on the edge, and high-complexity repair (such as generating details in overexposed areas) is performed on the cloud, reducing the computing power requirements of the terminal.
[0415] The embodiment of the present application proposes an image processing method that can dynamically determine one or more exposure parameters for the current scene based on the first preview image of the current scene and / or the multimodal sensor data of the current scene, and after using one or more exposure parameters for image acquisition, perform layered image fusion on the acquired image. In summary, the present application can realize dynamic partitioned exposure and dynamic partitioned fusion through preview images and / or multimodal sensor data, thereby optimizing the processing effect and processing efficiency of HDR, and achieving significant improvements in multiple dimensions such as dynamic range, energy efficiency, and application scenarios.
[0416] Based on the above embodiment, in another embodiment of the present application, Figure 9 This is a schematic diagram of the structure of the image processing device proposed in the embodiment of the present application. Figure 9 As shown, the image processing device 110 proposed in this embodiment of the application may include:
[0417] An acquiring unit 1101 is configured to acquire a first preview image of a current scene and / or multimodal sensor data of the current scene;
[0418] a determining unit 1102, configured to determine one or more exposure parameters of the current scene based on the first preview image and / or the multimodal sensor data;
[0419] The acquiring unit 1101 is further configured to acquire one or more first images of the current scene based on one or more exposure parameters of the current scene;
[0420] The determining unit 1102 is further configured to perform layered image fusion based on the one or more first images of the current scene to determine a second image of the current scene.
[0421] In the embodiments of the present application, further, Figure 10 This is a schematic diagram of the structure of the electronic device proposed in the embodiment of the present application, such as Figure 10 As shown, the electronic device 120 proposed in the embodiment of the present application may include a processor 1201, a memory 1202, a communication interface 1203, and a bus 1204 for connecting the processor 1201, the memory 1202 and the communication interface 1203.
[0422] In an embodiment of the present application, the processor 1201 may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It is understandable that for different devices, the electronic device used to implement the above-mentioned processor function may also be other, and the embodiment of the present application is not specifically limited. The electronic device 120 may further include a memory 1202, which may be connected to the processor 1201, wherein the memory 1202 is used to store executable program code, the program code including computer operating instructions, and the memory 1202 may include a high-speed RAM memory, and may also include a non-volatile memory, for example, at least two disk memories.
[0423] In the embodiment of the present application, the bus 1204 is used to connect the communication interface 1203, the processor 1201 and the memory 1202, as well as the mutual communication between these devices.
[0424] In actual applications, the above-mentioned memory 1202 can be a volatile memory (volatile memory), such as random-access memory (Random-Access Memory, RAM); or a non-volatile memory (non-volatile memory), such as read-only memory (Read-Only Memory, ROM), flash memory (flash memory), hard disk drive (Hard Disk Drive, HDD) or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 1201.
[0425] Furthermore, in an embodiment of the present application, processor 1201 is used to: obtain a first preview image of the current scene and / or multimodal sensor data of the current scene; determine one or more exposure parameters of the current scene based on the first preview image and / or multimodal sensor data; obtain one or more first images of the current scene based on the one or more exposure parameters of the current scene; and perform layered image fusion based on the one or more first images of the current scene to determine a second image of the current scene.
[0426] In addition, the functional modules in this embodiment may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or software functional modules.
[0427] If the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0428] An embodiment of the present application provides a computer-readable storage medium having a program stored thereon, which implements the image processing method described above when executed by a processor.
[0429] Specifically, the program instructions corresponding to an image processing method in this embodiment may be stored on a storage medium such as a CD, a hard disk, or a USB flash drive. When the program instructions corresponding to an image processing method in the storage medium are read or executed by an electronic device, the following steps are included:
[0430] Acquire a first preview image of the current scene and / or multimodal sensor data of the current scene;
[0431] determining one or more exposure parameters for the current scene based on the first preview image and / or the multimodal sensor data;
[0432] acquiring one or more first images of the current scene based on one or more exposure parameters of the current scene;
[0433] A second image of the current scene is determined by performing layered image fusion based on one or more first images of the current scene.
[0434] The embodiment of the present application also provides a computer program product.
[0435] In some embodiments, the computer program product may include a computer program or instructions.
[0436] In some embodiments, the computer program product can be applied to the computer device in the embodiments of the present application, and the computer program instructions enable the computer to execute the corresponding processes implemented by the computer device in the various methods of the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0437] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0438] The present application is described with reference to the implementation flow charts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flow charts and / or block diagrams, as well as the combination of processes and / or boxes in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the implementation flow charts. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0439] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which is implemented in the implementation flow diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0440] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process described in the flowchart. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0441] The above description is merely a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application.
Claims
1. An image processing method, characterized in that: The method comprises: Acquiring a first preview image of a current scene and / or multimodal sensor data of the current scene; determining one or more exposure parameters of the current scene based on the first preview image and / or the multimodal sensor data; acquiring one or more first images of the current scene based on one or more exposure parameters of the current scene; Perform layered image fusion based on one or more first images of the current scene to determine a second image of the current scene.
2. The method according to claim 1, characterized in that The determining one or more exposure parameters of the current scene based on the first preview image and / or the multimodal sensor data includes: determining, based on the first preview image and / or the multimodal sensor data, regionalization information of one or more image regions corresponding to the first preview image; Determining an exposure strategy for the current scene based on the regionalized information of the one or more image regions; wherein the exposure strategy is used to determine a calculation method for exposure parameters and / or a number of exposure parameters for the current scene; Based on the exposure strategy of the current scene, one or more exposure parameters of the current scene are determined.
3. The method according to claim 2, characterized in that The method further comprises: Determining, based on the first preview image and / or the multimodal sensor data, one or more image regions corresponding to the first preview image and region parameters of the one or more image regions using a region segmentation model; The regional parameters include at least one or more of the following: a region mask of one or more image regions, a region type of one or more image regions, and a region boundary of one or more image regions.
4. The method according to claim 3, characterized in that The method further comprises: determining, based on the first preview image and / or the multimodal sensor data, a brightness difference index of one or more image regions corresponding to the first preview image; Based on the brightness difference index of the one or more image regions and a first index threshold, it is determined whether to determine the one or more exposure parameters.
5. The method according to claim 4, characterized in that The determining the exposure strategy of the current scene based on the regionalized information of the one or more image regions includes: determining a scene type corresponding to the current scene based on the regional parameters of the one or more image regions, and / or the brightness difference index of the one or more image regions, and / or the regionalization information of the one or more image regions; Based on the scene type, an exposure strategy for the current scene is determined.
6. The method according to any one of claims 1 to 5, characterized in that The obtaining of the first preview image of the current scene includes: Acquire an initial preview image of the current scene through a configured first camera, and downsample the initial preview image to obtain the first preview image; and / or, A first preview image of the current scene is acquired through the configured second camera.
7. The method according to any one of claims 1 to 5, characterized in that The acquiring multimodal sensor data of the current scene includes: Acquiring the multimodal sensor data via the configured multimodal sensor; in, The multimodal sensor includes at least one or more of the following: First camera; second camera; time-of-flight TOF sensor, ambient light sensor, Danxia original color lens.
8. The method according to any one of claims 1 to 5, characterized in that The acquiring one or more first images of the current scene based on one or more exposure parameters of the current scene includes: One or more first images of the current scene are acquired based on the one or more exposure parameters through the configured first camera.
9. The method according to claim 8, characterized in that The method further comprises: acquiring a second preview image of the current scene through a configured first camera, and determining a third image of the current scene based on the second preview image; and / or, A third image of the current scene is acquired through the configured second camera.
10. The method according to claim 1 or 9, characterized in that The performing layered image fusion based on the one or more first images of the current scene to determine the second image of the current scene includes: determining a first type of image and a second type of image of the current scene based on the one or more first images of the current scene; Image fusion is performed based on the first type of image and the second type of image to determine a second image of the current scene.
11. The method according to claim 10, characterized in that The performing image fusion based on the first type of image and the second type of image to determine the second image of the current scene includes: Performing fusion based on the first type of images to determine a first fused image; A second image of the current scene is determined based on the first fused image and the second type of image.
12. The method according to claim 11, characterized in that The first type of images includes a plurality of static images, and the fusing based on the first type of images to determine a first fused image includes: determining a fusion weight corresponding to the first type of image based on image information of the first type of image; The first type of images are fused based on the fusion weight to determine a first fused image.
13. The method according to claim 12, characterized in that The second-type image includes a plurality of dynamic images, and determining the second-type image of the current scene based on the one or more first images of the current scene includes: determining motion information corresponding to the current scene based on a third image of the current scene; performing alignment processing on the one or more first images based on the motion information to obtain one or more aligned images; The second type of image is determined based on the aligned one or more images.
14. The method according to any one of claims 11 to 13, characterized in that The determining, based on the first fused image and the second type of image, the second image of the current scene includes: performing preset processing on the second type of image based on the first fused image to obtain a processed image; determining a second image of the current scene based on the first fused image and the processed image; The preset processing includes at least one or more of the following: Local contrast enhancement processing; deblurring processing; brightening processing; trusted multi-view classification processing.
15. The method according to any one of claims 2 to 4, wherein: The method further comprises: determining pixel value distribution probabilities for one or more image regions based on the first preview image and / or the multimodal sensor data; Regionalization information of the one or more image regions is determined based on a pixel value distribution probability of the one or more image regions.
16. The method according to claim 4 or 5, wherein: The method further comprises: Determining brightness difference data of one or more image regions corresponding to the first preview image based on the first preview image and / or the multimodal sensor data; wherein the brightness difference data includes at least one or more of the following: brightness mean difference; variance contrast; and gradient discontinuity parameter; Based on the brightness difference data of the one or more image regions, a brightness difference index of the one or more image regions is determined.
17. An image processing device, characterized in that: The image processing device comprises: an acquiring unit, configured to acquire a first preview image of a current scene and / or multimodal sensor data of the current scene; a determining unit, configured to determine one or more exposure parameters of the current scene based on the first preview image and / or the multimodal sensor data; The acquisition unit is further configured to acquire one or more first images of the current scene based on one or more exposure parameters of the current scene; The determining unit is further configured to perform layered image fusion based on the one or more first images of the current scene to determine a second image of the current scene.
18. An electronic device, characterized in that: The electronic device includes a processor and a memory storing instructions executable by the processor. When the instructions are executed by the processor, the method according to any one of claims 1 to 16 is implemented.
19. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 16 is implemented.