Image processing method and device, electronic equipment and computer readable medium
By performing feature fusion and enhancement factor prediction on visible and non-visible light images, the problem of low image synthesis efficiency in low-light scenes is solved, and real-time image processing in embedded systems is realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ADDX (BEIJING) TECH CO LTD
- Filing Date
- 2025-01-22
- Publication Date
- 2026-07-24
AI Technical Summary
In low-light or black-light scenarios, image synthesis is computationally inefficient and cannot be processed in real time in embedded systems or other resource-constrained environments.
By acquiring visible light images and corresponding non-visible light images, feature fusion processing is performed to generate image fusion features. Based on these features, enhancement factors are predicted for image enhancement, thus decoupling the prediction of enhancement factors from the image enhancement process.
It improves the computational efficiency of image synthesis in low-light and black-light scenes, ensuring real-time processing in embedded systems or other resource-constrained environments.
Smart Images

Figure CN122453671A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to the field of computer technology, and more particularly to image processing methods, apparatus, electronic devices, and computer-readable media. Background Technology
[0002] Advances in computer vision technology have made high-quality image synthesis possible.
[0003] However, the following technical problems often arise when compositing images:
[0004] In low-light or black-light scenarios, image synthesis is computationally inefficient and cannot be processed in real time in embedded systems or other resource-constrained environments.
[0005] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0007] Some embodiments of this disclosure provide image processing methods, apparatuses, electronic devices, and computer-readable media to address one or more of the technical problems mentioned in the background section above.
[0008] In a first aspect, some embodiments of this disclosure provide an image processing method, the method comprising: acquiring a visible light image and a corresponding non-visible light image in the same scene; performing feature fusion processing on the visible light image and the non-visible light image to generate image fusion features; performing prediction processing on an enhancement factor corresponding to the visible light image based on the image fusion features to generate a predicted enhancement factor, wherein the predicted enhancement factor includes a nonlinear transformation parameter; and performing at least one image enhancement processing on the visible light image based on the predicted enhancement factor to generate an enhanced visible light image.
[0009] Secondly, some embodiments of this disclosure provide an image processing apparatus, comprising: an acquisition unit configured to acquire a visible light image and a non-visible light image corresponding to the same scene; a feature fusion unit configured to perform feature fusion processing on the visible light image and the non-visible light image to generate image fusion features; a prediction unit configured to perform prediction processing on an enhancement factor corresponding to the visible light image based on the image fusion features to generate a predicted enhancement factor, wherein the predicted enhancement factor includes a nonlinear transformation parameter; and an image enhancement unit configured to perform at least one image enhancement processing on the visible light image based on the predicted enhancement factor to generate an enhanced visible light image.
[0010] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0011] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0012] The above embodiments of this disclosure have the following beneficial effects: the image processing methods of some embodiments of this disclosure improve the computational efficiency of image synthesis in low-light and black-light scenes, ensuring real-time processing in embedded systems or other resource-constrained environments. Specifically, the reason for the low computational efficiency of image synthesis is that the computational efficiency of image synthesis is low in low-light or black-light scenes, making real-time processing impossible in embedded systems or other resource-constrained environments. Based on this, the image processing methods of some embodiments of this disclosure first acquire a visible light image and a corresponding non-visible light image in the same scene. Thus, visible light and non-visible light images of the same scene can be acquired. Second, feature fusion processing is performed on the visible light image and the aforementioned non-visible light image to generate image fusion features. Thus, the features of the visible light image and the infrared spectrum image can be fused at multiple levels. Then, based on the aforementioned image fusion features, the enhancement factor corresponding to the aforementioned visible light image is predicted to generate a predicted enhancement factor. Thus, the degree factor of enhancement of the visible light image can be determined through the fused features. Finally, based on the aforementioned predicted enhancement factor, the aforementioned visible light image is subjected to at least one image enhancement processing to generate an enhanced visible light image. Therefore, by decoupling the prediction of enhancement factors and the image enhancement process, the computational efficiency and scalability of the method are improved, the deployment difficulty is reduced, and the computational efficiency of image synthesis in low-light and black-light scenarios is improved, ensuring real-time processing in embedded systems or other resource-constrained environments. Attached Figure Description
[0013] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0014] Figure 1 This is a schematic diagram illustrating an application scenario of the image processing method according to some embodiments of this disclosure;
[0015] Figure 2 This is a flowchart of some embodiments of the image processing method according to the present disclosure;
[0016] Figure 3 This is a flowchart of some other embodiments of the image processing method according to the present disclosure;
[0017] Figure 4 This is a flowchart of some further embodiments of the image processing method according to the present disclosure;
[0018] Figure 5 These are schematic diagrams illustrating the structure of some embodiments of the image processing apparatus according to the present disclosure;
[0019] Figure 6 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0020] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0021] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0022] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0023] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0024] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0025] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0026] Figure 1 This is a schematic diagram illustrating an application scenario of an image processing method according to some embodiments of the present disclosure.
[0027] exist Figure 1In the application scenario, firstly, the computing device 101 can acquire a visible light image 102 and a corresponding non-visible light image 103 in the same scene. Secondly, the computing device 101 can perform feature fusion processing on the visible light image 102 and the non-visible light image 103 to generate image fusion features 104. Then, based on the image fusion features 104, the computing device 101 can perform prediction processing on the enhancement factor corresponding to the visible light image 102 to generate a predicted enhancement factor 105. Finally, based on the predicted enhancement factor 105, the computing device 101 can perform at least one image enhancement processing on the visible light image 102 to generate an enhanced visible light image 106.
[0028] It should be noted that the aforementioned computing device 101 can be either hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed within the hardware devices listed above. It can be implemented as, for example, multiple software programs or software modules used to provide distributed services, or as a single software program or software module. No specific limitations are made here.
[0029] It should be understood that Figure 1 The number of computing devices shown is merely illustrative. Any number of computing devices can be used depending on implementation needs.
[0030] Continue to refer to Figure 2 The diagram illustrates a flow 200 of some embodiments of an image processing method according to the present disclosure. This image processing method includes the following steps:
[0031] Step 201: Obtain the visible light image and the corresponding non-visible light image in the same scene.
[0032] In some embodiments, the entity executing the image processing method (e.g. Figure 1 The computing device 101 shown can acquire visible light images and corresponding non-visible light images of the same scene. Here, visible light and infrared spectral images of the target scene can be acquired in real time by an imaging device, or the visible light and infrared spectral images can be captured by the imaging device and then transmitted through a associated terminal. The target scene can represent a low-brightness scene.
[0033] Optionally, after step 201, the above visible light image is mapped to a preset color space to generate a mapped visible light image.
[0034] In some embodiments, the execution entity can map the visible light image to a preset color space to generate a mapped visible light image. The preset color space can be the RGB color space, used to divide the visible light image into different channels.
[0035] Step 202: Perform feature fusion processing on the visible light image and the non-visible light image to generate image fusion features.
[0036] In some embodiments, the aforementioned execution entity may perform feature fusion processing on visible light images and non-visible light images to generate image fusion features.
[0037] Optionally, before step 202, the visible light image and the non-visible light image are normalized to generate a normalized visible light image and a normalized non-visible light image.
[0038] In some embodiments, the execution entity may perform normalization processing on the visible light image and the non-visible light image to generate a normalized visible light image and a normalized non-visible light image. Here, since the value range of each pixel in the eight-bit stored RGB image is 0 to 255, the normalization process is to adjust the pixel value of each pixel in the visible light image to 1 / 255 of the original pixel value.
[0039] Step 203: Based on image fusion features, perform prediction processing on the enhancement factors corresponding to the visible light image to generate predicted enhancement factors.
[0040] In some embodiments, the execution entity may perform prediction processing on the enhancement factor corresponding to the visible light image based on the image fusion features to generate a predicted enhancement factor.
[0041] In practice, the enhancement factor corresponding to a visible light image can be predicted using the following steps:
[0042] The first step is to perform convolution processing on the above image fusion features to generate convolutional image fusion features.
[0043] The second step is to input the image fusion features after convolution into the hyperbolic tangent activation function to predict the enhancement factor and obtain the predicted enhancement factor.
[0044] In some alternative implementations of certain embodiments, the overall nonlinear transformation parameters for each channel of the visible light image can be pre-set as a prediction enhancement factor. The nonlinear transformation parameters can be obtained in advance based on different environmental and lighting conditions.
[0045] In some alternative implementations of certain embodiments, the visible light image can be downsampled to predict the transformation parameters at a lower resolution, and then upsampled to the original resolution to obtain the final nonlinear transformation parameters as a prediction enhancement factor.
[0046] Step 204: Based on the predicted enhancement factor, perform at least one image enhancement process on the visible light image to generate an enhanced visible light image.
[0047] In some embodiments, the execution entity may perform at least one image enhancement process on the visible light image based on the prediction enhancement factor to generate an enhanced visible light image.
[0048] In practice, based on the aforementioned prediction enhancement factor, each pixel in the visible light image can be iteratively processed to generate an iteratively processed visible light image as an enhanced visible light image. Here, the visible light image can be enhanced using the following formula:
[0049] F(x,α)=x+αx(1-x),α∈[-1,1], x∈[0,1],
[0050] Where x represents the pixels of the visible light image. α represents the prediction enhancement factor. In response to the completion of a single image enhancement process, the processed visible light image is used as the new visible light image, and image enhancement is performed again until the preset number of image enhancement processes is completed.
[0051] Optionally, after step 204, the following steps are also included:
[0052] The first step is to input the enhanced visible light image into a pre-trained image recognition model to obtain the image recognition result.
[0053] In some embodiments, the aforementioned executing entity may input the enhanced visible light image into a pre-trained image recognition model to obtain an image recognition result. The image recognition model may be a pre-trained model for content recognition of visible light images of a target scene. As an example, the image recognition model may be a pre-trained large-scale security recognition model.
[0054] The second step is to control the associated devices to perform the corresponding device operations in response to the image recognition results meeting the preset control conditions.
[0055] In some embodiments, the execution entity may control an associated device to perform a corresponding device operation in response to the image recognition result satisfying a preset control condition. The preset control condition indicates that the content displayed by the visible light image satisfies the control conditions of an associated device.
[0056] As an example, in response to the image recognition result indicating a fire, the device can be a fire sprinkler, and the preset control condition can be to identify the fire and control the fire sprinkler to perform a spraying operation for fire extinguishing.
[0057] The various embodiments of this disclosure have the following beneficial effects: the image processing methods of some embodiments of this disclosure improve the computational efficiency of image synthesis in low-light and black-light scenes, ensuring real-time processing in embedded systems or other resource-constrained environments. Specifically, the reason for the low computational efficiency of image synthesis is that the computational efficiency of image synthesis is low in low-light or black-light scenes, making real-time processing impossible in embedded systems or other resource-constrained environments. Based on this, the image processing methods of some embodiments of this disclosure first perform feature fusion processing on a visible light image and its corresponding infrared spectral image to generate image fusion features. This allows for multi-level fusion of features between the visible light image and the infrared spectral image. Second, based on the aforementioned image fusion features, a prediction processing is performed on the enhancement factor corresponding to the visible light image to generate a predicted enhancement factor. This allows the degree of enhancement of the visible light image to be determined through the fused features. Finally, based on the aforementioned predicted enhancement factor, at least one image enhancement processing is performed on the visible light image to generate an enhanced visible light image. Therefore, by decoupling the prediction of enhancement factors and the image enhancement process, the computational efficiency and scalability of the method are improved, the deployment difficulty is reduced, and the computational efficiency of image synthesis in low-light and black-light scenarios is improved, ensuring real-time processing in embedded systems or other resource-constrained environments.
[0058] Further reference Figure 3 This illustrates a flow 300 of another embodiment of the image processing method. Flow 300 of the image processing method includes the following steps:
[0059] Step 301: Obtain the visible light image and the corresponding non-visible light image in the same scene.
[0060] In some embodiments, the specific implementation of step 301 and its resulting technical effects can be found in [reference needed]. Figure 2 Step 201 in the corresponding embodiment will not be repeated here.
[0061] Step 302: Perform feature extraction processing on the visible light image and the non-visible light image respectively to generate the first image feature and the second image feature.
[0062] In some embodiments, the execution entity may perform feature extraction processing on the visible light image and the non-visible light image respectively to generate a first image feature and a second image feature.
[0063] Step 303: Input the first image features and the second image features into the pre-trained feature fusion model to obtain image fusion features.
[0064] In some embodiments, the execution entity can input the first image features and the second image features into a pre-trained feature fusion model to obtain image fusion features. The feature fusion model can be a depthwise separable convolutional model. The feature fusion model can include a first convolutional layer and a second convolutional layer. The first convolutional layer can be a 3x3 depthwise convolutional layer. The second convolutional layer can be a 1x1 pointwise convolutional layer. The stride and padding parameters of the 3x3 convolutional kernel are both set to 1, ensuring that the feature map size remains unchanged during calculation, and that the final predicted enhancement factor corresponds to the input image pixels, while also ensuring the alignment of infrared image features and RGB image features.
[0065] In practice, image fusion features can be obtained through the following steps:
[0066] The first step, based on the first image features and the second image features, is to perform the following fusion steps:
[0067] In the first fusion step, the first image features and the second image features are respectively input into the first convolutional layer of the feature fusion model to obtain the first convolutional features and the second convolutional features. Here, the calculation formula for the first convolutional layer can be expressed by the following formula:
[0068]
[0069] Where i and j represent pixel position indices, and c represents the channel position. H k W This represents the length and width of the convolution kernel. X represents the input feature map. This represents the convolution kernel applied to the c-th channel.
[0070] The second fusion step involves fusing the first and second convolutional features to generate fused convolutional features. This fusion process can be pixel-level channel-dimensional feature fusion.
[0071] The third fusion step involves inputting the fused convolutional features and the second convolutional features into the second convolutional layer of the feature fusion model to obtain the third and fourth convolutional features. Here, the calculation formula for the second convolutional layer can be expressed as follows:
[0072]
[0073] Where c′ represents the channel position index of the output result. Kp This represents a 1x1 convolution kernel matrix.
[0074] The fourth fusion step involves fusing the third and fourth convolutional features to generate a second-fusion convolutional feature.
[0075] The fifth fusion step updates the number of times the above fusion steps have been executed.
[0076] In the sixth fusion step, in response to the above-mentioned number of executions meeting the preset execution number threshold, the above-mentioned convolutional features after secondary fusion are determined as image fusion features.
[0077] The second step is to respond to the fact that the number of executions does not meet the preset execution threshold, and to use the convolutional features after secondary fusion as the first image features and the second convolutional features as the second image features, and to execute the fusion step again.
[0078] In the process of adopting technical solutions to solve the above-mentioned technical problems, the following technical problems often arise: during the training of the feature fusion model, due to the presence of a lot of noise and color cast in the visible light image, the fused features still have noise and color cast problems, which in turn leads to poor performance of the trained feature fusion model.
[0079] Alternatively, the feature fusion model described above can be trained using the following loss functions:
[0080] Image structure loss function for fusion:
[0081]
[0082] Among them, Loss spa Let Y represent the image structure loss function. Y represents the average intensity map of the enhanced visible light image. V represents the average intensity map of the visible light image. I represents the infrared image. i represents the current pixel. j represents the four pixels adjacent to the current pixel (top, bottom, left, and right). K represents the number of parts the image is divided into.
[0083] Exposure loss function:
[0084]
[0085] Among them, Loss exp Let M represent the exposure loss function. M represents the number of parts the image is divided into; here, 4×4 average pooling is used to downsample the image.
[0086] Color loss function:
[0087]
[0088] Where J represents the average value of a certain channel of the enhanced image.
[0089] Smoothing loss function:
[0090]
[0091] in, and These represent the gradients of the generated nonlinear transformation parameter matrix in the x and y directions, respectively.
[0092] The aforementioned loss functions, as an inventive point of this disclosure, combined with the optional steps below, solve the technical problem that "during the training of the feature fusion model, due to the presence of significant noise and color cast in visible light images, the fused features still suffer from noise and color cast, resulting in poor performance of the trained feature fusion model." The reasons for the poor performance of the feature fusion model are as follows: during the training of the feature fusion model, the presence of significant noise and color cast in visible light images results in the fused features still suffering from noise and color cast, thus leading to poor performance of the trained feature fusion model. Solving these factors can improve the fusion effect of the feature fusion model. To achieve this effect, this disclosure trains the feature fusion model by setting four loss functions. Here, the image structure loss function constrains the generated image to retain high-frequency texture information from the two input images; the exposure loss function utilizes the reflection intensity information of the infrared image, providing a smoother brightening effect for RGB images in black light scenes compared to the set average exposure value; the color loss function controls the continuity of the color channels in the enhanced image, improving color cast issues; and the smoothing loss function controls the smoothness of the enhancement factor predicted by the network, suppressing noise in the enhancement result. Combined with the optional steps below, the enhanced visible light image is input into a pre-trained image recognition model to obtain an image recognition result; in response to the image recognition result satisfying preset control conditions, the associated device is controlled to perform the corresponding device operation. Thus, by using the above loss functions, the noise and color cast issues in visible light images can be suppressed, thereby improving the performance of the trained feature fusion model. This allows for effective recognition of visible light images, enabling the control of associated devices to perform corresponding device operations and avoiding waste of device operating resources.
[0093] Step 304: Based on image fusion features, perform prediction processing on the enhancement factors corresponding to the visible light image to generate predicted enhancement factors.
[0094] Step 305: Based on the predicted enhancement factor, perform at least one image enhancement process on the visible light image to generate an enhanced visible light image.
[0095] In some embodiments, the specific implementation of steps 303-304 and the resulting technical effects can be found in [reference needed]. Figure 2 Steps 202-203 in the corresponding embodiments will not be repeated here.
[0096] Optionally, after step 304, the following steps are also included:
[0097] The first step is to input the enhanced visible light image into a pre-trained image recognition model to obtain the image recognition result.
[0098] In some embodiments, the aforementioned executing entity may input the enhanced visible light image into a pre-trained image recognition model to obtain an image recognition result. The image recognition model may be a pre-trained model for content recognition of visible light images of a target scene. As an example, the image recognition model may be a pre-trained large-scale security recognition model.
[0099] The second step is to control the associated devices to perform the corresponding device operations in response to the image recognition results meeting the preset control conditions.
[0100] In some embodiments, the execution entity may control an associated device to perform a corresponding device operation in response to the image recognition result satisfying a preset control condition. The preset control condition indicates that the content displayed by the visible light image satisfies the control conditions of an associated device.
[0101] As an example, in response to the image recognition result indicating a fire, the device can be a fire sprinkler, and the preset control condition can be to identify the fire and control the fire sprinkler to perform a spraying operation for fire extinguishing.
[0102] from Figure 3 It can be seen from this that, with Figure 2 Compared to the description of some corresponding embodiments, Figure 3 In some corresponding embodiments, the image processing method flow 300 further emphasizes the specific steps of performing feature fusion processing on the aforementioned visible light image and the corresponding infrared spectral image to generate image fusion features. Thus, the schemes described in these embodiments generate multi-level feature fusion image features, combining the reflectance intensity and texture information of infrared images, and decoupling the prediction of enhancement factors and image enhancement, thereby improving the computational efficiency and scalability of the method and reducing deployment difficulty.
[0103] Further reference Figure 4 The diagram illustrates a flow 400 of another embodiment of the image processing method. Flow 400 of the image processing method includes the following steps:
[0104] Step 401: Obtain the visible light image and the corresponding non-visible light image in the same scene.
[0105] Step 402: Perform feature fusion processing on the visible light image and the non-visible light image to generate image fusion features.
[0106] In some embodiments, the specific implementation of steps 401-402 and the resulting technical effects can be found in [reference needed]. Figure 2 Steps 201-202 in the corresponding embodiments will not be repeated here.
[0107] Step 403: Based on the visible light image, determine the video frame sequence corresponding to the visible light image.
[0108] In some embodiments, the executing entity may determine a video frame sequence corresponding to the visible light image based on the visible light image. The video frame sequence may be a video containing the visible light image.
[0109] Step 404: Perform keyframe extraction processing on the video frame sequence to generate a keyframe sequence.
[0110] In some embodiments, the execution entity may perform keyframe extraction processing on the video frame sequence to generate a keyframe sequence.
[0111] Step 405: Determine whether the visible light image is a keyframe.
[0112] In some embodiments, the execution entity may determine whether the visible light image is a keyframe.
[0113] Step 406: In response to determining the visible light image as a keyframe, the visible light image is input into a pre-trained nonlinear transformation parameter prediction model to obtain the predicted nonlinear transformation parameters as prediction enhancement factors.
[0114] In some embodiments, the execution entity may, in response to determining that the visible light image is a keyframe, input the visible light image into a pre-trained nonlinear transformation parameter prediction model to obtain the predicted nonlinear transformation parameters as prediction enhancement factors.
[0115] Optionally, after step 406, the following steps are also included:
[0116] The first step is to determine the first image corresponding to the visible light image in response to the determination that the visible light image is not a keyframe.
[0117] In some embodiments, the execution entity may determine a first image corresponding to the visible light image in response to determining that the visible light image is not a keyframe. The first image may be the keyframe in the keyframe sequence that is closest to the visible light image.
[0118] The second step is to use the nonlinear transformation parameters corresponding to the first image as the prediction enhancement factor.
[0119] In some embodiments, the execution entity may use the nonlinear transformation parameters corresponding to the first image as a prediction enhancement factor.
[0120] Step 407: Based on the predicted enhancement factor, perform at least one image enhancement process on the visible light image to generate an enhanced visible light image.
[0121] In some embodiments, the specific implementation of step 407 and its resulting technical effects can be found in [reference needed]. Figure 2 Step 204 in the corresponding embodiment will not be repeated here.
[0122] Further reference Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an image processing apparatus, which are similar to... Figure 2 Corresponding to the method embodiments shown, this image processing apparatus can be specifically applied to various electronic devices.
[0123] like Figure 5 As shown, an image processing apparatus 500 in some embodiments includes: an acquisition unit 501, a feature fusion unit 502, a prediction unit 503, and an image enhancement unit 504. The acquisition unit 501 is configured to acquire a visible light image and a corresponding non-visible light image within the same scene; the feature fusion unit 502 is configured to perform feature fusion processing on the visible light image and the corresponding infrared spectral image to generate image fusion features; the prediction unit 503 is configured to perform prediction processing on an enhancement factor corresponding to the visible light image based on the image fusion features to generate a predicted enhancement factor; and the image enhancement unit 504 is configured to perform at least one image enhancement processing on the visible light image based on the predicted enhancement factor to generate an enhanced visible light image.
[0124] It is understandable that the units described in the image processing apparatus 500 and the reference Figure 2 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the image processing apparatus 500 and the units contained therein, and will not be repeated here.
[0125] The following is for reference. Figure 6It illustrates an electronic device 600 suitable for implementing some embodiments of the present disclosure (such as...). Figure 1 A schematic diagram of the structure of the computing device 101 shown. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0126] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0127] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 6 Each box shown can represent a device or multiple devices as needed.
[0128] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined above in the methods of some embodiments of this disclosure.
[0129] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0130] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0131] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a visible light image and a corresponding non-visible light image of the same scene; perform feature fusion processing on the visible light image and the aforementioned non-visible light image to generate image fusion features; based on the aforementioned image fusion features, perform prediction processing on the enhancement factor corresponding to the aforementioned visible light image to generate a predicted enhancement factor, wherein the predicted enhancement factor includes nonlinear transformation parameters; and based on the aforementioned predicted enhancement factor, perform at least one image enhancement processing on the aforementioned visible light image to generate an enhanced visible light image.
[0132] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0134] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a feature fusion unit, a prediction unit, and an image enhancement unit. The names of these units do not necessarily limit the specific unit; for example, the acquisition unit may also be described as "a unit that acquires a visible light image and a corresponding non-visible light image of the same scene."
[0135] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0136] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. An image processing method, comprising: Acquire visible light images and corresponding non-visible light images of the same scene; The visible light image and the non-visible light image are subjected to feature fusion processing to generate image fusion features; Based on the image fusion features, the enhancement factor corresponding to the visible light image is predicted to generate a predicted enhancement factor, wherein the predicted enhancement factor includes a nonlinear transformation parameter. Based on the predicted enhancement factor, the visible light image is subjected to at least one image enhancement process to generate an enhanced visible light image.
2. The method according to claim 1, wherein, The method further includes: The enhanced visible light image is input into a pre-trained image recognition model to obtain the image recognition result; In response to the image recognition result satisfying the preset control conditions, the associated device is controlled to perform the corresponding device operation.
3. The method according to claim 1, wherein, The non-visible light images include infrared images, near-infrared images, and multispectral images.
4. The method according to claim 1, wherein, Before performing feature fusion processing on the visible light image and the non-visible light image to generate image fusion features, the method further includes: The visible light image and the non-visible light image are normalized to generate a normalized visible light image and a normalized non-visible light image.
5. The method according to claim 1, wherein, After acquiring the visible light image and the corresponding non-visible light image in the same scene, the method further includes: The visible light image is mapped to a preset color space to generate a mapped visible light image.
6. The method according to claim 1, wherein, The feature fusion processing of the visible light image and the non-visible light image to generate image fusion features includes: Feature extraction processing is performed on the visible light image and the non-visible light image respectively to generate first image features and second image features; The first image features and the second image features are input into a pre-trained feature fusion model to obtain image fusion features.
7. The method according to claim 1, wherein, The step of predicting the enhancement factor corresponding to the visible light image based on the image fusion features to generate a predicted enhancement factor includes: The image fusion features are convolved to generate convolved image fusion features; The convolutional image fusion features are input into the hyperbolic tangent activation function to predict the enhancement factor, thus obtaining the predicted enhancement factor.
8. The method according to claim 6, wherein, The feature fusion model includes a first convolutional layer and a second convolutional layer; And, the step of inputting the first image features and the second image features into a pre-trained feature fusion model to obtain image fusion features includes: Based on the first image features and the second image features, the following fusion steps are performed: The first image features and the second image features are respectively input into the first convolutional layer of the feature fusion model to obtain the first convolutional features and the second convolutional features; The first convolutional features and the second convolutional features are fused to generate fused convolutional features. The fused convolutional features and the second convolutional features are respectively input into the second convolutional layer of the feature fusion model to obtain the third convolutional features and the fourth convolutional features; The third and fourth convolutional features are fused together to generate a second-fusion convolutional feature. Update the number of times the fusion step has been executed; In response to the execution count meeting a preset execution count threshold, the convolutional features after secondary fusion are determined as image fusion features.
9. The method according to claim 8, wherein, The method further includes: In response to the fact that the number of executions does not meet the preset execution threshold, the convolutional features after secondary fusion are used as the first image features, and the features after the second convolution are used as the second image features, and the fusion step is executed again.
10. The method according to claim 1, wherein, The step of predicting the enhancement factor corresponding to the visible light image based on the image fusion features to generate a predicted enhancement factor includes: Based on the visible light image, determine the video frame sequence corresponding to the visible light image; The video frame sequence is subjected to keyframe extraction processing to generate a keyframe sequence; Determine whether the visible light image is a keyframe; In response to determining that the visible light image is a keyframe, the visible light image is input into a pre-trained nonlinear transformation parameter prediction model to obtain the predicted nonlinear transformation parameters as prediction enhancement factors.
11. The method according to claim 10, wherein, The step of predicting the enhancement factor corresponding to the visible light image based on the image fusion features to generate a predicted enhancement factor further includes: In response to determining that the visible light image is not a keyframe, a first image corresponding to the visible light image is determined; The nonlinear transformation parameters corresponding to the first image are used as prediction enhancement factors.
12. An image processing apparatus, comprising: The acquisition unit is configured to acquire a visible light image and a corresponding non-visible light image in the same scene; The feature fusion unit is configured to perform feature fusion processing on the visible light image and the non-visible light image to generate image fusion features; The prediction unit is configured to perform prediction processing on the enhancement factor corresponding to the visible light image based on the image fusion features to generate a predicted enhancement factor, wherein the predicted enhancement factor includes nonlinear transformation parameters. An image enhancement unit is configured to perform at least one image enhancement process on the visible light image based on the predicted enhancement factor to generate an enhanced visible light image.
13. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 11.
14. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 11.