An image processing method and an electronic device
By obtaining multi-spectral data to predict the color information of light sources and performing color restoration, the image color casting problem of electronic devices under the influence of colored light sources is solved, and the accurate restoration of image colors is achieved, and the user experience and image quality are improved.
Patent Information
- Application Number
- CN202411791851.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Electronic devices are susceptible to colored light sources when shooting images, resulting in color casts in the image, affecting image quality and user experience.
By obtaining multispectral data of the target object, predicting the color information of the light source and performing color reduction, multispectral data is collected using multispectral sensors, and combining the target light source prediction network and color reduction network, to achieve accurate image color reduction.
Improve the accuracy of image color restoration, reduce image color casting, and improve user's visual experience and image quality.
Smart Images

Figure CN119364200B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to an image processing method and an electronic device. Background Art
[0002] With the development of electronic device (such as mobile phone) technologies, the shooting function of electronic devices has also developed rapidly, and more and more users like to use electronic devices to shoot images (such as photos or videos).
[0003] However, during the process of an electronic device shooting an image, it may be affected by a colored light source (such as an LED lamp), resulting in color cast in the image captured by the electronic device. This not only affects the image quality but also affects the user's shooting experience. Therefore, how to reduce the occurrence of color cast in the image and achieve accurate color restoration of the image to improve the image quality and the user's shooting experience is an urgent problem to be solved. Summary of the Invention
[0004] Embodiments of this application provide an image processing method and an electronic device for achieving accurate color restoration of an image, thereby improving the image quality and the user's shooting experience.
[0005] To achieve the above objective, the embodiments of this application adopt the following technical solutions:
[0006] In a first aspect, an image processing method is provided, which is applied to an electronic device. In this method, the electronic device obtains a first image frame and multispectral data corresponding to the region where a target object is located in the first image frame. The reflectivity of the target object is within a preset reflectivity range. Then, the electronic device can predict the light source of the shooting environment of the first image frame based on the multispectral data corresponding to the region where the target object is located in the first image frame, and obtain the light source color information corresponding to the first image frame. Then, the electronic device can perform color restoration on the first image frame based on the light source color information corresponding to the first image frame, and obtain a second image frame.
[0007] In the embodiments of this application, since multispectral data can reflect the reflectivity of the target object and the light source color information, if the reflectivity of the target object is within the preset reflectivity range, it indicates that the reflectivity of the target object is relatively stable. The electronic device can perform color restoration on the first image frame based on the light source color information predicted from the multispectral data. In this way, accurate color restoration of the image can be achieved, so that the second image frame after color restoration can reflect the real picture content, reducing the occurrence of a low color restoration degree of the image caused by the inability to distinguish the color of the target object from the color of the ambient light source, thereby improving the display effect of the second image frame and the user's visual experience.
[0008] In a possible implementation of the first aspect, the process by which the electronic device predicts the light source color information corresponding to the first image frame may specifically include: The electronic device performs light source prediction on the first image frame based on the multispectral data of the target object in the first image frame to obtain the XYZ tristimulus values in the first color space. Subsequently, the electronic device can extract coordinates from the XYZ tristimulus values to obtain the x-y chromaticity coordinates. Among them, the x-y chromaticity coordinates are the light source color information in the first color space.
[0009] In the embodiments of the present application, since the XYZ tristimulus values in the first color space include the light source intensity and the x-y chromaticity coordinates, and the light source intensity is affected by the distance between the ambient light source and the user. Therefore, in order to ensure that the x-y chromaticity coordinates are not affected by other factors (i.e., the distance between the ambient light source and the user), the electronic device can extract coordinates from the XYZ tristimulus values to obtain the x-y chromaticity coordinates that can represent the light source color information. In this way, accurate prediction of the light source color information can be achieved, providing a basis for subsequent accurate color restoration of the image.
[0010] In a possible implementation of the first aspect, the process by which the electronic device performs color restoration on the first image frame may specifically include: The electronic device performs spatial conversion on the light source color information in the first color space to obtain the light source color information in the second color space. Among them, the color mode corresponding to the second color space is the same as the color mode corresponding to the first image frame. Subsequently, the electronic device can perform color restoration on the first image frame according to the light source color information in the second color space to obtain the second image frame.
[0011] Among them, the first image frame is an RGB image, the first color space is the CIE color space, and the second color space is the RGB color space.
[0012] In the embodiments of the present application, in order to ensure that the light source color information is consistent with the presentation form of the first image frame, the electronic device can perform spatial conversion on the light source color information in the CIE color space to obtain the light source color information in the RGB color space. In this way, accurate restoration of the RGB image can be achieved, enabling the second image frame after color restoration to reflect the real picture content, reducing the situation where the color restoration degree of the image is relatively low due to the inability to distinguish the color of the target object from the color of the ambient light source, and further improving the display effect of the second image frame and enhancing the user's visual experience.
[0013] In a possible implementation of the first aspect, the process by which the electronic device obtains the multispectral data corresponding to the region where the target object is located in the first image frame may specifically include: The electronic device obtains the multispectral data of the first image frame. Among them, each pixel point in the first image frame corresponds to a multispectral data. After that, the electronic device can perform region segmentation on the first image frame to obtain the region image where the target object is located in the first image frame. After that, the electronic device can obtain the multispectral data of the pixel points corresponding to the region image where the target object is located from the multispectral data of the first image frame.
[0014] In the embodiments of the present application, when the electronic device performs region segmentation on the first image frame, the region image where the target object is located in the first image frame can be obtained. In this way, it is convenient for the electronic device to extract the multispectral data of the region where the target object is located, providing a basis for subsequent accurate prediction of the light source color information.
[0015] In a possible implementation of the first aspect, when the first image frame is a portrait image frame and the target object in the portrait image frame includes human skin, the process by which the electronic device obtains the multispectral data corresponding to the region where the target object is located in the first image frame may specifically include: The electronic device performs region segmentation on the portrait image frame to obtain a skin mask image corresponding to the region where the human skin is located. After that, the electronic device can obtain the multispectral data of the pixel points corresponding to the skin mask image from the multispectral data of the portrait image frame.
[0016] In the embodiments of the present application, through the method of region segmentation, accurate segmentation of the human skin region can be achieved, and then the multispectral data of the region where the human skin is located in the first image frame can be obtained, providing a basis for subsequent accurate prediction of the light source color information of the human skin region.
[0017] In a possible implementation of the first aspect, the process by which the electronic device predicts the light source color information corresponding to the first image frame may specifically include: When the multispectral data of the pixel points in the region where the target object is located in the first image frame meets the preset conditions, the electronic device can perform light source prediction on the shooting environment of the first image frame according to the multispectral data of any pixel point in the region where the target object is located in the first image frame, and obtain the light source color information corresponding to the first image frame.
[0018] Among them, the preset conditions include: the difference between the multispectral data of any two pixel points in the region where the target object is located is less than the preset difference, and / or, the ratio of the number of pixel points with the same multispectral data to the total number of pixel points in the region where the target object is located is greater than the preset ratio.
[0019] In the embodiments of the present application, if the multispectral data of the pixel points in the area where the target object is located in the first image frame meets the preset conditions, it indicates that there is only a single light source in the shooting environment of the first image frame. Therefore, in order to reduce the waste of computing resources, the electronic device can predict the light source of the shooting environment of the first image frame only according to the multispectral data of any pixel point in the area where the target object is located in the first image frame. In this way, the utilization rate of computing resources can be improved.
[0020] In a possible implementation manner of the first aspect, the process by which the electronic device predicts the light source color information corresponding to the first image frame may further include: when the multispectral data of the pixel points in the area where the target object is located in the first image frame does not meet the preset conditions, the electronic device predicts the light source of the shooting environment of the target object in the first image frame according to the multispectral data of each pixel point in the area where the target object is located, so as to obtain the light source color information corresponding to the target object in the first image frame. After that, the electronic device restores the color of the area where the target object is located in the first image frame according to the light source color information corresponding to the target object in the first image frame, so as to obtain a second image frame.
[0021] In the embodiments of the present application, if the multispectral data of the pixel points in the area where the target object is located in the first image frame meets the preset conditions, it indicates that there may be multiple light sources in the shooting environment of the first image frame. Therefore, in order to improve the accuracy of predicting the light source color information, the electronic device can predict the light source of the shooting environment of the target object in the first image frame according to the multispectral data of each pixel point in the area where the target object is located. In this way, it provides a basis for accurately restoring the image color subsequently. In addition, the electronic device only performs light source prediction on the target object in the first image frame, so that color restoration can be carried out in a targeted manner, further improving the color restoration degree of the first image frame, reducing the situation where the color restoration of the first image frame is affected by the interference of the image background on the light source prediction accuracy, and improving the display effect of the first image frame.
[0022] In a possible implementation manner of the first aspect, the process by which the electronic device predicts the light source color information corresponding to the first image frame may specifically include: the electronic device inputs the multispectral data corresponding to the area where the target object is located in the first image frame into a target light source prediction network, so as to obtain the light source color information corresponding to the first image frame. The target light source prediction network is used to predict the relationship between the light source color information of the image frame and the multispectral data corresponding to the area where the target object is located in the image frame.
[0023] In the embodiments of the present application, predicting the light source color information corresponding to the first image frame through the target light source prediction network can achieve accurate prediction of the light source color information, making the predicted light source color information close to the real light source color information, and thus providing a basis for the subsequent color restoration of the first image frame.
[0024] In a possible implementation of the first aspect, the process for the electronic device to train the target light source prediction network may specifically include: The electronic device obtains a first sample data set. The first sample data set includes multiple groups of first sample data, and the first sample data includes the multispectral data of the region where the target object is located in the sample image frame and the true light source color information. The acquisition time of the multispectral data is the same as the acquisition time of the true light source color information. For each group of first sample data, the electronic device may input the multispectral data of the region where the target object is located in the sample image frame into the pre-constructed light source prediction network to obtain the predicted light source color information. Then, the electronic device may adjust the parameters of the pre-constructed light source prediction network according to the predicted light source color information and the true light source color information to obtain the target light source prediction network.
[0025] In the embodiments of the present application, through the multispectral data of the region where the target object is located in the sample image frame and the true light source color information, the target light source prediction network can be trained, which provides a basis for accurately predicting the light source color information subsequently. In addition, since the acquisition time of the multispectral data is the same as the acquisition time of the true light source color information, the situation where the target light source prediction network has a poor prediction effect on the light source due to asynchronous acquisition time can be reduced, and the accuracy of predicting the light source color information is improved.
[0026] In a possible implementation of the first aspect, before the electronic device obtains the first sample data set, it may further include: The electronic device may obtain multiple sample image frames and the multispectral data of each sample image frame from the imaging device, and obtain the true light source color information from the spectrometer. The imaging device is a device including a multispectral sensor, and the multispectral sensor is used to collect the multispectral data of the sample image frame. The spectrometer is placed opposite to the imaging device, and the distance between the spectrometer and the target object is less than the preset distance. Then, for each sample image frame, the electronic device may align the sample image frame with the multispectral data of the sample image frame to obtain the multispectral data corresponding to each pixel point in the sample image frame. Then, the electronic device may obtain the multispectral data corresponding to the region where the target object is located in the sample image frame from the multispectral data corresponding to each pixel point in the sample image frame.
[0027] In the embodiments of the present application, in the data alignment method of the electronic device, each pixel point in the sample image frame can correspond to a piece of multispectral data. In this way, accurate extraction of the multispectral data corresponding to the region where the target object is located can be achieved, providing a basis for accurately training the target light source prediction network subsequently.
[0028] In a possible implementation of the first aspect, the color attributes of the shooting environments corresponding to the above-mentioned multiple sample image frames are different. Among them, the color attributes of the shooting environment may include at least one of brightness, hue, and saturation.
[0029] In the embodiments of the present application, the electronic device can collect multiple first sample image frames with different color attributes, that is, the color attributes of the environmental light sources corresponding to different groups of first sample data are different. In this way, the target light source prediction network can predict the environmental light sources with different color attributes, ensuring the diversity of the light source color information prediction, improving the accuracy of the light source prediction, and thus providing a basis for subsequent color restoration of the image.
[0030] In a possible implementation of the first aspect, the above-mentioned multiple sample image frames include multiple image frames collected by the shooting device when there is one light source in the shooting environment, including the target object at one or multiple positions, and / or, when there are multiple light sources in the shooting environment, the image frames collected by the shooting device at different shooting positions, including the target object at the same position.
[0031] In the embodiments of the present application, if there is only one light source in the shooting environment, the imaging device can collect multiple image frames of the target object at one or multiple positions. And if there are multiple light sources in the shooting environment, the imaging device collects the image frames of the target object at the same position at different shooting positions. In this way, the target light source prediction network can predict the light source for the sample image frames at different shooting positions, ensuring the diversity of the light source color information prediction, improving the accuracy of the light source prediction, and thus providing a basis for subsequent color restoration of the image.
[0032] In a possible implementation of the first aspect, the process of the electronic device restoring the color of the first image frame may specifically include: the electronic device inputs the light source color information corresponding to the first image frame and the first image frame into the target color restoration network to obtain a second image frame; among them, the target color restoration network is adjusted based on the predicted restored image and the real image during the iterative training process. The predicted restored image is obtained by restoring the color of the image to be restored according to the image to be restored and the light source color information, and the real image is an image not affected by the colored light source.
[0033] In the embodiments of the present application, by using the target color restoration network to restore the color of the first image frame, accurate restoration of the image color can be achieved, making the first image frame (i.e., the second image frame) after color restoration more in line with the real scene and improving the user's shooting experience.
[0034] In a possible implementation of the first aspect, the above-mentioned real image is obtained by performing color restoration on the image to be restored through a preset processing method. The image to be restored is acquired by an imaging device. The light source color information is acquired by a spectrometer when the imaging device acquires the image to be restored.
[0035] In the embodiments of the present application, the real image can be obtained by performing color restoration on the image to be restored by using a preset processing method. In this way, the real restored image can be made more in line with the real shooting scene, that is, the target color restoration network trained with the real image can restore a scene that conforms to the real shooting scene, improving the user's visual experience.
[0036] In a possible implementation of the first aspect, the above-mentioned real image is acquired by the imaging device under target light. The target light is used to simulate natural light. The image to be restored is rendered according to preset color information based on the real image. The light source color information is the preset color information.
[0037] In the embodiments of the present application, since the real image is an image frame acquired by the shooting device under target light, and the target light can simulate natural light. Therefore, the situation where color deviation occurs due to light difference can be reduced, enabling the real image to display real colors. Then, the electronic device can render the real image according to the preset color information to obtain the image to be restored. In this way, the acquisition efficiency of the second sample data (i.e., the image to be restored, the light source color information, and the real image) can be improved, and further the training efficiency of the color restoration network can be improved.
[0038] In a second aspect, the present application provides an electronic device, which includes a camera, a multispectral sensor, a memory, and one or more processors; the camera, the multispectral sensor, the memory, and the processor are coupled; the camera is used to acquire a first image frame, the multispectral sensor is used to acquire multispectral data, the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device is caused to execute the method as described above.
[0039] In a third aspect, the present application provides a computer-readable storage medium, including computer instructions, which when running on an electronic device, cause the electronic device to execute the method as described above.
[0040] In a fourth aspect, the present application provides a computer program product, which when running on an electronic device, causes the electronic device to execute the method as described above.
[0041] Fifth aspect, a chip is provided, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is configured to execute the code in the memory. When the code is executed, the processor is configured to execute the method as described above.
[0042] Among them, for the beneficial effects that can be achieved by the electronic device described in the second aspect provided above, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the chip described in the fifth aspect, reference can be made to the beneficial effects in the first aspect and any possible design thereof, which will not be elaborated here. Description of the Drawings
[0043] Figure 1 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application;
[0044] Figure 2 It is a schematic diagram showing the front camera and the rear camera of a mobile phone provided by an embodiment of the present application;
[0045] Figure 3 It is a flowchart of an image processing method provided by an embodiment of the present application;
[0046] Figure 4 It is a schematic diagram of the interface of a mobile phone taking a photo provided by an embodiment of the present application;
[0047] Figure 5 It is a schematic diagram of the interface of taking a photo under the color restoration function provided by an embodiment of the present application;
[0048] Figure 6 It is a schematic diagram of a mobile phone performing semantic segmentation on the first image frame provided by an embodiment of the present application;
[0049] Figure 7 It is a schematic diagram of the multi-spectral data of a target object provided by an embodiment of the present application;
[0050] Figure 8 It is a schematic diagram of the structure of a target light source prediction network provided by an embodiment of the present application;
[0051] Figure 9 It is a schematic diagram of training a target light source prediction network provided by an embodiment of the present application;
[0052] Figure 10 It is a schematic diagram of training a target color restoration network provided by an embodiment of the present application;
[0053] Figure 11 It is a schematic diagram of performing color restoration on the first image frame provided by an embodiment of the present application. Detailed implementation manners
[0054] The following describes the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification and claims of the present application, the singular forms "a", "the", "above-mentioned", "this" and "this one" are also intended to include expressions such as "one or more", unless there is a clear indication to the contrary in the context. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" mean one or more than two (including two). The term "and / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist; for example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally means that the associated objects before and after are an "or" relationship.
[0055] The reference to "one embodiment" or "some embodiments" etc. described in this specification means that a specific feature, structure or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways. The term "connection" includes direct connection and indirect connection, unless otherwise stated. "First" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features.
[0056] In the embodiments of the present application, words such as "exemplarily" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0057] In some embodiments, during the process of taking pictures by an electronic device, the captured image frames may have color deviations due to some reasons (such as lighting, scenes, etc.), resulting in poor shooting effects of the electronic device, thereby affecting the user's shooting experience.
[0058] Exemplarily, for a portrait shooting scenario, the electronic device may be affected by a colored light source (such as an LED lamp), resulting in color cast in the captured portrait image frame and unable to present a visually appealing portrait. That is to say, the portrait effect in the portrait image frame captured by the electronic device may be different from the characteristics of the person in the actual scene. For example, the skin color of the person in the portrait image frame captured by the electronic device is darker than the real skin color of the person. Another example is that there is a difference between the color of the person's clothes in the portrait image frame captured by the electronic device and the real color of the person's clothes (for example, a pure white shirt appears grayish white in the image frame). It can be understood that if the portrait image frame generated by the electronic device cannot truly reflect the characteristics of the person, it will affect the presentation effect of the portrait image frame and further affect the user's visual experience.
[0059] In some embodiments, in order to be able to restore the image color and truly present the picture content, the electronic device can collect a first image frame through an RGB image sensor. Among them, the first image frame is an RGB image, and the RGB image can be composed of three color channels. The color channels can be the red (R) channel, the green (G) channel, and the blue (B) channel. That is to say, the electronic device can obtain the color information of the first image frame by superimposing the three color channels on each other. After that, the electronic device can perform color restoration on the first image frame according to the three color channel information corresponding to the first image frame to obtain a second image frame. That is, the second image frame generated by the electronic device is an image frame obtained after color restoration based on the first image frame.
[0060] However, since the electronic device cannot determine whether the color information of the target object in the first image frame is the color information of the target object itself or the color information generated due to the environmental light source based on the three color channel information corresponding to the first image frame, this results in the electronic device being unable to specifically restore the color of the target object itself, that is, unable to accurately perform color restoration, thereby limiting the color restoration ability of the electronic device and ultimately affecting the user's shooting experience.
[0061] Therefore, in order to improve the color restoration degree of an image frame and enhance the user's shooting experience, the embodiments of the present application provide an image processing method. In this method, in response to the user's shooting operation, the electronic device acquires a first image frame and multispectral data corresponding to the area where the target object is located in the first image frame. Among them, the reflectivity of the target object is within a preset reflectivity range. After that, the electronic device can predict the light source of the shooting environment of the first image frame based on the multispectral data corresponding to the area where the target object is located in the first image frame, and obtain the light source color information corresponding to the first image frame. After that, the electronic device can perform color restoration on the first image frame according to the light source color information corresponding to the first image frame, and obtain a second image frame.
[0062] In the embodiments of the present application, since the multispectral data can reflect the reflectivity of the target object and the light source color information, therefore, if the reflectivity of the target object is within the preset reflectivity range, it indicates that the reflectivity of the target object is relatively stable. The electronic device can perform color restoration on the first image frame according to the light source color information predicted from the multispectral data. In this way, accurate restoration of the image color can be achieved, so that the second image frame after color restoration can reflect the real picture content, and reduce the situation where the color restoration degree of the image is low due to the inability to distinguish the color of the target object from the color of the ambient light source. Furthermore, the display effect of the second image frame is improved, and the user's visual experience is enhanced.
[0063] In some examples, the electronic device in the embodiments of the present application may be a mobile phone, a tablet computer, a smart watch, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, and a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, etc., which include a camera and a multispectral sensor. The embodiments of the present application do not impose special restrictions on the specific form of the electronic device.
[0064] Exemplarily, Figure 1 shows a schematic hardware structure diagram of the electronic device 200. As Figure 1As shown, the electronic device 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 211, a power management module 212, a battery 213, antenna 1, antenna 2, a mobile communication module 240, a wireless communication module 250, an audio module 270, a sensor module 280, keys 290, a motor 291, an indicator 292, cameras 1-N 293, a display screen 294, and subscriber identification module (SIM) card interfaces 1-N 295, etc.
[0065] It can be understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may include more or fewer components than those shown, or combine certain components, or split certain components, or have different component arrangements. The components shown may be implemented in hardware, software, or a combination of software and hardware.
[0066] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0067] In some embodiments, the electronic device 200 may complete the image processing method provided in this application through the processor 210.
[0068] The wireless communication function of the electronic device 200 may be implemented through antenna 1, antenna 2, the mobile communication module 240, the wireless communication module 250, the modem processor, and the baseband processor, etc.
[0069] The electronic device 200 implements the display function through the GPU, the display screen 294, the application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 294 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.
[0070] The display screen (or called the screen) 294 is used to display images, videos, etc. For example, the display screen can be a touch screen, etc. In some embodiments, the electronic device 200 may include 1 or N display screens 294, where N is a positive integer greater than 1. In the embodiments of the present application, the display screen 294 can be used to display a preview interface, a shooting interface, etc. in the video recording mode.
[0071] The electronic device 200 can implement the shooting function through the ISP, the camera 293, the video codec, the GPU, the display screen 294, the application processor, etc.
[0072] The ISP is used to process the data fed back by the camera 293. For example, when the electronic device takes a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive element (or called the image sensor). The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 293. In some embodiments, the camera 293 includes a shutter. The shutter is a device in the camera used to control the time for light to irradiate the photosensitive element.
[0073] The camera 293 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device 200 may include 1 or N cameras 293, where N is a positive integer greater than 1.
[0074] In some embodiments, the camera 293 may include a lens, which is an optical component for generating an image.
[0075] Exemplarily, the above-mentioned N cameras 293 may include: one or more front cameras and one or more rear cameras. For example, please refer to Figure 2 , taking the above-mentioned electronic device 200 as a mobile phone as an example. Figure 2 There is a front camera shown on the interface of (a) in Figure 2 , such as the front camera 20. There are three rear cameras shown on the interface of (b) in
[0076] , such as the rear camera 21, the rear camera 22, and the rear camera 23. Of course, the number of cameras in the above-mentioned mobile phone includes but is not limited to the number described in the above embodiments.
[0077] In some embodiments, the camera 293 may further include an image sensor, and the image sensor is an RGB image sensor with RGB three channels. That is to say, the electronic device 200 can collect RGB images (i.e., the first image frames) through this RGB image sensor.
[0078] The sensor module 280 may include a multispectral sensor, etc. Among them, the multispectral sensor can capture more information about the light source, so that the electronic device 200 can capture image frames with accurate colors. Specifically, the multispectral sensor can capture spectral information (or spectral signals) in multiple bands from ultraviolet light to near-infrared light. That is to say, the multispectral sensor is a device for sensing spectral information in multiple wavelength ranges such as infrared rays, visible light, and ultraviolet rays.
[0079] In some embodiments, the multispectral sensor may be used to collect multispectral data when the electronic device 200 captures the first image frame. Among them, the multispectral data includes spectral information in multiple bands.
[0080] The image processing method of the embodiments of the present application can be used in the scenario where an electronic device captures images. For example, the electronic device captures a photo through the front camera or the rear camera of the electronic device. Hereinafter, taking the electronic device as a mobile phone as an example, the method of the embodiments of the present application will be described. Specifically, as Figure 3 shown, the image processing method may include S301 to S306.
[0081] S301, in response to the user's shooting operation, the mobile phone collects the first image frame and multispectral data. Among them, the first image frame includes the target object.
[0082] In some embodiments, after detecting a user's shooting operation, the mobile phone can collect a first image frame through an RGB image sensor and collect multispectral data through a multispectral sensor at the same time. That is to say, the acquisition time of the first image frame is the same as the acquisition time of the multispectral data. Among them, the first image frame may include a target object. The target object is an object with a relatively stable reflectivity, such as an object with a reflectivity within a preset reflection range. The preset reflection range can be preset according to the reflectivity of the type to which the target object belongs. The multispectral data may include spectral information of multiple bands. The multiple bands may be infrared, visible light, ultraviolet light, etc. respectively.
[0083] Exemplarily, the above-mentioned target object may be a human skin area, a facial skin area, green plants, inner page papers of the same color system (such as white, yellow, etc.) in a paper book, etc. As long as it is an object with a relatively stable reflectivity, it can be used as a target object, and no specific limitation is made. It can be understood that since most areas of the human face are skin, the target object can also be a facial skin area.
[0084] It should be noted that in order to facilitate the subsequent mobile phone to distinguish the multispectral data of the area where the target object is located in the first image frame from the multispectral data of the background area in the first image frame, so as to improve the accuracy of predicting the light source color information, the resolution of the multispectral sensor in the mobile phone needs to be greater than or equal to the preset resolution. It can be understood that a high-resolution multispectral sensor can not only obtain more spectral and spatial information, thereby providing more accurate multispectral data, reducing errors and interference, improving the reliability and accuracy of the multispectral data, but also capture richer color information, providing a basis for subsequent restoration of image colors.
[0085] In some embodiments of the present application, considering that the reflectivities of different types of objects will vary greatly, while the reflectivities of the same type of objects have a certain similarity, that is, the difference between the reflectivities of two objects belonging to the same type will not be too large. Therefore, when the reflectivity of the above-mentioned target object is within the preset reflection range, the mobile phone can determine that the target object belongs to the type to which the preset reflection range belongs. That is to say, the preset reflection range corresponds to the type of the target object, and the reflectivities of objects of the same type will all be within the same preset reflection range. Among them, the preset reflection range can be preset according to the reflectivities of objects of the same type.
[0086] Exemplarily, taking a portrait shooting scenario as an example, where the target objects include target object A and target object B. If the reflectivity of target object A is within the preset reflection range corresponding to the human skin area, and the reflectivity of target object B is also within the preset reflection range corresponding to the human skin area, it indicates that the difference between the reflectivity of target object A and that of target object B is relatively small, that is, there is similarity between the reflectivity of target object A and that of target object B. Therefore, the mobile phone can determine that target object A and target object B are of the same type of object, that is, both target object A and target object B are human skin.
[0087] In some cases, taking the above target object as human skin as an example, the reflectivities of different human skins have a certain degree of similarity. However, considering that the skin colors of different users may be different, that is to say, a user may have yellow skin or white skin. And the reflectivities of human skins with different skin colors may have relatively large differences. For example, the difference between the reflectivity of white skin and that of black skin is even greater. Therefore, the mobile phone can set different preset reflection ranges for human skins with different skin colors.
[0088] In one case, the above shooting operation can be an operation where the user clicks the shooting control in the camera mode. For example, please refer to Figure 4 , the mobile phone is displaying Figure 4 the first shooting interface shown in (a) of . Among them, the first shooting interface includes a shooting control 4A. If the shooting control 4A is clicked by the user, it means that the user wants to take a photo in the current scenario, that is, the user triggers the shooting operation. Therefore, the mobile phone can take a photo of the current scenario in the camera mode to obtain the first image frame. In one example, the mobile phone can automatically turn on the color restoration function by default in the camera mode. That is to say, in response to the shooting operation of the user in the camera mode, the mobile phone can perform color restoration on the collected image frame so that the generated image frame meets the user's visual experience. In another example, the mobile phone can not turn on the color restoration function by default in the camera mode. That is to say, in response to the shooting operation of the user in the camera mode, the mobile phone can directly store the collected image frame in the album, that is, the mobile phone does not need to perform color restoration on the collected image frame.
[0089] In another case, the above shooting operation can also be an operation where the user clicks the shooting control in the camera mode with the color restoration function enabled. For example, please refer to Figure 5 , user C holds the mobile phone to take a photo of user D. That is to say, user C points the mobile phone camera at user D for shooting. That is, the mobile phone can display Figure 5The camera interface in the camera mode shown in (a). Among them, the camera interface in this camera mode includes a preview image and a "More" control. The preview image includes user D and light source 501. After that, when a touch operation of the user on the "More" control is detected, in response to this touch operation, the mobile phone can display Figure 5 The settings interface shown in (b). Among them, this settings interface includes multiple function controls, and the function control can be a "slow motion" control, a "time-lapse" control, a "color restoration" control, etc. After that, when a touch operation of the user on the "color restoration" control is detected, in response to this touch operation, the mobile phone can turn on the color restoration function, and after the color restoration function is turned on, the mobile phone can display Figure 5 The camera interface under the color restoration function shown in (c). Among them, the camera interface under this color restoration function may include a preview image and a capture control 5A. After that, when a touch operation of the user on the capture control 5A is detected, in response to this touch operation, the mobile phone can use the color restoration function to capture in this scenario to generate a color-restored image frame (that is, Figure 5 The image 510 included in the interface shown in (d). After that, the mobile phone can save the color-restored image frame to the mobile phone album so that the user can view the color-restored image frame.
[0090] In another case, the above shooting operation can also be an operation where the user clicks the recording control in the video recording mode. That is to say, if a click operation of the user on the recording control is detected, the mobile phone can perform a video recording operation to generate a corresponding video file. It can be understood that this video file can be composed of multiple image frames. That is to say, if a click operation of the user on the recording control is detected, the mobile phone can perform color restoration on each captured image frame, and obtain a video file based on the multiple color-restored image frames. That is, each image frame in this video file is a color-restored image frame.
[0091] In some embodiments, if the above target object is the user, that is, the shooting scenario of the mobile phone is a portrait scenario, the above shooting operation can also be an operation where the user clicks the capture control in the portrait mode. For example, please refer to Figure 4 again, the mobile phone is displaying Figure 4 The second camera interface shown in (b). Among them, this second camera interface includes a capture control 4B. If the capture control 4B is clicked by the user, the mobile phone can capture the current scene in the portrait mode to obtain a first image frame.
[0092] It can be understood that the above first image frame is an image frame including the target object directly generated after the mobile phone camera takes a picture, that is to say, this first image frame is an image frame without color restoration. Exemplarily, the camera can be the front camera or the rear camera of the mobile phone. The number of cameras used by the mobile phone to take the first image frame can be one or more, and no specific limitation is made.
[0093] Correspondingly, the above first image frame can also be an image frame including the target object stored in the mobile phone album, that is to say, this first image frame can be an image frame obtained by the mobile phone camera taking a picture before, or an image frame downloaded by the mobile phone from a target application (APP), etc., and no specific limitation is made. Exemplarily, the target APP can be a social type APP, a news type APP, an entertainment type APP, etc. It should be noted that if the first image frame is an image frame stored in the mobile phone album, the multispectral data of the first image frame can be the data carried by the first image frame, that is, the first image frame can be stored in the mobile phone album together with the multispectral data.
[0094] In one implementation, the mobile phone will perform color restoration on the collected first image frame only when it receives the user's shooting operation. That is to say, during the process of the mobile phone displaying the preview image, it will not perform color restoration on the preview image. In this way, unnecessary power consumption losses can be reduced and the usage duration of the mobile phone can be improved.
[0095] S302. The mobile phone aligns the first image frame and the multispectral data to obtain the multispectral data corresponding to the first image frame.
[0096] In some embodiments, after the above first image frame and multispectral data are collected, the mobile phone can align the first image frame and the multispectral data to obtain the multispectral data corresponding to the first image frame. In this way, each pixel point in the first image frame can correspond to a piece of multispectral data, thereby providing a basis for subsequent prediction of the light source color information.
[0097] Specifically, after obtaining the multispectral data corresponding to the above first image frame, the mobile phone can obtain the multispectral data of the area where the target object is located in the first image frame from the multispectral data corresponding to the first image frame.
[0098] S303. The mobile phone performs region segmentation on the first image frame to obtain the region image where the target object is located.
[0099] In some embodiments, after the above first image frame is acquired, the mobile phone can segment the area where the target object is located in the first image frame to obtain the area image of the target object. Among them, the area image of the target object is used to represent the area where the target object is located in the first image frame.
[0100] Specifically, the mobile phone can perform area segmentation on the above first image frame according to the area segmentation network to obtain the area image of the target object. Among them, the area image can be represented by a mask image. The mask image is an image used to mark and segment specific targets or areas in the image. Specifically, when the mobile phone inputs the first image frame into the area segmentation network, the mask image of the above target object can be obtained.
[0101] In some cases, if the above first image frame is a portrait image frame, that is, the first image frame is an image frame including a target object of human skin, the mobile phone can perform area segmentation on the portrait image frame to obtain a skin mask image corresponding to the area where the human skin is located. Or, if the first image frame is a portrait image frame including a target object of facial skin, the mobile phone can perform area segmentation on the portrait image frame to obtain a facial mask image corresponding to the area where the facial skin is located.
[0102] In other cases, if the above first image frame is a green plant image frame, that is, the first image frame is an image frame including a target object of green plants, the mobile phone can perform area segmentation on the green plant image frame to obtain a mask image corresponding to the area where the green plants are located.
[0103] Exemplarily, as Figure 6 shown, the first image frame 601 is an image frame including a target object of human skin. When the mobile phone inputs the first image frame 601 into the area segmentation network, a skin mask image 602 corresponding to the area where the human skin is located can be obtained. It can be understood that the skin mask image 602 replaces the background and clothes in the first image frame 601 with black, and replaces the face and skin in the first image frame 601 with white. The size of the skin mask image 602 is the same as that of the first image frame 601.
[0104] It can be understood that by performing area segmentation on the above first image frame, the mobile phone can separate the foreground and background of the first image frame, that is, implement a matting technology for separating the target object from the background, which provides a basis for the mobile phone to predict the light source color information in the future.
[0105] It should be noted that the above process of data alignment and the process of region segmentation can be executed simultaneously or sequentially, and no specific limitation is made. For example, the mobile phone can execute the steps of S302 and S303 simultaneously. Another example is that the mobile phone can first execute the step of S303 and then execute the step of S302.
[0106] S304. The mobile phone obtains the multispectral data corresponding to the region image of the target object from the multispectral data corresponding to the first image frame.
[0107] Specifically, after obtaining the above region image of the target object, the mobile phone can extract the multispectral data of the target object from the multispectral data adapted to the first image according to this region image, that is, extract the multispectral data of the pixel points in the region where the target object is located in the first image frame. For example, taking the target object as human skin as an example, the mobile phone can only extract the multispectral data of the region where human skin is located in the first image frame from the multispectral data.
[0108] It should be noted that this application takes into account that the multispectral data can reflect the reflectivity of the target object and the light source color information of the shooting environment, that is, this multispectral data can be the product of the reflectivity and the light source color information. Therefore, if the reflectivity of the target object in the above first image frame is relatively stable, the mobile phone can use the target object as a reference object with a known reflectivity to predict the light source color information corresponding to the first image frame according to the multispectral data of the region where the target object is located in this first image frame. In this way, the accuracy of predicting the light source color information can be improved, and thus a basis is provided for restoring the color of the first image frame subsequently.
[0109] Exemplarily, taking the multispectral data of the above target object as Figure 7 the multispectral data shown for example, Figure 7 the curve a in is used to represent the multispectral data of band 1, the curve b is used to represent the multispectral data of band 2, and the curve c is used to represent the multispectral data of band 3. That is to say, the multispectral data of the target object can include the multispectral data of band 1, the multispectral data of band 2, and the multispectral data of band 3. It can be seen that Figure 7 the abscissa corresponding to each curve in is the band length, and the ordinate is the band intensity.
[0110] In some embodiments, after obtaining the above multispectral data of the target object, the mobile phone can delete all the multispectral data in the multispectral data corresponding to the first image frame except the multispectral data of the target object. That is, only the multispectral data of the target object in the first image frame is included in the mobile phone. In this way, the situation where useless multispectral data occupies the mobile phone memory can be reduced, and the utilization rate of the mobile phone memory space can be improved.
[0111] S305. The mobile phone predicts the light source of the shooting environment of the first image frame based on the multispectral data corresponding to the regional image of the target object, and obtains the light source color information corresponding to the first image frame.
[0112] Specifically, after obtaining the multispectral data corresponding to the regional image of the above-mentioned target object, the mobile phone can predict the light source of the shooting environment of the first image frame according to the multispectral data of the target object, so as to obtain the light source color information corresponding to the first image frame. Among them, the light source color information is used to characterize the light source of the environment where the target object is located when the mobile phone shoots the first image frame, that is, to characterize the light source of the shooting environment of the first image frame. The light source color information can be characterized by x-y chromaticity coordinates. The x-y chromaticity coordinates are used to characterize the color information of the ambient light source. The x-y chromaticity coordinates are a way to represent colors in the CIE color space (or the first color space). The CIE color space is used to formulate standards for color measurement. The CIE color space can define colors according to the three primary colors (X, Y, Z), or can also define colors according to the human eye's perception of colors (i.e., luminance (L), red component (a), and green component (b)).
[0113] In some cases, when the mobile phone predicts the light source of the shooting environment of the first image frame according to the multispectral data corresponding to the regional image of the above-mentioned target object, it can obtain the tristimulus values of XYZ in the CIE color space. Among them, the tristimulus values of XYZ are used to characterize the degree of stimulation of the human eye by the three primary colors. The tristimulus values of XYZ can include the light source intensity and x-y chromaticity coordinates. The light source intensity is affected by the distance between the ambient light source and the user. That is to say, if the distance between the ambient light source and the user is closer, the light source intensity is stronger. Therefore, in order to ensure that the above x-y chromaticity coordinates are not affected by other factors (i.e., the distance between the ambient light source and the user), the mobile phone can extract coordinates from the tristimulus values of XYZ to obtain the x-y chromaticity coordinates, that is, to obtain the light source color information in the CIE color space.
[0114] In one implementation, the mobile phone can determine the number of light sources existing in the shooting environment of the first image frame according to the multispectral data of the pixel points in the regional image. After that, when the number of light sources is one, the mobile phone can perform light source prediction on the shooting environment of the first image frame according to the multispectral data of any pixel point in the regional image, and obtain a light source color information that can represent the first image frame. In this way, the waste of computing resources can be reduced, and then the utilization rate of computing resources can be improved. When the number of light sources is multiple, the mobile phone can perform light source prediction on the shooting environment of the first image frame according to the multispectral data of each pixel point in the regional image, and obtain the light source color information of each pixel point of the target object in the first image frame, that is, obtain multiple light source color information that can represent the target object in the first image frame. In this way, the accuracy of light source color information prediction can be improved, providing a basis for accurately restoring the image color subsequently.
[0115] In some embodiments, the mobile phone can determine whether the multispectral data of the pixel points in the regional image meet a preset condition. If the multispectral data of the pixel points in the regional image meet the preset condition, it means that there is only a single light source in the shooting environment of the first image frame, and the mobile phone can determine that the number of the above light sources is one. If the multispectral data of the pixel points in the regional image do not meet the preset condition, it means that there may be multiple light sources in the shooting environment of the first image frame, and the mobile phone can determine that the number of light sources is multiple. Wherein, the preset condition may include that the difference between the multispectral data of any two pixel points in the regional image is less than a preset difference, and / or, the ratio of the number of pixel points with the same multispectral data to the total number of pixel points in the regional image is greater than a preset ratio.
[0116] It can be understood that if the difference between the multispectral data of any two pixel points in the regional image is less than the preset difference, it means that the multispectral data of each pixel point in the regional image are relatively close, that is, the multispectral data of each pixel point in the regional image can be approximately the same. And, if the ratio of the number of pixel points with the same multispectral data to the total number of pixel points in the regional image is greater than the preset ratio, it means that the number of pixel points with the same multispectral data in the regional image is large, that is, there is only a single light source in the shooting environment of the first image frame. Therefore, in order to reduce the waste of computing resources, the mobile phone can perform light source prediction on the shooting environment of the first image frame only according to the multispectral data of any pixel point in the regional image, and obtain a light source color information that can represent the first image frame.
[0117] In some other embodiments, the mobile phone can calculate the multispectral data of the pixel points in the above-mentioned regional image through statistical methods to obtain the target multispectral data. Among them, the statistical methods can include the method for solving the average value, the method for solving the mode, and the method for solving the median, etc. After that, when the differences between the target multispectral data and the multispectral data of each pixel point in the regional image are all less than the preset difference, the mobile phone can determine that the number of the above light sources is 1. When the difference between the target multispectral data and the multispectral data of any pixel point in the regional image is greater than or equal to the preset difference, the mobile phone can determine that the number of light sources is multiple.
[0118] In some other embodiments, the mobile phone can form at least one set of pixel points with the same multispectral data in the above-mentioned regional image. After that, the mobile phone can determine the target set of pixel points with the largest number of pixel points from the at least one set of pixel points. After that, when the ratio between the number of pixel points in the target set of pixel points and the total number of pixel points in the regional image is greater than the preset ratio, and / or, the differences between the multispectral data of the target pixel points in the target set of pixel points and the multispectral data of each pixel point in the regional image are all less than the preset difference, the mobile phone can determine that the number of the above light sources is 1. When the ratio between the number of pixel points in the target set of pixel points and the total number of pixel points in the regional image is less than or equal to the preset ratio, or, the difference between the multispectral data of the target pixel points in the target set of pixel points and the multispectral data of any pixel point in the regional image is greater than or equal to the preset difference, the mobile phone can determine that the number of light sources is multiple.
[0119] In one implementation manner, the mobile phone can predict the light source of the shooting environment of the first image frame according to the target light source prediction network and the multispectral data corresponding to the area where the target object is located in the first image frame. Among them, the target light source prediction network is used to predict the relationship between the light source color information of the image frame and the multispectral data corresponding to the area where the target object is located in the image frame, so that the predicted light source color information is close to the real light source color information, thereby providing a basis for the subsequent color restoration of the first image frame. Exemplarily, as Figure 8 shown, the target light source prediction network can be a deep neural network (DNN), and the deep neural network can include an input layer, multiple hidden layers, and an output layer. The input layer is used to convert the input data into a format that can be processed inside the neural network, that is, to convert the first image frame into a vector form. The hidden layer is used to convert the input data into a higher-level feature representation. The output layer is used to output the processing result of the input data.
[0120] Specifically, the mobile phone can input the multispectral data of the target object corresponding to the area in the first image frame into the target light source prediction network to obtain the light source color information corresponding to the first image frame. It can be understood that if there is only a single light source in the shooting environment of the first image frame, it means that the multispectral data of each pixel point in the first image frame is the same, that is, the light source color information of the target object in the first image frame is the same as the light source color information of the first image frame. Therefore, the mobile phone can input the multispectral data of any pixel point in the region image into the target light source prediction network to obtain the light source color information of the first image frame. If there are multiple light sources in the shooting environment of the first image frame, it means that the multispectral data of the pixel points in the first image frame is different, that is, the light source color information of the target object in the first image frame may be different from the light source color information of other regions in the first image frame. Therefore, in order to achieve accurate restoration of image color, the mobile phone can input the multispectral data of each pixel point in the region image into the target light source prediction network to obtain the light source color information of each pixel point of the target object in the first image frame.
[0121] In some embodiments, such as Figure 9 shown, taking the above-mentioned target light source prediction network as an example of a neural network for predicting the light source color information of the human skin area in the first image frame, the training process of the target light source prediction network can specifically include: The mobile phone can obtain a first sample data set. Among them, the first sample data set includes multiple groups of first sample data, and each group of first sample data can include the multispectral data of the area where the human skin is located in the sample image frame and the true light source color information. The sample image frame is an image frame including the target object as the human skin. The multispectral data is collected by the shooting device through a multispectral sensor. The true light source color information is the light source color information collected by the spectrometer when the shooting device collects the sample image frame. The spectrometer is a scientific instrument that decomposes light with complex components into spectral lines. The spectrometer is used to record the true light source color information near the user when the shooting device collects the sample image frame.
[0122] It can be understood that the above-mentioned spectrometer is placed opposite to the shooting device, and the distance between the spectrometer and the target object can be less than a preset distance. That is to say, when the shooting device collects the sample image frame and the multispectral data, the spectrometer can collect the above-mentioned true light source color information at the same time. Among them, the shooting device is a device including a multispectral sensor, and the multispectral sensor is used to collect the multispectral data of the sample image frame. Exemplarily, the shooting device can be a mobile phone, a computer, etc., as long as it is a device equipped with a multispectral sensor, and there is no specific limitation.
[0123] It should be noted that if there is one or more light sources in the shooting environment of the above sample image frame, the shooting device can collect only one sample image frame and the corresponding true light source color information of the sample image frame, so as to generate a set of first sample data. If there are multiple light sources in the shooting environment of the first sample image frame, the shooting device can collect multiple sample image frames of the target object at the same position and the corresponding true light source color information of each sample image frame according to different shooting positions, so as to generate multiple sets of first sample data. In this way, the mobile phone can train the light source prediction network according to different light source color information, ensuring the diversity of the first sample data and providing a basis for the mobile phone to accurately predict the light source color information in the future. Among them, the number of collected sample image frames is positively correlated with the number of light sources in the shooting environment. That is, the more light sources there are in the shooting environment, the more sample image frames will be collected. For example, if the number of light sources in the shooting environment is 2, the number of collected sample image frames can be 4, 6, etc., as long as the number of collected frames is greater than the number of light sources, and the specific number is not limited.
[0124] In one implementation, before obtaining the first sample data set, the mobile phone can obtain multiple sample image frames and the multispectral data of the sample image frames from the shooting device. Then, for each sample image frame, the mobile phone can align the sample image frame with the multispectral data of the sample image frame to obtain multispectral data adapted to the sample image frame, that is, obtain the multispectral data corresponding to each pixel point in the sample image frame. Then, the mobile phone can obtain the multispectral data of the pixel points in the area where the human skin is located from the multispectral data adapted to the sample image frame, that is, obtain the multispectral data corresponding to the area where the human skin is located in the sample image frame.
[0125] In another implementation, the multispectral data corresponding to the area where the human skin is located in the above sample image frame can also be directly obtained from the shooting device. That is, after the shooting device extracts the multispectral data corresponding to the area where the human skin is located in the sample image frame from the multispectral data adapted to the sample image frame, it can directly send the multispectral data corresponding to the area where the human skin is located in the sample image frame to the mobile phone.
[0126] In some cases, to ensure the diversity of the prediction of the light source color information, the mobile phone can collect multiple first sample image frames with different color attributes, that is, the color attributes of the ambient light sources corresponding to different groups of first sample data are different. In this way, the target light source prediction network can perform predictions for ambient light sources with different color attributes, improving the accuracy of the light source prediction, and thus providing a basis for subsequent restoration of the image color. Among them, the color attributes can include brightness, hue, saturation, etc. Among them, brightness is used to represent the light and dark degree of the color. Hue is used to represent the type of color (such as red, green, blue, etc.). Saturation is used to represent the purity and intensity of the color in the color.
[0127] In other cases, to ensure the diversity of the prediction of the light source color information, the mobile phone can also obtain the sample image frames and the real light source color information at different shooting positions. That is, the area of the human skin in the sample image frames in different groups of first sample data is the same, but the shooting positions of the imaging devices for collecting the sample image frames are different. In this way, the target light source prediction network can perform light source predictions for the sample image frames at different shooting positions, improving the accuracy of the light source prediction, and thus providing a basis for subsequent restoration of the image color.
[0128] Specifically, as Figure 9 shown, after obtaining the above first sample database, for each group of first sample data in the first sample database, the mobile phone can input the multispectral data of the area of the human skin in the sample image frame into the pre-constructed light source prediction network to obtain the predicted light source color information. Then, the mobile phone can adjust the parameters of the light source prediction network according to the predicted light source color information and the above real light source color information to obtain a trained light source prediction network (or called the target light source prediction network).
[0129] In some embodiments, as Figure 9 shown, the mobile phone can determine the loss value of the light source prediction network according to the predicted light source color information and the real light source color information. Then, the mobile phone adjusts the parameters of the light source prediction network according to the loss value of the light source prediction network until the light source prediction network meets the first training requirement to obtain the target light source prediction network. Among them, the first training requirement can include that the loss value of the light source prediction network is less than the second preset loss value, or the predicted light source color information is the same as the real light source color information, or the first training times reach the first preset times. The first training times are the times when the mobile phone trains the light source prediction network.
[0130] It should be noted that the training process of the above-mentioned target light source prediction network can also be executed by a cloud service. That is to say, after the cloud service finishes training the light source prediction network, that is, obtains the target light source prediction network, it can send the target light source prediction network to the mobile phone. After that, the mobile phone can predict the light source color information of the first image frame according to the target light source prediction network.
[0131] S306. The mobile phone performs color restoration on the first image frame according to the light source color information corresponding to the first image frame to obtain a second image frame.
[0132] Specifically, after obtaining the light source color information corresponding to the first image frame, the mobile phone can perform color restoration on the first image frame according to the light source color information corresponding to the first image frame to obtain the first image frame after color restoration, that is, obtain the second image frame.
[0133] In some embodiments, when obtaining one light source color information, it indicates that there is only a single light source in the shooting environment of the first image frame. The mobile phone can perform color restoration on the entire first image frame according to the one light source color information. It can be understood that if there is only a single light source in the shooting environment of the first image frame, it means that the multispectral data of each pixel point in the first image frame is the same, that is, the light source color information of the target object can be equivalent to the light source color information of the first image frame. Therefore, the mobile phone can perform color restoration of the entire image only according to one light source color information, improving the color restoration accuracy of the image.
[0134] In other embodiments, when obtaining multiple light source color information, it indicates that there are multiple light sources in the shooting environment of the first image frame. The mobile phone can perform color restoration on the target object in the first image frame according to the multiple light source color information. It can be understood that if there are multiple light sources in the shooting environment of the first image frame, it means that the multispectral data of each pixel point in the first image frame may be different, that is, the light source color information of the target object may be different from the light source color information of other regions in the first image frame. Therefore, to ensure the accuracy of color restoration, the mobile phone can perform color restoration on the target object in the first image frame according to the multiple light source color information. In this way, the color restoration accuracy of the target object in the first image frame can be achieved, reducing the occurrence of color restoration errors in other regions of the first image frame, ensuring the color restoration accuracy of the first image frame, and thus improving the user's visual experience.
[0135] It can be understood that the above light source color information is a representation form in the CIE color space, and the above first image frame is an RGB image. Therefore, in order to ensure the consistency of the light source color information and the representation form of the first image frame, the mobile phone can perform a spatial conversion on the light source color information in the CIE color space to obtain the light source color information in the RGB color space (or referred to as the second color space). Among them, this spatial conversion refers to converting the light source color information from the CIE color space to the RGB color space. This RGB color space is used to generate colors. The CIE color space can generate colors by superimposing different proportions of red, green, and blue colors. After that, the mobile phone can perform color restoration on the first image frame according to the light source color information in this RGB color space.
[0136] In some cases, if there is a single light source in the shooting environment of the above first image frame, the light source color information in the above RGB color space can be characterized by the RGB values of any pixel point in the area where the target object is located in the first image frame. And if there are multiple light sources in the shooting environment of the first image frame, the light source color information in the RGB color space can be characterized by an RGB image with the same resolution as the first image frame. Among them, each pixel point in the area where the target object is located in this RGB image corresponds to an RGB value, and the RGB values of all pixel points except the area where the target object is located in this RGB image are empty.
[0137] In one implementation manner, the mobile phone can perform color restoration on the first image frame according to the target color restoration network and the above light source color information. Among them, this target color restoration network is used to perform color restoration on the first image frame, so that the color-restored first image frame is more in line with the real scene, improving the accuracy of image color restoration. Specifically, the mobile phone can input the above light source color information and the first image frame into the target color restoration network to obtain a second image frame.
[0138] In an example, the above target color restoration network can be an encoder-decoder. This encoder-decoder can include an encoder and a decoder. In another example, the above target color restoration network can also be a deep neural network, etc., which is not specifically limited.
[0139] In some embodiments, such as Figure 10As shown above, the training process of the above-mentioned target color restoration network may specifically include: The mobile phone can obtain a second sample data set. Among them, the second sample data set includes multiple groups of second sample data, and each group of second sample data may include an image to be restored, light source color information, and a real image. Both the image to be restored and the real image include the target object, and the image to be restored is an image frame without color restoration, and the real image is an image frame not affected by a colored light source. After that, for each group of second sample data in the second sample data set, the mobile phone can input the image to be restored and the light source color information into the pre-constructed color restoration network to obtain a predicted restored image. After that, the mobile phone can adjust the parameters of the color restoration network according to the predicted restored image and the real image to obtain a trained color restoration network (or called the target color restoration network).
[0140] Specifically, as Figure 10 shown, the mobile phone can determine the loss value of the color restoration network according to the predicted restored image and the real image. After that, the mobile phone can adjust the parameters of the color restoration network according to the loss value of the color restoration network until the color restoration network meets the second training requirement to obtain a trained color restoration network (or called the target color restoration network). Among them, the second training requirement may include that the loss value of the color restoration network is less than the second preset loss value, or the predicted restored image is the same as the real image, or the second training times reach the second preset times. The second training times are the times when the mobile phone trains the color restoration network.
[0141] In some cases, the real image in the above-mentioned second sample data can be obtained by performing color restoration on the image to be restored in the second sample data in a preset processing method. Among them, the preset processing method is a retouching method. The image to be restored can be collected by a camera device. The light source color information in the second sample data can be collected by a spectrometer when the camera device collects the image to be restored. In this way, the real restored image can be made more in line with the real shooting scene, that is, the target color restoration network trained by the real image can restore a scene that conforms to the real shooting scene, improving the user's visual experience.
[0142] In some other cases, the real image in the above-mentioned second sample data may be an image frame captured by the imaging device under the target light. Among them, the target light is D65 light. The spectrum of the D65 light is approximately the same as that of natural sunlight. That is to say, the D65 light can simulate natural light, which can reduce the color deviation caused by the illumination difference, so that the real image can show the real color. Then, the mobile phone can perform degradation processing on the real image to obtain the image to be restored in the above-mentioned second sample data. Specifically, the mobile phone can render the real image according to the preset color information to obtain the image to be restored. It can be understood that the preset color information is the light source color information in the second sample data. In this way, the acquisition efficiency of the second sample data can be improved, and further the training efficiency of the color restoration network can be improved.
[0143] In one implementation, the above-mentioned second sample data set and the first sample data set can be the same sample data set. That is to say, the sample data in the sample data set can include the real light source color information, the image to be restored, the real restored image, and the multispectral data corresponding to the area of the human skin in the image to be restored at the same time. Then, when training the above-mentioned target light source prediction network, the mobile phone can obtain the real light source color information in the sample data and the multispectral data corresponding to the area of the human skin in the image to be restored. That is, the sample image frame in the first sample data can be equivalent to the image to be restored in the second sample data. When training the above-mentioned target color restoration network, the mobile phone can obtain the real light source color information, the image to be restored, and the real image in the sample data. That is, the real light source color information in the first sample data can be equivalent to the light source color information in the second sample data.
[0144] It should be noted that the training process of the above-mentioned target color restoration network can also be executed by the cloud service. That is to say, after the cloud service completes the training of the color restoration network, that is, obtains the target color restoration network, it can send the target color restoration network to the mobile phone. Then, the mobile phone can perform color restoration on the first image frame according to the target color restoration network to obtain the second image frame.
[0145] The above describes the process of how the mobile phone performs color restoration on the first image frame. Next, a possible implementation process of the mobile phone performing color restoration on the first image frame will be specifically introduced. As Figure 11 shown, the color restoration process may specifically include: The mobile phone captures the first image frame and multispectral data. Among them, the first image frame includes the target object (i.e., the user's skin). Then, the mobile phone can perform region segmentation on the first image frame to obtain the region image where the target object is located (i.e., the region image of the user's skin S). Then, the mobile phone can obtain the multispectral data corresponding to the region image of the target object from the multispectral data of the first image frame according to the region image of the target object.
[0146] After that, the mobile phone can input the multi-spectral data corresponding to the regional image of the above target object into the target light source prediction network to obtain the XYZ tristimulus values. After that, the mobile phone can extract coordinates from the XYZ tristimulus values to obtain the x-y color coordinates. After that, the mobile phone can perform spatial transformation on the x-y color coordinates to obtain the light source color information. Among them, the light source color information is used to represent the color information in the RGB color space. After that, the mobile phone can input the light source color information and the above first image frame into the target color restoration network to obtain the second image frame. Among them, the second image frame is the first image frame after color restoration.
[0147] In one implementation, since the above multi-spectral data is composed based on the reflectivity of the target object and the light source color information of the shooting environment, when the reflectivity of the target object is in a stable state, the mobile phone can predict the light source color information of the shooting environment. And when the light source color information of the shooting environment is in a stable state, the mobile phone can also predict the reflectivity of the target object. After that, the mobile phone can determine the type of the target object according to the reflectivity of the target object and the preset reflection interval. For example, if the reflectivity of the target object is within the preset reflection interval of the user, the mobile phone can determine that the target object is human skin.
[0148] In some embodiments, the mobile phone can collect a third image frame through an image sensor and multi-spectral data through a multi-spectral sensor under D65 light. Among them, the third image frame includes human skin. After that, the mobile phone can perform regional segmentation on the third image frame to obtain the multi-spectral data corresponding to the region where the human skin is located. After that, the mobile phone can perform skin texture detection on the user according to the multi-spectral data corresponding to the region where the human skin is located to obtain the skin state of the user. For example, the mobile phone can input the multi-spectral data corresponding to the region where the human skin is located into the skin texture detection network to obtain the skin state of the user.
[0149] In some embodiments, the present application provides a computer storage medium, including computer instructions, which when running on an electronic device, enable the electronic device to execute the method for adjusting usage parameters as described above.
[0150] In some embodiments, the present application provides a computer program product, which when running on an electronic device, enables the electronic device to execute the method for adjusting usage parameters as described above.
[0151] From the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0152] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0153] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it can be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0154] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0155] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: USB flash drive, mobile hard disk, read only memory (ROM), random access memory (RAM), magnetic disk or optical disc and other various media that can store program codes.
[0156] The above content is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application shall be covered by the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
Claims
1. An image processing method, characterized in that, Applied to an electronic device, the method includes: The electronic device acquires a first image frame and multispectral data corresponding to the region where a target object is located in the first image frame; wherein, the reflectivity of the target object is within a preset reflection range; The electronic device predicts the light source of the shooting environment of the first image frame based on the multispectral data corresponding to the region where the target object is located in the first image frame and a target light source prediction network, and obtains the light source color information corresponding to the first image frame; wherein, the target light source prediction network is used to predict the relationship between the light source color information of an image frame and the multispectral data corresponding to the region where the target object is located in the image frame, and the light source color information corresponding to the first image frame is represented by x-y color coordinates; The electronic device performs color restoration on the first image frame according to the light source color information corresponding to the first image frame, and obtains a second image frame.
2. The method according to claim 1, wherein The electronic device predicts the light source of the shooting environment of the first image frame based on the multispectral data corresponding to the region where the target object is located in the first image frame and a target light source prediction network, and obtains the light source color information corresponding to the first image frame, including: The electronic device inputs the multispectral data corresponding to the region where the target object is located in the first image frame into the target light source prediction network, and obtains the XYZ tristimulus values in a first color space; The electronic device extracts coordinates from the XYZ tristimulus values, and obtains x-y color coordinates; wherein, the x-y color coordinates are the light source color information in the first color space.
3. The method according to claim 2, characterized in that, The electronic device performs color restoration on the first image frame according to the light source color information corresponding to the first image frame, and obtains a second image frame, including: The electronic device performs a space conversion on the light source color information in the first color space, and obtains the light source color information in a second color space; wherein, the color mode corresponding to the second color space is the same as the color mode corresponding to the first image frame; The electronic device performs color restoration on the first image frame according to the light source color information in the second color space, and obtains the second image frame.
4. The method according to claim 3, characterized in that, The first image frame is an RGB image, the first color space is the CIE color space, and the second color space is the RGB color space.
5. The method according to any one of claims 1-4, characterized in that, The electronic device acquires a first image frame and multispectral data corresponding to the region where a target object is located in the first image frame, including: The electronic device acquires the first image frame and the multispectral data of the first image frame; wherein, each pixel point in the first image frame corresponds to a multispectral data; The electronic device performs region segmentation on the first image frame, and obtains the region image where the target object is located in the first image frame; The electronic device acquires the multispectral data of the pixel points corresponding to the region image where the target object is located from the multispectral data of the first image frame.
6. The method according to claim 5, wherein The electronic device predicts the light source of the shooting environment of the first image frame based on the multispectral data corresponding to the area where the target object is located in the first image frame and the target light source prediction network, and obtains the light source color information corresponding to the first image frame, including: When the multispectral data of the pixel points in the area where the target object is located in the first image frame meets the preset conditions, the electronic device predicts the light source of the shooting environment of the first image frame based on the multispectral data of any pixel point in the area where the target object is located in the first image frame and the target light source prediction network, and obtains the light source color information corresponding to the first image frame; Wherein, the preset conditions include: the difference between the multispectral data of any two pixel points in the area where the target object is located is less than the preset difference, and / or, the ratio between the number of pixel points with the same multispectral data and the total number of pixel points in the area where the target object is located is greater than the preset ratio.
7. The method according to claim 5, characterized in that, The electronic device predicts the light source of the shooting environment of the first image frame based on the multispectral data corresponding to the area where the target object is located in the first image frame and the target light source prediction network, and obtains the light source color information corresponding to the first image frame, including: When the multispectral data of the pixel points in the area where the target object is located in the first image frame does not meet the preset conditions, the electronic device predicts the light source of the shooting environment of the target object in the first image frame based on the multispectral data of each pixel point in the area where the target object is located and the target light source prediction network, and obtains the light source color information corresponding to the target object in the first image frame; The electronic device performs color restoration on the first image frame according to the light source color information corresponding to the first image frame, and obtains a second image frame, including: The electronic device performs color restoration on the area where the target object is located in the first image frame according to the light source color information corresponding to the target object in the first image frame, and obtains the second image frame.
8. The method according to claim 5, wherein The first image frame is a portrait image frame, and the target object of the portrait image frame includes human skin. The electronic device performs region segmentation on the first image frame to obtain the region image where the target object is located in the first image frame, including: The electronic device performs region segmentation on the portrait image frame to obtain a skin mask map corresponding to the area where the human skin is located; The electronic device obtains the multispectral data of the pixel points corresponding to the region image of the target object from the multispectral data of the first image frame, including: The electronic device obtains the multispectral data of the pixel points corresponding to the skin mask map from the multispectral data of the portrait image frame.
9. The method according to claim 1, characterized in that, The target light source prediction network is trained through the following process: Obtain a first sample data set; wherein, the first sample data set includes multiple groups of first sample data, and the first sample data includes the multispectral data of the area where the target object is located in the sample image frame and the true light source color information, and the acquisition time of the multispectral data is the same as the acquisition time of the true light source color information; For each group of the first sample data, input the multispectral data of the region where the target object is located in the sample image frame into a pre-constructed light source prediction network to obtain predicted light source color information; According to the predicted light source color information and the true light source color information, adjust the parameters of the pre-constructed light source prediction network to obtain the target light source prediction network.
10. The method according to claim 9, characterized in that, Before obtaining the first sample data set, the method further includes: Obtain multiple sample image frames and the multispectral data of each sample image frame from a photographing device, and obtain the true light source color information from a spectrometer; wherein, the photographing device is a device including a multispectral sensor for collecting the multispectral data of the sample image frame, the spectrometer is placed opposite to the photographing device, and the distance between the spectrometer and the target object is less than a preset distance; For each sample image frame, align the sample image frame with the multispectral data of the sample image frame to obtain the multispectral data corresponding to each pixel point in the sample image frame; Obtain the multispectral data corresponding to the region where the target object is located in the sample image frame from the multispectral data corresponding to each pixel point in the sample image frame.
11. The method according to claim 10, wherein The color attributes of the shooting environments corresponding to the multiple sample image frames are different, and the color attributes of the shooting environment include at least one of brightness, hue, and saturation; and / or, The multiple sample image frames include multiple image frames obtained by the photographing device when there is one light source in the shooting environment, including the target object at one position or multiple positions; and / or, The multiple sample image frames include image frames obtained by the photographing device at different shooting positions when there are multiple light sources in the shooting environment, including the target object at the same position.
12. The method according to claim 1, characterized in that, The electronic device performs color restoration on the first image frame according to the light source color information corresponding to the first image frame to obtain a second image frame, including: The electronic device inputs the light source color information corresponding to the first image frame and the first image frame into a target color restoration network to obtain the second image frame; wherein, the target color restoration network is obtained by adjusting parameters based on a predicted restored image and a true image during iterative training, the predicted restored image is obtained by performing color restoration on the image to be restored according to the image to be restored and the light source color information, and the true image is an image not affected by a colored light source.
13. The method according to claim 12, wherein The true image is obtained by performing color restoration on the image to be restored through a preset processing method, the image to be restored is collected by a camera device, and the light source color information is collected by a spectrometer when the camera device collects the image to be restored; and / or, The true image is collected by the camera device under target light, the target light is used to simulate natural light, the image to be restored is rendered according to preset color information based on the true image, and the light source color information is the preset color information.
14. An electronic device, characterized in that, The electronic device includes a camera, a multispectral sensor, a memory, and one or more processors; the camera, the multispectral sensor, the memory, and the processor are coupled; the camera is configured to capture a first image frame, the multispectral sensor is configured to capture multispectral data, the memory is configured to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device is caused to execute the method according to any one of claims 1 to 13.
15. A computer-readable storage medium, characterized in that, including computer instructions that, when run on an electronic device, cause the electronic device to execute the method according to any one of claims 1 to 13.
16. A computer program product, characterized in that, including computer instructions that, when run on an electronic device, cause the electronic device to execute the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Image processing system and method, computer readable medium, and electronic device
CN115314617A
Image color processing method and device
CN116684743A
Image processing method and electronic equipment
CN117745620A
Image processing method, model training method and related equipment
CN118509718A