Image processing method and device, electronic equipment, storage medium and program product

By acquiring image sample pairs under dark and bright conditions, using preset image processing models for feature extraction and model convergence, the problem of low accuracy of three-dimensional feature information under dark conditions is solved, and higher feature information accuracy and depth information simulation are achieved.

CN120020897APending Publication Date: 2025-05-20ORIENTAL POWER HOLDINGS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311540021.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

When the prior art extracts three-dimensional feature information under dark light conditions, it is difficult for the image processing model to accurately capture detailed information, resulting in low accuracy of the three-dimensional feature information.

Method used

通过获取暗光和明亮条件下的图像样本对,采用预设图像处理模型进行多粒度特征提取,提取频域特征,并根据目标损失对模型进行收敛,得到图像处理模型以识别三维特征信息。

Benefits of technology

The accuracy of the image processing model extracting three-dimensional feature information under dark light conditions is improved, and the depth information of the target object under bright light conditions can be more accurately simulated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020897A_ABST
    Figure CN120020897A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device, electronic equipment, a storage medium and a program product. The image processing method and device can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, Internet of Vehicles, games and virtual reality. According to the invention, an image sample pair can be obtained, wherein the image sample pair comprises a dark image sample and a bright image sample; performing multi-granularity feature extraction on the image sample pair by adopting a preset image processing model to obtain an image feature set corresponding to feature granularity; dark light frequency domain features are extracted from the dark light image features, and bright frequency domain features are extracted from the bright light image features; determining a target loss of the image sample pair based on the dark light frequency domain feature and the bright frequency domain feature, and converging a preset image processing model according to the target loss to obtain an image processing model; and adopting the image processing model to identify three-dimensional feature information in the dark light image containing the target object. According to the invention, the accuracy of the three-dimensional feature information extracted by the image processing model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to an image processing method, apparatus, electronic device, storage medium, and program product. Among them, the storage medium may refer to a computer-readable storage medium, and the program product may refer to a computer program product. Background Art

[0002] In the process of creating a virtual three-dimensional scene, the actual image of the real scene is acquired, and a three-dimensional feature information is extracted from the actual image by using an image processing model. Then, the virtual three-dimensional scene is created by using the three-dimensional feature information of the actual image.

[0003] Among them, in the virtual three-dimensional scene, there is a situation where it is necessary to create a virtual scene under low-light conditions such as at night. For a virtual scene under low-light conditions, it is generally created from the three-dimensional feature information extracted from a low-light image such as a night image. However, the details in the low-light image are not clear enough, which makes it difficult for the image processing model to accurately capture the detail information in the low-light image, and further makes it difficult for the image processing model to accurately extract the three-dimensional feature information in the low-light image.

[0004] In summary, there is currently a problem that the accuracy of the three-dimensional feature information extracted by the image processing model is relatively low. Summary of the Invention

[0005] Embodiments of this application provide an image processing method, apparatus, electronic device, storage medium, and program product, which can improve the accuracy of the three-dimensional feature information extracted by the image processing model.

[0006] An image processing method includes:

[0007] Obtaining at least one pair of image samples, where the pair of image samples includes a low-light image sample taken under low-light conditions and a bright image sample taken under bright illumination conditions;

[0008] Performing multi-granularity feature extraction on the pair of image samples by using a preset image processing model to obtain an image feature set corresponding to each feature granularity, where the image feature set includes a low-light image feature of the low-light image sample and a bright image feature of the bright image sample;

[0009] Extracting features of at least one frequency domain from the low-light image features to obtain low-light frequency domain features, and extracting features corresponding to the frequency domain from the bright image features to obtain bright frequency domain features;

[0010] Determining a target loss of the pair of image samples based on the low-light frequency domain features and the bright frequency domain features, and converging the preset image processing model according to the target loss to obtain an image processing model;

[0011] The three-dimensional feature information is recognized in a low-light image containing a target object by using an image processing model, and the three-dimensional feature information includes the depth information simulating the target object under bright illumination conditions.

[0012] Correspondingly, an embodiment of the present application provides an image processing apparatus, including:

[0013] An acquisition unit, which can be used to acquire at least one pair of image samples, and the pair of image samples includes a low-light image sample taken under low-light conditions and a bright image sample taken under bright illumination conditions;

[0014] A first extraction unit, which can be used to perform multi-granularity feature extraction on the pair of image samples by using a preset image processing model to obtain an image feature set corresponding to each feature granularity, and the image feature set includes the low-light image features of the low-light image sample and the bright image features of the bright image sample;

[0015] A second extraction unit, which can be used to extract the features of at least one frequency domain from the low-light image features to obtain the low-light frequency domain features, and extract the features corresponding to the frequency domain from the bright image features to obtain the bright frequency domain features;

[0016] A determination unit, which can be used to determine the target loss of the pair of image samples based on the low-light frequency domain features and the bright frequency domain features, and converge the preset image processing model according to the target loss to obtain an image processing model;

[0017] An identification unit, which can be used to recognize the three-dimensional feature information in a low-light image containing a target object by using the image processing model, and the three-dimensional feature information includes the depth information simulating the target object under bright illumination conditions.

[0018] Optionally, in some embodiments, the determination unit can specifically be used to recognize the target frequency domain features corresponding to each frequency domain in the low-light frequency domain features, and screen out the candidate frequency domain features corresponding to each frequency domain in the bright frequency domain features; calculate the loss between the target frequency domain features and the candidate frequency domain features to obtain the initial loss corresponding to the pair of image samples under each target frequency domain; fuse each initial loss to obtain the target loss of the pair of image samples.

[0019] Optionally, in some embodiments, the target frequency domain includes a high-frequency domain and a low-frequency domain; the determining unit may specifically be configured to extract, from the target frequency domain features, target high-frequency features corresponding to the low-light image samples in the high-frequency domain and target low-frequency features corresponding to the low-light image samples in the low-frequency domain; screen out, from the candidate frequency domain features, candidate high-frequency features corresponding to the bright image samples in the high-frequency domain and candidate low-frequency features corresponding to the bright image samples in the low-frequency domain; calculate an initial high-frequency loss of the image sample pair in the high-frequency domain based on the target high-frequency features and the candidate high-frequency features, and calculate an initial low-frequency loss of the target image sample pair in the low-frequency domain based on the target low-frequency features and the candidate low-frequency features; determine the initial loss corresponding to the image sample pair in each target frequency domain based on the initial high-frequency loss and the initial low-frequency loss.

[0020] Optionally, in some embodiments, the target high-frequency features include a plurality of sub-target high-frequency features, and the candidate high-frequency features include a plurality of sub-candidate high-frequency features; the determining unit may specifically be configured to fuse the sub-target high-frequency features to obtain a target fused high-frequency feature corresponding to the low-light image sample in the high-frequency domain; fuse the sub-candidate high-frequency features to obtain a candidate fused high-frequency feature corresponding to the bright image sample in the high-frequency domain; calculate an initial high-frequency loss of the image sample pair in the high-frequency domain based on the target fused high-frequency feature and the candidate fused high-frequency feature.

[0021] Optionally, in some embodiments, the determining unit may specifically be configured to obtain a weight corresponding to the high-frequency domain and a candidate weight corresponding to the low-frequency domain, where the weight is greater than the candidate weight; perform weighting on the initial high-frequency loss and the initial low-frequency loss based on the weight and the candidate weight to obtain the target loss of the image sample pair.

[0022] Optionally, in some embodiments, the first extraction unit may specifically be configured to, when the feature granularity includes an encoded feature granularity, encode the low-light image sample by using a low-light image processing sub-model of a preset image processing model to obtain a low-light encoded feature corresponding to the low-light image sample; encode the bright image sample by using a bright image processing sub-model of the preset image processing model to obtain a bright encoded feature corresponding to the bright image sample; generate an encoded feature set corresponding to the encoded feature granularity based on the low-light encoded feature and the bright encoded feature, and generate an image feature set corresponding to each feature granularity based on the encoded feature set.

[0023] Optionally, in some embodiments, the first extraction unit may specifically be configured to, when the feature granularity further includes a decoding feature granularity, screen out target low-light encoded features from the low-light encoded features, and use the low-light image processing sub-model to decode the target low-light encoded features to obtain low-light decoded features; extract target bright encoded features from the bright encoded features, and use the bright image processing sub-model to decode the target bright encoded features to obtain bright decoded features; generate a decoded feature set corresponding to the decoding feature granularity based on the low-light decoded features and the bright decoded features, and generate an image feature set corresponding to each feature granularity based on the encoded feature set and the decoded feature set.

[0024] Optionally, in some embodiments, the first extraction unit may specifically be configured to, when the feature granularity further includes a convolutional feature granularity, obtain target low-light decoded features from the low-light decoded features, and use the low-light image processing sub-model to perform convolution on the target low-light decoded features to obtain low-light convolutional features; extract target bright decoded features from the bright decoded features, and use the bright image processing sub-model to perform convolution on the target bright decoded features to obtain bright convolutional features; generate a convolutional feature set corresponding to the convolutional feature granularity based on the low-light convolutional features and the bright convolutional features, and generate an image feature set corresponding to each feature granularity according to the convolutional feature set, the encoded feature set, and the decoded feature set.

[0025] Optionally, in some embodiments, the encoded feature granularity includes multiple sub-encoded feature granularities; the first extraction unit may specifically be configured to use the low-light image processing sub-model to encode the low-light image sample to obtain initial low-light encoded features corresponding to each encoding level of the low-light image processing sub-model; screen out candidate low-light encoded features corresponding to each sub-encoded feature granularity from the initial low-light encoded features; generate low-light encoded features corresponding to the low-light image sample based on the candidate low-light encoded features.

[0026] Optionally, in some embodiments, the first extraction unit may specifically be configured to determine a target encoding level corresponding to each sub-encoded feature granularity in the encoding level; screen out reference low-light encoded features corresponding to each target encoding level from the initial low-light encoded features; use the reference low-light encoded features as candidate low-light encoded features corresponding to the sub-encoded feature granularity.

[0027] Optionally, in some embodiments, the second extraction unit may specifically be configured to perform frequency domain decomposition on the low-light image features to obtain target high-frequency features corresponding to the high-frequency domain of the low-light image sample and target low-frequency features corresponding to the low-frequency domain; generate low-light frequency domain features based on the target high-frequency features and the target low-frequency features.

[0028] Optionally, in some embodiments, the determining unit may specifically be further configured to obtain an initial bright image sample, where the initial bright image sample includes candidate objects under bright illumination conditions; use an initial image processing model to predict the initial bright image sample to obtain initial three-dimensional feature information corresponding to the initial bright image sample, where the initial three-dimensional feature information includes the depth information of the candidate objects under bright illumination conditions; and converge the initial image processing model based on the initial three-dimensional feature information to obtain a preset image processing model.

[0029] Optionally, in some embodiments, the determining unit may specifically be configured to encode an initial image sample using an initial bright image processing sub-model of the initial image processing model to obtain initial encoded features; decode the initial encoded features using the initial bright image processing sub-model to obtain initial three-dimensional feature information corresponding to the initial image sample; converge the initial bright image processing sub-model based on the initial three-dimensional feature information to obtain a bright image processing sub-model, and generate a preset image processing model based on the bright image processing sub-model.

[0030] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory, where the memory stores an application program, and the processor is configured to run the application program in the memory to implement the image processing method provided by the embodiment of the present invention.

[0031] In addition, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute any one of the image processing methods provided by the embodiment of the present application.

[0032] In addition, an embodiment of the present application further provides a computer program product, including a computer program, where the computer program, when executed by a processor, implements any one of the image processing methods provided by the embodiment of the present application.

[0033] The present application can obtain at least one pair of image samples. The pair of image samples includes a low-light image sample taken under low-light conditions and a bright image sample taken under bright illumination conditions. A preset image processing model is used to perform multi-granularity feature extraction on the pair of image samples to obtain an image feature set corresponding to each feature granularity. The image feature set includes the low-light image features of the low-light image sample and the bright image features of the bright image sample. At least one frequency-domain feature is extracted from the low-light image features to obtain low-light frequency-domain features, and features corresponding to the frequency domain are extracted from the bright image features to obtain bright frequency-domain features. Based on the low-light frequency-domain features and the bright frequency-domain features, the target loss of the pair of image samples is determined, and the preset image processing model is converged according to the target loss to obtain an image processing model. The image processing model is used to identify three-dimensional feature information in the low-light image containing the target object. The three-dimensional feature information includes depth information simulating the target object under bright illumination conditions. Since the present application can use a preset image processing model to perform multi-granularity feature extraction on a pair of image samples to obtain the low-light image features of the low-light image sample and the bright image features of the bright image sample at each feature granularity, in this way, the present application can extract low-light frequency-domain features from the low-light image features and extract bright frequency-domain features from the bright image features. Based on this, the present application can use the target loss obtained based on the low-light frequency-domain features and the bright frequency-domain features to converge the preset image processing model, thereby improving the prediction accuracy of the converged image processing model. Based on this, the present application can use the converged image processing model to predict the low-light image to obtain depth information simulating the target object under bright illumination conditions, thus improving the accuracy of the three-dimensional feature information extracted by the image processing model. Description of the Drawings

[0034] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0035] Figure 1 It is a schematic diagram of the scenario of the image processing method provided by the embodiment of the present application;

[0036] Figure 2 It is a first schematic flowchart of the image processing method provided by the embodiment of the present application;

[0037] Figure 3 It is a schematic structural diagram of the preset image processing model provided by the embodiment of the present application;

[0038] Figure 4 It is a schematic diagram of obtaining the initial high-frequency loss provided by the embodiment of the present application;

[0039] Figure 5 is the effect diagram of the three-dimensional feature information provided by the embodiment of the present application;

[0040] Figure 6 is the second flowchart of the image processing method provided by the embodiment of the present application;

[0041] Figure 7 is the structural schematic diagram of the image processing device provided by the embodiment of the present application;

[0042] Figure 8 is the structural schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0043] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0044] The embodiment of the present application provides an image processing method, device, electronic device, computer-readable storage medium, and computer program product. Among them, the image processing device can be integrated in the electronic device, and the electronic device can be a server or a terminal device, etc.

[0045] Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, a virtual reality device, a game terminal, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.

[0046] For example, refer to Figure 1, taking the example that an image processing device is integrated in an electronic device, the electronic device can obtain at least one pair of image samples, where the pair of image samples includes a low-light image sample taken under low-light conditions and a bright image sample taken under bright light conditions; use a preset image processing model to perform multi-granularity feature extraction on the pair of image samples to obtain an image feature set corresponding to each feature granularity, and the image feature set includes the low-light image features of the low-light image sample and the bright image features of the bright image sample; extract at least one frequency-domain feature from the low-light image features to obtain low-light frequency-domain features, and extract the features corresponding to the frequency domain from the bright image features to obtain bright frequency-domain features; based on the low-light frequency-domain features and the bright frequency-domain features, determine the target loss of the pair of image samples, and converge the preset image processing model according to the target loss to obtain an image processing model; use the image processing model to identify three-dimensional feature information in a low-light image containing a target object, and the three-dimensional feature information includes depth information simulating the target object under bright light conditions.

[0047] Among them, the present application provides an image processing method that can process a low-light image sample to obtain the low-light frequency-domain features of the low-light image sample, and can process a bright image sample to obtain the bright frequency-domain features corresponding to the bright image sample. In this way, the preset image processing model can be converged using the target loss obtained based on the low-light frequency-domain features and the bright frequency-domain features to obtain an image processing model, thereby improving the prediction accuracy of the image processing model. Based on this, the present application can use the image processing model to identify the depth information simulating the target object under bright light conditions in a low-light image.

[0048] Among them, it can be understood that in the specific implementation of the present application, it involves relevant data such as low-light image samples, bright image samples, and low-light images. When the following embodiments of the present application are applied to specific products or technologies, permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0049] Among them, the present application relates to Artificial Intelligence (AI). Artificial Intelligence is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, Artificial Intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial Intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machine to have the functions of perception, reasoning, and decision-making.

[0050] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0051] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0052] This embodiment will be described from the perspective of an image processing device, which can be specifically integrated in an electronic device. The electronic device can be a server or a terminal device, etc.; among them, the terminal can include devices such as a tablet computer, a laptop computer, a personal computer (PC, Personal Computer), a wearable device, a virtual reality device, or other intelligent devices that can obtain data.

[0053] This application provides an image processing method, including:

[0054] Obtain at least one pair of image samples, where the pair of image samples includes a low-light image sample taken under low-light conditions and a bright image sample taken under bright light conditions; use a preset image processing model to perform multi-granularity feature extraction on the pair of image samples to obtain an image feature set corresponding to each feature granularity, and the image feature set includes the low-light image features of the low-light image sample and the bright image features of the bright image sample; extract at least one frequency-domain feature from the low-light image features to obtain low-light frequency-domain features, and extract the features corresponding to the frequency domain from the bright image features to obtain bright frequency-domain features; based on the low-light frequency-domain features and the bright frequency-domain features, determine the target loss of the pair of image samples, and converge the preset image processing model according to the target loss to obtain an image processing model; use the image processing model to identify three-dimensional feature information in a low-light image containing a target object, and the three-dimensional feature information includes depth information simulating the target object under bright light conditions.

[0055] Specifically, as Figure 2 shown, the specific process of this image processing method is as shown in steps S201 to S205:

[0056] It should be noted here that this application can be applied to many fields such as games, rendering, augmented reality (AR), virtual reality technology (VR), and AI generated content (AIGC). It can assist in generating three-dimensional data under low-light conditions based on real scenes, including scenes, people, objects, etc., improving the production efficiency of three-dimensional data, and effectively supporting the research and development and advancement of the above fields. For example, this application can be applied to the game business. The three-dimensional data, i.e., three-dimensional feature information, generated by this application can quickly realize three-dimensional modeling in the game business and improve the development efficiency of the game.

[0057] S201. Obtain at least one pair of image samples, where the pair of image samples includes a low-light image sample taken under low-light conditions and a bright image sample taken under bright light conditions.

[0058] Among them, the low-light conditions may refer to conditions where the intensity of visible light is low and insufficient to provide enough light for the camera or the human eye. For example, the low-light conditions may be the conditions at night.

[0059] Among them, the bright light conditions may refer to conditions where the intensity of visible light is high and sufficient to provide enough light for the camera or the human eye. For example, the bright light conditions may refer to the conditions during the day.

[0060] Among them, the low-light image sample may include candidate objects, and the candidate objects may include, but are not limited to, at least one of people, animals, objects, and plants. The bright image sample may also include candidate objects.

[0061] It should be noted here that the candidate objects included in the low-light image sample and the candidate objects included in the bright image sample may be different. For example, the candidate objects included in the low-light image sample may be animals and plants, and the candidate objects included in the bright image sample may be people and plants.

[0062] For step S201, the method of the step "obtain at least one pair of image samples" may be: the electronic device sends a sample acquisition request to the storage server, so that the storage server extracts the low-light image sample and the bright image sample from the storage space of the storage server based on the sample acquisition request and returns them to the electronic device; the electronic device receives the low-light image sample and the bright image sample and generates at least one pair of image samples based on the low-light image sample and the bright image sample.

[0063] Among them, the storage server and the electronic device may both be set in the same local area network. In this way, the electronic device can quickly obtain the low-light image sample and the bright image sample from the storage server through the local area network.

[0064] Specifically, the method for the step of "generating at least one pair of image samples based on the low-light image samples and the bright image samples" can be as follows: identifying the target image category of the low-light image samples in the low-light image samples; identifying the image category of the bright image samples in the bright image samples; based on the image category, respectively extracting the target low-light image samples and the target bright image samples belonging to the same target image category from the low-light image samples and the bright image samples; and generating a pair of image samples based on the target low-light image samples and the target bright image samples. Alternatively, candidate low-light image samples can be randomly selected from the low-light image samples, candidate bright image samples can be randomly extracted from the bright image samples, and a pair of image samples can be generated based on the candidate low-light image samples and the candidate bright image samples.

[0065] Regarding step S201, the method for the step of "obtaining at least one pair of image samples" can be as follows: obtaining the initial low-light image samples and the initial bright image samples; augmenting the initial low-light image samples to obtain low-light image samples; augmenting the initial bright image samples to obtain bright image samples; and constructing at least one pair of image samples based on the low-light image samples and the bright image samples.

[0066] Specifically, the method for the step of "constructing at least one pair of image samples based on the low-light image samples and the bright image samples" can refer to the method of "generating at least one pair of image samples based on the low-light image samples and the bright image samples" described above, and will not be elaborated here.

[0067] Among them, the augmentation methods for the initial low-light image samples and the initial bright image samples in this application can include but are not limited to at least one of mirror flipping, translation, rotation, scaling, shearing, contrast adjustment, image deformation, and color transformation, and can be specifically selected according to requirements.

[0068] S202. Using a preset image processing model to perform multi-granularity feature extraction on the pair of image samples to obtain an image feature set corresponding to each feature granularity.

[0069] Among them, the image feature set includes the low-light image features of the low-light image samples and the bright image features of the bright image samples.

[0070] Among them, the feature granularity can be the granularity level representing the low-light image features or the bright image features. For example, the feature granularity can include the encoded feature granularity, the decoded feature granularity, and the convolutional feature granularity.

[0071] The so-called encoding feature granularity may refer to the granularity level obtained by encoding the low-light image features or bright image features; the so-called decoding feature granularity may refer to the granularity level obtained by decoding the low-light image features or bright image features; the so-called convolutional feature granularity may refer to the granularity level obtained by convolving the low-light image features or bright image features.

[0072] After obtaining the image sample pair, this application can perform feature extraction on the image sample pair. Specifically, for step S202, the step of "using a preset image processing model to perform multi-granularity feature extraction on the image sample pair to obtain an image feature set corresponding to each feature granularity" can be as shown in steps S2021 to S2023:

[0073] S2021. When the feature granularity includes the encoding feature granularity, use the low-light image processing sub-model of the preset image processing model to encode the low-light image sample to obtain the low-light encoding feature corresponding to the low-light image sample.

[0074] Among them, the encoding feature granularity may include multiple sub-encoding feature granularities; correspondingly, the low-light encoding feature may include the initial low-light encoding features corresponding to each sub-encoding feature granularity.

[0075] It can be understood here that the information included in the initial low-light encoding features corresponding to different sub-encoding feature granularities is different.

[0076] For step S2021, the method of the step of "using the low-light image processing sub-model of the preset image processing model to encode the low-light image sample to obtain the low-light encoding feature corresponding to the low-light image sample" can be as shown in steps S21 to S23:

[0077] S21. Use the low-light image processing sub-model to encode the low-light image sample to obtain the initial low-light encoding features corresponding to each encoding layer of the low-light image processing sub-model.

[0078] Among them, for example, as Figure 3 shown, the low-light image processing sub-model includes a first encoder. Based on this, for step S21, this application can use the first encoder of the low-light image processing sub-model to encode the low-light image sample to obtain the initial low-light encoding features corresponding to each encoding layer of the low-light image processing sub-model.

[0079] Specifically, the first encoder may include multiple encoding layers, and this application can determine the encoding order of each encoding layer; based on the encoding order, determine the target encoding layer ranked first among the multiple encoding layers; use the target encoding layer to encode the low-light image sample to obtain the initial low-light encoding feature of the target encoding layer; use the initial low-light encoding feature of the target encoding layer as the input feature of the next encoding layer of the target encoding layer, and use the next encoding layer to extract features from the input feature to obtain the initial low-light encoding feature of the next encoding layer; and so on until each encoding layer is encoded, obtaining the initial low-light encoding features corresponding to each encoding level respectively.

[0080] Among them, the parameters and activation functions of each encoding layer of the first encoder can be different.

[0081] For example, the first encoder includes a first encoding layer, a second encoding layer, and a third encoding layer arranged in sequence. This application can use the first encoding layer to encode the low-light image sample to obtain the first initial low-light encoding feature; use the second encoding layer to encode the first initial low-light encoding feature to obtain the second initial low-light encoding feature; use the third encoding layer to encode the second initial low-light encoding feature to obtain the third initial low-light encoding feature; and use the first initial low-light encoding feature, the second initial low-light encoding feature, and the third initial low-light encoding feature as the initial low-light encoding features.

[0082] S22. Screen out the candidate low-light encoding features corresponding to each sub-encoding feature granularity from the initial low-light encoding features.

[0083] For step S22, the method of the step "screen out the candidate low-light encoding features corresponding to each sub-encoding feature granularity from the initial low-light encoding features" can be: in the encoding level, determine the target encoding level corresponding to each sub-encoding feature granularity; in the initial low-light encoding features, screen out the reference low-light encoding features corresponding to each target encoding level; and use the reference low-light encoding features as the candidate low-light encoding features corresponding to the sub-encoding feature granularity.

[0084] Among them, the sub-encoding feature granularity and the encoding level are in one-to-one correspondence. Based on this, this application can determine the target encoding level corresponding to the sub-encoding feature granularity among multiple encoding levels based on the preset correspondence between the sub-encoding feature granularity and the encoding level.

[0085] S23. Generate the low-light encoding feature corresponding to the low-light image sample based on the candidate low-light encoding features.

[0086] For step S23, among them, this application can use the candidate low-light encoding features corresponding to each sub-encoding feature granularity as the low-light encoding features.

[0087] S2022. Use the bright image processing sub-model of the preset image processing model to encode the bright image sample, and obtain the bright encoding feature corresponding to the bright image sample.

[0088] Regarding step S2022, the method of the step "Use the bright image processing sub-model of the preset image processing model to encode the bright image sample, and obtain the bright encoding feature corresponding to the bright image sample" can be: Use the bright image processing sub-model to encode the bright image sample, and obtain the initial bright encoding feature corresponding to each candidate encoding level of the bright image processing sub-model; In the initial bright encoding feature, screen out the candidate bright encoding features corresponding to each sub-encoding feature granularity; Based on the candidate bright encoding features, generate the bright encoding feature corresponding to the bright image sample.

[0089] Among them, as Figure 3 shown, the bright image processing sub-model includes a second encoder. Regarding the step "Use the bright image processing sub-model to encode the bright image sample, and obtain the initial bright encoding feature corresponding to each candidate encoding level of the bright image processing sub-model", this application can use the second encoder of the bright image processing sub-model to encode the bright image sample, and obtain the initial bright encoding features corresponding to each candidate encoding level of the bright image processing sub-model respectively.

[0090] Specifically, the second encoder can include multiple candidate encoding layers. This application can determine the encoding order of each candidate encoding layer; Based on the candidate encoding order, determine the initial encoding layer ranked first among the multiple candidate encoding layers; Use the initial encoding layer to encode the bright image sample, and obtain the initial bright encoding feature of the initial encoding layer; Use the initial bright encoding feature of the initial encoding layer as the candidate input feature of the next candidate encoding layer of the initial encoding layer, and use the next candidate encoding layer to extract features from the candidate input feature, and obtain the initial bright encoding feature of the next candidate encoding layer; And so on, until each candidate encoding layer is encoded, and obtain the initial bright encoding features corresponding to each candidate encoding level respectively.

[0091] Among them, the parameters and activation functions of each candidate encoding layer of the second encoder can be different.

[0092] Specifically, the method of the step "In the initial bright encoding feature, screen out the candidate bright encoding features corresponding to each sub-encoding feature granularity" can specifically refer to the method of the foregoing "In the initial dark light encoding feature, screen out the candidate dark light encoding features corresponding to each sub-encoding feature granularity", which will not be elaborated here.

[0093] Specifically, for the step of "generating the bright encoding features corresponding to the bright image samples based on the candidate bright encoding features", reference can be made to the aforementioned step of "generating the dark encoding features corresponding to the dark image samples based on the candidate dark encoding features", which will not be elaborated here.

[0094] S2023. Generate an encoding feature set corresponding to the encoding feature granularity based on the dark encoding features and the bright encoding features, and generate an image feature set corresponding to each feature granularity based on the encoding feature set.

[0095] Regarding step S2023, it can be understood that the encoding feature set may include the dark encoding features of the dark image samples and the bright encoding features of the bright image samples.

[0096] After obtaining the encoding feature set in this application, an image feature set can be generated. Specifically,

[0097] There are various ways for the step of "generating an image feature set corresponding to each feature granularity based on the encoding feature set". For example, this application can use the encoding feature set as the image feature set. In this case, this application can use the dark encoding features of the dark image samples as the dark image features and the bright encoding features of the bright image samples as the bright image features. Another example is that when the feature granularity includes the encoding feature granularity and the decoding feature granularity, generating the image feature set can be as shown in steps S31 to S33; another example is that when the feature granularity includes the encoding feature granularity, the decoding feature granularity, and the convolutional feature granularity, generating the image feature set can be as shown in steps S61 to S63.

[0098] Among them, steps S31 to S33 can be as follows:

[0099] S31. When the feature granularity further includes the decoding feature granularity, screen out the target dark encoding features from the dark encoding features, and use the dark image processing sub-model to decode the target dark encoding features to obtain the dark decoding features.

[0100] Regarding step S31, the way of the step of "screening out the target dark encoding features from the dark encoding features" can be: extracting the dark encoding features corresponding to the last encoding level from the dark encoding features; using the dark encoding features corresponding to the last encoding level as the target dark encoding features.

[0101] After obtaining the target low-light encoded feature, the present application can decode the target low-light encoded feature. Among them, the decoding feature granularity can include multiple sub-decoding feature granularities. Specifically, for step S31, the method of the step "using the low-light image processing sub-model to decode the target low-light encoded feature to obtain the low-light decoded feature" can be: using the low-light image processing sub-model to decode the target low-light encoded feature to obtain the initial low-light decoded feature corresponding to each decoding level of the low-light image processing sub-model; in the initial low-light decoded feature, screening out the candidate low-light decoded features corresponding to each sub-decoding feature granularity; and generating the low-light decoded feature corresponding to the low-light image sample based on the candidate low-light decoded features.

[0102] Specifically, the method of the step "using the low-light image processing sub-model to decode the target low-light encoded feature to obtain the initial low-light decoded feature corresponding to each decoding level of the low-light image processing sub-model" can be: as Figure 3 shown, using the first decoder included in the low-light image processing sub-model to decode the low-light image sample to obtain the initial low-light decoded feature corresponding to each decoding level of the low-light image processing sub-model.

[0103] Among them, the first decoder can include multiple decoding layers, and among them, the parameters or activation functions included in different decoding layers can be different.

[0104] The present application can determine the decoding order of each decoding layer; based on the decoding order, determine the target decoding layer ranked first among the multiple decoding layers; use the target decoding layer to decode the target low-light encoded feature to obtain the initial low-light decoded feature of the target decoding layer; use the initial low-light decoded feature of the target decoding layer as the input feature of the next decoding layer of the target decoding layer; and use the next decoding layer to perform feature extraction on the input feature to obtain the initial low-light decoded feature of the next decoding layer; and so on until each decoding layer is decoded, and the initial low-light decoded feature corresponding to each decoding level is obtained.

[0105] Specifically, the method of the step "in the initial low-light decoded feature, screening out the candidate low-light decoded features corresponding to each sub-decoding feature granularity" can be: in the decoding level, determine the target decoding level corresponding to each sub-decoding feature granularity; in the initial low-light decoded feature, screen out the reference low-light decoded features corresponding to each target decoding level, and use the reference low-light decoded features as the candidate low-light decoded features corresponding to the sub-decoding feature granularity.

[0106] Among them, the sub-decoding feature granularity and the decoding level are in one-to-one correspondence. Based on this, the present application can determine the target decoding level corresponding to each sub-decoding feature granularity among the multiple decoding levels based on the preset mapping relationship between the sub-decoding feature granularity and the decoding level.

[0107] Specifically, the present application may use the candidate low-light decoding features corresponding to each sub-decoding feature granularity as the low-light decoding features corresponding to the low-light image sample.

[0108] S32. Extract the target bright encoding features from the bright encoding features, and use the bright image processing sub-model to decode the target bright encoding features to obtain bright decoding features.

[0109] For step S32, the method of the step "Extract the target bright encoding features from the bright encoding features" may be: extract the bright encoding features corresponding to the last candidate encoding level from the bright encoding features; use the bright encoding features corresponding to the last candidate encoding level as the target bright encoding features.

[0110] For step S32, the method of the step "Use the bright image processing sub-model to decode the target bright encoding features to obtain bright decoding features" may be: use the bright image processing sub-model to decode the target bright encoding features to obtain the initial bright decoding features corresponding to each candidate decoding level of the bright image processing sub-model; among the initial bright decoding features, screen out the candidate bright decoding features corresponding to each sub-decoding feature granularity; based on the candidate bright decoding features, generate the bright decoding features corresponding to the bright image sample.

[0111] Specifically, the method of the step "Use the bright image processing sub-model to decode the target bright encoding features to obtain the initial bright decoding features corresponding to each candidate decoding level of the bright image processing sub-model" may be: as Figure 3 shown, use the second decoder included in the bright processing sub-model to decode the target bright encoding features to obtain the initial bright decoding features corresponding to each candidate decoding level of the bright image processing sub-model.

[0112] Among them, the second decoder may include multiple candidate decoding layers, and the parameters or activation functions included in different candidate decoding layers may also be different. Specifically, for the method of using the multiple candidate decoding layers of the second decoder to obtain the initial bright decoding features corresponding to each candidate decoding level, reference may be specifically made to the relevant explanations of the foregoing "Use the first decoder included in the low-light image processing sub-model to decode the low-light image sample to obtain the initial low-light decoding features corresponding to each decoding level of the low-light image processing sub-model", which will not be elaborated here.

[0113] Specifically, the method of the step "in the initial bright decoding features, screen out the candidate bright decoding features corresponding to each sub-decoding feature granularity" can refer to the description of the aforementioned "in the initial dark decoding features, screen out the candidate dark decoding features corresponding to each sub-decoding feature granularity", which will not be elaborated here.

[0114] Specifically, this application can use the candidate bright decoding features corresponding to each sub-decoding feature granularity as the bright decoding features corresponding to the bright image samples.

[0115] S33. Based on the dark decoding features and the bright decoding features, generate a decoding feature set corresponding to the decoding feature granularity, and based on the encoding feature set and the decoding feature set, generate an image feature set corresponding to each feature granularity.

[0116] Regarding step S33, it can be understood here that the decoding feature set of this application can include the dark decoding features corresponding to the dark image samples and the bright decoding features corresponding to the bright image samples.

[0117] It should be noted here that regarding step S33, it can be understood here that this application can use the dark decoding features and the dark encoding features corresponding to the dark image samples as the dark image features; and use the bright decoding features and the bright encoding features corresponding to the bright image samples as the bright image features.

[0118] Regarding step S33, the step of "based on the encoding feature set and the decoding feature set, generate an image feature set corresponding to each feature granularity" can specifically be as shown in steps S61 to S63:

[0119] S61. When the feature granularity further includes a convolutional feature granularity, obtain the target dark decoding feature in the dark decoding features, and use the dark image processing sub-model to perform convolution on the target dark decoding feature to obtain a dark convolutional feature.

[0120] Regarding step S61, the method of the step "obtain the target dark decoding feature in the dark decoding features" can be: extract the initial dark decoding feature corresponding to the last decoding level in the dark decoding features, and use the initial dark decoding feature corresponding to the last decoding level as the target dark decoding feature.

[0121] Regarding step S61, the method of the step "use the dark image processing sub-model to perform convolution on the target dark decoding feature to obtain a dark convolutional feature" can be: use the first convolutional layer of the dark image processing sub-model to perform convolution on the target dark decoding feature to obtain a dark convolutional feature.

[0122] Among them, it can be understood that the low-light convolution feature can be the feature finally output by the low-light image processing sub-model, that is, the low-light convolution feature represents the three-dimensional feature of the low-light image sample.

[0123] S62. Extract the target bright decoding feature from the bright decoding feature, and use the bright image processing sub-model to perform convolution on the target bright decoding feature to obtain the bright convolution feature.

[0124] Regarding step S62, similar to step S61, the present application can also use the initial bright decoding feature corresponding to the last candidate decoding layer as the target bright decoding feature; and use the second convolution layer of the bright image processing sub-model to perform convolution on the target bright decoding feature to obtain the bright convolution feature.

[0125] Among them, it can be understood that the bright convolution feature can be the feature finally output by the bright image processing sub-model, that is, the bright convolution feature represents the three-dimensional feature of the bright image sample.

[0126] S63. Generate a convolution feature set corresponding to the convolution feature granularity based on the low-light convolution feature and the bright convolution feature, and generate an image feature set corresponding to each feature granularity based on the convolution feature set, the encoding feature set, and the decoding feature set.

[0127] Regarding step S202, here, the case where the feature granularity includes the encoding feature granularity, the decoding feature granularity, and the convolution feature granularity is taken as an example for elaboration.

[0128] For example, as Figure 3 shown, the preset image processing model can include a low-light image processing sub-model and a bright image processing sub-model; the present application can use the low-light image processing sub-model to process the low-light image sample to obtain the low-light image feature; use the bright image processing sub-model to process the bright image sample to obtain the bright image feature.

[0129] (1) For the low-light image sample, the present application can use the first encoder of the low-light image processing sub-model to encode the low-light image sample to obtain the low-light encoding feature corresponding to the low-light image sample, and this low-light encoding feature is the feature of the low-light image sample at the encoding feature granularity.

[0130] Among them, the low-light encoding feature can include the candidate low-light encoding features corresponding to each sub-encoding feature granularity. The low-light encoding feature can be expressed as en = [en1,..., enn], where enn represents the nth candidate low-light encoding feature, and the n candidate low-light encoding features are sorted in descending order of size.

[0131] Then, the present application can screen out the target low-light encoding feature from the low-light encoding features, and use the first decoder of the low-light image processing sub-model to decode the target low-light encoding feature to obtain a low-light decoded feature, which is the feature of the low-light image sample at the decoded feature granularity.

[0132] Among them, the low-light decoded feature can include candidate low-light decoded features corresponding to each sub-decoded feature granularity. The low-light decoded feature can be expressed as dn = [dn1, …, dnn], where dnn represents the nth candidate low-light decoded feature; the n candidate low-light decoded features are sorted in descending order of size.

[0133] Next, the present application can obtain the target low-light decoded feature from the low-light decoded feature, and use the first convolutional layer of the low-light image processing sub-model to perform convolution on the target low-light decoded feature to obtain a low-light convolutional feature, which is the feature of the low-light image sample at the convolutional feature granularity. The low-light convolutional feature can be denoted as d_night.

[0134] The present application can use the low-light encoding feature, the low-light decoded feature, and the low-light convolutional feature as the low-light image features included in the image feature set.

[0135] (2) For a bright image sample, the present application can use the second encoder of the bright image processing sub-model to encode the bright image sample to obtain a bright encoding feature corresponding to the bright image sample, which is the feature of the bright image sample at the encoding feature granularity.

[0136] Among them, the bright encoding feature can include candidate bright encoding features corresponding to each sub-encoding feature granularity; the candidate bright encoding feature can be expressed as ed = [ed1, …, edn], where edn is the nth candidate bright encoding feature, and the n candidate bright encoding features are sorted in descending order of size.

[0137] Then, the present application can extract the target bright encoding feature from the bright encoding feature, and use the second decoder of the bright image processing sub-model to decode the target bright encoding feature to obtain a bright decoded feature, which is the feature of the bright image sample at the decoded feature granularity.

[0138] Among them, the bright decoded feature can include candidate bright decoded features corresponding to each sub-decoded feature granularity; the bright decoded feature can be expressed as dd = [dd1, …, ddn], where ddn can refer to the nth candidate bright decoded feature, and the n candidate bright decoded features are sorted in descending order of size.

[0139] Next, the present application can extract the target bright decoding feature from the bright decoding feature, and use the second convolutional layer of the bright image processing sub-model to perform convolution on the target bright decoding feature to obtain a bright convolution feature, which is the feature of the bright image sample at the convolutional feature granularity. The bright convolution feature can be denoted as d_day.

[0140] The present application can use the bright encoding feature, the bright decoding feature, and the bright convolution feature as the bright image features included in the image feature set.

[0141] S203: Extract the features in at least one frequency domain from the low-light image features to obtain low-light frequency domain features, and extract the features corresponding to the frequency domain from the bright image features to obtain bright frequency domain features.

[0142] After the present application obtains the low-light image features, it can extract features from the low-light image features. Specifically, for step S203, the method of the step "Extract the features in at least one frequency domain from the low-light image features to obtain low-light frequency domain features" can be: perform frequency domain decomposition on the low-light image features to obtain the target high-frequency feature corresponding to the high-frequency domain of the low-light image sample and the target low-frequency feature corresponding to the low-frequency domain; generate low-light frequency domain features based on the target high-frequency feature and the target low-frequency feature.

[0143] Specifically, for the step "Perform frequency domain decomposition on the low-light image features to obtain the target high-frequency feature corresponding to the high-frequency domain of the low-light image sample and the target low-frequency feature corresponding to the low-frequency domain", the method can be: use a transformation function to perform frequency domain decomposition on the low-light image features to obtain the target high-frequency feature corresponding to the high-frequency domain of the low-light image sample and the target low-frequency feature corresponding to the low-frequency domain.

[0144] Among them, the transformation function can include a discrete wavelet transform function; among them, the discrete wavelet transform function can include at least one of Haar wavelet, Daubechies wavelet, and Symlet wavelet.

[0145] Among them, the method of the step "Generate low-light frequency domain features based on the target high-frequency feature and the target low-frequency feature" can be: use the target high-frequency feature and the target low-frequency feature as the low-light frequency domain features.

[0146] For step S203, the method of the step "Extract the features corresponding to the frequency domain from the bright image features to obtain bright frequency domain features" can be: perform frequency domain decomposition on the bright image features to obtain the candidate high-frequency feature corresponding to the high-frequency domain of the bright image sample and the candidate low-frequency feature corresponding to the low-frequency domain; generate bright frequency domain features based on the candidate high-frequency feature and the candidate low-frequency feature.

[0147] Among them, the method of the step "performing frequency-domain decomposition on the bright image features to obtain the candidate high-frequency features corresponding to the bright image samples in the high-frequency domain and the candidate low-frequency features corresponding to the bright image samples in the low-frequency domain" can specifically refer to the method of the above "performing frequency-domain decomposition on the low-light image features to obtain the target high-frequency features corresponding to the low-light image samples in the high-frequency domain and the target low-frequency features corresponding to the low-light image samples in the low-frequency domain", which will not be elaborated here.

[0148] Among them, the candidate high-frequency features and the candidate low-frequency features can be used as the bright frequency-domain features.

[0149] S204. Based on the low-light frequency-domain features and the bright frequency-domain features, determine the target loss of the image sample pair, and converge the preset image processing model according to the target loss to obtain the image processing model.

[0150] After obtaining the low-light frequency-domain features and the bright frequency-domain features in this application, the target loss can be determined to facilitate the convergence of the preset image processing model. Specifically, for step S204, the method of the step "based on the low-light frequency-domain features and the bright frequency-domain features, determine the target loss of the image sample pair" can refer to steps S41 to S43:

[0151] S41. Identify the target frequency-domain features corresponding to each frequency domain in the low-light frequency-domain features, and screen out the candidate frequency-domain features corresponding to each frequency domain in the bright frequency-domain features.

[0152] Among them, the low-light frequency-domain features can include multiple sub-low-light frequency-domain features. For example, the sub-low-light frequency-domain features can include the target high-frequency features and can also include the target low-frequency features.

[0153] For step S41, the method of the step "identify the target frequency-domain features corresponding to each frequency domain in the low-light frequency-domain features" can be: obtain the sorting serial numbers of each sub-low-light frequency-domain feature included in the low-light frequency-domain features; based on the sorting serial numbers, extract the target frequency-domain features corresponding to each frequency domain from the sub-low-light frequency-domain features.

[0154] Among them, when generating the low-light frequency-domain features, each sub-low-light frequency-domain feature has a corresponding sorting serial number.

[0155] Specifically, the method of the step "based on the sorting serial numbers, extract the target frequency-domain features corresponding to each frequency domain from the sub-low-light frequency-domain features" can be: determine the target sorting serial numbers corresponding to the high-frequency domain and the candidate sorting serial numbers corresponding to the low-frequency domain in the sorting serial numbers; extract the sub-low-light frequency-domain features corresponding to the target sorting serial numbers from the sub-low-light frequency-domain features, and use the sub-low-light frequency-domain features corresponding to the target sorting serial numbers as the target frequency-domain features corresponding to the high-frequency domain; extract the sub-low-light frequency-domain features corresponding to the candidate sorting serial numbers from the sub-low-light frequency-domain features, and use the sub-low-light frequency-domain features corresponding to the candidate sorting serial numbers as the target frequency-domain features corresponding to the low-frequency domain.

[0156] Among them, the bright frequency domain features may include multiple sub-bright frequency domain features. For example, the sub-dark frequency domain features may include candidate high-frequency features and may also include candidate low-frequency features.

[0157] Regarding step S41, the method of the step "screening out the candidate frequency domain features corresponding to each frequency domain in the bright frequency domain features" may be: obtaining the candidate sorting numbers of each sub-bright frequency domain feature included in the bright frequency domain features; based on the candidate sorting numbers, extracting the target frequency domain features corresponding to each frequency domain from the sub-bright frequency domain features.

[0158] Among them, when generating the bright frequency domain features, each sub-bright frequency domain feature has a corresponding candidate sorting number.

[0159] Among them, the method of the step "extracting the target frequency domain features corresponding to each frequency domain from the sub-bright frequency domain features based on the candidate sorting numbers" can specifically refer to the above step of "extracting the target frequency domain features corresponding to each frequency domain from the sub-dark frequency domain features based on the sorting numbers", which will not be elaborated here.

[0160] S42. Calculate the loss between the target frequency domain features and the candidate frequency domain features to obtain the initial loss corresponding to the image sample pair in each target frequency domain.

[0161] After the present application obtains the target frequency domain features and the candidate frequency domain features, it can calculate the loss between the target frequency domain features and the candidate frequency domain features. Specifically, regarding step S42, when the target frequency domain includes a high-frequency domain and a low-frequency domain, the method of the step "calculating the loss between the target frequency domain features and the candidate frequency domain features to obtain the initial loss corresponding to the image sample pair in each target frequency domain" may be: extracting the target high-frequency features corresponding to the dark image sample in the high-frequency domain and the target low-frequency features corresponding to the dark image sample in the low-frequency domain from the target frequency domain features; screening out the candidate high-frequency features corresponding to the bright image sample in the high-frequency domain and the candidate low-frequency features corresponding to the bright image sample in the low-frequency domain from the candidate frequency domain features; calculating the initial high-frequency loss of the image sample pair in the high-frequency domain based on the target high-frequency features and the candidate high-frequency features, and calculating the initial low-frequency loss of the target image sample pair in the low-frequency domain according to the target low-frequency features and the candidate low-frequency features; determining the initial loss corresponding to the image sample pair in each target frequency domain based on the initial high-frequency loss and the initial low-frequency loss.

[0162] Among them, the loss function can be used to calculate the initial high-frequency loss of the image sample pair in the high-frequency domain based on the target high-frequency features and the candidate high-frequency features; the loss function can be used to calculate the initial low-frequency loss of the target image sample pair in the low-frequency domain according to the target low-frequency features and the candidate low-frequency features. Among them, the loss function can be the cross-entropy loss function.

[0163] Specifically, the target high-frequency features include multiple sub-target high-frequency features, and the candidate high-frequency features include multiple sub-candidate high-frequency features. Based on this, the method of the step "calculating the initial high-frequency loss of the image sample pair in the high-frequency domain based on the target high-frequency features and the candidate high-frequency features" can be: fusing the sub-target high-frequency features to obtain the target fusion high-frequency features corresponding to the low-light image sample in the high-frequency domain; fusing the sub-candidate high-frequency features to obtain the candidate fusion high-frequency features corresponding to the bright image sample in the high-frequency domain; calculating the initial high-frequency loss corresponding to the low-light image sample and the bright image sample in the high-frequency domain based on the target fusion high-frequency features and the candidate fusion high-frequency features.

[0164] Among them, the present application can merge the sub-target high-frequency features to obtain the target fusion high-frequency features corresponding to the low-light image sample in the high-frequency domain; similarly, the sub-candidate high-frequency features can be merged to obtain the candidate fusion high-frequency features corresponding to the bright image sample in the high-frequency domain.

[0165] Among them, the initial high-frequency loss and the initial low-frequency loss can be used as the initial loss.

[0166] S43. Fuse each initial loss to obtain the target loss of the image sample pair.

[0167] For step S43, the method of the step "fusing each initial loss to obtain the target loss of the image sample pair" can be: obtaining the weight corresponding to the high-frequency domain and the candidate weight corresponding to the low-frequency domain, and the weight is greater than the candidate weight; weighting the initial high-frequency loss and the initial low-frequency loss based on the weight and the candidate weight to obtain the target loss of the image sample pair.

[0168] It can be understood here that the present application sets the weight corresponding to the high-frequency domain to be greater than the candidate weight corresponding to the low-frequency domain, which can make the preset image processing model pay more attention to the features in the high-frequency domain such as the target high-frequency features and the candidate high-frequency features; since the features in the high-frequency domain contain the edge information of the low-light image sample or the bright image sample, the trained image processing model can make more accurate predictions.

[0169] For steps S203 to S204, a specific example is described here. For example, the present application can perform frequency domain decomposition on the bright coding feature, the bright decoding feature, and the bright convolution feature included in the low-light image feature respectively to obtain the target high-frequency feature and the target low-frequency feature corresponding to the bright coding feature, the target high-frequency feature and the target low-frequency feature corresponding to the bright decoding feature, and the target high-frequency feature and the target low-frequency feature corresponding to the bright convolution feature. Among them, each target high-frequency feature can include multiple sub-target high-frequency features.

[0170] Here, taking the nth candidate low-light encoding feature enn in the low-light encoding feature en = [en1, …, enn] as an example, the frequency-domain decomposition is described as shown in the following formula (1):

[0171] f_engpi1, f_engpi2, f_engpi3, f_enlpi = F_dwt(enn) Formula (1)

[0172] Among them, F_dwt() may refer to the discrete wavelet transform function; f_engpi1, f_engpi2, and f_engpi3 may refer to the three sub-target high-frequency features included in the target high-frequency feature corresponding to the nth candidate low-light encoding feature enn; f_enlpi may refer to the target low-frequency feature corresponding to the nth candidate low-light encoding feature enn.

[0173] Then, the present application can extract the target high-frequency feature in the high-frequency domain and the target low-frequency feature in the low-frequency domain; for each low-light image feature of the present application, the corresponding target high-frequency feature and target low-frequency feature can be obtained through the above formula (1); for each bright image feature of the present application, the corresponding candidate high-frequency feature and candidate low-frequency feature can also be obtained through the above formula (1). Then, as Figure 3 shown, based on the target high-frequency feature and the candidate high-frequency feature, the initial high-frequency loss in the high-frequency domain of the image sample pair can be calculated; based on the target low-frequency feature and the candidate low-frequency feature, the initial low-frequency loss in the low-frequency domain of the target image sample pair can be calculated.

[0174] Among them, for the three sub-target high-frequency features f_engpi1, f_engpi2, and f_engpi3 included in the target high-frequency feature, the present application can merge the three sub-target high-frequency features f_engpi1, f_engpi2, and f_engpi3 to obtain the target fused high-frequency feature, as shown in formula (2):

[0175] f_engpi = concat(f_engpi1, f_engpi2, f_engpi3) Formula (2)

[0176] Among them, f_engpi may refer to the target fused high-frequency feature; concat() may refer to the merging operation.

[0177] For each target high-frequency feature of the present application, the corresponding target fused high-frequency feature can be obtained through the above formula (2); for each candidate high-frequency feature of the present application, the corresponding candidate fused high-frequency feature can also be obtained through the above formula (2). Based on this, for the initial high-frequency loss, the present application can calculate the initial high-frequency loss corresponding to the low-light image sample and the bright image sample in the high-frequency domain based on the target fused high-frequency feature and the candidate fused high-frequency feature.

[0178] Next, the initial high-frequency loss and the initial low-frequency loss can be fused to obtain the target loss of the image sample pair, so as to converge the preset image processing model based on the target loss to obtain the image processing model.

[0179] Among them, for the initial high-frequency loss, the present application can calculate it using a loss function. Here, the nth candidate low-light encoded feature enn and the nth candidate bright encoded feature are taken as examples.

[0180] For example, the present application can perform frequency domain decomposition on the nth candidate low-light encoded feature enn to obtain the target fused high-frequency feature and the target low-frequency feature corresponding to the nth candidate low-light encoded feature enn; perform frequency domain decomposition on the nth candidate bright encoded feature edn to obtain the candidate fused high-frequency feature and the candidate low-frequency feature corresponding to the nth candidate bright encoded feature edn. As Figure 4 shown, taking the target fused high-frequency feature and the candidate fused high-frequency feature as examples, the present application can input the target fused high-frequency feature and the candidate fused high-frequency feature into the discriminator, and the discriminator of the preset image processing model uses the loss function to calculate the initial high-frequency loss between the target fused high-frequency feature and the candidate fused high-frequency feature.

[0181] Among them, after generating the candidate fused high-frequency feature and the candidate low-frequency feature of the bright image sample, the candidate fused high-frequency feature and the candidate low-frequency feature can be stored in the cache space. When calculating the initial loss, the candidate fused high-frequency feature and the candidate low-frequency feature are extracted from the cache space to calculate the initial loss.

[0182] Here, it can be understood that the present application uses the loss function to make the target fused high-frequency feature and the candidate fused high-frequency feature close, so that the discriminator cannot distinguish the target fused high-frequency feature and the candidate fused high-frequency feature; similarly, the target low-frequency feature and the candidate low-frequency feature can be made close, so that the discriminator cannot distinguish the target low-frequency feature and the candidate low-frequency feature, thereby making the information under low-light conditions close to the information under bright illumination conditions.

[0183] For step S204, the method of the step "converging the preset image processing model according to the target loss to obtain the image processing model" can be: converging the low-light image processing sub-model of the preset image processing model according to the target loss to obtain the target low-light image processing sub-model, and generating the preset image processing model based on the target low-light image processing sub-model.

[0184] It can be understood here that in this application, the bright image processing sub-model is used to assist in the training of the low-light image processing sub-model, and the model parameters of the bright image processing sub-model are not updated.

[0185] Furthermore, before step S202 in this application, the initial image processing model can also be pre-trained to obtain the preset image processing model, specifically as steps S51 to S53:

[0186] S51. Obtain the initial bright image sample, where the initial bright image sample includes candidate objects under bright lighting conditions.

[0187] For step S51, the method of the step "obtain the initial bright image sample" can be: the electronic device can send an initial sample acquisition request to the target storage server, so that the target storage server extracts the initial bright image sample from the storage space of the target storage server based on the initial sample acquisition request and returns it to the electronic device; the electronic device receives the initial bright image sample. Or, the initial bright image sample can be extracted from the local database of the electronic device.

[0188] Among them, the target storage server can be the same server as the aforementioned storage server; or, the target storage server is a different server from the aforementioned storage server, and both the target storage server and the electronic device can be set in the same local area network.

[0189] S52. Use the initial image processing model to predict the initial bright image sample to obtain the initial three-dimensional feature information corresponding to the initial bright image sample.

[0190] Among them, the initial three-dimensional feature information includes the depth information of the candidate object under bright lighting conditions.

[0191] For step S52, the method of "using the initial image processing model to predict the initial bright image sample to obtain the initial three-dimensional feature information corresponding to the initial bright image sample" can be: using the initial bright image processing sub-model of the initial image processing model to encode the initial image sample to obtain the initial encoded feature; using the initial bright image processing sub-model to decode the initial encoded feature to obtain the initial three-dimensional feature information corresponding to the initial image sample.

[0192] Among them, the initial bright image processing sub-model may include a second initial encoder and a second initial decoder. Based on this, the present application may use the second initial encoder to encode the initial bright image sample to obtain an initial encoded feature; and use the second initial decoder to decode the initial encoded feature to obtain initial three-dimensional feature information.

[0193] Specifically, the second initial encoder of the present application encodes the initial bright image sample to obtain an initial encoded feature. Among them, the initial encoded feature may be expressed as ed’ = [ed1’, …, edn’]. The initial encoded feature may include multiple sub-initial encoded features. Among them, edn may refer to the nth sub-initial encoded feature, and the order of the n sub-initial encoded features may be sorted from largest to smallest in terms of size. When the initial encoded feature of the present application is decoded by the second initial decoder, an initial sample decoded feature is obtained; the second initial convolutional layer of the initial bright image processing sub-model is used to perform convolution on the initial sample decoded feature to obtain initial three-dimensional feature information.

[0194] S53. Based on the initial three-dimensional feature information, converge the initial image processing model to obtain a preset image processing model.

[0195] For step S53, the method of the step “Based on the initial three-dimensional feature information, converge the initial image processing model to obtain a preset image processing model” may be: obtain the three-dimensional feature label corresponding to the initial bright image sample; calculate the candidate loss between the initial three-dimensional feature information and the three-dimensional feature label; and based on the candidate loss, converge the initial image processing model to obtain a preset image processing model.

[0196] Specifically, for step S53, the present application may converge the initial bright image processing sub-model based on the initial three-dimensional feature information to obtain a bright image processing sub-model, and generate a preset image processing model based on the bright image processing sub-model.

[0197] It can be understood here that the bright image processing sub-model includes a second encoder and a second decoder. The second encoder may be the encoder corresponding to the convergence of the second initial encoder, and the second decoder may be the decoder corresponding to the convergence of the second initial decoder.

[0198] S205. Use the image processing model to identify three-dimensional feature information in a low-light image containing a target object.

[0199] Among them, the three-dimensional feature information includes depth information simulating the target object under bright illumination conditions.

[0200] Specifically, after the preset image processing model converges, the obtained image processing model includes a target low-light image processing sub-model, where the target low-light image processing sub-model can be the sub-model corresponding to the converged low-light image processing sub-model.

[0201] Regarding step S205, the present application can use the target low-light image processing sub-model of the image processing model to predict a low-light image containing a target object to obtain three-dimensional feature information.

[0202] The target low-light image processing sub-model can include a first target encoder, a first target decoder, and a first target convolutional layer. Among them, the first target encoder can be the encoder corresponding to the converged first encoder, the first target decoder can be the decoder corresponding to the converged first decoder, and the first target convolutional layer can be the convolutional layer corresponding to the converged first convolutional layer.

[0203] Specifically, the present application can use the first target encoder to encode a low-light image containing a target object to obtain low-light encoded image features; use the first target decoder to decode the low-light encoded image features to obtain low-light decoded image features; use the first convolutional layer to perform convolution on the low-light decoded image features to obtain three-dimensional feature information.

[0204] Among them, the effect diagram of the present application using the image processing model to identify three-dimensional feature information in a low-light image can be as Figure 5 shown.

[0205] In summary, the present application can achieve: 1) performing high-frequency and low-frequency decomposition on information to make the method more focused on information alignment in the high-frequency part; 2) since the low-light convolutional features and bright convolutional features corresponding to the convolutional feature granularity can both be characterized as images, based on this, the present application can achieve aligning and constraining the low-light convolutional features and bright convolutional features in the image space; in addition, the present application can also align and constrain the low-light encoded features and bright encoded features corresponding to the encoded feature granularity, and align and constrain the low-light decoded features and bright decoded features corresponding to the decoded feature granularity to achieve feature alignment and constraint in the feature space; 3) optionally applying alignment in both the high-frequency domain and the low-frequency domain during the encoder and decoder processes of the features.

[0206] The present application can obtain at least one pair of image samples. The pair of image samples includes a low-light image sample captured under low-light conditions and a bright image sample captured under bright light conditions. The preset image processing model is used to perform multi-granularity feature extraction on the pair of image samples to obtain an image feature set corresponding to each feature granularity. The image feature set includes the low-light image features of the low-light image sample and the bright image features of the bright image sample. At least one frequency-domain feature is extracted from the low-light image features to obtain low-light frequency-domain features, and features corresponding to the frequency domain are extracted from the bright image features to obtain bright frequency-domain features. Based on the low-light frequency-domain features and the bright frequency-domain features, the target loss of the pair of image samples is determined, and the preset image processing model is converged according to the target loss to obtain an image processing model. The image processing model is used to identify three-dimensional feature information in the low-light image containing the target object. The three-dimensional feature information includes depth information simulating the target object under bright light conditions. Since the present application can perform multi-granularity feature extraction on the pair of image samples using the preset image processing model to obtain the low-light image features of the low-light image sample and the bright image features of the bright image sample at each feature granularity, in this way, the present application can extract low-light frequency-domain features from the low-light image features and extract bright frequency-domain features from the bright image features. Based on this, the present application can use the target loss obtained based on the low-light frequency-domain features and the bright frequency-domain features to converge the preset image processing model, thereby improving the prediction accuracy of the converged image processing model. Based on this, the present application can use the converged image processing model to predict the low-light image to obtain depth information simulating the target object under bright light conditions, thus improving the accuracy of the three-dimensional feature information extracted by the image processing model.

[0207] According to the method described in the above embodiments, the following will give further detailed examples.

[0208] In this embodiment, it is assumed that the image processing device is specifically integrated in an electronic device, and the electronic device is a server.

[0209] It should be noted here that the present application can be applied to the game business. The three-dimensional data, that is, the three-dimensional feature information, generated by the present application can quickly realize three-dimensional modeling in the game business and improve the development efficiency of the game.

[0210] Such as Figure 6 As shown, an image processing method, the specific process is as follows in steps S501 to S508:

[0211] S501. The electronic device obtains at least one pair of image samples. The pair of image samples includes a low-light image sample captured under low-light conditions and a bright image sample captured under bright light conditions.

[0212] Specifically, the electronic device may send a sample acquisition request to the storage server, so that the storage server extracts low-light image samples and bright image samples from the storage space of the storage server based on the sample acquisition request and returns them to the electronic device; the electronic device receives the low-light image samples and bright image samples and generates at least one image sample pair based on the low-light image samples and bright image samples.

[0213] S502. The electronic device uses a preset image processing model to perform multi-granularity feature extraction on the image sample pair to obtain an image feature set corresponding to each feature granularity.

[0214] Among them, the image feature set includes the low-light image features of the low-light image sample and the bright image features of the bright image sample.

[0215] Specifically, when the feature granularity includes the encoding feature granularity, the electronic device may use the low-light image processing sub-model of the preset image processing model to encode the low-light image sample to obtain the low-light encoding feature corresponding to the low-light image sample; use the bright image processing sub-model of the preset image processing model to encode the bright image sample to obtain the bright encoding feature corresponding to the bright image sample; generate an encoding feature set corresponding to the encoding feature granularity based on the low-light encoding feature and the bright encoding feature, and generate an image feature set corresponding to each feature granularity based on the encoding feature set.

[0216] Further, when the feature granularity further includes the encoding feature granularity, the electronic device may screen out the target low-light encoding feature from the low-light encoding features and use the low-light image processing sub-model to decode the target low-light encoding feature to obtain the low-light decoded feature; extract the target bright encoding feature from the bright encoding features and use the bright image processing sub-model to decode the target bright encoding feature to obtain the bright decoded feature; generate a decoded feature set corresponding to the decoding feature granularity based on the low-light decoded feature and the bright decoded feature, and generate an image feature set corresponding to each feature granularity based on the encoding feature set and the decoded feature set.

[0217] Further, when the feature granularity further includes the convolutional feature granularity, the electronic device may obtain the target low-light decoded feature from the low-light decoded features and use the low-light image processing sub-model to perform convolution on the target low-light decoded feature to obtain the low-light convolutional feature; extract the target bright decoded feature from the bright decoded features and use the bright image processing sub-model to perform convolution on the target bright decoded feature to obtain the bright convolutional feature; generate a convolutional feature set corresponding to the convolutional feature granularity based on the low-light convolutional feature and the bright convolutional feature, and generate an image feature set corresponding to each feature granularity based on the convolutional feature set, the encoding feature set, and the decoded feature set.

[0218] S503. The electronic device extracts the features of at least one frequency domain from the low-light image features to obtain the low-light frequency domain features, and extracts the features corresponding to the frequency domain from the bright image features to obtain the bright frequency domain features.

[0219] Specifically, the electronic device can perform frequency domain decomposition on the low-light image features to obtain the target high-frequency features corresponding to the high-frequency domain of the low-light image samples and the target low-frequency features corresponding to the low-frequency domain; based on the target high-frequency features and the target low-frequency features, generate the low-light frequency domain features.

[0220] Specifically, the electronic device can perform frequency domain decomposition on the bright image features to obtain the candidate high-frequency features corresponding to the high-frequency domain of the bright image samples and the candidate low-frequency features corresponding to the low-frequency domain; based on the candidate high-frequency features and the candidate low-frequency features, generate the bright frequency domain features.

[0221] S504. The electronic device identifies the target frequency domain features corresponding to each frequency domain in the low-light frequency domain features, and filters out the candidate frequency domain features corresponding to each frequency domain in the bright frequency domain features.

[0222] S505. The electronic device calculates the loss between the target frequency domain features and the candidate frequency domain features to obtain the initial loss corresponding to each target frequency domain of the image sample pair.

[0223] Specifically, the electronic device can extract the target high-frequency features corresponding to the high-frequency domain of the low-light image samples and the target low-frequency features corresponding to the low-frequency domain from the target frequency domain features; filter out the candidate high-frequency features corresponding to the high-frequency domain of the bright image samples and the candidate low-frequency features corresponding to the low-frequency domain from the candidate frequency domain features; based on the target high-frequency features and the candidate high-frequency features, calculate the initial high-frequency loss of the image sample pair in the high-frequency domain, and according to the target low-frequency features and the candidate low-frequency features, calculate the initial low-frequency loss of the target image sample pair in the low-frequency domain; based on the initial high-frequency loss and the initial low-frequency loss, determine the initial loss corresponding to each target frequency domain of the image sample pair.

[0224] S506. The electronic device fuses each initial loss to obtain the target loss of the image sample pair.

[0225] Specifically, the electronic device can obtain the weight corresponding to the high-frequency domain and the candidate weight corresponding to the low-frequency domain, and the weight is greater than the candidate weight; based on the weight and the candidate weight, weight the initial high-frequency loss and the initial low-frequency loss to obtain the target loss of the image sample pair.

[0226] S507. The electronic device converges the preset image processing model according to the target loss to obtain the image processing model.

[0227] Specifically, the electronic device may converge the low-light image processing sub-model of the preset image processing model according to the target loss to obtain the target low-light image processing sub-model, and generate the preset image processing model based on the target low-light image processing sub-model.

[0228] S508. The electronic device uses the image processing model to identify three-dimensional feature information in a low-light image including a target object.

[0229] Specifically, the electronic device may use the target low-light image processing sub-model of the image processing model to predict a low-light image including a target object to obtain three-dimensional feature information. Among them, the three-dimensional feature information includes depth information simulating the target object under bright illumination conditions.

[0230] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.

[0231] This application can obtain at least one pair of image samples. The pair of image samples includes a low-light image sample taken under low-light conditions and a bright image sample taken under bright illumination conditions. The preset image processing model is used to perform multi-granularity feature extraction on the pair of image samples to obtain an image feature set corresponding to each feature granularity. The image feature set includes the low-light image features of the low-light image sample and the bright image features of the bright image sample. At least one frequency domain feature is extracted from the low-light image features to obtain low-light frequency domain features, and the features corresponding to the frequency domain are extracted from the bright image features to obtain bright frequency domain features. Based on the low-light frequency domain features and the bright frequency domain features, the target loss of the pair of image samples is determined, and the preset image processing model is converged according to the target loss to obtain an image processing model. The image processing model is used to identify three-dimensional feature information in a low-light image including a target object. The three-dimensional feature information includes depth information simulating the target object under bright illumination conditions. Since this application can perform multi-granularity feature extraction on the pair of image samples using the preset image processing model to obtain the low-light image features of the low-light image sample and the bright image features of the bright image sample at each feature granularity, in this way, this application can extract low-light frequency domain features from the low-light image features and extract bright frequency domain features from the bright image features. Based on this, this application can use the target loss obtained based on the low-light frequency domain features and the bright frequency domain features to converge the preset image processing model, thereby improving the prediction accuracy of the converged image processing model. Based on this, this application can use the converged image processing model to predict the low-light image to obtain depth information simulating the target object under bright illumination conditions, thus improving the accuracy of the three-dimensional feature information extracted by the image processing model.

[0232] To better implement the above method, an embodiment of the present application further provides an image processing apparatus, which can be integrated in an electronic device, such as a server or a terminal, etc. The terminal may include a tablet computer, a laptop computer, and / or a personal computer, etc.

[0233] For example, as Figure 7 shown, the image processing apparatus may include an acquisition unit 301, a first extraction unit 302, a second extraction unit 303, a determination unit 304, and an identification unit 305, as follows:

[0234] (1) Acquisition unit 301;

[0235] The acquisition unit 301 can be used to acquire at least one pair of image samples. The pair of image samples includes a low-light image sample taken under low-light conditions and a bright image sample taken under bright light conditions.

[0236] (2) First extraction unit 302;

[0237] The first extraction unit 302 can be used to perform multi-granularity feature extraction on the pair of image samples using a preset image processing model to obtain an image feature set corresponding to each feature granularity. The image feature set includes low-light image features of the low-light image sample and bright image features of the bright image sample.

[0238] For example, the first extraction unit 302 can be used to, when the feature granularity includes an encoded feature granularity, encode the low-light image sample using the low-light image processing sub-model of the preset image processing model to obtain the low-light encoded feature corresponding to the low-light image sample; encode the bright image sample using the bright image processing sub-model of the preset image processing model to obtain the bright encoded feature corresponding to the bright image sample; generate an encoded feature set corresponding to the encoded feature granularity based on the low-light encoded feature and the bright encoded feature, and generate an image feature set corresponding to each feature granularity based on the encoded feature set.

[0239] (3) Second extraction unit 303;

[0240] The second extraction unit 303 can be used to extract at least one frequency-domain feature from the low-light image features to obtain low-light frequency-domain features, and extract the features corresponding to the frequency domain from the bright image features to obtain bright frequency-domain features.

[0241] For example, the second extraction unit 303 can be used to perform frequency-domain decomposition on the low-light image features to obtain the target high-frequency features corresponding to the high-frequency domain of the low-light image sample and the target low-frequency features corresponding to the low-frequency domain; generate low-light frequency-domain features based on the target high-frequency features and the target low-frequency features.

[0242] (4) Determination unit 304;

[0243] A determination unit 304, which can be used to determine the target loss of an image sample pair based on dark-light frequency domain features and bright-light frequency domain features, and converge a preset image processing model according to the target loss to obtain the image processing model.

[0244] For example, the determination unit 304 can be used to identify the target frequency domain feature corresponding to each frequency domain in the dark-light frequency domain features, and screen out the candidate frequency domain features corresponding to each frequency domain in the bright-light frequency domain features; calculate the loss between the target frequency domain feature and the candidate frequency domain feature to obtain the initial loss corresponding to the image sample pair under each target frequency domain; fuse each initial loss to obtain the target loss of the image sample pair.

[0245] Another example is that the determination unit 304 can be used to obtain an initial bright image sample, where the initial bright image sample includes candidate objects under bright light conditions; use the initial image processing model to predict the initial bright image sample to obtain the initial three-dimensional feature information corresponding to the initial bright image sample, where the initial three-dimensional feature information includes the depth information of the candidate objects under bright light conditions; converge the initial image processing model based on the initial three-dimensional feature information to obtain the preset image processing model.

[0246] (5) An identification unit 305;

[0247] The identification unit 305 can be used to identify three-dimensional feature information in a dark-light image containing a target object by using the image processing model, where the three-dimensional feature information includes the depth information of the simulated target object under bright light conditions.

[0248] As can be seen from the above, the acquisition unit 301 of the present application can be used to acquire at least one pair of image samples. The pair of image samples includes a low-light image sample taken under low-light conditions and a bright image sample taken under bright light conditions. The first extraction unit 302 can be used to perform multi-granularity feature extraction on the pair of image samples by using a preset image processing model, and obtain an image feature set corresponding to each feature granularity. The image feature set includes a low-light image feature of the low-light image sample and a bright image feature of the bright image sample. The second extraction unit 303 can be used to extract features in at least one frequency domain from the low-light image features to obtain low-light frequency domain features, and extract features corresponding to the frequency domain from the bright image features to obtain bright frequency domain features. The determination unit 304 can be used to determine the target loss of the pair of image samples based on the low-light frequency domain features and the bright frequency domain features, and converge the preset image processing model according to the target loss to obtain an image processing model. The recognition unit 305 can be used to use the image processing model to recognize three-dimensional feature information in a low-light image including a target object. The three-dimensional feature information includes depth information simulating the target object under bright light conditions. Since the present application can perform multi-granularity feature extraction on a pair of image samples by using a preset image processing model to obtain a low-light image feature of a low-light image sample and a bright image feature of a bright image sample under each feature granularity, in this way, the present application can extract low-light frequency domain features from the low-light image features and extract bright frequency domain features from the bright image features. Based on this, the present application can use the target loss obtained based on the low-light frequency domain features and the bright frequency domain features to converge the preset image processing model, so as to improve the prediction accuracy of the converged image processing model. Based on this, the present application can use the converged image processing model to predict a low-light image to obtain depth information simulating the target object under bright light conditions, thereby improving the accuracy of the three-dimensional feature information extracted by the image processing model.

[0249] The embodiment of the present application also provides an electronic device, as Figure 8 shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of the present invention. Specifically:

[0250] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404 and other components. Those skilled in the art can understand that Figure 8 the structure of the electronic device shown in

[0251] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 402, and by invoking the data stored in the memory 402, it executes various functions of the electronic device and processes data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.

[0252] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store the data created according to the use of the electronic device. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0253] The electronic device also includes a power supply 403 that powers each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0254] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0255] Although not shown, the electronic device may also include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to realize various functions as follows:

[0256] Obtain at least one pair of image samples, where the pair of image samples includes a low-light image sample taken under low-light conditions and a bright image sample taken under bright light conditions; use a preset image processing model to perform multi-granularity feature extraction on the pair of image samples to obtain an image feature set corresponding to each feature granularity, and the image feature set includes the low-light image features of the low-light image sample and the bright image features of the bright image sample; extract at least one frequency-domain feature from the low-light image features to obtain low-light frequency-domain features, and extract the features corresponding to the frequency domain from the bright image features to obtain bright frequency-domain features; based on the low-light frequency-domain features and the bright frequency-domain features, determine the target loss of the pair of image samples, and converge the preset image processing model according to the target loss to obtain an image processing model; use the image processing model to identify three-dimensional feature information in the low-light image containing the target object, and the three-dimensional feature information includes the depth information simulating the target object under bright light conditions.

[0257] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.

[0258] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a computer program or by controlling related hardware through a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0259] Therefore, an embodiment of the present application provides a computer-readable storage medium in which a computer program is stored, and the computer program can be loaded by a processor to execute any one of the image processing methods provided by the embodiments of the present application.

[0260] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.

[0261] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), a magnetic disk or an optical disc, etc.

[0262] Since the instructions stored in the computer-readable storage medium can execute the steps in any one of the image processing methods provided by the embodiments of the present application, the beneficial effects achievable by any one of the image processing methods provided by the embodiments of the present application can be achieved. For details, reference may be made to the previous embodiments and will not be elaborated here.

[0263] Among them, according to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in the various alternative implementations provided in the above embodiments.

[0264] The above has introduced in detail an image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. An image processing method, characterized in that: include: Acquire at least one image sample pair, the image sample pair comprising a dark-light image sample taken under a dark-light condition and a bright-light image sample taken under a bright-light condition; Using a preset image processing model to perform multi-granularity feature extraction on the image sample pair, to obtain an image feature set corresponding to each feature granularity, the image feature set including dark-light image features of the dark-light image sample and bright-light image features of the bright-light image sample; Extracting at least one frequency domain feature from the dark light image features to obtain a dark light frequency domain feature, and extracting a feature corresponding to the frequency domain from the bright light image features to obtain a bright frequency domain feature; Determine the target loss of the image sample pair based on the dark light frequency domain feature and the bright light frequency domain feature, and converge the preset image processing model according to the target loss to obtain an image processing model; The image processing model is used to identify three-dimensional feature information in a dark light image containing a target object, wherein the three-dimensional feature information includes depth information simulating the target object under bright light conditions.

2. The image processing method according to claim 1, characterized in that: The determining the target loss of the image sample pair based on the dark light frequency domain feature and the bright light frequency domain feature comprises: Identify the target frequency domain features corresponding to each frequency domain in the dark frequency domain features, and select the candidate frequency domain features corresponding to each frequency domain in the bright frequency domain features; Calculating the loss between the target frequency domain feature and the candidate frequency domain feature to obtain the initial loss corresponding to the image sample pair in each target frequency domain; Each of the initial losses is fused to obtain the target loss of the image sample pair.

3. The image processing method according to claim 2, characterized in that: The target frequency domain includes a high frequency domain and a low frequency domain; the calculating the loss between the target frequency domain feature and the candidate frequency domain feature to obtain the initial loss corresponding to the image sample pair in each target frequency domain includes: Extracting the target high-frequency features corresponding to the dark-light image sample in the high-frequency domain and the target low-frequency features corresponding to the low-frequency domain from the target frequency-domain features; Screening out candidate high-frequency features corresponding to the bright image sample in the high-frequency domain and candidate low-frequency features corresponding to the bright image sample in the low-frequency domain from the candidate frequency-domain features; Based on the target high-frequency feature and the candidate high-frequency feature, calculate the initial high-frequency loss of the image sample pair in the high-frequency domain, and according to the target low-frequency feature and the candidate low-frequency feature, calculate the initial low-frequency loss of the target image sample pair in the low-frequency domain; Based on the initial high-frequency loss and the initial low-frequency loss, an initial loss corresponding to the image sample pair in each target frequency domain is determined.

4. The image processing method according to claim 3, characterized in that: The target high-frequency feature includes a plurality of sub-target high-frequency features, and the candidate high-frequency feature includes a plurality of sub-candidate high-frequency features; and the initial high-frequency loss of the image sample pair in the high-frequency domain is calculated based on the target high-frequency feature and the candidate high-frequency feature, including: Fusing the sub-target high-frequency features to obtain the target fused high-frequency features corresponding to the dark-light image sample in the high-frequency domain; Fusing the sub-candidate high-frequency features to obtain candidate fused high-frequency features corresponding to the bright image sample in the high-frequency domain; Based on the target fused high-frequency features and the candidate fused high-frequency features, an initial high-frequency loss of the image sample pair in the high-frequency domain is calculated.

5. The image processing method according to claim 3, characterized in that: The fusing of each of the initial losses to obtain the target loss of the image sample pair includes: Obtaining a weight corresponding to the high frequency domain and a candidate weight corresponding to the low frequency domain, wherein the weight is greater than the candidate weight; Based on the weight and the candidate weight, the initial high-frequency loss and the initial low-frequency loss are weighted to obtain a target loss for the image sample pair.

6. The image processing method according to claim 1, characterized in that: The method of using a preset image processing model to perform multi-granularity feature extraction on the image sample pair to obtain an image feature set corresponding to each feature granularity includes: When the feature granularity includes a coding feature granularity, a dark-light image processing sub-model of the preset image processing model is used to encode the dark-light image sample to obtain a dark-light coding feature corresponding to the dark-light image sample; Using the bright image processing sub-model of the preset image processing model, encoding the bright image sample to obtain a bright coding feature corresponding to the bright image sample; Based on the dark light coding feature and the bright coding feature, a coding feature set corresponding to the coding feature granularity is generated, and based on the coding feature set, an image feature set corresponding to each feature granularity is generated.

7. The image processing method according to claim 6, characterized in that: The step of generating an image feature set corresponding to each feature granularity based on the encoding feature set includes: When the feature granularity also includes a decoding feature granularity, a target dark light coding feature is screened out from the dark light coding feature, and the target dark light coding feature is decoded using the dark light image processing sub-model to obtain a dark light decoding feature; Extracting a target bright coding feature from the bright coding feature, and decoding the target bright coding feature using the bright image processing sub-model to obtain a bright decoding feature; Based on the dark light decoding feature and the bright light decoding feature, a decoding feature set corresponding to the decoding feature granularity is generated, and based on the encoding feature set and the decoding feature set, an image feature set corresponding to each feature granularity is generated.

8. The image processing method according to claim 7, characterized in that: The step of generating an image feature set corresponding to each feature granularity based on the encoding feature set and the decoding feature set includes: When the feature granularity also includes a convolution feature granularity, obtaining a target dark light decoding feature from the dark light decoding feature, and convolving the target dark light decoding feature using the dark light image processing sub-model to obtain a dark light convolution feature; Extracting a target bright decoding feature from the bright decoding feature, and convolving the target bright decoding feature using the bright image processing sub-model to obtain a bright convolution feature; Based on the dark-light convolution features and the bright-light convolution features, a convolution feature set corresponding to the convolution feature granularity is generated, and based on the convolution feature set, the encoding feature set and the decoding feature set, an image feature set corresponding to each feature granularity is generated.

9. The image processing method according to claim 6, characterized in that: The coding feature granularity includes a plurality of sub-coding feature granularities; the dark-light image processing sub-model of the preset image processing model is adopted to encode the dark-light image sample to obtain the dark-light coding feature corresponding to the dark-light image sample, including: Using the dark-light image processing sub-model, encoding the dark-light image sample, and obtaining initial dark-light coding features corresponding to each coding level of the dark-light image processing sub-model; In the initial dark light coding features, a candidate dark light coding feature corresponding to each sub-coding feature granularity is selected; Based on the candidate dark-light coding features, a dark-light coding feature corresponding to the dark-light image sample is generated.

10. The image processing method according to claim 9, characterized in that: The step of selecting candidate dark light coding features corresponding to each sub-coding feature granularity from the initial dark light coding features includes: In the coding level, determining a target coding level corresponding to each of the sub-coding feature granularities; Selecting reference dark light coding features corresponding to each target coding level from the initial dark light coding features; The reference dark light coding feature is used as a candidate dark light coding feature corresponding to the sub-coding feature granularity.

11. The image processing method according to claim 1, characterized in that: The step of extracting at least one frequency domain feature from the dark light image feature to obtain the dark light frequency domain feature comprises: Decomposing the dark-light image features in the frequency domain to obtain target high-frequency features corresponding to the dark-light image samples in the high-frequency domain and target low-frequency features corresponding to the dark-light image samples in the low-frequency domain; The dark light frequency domain feature is generated based on the target high frequency feature and the target low frequency feature.

12. The image processing method according to claim 1, characterized in that: Before extracting multi-granularity features from the image sample pair using a preset image processing model to obtain an image feature set corresponding to each feature granularity, the method further includes: Acquire an initial bright image sample, the initial bright image sample including a candidate object under bright lighting conditions; Using an initial image processing model to predict the initial bright image sample, to obtain initial three-dimensional feature information corresponding to the initial bright image sample, the initial three-dimensional feature information including depth information of the candidate object under the bright lighting condition; Based on the initial three-dimensional feature information, the initial image processing model is converged to obtain the preset image processing model.

13. The image processing method according to claim 12, characterized in that: The using the initial image processing model to predict the initial bright image sample to obtain the initial three-dimensional feature information corresponding to the initial bright image sample includes: Encoding the initial image sample using the initial bright image processing sub-model of the initial image processing model to obtain an initial encoding feature; Decoding the initial coded features using the initial bright image processing sub-model to obtain initial three-dimensional feature information corresponding to the initial image sample; The method of converging the initial image processing model based on the initial three-dimensional feature information to obtain the preset image processing model includes: converging the initial bright image processing sub-model based on the initial three-dimensional feature information to obtain a bright image processing sub-model, and generating the preset image processing model based on the bright image processing sub-model.

14. An image processing device, characterized in that: include: an acquisition unit, configured to acquire at least one image sample pair, wherein the image sample pair comprises a dark-light image sample taken under a dark-light condition and a bright-light image sample taken under a bright-light condition; A first extraction unit is used to perform multi-granularity feature extraction on the image sample pair using a preset image processing model to obtain an image feature set corresponding to each feature granularity, wherein the image feature set includes dark-light image features of the dark-light image sample and bright-light image features of the bright-light image sample; A second extraction unit is used to extract at least one frequency domain feature from the dark light image feature to obtain a dark light frequency domain feature, and to extract a feature corresponding to the frequency domain from the bright light image feature to obtain a bright frequency domain feature; a determining unit, configured to determine a target loss of the image sample pair based on the dark light frequency domain feature and the bright light frequency domain feature, and converge the preset image processing model according to the target loss to obtain an image processing model; The recognition unit is used to recognize three-dimensional feature information in a dark light image containing a target object by using the image processing model, wherein the three-dimensional feature information includes depth information simulating the target object under bright light conditions.

15. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores an application program, and the processor is used to run the application program in the memory to execute the steps in the image processing method according to any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the image processing method according to any one of claims 1 to 13.

17. A computer program product, characterized in that The computer program product stores a computer program, and the computer program is suitable for being loaded by a processor to execute the image processing method according to any one of claims 1 to 13.