Image normalization method, model training method, electronic equipment and storage medium
By determining the depth interval of the target object in the depth image and performing specific normalization processing, the problems of poor contrast and strong background noise in the depth image are solved, and the reliability of the image is improved.
Patent Information
- Application Number
- CN202311705237.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-13
AI Technical Summary
The image contrast of depth images is poor and has strong background noise, resulting in poor reliability in subsequent applications.
By determining the depth interval of the target object in the depth image, the pixel values within the first depth interval are mapped to the gray value interval [0,255], and the pixel values within the second depth interval are set to zero, and then uniformly normalized to [0,1].
The target object information in the depth image is effectively retained, the influence of interference information is eliminated, and the reliability in subsequent applications is improved.
Smart Images

Figure CN120147653A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of deep learning, and in particular, to an image normalization method, a model training method, an electronic device, and a storage medium.
Background Art
[0002] Generally, the depth images collected by a depth camera can be used for visual display or depth model training. However, due to the limitations of the imaging principle of the depth images themselves, their image contrast is often poor, and there is strong background noise. Therefore, when performing subsequent applications on depth images, it is necessary to perform normalization processing on them.
[0003] In related technologies, normalization is often performed on the entire depth image. Then, there may be a large amount of invalid information in the normalized image, resulting in poor reliability of the depth image in subsequent application processes.
Summary of the Invention
[0004] The embodiments of the present application provide an image normalization method, a model training method, an electronic device, and a storage medium, which can mainly retain the effective information part of the original depth image in the normalized image, thereby ensuring the reliability in subsequent application processes.
[0005] In a first aspect, an image normalization method is proposed in the embodiments of the present application. The method includes:
[0006] Obtain a target depth image including a target object;
[0007] Determine a first depth interval where the target object is located in the target depth image, and a second depth interval that does not include the target object;
[0008] Map the depth values of each pixel point in the first depth interval to the gray value interval [0, 255];
[0009] Set the depth values of each pixel point in the second depth interval to zero;
[0010] Normalize each pixel point whose depth value in the first depth interval is mapped to the gray value interval [0, 255], and each pixel point whose depth value in the second depth interval is set to zero, to [0, 1].
[0011] In the embodiments of the present application, the depth value corresponding to the target object in the target depth image can be considered to be within a certain depth interval range, and outside the above depth interval range, it can be considered as interference information, such as background interference and foreground interference. Therefore, when performing normalization, on the one hand, the pixel values of the pixel points within the depth interval range where the target object is located can be normalized to [0, 255], and on the other hand, the pixel values of the pixel points outside the above depth interval range can be directly set to zero, so as to eliminate the influence of interference information. On this basis, it is then uniformly normalized to [0, 1]. This method can enable the normalized image to mainly retain the effective information part in the original depth image, that is, the information of the target object, thus ensuring the reliability in the subsequent application process.
[0012] Optionally, determining the first depth interval where the target object is located in the target depth image includes:
[0013] Performing key point detection on the target depth image to obtain multiple key points corresponding to the target object and the depth value of each key point;
[0014] Based on the depth values of each key point, determining the average depth value of the multiple key points;
[0015] Taking the difference between the average depth value and the preset first depth value in the first direction as the lower limit value, and taking the sum of the average depth value and the preset second depth value in the second direction as the upper limit value. The lower limit value and the upper limit value together constitute the first depth interval. The first direction is the side close to the depth camera that collects the target depth image, and the second direction is the side far from the depth camera that collects the depth image.
[0016] In the embodiments of the present application, the target object can be considered to include multiple key points. Then, through the above multiple key points, the target object can be located in the target depth image. Thus, based on the depth values of each of the above multiple key points, the average depth value of the multiple key points is determined, and taking this average depth value as a reference, a preset depth is extended respectively towards the side close to the depth camera and away from the depth camera, so as to more accurately form the depth interval range where the target object is located.
[0017] Optionally, the target object is a human face, and the first depth value is less than the second depth value.
[0018] In the embodiments of the present application, when the target object is a human face, when extending a preset depth respectively towards the side close to the depth camera and away from the depth camera with this average depth value as a reference, since the depth range of the human face is smaller than the depth range of the back of the head, the depth extended towards the side close to the depth camera is less than the depth extended towards the side far from the depth camera, so as to more accurately form the depth interval range where the human face is located.
[0019] Optionally, the first depth value is in the range of [50 mm, 100 mm], and the second depth value is in the range of [150 mm, 250 mm].
[0020] In the embodiment of the present application, when the target object is a human face, taking this depth average value as a reference, it extends [50 mm, 100 mm] respectively towards the side close to the depth camera and extends [150 mm, 250 mm] away from the depth camera, so as to more accurately form the depth interval range where the human face is located.
[0021] Optionally, the first depth value is 100 mm, and the second depth value is 200 mm.
[0022] In the embodiment of the present application, when the target object is a human face, taking this depth average value as a reference, it extends 100 mm respectively towards the side close to the depth camera and extends 200 mm away from the depth camera, retaining as much information of the target object as possible and excluding as much interference information as possible, so as to more accurately form the depth interval range where the human face is located.
[0023] Optionally, before normalizing to [0, 1] each pixel point that maps the depth value in the first depth interval to the gray value interval [0, 255] and each pixel point that sets the depth value in the second depth interval to zero, the method further includes:
[0024] Setting each pixel point that sets the depth value in the second depth interval to zero to 255;
[0025] Performing local histogram equalization processing on the common of each pixel point that maps the depth value in the first depth interval to the gray value interval [0, 255] and each pixel point that sets the second depth interval to 255;
[0026] Setting each pixel point that sets the second depth interval to 255 to zero;
[0027] Normalizing to [0, 1] each pixel point that maps the depth value in the first depth interval to the gray value interval [0, 255] and each pixel point that sets the depth value in the second depth interval to zero, includes:
[0028] Normalizing to [0, 1] each pixel point that maps the depth value in the first depth interval to the gray value interval [0, 255] and has undergone the local histogram equalization processing and each pixel point that sets the depth value in the second depth interval to zero.
[0029] In the embodiment of the present application, histogram equalization is performed on each pixel in the first depth interval normalized to [0,255], thereby increasing the contrast between the pixels representing the effective information part, and each pixel in the second depth interval that is set to zero is set to 255. Even if histogram equalization is performed on each pixel in the second depth interval, there will be no change in each pixel in the second depth interval, and after the histogram equalization is completed, the pixel value of each pixel in the second depth interval is reset to zero. On this basis, unified normalization to [0,1] is performed. This method can make the normalized image mainly retain the effective information part of the original depth image, while improving the contrast of the effective information part as much as possible, without amplifying the interference information, thereby ensuring reliability in subsequent applications.
[0030] Optionally, mapping the depth value of each pixel in the first depth interval to a grayscale value interval [0, 255] includes:
[0031] For an i-th pixel in the first depth interval, calculating a first depth difference between a depth value of the i-th pixel and the depth lower limit;
[0032] Calculating a second depth difference between the upper depth limit and the lower depth limit;
[0033] Calculating a depth difference ratio between the first depth difference and the second depth difference;
[0034] The product of the depth difference ratio and 255 is calculated to obtain the grayscale value corresponding to the i-th pixel.
[0035] In an embodiment of the present application, for any pixel point within the first depth interval, the first depth difference between the depth value of the pixel point and the lower limit value of the first depth interval, and the second depth difference between the upper limit value and the lower limit value of the first depth interval are first calculated, and then the ratio of the above-mentioned first depth difference to the second depth difference is calculated. The ratio can be considered to be used to measure the extent to which the depth of the above-mentioned pixel point is located in the first depth interval. Finally, the ratio is multiplied by 255 to obtain the grayscale value corresponding to the pixel point. Through the above method, the grayscale value corresponding to each pixel point in the first depth interval can be calculated.
[0036] In a second aspect, an embodiment of the present application provides a model training method, the method comprising:
[0037] Acquire multiple target depth images containing the target object;
[0038] Preprocessing each of the target depth images, wherein the preprocessing at least includes the normalization processing described in any embodiment of the first aspect;
[0039] Train a reference model with the pre - processed multiple target depth images to obtain a target model.
[0040] In the embodiments of the present application, first, multiple target depth images containing a target object are collected. Then, each of the above - mentioned target depth images is pre - processed. For example, at least the normalization method described in the first aspect is used for normalization processing. Then, the pre - processed target depth images are input into the reference model for training, so as to obtain an available target model. Since each target depth image mainly contains the effective information of the target object, it can be considered that the target model has high performance.
[0041] In a third aspect, embodiments of the present application provide a normalization device for images. The device includes:
[0042] An acquisition unit for acquiring a target depth image containing a target object;
[0043] A determination unit for determining a first depth interval where the target object is located in the target depth image;
[0044] The determination unit is further configured to determine a second depth interval in the target depth image that does not contain the target object;
[0045] A mapping unit for mapping the depth values of each pixel point in the first depth interval to the gray - scale value interval [0, 255];
[0046] The processing unit further sets the depth values of each pixel point in the second depth interval to zero;
[0047] A normalization unit for normalizing each pixel point whose depth value in the first depth interval is mapped to the gray - scale value interval [0, 255], and each pixel point whose depth value in the second depth interval is set to zero, to [0, 1].
[0048] Optionally, the determination unit is specifically configured to:
[0049] Perform key - point detection on the target depth image to obtain multiple key points corresponding to the target object and the depth value of each key point;
[0050] Based on the depth values of each key point, determine the average depth value of the multiple key points;
[0051] The difference between the depth average value and a preset first depth value in a first direction is used as a lower depth limit value, and the sum of the depth average value and a preset second depth value in a second direction is used as an upper depth limit value. The lower depth limit value and the upper depth limit value together form the first depth interval. The first direction is the side close to the depth camera that acquires the target depth image, and the second direction is the side far from the depth camera acquisition side.
[0052] Optionally, the target object is a human face, and the first depth value is less than the second depth value.
[0053] Optionally, the first depth value is in the range of [50 mm, 100 mm], and the second depth value is in the range of [150 mm, 250 mm].
[0054] Optionally, the first depth value is 100 mm, and the second depth value is 200 mm.
[0055] Optionally, the processing unit is further configured to:
[0056] Set each pixel point with a depth value of zero in the second depth interval to 255;
[0057] Perform local histogram equalization processing on each pixel point mapped to the gray value interval [0, 255] in the first depth interval and each pixel point set to 255 in the second depth interval;
[0058] Set each pixel point set to 255 in the second depth interval to zero;
[0059] The normalization unit is specifically configured to:
[0060] Normalize each pixel point mapped to the gray value interval [0, 255] and processed by the local histogram equalization in the first depth interval, and each pixel point with a depth value of zero in the second depth interval to [0, 1].
[0061] Optionally, the mapping unit is specifically configured to:
[0062] For the i-th pixel point in the first depth interval, calculate a first depth difference between the depth value of the i-th pixel point and the lower limit value of the first depth interval;
[0063] Calculate a second depth difference between the upper depth limit value and the lower depth limit value;
[0064] Calculate a depth difference ratio of the first depth difference to the second depth difference;
[0065] Calculate the product of the depth difference ratio and 255 to obtain the gray value corresponding to the i-th pixel point.
[0066] In a fourth aspect, an embodiment of the present application provides a model training device, which includes:
[0067] An acquisition unit, configured to acquire multiple target depth images including a target object;
[0068] A preprocessing unit, configured to preprocess each of the target depth images, and the preprocessing at least includes the normalization processing described in any embodiment of the first aspect;
[0069] A training unit, configured to train a benchmark model with the multiple preprocessed target depth images to obtain a target model.
[0070] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes:
[0071] At least one processor;
[0072] A memory coupled to the processor;
[0073] When the at least one processor executes a computer program stored in the memory, the steps of the method described in any embodiment of the first aspect or the second aspect are implemented.
[0074] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any embodiment of the first aspect are implemented.
[0075] It should be understood that the technical solutions of the third to sixth aspects of the embodiments of the present application are consistent with those of the first and second aspects of the embodiments of the present application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar, and will not be repeated here.
Description of the Drawings
[0076] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this specification, and those of ordinary skill in the art can also obtain other drawings according to these drawings without creative efforts.
[0077] Figure 1 It is a flowchart of a method for normalizing an image provided by an embodiment of the present application;
[0078] Figure 2 It is a schematic diagram of dividing a first depth interval and a second depth interval provided by an embodiment of the present application;
[0079] Figure 3Flow chart of a method for determining a first depth interval where a target object is located provided by an embodiment of the present application;
[0080] Figure 4 Schematic diagram of determining a first depth interval based on depth average provided by an embodiment of the present application;
[0081] Figure 5 Flow chart of a method for normalizing a target depth image to [0, 1] provided by an embodiment of the present application;
[0082] Figure 6 Flow chart of a method for mapping pixel points in a first depth interval to a gray scale interval provided by an embodiment of the present application;
[0083] Figure 7 Flow chart of a model training method provided by an embodiment of the present application;
[0084] Figure 8 Schematic diagram of a structure of a normalization device for an image provided by an embodiment of the present application;
[0085] Figure 9 Schematic diagram of a structure of a model training device provided by an embodiment of the present application;
[0086] Figure 10 Schematic diagram of a structure of an electronic device provided by an embodiment of the present application.
Detailed implementation manners
[0087] To better understand the technical solutions of this specification, the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0088] It should be clear that the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this specification without creative efforts belong to the scope protected by this specification.
[0089] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit this specification. The singular forms "a", "the" and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0090] A depth image refers to an image that uses the distance (depth) from a depth camera to each point in a scene as a pixel value, and it can directly reflect the geometric shape of the visible surface of an object.
[0091] At present, depth images are widely used. For example, the collected depth images can be directly visualized, or used in depth model training. Exemplarily, depth images can be used to train face detection models, face recognition models, live detection models, and 3D face reconstruction, etc. However, due to the limitations of the imaging principle of depth images themselves, the contrast of depth images is often poor and has strong background noise. Therefore, when applying the above-mentioned applications to depth images, it is necessary to perform normalization processing on them.
[0092] It has been found by the inventors of the present application that in related technologies, normalization is often performed on the entire depth image. However, in addition to including some valid information, depth images may also include a large amount of invalid information. Then, there may be a large amount of invalid information in the normalized image, resulting in poor reliability of the depth image in subsequent application processes.
[0093] In view of this, the embodiments of the present application provide a method for normalizing an image. In this method, by determining the depth range of the target object in the target depth image, and then normalizing the pixel values of each pixel point located in the above depth range to [0, 255], while directly setting the pixel values of each pixel point outside the above depth range to zero, so as to eliminate the influence of interference information. On this basis, it is then uniformly normalized to [0, 1], so as to be able to mainly retain the valid information part in the original depth image, that is, the information of the target object in the normalized image, thus ensuring the reliability in subsequent application processes.
[0094] The technical solutions provided by the embodiments of the present application will be introduced below with reference to the accompanying drawings. Please refer to Figure 1 , the embodiments of the present application provide a method for normalizing an image, and the process of this method is described as follows:
[0095] Step 101: Obtain a target depth image containing a target object.
[0096] In the embodiments of the present application, a target depth image containing a target object can be collected by a depth camera. For example, the depth camera can be a TOF camera, an RGBD camera, or a structured light camera, and the type of the depth camera is not particularly limited here. At the same time, the type of the target object can be the whole or part of different objects, people, animals in different scenarios, etc., and the type of the target object is also not particularly limited here.
[0097] Step 102: Determine the first depth range where the target object is located in the target depth image.
[0098] Step 103: Determine the second depth range in the target depth image that does not contain the target object.
[0099] In the embodiments of the present application, it can be considered that the depth camera has a corresponding measured depth, that is, any object within the measured depth will be detected by the depth camera and presented in the formed target depth image. Usually, the depth range where the target object itself is located is less than the measured depth supported by the depth camera. Therefore, there are some useless or even interfering depth ranges. Therefore, in order to more accurately normalize the effective information (i.e., the target object) in the target depth image, it is necessary to effectively distinguish the first depth range where the target object is located in the target depth image and the second depth range that does not include the target object, so as to adopt different processing strategies for the first depth range and the second depth range subsequently.
[0100] For example, please refer to Figure 2 , within the ranging depth range of the depth camera, the foreground part containing the target object can be considered as the first depth range; while the super foreground and background parts that do not contain the target object can be considered as the second depth range.
[0101] Step 104: Map the depth values of the pixel points in the first depth range to the gray value range [0, 255].
[0102] Step 105: Set the depth values of the pixel points in the second depth range to zero.
[0103] In the embodiments of the present application, it can be considered that the pixel points in the first depth range can represent the effective information of the target object. Then, the depth values (i.e., pixel values) of the pixel points in the first depth range can be mapped to the corresponding gray values in the gray value range [0, 255], so that the above effective information can be fully utilized in the subsequent process; while the pixel points in the second depth range can be considered not to provide effective information for the target object, so the depth values (i.e., pixel values) of the pixel points in the second depth range can be directly set to zero, thereby eliminating the interference of the pixel points in the second depth range on the pixel points representing the target object in the first depth range.
[0104] Step 106: Normalize the pixel points whose depth values are mapped to the gray value range [0, 255] in the first depth range and the pixel points whose depth values are set to zero in the second depth range to [0, 1].
[0105] In the embodiments of the present application, the superposition of the first depth interval and the second depth interval can be regarded as the target depth image. Then, after processing the first depth interval in step 103 and the second depth interval in step 104 using different strategies respectively, whether it is each pixel point in the first depth interval mapped to the gray value interval [0, 255], or each pixel point in the second depth interval with the depth value (pixel value) set to zero, the pixel value of each pixel point is divided by 255, and the entire target depth image can be normalized to between [0, 1].
[0106] It should be understood that in the above embodiments, pixel value normalization is mainly performed. Of course, image specification (i.e., image size) and image content normalization can also be performed. Or rather, the target depth image obtained in step 101 can be regarded as a depth image that has been normalized in terms of image specification and image content.
[0107] Please refer to Figure 3 , which is a schematic flowchart of a method for determining the first depth interval where the target object is located in the target depth image provided by the embodiments of the present application. Step 102 can be specifically implemented by executing sub-step 1021 and sub-step 1023:
[0108] Step 1021: Perform key point detection on the target depth image to obtain multiple key points corresponding to the target object and the depth value of each of the key points.
[0109] Step 1022: Based on the depth value of each key point, determine the average depth value of the multiple key points.
[0110] Step 1023: Take the difference between the average depth value and the preset first depth value in the first direction as the depth lower limit value, and take the sum of the average depth value and the preset second depth value in the second direction as the depth upper limit value. The depth lower limit value and the depth upper limit value together constitute the first depth interval. The first direction is the side close to the depth camera that collects the target depth image, and the second direction is the side far from the depth camera that collects the depth image.
[0111] In the embodiments of the present application, the target object can be considered to include multiple key points. Then, by performing key point detection on the target depth image, the target object can be located in the target depth image. Of course, when the types of the target objects are different, the types and quantities of the key points used to locate the target object may also be different. The types and quantities of the key points are not particularly limited here.
[0112] After determining multiple key points from the target depth image, the depth value of each key point among the multiple key points can be obtained. Then, based on the depth value of each key point, the average depth value of the multiple key points is determined, and this average depth value is used as a reference to extend a preset depth respectively towards the side closer to the depth camera and away from the depth camera, obtaining the corresponding lower depth limit value and upper depth limit value, thereby relatively accurately forming the first depth interval where the target object is located.
[0113] It should be understood that in the above embodiments, when determining the average depth value of multiple key points of the target object, a real-time calculation method is adopted. Of course, it is also possible to calculate the average depth value of the key points of different types of target objects through multiple experiments in history, and then establish the corresponding relationship between the target object type and the average depth value. In the actual application process, only the actual target object type of the target object included in the target depth image needs to be identified, and the actual average depth value corresponding to the actual target object type can be obtained based on the corresponding relationship between the target object type and the average depth value.
[0114] For example, please refer to Figure 4 , after obtaining the average depth value of the key points in the target object based on real-time calculation or historical experience, a first depth value can be extended towards the side closer to the depth camera and a second depth value can be extended towards the side away from the depth camera, thereby determining the first depth interval where the target object is located.
[0115] In the embodiments of the present application, when the target object is a human face, when extending a preset depth respectively towards the side closer to the depth camera and away from the depth camera with the average depth value as a reference, since the depth range of the human face is smaller than that of the back of the head, therefore, the first depth value is less than the second depth value, that is, the depth extended towards the side closer to the depth camera is less than the depth extended towards the side away from the depth camera, thereby relatively accurately determining the first depth interval where the human face is located.
[0116] In the embodiments of the present application, when the target object is a human face, the first depth value is [50mm, 100mm], and the second depth value is [150m, 250mm], that is, extend [50mm, 100mm] towards the side closer to the depth camera respectively, and extend [150m, 250mm] away from the depth camera, thereby relatively accurately determining the first depth interval where the human face is located.
[0117] Furthermore, in the embodiments of the present application, when the target object is a human face, the first depth value is 100mm, and the second depth value is 200mm, that is, extend 100mm towards the side closer to the depth camera respectively, and extend 200mm away from the depth camera, retaining as much information of the target object as possible and excluding as much interference information as possible, thereby relatively accurately determining the first depth interval where the human face is located.
[0118] It should be understood that when the target object is a specific object or an animal, the above-mentioned average depth value, the first depth value and the second depth value may change according to actual conditions.
[0119] See also Figure 5 , which is a flow chart of a method for normalizing to [0,1] provided in an embodiment of the present application. Before executing step 106, steps 107-109 may also be executed:
[0120] Step 107: setting the depth value of each pixel in the second depth interval to zero to 255.
[0121] Step 108: Local histogram equalization is performed on all pixels in the first depth interval that are mapped to the grayscale value interval [0, 255] and all pixels in the second depth interval that are set to 255.
[0122] Step 109: Set each pixel point in the second depth interval that is set to 255 to zero.
[0123] Step 106 can be specifically implemented by executing sub-step 110:
[0124] Step 110: normalize all pixels in the first depth interval that are mapped to the grayscale value interval [0, 255] and processed by local histogram equalization, and all pixels in the second depth interval that have depth values set to zero to [0, 1].
[0125] In the embodiment of the present application, local histogram equalization is performed on each pixel in the first depth interval mapped to the grayscale value interval [0,255], thereby increasing the contrast between the pixels representing the effective information. If each pixel whose depth value is set to zero in the second depth interval is set to 255, then even if local histogram equalization is performed on each pixel in the second depth interval, there will be no change in each pixel in the second depth interval, that is, it is equivalent to that each pixel set to 255 in the second depth interval does not participate in the above-mentioned local histogram equalization process, thereby reserving a larger mapping space for each pixel in the grayscale value interval [0,255] in the first depth interval, thereby having a better stretching effect.
[0126] After the histogram equalization process is completed, the pixel value of each pixel in the second depth interval is reset to zero. On this basis, it is normalized to [0,1]. This method can make the normalized image mainly retain the effective information part of the original depth image, while improving the contrast of the effective information part as much as possible, without amplifying the interference information, so as to ensure the reliability in the subsequent application process.
[0127] See also Figure 6, which is a schematic flowchart of the method for mapping the pixel points in the first depth interval to the gray scale interval in the embodiment of the present application. Step 104 can be specifically implemented by executing sub-steps 1041-1044:
[0128] Step 1041: For the i-th pixel point in the first depth interval, calculate the first depth difference between the depth value of the i-th pixel point and the depth lower limit value.
[0129] Step 1042: Calculate the second depth difference between the depth upper limit value and the depth lower limit value.
[0130] Step 1043: Calculate the depth difference ratio of the first depth difference to the second depth difference.
[0131] Step 1044: Calculate the product of the depth difference ratio and 255 to obtain the gray scale value corresponding to the i-th pixel point.
[0132] In the embodiment of the present application, for any pixel point in the first depth interval, first calculate the first depth difference between the depth value of the pixel point and the depth lower limit value of the first depth interval, and the second depth difference between the depth upper limit value and the depth lower limit value of the first depth interval, then calculate the ratio of the above first depth difference to the second depth difference, which can be considered to measure the degree to which the depth of the above pixel point is located in the first depth interval, and finally multiply the ratio by 255 to obtain the gray scale value corresponding to the pixel point. Through the above method, the gray scale value corresponding to each pixel point in the range of the first depth interval can be calculated.
[0133] For example, if the depth value of the i-th pixel point in the first depth interval is 300, the depth upper limit value of the first depth interval is 500, and the depth lower limit value of the first depth interval is 200, then the first depth difference is 300 - 200 = 100, the second depth difference is 500 - 200 = 300, the ratio of the first depth difference to the second depth difference is 100 / 300 = 1 / 3, and finally, the gray scale value mapped by the i-th pixel point is 1 / 3 * 255 = 85.
[0134] Please refer to Figure 7 , which is a schematic flowchart of a model training method provided by the embodiment of the present application. The method process is as follows:
[0135] Step 201: Obtain multiple target depth images including the target object.
[0136] Step 202: Perform preprocessing on each target depth image, and the preprocessing includes at least Figure 1 , Figure 3 , Figures 5 - 6 the normalization processing described in any one of the embodiments of
[0137] Step 203: Train the pre - processed multiple target depth images on a benchmark model to obtain a target model.
[0138] In the embodiments of the present application, first, multiple target depth images containing a target object are collected, and then each of the above - mentioned target depth images is pre - processed. For example, at least the normalization method described in the first aspect is used for normalization processing. Of course, it may also include normalization in terms of image specifications and image content, as well as data augmentation processing, etc. The specific content of the pre - processing is not particularly limited here. Then, the pre - processed target depth images are input into the benchmark model for training, so as to obtain a usable target model. Since each target depth image mainly contains the effective information of the target object, it can be considered that the target model has high performance.
[0139] It should be understood that the above - mentioned target models include, but are not limited to, face detection models, face recognition models, and liveness detection models, etc.
[0140] Please refer to Figure 8 , based on the same inventive concept, the embodiments of the present application also provide a normalization device for images. The device includes: an acquisition unit 301, a determination unit 302, a mapping unit 303, a processing unit 304, and a normalization unit 305.
[0141] The acquisition unit 301 is used to acquire a target depth image containing a target object;
[0142] The determination unit 302 is used to determine the first depth interval where the target object is located in the target depth image;
[0143] The determination unit 302 is further used to determine the second depth interval where the target object is not included in the target depth image;
[0144] The mapping unit 303 is used to map the depth values of each pixel point in the first depth interval to the gray - scale value interval [0, 255];
[0145] The processing unit 304 further sets the depth values of each pixel point in the second depth interval to zero;
[0146] The normalization unit 305 is used to normalize each pixel point whose depth value in the first depth interval is mapped to the gray - scale value interval [0, 255], and each pixel point whose depth value in the second depth interval is set to zero, to [0, 1].
[0147] Optionally, the determination unit 302 is specifically used for:
[0148] Perform key - point detection on the target depth image to obtain multiple key points corresponding to the target object and the depth value of each of the key points;
[0149] Determine the average depth of multiple key points based on the depth values of each key point;
[0150] Take the difference between the average depth and the preset first depth value in the first direction as the lower depth limit value, and take the sum of the average depth and the preset second depth value in the second direction as the upper depth limit value. The lower depth limit value and the upper depth limit value together constitute the first depth interval. The first direction is the side of the depth camera close to the acquisition target depth image, and the second direction is the side away from the acquisition depth camera.
[0151] Optionally, the target object is a human face, and the first depth value is less than the second depth value.
[0152] Optionally, the first depth value is in the range of [50mm, 100mm], and the second depth value is in the range of [150m, 250mm].
[0153] Optionally, the first depth value is 100mm, and the second depth value is 200mm.
[0154] Optionally, the processing unit 304 is further configured to:
[0155] Set each pixel point with a depth value of zero in the second depth interval to 255;
[0156] Perform local histogram equalization processing on each pixel point mapped to the gray value interval [0, 255] in the first depth interval and each pixel point set to 255 in the second depth interval;
[0157] Set each pixel point set to 255 in the second depth interval to zero;
[0158] The normalization unit 305 is specifically configured to:
[0159] Normalize each pixel point mapped to the gray value interval [0, 255] and processed by local histogram equalization in the first depth interval, and each pixel point with a depth value of zero in the second depth interval to [0, 1].
[0160] Optionally, the mapping unit 303 is specifically configured to:
[0161] For the i-th pixel point in the first depth interval, calculate the first depth difference between the depth value of the i-th pixel point and the lower depth limit value;
[0162] Calculate the second depth difference between the upper depth limit value and the lower depth limit value;
[0163] Calculate the depth difference ratio of the first depth difference to the second depth difference;
[0164] Calculate the product of the depth difference ratio and 255 to obtain the gray value corresponding to the i-th pixel point.
[0165] Please refer to Figure 9 , based on the same inventive concept, an embodiment of the present application further provides a model training device, which includes: an acquisition unit 401, a preprocessing unit 402, and a training unit 403.
[0166] The acquisition unit 401 is configured to acquire multiple target depth images including a target object;
[0167] The preprocessing unit 402 is configured to preprocess each target depth image, and the preprocessing at least includes Figure 1 , Figure 3 , Figures 5 - 6 the normalization processing described above;
[0168] The training unit 403 is configured to train a benchmark model with the multiple preprocessed target depth images to obtain a target model.
[0169] Please refer to Figure 10 , based on the same inventive concept, an embodiment of the present application provides an electronic device 100, which includes at least one processor 501. The processor 501 is configured to execute a computer program stored in a memory to implement the steps of the image normalization method provided in the embodiments of the present application as shown in Figure 1 , 3 , 5-6 or Figure 7 shown in the model training method shown in.
[0170] Optionally, the processor 501 may specifically be a central processing unit or a specific ASIC, and may be one or more integrated circuits for controlling program execution.
[0171] Optionally, the electronic device 100 may further include a memory 502 coupled to at least one processor 501. The memory 502 may include a ROM, a RAM, and a disk memory. The memory 502 is used to store data required for the operation of the processor 501, that is, instructions that can be executed by at least one processor 501. The at least one processor 501 executes the instructions stored in the memory 502 to execute the method as shown in Figure 1 , 3 , 5-6 or 7. Among them, the number of the memories 502 is one or more.
[0172] Among them, the entity devices corresponding to the acquisition unit 301, the determination unit 302, the mapping unit 303, the processing unit 304, and the normalization unit 305, or the acquisition unit 401, the preprocessing unit 402, and the training unit 403 may all be the aforementioned processor 501. The electronic device 100 can be used to execute Figure 1 , 3, the method provided by the embodiment shown in 5-6 or 7. Therefore, for the functions that can be realized by each functional module in the electronic device 100, reference may be made to the corresponding descriptions in the embodiments shown in Figure 1 , 3 , 5-6 or 7, and details are not repeated here.
[0173] The embodiment of the present application further provides a computer storage medium. The computer storage medium stores computer instructions. When the computer instructions run on a computer, the computer is caused to execute the method as described in Figure 1 , 3 , 5-6 or 7.
[0174] The above are only the preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the scope of protection of this specification.
Claims
1. A method for normalizing an image, characterized in that, the method comprises: acquiring a target depth image containing a target object; determining a first depth interval where the target object is located in the target depth image; determining a second depth interval in the target depth image that does not contain the target object; mapping the depth values of each pixel point in the first depth interval to the gray value interval [0, 255]; setting the depth values of each pixel point in the second depth interval to zero; normalizing each pixel point whose depth value in the first depth interval is mapped to the gray value interval [0, 255] and each pixel point whose depth value in the second depth interval is set to zero to [0, 1].
2. The method according to claim 1, characterized in that, determining a first depth interval where the target object is located in the target depth image includes: performing key point detection on the target depth image to obtain a plurality of key points corresponding to the target object and the depth value of each key point; determining the average depth value of the plurality of key points based on the depth value of each key point; using the difference between the average depth value and a preset first depth value in the first direction as the depth lower limit value, and using the sum of the average depth value and a preset second depth value in the second direction as the depth upper limit value, where the depth lower limit value and the depth upper limit value together constitute the first depth interval, the first direction is the side close to the depth camera that acquires the target depth image, and the second direction is the side far from the depth camera side.
3. The method according to claim 2, characterized in that, the target object is a human face, and the first depth value is less than the second depth value.
4. The method according to claim 3, characterized in that, the first depth value is [50mm, 100mm], and the second depth value is [150m, 250mm].
5. The method according to claim 4, characterized in that, the first depth value is 100mm, and the second depth value is 200mm.
6. The method according to claim 1, characterized in that, before normalizing each pixel point whose depth value in the first depth interval is mapped to the gray value interval [0, 255] and each pixel point whose depth value in the second depth interval is set to zero to [0, 1], the method further includes: setting each pixel point whose depth value in the second depth interval is set to zero to 255; performing local histogram equalization processing on each pixel point whose depth value in the first depth interval is mapped to the gray value interval [0, 255] and each pixel point whose depth value in the second depth interval is set to 255; setting each pixel point whose depth value in the second depth interval is set to 255 to zero; normalizing each pixel point whose depth value in the first depth interval is mapped to the gray value interval [0, 255] and each pixel point whose depth value in the second depth interval is set to zero to [0, 1], including: Normalize all pixel points in the first depth interval that are mapped to the gray value interval [0, 255] and processed by the local histogram equalization, and all pixel points with the depth value set to zero in the second depth interval to [0, 1].
7. The method according to claim 2, wherein, Mapping the depth values of the pixel points in the first depth interval to the gray value interval [0, 255] includes: For the i-th pixel point in the first depth interval, calculate a first depth difference between the depth value of the i-th pixel point and the lower limit value of the first depth interval; Calculate a second depth difference between the upper limit value and the lower limit value of the first depth interval; Calculate a depth difference ratio of the first depth difference to the second depth difference; Calculate the product of the depth difference ratio and 255 to obtain the gray value corresponding to the i-th pixel point.
8. A model training method, wherein, The method includes: Obtain multiple target depth images including a target object; Perform preprocessing on each of the target depth images, and the preprocessing at least includes the normalization processing according to any one of claims 1-7; Train a benchmark model with the preprocessed multiple target depth images to obtain a target model.
9. An electronic device, wherein, The electronic device includes: At least one processor; A memory coupled to the at least one processor; The at least one processor is configured to implement the steps of the method according to any one of claims 1-7 or 8 when executing a computer program stored in the memory.
10. A computer-readable storage medium, on which a computer program is stored, wherein, The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1-7 or 8.