Image processing method and electronic device
By acquiring frame image sequences in electronic devices and using depth maps and location information to calculate the distance between the subject and the device, the problem of recognition errors in existing technologies is solved, and a more accurate subject recognition sequence is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LENOVO (BEIJING) LTD
- Filing Date
- 2022-07-29
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, electronic devices cannot accurately determine the distance order of objects in the case of multiple objects being photographed based on the size of the detection frame, leading to recognition errors.
By acquiring a sequence of frame images, the depth map of the subject is determined using depth map and location information. The distance between the subject and the device is calculated based on the depth value and weight value of the image region, thereby determining the order of recognition.
It improves the accuracy of identifying the order of multiple captured objects and reduces the impact of detection box size and occlusion factors.
Smart Images

Figure CN115272453B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electronic devices, and relates to but is not limited to an image processing method and an electronic device. BACKGROUND
[0002] In a case where multiple shooting objects face the electronic device, the electronic device detects each shooting object, and performs sequential identification on the shooting objects after detection. Shooting objects closer to the electronic device are preferentially identified, and shooting objects farther away wait until shooting objects closer to the electronic device pass through, and then are identified. In the related art, the distance from the electronic device is determined by distinguishing the size of a shooting object detection box, and the identification sequence is determined. However, the size of the detection box sometimes cannot accurately represent the sequence of the shooting objects, and identification errors may occur. SUMMARY
[0003] The present application provides an image processing method and an electronic device.
[0004] The technical solution of the present application embodiment is implemented as follows:
[0005] The present application provides an image processing method, which comprises: acquiring a frame image sequence captured by an electronic device; each frame image in the frame image sequence comprises at least one shooting object; determining a depth map of a target shooting object based on a depth map of a target frame image and position information of at least one shooting object in the target frame image; determining a weight value of each image region in the depth map of the target shooting object based on a depth value of at least one image region in the depth map of the target shooting object; wherein the weight value is used to represent a reference coefficient of the depth value of each image region to the depth value of the depth map of the target shooting object; and determining a distance between the target shooting object and the electronic device based on the depth value and the weight value of each image region.
[0006] The present application provides an image processing device, which comprises an acquisition module and a determination module, wherein: the acquisition module is configured to acquire a frame image sequence captured by an electronic device; each frame image in the frame image sequence comprises at least one shooting object; and the determination module is configured to determine a depth map of a target shooting object based on a depth map of a target frame image and position information of at least one shooting object in the target frame image; determine a weight value of each image region in the depth map of the target shooting object based on a depth value of at least one image region in the depth map of the target shooting object; wherein the weight value is used to represent a reference coefficient of the depth value of each image region to the depth value of the depth map of the target shooting object; and determine a distance between the target shooting object and the electronic device based on the depth value and the weight value of each image region.
[0007] The embodiment of the present application provides an electronic device, comprising a memory and a processor, the memory stores a computer program which can run on the processor, and the processor implements the steps in the above method when executing the program.
[0008] The embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps in the above method.
[0009] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:
[0010] In the embodiment of the present application, the weight value of each image region is determined based on the depth value of at least one image region in the depth map of the target shooting object; and the distance between the target shooting object and the electronic device is determined based on the depth value and the weight value of each image region. In this way, in the process of determining the identification sequence of multiple shooting objects, the distance between the shooting object and the electronic device is determined based on the depth value of each region of the shooting object, and the identification sequence of the multiple shooting objects is determined based on the distance, thereby improving the accuracy of determining the identification sequence. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical scheme in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0012] Figure 1A An optional architecture schematic diagram of an execution system of the image processing method provided by the embodiment of the present application;
[0013] Figure 1B A flowchart schematic diagram of the image processing method provided by the embodiment of the present application;
[0014] Figure 1C An application scenario schematic diagram of the image processing method provided by the embodiment of the present application;
[0015] Figure 2 A flowchart schematic diagram of multi-target tracking in the image processing method provided by the embodiment of the present application;
[0016] Figure 3 A flowchart schematic diagram of the image processing method provided by the embodiment of the present application;
[0017] Figure 4 A flowchart schematic diagram of the image processing method provided by the embodiment of the present application;
[0018] Figure 5 A schematic diagram of a composition structure of an image processing device provided by an embodiment of the present application is shown in FIG. 1.
[0019] Figure 6 A schematic diagram of a hardware entity of an electronic device provided by an embodiment of the present application is shown in FIG. 2. DETAILED DESCRIPTION
[0020] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application but not all of the embodiments. The following embodiments are used to explain the present application but not to limit the scope of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of the present application.
[0021] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0022] It should be noted that the terms "first", "second", "third" involved in the embodiments of the present application are only to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first", "second", "third" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0023] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which the embodiments of the present application belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have meanings consistent with meanings in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as such.
[0024] The embodiment of the present disclosure provides an image processing method, which can be executed by a processor of an electronic device. Wherein, the electronic device can also be connected with an image acquisition device through a network to acquire a frame image sequence, and the electronic device can be a server, a notebook computer, a tablet computer, a desktop computer, a smart television, a face recognition gate, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated message device, a portable game device) and the like with image processing capability. An application program can be run on the electronic device to interact with the user, and the application program can be understood as a front end, which can be a browser in a browser / server (B / S) mode or a client in a client / server (C / S) structure when implemented.
[0025] Figure 1A An optional architecture schematic diagram of an execution system of the image processing method provided by the embodiment of the present disclosure is shown in Figure 1A The electronic device 100 is connected with the image acquisition device 101 or the server 300 through the network 200. The network 200 can be a wide area network or a local area network, or a combination of the two. The electronic device 100 and the image acquisition device 101 can be physically separate or integrated. The image acquisition device 101 can send or store the acquired frame image sequence to the electronic device 100 through the network 200. The electronic device 100 implements the technical solution of image comparison in the embodiment of the present disclosure.
[0026] In some embodiments, Figure 1A The system shown in the figure can also have no network 200 and image acquisition device 101, and only have the electronic device 100, so that the electronic device acquires the frame image sequence from the local and implements the image processing. In other embodiments, the image acquisition device 101 and the electronic device 100 can be integrated, such as a mobile phone with a camera, or a face recognition gate with a camera, so that the image acquisition device 101 and the electronic device 100 can adopt a bus connection mode instead of a network.
[0027] Based on Figure 1A The architecture schematic diagram shown in the figure, the present application provides an image processing method, Figure 1B A flowchart of an image processing method provided by the embodiment of the present disclosure is shown in Figure 1B The method comprises at least the following steps:
[0028] Step S101, acquiring a frame image sequence photographed by an electronic device; each frame image in the frame image sequence comprises at least one photographed object.
[0029] Here, as shown in Figure 1A The electronic device can be an image acquisition device 101 or an electronic device 100 integrated with the image acquisition device 101. For example, the electronic device 100 is a face recognition gate, and the image acquisition device 101 is a camera that captures a sequence of frame images and transmits the sequence of frame images to the face recognition gate through Bluetooth or a network connection.
[0030] Here, the sequence of frame images can be frame images at consecutive time points or frame images at intervals of a preset number of frames. For example, the frame images can be at intervals of one frame at time points 1, 3, and 5, or the frame images can be at consecutive time points 1, 2, and 3.
[0031] Here, each frame image can be a portrait image including multiple persons or a group image including multiple non-persons. For example, a landscape image including multiple buildings or a group image including multiple animals. For example, the frame image can be a group portrait as shown in Figure 1C
[0032] Here, the photographed object can be a person or a non-person, for example, a person or a natural landscape or a building or an animal. For example, the photographed object can be a person 11 as shown in Figure 1C
[0033] In step S102, a depth map of a target photographed object is determined based on a depth map of a target frame image and position information of at least one photographed object in the target frame image.
[0034] Here, the position information can be position coordinates, an aspect ratio, and a height of the photographed object obtained based on a multi-target tracking method.
[0035] Here, the target frame image can be an image whose resolution and brightness respectively meet a resolution threshold and a brightness threshold, or an image whose at least one of the resolution and the brightness does not meet the resolution threshold or the brightness threshold. In the case where the target frame image is an image whose resolution and brightness respectively meet a resolution threshold and a brightness threshold, the depth map of the target frame image can be obtained by performing depth estimation on the target frame image. Here, the position information of the target photographed object is obtained according to the position information of the at least one photographed object and the target photographed object determined from the at least one photographed object, and the depth map of the target photographed object is obtained from the target frame image according to the position information of the target photographed object. For example, as shown in Figure 1C the target face 11 in the group portrait is selected as the target photographed object, the upper left corner of the group portrait is taken as the coordinate origin, the position coordinates of the face 11 are (5, 5), the aspect ratio is 1, and the height is 1. The depth map of the face 11 is obtained from the depth map of the group portrait based on the above position information of the face 11.
[0036] In an implementable manner, in the application scenario where the target frame image does not meet the definition threshold or the brightness threshold in at least one of definition and brightness, the method further comprises: determining an image of at least one photographed object based on the target frame image in the sequence of frame images and position information of the at least one photographed object in the target frame image; determining at least one image of the photographed object after brightness correction and / or definition correction based on the image of each photographed object and a corresponding template image; wherein each template image comprises at least one photographed object; and determining a depth map of the target frame image based on the image of the target photographed object after brightness correction and / or definition correction. As shown in Figure 1C the image of the face 11 in the dashed box is the image of the face 11. The image of the face 11 is compared with the template image of the face with standard definition and brightness to obtain the image of the face after definition and brightness correction, and the depth of the face 11 is estimated to obtain the depth map of the face 11.
[0037] In an implementable manner, the determination of the at least one image of the photographed object after brightness correction and / or definition correction based on the image of each photographed object and the corresponding template image comprises: performing color space conversion on the image of each photographed object and the corresponding template image to obtain a template image in at least one target color space and an image of the photographed object in at least one target color space; the target color space comprises at least brightness information of the image; and performing brightness correction and / or definition correction on the image of each photographed object based on the brightness of each template image and the brightness information of each image of the photographed object.
[0038] Here, before the color space conversion, the method further comprises: normalizing the size of each image of the photographed object and the corresponding template image. As an example, the image of the face 11 and the corresponding template image are regularly cropped to 128*128. The image of the face 11 and the template image are aligned through the normalization, thereby improving the accuracy of the comparison result.
[0039] Here, the color space can include Red Green Blue (RGB), and can also include Luminance Chrominance (YUV). Illustratively, according to the lightness contrast of the V channel, it is determined whether the face 11 is a black face. If so, gamma correction is performed, the gamma curve of the face 11 image is edited, the face 11 image is edited in a non-linear color tone, the dark and light parts in the face 11 image are detected, and the proportion of the two is increased, thereby improving the contrast of the face 11 image. After the gamma correction, depth estimation is performed to obtain a depth map, and it is determined whether to blur. If so, deblurring processing is performed to realize sharpness correction.
[0040] In step S103, a weight value of each image region is determined based on a depth value of at least one image region in the depth map of the target photographic object. The weight value is used to represent a reference coefficient of the depth value of each image region to the depth value of the depth map of the target photographic object.
[0041] Here, in the case where the image of the target photographic object is a face image, the image region can be divided based on the following steps: based on the image of the target photographic object, a 3D face is obtained by using a 3D face key point model to obtain 3D face key point information; the 3D face is divided into at least one image region based on the 3D face key point information; for example, the face image is divided into forehead, nose, left and right cheeks, and mouth region. The depth map of the target photographic object is divided into at least one image region based on the region divided based on the 3D face key point information.
[0042] In an implementable manner, the step S103 of determining the weight value of each image region based on the depth value of at least one image region in the depth map of the target photographic object includes: step S1031 of determining a difference value between the depth value of each image region and the depth value of a target image region to obtain a proportion parameter between each difference value; step S1032 of determining a change amount of the weight value of each image region based on the proportion parameter between each difference value; and step S1033 of determining the weight value of each image region based on the change amount of each weight value and the corresponding initial weight value.
[0043] Here, before the step of determining the difference value between the depth value of each image region and the depth value of a target image region to obtain a proportion parameter between each difference value, the method further includes: setting an initial weight value for each image region, and illustratively, setting different initial weight values for these regions as 0.2, 0.4, 0.05, 0.05, and 0.3, respectively.
[0044] Here, the target image region can be any of the at least one image region, which is not limited here. Exemplarily, in the case that the at least one image region is different regions of the forehead, nose, left and right cheeks, and mouth of a face, and the target image region is the nose region, the difference between the depth value of the nose region and 80% (percentage) of the depth value of the forehead region, nose region, left and right cheek region, and mouth region is: a, 0, b, c, d, and the difference ratio is a:0:b:c:d; using the ratio parameter of the difference ratio, the total of the split weight value 1 is obtained to obtain the change amount of the weight value of each face region, for example, w1, w2, w3, w4, and w5. The change amount of the weight value is summed with the initial weight value of the corresponding region to obtain the changed weight value of each face region.
[0045] Here, the reference coefficient can be the contribution rate, i.e., the contribution rate of the depth value of each image region to the depth value of the depth map of the target shooting object, for example, the reference coefficient of the depth value of the nose region to the depth value of the depth map of the target shooting object is 0.15, i.e., the contribution rate is 15%, and after multiplying the depth value of the nose region by 15%, the weighted depth value of the other regions is added to obtain the depth value of the depth map of the target shooting object.
[0046] Step S104, determining the distance between the target shooting object and the electronic device based on the depth value and the weight value of each image region.
[0047] Here, the depth value of the target shooting object is the distance between the target shooting object and the electronic device, and the distances at different time points form a dynamic distance sequence, and the distances of multiple shooting objects at different time points form multiple dynamic distance sequences.
[0048] In an implementable manner, the step S104 of determining the distance between the target shooting object and the electronic device based on the depth value and the weight value of each image region comprises: step S1041, processing the depth value of each image region based on the weight value of each image region to obtain a weighted depth value of each image region; and step S1042, fusing each weighted depth value to obtain the distance between the target shooting object and the electronic device.
[0049] Here, the weight value of each image region can be multiplied by the depth value of the corresponding image region to obtain a weighted depth value of each image region, and the weighted depth values of each image region are added to obtain a weighted depth value of the target shooting object. Exemplarily, the depth value of the depth map represents the actual distance between the camera and the shooting object. In the case that the target shooting object is a face and the camera shoots the face, the weighted depth value of the face is used as the actual distance between the face and the camera.
[0050] In the above embodiment, the weight value of each image region is determined based on the depth value of at least one image region in the depth map of the target shooting object, and the distance between the target shooting object and the electronic device is determined based on the depth value and the weight value of each image region. In this way, in the process of determining the recognition sequence of multiple shooting objects, the distance between the shooting object and the electronic device is determined based on the depth value of each region of the shooting object, the recognition sequence of multiple shooting objects is determined based on the distance, the accuracy of determining the recognition sequence is improved, and the process of determining the recognition sequence is not affected by the shooting object detection, the size of the shooting object, and the like.
[0051] The image processing method provided in the embodiment of the present application comprises at least the following steps:
[0052] In step S201, a frame image sequence captured by an electronic device is obtained, and each frame image in the frame image sequence comprises at least one shooting object.
[0053] In step S202, a detection box of each shooting object in the target frame image is determined.
[0054] Here, the detection box of the shooting object in the image can be obtained through image detection. For example, as shown in the detection box 11. Figure 1C
[0055] In step S203, a relationship matrix between the target frame image and the previous frame image is determined to obtain an initialized tracking chain.
[0056] Here, the image information of the target frame image can be described through a matrix, and the relationship matrix is determined as the motion trajectory of each shooting object in the target frame image, i.e., the initialized tracking chain.
[0057] In an implementable manner, the step S203 of determining the relationship matrix between the target frame image and the previous frame image to obtain the initialized tracking chain comprises:
[0058] In step S2031, the occluded state of at least one shooting object in the target frame image is determined based on the relationship matrix.
[0059] For example, in the case where the relationship matrix is less than a threshold value, it is determined that at least one person in the portrait is in an occluded state.
[0060] In step S2032, in the case where the occluded state is not occluded, the width-height ratio information and the height information in the relationship matrix are updated to three-dimensional information to obtain a three-dimensional relationship matrix.
[0061] Here, the aspect ratio information and the height information in the relationship matrix can be updated into three-dimensional information by means of Kalman filter updating, to obtain a three-dimensional relationship matrix. Here, the Kalman filter updating can realize tracking of solving occlusion by means of depth information in a three-dimensional perspective. Exemplarily, the position coordinate of a center in a two-dimensional perspective of any photographed object is (x, y), the aspect ratio information is a, and the height information is h. A constant velocity model of Kalman filter with Gaussian noise in a two-dimensional perspective is shown in formula (2-1):
[0062]
[0063] wherein x t is the horizontal coordinate of the center position of the photographed object in the target frame image, x t-1 is the horizontal coordinate of the center position of the photographed object in the previous frame image, is the estimated horizontal coordinate of the center position of the photographed object in the previous frame image, and ε x is Gaussian noise.
[0064] The mapping relationship between the position coordinate of the center in the two-dimensional perspective and the position coordinate of the center in the three-dimensional perspective is shown in formula (2-2):
[0065]
[0066] wherein the position coordinate of the center of the photographed object in the three-dimensional perspective is (X t , Y t , Z t ), and f is the focal length of the camera for photographing the photographed object.
[0067] The constant velocity model of Kalman filter with Gaussian noise in a three-dimensional perspective is shown in formula (2-3):
[0068]
[0069] wherein x t is the horizontal coordinate of the center position of the photographed object in the target frame image, x t-1 is the horizontal coordinate of the center position of the photographed object in the previous frame image, is the estimated horizontal coordinate of the center position of the photographed object in the previous frame image, and ε X is Gaussian noise.
[0070] The velocity model of the relationship between the two-dimensional and the three-dimensional can be obtained by combining formula (2-1), (2-2), and (2-3), as shown in formula (2-4):
[0071]
[0072] Here, the aspect ratio information and the height information in the relationship matrix are updated into three-dimensional information based on Kalman filtering, which can be achieved by using formula (2-4); and The conversion of the two-dimensional perspective information and the three-dimensional perspective information is obtained, and the tracking problem in the occlusion scene is solved.
[0073] Exemplarily, the relationship matrix is represented as a wrap matrix; the aspect ratio information in the three-dimensional perspective is identified as A t ; and the height information is represented as H. The three-dimensional relationship matrix can be obtained by using the following pseudo code:
[0074] The prediction function ():
[0075] The wrap matrix between the current frame image and the past frame image is found
[0076] For all active tracking chains do:
[0077] Wrap the state of the current tracking chain with the wrap matrix:
[0078] if the tracking is occluded, it is assumed that there is no velocity model update for a and h
[0079] Else, there is a velocity model update for A t and H
[0080] The wrap state is updated by Kalman filtering KF
[0081] In step S2033, based on the depth information in the three-dimensional relationship matrix and the predicted depth information, the occluded tracking and the occluded tracking are determined, and the initialized tracking chain is obtained.
[0082] Here, an occlusion coefficient a sup can be set, and in the case that the predicted depth information Z f is less than the depth information Z v in the three-dimensional relationship matrix, it is determined that the actual depth is greater than or equal to the predicted depth, and there is no occlusion. In the case that the predicted depth information Z f is greater than the depth information Z v in the three-dimensional relationship matrix, it is determined that the actual depth is less than the predicted depth, and there is an object occlusion.
[0083] In step S204, based on the active tracking in the initialized tracking chain and each detection frame, the first tracking matched with the detection frame, the tracking not matched with the detection frame, and the detection frame not matched with the tracking in the initialized tracking chain are determined.
[0084] Here, the active tracking is the tracking to be matched in the initialized tracking chain.
[0085] Here, steps S204 to S206 are a multi-target tracking method, Figure 2 A flowchart of a multi-target tracking process in an image processing method provided by an embodiment of the present application is shown in Figure 2 As shown, the activated tracking T1 in the initialized tracking chain T and each detection frame D can be matched by a matching function, where the matching function can be realized by comparing the appearance information (appearance information) in the correlation matrix to obtain the first tracking X matched with the detection frame, the tracking T2 not matched with the detection frame, and the detection frame Z not matched with the tracking.
[0086] In step S205, the second tracking matched with the detection frame in the tracking chain is determined based on the tracking not matched with the detection frame in the tracking chain and the detection frame not matched with the tracking.
[0087] Exemplarily, as shown in Figure 2 The matching of the tracking T2 not matched with the detection frame and the detection frame Z not matched with the tracking can be realized by IOU matching.
[0088] Here, in the process of determining the second tracking matched with the detection frame in the tracking chain based on the tracking not matched with the detection frame in the tracking chain and the detection frame not matched with the tracking, the tracking T3 matched with the detection frame can be generated. The occluded tracking and the non-occluded tracking in T3 can be determined according to the visibility condition C1 to obtain the non-occluded (visible) tracking chain Y1 and the occluded tracking chain Y2. For the occluded tracking chain Y2, it can be determined whether to delete the occluded tracking according to the deletion condition C2.
[0089] Exemplarily, for the two times of matching in steps S204 and S205, the following pseudo code can be used for implementation, where the predicted depth information Z f , the depth information Z v in the three-dimensional correlation matrix, the occlusion coefficient a sup , the deletion coefficient a d :
[0090] Matching function ():
[0091] Predicted depth Z f and horizontal depth Z v
[0092] X is the tracking matched with the detection
[0093] Y is the tracking not matched
[0094] Z is the detection not matched
[0095] If Z f <a sup *Z v, indicates that this track is an unoccluded track
[0096] Else, this track is an occluded track
[0097] Match undefined visible tracks T2 with undefined detection boxes based on IOU
[0098] Match detection boxes with active tracks bidirectionally based on appearance information
[0099] Divide Y into visible Y1 track chain and occluded Y2 track chain by whether occluded or not
[0100] For all tracks in Y2 do:
[0101] If Z f <a d *Z v , delete track
[0102] Output X, Y1, Y2, Z
[0103] In step S206, the position information of the at least one photographed object in the target frame image is determined based on the first track and the second track.
[0104] Here, the position information of the photographed object in the target frame image can be updated by Kalman filtering of the first track and the second track, so as to realize prediction of the motion trajectory of the photographed object in the target frame image by the previous frame image. Here, for the track Y1 which is not matched but not occluded, the number of consecutive matching failures can be increased, so as to realize multiple matching of Y1 and increase the probability of successful matching.
[0105] Exemplarily, the above process can be implemented by the following pseudo code:
[0106] Update function():
[0107] X, Y1, Y2, Z = match function();
[0108] Update all tracks on matches in X with Kalman filter KF
[0109] Increase the number of consecutive matching failures ages of all Y1:
[0110] Add Y2 to the occluded track data set
[0111] In step S207, the depth map of the target photographed object is determined based on the depth map of the target frame image and the position information of the at least one photographed object in the target frame image.
[0112] Step S208, determining a weight value of each image region based on the depth value of at least one image region in the depth map of the target shooting object; wherein the weight value is used to represent a reference coefficient of the depth value of each image region to the depth value of the depth map of the target shooting object.
[0113] Step S209, determining the distance between the target shooting object and the electronic device based on the depth value and the weight value of each image region.
[0114] In the above embodiment, by updating the aspect ratio information and the height information in the relationship matrix as three-dimensional information in the case where the occluded state is unoccluded, a three-dimensional relationship matrix is obtained, so that the tracking problem of occlusion can be solved by using three-dimensional perspective, and the tracking of the shooting object in the case of occlusion is improved, and the robustness of the process of predicting the position information of the shooting object is improved.
[0115] Figure 3 A flowchart of an image processing method provided by an embodiment of the present application is shown in FIG. 3. Figure 3 The method includes the following steps.
[0116] Step S301, obtaining a frame image sequence shot by an electronic device; each frame image in the frame image sequence includes at least one shooting object.
[0117] Step S302, determining a depth map of a target shooting object based on a depth map of a target frame image and position information of at least one shooting object in the target frame image.
[0118] Step S304, determining a weight value of each image region based on the depth value of at least one image region in the depth map of the target shooting object; wherein the weight value is used to represent a reference coefficient of the depth value of each image region to the depth value of the depth map of the target shooting object.
[0119] In an implementable manner, before the step S304, the method further includes:
[0120] Step S303, determining an average depth value and a variance threshold value of each image region; in the case where the variance value of any depth value in each image region to the average depth value is greater than the variance threshold value, performing smoothing processing on the depth value based on the depth value of a neighboring region of the image region.
[0121] Herein, each of the image regions can be an image based on a target photographed object, a 3D face is obtained by using a 3D face key point model, and 3D face key point information is obtained; the 3D face is divided into at least one image region based on the 3D face key point information; and the depth map of the target photographed object is divided into different image regions after the depth map of the target photographed object is divided into at least one image region based on the regions divided based on the 3D face key point information.
[0122] Exemplarily, the image region is a face region, an average value and a variance value are calculated on the face depth map estimated by depth on the regions divided based on the face, a variance threshold value is set, and abnormality repair is performed in units of regions. If the variance between the depth value and the average value is greater than the variance threshold value, the region is smoothed according to the depth value of the adjacent region.
[0123] In an implementable manner, after the depth value is smoothed based on the depth value of the adjacent region of the image region, the method further comprises: determining the height of the target photographed object based on the image region after the smoothing; and adjusting a parameter for recognizing the target photographed object based on the height of the target photographed object and the distance between the target photographed object and the electronic device.
[0124] Herein, the parameter for recognizing the target photographed object can be a shooting pitch angle.
[0125] Herein, in a case where the electronic device is a gate camera and the target photographed object is a target person, the parameters such as the angle of the gate camera and the focal length are known, the height of the target person can be determined according to the angle of the gate camera, the focal length, and the vertical distance between the horizontal plane of the camera and the ground. The pitch angle of the gate camera is adjusted according to the height of the target person, and the accuracy of the recognition of the target person is increased.
[0126] In step S305, the distance between the target photographed object and the electronic device is determined based on the depth value and the weight value of each of the image regions.
[0127] In the above implementation process, on the one hand, the depth value is smoothed based on the depth value of the adjacent region of the image region, so that the abnormal depth value in the image can be repaired, thereby correcting the abnormal depth value, obtaining more accurate weight value of each image region, and improving the accuracy of determining the distance between the target photographed object and the electronic device based on the depth value and the weight value of each of the image regions. On the other hand, the parameter for recognizing the target photographed object is adjusted based on the height of the target photographed object and the distance between the target photographed object and the electronic device, so that the pitch angle of the electronic device can be adjusted according to the height of the target person, and the accuracy of the recognition of the target person is increased.
[0128] When multiple people (photographed objects) walk towards the gate machine (electronic device) at the same time and in the same direction, the gate machine detects each person, and after detection, performs face recognition on the photographed objects in sequence. The person closer to the gate machine is given priority in face recognition, and the person farther away waits until the person closer to the gate machine passes through the gate machine, and then the person farther away performs face recognition and passes through the gate machine. When multiple people walk towards the gate machine at the same time, the photographed objects may block each other, affecting face detection and thus affecting face recognition of the gate machine. In related technologies, face detection is achieved by two methods: 1) distinguishing the distance of a person from the gate machine according to the size of the face detection box; and 2) predicting the position information of the target photographed object in a short time by some offline image tracking method. In the offline tracking algorithm, each frame of the detected target (target photographed object) is used as a node of a map, and motion, appearance, and other reference quantities are used as the cost or weight value (similarity measure) of the edge between the nodes, to construct a global graph structure, and to solve the shortest path, minimum cost flow, and other optimal results.
[0129] However, the above process has the following problems: 1) First, because the areas of different faces are of different sizes, a person farther away from the gate machine may be identified first because the area of the face is larger than that of a person closer to the gate machine. Second, the above method relies on the accuracy of the face detection algorithm for face detection positioning. 2) If multiple people walk towards the gate machine and there is a long period of mutual blocking, the above method is inefficient in positioning the position information of the target photographed object.
[0130] To solve the above problems, the present application provides an image processing method, which combines the depth map of the target frame image obtained by depth estimation, the position information of multiple photographed objects obtained by multi-face (target) tracking, and at least one image region of the depth map of the target photographed object obtained by 3d (three-dimensional) face key point alignment, to dynamically estimate the relative distance between the multiple photographed objects and the gate machine, determine the photographed object closer to the gate machine, and determine the order of identification.
[0131] The present application provides an image processing method, as described above, Figure 4 The method comprises:
[0132] Step S401: acquiring each frame image sequence by a front-end camera (image acquisition device);
[0133] Step S402: acquiring single or multiple face position information by target detection and multi-target tracking on the target frame image in the frame image sequence;
[0134] Step S403: performing brightness correction and / or sharpness correction on the target frame image;
[0135] Here, the face image is cropped from the original image (target frame image) and resized to 128*128, and the template face image is converted from RGB to YUV. If the illumination intensity of the V channel is low, the face is determined to be black. If so, gamma correction is performed, and then depth estimation is performed to obtain a depth map. If the depth map is blurred, deblurring is performed.
[0136] In step S404, depth estimation is performed on the target frame image after the brightness correction and / or the sharpness correction, to obtain a depth map.
[0137] Here, the original image containing the face is subjected to depth estimation to obtain a depth map. The position information of multiple faces is obtained based on a tracking method based on depth information, so that the depth map of multiple faces is obtained.
[0138] In step S405, the face depth map is divided into face regions.
[0139] A 3D face is generated from a 2D face based on a 3D face key point model, and 3D face key point information is obtained. Next, the face is divided into regions based on the 3D face key point information, including forehead, nose, left and right cheeks, and mouth regions.
[0140] In step S406, based on at least one face region, the face depth map is repaired for abnormalities.
[0141] For example, the image region is a face region, and the average value and the variance value are calculated on the face depth map obtained by depth estimation based on the regions divided from the face. A variance threshold is set, and the abnormality is repaired in units of regions. If the variance between the depth value and the average value is greater than the variance threshold, the region is smoothed based on the depth value of the adjacent region.
[0142] In step S407, based on the depth map after the abnormality repair, the weight values of the face regions are determined.
[0143] Here, the regions are set with different initial weights, which are 0.2, 0.4, 0.05, 0.05, and 0.3, respectively. The 80% of the depth values of the face regions are taken to obtain the corresponding ratios, and the gain of the weight is determined according to the different values of the difference ratios.
[0144] In step S408, based on the depth value and the weight value of each image region, the distance between the target shooting object and the electronic device is determined, and multiple dynamic distance sequences are obtained.
[0145] In an implementable manner, in a case where the electronic device is a gate camera and the target photographing object is a target person, the method further comprises: determining a height of the target person according to a gate camera angle, a focal length, and a vertical distance between a horizontal plane of the camera and the ground (step S409); and adjusting a pitch angle of the gate camera according to the height of the target person to increase the accuracy of target task recognition (step S410).
[0146] In an implementable manner, in a case where the position coordinates of the two-dimensional center of any one photographing object are (x, y), the aspect ratio information is a, the height information is h, the constant speed model of Kalman filtering with Gaussian noise in a two-dimensional perspective is as shown in formula (2-1), the mapping relationship between the position coordinates of the two-dimensional center and the position coordinates of the three-dimensional center is as shown in formula (2-2), the constant speed model of Kalman filtering with Gaussian noise in a three-dimensional perspective is as shown in formula (2-3), and the speed model of the relationship between the two-dimensional and the three-dimensional can be obtained by combining formulas (2-1), (2-2), and (2-3), as shown in formula (2-4), where the aspect ratio information and the height information in the relationship matrix are updated to three-dimensional information based on Kalman filtering, and formula (2-4) can be used to achieve the relationship. The conversion of the two-dimensional perspective information and the three-dimensional perspective information is obtained, and the tracking problem in a blocking scene is realized. The optimal matching between two adjacent frames can be solved by using Kalman filtering in real-time tracking (online). The real-time tracking is realized by a multi-target tracking method, where the multi-target tracking method is realized by the following pseudo code.
[0147] Prediction function ():
[0148] Finding the wrap matrix between the current frame image and the past frame image
[0149] For all active tracking chains do:
[0150] Wrap the state of the current tracking chain with the wrap matrix:
[0151] if the tracking is blocked, assuming no velocity model update for a and h
[0152] Else, there is A t and H velocity model update
[0153] Update the wrap state by Kalman filter KF
[0154] Update function ():
[0155] X, Y1, Y2, Z = matching function ();
[0156] Update the tracking on all matches in X with Kalman filter KF update step
[0157] ages: the number of consecutive missed matches for all Y1
[0158] Y2 is added to the occluded tracklet data set
[0159] Match function ():
[0160] Predicted depth Z f and horizontal depth Z v
[0161] X is a matched tracklet
[0162] Y is an unmatched tracklet
[0163] Z is an unmatched detection
[0164] If Z f <a sup *Z v , this tracklet is an unoccluded tracklet
[0165] Else, this tracklet is an occluded tracklet
[0166] Match undefined visible tracklets T2 and undefined detection boxes based on IOU
[0167] Match detection boxes and active tracklets bidirectionally based on appearance information
[0168] Divide Y into visible Y1 tracklets and Y2 occluded tracklets by whether they are occluded
[0169] For all tracklets in Y2 do:
[0170] If Z f <a d *Z v , delete the tracklet
[0171] Output X, Y1, Y2, Z
[0172] def main function():
[0173] for each frame do:
[0174] Predict function():
[0175] Update function()
[0176] Output, tracklets
[0177] Wherein, the input of the multi-target tracking method is a detection box D of a current frame image.
[0178] In the above embodiments, on the one hand, in determining the order of face recognition for multiple people, the distance between the subject and the electronic device is determined based on the depth value of each region of the subject's face. Based on this distance, the order of face recognition is determined, improving the accuracy of determining the recognition order and making it unaffected by factors such as face detection and face size. On the other hand, in cases of occlusion, the position information of the subject is obtained based on face tracking and localization, improving the robustness of predicting the subject's position. Combining these two aspects improves tracking performance when multiple subjects are mutually occluded, thereby improving the efficiency of face recognition when multiple subjects pass through the gate, and ultimately improving passage efficiency in multi-person scenarios.
[0179] Based on the foregoing embodiments, this application provides another image processing device, which includes the modules included, and can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0180] Figure 5 This is a schematic diagram of the composition structure of an image processing device provided in an embodiment of this application, as shown below. Figure 5 As shown, the device 500 includes an acquisition module 501 and a determination module 502, wherein:
[0181] The acquisition module 501 is used to acquire a sequence of frame images captured by an electronic device; each frame image in the sequence includes at least one subject being captured; the determination module 502 is used to determine a depth map of a target subject being captured based on a depth map of a target frame image and the position information of at least one subject being captured in the target frame image; determine a weight value for each image region based on the depth value of at least one image region in the depth map of the target subject being captured; wherein the weight value is used to represent a reference coefficient of the depth value of each image region to the depth value of the depth map of the target subject being captured; and determine the distance between the target subject being captured and the electronic device based on the depth value and weight value of each image region.
[0182] In some possible embodiments, the determining module 502 is further configured to: determine a detection box of each shooting object in the target frame image; determine a relationship matrix between the target frame image and a previous frame image, to obtain an initialized tracking chain; determine, based on an activated tracking in the initialized tracking chain and each detection box, a first tracking matched with a detection box, a tracking not matched with a detection box, and a detection box not matched with a tracking in the tracking chain; determine, based on the tracking not matched with a detection box and the detection box not matched with a tracking in the tracking chain, a second tracking matched with a detection box in the tracking chain; and determine, based on the first tracking and the second tracking, position information of at least one shooting object in the target frame image.
[0183] In some possible embodiments, the determining module 502 is further configured to: determine, based on the relationship matrix, an occlusion state of at least one shooting object in the target frame image; update, in a case where the occlusion state is not occluded, width-height ratio information and height information in the relationship matrix as three-dimensional information, to obtain a three-dimensional relationship matrix; and determine, based on depth information in the three-dimensional relationship matrix and predicted depth information, a tracking not occluded and a tracking occluded, to obtain the initialized tracking chain.
[0184] In some possible embodiments, the determining module 502 is further configured to: determine a difference value between a depth value of each image region and a depth value of a target image region, to obtain a proportion parameter between each difference value; determine, based on the proportion parameter between each difference value, a change amount of a weight value of each image region; and determine, based on the change amount of each weight value and a corresponding initial weight value, the weight value of each image region.
[0185] In some possible embodiments, the determining module 502 is further configured to: process, based on the weight value of each image region, the depth value of the image region, to obtain a weighted depth value of each image region; and fuse each weighted depth value, to obtain a distance between the target shooting object and the electronic device.
[0186] In some possible embodiments, the determining module 502 is further configured to: determine, based on a target frame image in the sequence of frame images and position information of at least one shooting object in the target frame image, an image of at least one shooting object; determine, based on each image of the shooting object and a corresponding template image, at least one image of the shooting object after brightness correction and / or definition correction, wherein each template image includes at least one shooting object; and determine, based on the image of the target shooting object after brightness correction and / or definition correction, a depth map of the target frame image.
[0187] In some possible embodiments, the determining module 502 is further configured to: perform color space conversion on each of the photographed images and the corresponding template images to obtain template images in at least one target color space and photographed images in at least one target color space; the target color space includes at least luminance information of the images; and perform luminance correction and / or sharpness correction on each of the photographed images based on the luminance of each of the template images and the luminance information of each of the photographed images.
[0188] In some possible embodiments, the determining module 502 is further configured to: determine an average depth value and a variance threshold value of each of the image regions; and in a case where a variance value of any depth value in each of the image regions and the average depth value is greater than the variance threshold value, perform smoothing processing on the depth value based on depth values of adjacent regions of the image region.
[0189] In some possible embodiments, the determining module 502 is further configured to: determine a height of the target photographed object based on the image region after the smoothing processing; and adjust a parameter for recognizing the target photographed object based on the height of the target photographed object and a distance between the target photographed object and the electronic device.
[0190] It should be noted that the above description of the device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application.
[0191] It should be noted that in the embodiments of the present application, if the above image processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a plurality of instructions for causing an electronic device (which can be a smart phone, a tablet computer, etc. with a camera) to execute all or part of the method described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read Only Memory, ROM), a magnetic disk or an optical disk, and various media that can store program codes. Thus, the embodiments of the present application are not limited to any specific hardware and software combination.
[0192] Correspondingly, the embodiments of the present application provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps of the image processing method in any of the above embodiments.
[0193] Correspondingly, in this embodiment of the application, a chip is also provided, the chip including programmable logic circuits and / or program instructions, which, when the chip is running, are used to implement the steps in any of the image processing methods described in the above embodiments.
[0194] Correspondingly, in this embodiment of the application, a computer program product is also provided, which, when executed by the processor of an electronic device, is used to implement the steps in any of the image processing methods described in the above embodiments.
[0195] Based on the same technical concept, this application provides an electronic device for implementing the image processing method described in the above method embodiments. Figure 6 This is a hardware entity diagram of an electronic device provided in an embodiment of this application, such as... Figure 6 As shown, the electronic device 600 includes a memory 610 and a processor 620. The memory 610 stores a computer program that can run on the processor 620. When the processor 620 executes the program, it implements the steps in any of the image processing methods described in the embodiments of this application.
[0196] The memory 610 is configured to store instructions and applications executable by the processor 620, and can also cache data to be processed or already processed by the processor 620 and various modules in the electronic device (e.g., image data, audio data, voice communication data and video communication data), which can be implemented by flash memory or random access memory (RAM).
[0197] When the processor 620 executes the program, it implements the steps of any of the above-mentioned image processing methods. The processor 620 typically controls the overall operation of the electronic device 600.
[0198] The aforementioned processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that other electronic devices can also implement the functions of the aforementioned processor, and this application does not specifically limit the specific implementation.
[0199] The computer storage medium / memory can be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Ferromagnetic Random Access Memory (FRAM), a Flash Memory, a magnetic surface storage, an optical disc, a Compact Disc Read-Only Memory (CD-ROM), or the like memory; or can be various electronic devices including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, and the like.
[0200] It is noted that the above description of the storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the description of the method embodiments of the present application.
[0201] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that the size of the sequence number of each process in various embodiments of the present application does not mean the execution order, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The sequence number of the above embodiments of the present application is only for description, not representing the advantages and disadvantages of the embodiments.
[0202] It should be noted that, in the present document, the terms "comprising", "containing", or any other similar term are intended to encompass non-exclusive inclusions, such that a process, method, article, or apparatus that comprises a list of elements does not necessarily include those elements only, but can include other elements not expressly listed, or can include elements inherent in such process, method, article, or apparatus. Without more limitations, an element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0203] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative, for example, the division of the units is only a logical functional division, and actual implementation can have another division manner, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed components can be through some interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0204] The units described above as separate components can or can not be physically separated, and the components shown as units can or can not be physical units; they can be located in one place or distributed on multiple network units; and part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0205] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.
[0206] Alternatively, the integrated unit of the present application, if implemented in the form of a software function module and sold or used as an independent product, can also be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for causing an apparatus to perform all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: mobile storage devices, ROM, magnetic or optical disks, and various other media that can store program codes.
[0207] The methods disclosed in the several method embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments.
[0208] The features disclosed in the several method or device embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.
[0209] The above merely provides an implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An image processing method, the method comprising: obtaining a sequence of frame images captured by an electronic device; each frame image in the sequence of frame images comprising at least one captured object; determining a bounding box of each captured object in a target frame image, and determining an initialized tracking chain according to adjacent image changes between the target frame image and a previous frame image; the initialized tracking chain being a motion trajectory of a captured object; determining a first tracking and a second tracking according to the bounding boxes and the initialized tracking chain, wherein the first tracking is a tracking in the tracking chain that matches a bounding box, and the second tracking is a tracking that matches a bounding box determined according to a tracking that does not match a bounding box in the tracking chain and a bounding box that does not match a tracking in the tracking chain; determining position information of at least one captured object in the target frame image based on the first tracking and the second tracking; determining a depth map of a target captured object based on a depth map of the target frame image and the position information of at least one captured object in the target frame image; determining a weight value of each image region in the depth map of the target captured object based on a depth value of the image region, wherein the weight value is used to represent a reference coefficient of the depth value of each image region to a depth value of the depth map of the target captured object; and determining a distance between the target captured object and the electronic device based on the depth value and the weight value of each image region.
2. The method of claim 1, wherein determining an initialized tracking chain according to adjacent image changes between the target frame image and a previous frame image comprises: determining a relationship matrix of changes in image information between the target frame image and the previous frame image, and determining a motion trajectory of each captured object in the target frame image according to the relationship matrix, and determining the motion trajectory of each captured object as an initialized tracking chain of each captured object; wherein determining a first tracking and a second tracking according to the bounding boxes and the initialized tracking chain comprises: determining a first tracking that matches a bounding box in the tracking chain, a tracking that does not match a bounding box, and a bounding box that does not match a tracking in the tracking chain based on an activated tracking in the initialized tracking chain and each bounding box; determining a second tracking that matches a bounding box in the tracking chain based on the tracking that does not match a bounding box and the bounding box that does not match a tracking in the tracking chain; and wherein the activated tracking is a tracking to be matched in the initialized tracking chain.
3. The method of claim 2, wherein determining a relationship matrix between the target frame image and a previous frame image to obtain an initialized tracking chain comprises: determining an occlusion state of at least one captured object in the target frame image based on the relationship matrix; updating width-height ratio information and height information in the relationship matrix to three-dimensional information to obtain a three-dimensional relationship matrix in a case where the occlusion state is unoccluded; determining an unoccluded tracking and an occluded tracking based on depth information in the three-dimensional relationship matrix and predicted depth information to obtain the initialized tracking chain. 4. The method of claim 1, wherein the determining the weight value of each image region based on the depth value of at least one image region in the depth map of the target subject comprises: determining a difference between the depth value of each image region and the depth value of a target image region, to obtain a ratio parameter between each difference; determining a variation of the weight value of each image region based on the ratio parameter between each difference; and determining the weight value of each image region based on the variation of the weight value and a corresponding initial weight value.
5. The method of claim 1, wherein the determining the distance between the target subject and the electronic device based on the depth value and the weight value of each image region comprises: processing the depth value of each image region based on the weight value of the image region, to obtain a weighted depth value of each image region; and fusing each weighted depth value to obtain the distance between the target subject and the electronic device.
6. The method of any one of claims 1 to 5, further comprising: determining an image of at least one subject based on a target frame image in the sequence of frame images and position information of at least one subject in the target frame image; determining at least one image of a subject after brightness correction and / or sharpness correction based on the image of each subject and a corresponding template image, wherein each template image comprises at least one subject; and determining a depth map of the target frame image based on the image of the target subject after brightness correction and / or sharpness correction.
7. The method of claim 6, wherein the determining at least one image of a subject after brightness correction and / or sharpness correction based on the image of each subject and a corresponding template image comprises: performing color space conversion on the image of each subject and the corresponding template image, to obtain a template image in at least one target color space and an image of a subject in at least one target color space; wherein the target color space comprises at least brightness information of an image; and performing brightness correction and / or sharpness correction on the image of each subject based on the brightness of each template image and the brightness information of the image of each subject.
8. The method of any one of claims 1 to 5, wherein the method further comprises, before the determining the weight value of each image region based on the depth value of at least one image region in the depth map of the target subject, the method further comprises: determining an average depth value of each image region and a variance threshold; and performing smoothing processing on a depth value in each image region based on the depth value of a neighboring region of the image region, in a case where a variance value of the depth value and the average depth value is greater than the variance threshold.
9. The method of claim 8, further comprising: determining a height of the target subject based on the image region after the smoothing processing. Adjust a parameter for recognizing the target photographic object based on the height of the target photographic object and the distance between the target photographic object and the electronic device.
10. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, the processor implementing the steps of the method of any one of claims 1 to 9 when executing the program.
Citation Information
Patent Citations
Imaging method and device and mobile terminal
CN105827933A
Image processing method and device, storage medium and terminal
CN111524087A