Image processing method, device, equipment and storage medium
By acquiring and processing image depth information, the size of the target object is adjusted to solve the problem of unstable portraits in video conferencing, achieving stable display and beautiful effects of the target object in the image.
Patent Information
- Application Number
- CN202210398467.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-04-15
AI Technical Summary
In a video conferencing scenario, the change in the distance between a person and the camera causes the person's portrait to appear too large or too small. When a person moves back and forth from the camera, the person's portrait in the video will appear larger or smaller, affecting the appearance.
By acquiring an image and its depth information, the depth of the target object is extracted, and the image is processed based on the depth of the target object and a reference depth, and the size of the target object is adjusted to maintain a suitable ratio and stable display.
This prevents the target object from being too large or too small in the image, reduces the feeling of dizziness, and improves the display effect of the target object in the image, especially maintaining a stable size and proportion of the portrait in video conferencing.
Smart Images

Figure CN114882089B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and is related to but not limited to image processing methods, devices, equipment and storage media. Background Art
[0002] With the continuous development of image processing technology, image and video related applications have been widely developed.
[0003] In a video conferencing scenario, if the distance between a person and the camera changes, it will affect the person's imaging, and the displayed person's portrait may be too large or too small.
[0004] If a person moves back and forth in front of the camera, the person's image in the video will become larger and smaller, which is not very beautiful. Summary of the Invention
[0005] The present application provides an image processing method, apparatus, device, and storage medium, which improve the display effect of a target object in a target image.
[0006] The technical solution of this application is achieved as follows:
[0007] The present application provides an image processing method, the method comprising:
[0008] Acquire a first image and depth information corresponding to the first image;
[0009] extracting a depth of a target object in the first image based on the first image and depth information corresponding to the first image;
[0010] The first image is processed based at least on the depth of the target object and a reference depth to obtain a target image.
[0011] The present application provides an image processing device, comprising:
[0012] an acquiring unit, configured to acquire a first image and depth information corresponding to the first image;
[0013] an extraction unit, configured to extract a depth of a target object in the first image based on the first image and depth information corresponding to the first image;
[0014] A processing unit is configured to process the first image based at least on the depth of the target object and a reference depth to obtain a target image.
[0015] The present application also provides an electronic device, comprising: a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor implements the above-mentioned image processing method when executing the program.
[0016] The present application also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned image processing method is implemented.
[0017] The image processing method, apparatus, device and storage medium provided in the present application include: obtaining a first image and depth information corresponding to the first image; extracting the depth of the target object in the first image based on the first image and the depth information corresponding to the first image; processing the first image based on at least the depth of the target object and the reference depth to obtain a target image. For the scheme of the present application, by obtaining the depth of the target object and processing the first image according to the depth of the target object and the reference depth, the display effect of the target object in the target image is improved. For example, in an image processing scenario, the target image obtained after processing according to the depth of the target object and the reference depth can avoid the situation where the target object in the target object is too large or too small; for another example, in a video or conference scenario, since each image is processed based on the reference depth, the dizziness caused by the sudden increase or decrease of the target object's imaging when the distance between the target object and the camera changes can be avoided. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A schematic diagram of an optional structure of the image processing system provided in an embodiment of the present application;
[0019] Figure 2 An optional structural diagram of the image processing system provided in the embodiment of the present application
[0020] Figure 3 A schematic diagram of an optional flow chart of the image processing method provided in an embodiment of the present application;
[0021] Figure 4 A schematic diagram of an optional flow chart of the image processing method provided in an embodiment of the present application;
[0022] Figure 5 A schematic diagram of an optional flow chart of the image processing method provided in an embodiment of the present application;
[0023] Figure 6 A schematic diagram of an optional flow chart of the image processing method provided in an embodiment of the present application;
[0024] Figure 7 A schematic diagram of an optional flow chart of the image processing method provided in an embodiment of the present application;
[0025] Figure 8 A schematic diagram of an optional flow chart of the image processing method provided in an embodiment of the present application;
[0026] Figure 9 A schematic diagram of an optional flow chart of the image processing method provided in an embodiment of the present application;
[0027] Figure 10 An optional schematic diagram of an image before processing provided in an embodiment of the present application;
[0028] Figure 11 An optional structural diagram of the adjustment process provided in an embodiment of the present application;
[0029] Figure 12 An optional schematic diagram of a processed image provided in an embodiment of the present application;
[0030] Figure 13 An optional schematic diagram of an image before processing provided in an embodiment of the present application;
[0031] Figure 14 An optional structural diagram of the adjustment process provided in an embodiment of the present application;
[0032] Figure 15 An optional schematic diagram of a processed image provided in an embodiment of the present application;
[0033] Figure 16 A schematic diagram of an optional structure of the image processing device provided in an embodiment of the present application;
[0034] Figure 17 This is a schematic diagram of an optional structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0035] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the specific technical solutions of the application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.
[0036] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0037] In the following description, the terms "first, second, and third" are used merely as examples to distinguish between different objects and do not represent a specific order or precedence for the objects. It is understood that the specific order or precedence of "first, second, and third" can be interchanged where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0039] The embodiments of the present application may provide an image processing method, apparatus, device, and storage medium. In practical applications, the image processing method may be implemented by an image processing apparatus, and the functional entities in the image processing apparatus may be collaboratively implemented by hardware resources of the electronic device, such as computing resources such as a processor, and communication resources (e.g., for supporting various communication methods such as optical cables and cellular communications).
[0040] The image processing method provided in the embodiment of the present application is applied to an image processing system, and the image processing system includes: an image processing device.
[0041] The image processing device is used to perform the following operations: obtaining a first image and depth information corresponding to the first image; extracting the depth of a target object in the first image based on the first image and the depth information corresponding to the first image; and processing the first image based at least on the depth of the target object and a reference depth to obtain a target image.
[0042] Optionally, the image processing system may further include an image acquisition device, wherein the image acquisition device may be used to acquire the first image and depth information corresponding to the first image.
[0043] Optionally, the image processing system may further include an image display device, wherein the image display device is used to display images. For example, the image display device may display the first image and the target image.
[0044] It should be noted that the image processing device, image acquisition device and image display device can be integrated into the same electronic device; they can also be independently deployed on different electronic devices; or two of the image processing device, image acquisition device and image display device can be integrated into one electronic device, and the other one can be independently deployed on a different electronic device.
[0045] As an example, the image processing system is applied to image processing scenarios such as Figure 1 As shown, the image processing system includes: an image processing device 101 and an image acquisition device 102. The image processing device 101 and the image acquisition device 102 can transmit data.
[0046] The image acquisition device 102 is configured to acquire a first image and depth information corresponding to the first image, and send the acquired first image and the depth information corresponding to the first image to the image processing device 101 .
[0047] The image acquisition device 102 may be an electronic device capable of acquiring relevant images and depth information.
[0048] Exemplarily, the image acquisition device 102 may include a red, green, blue (RGB) camera 1021 and a depth camera 1022. Accordingly, the first image is acquired by the RGB camera, and the depth information corresponding to the first image is acquired by the depth camera.
[0049] The image processing device 101 is used to perform the following steps: obtaining a first image and depth information corresponding to the first image; extracting the depth of a target object in the first image based on the first image and the depth information corresponding to the first image; and processing the first image based at least on the depth of the target object and a reference depth to obtain a target image.
[0050] The image processing device 101 may be an electronic device with relevant image processing capabilities. For example, the image processing device 101 may include, but is not limited to: terminal devices such as mobile phones and tablet computers; processing devices such as servers and desktop computers; smart wearable devices such as smart watches; or virtual devices with relevant image processing capabilities.
[0051] As another example, the image processing system is applied to a two-party video processing scenario, such as Figure 2 As shown, the image processing system includes a user 1 side and a user 2 side. The user 1 side includes a first image processing device 201 and a first image acquisition device 202 , and the user 2 side includes a second image processing device 203 and a second image acquisition device 204 .
[0052] Communication can be performed between the first image processing device 201 and the first image acquisition device 202 , between the second image processing device 203 and the second image acquisition device 204 , and between the first image processing device 201 and the second image processing device 203 .
[0053] The first image acquisition device 202 is used to acquire images and depth information of user 1 , and the second image acquisition device 204 is used to acquire images and depth information of user 2 .
[0054] It should be noted here that when the image processing process of the embodiment of the present application is performed on the collected image of user 1, the image processing process can be completed on the first image processing device 201 side; or the collected image of user 1 can be sent to the second image processing device 203, and the image processing process can be completed on the second image processing device 203 side.
[0055] It should be noted that the embodiments of this application do not limit the specific number of image processing devices and image acquisition devices included in the image processing system, and can be configured according to specific scenario requirements. For example, the image processing system can also be used in multi-person conferences or video scenarios, and the specific number of image processing devices and image acquisition devices can be determined based on the number of users in the conference or video scenario.
[0056] Next, combine Figure 1 and Figure 2 The schematic diagram of the image processing system shown in FIG. 1 illustrates various embodiments of the image processing method, apparatus, device, and storage medium provided in the embodiments of the present application.
[0057] In a first aspect, an embodiment of the present application provides an image processing method, which is applied to an image processing device; wherein the image processing device can be deployed in Figure 1 The image processing device 101 in Figure 2 The first image processing device 201 is enclosed in the brackets, or Figure 2 The second image processing device 203 in the embodiment of the present application is described below with the image processing device as the execution subject.
[0058] Figure 3 A flow chart showing an optional image processing method is shown in FIG. Figure 3 The image processing method may include but is not limited to the following: Figure 3 S301 to S303 shown.
[0059] S301: An image processing device obtains a first image and depth information corresponding to the first image.
[0060] The first image refers to any image to be processed.
[0061] Exemplarily, the first image may be a captured original image, or may be an image after the original image has been beautified (eg, by smoothing the skin, slimming the face, etc.).
[0062] The first image includes the target object.
[0063] The target object refers to the subject in the first image. It can also be considered as the subject captured by the image acquisition device. Exemplarily, the first image may include the target object and the background.
[0064] For example, when photographing person A, the target object in the first image is person A.
[0065] The embodiment of the present application does not limit the specific object type of the target object, which can be determined according to the specific scenario.
[0066] In a possible implementation, the target object is an object that can move autonomously. For example, the target object may include but is not limited to: a person, a kitten, or a puppy, etc.
[0067] In another possible embodiment, the target object is an object that cannot move autonomously. For example, the target object may include, but is not limited to, a cup, clothing, or a book. In this case, although these objects cannot move autonomously, they can be passively moved by people or other objects, so these objects can still serve as target objects in the embodiments of this application.
[0068] The depth information corresponding to the first image includes a depth value corresponding to each pixel in the first image.
[0069] In a possible implementation, S301 may be implemented as follows: the image processing device captures the first image through its own RGB camera, and captures depth information corresponding to the first image through its own depth camera.
[0070] In another possible implementation, the first image and the depth information corresponding to the first image are captured by an image capture device, and the captured first image and the depth information corresponding to the first image are sent to an image processing device. Correspondingly, S301 can be implemented as follows: the image processing device receives the first image and the depth information corresponding to the first image sent by the image capture device.
[0071] S302: The image processing device extracts the depth of the target object in the first image based on the first image and depth information corresponding to the first image.
[0072] S302 can be implemented as follows: the image processing device extracts the target object in the first image, obtains multiple target pixels corresponding to the target object, then determines the depth values corresponding to the multiple target pixels in the depth information corresponding to the first image, and determines the target object based on the multiple depth values.
[0073] S303: The image processing device processes the first image based at least on the depth of the target object and a reference depth to obtain a target image.
[0074] The specific value of the reference depth in the embodiment of the present application is not limited and can be configured according to actual needs. Here, for a target object, it can correspond to one reference depth or multiple reference depths.
[0075] The embodiments of the present application do not limit the specific operations performed on the first image and can be configured according to actual needs and scenarios. For example, the first processing may include but is not limited to: adjusting the target object in the first image, blurring or replacing the background image corresponding to the first image, etc.
[0076] In Example 1, S303 may be implemented as follows: the image processing device adjusts the target object in the first image based at least on the depth of the target object and the reference depth to obtain a target image.
[0077] Example 2, S303 can be implemented as follows: the image processing device adjusts the target object in the first image based at least on the depth of the target object and the reference depth; then processes the background image corresponding to the first image, and superimposes the adjusted image of the target object with the background image to obtain the target image.
[0078] The image processing solution provided by the embodiment of the present application includes: obtaining a first image and depth information corresponding to the first image; extracting the depth of the target object in the first image based on the first image and the depth information corresponding to the first image; processing the first image based on at least the depth of the target object and the reference depth to obtain a target image. For the solution of the present application, by obtaining the depth of the target object and processing the first image according to the depth of the target object and the reference depth, the display effect of the target object in the target image is improved. For example, in an image processing scenario, the target image obtained after processing according to the depth of the target object and the reference depth can avoid the situation where the target object in the target object is too large or too small; for another example, in a video or conference scenario, since each image is processed based on the reference depth, the dizziness caused by the sudden increase or decrease of the target object's imaging when the distance between the target object and the camera changes can be avoided.
[0079] Next, a description will be given of a process in which the image processing device extracts the depth of the target object in the first image based on the first image and the depth information corresponding to the first image at S302 .
[0080] Specific methods include but are not limited to the following two:
[0081] Method A: Get the depth of the target object based on all target pixels corresponding to the target object;
[0082] Method B: obtaining the depth of the target object based on a portion of target pixels corresponding to the target object.
[0083] The following describes the implementation process of method A. Figure 4 The process may include but is not limited to the following S3021 to S3023.
[0084] S3021. The image processing device extracts N target pixels corresponding to the target object in the first image.
[0085] S3021 can be implemented as follows: the image processing device uses an image recognition algorithm to identify the target object in the first image, and then uses an edge extraction algorithm to determine the edge pixels of the target object in the first image, and uses the edge pixels and other pixels inside the edge pixels as the N target pixels corresponding to the target object.
[0086] It should be noted that the N target pixels are all target pixels corresponding to the target object.
[0087] S3022. The image processing device determines, from the depth information corresponding to the first image, depth values corresponding to the N target pixels.
[0088] Since the first image and the depth information corresponding to the first image correspond to each other in a one-to-one manner based on pixel granularity, S3022 may be implemented as follows: the image processing device searches for depth values corresponding to the N target pixels in the depth information corresponding to the first image.
[0089] S3023: The image processing device determines the depth of the target object based on the depth values corresponding to the N target pixels.
[0090] The embodiment of the present application does not limit the specific method of determining the depth of the target object based on the depth values corresponding to the N target pixels, and can be configured according to actual needs.
[0091] The depth of the target object may be an average value, a mode value, or a weighted average value of depth values corresponding to N target pixels.
[0092] For the weighted average approach, the depth values corresponding to the N target pixels can be divided into different parts, and then a corresponding weight can be assigned to each part. For example, if the target object is a person, the weight of the head can be assigned to 0.8, the weight of the torso to 0.95, and the weight of the hand to 0.3.
[0093] Compared with method B, the solution of method A is simpler to implement.
[0094] The following describes the implementation process of Method B. The difference between Method A and Method B is that in Method B,
[0095] Therefore, in Method B, the depth of the target object is determined based on a portion of the target pixels corresponding to the target object. Accordingly, in Method B, the depth values corresponding to L target pixels must be selected from the original N depth values corresponding to the target pixels. The depth of the target object is then determined based on the depth values corresponding to the L target pixels. Where L is less than N.
[0096] The implementation process of determining the depth of the target object based on the depth values corresponding to the L target pixels is similar to that of method A. For specific implementation, please refer to method A and will not be described in detail here.
[0097] Regarding selecting the depth values corresponding to L target pixels from the depth values corresponding to the original N target pixels, the embodiment of the present application does not limit the specific selection method and can be configured according to actual needs.
[0098] For example, when the target object is a person, selection can be made based on the part of the person; or selection can be made based on the concentration of depth values corresponding to N target pixels, that is, depth values with high dispersion are filtered out.
[0099] Compared with method A, the depth value of the target object determined by the solution of method B is more accurate.
[0100] Next, a description will be given of a process in which the image processing device processes the first image based on at least the depth of the target object and the reference depth to obtain the target image in S303.
[0101] This process may include but is not limited to:
[0102] Method A1: Process the target object in the first image to obtain a target image;
[0103] Mode B1: Process the target object in the first image and the background image in the first image to obtain a target image.
[0104] Next, the process of processing the target object in the first image to obtain the target image in method A1 is described. Figure 5 The process may include but is not limited to the following S3031 and S3032.
[0105] S3031. The image processing device adjusts the target object based at least on the depth of the target object and a reference depth.
[0106] S3031 may be implemented by: the image processing device enlarging or reducing the target object based at least on the depth of the target object and the reference depth.
[0107] Whether to scale up or scale down can be configured based on actual needs. For example, when the size of the target object is larger than a first size threshold, the target object can be scaled down based at least on the depth of the target object and the reference depth; and when the size of the target object is smaller than a second size threshold, the target object can be scaled up based at least on the depth of the target object and the reference depth.
[0108] The second size threshold is smaller than the first size threshold.
[0109] The embodiment of the present application does not limit the specific method of adjusting the target object, and can be configured according to actual needs.
[0110] In a possible implementation, the image processing device may adjust the two-dimensional image corresponding to the target object to achieve adjustment of the target object.
[0111] In another possible implementation, the image processing device may adjust the three-dimensional image corresponding to the target object, and then map the three-dimensional model into a two-dimensional image to achieve adjustment of the target object.
[0112] S3032. The image processing device superimposes the adjusted image of the target object with the background image corresponding to the first image to obtain the target image.
[0113] The background image corresponding to the first image is used to limit the background of the target object. The embodiment of the present application does not limit the specific image content of the background image corresponding to the first image and can be configured according to actual needs.
[0114] In a possible implementation, the background image corresponding to the first image may be an image of the background portion of the first image excluding the target object.
[0115] In another possible implementation, the background image corresponding to the first image may be an image of the background where the target object is located that is captured in advance.
[0116] S3032 can be implemented as follows: the image processing device uses the adjusted image of the target object as the foreground, uses the background image corresponding to the first image as the background corresponding to the foreground, and then superimposes the foreground and the background to obtain the target image.
[0117] Comparing method A1 with method B1, the solution of method A1 is simple to implement, while the solution of method B1 has a better display effect of the target object in the target image.
[0118] Next, the process of adjusting the target object based on at least the depth of the target object and the reference depth by the image processing device at step S3031 will be described. This process may include but is not limited to the following method A2 or method B2.
[0119] Method A2: Using the two-dimensional image as the dimension, adjust the target object;
[0120] Method B2: Using the three-dimensional model as the dimension, adjust the target object.
[0121] Next, the process of adjusting the target object using a two-dimensional image as the dimension in method A2 is described.
[0122] The implementation of this process may include: the image processing device determines the size relationship between the depth of the target object and the reference depth; if the depth of the target object is greater than the reference depth, the portion corresponding to the target object in the first image is enlarged; if the depth of the target object is less than the reference depth, the portion corresponding to the target object in the first image is reduced.
[0123] Exemplarily, the image processing device determines a first ratio between the depth of the target object and the reference depth, and judges the size relationship between the first ratio and one. If the first ratio is greater than 1, the target object is enlarged by the first ratio times; if the first ratio is less than 1, the target object is reduced by the first ratio times.
[0124] Method B2: Using the three-dimensional model as the dimension, adjust the target object.
[0125] Next, the process of adjusting the target object using the three-dimensional model as the dimension in method B2 will be described.
[0126] like Figure 6 As shown, the implementation of this process may include but is not limited to the following S601 to S603.
[0127] S601: An image processing device constructs a three-dimensional model based on depth values corresponding to N target pixels in the target object and color values corresponding to the N target pixels.
[0128] The three-dimensional model is used to represent the three-dimensional information of the target object.
[0129] S601 can be implemented as follows: the image processing device determines, for each target pixel among N target pixels, the three-dimensional point corresponding to the pixel in the three-dimensional model based on the depth value corresponding to the target pixel and the color value corresponding to the target pixel, thereby obtaining N three-dimensional points, and then determining the three-dimensional model based on the N three-dimensional points.
[0130] Here, the coordinate value of the three-dimensional point corresponding to a pixel includes: the pixel value of the pixel in two directions, and the depth value corresponding to the pixel.
[0131] The color value of a three-dimensional point corresponding to a pixel may be an RGB value passing through the pixel.
[0132] S602: The image processing device adjusts the imaging distance of the target object based on the depth of the target object and a reference depth.
[0133] In a possible implementation, if the depth of the target object is greater than the reference depth, S602 may be implemented as: the image processing device reduces the imaging distance of the target object.
[0134] In another possible implementation, if the depth of the target object is less than the reference depth, S602 may be implemented as: the image processing device increases the imaging distance of the target object.
[0135] Here, there is no limitation on the specific scale of increasing or decreasing, and it can be configured according to actual needs.
[0136] S603: The image processing device maps the three-dimensional model based on the adjusted imaging distance to obtain an adjusted image of the target object.
[0137] S603 may be implemented as follows: the image processing device maps the three-dimensional model into a two-dimensional image based on the adjusted imaging distance mapping, thereby obtaining an adjusted image of the target object.
[0138] Comparing Method A2 with Method B2, Method A2 is simple to implement. Since the target objects are not in absolute proportional relationship, Method B2 adjusts the target objects through a three-dimensional model, and the result obtained is more consistent with reality.
[0139] The data processing method provided in the embodiment of the present application can also adjust the local parts of the target object. The following takes the first part as an example to illustrate the local adjustment process. Figure 7 As shown, the process may include but is not limited to the following S701 and S702.
[0140] S701: An image processing device determines that a ratio of a first part of a target object to a first size of the target object is outside a preset ratio range.
[0141] S701 can be implemented as follows: the image processing device detects the size ratio of each part in the target image to the target object, obtains multiple size ratios, and determines whether the size ratio corresponding to each part is within a preset ratio range; if the image processing device determines that the size ratio of a first part in the target object to a first size of the target object is outside the preset ratio range corresponding to the first part, the following S702 is executed on the first part.
[0142] The embodiment of the present application does not limit the specific value of the preset ratio range, and can be configured according to actual needs.
[0143] It should be noted that for different parts, the corresponding size ratio range may be different. The specific size ratio range is determined according to the object type and part type of the target object.
[0144] S702: The image processing device adjusts the size of the first part of the target object so that the adjusted ratio of the first part to the second size of the target object is within the preset ratio range.
[0145] The embodiment of the present application does not limit the granularity of adjusting the size of the first part, and can be configured according to actual needs.
[0146] Since the first part is close to the camera, distortion may occur. Through the adjustment process of the first part in the embodiment of the present application, the distortion can be corrected, thereby further improving the display effect of the target object.
[0147] Next, the process of obtaining the background image corresponding to the first image in advance is described. Figure 8 As shown, the process may include but is not limited to the following S801 and S802.
[0148] S801: An image processing device obtains M frames of second images that meet a first condition.
[0149] There is no target object in the second image.
[0150] S801 may be implemented as follows: the image processing device obtains the first J frames of images of the first image, and selects M frames of second images that meet the first condition and do not contain the target object from the J frames of images.
[0151] The embodiment of the present application does not specifically limit the values of J and M, and can be configured according to actual needs.
[0152] The first condition is used to limit the number of second images in the M frames. The embodiment of the present application does not limit the specific content of the first condition.
[0153] Exemplarily, the first condition may include: the similarity between the images is greater than or equal to a first similarity threshold. In simple terms, the target image is not only absent from the M frames of second images, but also has similar content.
[0154] S802: The image processing device synthesizes the background image based on the M frames of second images.
[0155] S802 may be implemented as follows: the image processing device fuses the M frames of second images, and uses the fused image as the background image corresponding to the first image.
[0156] Acquiring the background image in advance can, on the one hand, improve the display effect of the background image; on the other hand, it can reduce the amount of data processing and improve processing efficiency.
[0157] Next, the process of processing the target object in the first image and the background image in the first image in mode B1 will be described. Figure 9 As shown, the process may include but is not limited to the following S901 to S903.
[0158] S901: The image processing device adjusts the size of the target object.
[0159] The implementation of S901 may refer to the process of adjusting the target object by the image processing device at least based on the depth of the target object and the reference depth in S3031, which will not be described in detail here.
[0160] S902: The image processing device determines that a size change of the target object is greater than a first threshold.
[0161] The size change of the target object refers to the size change of the target object after adjustment compared with that before adjustment, that is, the size change obtained by comparing the size of the target object in the first image with the size of the target object in the target image.
[0162] The embodiment of the present application does not limit the value of the first threshold and can be configured according to actual needs.
[0163] S902 can be implemented as follows: the image processing device determines the difference between the image size corresponding to the adjusted target object and the image size corresponding to the target object before adjustment, takes the difference as a size change, and determines whether the size change is greater than a first threshold. If the size change is greater than the first threshold, execute the following S903.
[0164] It should be noted that if the size change is less than or equal to the first threshold, the adjusted image of the target object is directly superimposed on the background image to obtain the target image.
[0165] S903: The image processing device performs first processing on the background image corresponding to the first image.
[0166] The first processing includes background blurring or background replacement.
[0167] S903 may be implemented as follows: the image processing device starts a first function to perform a first process on the background image corresponding to the first image. The first function may be a background blur function or a background replacement function.
[0168] In this embodiment, when the size change of the target object is greater than a first threshold, the background image is blurred or replaced, which can further improve the display effect of the target object.
[0169] Below, taking a video conferencing scenario as an example, the image processing method provided by this application is described through an embodiment.
[0170] During a video conference, the changing distance between the person and the camera can cause image disproportion. For example, if someone waves their arms, their hands appear oversized. Disproportionate proportions between the hands and the body, or between the hands and the face, can create a comical effect. When someone moves back and forth in front of the camera, their image in the video can appear disproportionately large and small, creating an unsightly effect.
[0171] In response to the above problems, an embodiment of the present application provides an image processing method that can improve the above problems.
[0172] Below, the principle of the image processing method provided in the embodiment of the present application is briefly described.
[0173] The camera module (RGB camera and depth camera) is used to capture images of the human body and the corresponding depth information. Combined with human body segmentation technology, the captured human portrait is divided into different parts. The size of the human hand is then adjusted to make the proportions of various parts of the human body (such as the human hand) more appropriate to the size of the human body / face.
[0174] The specific processing method may include but is not limited to the following steps 1 to 4.
[0175] Step 1: Use an RGB camera and a depth camera during video conferencing to capture RGB images and the depth information corresponding to the RGB images.
[0176] Step 2: Use RGB images and depth information, and use deep learning / computer vision methods to perform image segmentation on faces / bodies / hands, and extract images and depth information of faces, bodies, and hands.
[0177] Step 3: Based on the extracted images and depth information of the face, body, and hands, a three-dimensional model is created for the portraits of the people in the current meeting.
[0178] Step 4: Adjust the 3D model of the segmented face / body / hand, perform 2D mapping at an optimal camera distance (e.g., 1 meter), and regenerate a 2D image.
[0179] The adjustments in step 4 may include but are not limited to the following three situations:
[0180] Case 1: Bring the original portrait 3D model that is too far away closer and enlarge the portrait.
[0181] Case 2: Move the original portrait 3D model that is too close further away to reduce the portrait size.
[0182] Case 3: According to the proportion of the portrait, adjust the size of the abnormal part (such as the hand) so that the abnormal part accounts for a proper proportion of the human body / face.
[0183] It should be noted that the abnormal parts here can be one or more, and the embodiments of the present application do not limit this.
[0184] For example, at the first moment, the camera captures image A (equivalent to the first image), such as Figure 10 As shown, it can be seen that in image A, the target object is person H, the portrait proportion of person H is too large (the size ratio of person H to image A is not appropriate), and the hand of person H is too large (that is, the size ratio of person H's hand to the overall size of person H is not appropriate).
[0185] Image A is processed using the image processing method of the present application, the size of the person H in image A is adjusted, and the size of the hand of the person H in image A is adjusted to obtain image B (equivalent to the target image).
[0186] The process of adjusting the size of the character H in image A can be referred to Figure 11 In the example shown, a 3D model A is created based on the image and depth information of a person H. The depth of person H is 30 cm (i.e., the original imaging distance is 30 cm). Since the reference depth is 1 meter, the imaging distance of 3D model A is increased, and the image is re-imaged. The hand-to-face ratio of the person is adjusted to an appropriate level. After adjustment, the size of person H in image A is 2 / 3, and the hand-to-face ratio is also appropriate.
[0187] Image B can be referenced Figure 12 The content shown, from Figure 12 It can be seen from the figure that in image B, the size ratio of person H to image B is 2 / 3, that is, the size of person H is appropriate, and the ratio of human hands to human faces is also appropriate.
[0188] For example, at the second moment, the distance between the person H and the camera moves, and the camera captures an image C (equivalent to the first image), such as Figure 13 As shown, it can be seen that in image C, the target object is person H, and the portrait proportion of person H is too small (the size ratio of person H to image C is not appropriate) and cannot be seen clearly.
[0189] Image A is processed using the image processing method of the present application, and the size of the person H in image C is adjusted to obtain image D (equivalent to the target image).
[0190] The process of adjusting the size of the character H in the image C can be referred to Figure 14In the example shown, 3D model B is created based on the image and depth information of person H. Person H's depth is 3 meters (i.e., the original imaging distance is 3 meters). Since the reference depth is 1 meter, the imaging distance is too far. Therefore, the imaging distance of 3D model A is shortened and re-imaged. After adjustment, the size of person H in image C is now 2 / 3, which is an appropriate proportion.
[0191] Image D can be referenced Figure 15 The content shown, from Figure 15 It can be seen from the figure that in image D, the size ratio of character H to image D is 2 / 3, that is, the size of character H is appropriate.
[0192] The technical effects of adopting this solution may include but are not limited to:
[0193] First, the proportions of the human body and hands are more coordinated (appropriate) and more beautiful.
[0194] Second, the proportion of portrait size to image size is more appropriate, and the distance between the human body and the camera is more appropriate.
[0195] Third, there is no problem of the size of local parts being inappropriately proportioned to the body size (i.e. there is no problem of big hands).
[0196] Fourthly, when a person moves back and forth in front of the camera, the image in the video will appear larger or smaller.
[0197] In the second aspect, in order to implement the above-mentioned image processing method, an image processing device according to an embodiment of the present application is provided. Figure 16 The structure of the image processing device is described with reference to FIG.
[0198] like Figure 16 As shown, the image processing device 160 includes: an acquisition unit 1601, an extraction unit 1602, and a processing unit 1603.
[0199] An acquiring unit 1601 is configured to acquire a first image and depth information corresponding to the first image;
[0200] an extraction unit 1602 for extracting a depth of a target object in the first image based on the first image and depth information corresponding to the first image;
[0201] The processing unit 1603 is configured to process the first image based at least on the depth of the target object and a reference depth to obtain a target image.
[0202] In some embodiments, the extraction unit 1602 is specifically configured to:
[0203] Extracting N target pixels corresponding to the target object in the first image;
[0204] Determine, in the depth information corresponding to the first image, depth values corresponding to the N target pixels respectively;
[0205] The depth of the target object is determined based on the depth values respectively corresponding to the N target pixels.
[0206] In some embodiments, the processing unit 1603 is specifically configured to:
[0207] adjusting the target object based at least on the depth of the target object and a reference depth;
[0208] The adjusted image of the target object is superimposed on a background image corresponding to the first image to obtain the target image.
[0209] In some embodiments, the processing unit 1603 is further specifically configured to:
[0210] Constructing a three-dimensional model based on the depth values corresponding to the N target pixels in the target object and the color values corresponding to the N target pixels; the three-dimensional model is used to represent the three-dimensional information of the target object;
[0211] adjusting an imaging distance of the target object based on the depth of the target object and a reference depth;
[0212] The three-dimensional model is mapped based on the adjusted imaging distance to obtain an adjusted image of the target object.
[0213] In some embodiments, the image processing device 160 includes an adjustment unit configured to:
[0214] Determining that a ratio of a first portion of the target object to a first size of the target object is outside a preset ratio range;
[0215] The size of the first part of the target object is adjusted so that the adjusted ratio of the first part to the second size of the target object is within the preset ratio range.
[0216] In some embodiments, the image processing device 160 includes an obtaining unit, which is configured to, before superimposing the image after the target object is adjusted with the background image corresponding to the first image to obtain the target image, perform the following steps: obtaining M frames of second images that meet a first condition; wherein the target object does not exist in the second images;
[0217] The background image is synthesized based on the M frames of second images.
[0218] In some embodiments, the processing unit 1603 is further specifically configured to:
[0219] adjusting the size of the target object;
[0220] Determining that a size change of the target object is greater than a first threshold;
[0221] Performing a first processing on a background image corresponding to the first image; wherein the first processing includes background blurring or background replacement.
[0222] It should be noted that the image processing device provided in the embodiment of the present application includes the various units included, which can be implemented by a processor in an electronic device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0223] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.
[0224] It should be noted that, in the embodiment of the present application, if the above-mentioned image processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0225] In the third aspect, in order to implement the above-mentioned image processing method, an embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps in the image processing method provided in the above-mentioned embodiment are implemented.
[0226] The following combination Figure 17 The electronic device 170 shown is a block diagram of the electronic device.
[0227] In one example, the electronic device 170 may be the electronic device described above. Figure 17 As shown, the electronic device 170 includes: a processor 1701, at least one communication bus 1702, a user interface 1703, at least one external communication interface 1704, and a memory 1705. The communication bus 1702 is configured to enable communication between these components. The user interface 1703 may include a display screen, and the external communication interface 1704 may include a standard wired interface and a wireless interface.
[0228] The memory 1705 is configured to store instructions and applications executable by the processor 1701, and can also cache data to be processed or processed by the processor 1701 and various modules in the electronic device (for example, image data, audio data, voice communication data and video communication data), which can be implemented through flash memory (FLASH) or random access memory (RAM).
[0229] In a fourth aspect, an embodiment of the present application provides a storage medium, that is, a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the image processing method provided in the above embodiment are implemented.
[0230] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0231] It should be understood that “one embodiment” or “an embodiment” mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, “in one embodiment” or “in some embodiments” appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0232] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0233] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0234] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0235] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0236] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.
[0237] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0238] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An image processing method, comprising: Acquire a first image and depth information corresponding to the first image; extracting a depth of a target object in the first image based on the first image and depth information corresponding to the first image; The first image is processed based at least on the depth and reference depth of the target object to obtain a target image; wherein, the processing of the first image includes: adjusting the target object based at least on the depth and reference depth of the target object; superimposing the adjusted image of the target object with the background image corresponding to the first image to obtain the target image; wherein, adjusting the target object based at least on the depth and reference depth of the target object includes: constructing a three-dimensional model based on the depth values corresponding to N target pixels in the target object and the color values corresponding to the N target pixels; the three-dimensional model is used to characterize the three-dimensional stereo information of the target object; adjusting the imaging distance of the target object based on the depth and reference depth of the target object; mapping the three-dimensional model based on the adjusted imaging distance to obtain the adjusted image of the target object.
2. The method according to claim 1, wherein extracting the depth of the target object in the first image based on the first image and the depth information corresponding to the first image comprises: Extracting N target pixels corresponding to the target object in the first image; Determine, in the depth information corresponding to the first image, depth values corresponding to the N target pixels respectively; The depth of the target object is determined based on the depth values respectively corresponding to the N target pixels.
3. The method according to claim 1, further comprising: Determining that a ratio of a first portion of the target object to a first size of the target object is outside a preset ratio range; The size of the first part of the target object is adjusted so that the adjusted ratio of the first part to the second size of the target object is within the preset ratio range.
4. The method according to claim 1, before superimposing the adjusted image of the target object with the background image corresponding to the first image to obtain the target image, the method further comprises: Obtaining M frames of second images that meet a first condition; wherein the target object does not exist in the second images; The background image is synthesized based on the M frames of second images.
5. The method according to claim 1, wherein processing the first image further comprises: adjusting the size of the target object; Determining that a size change of the target object is greater than a first threshold; Performing a first processing on the background image corresponding to the first image; wherein the first processing includes background blurring or background replacement.
6. An image processing device, comprising: an acquiring unit, configured to acquire a first image and depth information corresponding to the first image; an extraction unit, configured to extract a depth of a target object in the first image based on the first image and depth information corresponding to the first image; A processing unit, configured to process the first image based at least on the depth and reference depth of the target object to obtain a target image; wherein processing the first image comprises: adjusting the target object based at least on the depth and reference depth of the target object; and superimposing the adjusted image of the target object with a background image corresponding to the first image to obtain the target image; wherein adjusting the target object based at least on the depth and reference depth of the target object comprises: constructing a three-dimensional model based on depth values corresponding to N target pixels in the target object and color values corresponding to the N target pixels; the three-dimensional model is used to characterize three-dimensional stereo information of the target object; adjusting the imaging distance of the target object based on the depth and reference depth of the target object; and mapping the three-dimensional model based on the adjusted imaging distance to obtain an adjusted image of the target object.
7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the image processing method according to any one of claims 1 to 5 is implemented.
8. A storage medium storing a computer program, wherein when the computer program is executed by a processor, the image processing method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Visual fatigue relieving system and method based on significant target depth dynamic adjustment
CN111695573A