Image background recognition method, device and electronic equipment
Patent Information
- Application Number
- CN202211004641.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-08-22
AI Technical Summary
[0004]然而上述两个方法都仅仅局限于图像的原始输入,且往往容易受摄像机角度、光照以及人脸姿态的干扰,仍然存在较高的错误识别率
[0019]本申请实施例提供的一种图像背景识别方法、装置及电子设备中,首先获取目标人脸深度图像;然后对目标人脸深度图像进行前景背景分割,得到背景图像;根据背景图像对应的深度信息对背景图像进行角度矫正;对角度矫正后的背景图像进行特征提取,得到背景图像特征;基于背景图像特征进行背景图像识别。该实施例基于人脸深度图像进行背景分割,可以提高分割的准确性,并利用背景图像的深度信息即D通道图像对背景图像进行角度矫正,基于矫正之后的图像进行特征提取和背景识别,可以进一步提高图像背景识别的准确率。
Smart Images

Figure CN115359531B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image background recognition method, apparatus and electronic device. Background Technology
[0002] In banking operations, certain individuals may use various methods to lure users to specific locations for facial recognition procedures. Banks refer to this type of malicious activity, where the background used for facial recognition is described as a "black background." Black background recognition can also be understood as background similarity recognition within a specific banking context.
[0003] For identifying similar backgrounds, some current methods extract background regions, use CNNs to extract features from these regions, and then calculate the distance between features of different background regions (such as cosine distance or Euclidean distance). A threshold is set; if the feature distance between two background images is less than the threshold, they are considered the same background. However, due to the rigidity of background extraction and threshold setting, this method often has a high false recognition rate (classifying dissimilar backgrounds as the same). Other methods use foreground-background segmentation networks in the extracted background region to finely segment the foreground and background, preserving as much background information as possible, thus reducing the false recognition rate to some extent.
[0004] However, both of the above methods are limited to the raw image input and are often susceptible to interference from camera angle, lighting, and facial pose, resulting in a high false recognition rate. Summary of the Invention
[0005] The purpose of this application is to provide an image background recognition method, apparatus, and electronic device. Background segmentation based on face depth images can improve the accuracy of segmentation. Furthermore, the background image is angle-corrected using the depth information of the background image, i.e., the D-channel image. Feature extraction and background recognition are then performed based on the corrected image, which can further improve the accuracy of image background recognition.
[0006] In a first aspect, embodiments of this application provide an image background recognition method, the method comprising: acquiring a depth image of a target face; performing foreground-background segmentation on the depth image of the target face to obtain a background image; performing angle correction on the background image based on the depth information corresponding to the background image; extracting features from the angle-corrected background image to obtain background image features; and performing background image recognition based on the background image features.
[0007] In a preferred embodiment of this application, the step of acquiring the target face depth image includes: acquiring the target face depth image corresponding to the target face through the iOS front-facing camera.
[0008] In a preferred embodiment of this application, the step of performing foreground-background segmentation on the target face depth image to obtain a background image includes: inputting the target face depth image into a preset foreground-background segmentation model and outputting the background image corresponding to the target face depth image.
[0009] In a preferred embodiment of this application, after the step of performing foreground-background segmentation on the depth image of the target face to obtain a background image, the method further includes: setting each pixel value of the segmented foreground image to the same specified pixel value.
[0010] In a preferred embodiment of this application, the training process of the foreground-background segmentation model is as follows: obtaining a training sample set; each sample in the training sample set is a standardized face depth image labeled with background image identifiers; using the samples in the training sample set to perform multi-scale training on a preset semantic segmentation model to obtain the foreground-background segmentation model.
[0011] In a preferred embodiment of this application, the step of obtaining the training sample set includes: obtaining multiple initial face depth images; for each initial face depth image, standardizing the RGB image and D channel image in the initial face depth image to obtain a standard face depth image; and labeling each standard face depth image with background image tags to obtain a training sample set.
[0012] In a preferred embodiment of this application, the step of standardizing the RGB image and D channel image in the initial face depth image to obtain a standard face depth image includes: standardizing the RGB image in the initial face depth image using the mean and variance of the ImageNet image set; and subtracting a specified value from the corresponding pixel value of the D channel in the initial face depth image to obtain the standard face depth image.
[0013] In a preferred embodiment of this application, the step of correcting the angle of the background image based on the depth information corresponding to the background image includes: obtaining the D-channel image corresponding to the background image; inputting the D-channel image into a preset parameter extractor to obtain six parameters; and performing an affine transformation on the background image based on the six parameters to achieve angle correction of the background image.
[0014] In a preferred embodiment of this application, the preset parameter extractor includes a 6-dimensional convolutional neural network called Mobilenet.
[0015] In a preferred embodiment of this application, the step of background image recognition based on background image features includes: calculating the similarity between the background image features and the black background features in a preset black background library; when the similarity exceeds a threshold, determining that the background image corresponding to the background image features is a suspicious black background.
[0016] Secondly, embodiments of this application also provide an image background recognition device, the device comprising: a depth image acquisition module for acquiring a depth image of a target face; a background segmentation module for performing foreground-background segmentation on the depth image of the target face to obtain a background image; a background correction module for performing angle correction on the background image based on the depth information corresponding to the background image; a feature extraction module for extracting features from the corrected background image to obtain background image features; and a background recognition module for performing background image recognition based on the background image features.
[0017] Thirdly, embodiments of this application also provide an electronic device, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method described in the first aspect above.
[0018] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are invoked and executed by a processor, the computer-executable instructions cause the processor to implement the method described in the first aspect above.
[0019] This application provides an image background recognition method, apparatus, and electronic device. First, a depth image of a target face is acquired. Then, foreground and background segmentation is performed on the target face depth image to obtain a background image. Angle correction is applied to the background image based on the depth information corresponding to the background image. Feature extraction is performed on the angle-corrected background image to obtain background image features. Background image recognition is then performed based on these features. This embodiment uses a face depth image for background segmentation, which improves segmentation accuracy. Furthermore, by utilizing the depth information of the background image (i.e., the D-channel image) to correct the angle of the background image, and then performing feature extraction and background recognition based on the corrected image, the accuracy of image background recognition can be further improved. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0021] Figure 1 A flowchart illustrating an image background recognition method provided in this application embodiment;
[0022] Figure 2 An example of foreground / background segmentation provided in this application embodiment;
[0023] Figure 3 A flowchart illustrating model training in an image background recognition method provided in this application embodiment;
[0024] Figure 4 A flowchart illustrating background image correction in an image background recognition method provided in this application embodiment;
[0025] Figure 5 A structural block diagram of an image background recognition device provided in an embodiment of this application;
[0026] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] The technical solutions of this application will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] Currently, the following two methods are commonly used for identifying similar backgrounds:
[0029] The first method involves segmenting the background region of an RGB image. It extracts the features of the background region image using a CNN and then calculates the distance between the features of different background regions (such as cosine distance and Euclidean distance). By setting a threshold, if the feature distance between two background images is less than the threshold, the two backgrounds are determined to be the same background. However, due to the rigidity of the background extraction and threshold setting, this method often has a high error rate.
[0030] The second approach, building upon the first, uses a foreground-background segmentation network for background region segmentation. By finely segmenting the foreground and background, it preserves background information as much as possible, thereby reducing the false recognition rate to some extent.
[0031] However, both of the above methods are limited to the raw image input and are often susceptible to interference from camera angle, lighting, and facial pose, resulting in a high false recognition rate.
[0032] Based on this, embodiments of this application provide an image background recognition method, apparatus, and electronic device. Background segmentation based on face depth images can improve the accuracy of segmentation, and the background image is angle-corrected by using the depth information of the background image, i.e., the D-channel image. Feature extraction and background recognition are performed based on the corrected image, which can further improve the accuracy of image background recognition.
[0033] To facilitate understanding of this embodiment, a method for image background recognition disclosed in this application will first be described in detail.
[0034] Figure 1 A flowchart of an image background recognition method provided in this application embodiment, the method specifically includes the following steps:
[0035] Step S102: Obtain the depth image of the target face.
[0036] In practical applications, the aforementioned target face depth image is usually a depth image obtained by capturing the target face with the iOS front-facing camera, i.e., an RGB-D image, which includes RGB images and D channel images.
[0037] Step S104: Perform foreground and background segmentation on the depth image of the target face to obtain the background image.
[0038] Foreground-background segmentation can be handled by a foreground-background segmentation model. This segmentation model is trained based on a pre-prepared training sample set, which consists of standardized face depth images labeled with background image identifiers.
[0039] Step S106: Correct the angle of the background image based on the depth information corresponding to the background image.
[0040] The depth information corresponding to the background image mentioned above refers to the D-channel image of the background image. Correcting the angle of the background image to a certain extent using this D-channel image can improve the accuracy of background recognition. This angle correction process can be achieved by performing an affine transformation based on parameters determined from the depth information of the background image.
[0041] Step S108: Feature extraction is performed on the background image after angle correction to obtain background image features. Feature extraction in this step can be achieved through various feature extraction models. For example, in this embodiment, a ResNet101 residual network is used to pre-train the background image after angle correction. After training, the feature map of the last layer of the model is taken and followed by a max pooling layer to extract 2048-dimensional features.
[0042] Step S110: Background image recognition is performed based on background image features.
[0043] In practice, the similarity between background image features and black background features in a preset black background library can be calculated. When the similarity exceeds a threshold, the background image corresponding to the background image feature is determined to be a suspicious black background. The similarity can be calculated using various distance formulas such as cosine distance and Euclidean distance, which will not be elaborated here. The threshold for determining whether two images share the same background can be set to 0.95.
[0044] In an image background recognition method provided in this application embodiment, a target face depth image is first acquired; then, foreground and background segmentation is performed on the target face depth image to obtain a background image; the background image is angle-corrected based on the depth information corresponding to the background image; features are extracted from the angle-corrected background image to obtain background image features; and background image recognition is performed based on the background image features. This embodiment performs background segmentation based on the face depth image, which can improve the accuracy of segmentation, and uses the depth information of the background image, i.e., the D-channel image, to perform angle correction on the background image, and performs feature extraction and background recognition based on the corrected image, which can further improve the accuracy of image background recognition.
[0045] This application also provides another image background recognition method, which is implemented based on the above embodiments; this embodiment focuses on describing the training process of the foreground-background segmentation model and the angle correction process of the background image.
[0046] The specific process of performing foreground-background segmentation on the target face depth image to obtain the background image is as follows:
[0047] The target face depth image is input into a pre-defined foreground-background segmentation model, which outputs the background image corresponding to the target face depth image. This foreground-background segmentation model can be implemented using DeepLabV3+. The resulting image after foreground-background segmentation is shown below. Figure 2 As shown, the foreground and background colors are different. In reality, the background image is blue and the foreground image is yellow, shown in different shades of gray in the figure. To avoid interference from the foreground image in the recognition of the background image, in this embodiment, each pixel value of the segmented foreground image can be set to the same specified pixel value, such as 0, which can further improve the accuracy of background image recognition.
[0048] See Figure 3 As shown, the training process of the foreground / background segmentation model described above is as follows:
[0049] Step S302: Obtain the training sample set; each sample in the training sample set is a standardized face depth image labeled with background image identifiers.
[0050] (1) Obtain multiple initial face depth images;
[0051] (2) For each initial face depth image, the RGB image and D channel image in the initial face depth image are standardized to obtain a standard face depth image.
[0052] Specifically, for the RGB images in the initial face depth image, the mean and variance of the ImageNet image set are used for standardization. For example, the pixel mean and variance of the ImageNet image set are mean = (0.485, 0.456, 0.406) and std = (0.229, 0.224, 0.225), respectively. For the pixels in the RGB three channels, the standardization of the RGB image is completed by dividing by 255 and then subtracting the mean divided by std.
[0053] For the D-channel image in the initial face depth image, a specified value needs to be subtracted from the corresponding pixel value of the D-channel to obtain the standardized D-channel image. Since the pixel values of the RGB image are approximately between -3 and 3 after standardization, while the value range of each pixel in the D-channel image is between 0 and 3, the standardization of the D-channel image can be performed by subtracting a specified value. In this embodiment, the specified value is set to 1.5, and the pixel value range of the standardized D-channel image is between -1.5 and 1.5.
[0054] (3) Mark the background image for each standard face depth image to obtain the training sample set.
[0055] Step S304: Use samples from the training sample set to train the preset semantic segmentation model at multiple scales to obtain a foreground-background segmentation model.
[0056] In this embodiment, since the input size of the acquired face depth image is 640*480 pixels, padding (edge padding) can be applied to the image width to expand it to 640*640 pixels, with a padding value of 0. A multi-scale training strategy can be used during model training, namely training at three scales: 320*320, 640*640, and 960*960. The initial learning rate is set to 0.0001, and the learning rate is reduced to 0.6 every 1 epoch. The learning rate is restarted every 5 epochs, returning to 0.0001.
[0057] The above steps for angle correction of the background image based on the depth information corresponding to the background image are described in [reference needed]. Figure 4 As shown, the specific steps include:
[0058] Step S402: Obtain the D channel image corresponding to the background image;
[0059] Step S404: Input the D channel image into the preset parameter extractor to obtain six parameters.
[0060] The aforementioned preset parameter extractor is implemented using MobileNet, a 6-dimensional convolutional neural network. By solving the D-channel image using MobileNet v3, the six parameters required for the affine transformation can be obtained.
[0061] Step S406 involves performing an affine transformation on the background image based on six parameters to correct the angle of the background image. Specifically, this can be achieved by calling cv2.WrapAffine.
[0062] Since both the face depth image captured by the camera and the segmented background image are in camera coordinate system, solving for the six parameters can convert the image in camera coordinate system to an image in world coordinate system, thereby reducing the phenomenon of inaccurate background recognition caused by factors such as camera angle.
[0063] The principle of image affine transformation is to transfer the original space (x,y) to the (u,v) space. The transformation process uses a camera model to solve for 6 parameters. The formula for transforming (x,y) space to (u,v) space is as follows:
[0064]
[0065]
[0066] Here, A is a 2*2 matrix and B is a 2*1 matrix, containing a total of 6 parameters.
[0067] The image background recognition method provided in this application utilizes the depth-sensing camera inherent in iOS devices to obtain a facial depth image, which includes D-channel depth information. This information can enhance the accuracy of background segmentation, and the D-channel information can be used to a certain extent to solve the parameters of the affine transformation, thereby correcting the original background to a certain degree, transforming the camera coordinate system image into a world coordinate system image, resisting interference from the angles between the camera, the person, and the background, improving the accuracy of background recognition, as well as the recall and precision of black background recognition.
[0068] Based on the above method embodiments, this application also provides an image background recognition device, see [link to relevant documentation]. Figure 5 As shown, the device includes:
[0069] The depth image acquisition module 502 is used to acquire the depth image of the target face; the background segmentation module 504 is used to segment the foreground and background of the depth image of the target face to obtain the background image; the background correction module 506 is used to correct the angle of the background image according to the depth information corresponding to the background image; the feature extraction module 508 is used to extract features from the corrected background image to obtain the background image features; and the background recognition module 510 is used to recognize the background image based on the background image features.
[0070] The aforementioned depth image acquisition module 502 is also used to acquire the target face depth image corresponding to the target face through the iOS front-facing camera.
[0071] The aforementioned background segmentation module 504 is also used to input the target face depth image into a preset foreground-background segmentation model and output the background image corresponding to the target face depth image.
[0072] The aforementioned background segmentation module 504 is also used to set each pixel value of the segmented foreground image to the same specified pixel value.
[0073] The aforementioned device also includes a model training module for performing the following training process for a foreground-background segmentation model: acquiring a training sample set; each sample in the training sample set is a standardized face depth image labeled with background image identifiers; and applying the samples in the training sample set to perform multi-scale training on a preset semantic segmentation model to obtain a foreground-background segmentation model.
[0074] The aforementioned model training module is also used to acquire multiple initial face depth images; for each initial face depth image, the RGB image and D channel image in the initial face depth image are standardized to obtain a standard face depth image; and background image labeling is performed on each standard face depth image to obtain a training sample set.
[0075] The aforementioned model training module is also used to standardize the RGB images in the initial face depth image using the mean and variance of the ImageNet image set; and to subtract a specified value from the corresponding pixel value of the D channel in the initial face depth image to obtain a standard face depth image.
[0076] The aforementioned background correction module 506 is also used to acquire the D-channel image corresponding to the background image; input the D-channel image into a preset parameter extractor to obtain six parameters; and perform an affine transformation on the background image based on the six parameters to achieve angle correction of the background image.
[0077] The aforementioned preset parameter extractor includes Mobilenet, a 6-dimensional convolutional neural network.
[0078] The aforementioned background recognition module 510 is also used to calculate the similarity between the background image features and the black background features in the preset black background library. When the similarity exceeds the threshold, the background image corresponding to the background image features is determined to be a suspicious black background.
[0079] The device provided in this application embodiment has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts of the device embodiment not mentioned can be referred to the corresponding content in the aforementioned method embodiment.
[0080] This application also provides an electronic device, such as... Figure 6 The diagram shows the structure of the electronic device, which includes a processor 61 and a memory 60. The memory 60 stores computer-executable instructions that can be executed by the processor 61. The processor 61 executes the computer-executable instructions to perform the following steps:
[0081] Acquire the depth image of the target face; perform foreground-background segmentation on the depth image of the target face to obtain the background image; perform angle correction on the background image based on the depth information corresponding to the background image; extract features from the angle-corrected background image to obtain the background image features; and perform background image recognition based on the background image features.
[0082] In a preferred embodiment of this application, the step of acquiring the target face depth image includes: acquiring the target face depth image corresponding to the target face through the iOS front-facing camera.
[0083] In a preferred embodiment of this application, the step of performing foreground-background segmentation on the target face depth image to obtain a background image includes: inputting the target face depth image into a preset foreground-background segmentation model and outputting the background image corresponding to the target face depth image.
[0084] In a preferred embodiment of this application, after the step of performing foreground-background segmentation on the depth image of the target face to obtain a background image, the method further includes: setting each pixel value of the segmented foreground image to the same specified pixel value.
[0085] In a preferred embodiment of this application, the training process of the foreground-background segmentation model is as follows: obtaining a training sample set; each sample in the training sample set is a standardized face depth image labeled with background image identifiers; using the samples in the training sample set to perform multi-scale training on a preset semantic segmentation model to obtain the foreground-background segmentation model.
[0086] In a preferred embodiment of this application, the step of obtaining the training sample set includes: obtaining multiple initial face depth images; for each initial face depth image, standardizing the RGB image and D channel image in the initial face depth image to obtain a standard face depth image; and labeling each standard face depth image with background image tags to obtain a training sample set.
[0087] In a preferred embodiment of this application, the step of standardizing the RGB image and D channel image in the initial face depth image to obtain a standard face depth image includes: standardizing the RGB image in the initial face depth image using the mean and variance of the ImageNet image set; and subtracting a specified value from the corresponding pixel value of the D channel in the initial face depth image to obtain the standard face depth image.
[0088] In a preferred embodiment of this application, the step of correcting the angle of the background image based on the depth information corresponding to the background image includes: obtaining the D-channel image corresponding to the background image; inputting the D-channel image into a preset parameter extractor to obtain six parameters; and performing an affine transformation on the background image based on the six parameters to achieve angle correction of the background image.
[0089] In a preferred embodiment of this application, the preset parameter extractor includes a 6-dimensional convolutional neural network called Mobilenet.
[0090] In a preferred embodiment of this application, the step of background image recognition based on background image features includes: calculating the similarity between the background image features and the black background features in a preset black background library; when the similarity exceeds a threshold, determining that the background image corresponding to the background image features is a suspicious black background.
[0091] This embodiment utilizes the depth-sensing camera inherent in iOS devices to obtain facial depth images, which include D-channel depth information. This information can enhance the accuracy of background segmentation, and the D-channel information can be used to a certain extent to resolve the parameters of the affine transformation, thereby correcting the original background to a certain degree. This transforms the camera coordinate system image into a world coordinate system image, eliminating interference from the angles between the camera, the person, and the background, and improving the accuracy of background recognition, as well as the recall and precision for black backgrounds.
[0092] exist Figure 6 In the illustrated embodiment, the electronic device further includes a bus 62 and a communication interface 63, wherein the processor 61, the communication interface 63, and the memory 60 are connected via the bus 62.
[0093] The memory 60 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 63 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 62 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 62 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0094] Processor 61 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 61 or by instructions in software form. Processor 61 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory, and the processor 61 reads the information in the memory and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiment.
[0095] This application also provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above-described method. For specific implementation details, please refer to the foregoing method embodiments, which will not be repeated here.
[0096] The computer program products of the methods, apparatus, and electronic devices provided in the embodiments of this application include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementations, please refer to the method embodiments, which will not be repeated here.
[0097] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of this application.
[0098] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0099] In the description of this application, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0100] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the technical scope disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. An image background recognition method, characterized in that, The method includes: Acquire a depth image of the target face; the depth image of the target face includes: an RGB-D channel image; The target face depth image is input into a preset foreground-background segmentation model, and the background image corresponding to the target face depth image is output; each pixel value of the segmented foreground image is set to the same specified pixel value; Angle correction of the background image based on the depth information corresponding to the background image includes: acquiring the D-channel image corresponding to the background image; inputting the D-channel image into a preset parameter extractor to obtain six parameters; performing an affine transformation on the background image based on the six parameters to achieve angle correction of the background image; the preset parameter extractor includes a MobileNet convolutional neural network with a 6-dimensional output layer; the affine transformation is used to transfer the original space (x,y) to the (u,v) space, and the transformation process uses a camera model to solve for the six parameters; the formula for transforming the (x,y) space to the (u,v) space is as follows: Where A is a 2*2 matrix and B is a 2*1 matrix, there are a total of 6 parameters; Feature extraction is performed on the background image after angle correction to obtain background image features; Background image recognition is performed based on the background image features.
2. The method according to claim 1, characterized in that, The steps to obtain a depth image of a target face include: The target face depth image is captured using the iOS front-facing camera.
3. The method according to claim 1, characterized in that, The training process of the foreground / background segmentation model is as follows: Obtain a training sample set; each sample in the training sample set is a standardized face depth image labeled with background image identifiers; The foreground-background segmentation model is obtained by using samples from the training sample set to train the preset semantic segmentation model at multiple scales.
4. The method according to claim 3, characterized in that, The steps to obtain the training sample set include: Acquire multiple initial face depth images; For each initial face depth image, the RGB image and D channel image in the initial face depth image are standardized to obtain a standard face depth image; Background image labeling is performed on each of the standard face depth images to obtain a training sample set.
5. The method according to claim 4, characterized in that, The step of standardizing the RGB image and D channel image in the initial face depth image to obtain a standard face depth image includes: For the RGB images in the initial face depth image, the mean and variance of the ImageNet image set are used for standardization. For the D channel image in the initial face depth image, a specified value is subtracted from the corresponding pixel value of the D channel to obtain the standard face depth image.
6. The method according to claim 1, characterized in that, The steps for background image recognition based on the background image features include: Calculate the similarity between the background image features and the black background features in the preset black background library. When the similarity exceeds a threshold, determine that the background image corresponding to the background image features is a suspicious black background.
7. An image background recognition device, characterized in that, The device includes: A depth image acquisition module is used to acquire a depth image of a target face; the depth image of the target face includes an RGB-D channel image; The background segmentation module is used to input the target face depth image into a preset foreground-background segmentation model and output the background image corresponding to the target face depth image; and to set each pixel value of the segmented foreground image to the same specified pixel value. The background correction module is used to correct the angle of the background image based on the depth information corresponding to the background image. This includes: acquiring the D-channel image corresponding to the background image; inputting the D-channel image into a preset parameter extractor to obtain six parameters; and performing an affine transformation on the background image based on the six parameters to achieve angle correction. The preset parameter extractor includes a MobileNet convolutional neural network with a 6-dimensional output layer. The affine transformation is used to transfer the original space (x,y) to the (u,v) space. The transformation process utilizes a camera model to solve for the six parameters. The formula for transforming the (x,y) space to the (u,v) space is as follows: Where A is a 2*2 matrix and B is a 2*1 matrix, there are a total of 6 parameters; The feature extraction module is used to extract features from the angle-corrected background image to obtain background image features; The background recognition module is used to perform background image recognition based on the background image features.
8. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the method of any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Foreground and background separation method based on RGB-D image, and foreground and background separation device
CN108447060A
Photo background similarity clustering method based on convolutional neural network and computer
CN110569878A
Target detection method and device based on depth image and storage medium
CN111178190A
Background extraction method with depth information and color information
CN111915687A