Image Processing Method, Apparatus, Electronic Device, and Computer-Readable Storage Medium
By identifying and filtering object poses in the image, and building a streamlined reference image collection, the redundancy and noise problems of image collection are solved, and the efficiency and accuracy of image processing are improved.
Patent Information
- Application Number
- CN202011074650.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-09
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-10-09
AI Technical Summary
The collected image collection has data redundancy problems, and some images have severe noise, which reduces the efficiency and accuracy of image processing applications.
By identifying the object poses in the image, adding the image to the corresponding pose set, and identifying the reference image from the pose set, building a streamlined reference image set, using the pose processing module and the integration module for image filtering and matching, and deleting redundant images and noise images.
It reduces the redundancy of the image set, improves the image quality of the image set, and improves the efficiency and accuracy of image processing applications.
Smart Images

Figure CN112132107B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to artificial intelligence technology, and in particular to an image processing method, device, electronic device, and computer-readable storage medium. Background Art
[0002] Artificial Intelligence (AI) is a comprehensive field of computer science that studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. With technological advancements, AI will be applied in more fields and play an increasingly important role.
[0003] Image processing is an important research direction in the field of artificial intelligence. Image processing refers to the technology of using computers to analyze images to achieve the desired results. It is widely used in various types of mutually beneficial network scenarios, such as social applications and online games.
[0004] However, during the implementation of the embodiments of the present application, the applicant discovered that the collected image set had data redundancy problems and some images had severe noise, which reduced the efficiency and accuracy of subsequent image processing applications based on this image set. Summary of the Invention
[0005] The embodiments of the present application provide an image processing method, apparatus, electronic device, and computer-readable storage medium, which can quickly and accurately extract images for subsequent applications from a collection of massive images.
[0006] The technical solution of the embodiment of the present application is implemented as follows:
[0007] The present invention provides an image processing method, including:
[0008] Recognizing an object pose in a plurality of images collected for the object, so as to add the plurality of images to an image set corresponding one-to-one to the plurality of object poses;
[0009] identifying a reference image of a first posture from a set of images of the first posture, wherein the first posture is any one of the plurality of object postures;
[0010] identifying a reference image of the second posture from the image set of the second posture based on a degree of matching between projections of the reference image of the first posture and images in the image set of the second posture;
[0011] The second posture is any one of the plurality of object postures that is different from the first posture;
[0012] Construct a reference image set of the object according to the reference image of the first pose and the reference image of the second pose.
[0013] An embodiment of the present application provides an image processing apparatus, including:
[0014] An identification module, configured to identify the object pose in a plurality of images collected for an object, so as to add the plurality of images to image sets corresponding to the plurality of object poses one by one;
[0015] A first pose processing module, configured to identify a reference image of the first pose from the image set of the first pose, where the first pose is any one of the plurality of object poses;
[0016] A second pose processing module, configured to identify a reference image of the second pose from the image set of the second pose according to the matching degree between the reference image of the first pose and the projection of the image in the image set of the second pose; where the second pose is any one of the plurality of object poses different from the first pose;
[0017] An integration module, configured to construct a reference image set of the object according to the reference image of the first pose and the reference image of the second pose.
[0018] In the above technical solution, the identification module is further configured to match each of the plurality of images collected for the object with a pose template respectively, so as to identify the object pose in each of the images; and add each of the images to the corresponding image set according to the object pose in each of the images.
[0019] In the above technical solution, the identification module is further configured to perform the following processing in real time for each collected image: determine key points in the image; determine a rotation value of the key points relative to corresponding points in the pose template; and determine the object pose in the image according to the rotation value.
[0020] In the above technical solution, the identification module is further configured to obtain a rotation angle interval corresponding to each object pose among the plurality of object poses; and determine the object pose corresponding to the rotation angle interval including the rotation value as the object pose in the image.
[0021] In the above technical solution, the identification module is further configured to determine a first distance between a key point representing the upper part of the eye and a key point representing the lower part of the eye among the key points; determine a second distance between a key point representing the left end of the eye and a key point representing the right end of the eye among the key points; and when the ratio of the first distance to the second distance is less than a closed-eye distance threshold, determine that the image has a closed-eye phenomenon and delete the image.
[0022] In the above technical solution, the image processing device further includes: a reminder module, configured to output information for reminding to re-collect images when the number of images in any one of the image sets is less than a number threshold, so as to collect images of the corresponding object posture.
[0023] In the above technical solution, the first posture processing module is further configured to determine the variance of each image in the image set of the first posture; determine the motion blur score of the corresponding image according to the variance of each image in the image set of the first posture; and use the image with the minimum motion blur score in the image set of the first posture as the reference image of the first posture.
[0024] In the above technical solution, the first posture processing module is further configured to perform the following processing on each image in the image set of the first posture: determine the grayscale image of the image; perform convolution processing on the grayscale image of the image to obtain the gradient image corresponding to the image; and determine the variance of the gradient image of the image and use it as the variance of the image.
[0025] In the above technical solution, the second posture processing module is further configured to determine the rigid inspection score of the images in the image set of the second posture according to the matching degree between the reference image of the first posture and the projections of the images in the image set of the second posture; delete the images in the image set of the second posture with a rigid inspection score less than an error threshold, and identify the reference image of the second posture from the image set of the second posture according to the rigid inspection score.
[0026] In the above technical solution, the second posture processing module is further configured to determine the variance of each image in the image set of the second posture; determine the motion blur score of the corresponding image according to the variance of each image in the image set of the second posture; and identify the reference image of the second posture according to the rigid inspection score and the motion blur score of each image in the image set of the second posture.
[0027] In the above technical solution, the second posture processing module is further configured to project the reference image of the first posture to obtain a reference three-dimensional image of the first posture; perform the following processing on each image in the image set of the second posture: project each image in the image set of the second posture to obtain a corresponding three-dimensional image of the second posture; determine the mapping relationship between the reference three-dimensional image of the first posture and the three-dimensional images of the second posture; and determine the matching degree between the reference image of the first posture and the projections of the images in the image set of the second posture according to the number of image points in the three-dimensional image of the second posture that satisfy the mapping relationship.
[0028] In the above technical solution, the second attitude processing module is further configured to determine the image depth information of the reference image of the first attitude; project the reference image of the first attitude according to the image depth information, the key points of the reference image of the first attitude, and the image acquisition parameters, so as to obtain a reference stereoscopic image of the first attitude.
[0029] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to implement the image processing method provided by the embodiment of the present application when executed.
[0030] The embodiment of the present application has the following beneficial effects:
[0031] By identifying the object attitude in the acquired image and adding the image to the image set corresponding to the attitude, the preliminary classification of the image is realized. Then, according to the reference images of multiple attitudes, a reference image set of the object is constructed, realizing fully automatic image screening. Moreover, a reference image is selected for each object attitude, and a refined reference image set is constructed based on this, reducing the redundancy of the image set, improving the picture quality in the image set, and enhancing the efficiency and accuracy of the image processing application implemented based on this reference image set. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a schematic structural diagram of an image processing system provided by an embodiment of the present application;
[0033] Figure 2 is a schematic structural diagram of an electronic device for image processing provided by an embodiment of the present application;
[0034] Figure 3 is a schematic flowchart of an image processing method provided by an embodiment of the present application;
[0035] Figure 4 is a schematic flowchart of an image processing method provided by an embodiment of the present application;
[0036] Figure 5 is a schematic diagram of facial image processing provided by an embodiment of the present application;
[0037] Figures 6A - 6B is a schematic diagram of image acquisition of an image processing method provided by an embodiment of the present application;
[0038] Figure 7 is a schematic diagram of object attitude grouping provided by an embodiment of the present application;
[0039] Figure 8 is a schematic diagram of facial key points provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To make the objectives, technical solutions and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0041] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0042] If similar descriptions such as "first / second" appear in the application documents, the following explanation shall be added. In the following description, the terms "first / second / third" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.
[0043] In the practical application of the relevant data collection and processing in the embodiments of this application, the informed consent or separate consent of the personal information subject should be obtained strictly in accordance with the requirements of relevant laws and regulations, and subsequent data use and processing behaviors should be carried out within the scope authorized by laws and regulations and the personal information subject.
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0045] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are subject to the following explanations.
[0046] 1) Image acquisition parameters: The parameters used in the process of projecting three-dimensional points in a scene onto a two-dimensional imaging plane by an image acquisition device to become image points. For example, the camera matrix used by a camera when taking a photo.
[0047] 2) Image depth information: Refers to the distance information from the image acquisition device to each point in the scene, reflecting the geometric shape of the visible surface of the object in the scene.
[0048] 3) Object: Refers to the target of image acquisition, such as a face, torso or limb.
[0049] 4) Object pose refers to the pose made by the target of image acquisition. For example, when the object of image acquisition is the face, the object pose can be a frontal face, face turning left, face turning right, or head raising; when the object of image acquisition is the torso, the object pose can be the torso standing upright, bending down, or taking a specific shape; when the object of image acquisition is a limb, the object pose can be the limb stretching, contracting, or taking a specific shape.
[0050] In the field of image processing, image datasets are of great value. However, constructing an image dataset with a large amount of data and high quality often requires a lot of manpower and material resources. Therefore, those skilled in the art tend to construct general image datasets for reuse in different application scenarios. Limited by this technical prejudice, those skilled in the art will try to retain a large amount of the original data collected. However, the applicant found during the implementation of the embodiments of the present application that there are a large number of data redundancy problems in the collected image set, and some images have serious noise, which reduces the efficiency and accuracy of image processing.
[0051] The embodiments of the present application provide an image processing method, apparatus, electronic device, and computer-readable storage medium, which can construct a refined reference image set, reduce the redundancy of the image set, improve the quality of the pictures in the image set, and improve the efficiency and accuracy of image processing applications implemented based on this reference image set.
[0052] The following describes the exemplary applications of the electronic device for image processing provided in the embodiments of the present application. The device provided in the embodiments of the present application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), etc., such as handheld terminals. It performs image acquisition for an object and recognizes the object postures in the multiple acquired images, adds the multiple images to an image set corresponding one-to-one to the multiple object postures, and respectively identifies reference images from the image sets of the first posture and the second posture to construct a reference image set of the object, so as to perform subsequent object reconstruction, expression basis construction, adding posture special effects, etc. based on the reference image set; it can also be implemented as a server or a server cluster, such as a server deployed in the cloud, which recognizes the object postures in the multiple images acquired for the object, adds the multiple images to an image set corresponding one-to-one to the multiple object postures, and respectively identifies reference images from the image sets of the first posture and the second posture to construct a reference image set of the object, so as to perform subsequent object reconstruction, expression basis construction, adding posture special effects, etc. based on the reference image set; it can also be implemented in a manner of cooperation between a user terminal and a server, such as cooperative processing between a handheld terminal and a cloud server. The handheld terminal performs image acquisition for the object and recognizes the object postures in the multiple acquired images, adds the multiple images to an image set corresponding one-to-one to the multiple object postures, and then the handheld terminal sends the image set to the cloud server. The cloud server respectively identifies reference images from the image sets of the first posture and the second posture to construct a reference image set of the object, performs subsequent object reconstruction, expression basis construction, adding posture special effects, etc. based on the reference image set, and feeds back the reconstructed object, the constructed expression basis, and the posture special effects to the handheld terminal.
[0053] Exemplarily, after the electronic terminal filters the images to obtain reference images of multiple postures, it can construct a set of drivable expression bases of the object according to the reference images of multiple postures, and in combination with a rendering engine, by setting different expression coefficients, make the expression bases generate different expression actions.
[0054] Exemplarily, after the electronic terminal recognizes the object postures in different images, it can use the images of different object postures as a training set and the object postures as annotations to train a posture recognition model. Through the posture recognition model, it can real-time recognize the corresponding object postures that appear in social applications and attach different special effects to enhance the fun of the application.
[0055] Exemplarily, when the object recognized by the electronic terminal is a face, 3D face reconstruction can be performed based on face reference images in different poses to meet the interactive entertainment needs in different scenarios. For example, in scenarios such as games and publicity, 3D models of human faces can be accurately reconstructed to build realistic game character attributes. In daily conversation scenarios, the reconstructed 3D face models can be used to customize exclusive emoticons to generate various subtle expressions. The reconstructed 3D face models can also be used in a wide range of application scenarios such as virtual makeup, virtual try-on, and virtual character images.
[0056] Next, exemplary applications when the electronic device is implemented as a server will be described.
[0057] Referring to Figure 1 , Figure 1 which is a schematic structural diagram of an image processing system provided by an embodiment of the present application. Taking the image processing system 100 as an example, to support an image processing application, terminals (exemplarily shown as terminals 400-1 and 400-2) are connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two. The server 200 stores data in the database 500, and the server 600 can obtain data from the database 500. Among them, the server 200 is a server for implementing the image processing method of the embodiment of the present application, and the server 600 is a server for performing specific applications such as face reconstruction and expression basis construction based on the selected images.
[0058] In some embodiments, terminals (exemplarily shown as terminals 400-1 and 400-2) can collect object images through applications on the terminals (exemplarily shown as applications 410-1 and 410-2), such as applications like cameras, social media, short videos, and video live broadcasts, and upload them to the server 200 for implementing the image processing method of the embodiment of the present application through the network 300. The server 200 can identify the object poses in the images, add the images to the image sets corresponding to the poses, then construct a reference image set of the object based on the reference images of multiple poses, and then store the reference image set in the database 500. When performing processing such as expression basis construction, adding pose special effects, and face reconstruction, the server 600 can obtain and use the stored reference image set from the database 500.
[0059] In some embodiments, the server 200 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminals (exemplarily shown as terminals 400-1 and 400-2) may be smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, etc., but are not limited thereto. The terminals and the server may be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.
[0060] See Figure 2 , Figure 2 is a schematic structural diagram of an electronic device for image processing provided by an embodiment of the present application. Taking the electronic device as a server as an example for illustration, Figure 2 The server 200 shown in the figure includes: at least one processor 210, a memory 250, and at least one network interface 220. Each component in the server 200 is coupled together through a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 240.
[0061] The processor 210 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0062] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 250 optionally includes one or more storage devices that are physically located far from the processor 210.
[0063] The memory 250 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM, Read Only Memory), and the volatile memory may be a random access memory (RAM, Random Access Memory). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.
[0064] In some embodiments, the memory 250 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are exemplarily described below.
[0065] The operating system 251 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0066] The network communication module 252 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.;
[0067] In some embodiments, the image processing apparatus provided by the embodiments of the present application can be implemented in software. Figure 2 Shown is an image processing apparatus 255 stored in the memory 250, which can be software in the form of programs and plugins, etc., including the following software modules: an identification module 2551, a first pose processing module 2552, a second pose processing module 2553, and an integration module 2554. These modules are logical, so they can be arbitrarily combined or further split according to the functions to be implemented. The functions of each module will be described below.
[0068] The image processing method provided by the embodiments of the present application can be provided as a cloud service. Any application (such as a social network application) can submit the image data collected in the application to the cloud service provider, and the cloud service performs image processing.
[0069] Next, the image processing method provided by the embodiments of the present application will be described. As mentioned above, the electronic device implementing the image processing method of the embodiments of the present application can be a terminal, a server, or a combination of both. Therefore, the execution subject of each step will not be repeated hereinafter.
[0070] It should be noted that in the examples of image processing hereinafter, the object is taken as the face for illustration. Those skilled in the art can apply the image processing method provided by the embodiments of the present application to the processing of an image set including other types of objects according to the understanding of the following text.
[0071] See Figure 3 , Figure 3 which is a schematic flowchart of the image processing method provided by the embodiments of the present application, and will be described in combination with the steps shown in Figure 3 shown below.
[0072] In step 101, the object poses in multiple images collected for the object are identified, so as to add the multiple images to image sets corresponding one by one to the multiple object poses.
[0073] Among them, the image can be a photo or a video frame. When the image is a photo, the image set can be continuously taken photos; when the image is a video frame, the multiple collected images can be video frames decoded from the captured video. The format of the image can be a three-channel (RGB, Red-Green-Blue) image. The RGB image obtains various colors through the changes of the three color channels of red (R), green (G), and blue (B) and their superposition with each other.
[0074] Exemplarily, when collecting images for the object, the object can be made to make various poses, and multiple images are taken for each pose. The object poses in the images are identified according to the image data, and the images are classified according to the identified object poses, that is, the images are added to the image sets corresponding to the object poses in the images. Finally, image sets corresponding one by one to multiple different object poses are formed, realizing the content classification of the images and improving the efficiency of subsequent use of the collected images.
[0075] In addition, after the terminal (exemplarily showing terminal 400-1 and terminal 400-2) collects pictures, it can choose to perform the above processing locally on the terminal instead of sending the collected images to the server 200. The object poses in the images are identified locally on the terminal, and the multiple images are added to image sets corresponding one by one to multiple different object poses to speed up the image processing speed and avoid the influence brought by network latency.
[0076] In some embodiments, referring to Figure 4 , Figure 4 shows a schematic flowchart of an image processing method provided by an embodiment of the present application. Figure 3 Step 101 in
[0077] can be implemented through steps 1011-1012. In step 1011, multiple images collected for the object are respectively matched with pose templates to identify the object poses in each image; in step 1012, according to the object poses in each image, each image is added to the image set corresponding to the object pose. The following is a specific description.
[0078] Exemplarily, in order to respectively match multiple images collected for an object with a pose template to identify the object pose in each image, the following processing can be performed in real time for each collected image: determining key points in the image; determining the rotation value of the key points relative to the corresponding points in the pose template; and determining the object pose in the image according to the rotation value.
[0079] After an image of the object is collected, key points of the image are determined in the image. The key points are points that can characterize the features of the object. Taking the object as the face as an example, the key points may include: face contour points, points representing the shapes of the five sense organs (eyebrows, eyes, ears, nose, mouth). The method for determining the key points is not limited in the embodiments of the present application. Subsequently, the corresponding points in the pose template to the key points are determined, and then the Perspective-n-Point (PnP) algorithm can be used to obtain the rotation value of the key points relative to the corresponding points in the pose template, and the object pose in the image is determined according to the rotation value.
[0080] Exemplarily, the object pose in the image can be determined according to the rotation value by obtaining the rotation angle interval corresponding to each object pose among multiple object poses; and determining the object pose corresponding to the rotation angle interval including the rotation value as the object pose in the image. When the rotation value in the image does not satisfy the rotation angle interval of any object, this image is deleted.
[0081] For example, after obtaining the rotation value of the key points relative to the corresponding points in the pose template, the rotation angle interval corresponding to each object pose among multiple object poses is obtained. For example, the rotation angle interval of object pose A is rotating 10 - 50 degrees around the z-axis, the rotation angle interval of object pose B is rotating -10 to -50 degrees around the z-axis, and the rotation angle interval of object pose C is rotating 10 to -10 degrees around the y-axis. When the rotation value of the image is rotating 25 degrees around the z-axis, the image is added to the image set of object pose A. When the rotation value of the image does not belong to any of the above three intervals, the image is deleted to save storage space.
[0082] By matching the image with the pose template in real time when collecting the image to identify the object pose in each image, on the one hand, the efficiency of image processing can be improved through real-time processing; on the other hand, by matching with the pose template, the object pose in the image can be identified simply and quickly, and there is no need to pre-train and deploy an identification model in advance, reducing the cost and being easy to implement.
[0083] In some other embodiments, in order to add multiple images to an image set corresponding to multiple object poses one by one, the multiple collected images may be classified respectively by a pre-trained machine learning model to determine the object pose in each image; and according to the object pose in each image, each image is added to the image set corresponding to the object pose.
[0084] Exemplarily, a machine learning model for object pose classification is pre-trained using an image data set. Among them, the samples in the image data set are object images labeled with object poses. When the image data set of a single object is insufficient, the images and labels of multiple objects can be jointly used to form the image data set for training. And the machine learning model for classifying images can be a logistic regression, a support vector machine, a neural network or an ensemble learning model.
[0085] By using a pre-trained machine learning model to identify the object pose in an image, a suitable classification model can be selected and fully trained to improve the accuracy of object pose recognition.
[0086] In some embodiments, when the object is a face, before determining the rotation value of the key points relative to the corresponding points in the pose template, it further includes: determining a first distance between the key points representing the upper part of the eyes and the key points representing the lower part of the eyes among the key points; determining a second distance between the key points representing the left end of the eyes and the key points representing the right end of the eyes among the key points; when the ratio of the first distance to the second distance is less than the closed-eye distance threshold, it is determined that the image shows a closed-eye phenomenon, and the image is deleted.
[0087] When the object is a face, if the object has a closed-eye behavior when the image is collected, the collected image cannot reflect the key features of the object and belongs to an invalid image, which needs to be identified and deleted as early as possible. For this purpose, a first distance between the key points representing the upper part of the eyes and the key points representing the lower part of the eyes among the key points of the image can be determined. The first distance represents the width when the eyes are open. For example, the key point representing the upper eyelid can be determined as the key point representing the upper part of the eyes, and the key point representing the lower eyelid can be determined as the key point representing the upper part of the eyes to determine the first distance; then, a second distance between the key points representing the left end of the eyes and the key points representing the right end of the eyes among the key points is determined. The second distance represents the length of both ends of the eyes; when the ratio of the first distance to the second distance is less than the closed-eye distance threshold, it indicates that the object in the image has a closed-eye phenomenon, and thus the image with the closed-eye phenomenon is deleted.
[0088] When the object is a face, by detecting and deleting the images with closed-eye behaviors, invalid images can be filtered out as early as possible, the redundancy of the image set can be reduced, and the efficiency of image processing can be improved.
[0089] In some embodiments, after identifying the object poses in multiple images collected for an object and adding the multiple images to image sets corresponding one-to-one to the multiple object poses, the method further includes: when the number of images in any image set is less than a quantity threshold, outputting information for prompting to re-collect images so as to collect images of the corresponding object pose.
[0090] Since the object pose may not be standard when collecting images, for example, the object closes its eyes, the rotation value is greater than the rotation angle range, etc., resulting in the number of images in the image set being less than the quantity threshold after adding the images to the image sets corresponding one-to-one to the multiple object poses. For example, the quantity threshold for each group of image sets is 15. When the number of images in the image set is less than 15, it cannot be ensured that valuable images of the object pose corresponding to this image set are collected. Therefore, information for prompting to re-collect images is output so as to collect images of the corresponding object pose. For example, the object can be prompted to re-collect images of the corresponding pose through applications (exemplarily shown as applications 410-1 and 410-2) on terminals (terminals 400-1 and 400-2).
[0091] By detecting the image sets with the number of images less than the quantity threshold and prompting the object to re-collect images of the corresponding pose, each image set of each object pose has a certain number of images, so as to ensure that high-quality images can be selected for each object pose.
[0092] In step 102, a reference image of the first pose is identified from the image set of the first pose, where the first pose is any one of the object poses.
[0093] After adding the multiple images to the image sets corresponding one-to-one to the multiple object poses, from the multiple image sets, the image set corresponding to the first pose is selected. The first pose can be the pose with the highest usage frequency, such as a frontal face image, a standing pose image, etc. Then, a reference image is selected from the image set corresponding to the first pose. The reference image is one or more images with the highest image quality in the image set corresponding to the first pose. By pre-selecting the image with the highest quality corresponding to the first pose, the speed and efficiency in using the image are improved.
[0094] In some embodiments, referring to Figure 4 , Figure 4 FIG. shows a schematic flowchart of an image processing method provided by an embodiment of the present application. Figure 3Step 102 in [the above] can be implemented through steps 1021 - 1023. In step 1021, the variance of each image in the image set of the first pose is determined; in step 1022, according to the variance of each image in the image set of the first pose, the motion blur score of the corresponding image is determined; in step 1023, the image with the minimum motion blur score in the image set of the first pose is used as the reference image of the first pose.
[0095] Among them, the motion blur score characterizes the degree of blurriness of the image. The lower the motion blur score of the image, the clearer the image; the higher the motion blur score of the image, the blurrier the image.
[0096] Exemplarily, the following processing can be performed on each image in the image set of the first pose: determining the grayscale image of the image; performing convolution processing on the grayscale image of the image to obtain the gradient image corresponding to the image; determining the variance of the gradient image of the image and using it as the variance of the image.
[0097] For example, by weighting the channels of the acquired image, the grayscale image of the image is obtained, then the Laplace operator is used to perform convolution processing on the grayscale image to obtain the processed gradient image, and then the variance of all gradients in the gradient image is obtained. The reciprocal of the variance can be used as the motion blur score of the corresponding image. Among them, the larger the variance of the image, the clearer the image boundary, and the lower the motion blur score of the image; the smaller the variance of the image, the blurrier the image boundary, and the higher the motion blur score of the image.
[0098] Judging the clarity of the image boundary through the variance of the image and selecting the reference image accordingly ensures the clarity of the reference image and improves the quality of the reference image set.
[0099] In step 103, according to the matching degree of the projection of the reference image of the first pose and the images in the image set of the second pose, the reference image of the second pose is identified from the image set of the second pose; where the second pose is any one of the multiple object poses different from the first pose.
[0100] In some embodiments, referring to Figure 4 , Figure 3 Step 103 in [the above] can be implemented through steps 1031 - 1033. In step 1031, according to the matching degree of the projection of the reference image of the first pose and the images in the image set of the second pose, the rigidity test score of the images in the image set of the second pose is determined; in step 1032, the images in the image set of the second pose with a rigidity test score less than the error threshold are deleted, and in step 1033, according to the rigidity test score, the reference image of the second pose is identified from the image set of the second pose. The following is a specific description.
[0101] Exemplarily, in order to identify a reference image of the second pose from the set of images of the second pose according to the rigidity test score, the variance of each image in the set of images of the second pose can be determined; according to the variance of each image in the set of images of the second pose, the motion blur score of the corresponding image can be determined; according to the rigidity test score and the motion blur score of each image in the set of images of the second pose, the reference image of the second pose can be identified.
[0102] For example, after determining the matching degree between the projection of the reference image of the first pose and the images in the set of images of the second pose, the rigidity test score of the images in the set of images of the second pose can be determined according to the matching degree. The higher the matching degree between the image and the reference image of the first pose, the higher the rigidity test score; the lower the matching degree between the image and the reference image of the first pose, the lower the rigidity test score. In order to make the selection of the reference image of the second pose more accurate, the motion blur score of each image in the set of images of the second pose can be determined. The method for determining the motion blur score is the same as that of the reference image of the first pose, and by combining the rigidity test score and the motion blur score of each image, the reference image of the second pose can be identified. For example, through the difference between the rigidity test score and the motion blur score, since the higher the rigidity test score, the higher the quality of the image, and the lower the motion blur score, the clearer the image, therefore, the greater the difference between the rigidity test score and the motion blur score, the higher the quality of the image; it is also possible to first filter out the images with obvious errors using the rigidity test score, and then select the reference image of the second pose through the motion blur score.
[0103] By combining the rigidity test score and the motion blur score of each image in the set of images of the second pose to identify the reference image of the second pose, both the number of valid points in the image and the clarity of the image are considered, effectively improving the accuracy of identifying the reference image of the second pose.
[0104] In some embodiments, in order to determine the matching degree between the projection of the reference image of the first pose and the images in the set of images of the second pose, before determining the rigidity test score of the images in the set of images of the second pose according to the matching degree between the projection of the reference image of the first pose and the images in the set of images of the second pose, it further includes: projecting the reference image of the first pose to obtain a reference three-dimensional image of the first pose; performing the following processing for each image in the set of images of the second pose: projecting each image in the set of images of the second pose to obtain a corresponding three-dimensional image of the second pose; determining the mapping relationship between the reference three-dimensional image of the first pose and the three-dimensional image of the second pose; and determining the matching degree between the projection of the reference image of the first pose and the images in the set of images of the second pose according to the number of image points in the three-dimensional image of the second pose that satisfy the mapping relationship.
[0105] Among them, in order to project the reference image in the first pose to obtain the reference stereo image in the first pose, the image depth information of the reference image in the first pose can be determined; according to the image depth information, the key points of the reference image in the first pose, and the image acquisition parameters, the reference image in the first pose is projected to obtain the reference stereo image in the first pose.
[0106] Among them, in order to project each image in the image set in the second pose to obtain the corresponding stereo image in the second pose, the image depth information of the image in the second pose can be determined; according to the image depth information, the key points of the reference image in the second pose, and the image acquisition parameters, the image in the second pose is projected to obtain the stereo image in the second pose.
[0107] For example, taking the object as the face, assuming that the object poses are four groups: frontal face, left side, right side, and looking up. Among them, the first pose is the frontal face, and the left side, right side, and looking up are all the second poses. When the reference image of the frontal face is recognized, the reference image of the frontal face can be projected into a stereo image according to the key points of the reference image of the frontal face, the image depth information of the reference image of the frontal face, and the image acquisition parameters, that is, the reference stereo image in the first pose. Then, the second pose stereo images of each image in the three groups of object poses of the left side, right side, and looking up can be determined in the same way. Suppose there is a point A in the stereo projection image of the reference image of the frontal face, and there is a point B in the stereo projection image of the sample image in the image set corresponding to the left side. In the real scene, point A and point B correspond to the same point on the object's face. Then, there is a rotation and translation relationship between point A and point B. Point B can coincide with point A through a certain rotation and translation. That is, point A and point B can establish the following equation:
[0108] A = R * B + T
[0109] Among them, R is the rotation coefficient and T is the translation coefficient.
[0110] Through multiple pairs of corresponding points similar to A and B in the stereo projection image of the reference image of the frontal face and the stereo projection image of the sample image, the above equation is fitted to determine the values of R and T, and then the mapping relationship between the reference stereo image of the frontal face and the stereo image of the sample image is determined. Then, according to the number of image points that satisfy the above equation, the matching degree of the projection between the sample image and the reference image of the frontal face is determined. For example, the number of key points of the reference image of the frontal face and the sample image is 100. Among them, the number of points that satisfy the above equation after fitting is 80, then the matching degree is 0.8. Then, the same processing is performed on each image in the image set corresponding to the left side to obtain the matching degree of each image with the reference image of the frontal face.
[0111] By obtaining the relational equation between the first pose reference image and the second pose image, the number of image points that can satisfy the relational equation between the first pose reference image and the second pose image can be clearly determined. The image points that satisfy the relational equation can be regarded as valid points in the image. Furthermore, the matching degree of the projection of the reference image of the first pose and the images in the image set of the second pose can be determined according to the number of image points that satisfy the relational equation, improving the accuracy and resolvability of the matching degree.
[0112] In step 104, according to the reference image of the first pose and the reference image of the second pose, a reference image set of the object is constructed.
[0113] By adding the reference images of each object pose, a reference image set of the object is constructed, which can be used for image processing applications in various scenarios. For example, 3D face reconstruction is performed based on the face reference images of different object poses to meet the interactive entertainment needs in different scenarios. In scenarios such as games and publicity, the 3D model of the human face can be accurately reconstructed to build realistic game character attributes. In the daily conversation scenario, the reconstructed 3D face model can be used to customize exclusive emoticons and expression bases to generate various subtle expressions and enhance the fun of chatting. The reconstructed 3D face model can also be used in a wide range of entertainment scenarios such as virtual makeup, virtual try-on, and virtual character images.
[0114] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.
[0115] Applications in the user terminal (such as social applications, short video applications, etc.) can collect images of different facial poses of the user through the terminal camera, and screen the collected images of different facial poses to obtain high-quality images (reference images) corresponding to each pose. According to the reference images of multiple facial poses, different expression actions of the user can be generated, such as opening the mouth, closing the mouth, blinking, etc., so that the user can use the generated simulated emoticons in social chats, and can also prompt the user to share the generated action expressions with friends and on the network platform.
[0116] See Figure 5 , Figure 5 shows a schematic diagram of facial image processing provided by the embodiments of the present application. The image processing method provided by the embodiments of the present application can be divided into two stages: 501, the rough screening stage, and 502, the fine screening stage. The processing is carried out from rough screening to fine screening. The rough screening stage can be processed in real time when collecting images, and the fine screening stage is processed offline after the image collection is completed. The following will be introduced separately.
[0117] See Figures 6A - 6B , Figure 6AThe following is a schematic diagram of image acquisition for the image processing method provided by an embodiment of the present application. Assume that a mobile phone capable of acquiring image depth information is used as the data acquisition device. When starting to acquire data, the user's face is facing the front camera. As shown in Figure 6A 604, the front camera can clearly display the entire face area. Then, turn the head in the order of turning left (as shown in Figure 6A 603), turning back to the middle, turning right (as shown in Figure 6A 602), turning back to the middle, and looking up (as shown in Figure 6A 601), while trying to keep the facial expression unchanged. Refer to Figure 6B . Figure 6B The following is a schematic diagram of image acquisition for the image processing method provided by an embodiment of the present application. After acquisition, save the sequence of depth images (RGBD, Red-Green-Blue + Depth Map) of the face during the entire head rotation process. Exemplarily, the acquired RGB images are as shown in Figure 6B 605 - 618, and the acquired image depth information is as shown in Figure 6B 619 - 632.
[0118] In the rough screening stage, while acquiring RGBD images, perform the following operations on each image: Step A: Detect facial key points (landmark). Step B: Perform closed-eye screening to remove the images with closed eyes. Step C: Calculate the rotation value and translation value relative to the preset general stereo template points through the PnP method. Step D: Group the current images according to the rotation value.
[0119] According to the requirements of subsequent image processing applications, set 4 object pose groups: front face, left side, right side, and looking up. Each group corresponds to an image set. Refer to Figure 7 . Figure 7 The following is a schematic diagram of object pose grouping provided by an embodiment of the present application. The judgment criterion is that the upward direction of the human head is the z-axis, and the front face is facing the x-axis. Then, a rotation of 10 - 50 degrees around the z-axis is considered the left side, a rotation of -10 to -50 degrees around the z-axis is considered the right side, and a rotation between 10 and -10 degrees around the y-axis is considered the front face. The number of divided groups can be adaptively determined according to actual needs. Among them, the images with rotation values that are too large are directly deleted. If the number of images in a certain group is too small (less than 3 images), it may be that the user did not take this angle, or all the images with closed eyes were deleted during the whole process. Then the entire data is invalid, and the user is prompted to reshoot.
[0120] Among them, refer to Figure 8 . Figure 8 The following is a schematic diagram of facial key points provided by an embodiment of the present application. Landmark can be used to detect closed eyes. Figure 8102 key points are marked out as shown.
[0121] According to the coordinate relationships among the face landmark points numbered 52, 53, 54, 55, 56, 57, 58, and 59, determine whether there is a closed-eye phenomenon in the left eye according to the following formula:
[0122]
[0123] where d left represents the ratio between the distance between the landmark points on the upper and lower eyelids of the left eye and the distance between the landmark points at both ends of the left eye corner, and l i (i = 52, 53, 54, 55, 56, 57, 58, 59) represents the coordinates of the landmark numbered i, and D threshold is the set closed-eye distance threshold. When d left is less than this threshold, it is determined that there is a closed-eye phenomenon.
[0124] Similarly, for the coordinate relationships among the face landmark points numbered 61, 62, 63, 64, 65, 66, and 67, the formula for determining whether there is a closed-eye phenomenon in the right eye is as follows:
[0125]
[0126] where d right represents the ratio between the distance between the landmark points on the upper and lower eyelids of the right eye and the distance between the landmark points at both ends of the right eye corner, and l i (i = 60, 61, 62, 63, 64, 65, 66, 67) represents the coordinates of the landmark numbered i, and D threshold is the set closed-eye distance threshold. When d right is less than this threshold, it is determined that there is a closed-eye phenomenon.
[0127] In the fine screening stage, the purpose of the fine screening stage is to select the best-quality image, that is, the reference image, from each of the four groups of object postures: frontal face, left side, right side, and looking up. This step can be carried out offline after the entire image acquisition is completed.
[0128] Step A: First, select the reference image of the frontal face. Perform motion blur scoring and sorting on each image in the frontal face group, and select the image with the smallest motion blur score as the reference image of the frontal face (the reference image of the first posture).
[0129] Step B: After selecting the reference image of the frontal face, for each group of the left and right side faces and the face looking up, select the image with the best quality. For all the images in each group, perform a rigid test with the reference image of the frontal face respectively, delete the images with obvious errors, calculate the motion blur score for the remaining images, and select the image with the best quality (the reference image of the second pose) by combining the scores of the rigid test and the motion blur score.
[0130] When the number of pictures required for subsequent applications exceeds 4, the top N images (N is a positive integer greater than 1) with the best comprehensive scores can be selected from the four groups of the frontal face, the left side, the right side, and the face looking up respectively, or the angles can be divided in each group. For example, the above method can be used to select the image with the best quality every 10 degrees.
[0131] Among them, the method for determining the motion blur score is as follows. For a color image, first convert it into a grayscale image, and then perform edge detection. Here, the Laplace-Gaussian method (Log, laplacian with Gaussian) can be directly used for edge detection. Calculate the variance of the image to obtain the blur value. Since it is difficult to have an accurate standard for judging whether an image is blurred or clear, it can be judged by a relative value. For example, the blurrier an image is, the blurrier its edges are, that is, the smaller the variance of the image, and the higher the motion blur score.
[0132] Among them, the method for rigid test is as follows. Taking the left side as an example, for the corresponding depth map, image landmark, and camera parameters, the corresponding stereo points of the image can be back-projected. Back-project the stereo points of all the images corresponding to the frontal face frame and the left side. The stereo point A of the frontal face frame and the stereo point B on the left side face. Since they correspond to the same point on the face, that is, there is a rotation and translation relationship between them, the following equation can be constructed: A = R * B + T, where R is the rotation coefficient R and T is the translation coefficient. For the projected stereo points of the 102 landmark points marked, the above formula is used to calculate by means of Random Sample Consensus (RAN SAC), so as to obtain the rotation coefficient R and the translation coefficient T. At the same time, the inliers can be screened out, and the number of inliers is recorded. The image with the largest number of inliers has the highest score in the rigid test.
[0133] After filtering the images to obtain the reference images of multiple facial poses, a reference image set can be constructed based on the reference images of multiple facial poses, and a set of drivable facial expression bases can be reconstructed based on the reference image set of the face. By combining with a rendering engine and setting different expression coefficients, different expression actions can be generated by the expression bases, such as actions of opening the mouth, closing the mouth, blinking, etc., so that users can use the generated simulated emoticons in social chats, such as customizing exclusive emoticons and adding expression special effects to the chat. It can also encourage users to share the generated action expressions with friends and on network platforms, and can also be used in a wide range of scenarios such as virtual makeup, virtual try-on, and virtual character images.
[0134] Next, the exemplary structure of the implementation of the image processing device 255 provided in the embodiments of the present application as a software module will be further described. In some embodiments, as Figure 2 shown, the software module in the image processing device 255 stored in the memory 250 may include:
[0135] An identification module 2551, configured to identify the object poses in a plurality of images collected for an object, so as to add the plurality of images to an image set corresponding to the plurality of object poses one by one;
[0136] A first pose processing module 2552, configured to identify the reference image of the first pose from the image set of the first pose, where the first pose is any one of the plurality of object poses;
[0137] A second pose processing module 2553, configured to identify the reference image of the second pose from the image set of the second pose according to the matching degree between the reference image of the first pose and the projection of the image in the image set of the second pose; where the second pose is any one of the plurality of object poses different from the first pose;
[0138] An integration module 2554, configured to construct a reference image set of the object according to the reference image of the first pose and the reference image of the second pose.
[0139] In some embodiments, the identification module is further configured to match each of the plurality of images collected for the object with a pose template to identify the object pose in each of the images; and add each of the images to an image set corresponding to the object pose according to the object pose in each of the images.
[0140] In some embodiments, the identification module is further configured to perform the following processing in real time for each collected image: determine the key points in the image; determine the rotation value of the key points relative to the corresponding points in the pose template; and determine the object pose in the image according to the rotation value.
[0141] In some embodiments, the recognition module is further configured to obtain a rotation angle interval corresponding to each object pose among the multiple object poses; and determine the object pose corresponding to the rotation angle interval including the rotation value as the object pose in the image.
[0142] In some embodiments, the recognition module is further configured to determine a first distance between a key point representing the upper part of the eye and a key point representing the lower part of the eye among the key points; determine a second distance between a key point representing the left end of the eye and a key point representing the right end of the eye among the key points; and when the ratio of the first distance to the second distance is less than a closed-eye distance threshold, determine that the image has a closed-eye phenomenon and delete the image.
[0143] In some embodiments, the image processing apparatus further includes: a reminder module 2555 ( Figure 2 not shown in the figure), which is configured to output information for prompting to re-acquire images when the number of images in any one of the image sets is less than a number threshold, so as to acquire images of the corresponding object pose.
[0144] In some embodiments, the first pose processing module is further configured to determine the variance of each image in the image set of the first pose; determine the motion blur score of the corresponding image according to the variance of each image in the image set of the first pose; and use the image with the minimum motion blur score in the image set of the first pose as the reference image of the first pose.
[0145] In some embodiments, the first pose processing module is further configured to perform the following processing on each image in the image set of the first pose: determine the grayscale image of the image; perform convolution processing on the grayscale image of the image to obtain the gradient image corresponding to the image; and determine the variance of the gradient image of the image and use it as the variance of the image.
[0146] In some embodiments, the second pose processing module is further configured to determine the rigidity test score of the images in the image set of the second pose according to the matching degree between the reference image of the first pose and the projections of the images in the image set of the second pose; delete the images in the image set of the second pose with a rigidity test score less than an error threshold, and identify the reference image of the second pose from the image set of the second pose according to the rigidity test score.
[0147] In some embodiments, the second pose processing module is further configured to determine the variance of each image in the image set of the second pose; determine the motion blur score of the corresponding image according to the variance of each image in the image set of the second pose; and identify the reference image of the second pose according to the rigidity test score and the motion blur score of each image in the image set of the second pose.
[0148] In some embodiments, the second pose processing module is further configured to project the reference image of the first pose to obtain a reference stereo image of the first pose; perform the following processing on each image in the image set of the second pose: project each image in the image set of the second pose to obtain a corresponding second pose stereo image; determine the mapping relationship between the reference stereo image of the first pose and the second pose stereo image; and determine the matching degree between the projection of the reference image of the first pose and the image in the image set of the second pose according to the number of image points in the second pose stereo image that satisfy the mapping relationship.
[0149] In some embodiments, the second pose processing module is further configured to determine the image depth information of the reference image of the first pose; and project the reference image of the first pose according to the image depth information, the key points of the reference image of the first pose, and the image acquisition parameters to obtain a reference stereo image of the first pose.
[0150] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to implement the image processing method provided by the embodiment of the present application when executed.
[0151] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0152] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0153] As an example, the executable instructions may or may not correspond to files in a file system, and may be stored as part of a file that holds other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program under discussion, or in multiple cooperating files (such as files that store one or more modules, subroutines, or code portions).
[0154] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one site, or on multiple computing devices distributed across multiple sites and interconnected via a communication network.
[0155] In summary, the embodiments of the present application have the following beneficial technical effects:
[0156] (1) By classifying the content of an image through identifying the object pose and selecting the reference images corresponding to each object pose to construct a reference image set, a fully automatic construction of the image set is achieved, without any manual participation, reducing the cost of image processing.
[0157] (2) By identifying the object pose in real time during acquisition and deleting images with poses not conforming to the rotation angle range and behaviors such as closing eyes, the speed of image processing is accelerated, and the efficiency of image processing is also improved.
[0158] (3) By first selecting the reference images of the first pose and then selecting the reference images of the second pose based on the reference images of the first pose, the accuracy of reference image selection is improved, and the computational amount during reference image selection is reduced at the same time.
[0159] (4) In the embodiments of the present application, the number of reference images for each object pose can be adaptively set, which can be coupled with the subsequent application requirements based on the reference image set, improving the intelligent level of image processing and also improving the efficiency of image processing.
[0160] The above description is only for the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. An image processing method, characterized in that, Including: Identifying the object postures in a plurality of images collected for an object, so as to add the plurality of images to an image set corresponding one-to-one to a plurality of object postures; Identifying a reference image of the first posture from the image set of the first posture, wherein the first posture is any one of the plurality of object postures; Projecting the reference image of the first posture to obtain a reference stereoscopic image of the first posture; Performing the following processing on each image in the image set of the second posture: Projecting each image in the image set of the second posture to obtain a corresponding stereoscopic image of the second posture; Determining a mapping relationship between the reference stereoscopic image of the first posture and the stereoscopic images of the second posture corresponding to each image in the image set of the second posture; Determining the matching degree of the projection of the reference image of the first posture and the images in the image set of the second posture according to the number of image points in the stereoscopic image of the second posture that satisfy the mapping relationship; Determining a rigidity test score of the images in the image set of the second posture according to the matching degree of the projection of the reference image of the first posture and the images in the image set of the second posture; Deleting the images in the image set of the second posture with the rigidity test score less than an error threshold, and identifying a reference image of the second posture from the image set of the second posture according to the rigidity test score; Wherein the second posture is any one of the plurality of object postures different from the first posture; Constructing a reference image set of the object according to the reference image of the first posture and the reference image of the second posture.
2. The method according to claim 1, characterized in that, The identifying the object postures in a plurality of images collected for an object, so as to add the plurality of images to an image set corresponding one-to-one to a plurality of object postures, includes: Matching the plurality of images collected for the object with a posture template respectively to identify the object postures in each of the images; Adding each of the images to a corresponding image set according to the object posture in each of the images.
3. The method according to claim 2, wherein The matching the plurality of images collected for the object with a posture template respectively to identify the object postures in each of the images includes: Performing the following processing on each collected image: Determining key points in the image; Determining a rotation value of the key points relative to the corresponding points in the posture template; Determining the object posture in the image according to the rotation value.
4. The method according to claim 3, characterized in that, The determining the object posture in the image according to the rotation value includes: Obtaining a rotation angle interval corresponding to each object posture among the plurality of object postures; Determining the object posture corresponding to the rotation angle interval including the rotation value as the object posture in the image.
5. The method according to claim 3, characterized in that, When the object is a face, before the determining the rotation value of the key points relative to the corresponding points in the posture template, it further includes: Determining a first distance between a key point representing the upper part of the eyes and a key point representing the lower part of the eyes among the key points; Determining a second distance between a key point representing the left end of the eyes and a key point representing the right end of the eyes among the key points; When the ratio of the first distance to the second distance is less than the eye-closure distance threshold, it is determined that the image exhibits an eye-closure phenomenon, and the image is deleted.
6. The method according to any one of claims 1 to 5, characterized in that, After identifying the object postures in the multiple images collected for the object to add the multiple images to the image sets corresponding to the multiple object postures one by one, it further includes: When the number of images in any one of the image sets is less than the number threshold, information for prompting to re-collect images is output to collect images of the corresponding object posture.
7. The method according to any one of claims 1 to 5, characterized in that The identifying the reference image of the first posture from the image set of the first posture includes: Determining the variance of each image in the image set of the first posture; Determining the motion blur score of the corresponding image according to the variance of each image in the image set of the first posture; Taking the image with the minimum motion blur score in the image set of the first posture as the reference image of the first posture.
8. The method according to claim 7, characterized in that, The determining the variance of each image in the image set of the first posture includes: Performing the following processing on each image in the image set of the first posture: Determining the grayscale image of the image; Performing convolution processing on the grayscale image of the image to obtain the gradient image corresponding to the image; Determining the variance of the gradient image of the image and taking it as the variance of the image.
9. The method according to claim 1, wherein The identifying the reference image of the second posture from the image set of the second posture according to the rigid inspection score includes: Determining the variance of each image in the image set of the second posture; Determining the motion blur score of the corresponding image according to the variance of each image in the image set of the second posture; Identifying the reference image of the second posture according to the rigid inspection score and the motion blur score of each image in the image set of the second posture.
10. The method according to claim 1, characterized in that, The projecting the reference image of the first posture to obtain the reference stereoscopic image of the first posture includes: Determining the image depth information of the reference image of the first posture; Projecting the reference image of the first posture according to the image depth information, the key points of the reference image of the first posture, and the image acquisition parameters to obtain the reference stereoscopic image of the first posture.
11. An image processing apparatus, characterized in that, It includes: An identification module, configured to identify the object postures in the multiple images collected for the object to add the multiple images to the image sets corresponding to the multiple object postures one by one; A first posture processing module, configured to identify the reference image of the first posture from the image set of the first posture, where the first posture is any one of the multiple object postures; A second pose processing module, configured to project the reference image of the first pose to obtain a reference stereo image of the first pose; perform the following processing on each image in the image set of the second pose: project each image in the image set of the second pose to obtain a corresponding second pose stereo image; determine a mapping relationship between the reference stereo image of the first pose and the second pose stereo image corresponding to each image in the image set of the second pose; determine the matching degree of the projection of the reference image of the first pose and the images in the image set of the second pose according to the number of image points in the second pose stereo image that satisfy the mapping relationship; determine the rigidity test score of the images in the image set of the second pose according to the matching degree of the projection of the reference image of the first pose and the images in the image set of the second pose; delete the images in the image set of the second pose with the rigidity test score less than the error threshold, and identify the reference image of the second pose from the image set of the second pose according to the rigidity test score; wherein, the second pose is any one of the multiple object poses different from the first pose. An integration module, configured to construct a reference image set of the object according to the reference image of the first pose and the reference image of the second pose.
12. An electronic device, characterized in that, Comprising: A memory, configured to store computer-executable instructions; A processor, when executing the computer-executable instructions stored in the memory, implements the image processing method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, Stored with computer-executable instructions, the computer-executable instructions, when executed, are used to implement the image processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Face image processing method and apparatus
CN105975935A
Image processing method, image processing device and terminal equipment
CN110084765A