Image generation method, apparatus, and electronic device
By generating facial images in different poses through image transformation processing, the problems of high cost and long time consumption in collecting training datasets are solved, the number of samples in the training dataset is increased, and the performance of the face recognition model is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Z-ONE TECH CO LTD
- Filing Date
- 2023-10-08
- Publication Date
- 2026-04-21
AI Technical Summary
Existing methods for collecting training datasets for human and animal face recognition are costly and time-consuming, making it difficult to efficiently acquire large-scale training data.
By determining the facial parameters in the first image and the parameters of a preset standard facial model, image transformation processing is performed to generate second images in different poses, thereby expanding the number of samples in the training dataset and reducing the need for manual data collection.
This effectively increased the number of samples in the training dataset for the face recognition model, reduced the cost and time of training dataset collection, and improved the performance of the face recognition model.
Smart Images

Figure CN117292424B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image generation method, apparatus and electronic device. Background Technology
[0002] Traditional face recognition methods rely on large-scale training datasets. Obtaining a good face recognition model often requires millions of face images to form the training dataset. Large manufacturers need to invest significant manpower and resources to collect millions of face data points, download, process, and label these data (e.g., face images) from the internet, and then integrate them into the training dataset for the face recognition model. Processing such a massive dataset is a difficult and expensive task. Therefore, existing methods for collecting training datasets for face recognition suffer from high costs and long processing times. Furthermore, similar problems exist in methods for collecting training datasets for other animal face recognition. Summary of the Invention
[0003] This application provides an image generation method, apparatus, and electronic device to solve the problems of high cost and long time consumption in existing methods for collecting training datasets for face recognition.
[0004] To address the aforementioned technical problems, in a first aspect, embodiments of this application provide an image generation method, the method comprising: determining a first image, the first image including a first face; obtaining a second image based on the first image, first parameters corresponding to the first face, and second parameters corresponding to a preset standard face model, the second image including a second face, the pose of the second face being different from the pose of the first face, the first parameters including first key point information of a first key point included in the first face, and pose information of the first face, and the second parameters including second key point information of a second key point included in the standard face model.
[0005] In this implementation, after determining the first image, image transformation processing is performed on the first image based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to the preset standard face model. This yields a processed second image, in which the pose of the second face differs from that of the first face. Therefore, by transforming the first image including the first face, second images corresponding to the same face in different poses can be obtained. In other words, two images are obtained from one image, and the poses of the faces in the two images are different. This increases the number of samples in the training dataset corresponding to the face recognition model. Since it is not necessary to manually collect every sample in the training dataset, the cost and time of collecting the training dataset are effectively reduced.
[0006] In one possible implementation of the first aspect described above, the first image and the second image are two-dimensional images. The second image is obtained based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to the preset standard face model. This includes: obtaining a third image based on the first key point information, the standard face model, and the second key point information, wherein the third image is a two-dimensional image; obtaining a fourth image based on the third image and pose information, wherein the fourth image includes the third face and is a two-dimensional image; and correcting the position of the third key point based on the third key point information of the third key point included in the third face and the index information corresponding to the third key point, thereby obtaining the second image.
[0007] In the implementation of this application, during the process of obtaining the second image from the first image, the first image is first converted into a two-dimensional third image, then the third image is converted into a two-dimensional fourth image. Finally, the second image is obtained based on the third key point information and corresponding index information of the third key points included in the third face in the fourth image. During the image conversion process, the first image undergoes multiple transformation and correction processes, making the information of the second face included in the obtained second image more accurate, facilitating the training of the face recognition model, and increasing the performance of the obtained face recognition model.
[0008] In one possible implementation of the first aspect above, the method further includes: determining the side to be repaired corresponding to the second face based on the cumulative amount of pixels of the second face; and performing symmetrical processing on the side to be repaired according to a preset symmetry weight to obtain a fifth image.
[0009] In this implementation, the second face included in the second image is symmetrically processed to obtain the fifth image. Thus, based on the second image, a fifth image with different facial information than the second face is obtained. That is, three images (including the first, second, and third images) are obtained from one image, and the poses and facial information included in the three images are all different. This effectively increases the number of samples in the training dataset corresponding to the face recognition model, and effectively reduces the cost and time of collecting the training dataset.
[0010] In one possible implementation of the first aspect above, obtaining a fourth image based on the third image and the pose information includes: obtaining a sixth image based on the third image and the pose information; determining the point to be corrected in the sixth image with an incorrect projection position; adjusting the position of the point to be corrected to the correct position corresponding to the point to be corrected, thereby obtaining the fourth image.
[0011] In this implementation, a sixth image is obtained from the third image. Then, points with incorrect projection positions on the sixth image are corrected, resulting in a fourth image. The fourth image contains more accurate information about the third face. Therefore, the second image obtained from the fourth image also contains more accurate information about the second face, facilitating the training of the face recognition model and increasing its performance.
[0012] In one possible implementation of the first aspect described above, obtaining a third image based on the first keypoint information, the standard face model, and the second keypoint information includes: obtaining a camera intrinsic matrix, a camera projection matrix, a rotation matrix, and a translation vector based on the first and second keypoint information; obtaining a model view matrix and an image projection matrix based on the camera intrinsic matrix, the camera projection matrix, the rotation matrix, and the translation vector; performing a first image transformation process on the standard face model based on the model view matrix and the image projection matrix to obtain the third image; and obtaining a sixth image based on the third image and pose information includes: performing a second image transformation process on the third image based on the pose information to obtain the sixth image.
[0013] In this implementation, the model-view matrix and image projection matrix obtained from the camera intrinsic matrix, camera projection matrix, rotation matrix, and translation vector are used to perform a first image transformation on the standard face model, ensuring the accuracy and completeness of the information in the resulting third image. Therefore, based on the pose information, a second image transformation is performed on the third image, resulting in a sixth image with even higher accuracy and completeness. This makes the information of the second face included in the final second image more accurate, facilitating the training of the face recognition model and increasing its performance.
[0014] In one possible implementation of the first aspect described above, the sixth image includes a fourth face and a background. Determining points on the sixth image whose projection positions are incorrect and need correction, and adjusting the positions of these points to their correct positions to obtain the fourth image, includes: determining a background template; separating projection points corresponding to the fourth face and the background based on the background template; determining a first point among the projection points corresponding to the fourth face whose projection position is incorrect and adjusting the position of this first point to its correct position to obtain a corrected fourth face; determining a second point among the projection points corresponding to the background whose projection position is incorrect and adjusting the position of this second point to its correct position to obtain a corrected background; and obtaining the fourth image based on the corrected fourth face and the corrected background.
[0015] In this implementation, the projection points corresponding to the fourth face and the projection points corresponding to the background region are separated based on the background template. Then, the first point to be corrected due to an incorrect projection position in the projection points corresponding to the fourth face and the second point to be corrected due to an incorrect projection position in the projection points corresponding to the background region are corrected respectively, resulting in a fourth image. This method can more accurately locate and correct the points to be corrected in the fourth face, and the information of the third face included in the resulting fourth image is also more accurate. Therefore, the information of the second face included in the second image obtained from the fourth image is also more accurate, facilitating the training of the face recognition model and increasing the performance of the obtained face recognition model.
[0016] In one possible implementation of the first aspect above, determining a first point to be corrected with an incorrect projection position among the projection points corresponding to the fourth face, and adjusting the position of the first point to be corrected to the correct position corresponding to the first point to be corrected, includes: if it is determined that the first point to be corrected with an incorrect projection position among the projection points corresponding to the fourth face is located outside the boundary of the image region corresponding to the fourth face, then normalizing the first point to be corrected to adjust the position of the first point to be corrected to the correct position corresponding to the first point to be corrected.
[0017] In the implementation of this application, the position of the first point to be corrected is corrected by normalization, so that the corrected position of the first point to be corrected is more accurate.
[0018] In one possible implementation of the first aspect above, the first key point information and pose information are obtained by inputting the first image into a preset recognition model for recognition processing to obtain the first key point information and pose information.
[0019] In this application, the first image is input into a preset recognition model for image recognition processing, which can obtain the first key point information and posture information more quickly and accurately.
[0020] Secondly, embodiments of this application provide an image generation apparatus for executing the image generation method described in the first aspect above. The apparatus includes: a first processing module for determining a first image, the first image including a first face; and a second processing module for obtaining a second image based on the first image, first parameters corresponding to the first face, and second parameters corresponding to a preset standard face model. The second image includes a second face, the pose of the second face being different from that of the first face. The first parameters include first key point information of first key points included in the first face and pose information corresponding to the first face. The second parameters include second key point information of second key points included in the preset standard face model.
[0021] Thirdly, embodiments of this application provide an electronic device, including: a memory for storing a computer program, the computer program including program instructions; and a processor for executing the program instructions to cause the electronic device to perform the image generation method provided by the first aspect and / or any possible implementation of the first aspect.
[0022] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, the computer program including program instructions, which are executed by an electronic device to perform the image generation method provided by the first aspect and / or any possible implementation of the first aspect.
[0023] Fifthly, embodiments of this application provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the image generation method provided by the first aspect and / or any possible implementation of the first aspect.
[0024] The beneficial effects of this application are:
[0025] The image generation method provided in this application, after determining a first image, performs image transformation processing on the first image based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to a preset standard face model, to obtain a processed second image. The pose of the second face in the second image is different from that of the first face. Therefore, by transforming the first image including the first face, it is possible to obtain second images corresponding to the same face in different poses, i.e., obtaining two images from one image, with different facial poses in the two images. This effectively increases the number of samples in the training dataset corresponding to the face recognition model. Since it is not necessary to manually collect every sample in the training dataset, it effectively reduces the cost and time of training dataset collection. Attached Figure Description
[0026] To more clearly illustrate the technical solution of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0027] Figure 1 This is a flowchart illustrating an image generation method based on some implementations of this application;
[0028] Figure 2 This is a schematic diagram illustrating a process for obtaining a second image, based on some implementations of this application.
[0029] Figure 3 This is a schematic diagram illustrating a process for obtaining a fourth image, based on some implementations of this application.
[0030] Figure 4 This is a schematic diagram illustrating a process for obtaining a third image, based on some implementations of this application.
[0031] Figure 5 This is a schematic diagram illustrating another process for obtaining the fourth image, based on some implementations of this application;
[0032] Figure 6 This is a flowchart illustrating another image generation method according to some implementations of this application;
[0033] Figure 7 This is a flowchart illustrating another image generation method according to some implementations of this application;
[0034] Figure 8 This is a schematic diagram of the structure of an image generation apparatus according to some implementations of this application;
[0035] Figure 9 This is a schematic diagram of the structure of an electronic device according to some implementations of this application. Detailed Implementation
[0036] The technical solution of this application will be described in further detail below with reference to the accompanying drawings.
[0037] As mentioned earlier, traditional face recognition methods rely on large-scale training datasets, and acquiring face images requires a significant investment of time and resources. Existing methods for acquiring training datasets for face recognition suffer from high costs and long processing times.
[0038] Based on this, this application provides an image generation method, apparatus, and electronic device that can increase the number of samples in the training dataset corresponding to the face recognition model, effectively reduce the cost and time of training dataset collection, and solve the problems of high cost and long time consumption in the collection method of training dataset for face recognition.
[0039] Next, please refer to the appendix. Figure 1 -Appendix Figure 7 This application provides a detailed explanation of the specific process and advantages of the image generation method provided.
[0040] The image generation method provided in this application is applied to electronic devices, such as computers, servers, and cluster servers, which are capable of image processing. In one implementation of this application, the image processing method, as follows... Figure 1 As shown, it includes the following steps:
[0041] S100: Determine a first image, the first image including a first face.
[0042] Specifically, the first image is a sample image that can be used for a face recognition model. Since face recognition models can be used for human face recognition, such as mobile phone screen recognition and access control recognition, as well as for animal type recognition, the first image needs to include a first face. The first face can be a side view image, a front view image, or an image from any angle of the object being captured.
[0043] S200: Obtain the second image based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to the preset standard face model.
[0044] The second image includes a second face, the pose of which is different from that of the first face. The first parameter includes the first key point information of the first key point included in the first face, as well as the pose information of the first face. The second parameter includes the second key point information of the second key point included in the standard face model.
[0045] The facial pose information in this application, also known as the Euler angles of the face, is based on a standard spatial coordinate system and includes vertical flip angles (i.e., pitch angle - rotation around the X-axis, pitch), horizontal flip angles (i.e., yaw angle - rotation around the Y-axis, raw), and in-plane rotation angles (i.e., roll angle - rotation around the Z-axis, roll), which correspond to the angles of head tilting up, tilting down, and turning.
[0046] The first keypoints, including the first face, are points located corresponding to key regions of the face, such as eyebrows, nose, mouth, and facial contours, obtained by keypoint detection on the first image including the first face. The information of the first keypoints can include their coordinates, number, and index. The index here refers to the label of the first keypoint. For example, if there are a total of 64 first keypoints, their labels would be 1-64, with keypoints for the eyes labeled 1-10, keypoints for the nose labeled 11-20, and so on. It should be noted that this application does not limit the number of keypoints or their corresponding indices; the above description is merely illustrative.
[0047] The second key point corresponding to the preset standard face model (a 3D face model) in this application corresponds one-to-one with the first key point included in the first face. It is also the point corresponding to the key area of the face. The number of key points and the position represented by the key points are the same.
[0048] The first key point information of the first face, including the first key point, and the pose information corresponding to the first face, can be identified and detected by a preset recognition model. In one implementation of this application, the first key point information and pose information are obtained by inputting the first image into a preset recognition model for image recognition processing to obtain the first key point information and pose information.
[0049] The preset recognition model can be a preset deep learning model, such as a landmark and pose model.
[0050] The image generation method provided in this application, after determining a first image, performs image transformation processing on the first image based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to a preset standard face model, to obtain a processed second image. The pose of the second face in the second image is different from that of the first face. Therefore, by transforming the first image including the first face, it is possible to obtain second images corresponding to the same face in different poses, i.e., obtaining two images from one image, with different facial poses in the two images. This effectively increases the number of samples in the training dataset corresponding to the face recognition model. Since it is not necessary to manually collect every sample in the training dataset, it effectively reduces the cost and time of training dataset collection.
[0051] Next, a detailed explanation will be given of step S200, which involves obtaining the specific content of the second image based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to the preset standard face model.
[0052] In one implementation of this application, the first image and the second image are two-dimensional images, such as... Figure 2 As shown, the second image is obtained based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to the preset standard face model, including the following steps:
[0053] S210: Based on the first key point information, the standard face model, and the second key point information, a third image is obtained. The third image is a two-dimensional image.
[0054] S220: Based on the third image and pose information, obtain the fourth image. The fourth image includes the third face and is a two-dimensional image.
[0055] S230: Based on the information of the third key point and the index information corresponding to the third key point included in the third face, the position of the third key point is corrected to obtain the second image.
[0056] The index of a key point refers to its label. In this application, the third key points included in the third face correspond to the first key points included in the first face. That is, the number of third key points included in the third face and the location of the area represented by key points with the same label should be the same. For example, in the first face, the key points of the eyes are labeled 1-10, and the key points of the nose are labeled 11-20. In the third face, the key points of the eyes should also be labeled 1-10, and the key points of the nose should also be labeled 11-20. Based on the third key point information and the corresponding index information of the third key points included in the third face, the position of the third key points is corrected. That is, firstly, based on the label and position of the third key points included in the third face, it is determined whether the projection position of the key points is incorrect. For example, the key points in the eye area are numbered 1-10, but key points numbered 5 and 6 are projected onto the nose in the third face. Therefore, it is necessary to determine whether the projection position of the key points is incorrect based on the index information of the key points known in advance. If it is incorrect, the position of the third key point is corrected, and the positions of key points numbered 5 and 6 are corrected to the correct position on the eyes to obtain an accurate fourth image.
[0057] When converting the first image to the second image, the standard face model is first converted into a third image, then the third image is converted into a two-dimensional fourth image. Finally, based on the third keypoint information and corresponding index information of the third keypoint included in the third face in the fourth image, the second image is obtained. During the image conversion process, the first image undergoes multiple transformation and correction processes, making the information of the second face in the resulting second image more accurate, facilitating the training of the face recognition model, and increasing the performance of the resulting face recognition model.
[0058] Furthermore, in one implementation of this application, during the acquisition of part of the training dataset, ordinary geometric transformations (such as mutual projection between two-dimensional and three-dimensional coordinate points) are performed on the acquired original images to increase the number of training dataset samples. However, images obtained through ordinary geometric transformations suffer from severe distortion, resulting in poor performance of the face recognition model. Compared to this method, the image generation method described above performs multiple transformations and corrections on the first image, making the second image containing more accurate information about the second face more suitable for training the face recognition model and increasing the performance of the obtained face recognition model.
[0059] In one implementation of this application, such as Figure 3 As shown, the fourth image is obtained based on the third image and pose information, including the following steps:
[0060] S221: Based on the third image and the pose information, obtain the sixth image.
[0061] S222: Determine the points on the sixth image whose projection positions are incorrect and need to be corrected.
[0062] S223: Adjust the position of the point to be corrected to the correct position corresponding to the point to be corrected, and obtain the fourth image.
[0063] After image conversion, due to insufficient computing power of electronic devices or errors during projection, the standard facial model may exhibit incorrect projection positions in the third image. Consequently, when the third image is transformed into the sixth image, points with incorrect projection positions will appear in the sixth image. Therefore, it is necessary to correct the positions of these incorrectly projected points in the sixth image and adjust them to their correct locations to obtain a clearer and more accurate fourth image.
[0064] The sixth image is derived from the third image. Then, points with incorrect projection positions on the sixth image are corrected, resulting in the fourth image. The fourth image contains more accurate information about the third face. Therefore, the second image derived from the fourth image also contains more accurate information about the second face, facilitating the training of the face recognition model and increasing its performance.
[0065] The first keypoint information and the second keypoint information can specifically be the coordinates of the first keypoint and the coordinates of the second keypoint. In one implementation of this application, a third image is obtained based on the first keypoint information, the standard face model, and the second keypoint information, such as... Figure 4 As shown, it includes the following steps:
[0066] S211: Based on the information of the first key point and the second key point, obtain the camera intrinsic parameter matrix, camera projection matrix, rotation matrix and translation vector.
[0067] S212: Based on the camera intrinsic parameter matrix, camera projection matrix, rotation matrix, and translation vector, obtain the model view matrix and image projection matrix.
[0068] S213: Based on the model view matrix and the image projection matrix, perform the first image transformation process on the standard facial model to obtain the third image.
[0069] The camera intrinsic matrix and camera projection matrix refer to the parameter matrices of the camera used to capture the first image. Based on the camera intrinsic matrix, camera projection matrix, rotation matrix, and translation vector, image transformations can be performed.
[0070] In one implementation of this application, obtaining a sixth image based on a third image and pose information includes: performing a second image transformation process on the third image based on the pose information to obtain the sixth image.
[0071] Specifically, as described above, the camera intrinsic matrix, camera projection matrix, rotation matrix, and translation vector are first obtained based on the first and second key point information. Then, the model view matrix and image projection matrix are calculated based on the camera intrinsic matrix, camera projection matrix, rotation matrix, and translation vector. Finally, the third image is transformed using the model view matrix, image projection matrix, and pose information to obtain a two-dimensional sixth image.
[0072] By using the camera intrinsic matrix, camera projection matrix, rotation matrix, and translation vector, the resulting model view matrix and image projection matrix are used to perform a first image transformation on the standard face model, ensuring the accuracy and completeness of the information in the resulting third image. Therefore, based on the pose information, a second image transformation is performed on the third image, resulting in a sixth image with even higher accuracy and completeness. This makes the information in the second face included in the final second image more accurate, facilitating the training of the face recognition model and increasing its performance.
[0073] In one implementation of this application, the sixth image includes a fourth face and a background. Points with incorrect projection positions on the sixth image are identified, and their positions are adjusted to the correct positions to obtain the fourth image. Figure 5 As shown, it includes the following steps:
[0074] S240: Determine the background template.
[0075] S250: Based on the background template, separate the projection points corresponding to the fourth face and the projection points corresponding to the background.
[0076] S260: Determine the first point to be corrected in the projection point corresponding to the fourth face that has an incorrect projection position, adjust the position of the first point to be corrected to the correct position corresponding to the first point to be corrected, and obtain the corrected fourth face.
[0077] S270: Determine the second point to be corrected in the projection point corresponding to the background part where the projection position is incorrect, adjust the position of the second point to be corrected to the correct position corresponding to the second point to be corrected, and obtain the corrected background part.
[0078] S280: Based on the corrected fourth face and the corrected background, obtain the fourth image.
[0079] A background template, also known as a background mask, is a standard background for an image. It can be designed using existing graphic design techniques. The background template in this application is a standard background template obtained from the sixth image, whose pixels match the pixels of the background area in the sixth image. Therefore, based on the background template, the projection points corresponding to the fourth face and the background can be accurately separated.
[0080] Steps S260 and S270 correct the positions of incorrectly positioned projection points in the projection points corresponding to the separated fourth face and background, respectively. Therefore, steps S260 and S270 do not have a fixed order; S260 can be executed first, followed by S270, or vice versa, or both can be executed simultaneously. Finally, the corrected fourth face and the corrected background are fused together to obtain a more accurate fourth image.
[0081] Based on the background template, the projection points corresponding to the fourth face and the projection points corresponding to the background region are separated. Then, the first point to be corrected with an incorrect projection position in the projection points corresponding to the fourth face and the second point to be corrected with an incorrect projection position in the projection points corresponding to the background are corrected respectively, resulting in the fourth image. This method can more accurately find and correct the points with incorrect projection positions in the fourth face, and the information of the third face included in the fourth image is also more accurate. Therefore, the information of the second face included in the second image obtained from the fourth image is also more accurate, which facilitates the training of the face recognition model and increases the performance of the obtained face recognition model.
[0082] In one implementation of this application, determining a first point to be corrected with an incorrect projection position among the projection points corresponding to the fourth face, and adjusting the position of the first point to be corrected to the correct position, includes: if it is determined that the first point to be corrected with an incorrect projection position among the projection points corresponding to the fourth face is located outside the boundary of the image area corresponding to the fourth face, then normalizing the first point to be corrected to adjust the position of the first point to be corrected to the correct position, so that the corrected position of the first point to be corrected is more accurate.
[0083] For points on the background that are projected in the wrong position, correction can be made by normalization or by other existing algorithms for correcting the position of projection points. This application will not elaborate on these methods.
[0084] In one implementation of this application, after step S200, as follows: Figure 6 As shown, it also includes the following steps:
[0085] S300: Determine the side of the second face to be repaired based on the accumulated amount of pixels of the second face.
[0086] S400: Based on the preset symmetry weights, the side to be repaired is symmetrically processed to obtain the fifth image.
[0087] In the process of symmetrical processing of the face, existing symmetrical processing methods can be used. Specifically, existing software can be used to perform symmetrical processing of the facial region in the image, such as AutoCAD. If the symmetrical option is enabled and an eye mask (i.e., a standard symmetrical eye template) is provided, soft symmetrical operation is performed. Based on the cumulative amount of visible pixels on the face (i.e., the second face), the side of the face that is more obscured is determined, and symmetrical processing is applied to use the visible part from the obscured side to symmetrically process the face, resulting in a fifth image with a symmetrical face. The soft symmetrical process provided in this application is as follows: based on the face image in the image, the side of the face image that is more obscured and the side with more dense facial pixels are determined. After determining the side of the face image that is more obscured and the side with more dense facial pixels, the pixels of the side with more dense facial pixels are symmetrically projected onto the side of the face image that is more obscured, resulting in a corresponding symmetrical face image.
[0088] The second face in the second image is symmetrically processed to obtain the fifth image. Thus, based on the second image, a fifth image with different facial information is obtained. In other words, three images (including the first, second, and fifth images) are obtained from one image, and the facial poses and facial information in the three images are all different. This effectively increases the number of samples in the training dataset corresponding to the face recognition model and effectively reduces the cost and time of collecting the training dataset.
[0089] In one implementation of this application, the image generation method is a fast and convenient way to obtain rich scene face data, such as... Figure 7 As shown, it includes the following steps:
[0090] The system uses cameras and other devices to capture images containing faces (i.e., the first image). The captured images are then input into the landmark and pose model. The landmark and pose model are used to obtain the landmark and pose labels of the captured images (i.e., the first key point information and pose information, the first key point information including the coordinates of the first key point). The landmark and pose model uses the optimal landmark and pose model in the target scene.
[0091] The coordinates of the second keypoint of the provided 3D model (i.e., the standard face model) are determined. Based on the coordinates of the first and second keypoints, the camera's projection matrix, intrinsic parameter matrix, rotation matrix, and translation vector are calculated. Using the camera's projection matrix, intrinsic parameter matrix, rotation matrix, and translation vector, the model-view matrix and projection matrix are calculated. Then, using the model-view matrix and projection matrix, and based on the vertex coordinates of the 3D model, the 3D model (i.e., the 3D face coordinates) is transformed to obtain the third image. Finally, based on the third image and pose information, the third image is transformed into a two-dimensional sixth image.
[0092] Next, the background mask (i.e., background template) corresponding to the third image is calculated, and the face projection points and background projection points of the sixth image are separated based on the background mask.
[0093] The system determines whether the projection positions of the face projection points are correct. If incorrect, it processes the positions of erroneous projection points located outside and inside the face region image, respectively. Specifically, erroneous projection points outside the face region image are normalized, their positions are adjusted according to certain rules, and finally, the points are remapped to the original image size. Then, points with incorrect projection positions in the background are processed according to the established rules. After processing both the erroneous projection points on the face and in the background, the fourth image is obtained.
[0094] Based on the information of the third key point corresponding to the third key point in the fourth image and the index corresponding to the third key point, the projection point (i.e., key point) and the index are combined to correct the position of the key point with incorrect projection position in the fourth image (i.e., to distort the fourth image) and obtain the second image.
[0095] The second image includes a face that is symmetrically processed by applying soft symmetry to create a frontal view. The results are returned as a fifth image, including frontal views with and without symmetry, face points within the image, points displayed outside the image, and symmetry weights.
[0096] Finally, the first, second, and fifth images are all output as samples for the training dataset used to train the face recognition model.
[0097] The image generation method provided in this application improves face recognition performance compared to traditional data augmentation methods. This method matches the performance of acquiring large datasets of millions of internet images by introducing facial appearance variations through synthetic data. Unlike general data augmentation methods, this method utilizes domain-specific techniques to generate synthetic images. By using synthetic data augmentation, the time and cost of collecting, processing, and labeling large numbers of images can be reduced, while simultaneously improving the accuracy of face recognition systems.
[0098] In one implementation of this application, an image generation apparatus is provided, such as... Figure 8 As shown, the device includes:
[0099] A first processing module is used to determine a first image, the first image including a first face.
[0100] The second processing module is used to obtain a second image based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to the preset standard face model. The second image includes a second face, the pose of which is different from that of the first face. The first parameters include the first key point information of the first key point included in the first face and the pose information of the first face. The second parameters include the second key point information of the second key point included in the standard face model.
[0101] For details on the specific operations that each processing module can perform, please refer to the above. Figure 1 The corresponding image generation method. Furthermore, based on the specific operation steps of the above image generation method, the image generation device may include more or fewer processing modules for processing the content in the above image generation method.
[0102] Please see Figure 9 , Figure 9 The diagram shown is a structural block diagram of an electronic device provided in this application. The electronic device may include one or more processors 1002, system control logic 1008 connected to at least one of the processors 1002, system memory 1004 connected to the system control logic 1008, non-volatile memory (NVM) 1006 connected to the system control logic 1008, and network interface 1010 connected to the system control logic 1008.
[0103] Processor 1002 may include one or more single-core or multi-core processors. Processor 1002 may include any combination of general-purpose processors and special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In the implementation herein, processor 1002 may be configured to perform the aforementioned image generation method.
[0104] In some implementations, system control logic 1008 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1002 and / or any suitable device or component communicating with system control logic 1008.
[0105] In some implementations, system control logic 1008 may include one or more memory controllers to provide an interface to system memory 1004. System memory 1004 may be used to load and store data and / or instructions. In some implementations, system memory 1004 of the electronic device may include any suitable volatile memory, such as suitable dynamic random access memory (DRAM).
[0106] NVM / memory 1006 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some implementations, NVM / memory 1006 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as at least one of a hard disk drive (HDD), a compact disc (CD) drive, and a digital versatile disc (DVD) drive.
[0107] NVM / Memory 1006 may include a portion of storage resources installed on an electronic device, or it may be accessible by the device, but is not necessarily part of the device. For example, NVM / Memory 1006 may be accessed over a network via network interface 1010.
[0108] Specifically, system memory 1004 and NVM / memory 1006 may each include a temporary copy and a permanent copy of instruction 1020. Instruction 1020 may include instructions that, when executed by at least one of processors 1002, cause the electronic device to perform the aforementioned image generation method. In some implementations, instruction 1020, hardware, firmware, and / or its software components may additionally / alternatively be located in system control logic 1008, network interface 1010, and / or processor 1002.
[0109] Network interface 1010 may include a transceiver for providing a radio interface for an electronic device, enabling communication with any other suitable device (such as a front-end module, antenna, etc.) via one or more networks. In some implementations, network interface 1010 may be integrated into other components of the electronic device. For example, network interface 1010 may be integrated into at least one of processor 1002, system memory 1004, NVM / memory 1006, and firmware device (not shown) with instructions that, when at least one of processor 1002 executes the instructions, enable the electronic device to implement the aforementioned image generation method.
[0110] The network interface 1010 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 1010 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.
[0111] In one implementation, at least one of the processors 1002 may be packaged together with the logic of one or more controllers for system control logic 1008 to form a system-in-a-package (SiP). In another implementation, at least one of the processors 1002 may be integrated on the same die with the logic of one or more controllers for system control logic 1008 to form a system-on-a-chip (SoC).
[0112] The electronic device may further include an input / output (I / O) device 1012. The I / O device 1012 may include a user interface enabling a user to interact with the electronic device; the peripheral component interface is designed to allow peripheral components to also interact with the electronic device. In some implementations, the electronic device may also include sensors for determining at least one type of environmental condition and location information relevant to the electronic device.
[0113] In some implementations, the user interface may include, but is not limited to, a display (e.g., a liquid crystal display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., a light-emitting diode flash), and a keyboard.
[0114] In some implementations, the peripheral component interface may include, but is not limited to, non-volatile memory ports, audio jacks, and power interfaces.
[0115] In some implementations, the sensors may include, but are not limited to, gyroscope sensors, accelerometers, proximity sensors, ambient light sensors, and positioning units. The positioning unit may also be part of or interact with network interface 1010 to communicate with components of the positioning network (e.g., Global Positioning System (GPS) satellites).
[0116] It is understood that the illustrative structures of the embodiments of the present invention do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0117] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.
[0118] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this paper are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0119] One or more aspects of at least one implementation can be implemented by representational instructions stored on a computer-readable storage medium, the instructions representing various logics in a processor, which, when read by a machine, cause the machine to create logic for performing the techniques described herein. These representations, referred to as “IP cores,” can be stored on tangible computer-readable storage media and made available to multiple customers or production facilities for loading into manufacturing machines that actually manufacture the logic or processor.
[0120] It should be noted that some structural or methodological features may be shown in the accompanying drawings in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some implementations, these features may be arranged in a different manner and / or order than those shown in the illustrative drawings. Furthermore, including structural or methodological features in a particular figure does not imply that such features are required in all implementations, and in some implementations, these features may be omitted or may be combined with other features.
[0121] It should be noted that the terms "first" and "second" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0122] It should be noted that some structural or methodological features may be shown in the accompanying drawings in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, including structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.
[0123] Although this application has been illustrated and described with reference to certain preferred embodiments, those skilled in the art should understand that the above description is a further detailed explanation of the application in conjunction with specific embodiments, and should not be construed as limiting the specific implementation of the application to these descriptions. Those skilled in the art can make various changes in form and detail, including some simple deductions or substitutions, without departing from the spirit and scope of this application.
Claims
1. An image generation method characterized by, The method includes: A first image is determined, the first image including a first face; A second image is obtained based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to a preset standard face model. The second image includes a second face, the pose of which differs from that of the first face. The first parameters include first key point information of the first key points included in the first face, and pose information of the first face. The second parameters include second key point information of the second key points included in the standard face model. The first image and the second image are two-dimensional images. The second image, obtained based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to the preset standard face model, includes: A third image is obtained based on the first key point information, the standard face model, and the second key point information. The third image is a two-dimensional image. Based on the pose information of the third image and the first face, a fourth image is obtained, the fourth image including the third face, and the fourth image is a two-dimensional image; Based on the third key point information of the third key point included in the third face and the index information corresponding to the third key point, the position of the third key point is corrected to obtain the second image.
2. The image generation method according to claim 1, characterized by, The method further includes: Based on the cumulative amount of pixels in the second face, determine the side of the second face that needs to be repaired; Based on preset symmetry weights, the side to be repaired is symmetrically processed to obtain the fifth image.
3. The image generation method according to claim 2, characterized by, Based on the third image and the pose information of the first face, a fourth image is obtained, including: Based on the pose information of the third image and the first face, a sixth image is obtained; Identify the points on the sixth image whose projection positions are incorrect and need to be corrected; The position of the point to be corrected is adjusted to the correct position corresponding to the point to be corrected, and the fourth image is obtained.
4. The image generation method according to claim 3, characterized in that, Based on the first key point information, the standard face model, and the second key point information, a third image is obtained, including: Based on the first key point information and the second key point information, the camera intrinsic parameter matrix, camera projection matrix, rotation matrix, and translation vector are obtained; Based on the camera intrinsic parameter matrix, the camera projection matrix, the rotation matrix, and the translation vector, the model view matrix and the image projection matrix are obtained; Based on the model view matrix and the image projection matrix, the standard face model is subjected to a first image transformation process to obtain the third image; Based on the pose information of the third image and the first face, a sixth image is obtained, including: Based on the pose information of the first face, the third image is subjected to a second image transformation process to obtain the sixth image.
5. The image generation method of claim 4, wherein, The sixth image includes a fourth face and a background. Points with incorrect projection positions on the sixth image are identified, and their positions are adjusted to the correct positions to obtain the fourth image, which includes: Determine the background template; Based on the background template, separate the projection points corresponding to the fourth face and the projection points corresponding to the background. Identify the first point to be corrected among the projection points corresponding to the fourth face, where the projection position is incorrect; adjust the position of the first point to be corrected to the correct position corresponding to the first point to be corrected, thus obtaining the corrected fourth face; and Determine the second point to be corrected in the projection points corresponding to the background part where the projection position is incorrect, and adjust the position of the second point to be corrected to the correct position corresponding to the second point to be corrected to obtain the corrected background part; The fourth image is obtained based on the corrected fourth face and the corrected background.
6. The image generation method of claim 5, wherein, Identifying the first point to be corrected among the projection points corresponding to the fourth face, where the projection position is incorrect, and adjusting the position of the first point to be corrected to the correct position, including: If it is determined that the first point to be corrected, whose projection position is incorrect, is located outside the boundary of the image region corresponding to the fourth face, then the first point to be corrected is normalized to adjust its position to the correct position.
7. The image generation method according to any one of claims 4 to 6, characterized by, The first key point information and the first face pose information are obtained in the following way: The first image is input into a preset recognition model for image recognition processing to obtain the first key point information and the pose information of the first face.
8. An image generation apparatus characterized by comprising: The device includes: A first processing module is configured to determine a first image, the first image including a first face; The second processing module is used to obtain a second image based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to a preset standard face model. The second image includes a second face, the pose of which differs from that of the first face. The first parameters include first key point information of the first key points included in the first face, and pose information of the first face. The second parameters include second key point information of the second key points included in the standard face model. The first image and the second image are two-dimensional images. Obtaining the second image based on the first image, the first parameters corresponding to the first face, and the second parameters corresponding to the preset standard face model includes: A third image is obtained based on the first key point information, the standard face model, and the second key point information. The third image is a two-dimensional image. Based on the pose information of the third image and the first face, a fourth image is obtained, the fourth image including the third face, and the fourth image is a two-dimensional image; Based on the third key point information of the third key point included in the third face and the index information corresponding to the third key point, the position of the third key point is corrected to obtain the second image.
9. An electronic device, comprising: include: A memory for storing computer programs, the computer programs including program instructions; A processor for executing the program instructions to cause the electronic device to perform the image generation method as described in any one of claims 1-7.