Face texture recognition model establishment method
Through feature extraction and visual transformation of neural network models, combined with upsampling blocks and attention gates, parameters are optimized, and the complex and costly parameter adjustment in the existing technology is solved, achieving efficient and accurate texture recognition.
Patent Information
- Application Number
- CN202410148946.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-01
- Publication Date
- 2025-08-01
AI Technical Summary
The existing face texture recognition methods have problems such as complex parameter adjustment, high cost and poor recognition effect. In particular, the texture-based feature extraction method requires manual design of filters. The edge detection operator-based method has poor detection results for wide textures, and the method based on three-dimensional scanning equipment is costly.
Using a neural network model, including an encoder and decoder, feature images are extracted and shaped through feature extraction layers and visual converters, combined with upsampling blocks and attention gates, the loss algorithm is used to optimize model parameters to achieve the recognition of specific textures.
Simplified parameter adjustment, reduced dependence on precision scanning equipment, improved the accuracy of recognition of wider textures, and enhanced the efficiency and accuracy of texture recognition.
Smart Images

Figure CN120412042A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for establishing a recognition model, and in particular to a method for establishing a face texture recognition model. Background Art
[0002] With technological advancements, more and more fields are researching and developing new products based on facial texture. For example, in the cosmetics field, researchers are developing cosmetics tailored to facial textures such as eye bags, tear marks, dimples, and wrinkles. In the field of beauty apps, researchers are applying different levels of beauty effects tailored to the individual textures of the face. Consequently, facial texture recognition technology is gaining increasing attention.
[0003] Existing facial texture recognition methods include texture-based feature extraction, edge detection operator-based feature extraction, and 3D scanning equipment. Texture-based feature extraction is achieved using filters such as Hessian filters, Frangi filters, and Gabor filters, while edge detection operator-based feature extraction uses operators such as the Canny operator, Laplace operator, and DoG operator.
[0004] However, facial texture detection methods based on texture feature extraction require manual filter design, which complicates parameter adjustment. Furthermore, facial texture recognition methods based on feature extraction using edge detection operators identify texture lines by detecting grayscale differences in the image, rather than texture depressions. Consequently, they produce poor results for wider textures. Finally, facial texture recognition methods based on 3D scanning devices require sophisticated scanning equipment, making them significantly more expensive than the other two methods. Summary of the Invention
[0005] The object of the present invention is to provide a method for establishing a facial texture recognition model that can recognize facial texture in a simple and accurate manner.
[0006] The present invention provides a method for establishing a facial texture recognition model. The facial texture recognition model is suitable for identifying a specific texture of a facial image to be recognized and is executed by a device. The device stores multiple training data, each training data including a facial training image and a reference truth image corresponding to the facial training image and marked with the specific texture. The method includes steps (A), (B), (C), (D), (E), (F), (G), and (H).
[0007] In step (A), the device provides a neural network model, which includes an encoder and a decoder, the encoder includes a feature extraction layer and a visual converter, and the decoder includes multiple upsampling blocks.
[0008] In step (B), the device extracts the features of the face training images of the training data by using the feature extraction layer of the encoder of the neural network model, so as to obtain multiple frames of training feature images respectively corresponding to the training data.
[0009] In step (C), the device reshapes the training feature images by using the vision transformer of the encoder of the neural network model, so as to obtain multiple frames of reshaped feature images respectively corresponding to the training feature images.
[0010] In step (D), the device uses the upsampling block of the decoder of the neural network model to obtain multiple frames of texture recognition images respectively corresponding to the reshaped feature images and labeled with the specific texture according to the reshaped feature images.
[0011] In step (E), the device uses a loss algorithm to obtain a training loss value according to the texture recognition images and the ground truth images of the training data.
[0012] In step (F), the device determines whether the training loss value is less than a threshold.
[0013] In step (G), when it is determined that the training loss value is not less than the threshold, the device adjusts multiple parameters of the neural network model, and repeats steps (B) to (F) by using the adjusted neural network model.
[0014] In step (H), when it is determined that the training loss value is less than the threshold, the device makes the neural network model become the face texture recognition model.
[0015] Preferably, in the method for establishing a face texture recognition model of the present invention, the device further stores multiple frames of face images respectively corresponding to the training data, and before step (A), the following steps are further included:
[0016] (I) For each face image, according to the face image, obtain annotation information including multiple annotation points;
[0017] (J) For each face image, according to multiple target annotation points in the annotation information corresponding to the face image, obtain an interested region related to the position where the specific texture is located; and
[0018] (K) For each face image, segment the interested region from the face image to be used as a face training image.
[0019] Preferably, in the method for establishing a face texture recognition model of the present invention, in step (J), the target annotation points include a first target annotation point (x1, y1) and a second target annotation point (x2, y2), the region of interest is a rectangle, and the horizontal axis x range and vertical axis y range of the region of interest in the face image are as follows:
[0020] The horizontal axis x range: min(point(x1), point(x2))
[0021] to
[0022] min(point(x1), point(x2)) + |point1(x1) – point 2(x2)|;
[0023] The vertical axis y range: min(point(y1), point(y2))
[0024] to
[0025] min(point(y1), point(y2)) + |point(y1) – point(y2)|;
[0026] Wherein, min(.) is to take the minimum value, and point(.) is to take the coordinate value.
[0027] Preferably, in the method for establishing a face texture recognition model of the present invention, in step (A), the feature extraction layer of the encoder includes a first convolutional layer, a pooling layer, a first block group layer with multiple blocks, a second block group layer with multiple blocks, and a third block group layer with multiple blocks. Each block is formed by connecting multiple convolutional layers in series. Step (B) includes the following steps:
[0028] (B-1) Using the first convolutional layer of the feature extraction layer, according to the face training images of the training data, obtain multiple frames of first encoded feature images respectively corresponding to the training data;
[0029] (B-2) Using the pooling layer of the feature extraction layer to reduce the dimension of the first encoded feature image;
[0030] (B-3) Using the first block group layer of the feature extraction layer, according to the first encoded feature image, obtain multiple frames of second encoded feature images respectively corresponding to the first encoded feature image;
[0031] (B-4) Using the second block group layer of the feature extraction layer, according to the second encoded feature image, obtain multiple frames of third encoded feature images respectively corresponding to the second encoded feature image; and
[0032] (B-5) Using the third block group layer of the feature extraction layer, obtain the training feature image according to the third encoded feature image.
[0033] Preferably, in the method for establishing a face texture recognition model of the present invention, step (D) includes the following steps:
[0034] (D-1) Using one of the upsampling blocks of the decoder of the neural network model, obtain multiple first decoded feature images respectively corresponding to the shaped feature image according to the shaped feature image and the third encoded feature image;
[0035] (D-2) Using another one of the upsampling blocks of the decoder of the neural network model, obtain multiple second decoded feature images respectively corresponding to the first decoded feature image according to the first decoded feature image and the second encoded feature image;
[0036] (D-3) Using another one of the upsampling blocks of the decoder of the neural network model, obtain multiple third decoded feature images respectively corresponding to the second decoded feature image according to the second decoded feature image and the first encoded feature image; and
[0037] (D-4) Using another one of the upsampling blocks of the decoder of the neural network model, obtain the texture recognition image according to the third decoded feature image.
[0038] Preferably, in the method for establishing a face texture recognition model of the present invention, in step (A), the decoder further includes a plurality of attention gates respectively corresponding to the upsampling blocks, and step (D) includes the following steps:
[0039] (D-1) Using one of the attention gates of the decoder of the neural network model, obtain multiple first attention feature images respectively corresponding to the shaped feature image according to the shaped feature image and the third encoded feature image;
[0040] (D-2) Using one of the upsampling blocks of the decoder of the neural network model, obtain multiple first decoded feature images respectively corresponding to the shaped feature image according to the shaped feature image and the first attention feature image;
[0041] (D-3) Using another one of the attention gates of the decoder of the neural network model, obtain multiple second attention feature images respectively corresponding to the first decoded feature image according to the first decoded feature image and the second encoded feature image;
[0042] (D-4) Using another one of the upsampling blocks of the decoder of the neural network model, obtain multiple frames of second decoded feature images respectively corresponding to the first decoded feature image according to the first decoded feature image and the second attention feature image;
[0043] (D-5) Using another one of the attention gates of the decoder of the neural network model, obtain multiple frames of third attention feature images respectively corresponding to the second decoded feature image according to the second decoded feature image and the first encoded feature image;
[0044] (D-6) Using another one of the upsampling blocks of the decoder of the neural network model, obtain multiple frames of third decoded feature images respectively corresponding to the second decoded feature image according to the second decoded feature image and the third attention feature image; and
[0045] (D-7) Using another one of the upsampling blocks of the decoder of the neural network model, obtain the texture recognition image according to the third decoded feature image.
[0046] Preferably, in the method for establishing a face texture recognition model of the present invention, step (C) includes the following steps:
[0047] (C-1) For each training feature image, use the vision transformer of the encoder of the neural network model to cut the training feature image into multiple image blocks of a fixed size;
[0048] (C-2) For each training feature image, use the vision transformer of the encoder of the neural network model to linearly map the image blocks into multiple vectors respectively corresponding to the image blocks;
[0049] (C-3) For each training feature image, use the vision transformer of the encoder of the neural network model to obtain multiple position encodings respectively corresponding to the vectors; and
[0050] (C-4) For each training feature image, use the vision transformer of the encoder of the neural network model to obtain a shaped feature image corresponding to the training feature image according to the vectors and the position encodings.
[0051] Preferably, in the method for establishing a face texture recognition model of the present invention, in step (E), the training loss value Dice loss is as follows:
[0052]
[0053] Wherein, X is the set of pixel points of the texture recognition image annotating the specific texture, and Y is the set of pixel points of the ground truth image of the training data annotating the specific texture.
[0054] The beneficial effects of the present invention are as follows: By using the device to extract the features of the face training images of the training data through the feature extraction layer, and reshaping the training feature images by using the vision transformer to enhance the recognition of wider textures, complex parameter adjustment and precise scanning equipment are not required. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Other features and effects of the present invention will be clearly presented in the embodiments with reference to the drawings, wherein:
[0056] Figure 1 is a block diagram illustrating a device for implementing an embodiment of the method for establishing a face texture recognition model of the present invention;
[0057] Figure 2 is a schematic diagram illustrating a neural network;
[0058] Figure 3 is a flowchart illustrating a training data establishment procedure of the embodiment of the method for establishing a face texture recognition model of the present invention;
[0059] Figure 4 is a flowchart illustrating a model establishment procedure of the embodiment of the method for establishing a face texture recognition model of the present invention;
[0060] Figure 5 is a flowchart for assisting in explaining Figure 4 the steps included in step 41;
[0061] Figure 6 is a flowchart for assisting in explaining Figure 4 the steps included in step 42;
[0062] Figure 7 is a flowchart for assisting in explaining Figure 4 the steps included in step 43; and
[0063] Figure 8 is a flowchart for assisting in explaining Figure 4 the steps included when step 43 does not use multiple attention gates. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] Before the present invention is described in detail, it should be noted that in the following description, similar components are denoted by the same reference numerals.
[0065] Refer to Figure 1, an apparatus 1 for implementing the method for establishing a face texture recognition model of the present invention is described. The apparatus 1 includes a storage unit 11, a central processing unit (CPU) 12 electrically connected to the storage unit 11, and a graphics processing unit (GPU) 13 electrically connected to the storage unit 11. The face texture recognition model is applicable to recognize a specific texture of a frame of face image to be recognized, and the specific texture is, for example, eye bags, tear stains, pits, and wrinkles.
[0066] The storage unit 11 stores multiple frames of face images and a neural network model (as Figure 2 shown). The neural network model includes an encoder 21 and a decoder 22. The encoder 21 includes a feature extraction layer 211 and a vision transformer (ViT) 212. The decoder 22 includes multiple upsampling blocks 221 and multiple attention gates 222. The feature extraction layer 211 includes a first convolutional layer 211a, a pooling layer 211b, a first block group layer 211c having multiple blocks, a second block group layer 211d having multiple blocks, and a third block group layer 211f having multiple blocks.
[0067] The convolutional kernel size of the first convolutional layer 211a is 7x7, the number of channels is 64, and the stride is 2. The convolutional kernel of the pooling layer 211b is 3x3, and the stride is 2.
[0068] The first block layer has 3 blocks, and each block is formed by connecting three convolutional layers in series. The convolutional kernel size of the first convolutional layer of the three layers is 1x1, and the number of channels is 64. The convolutional kernel size of the second convolutional layer is 3x3, and the number of channels is 64. The convolutional kernel size of the third convolutional layer is 1x1, and the number of channels is 256. The stride of the three convolutional layers is 1.
[0069] The second block group layer 211d has 4 blocks, and each block is formed by connecting three convolutional layers in series. The convolutional kernel size of the first convolutional layer of the three layers is 1x1, and the number of channels is 128. The convolutional kernel size of the second convolutional layer is 3x3, and the number of channels is 128. The convolutional kernel size of the third convolutional layer is 1x1, and the number of channels is 512. The stride of the first convolutional layer is 2, and the stride of the first and second convolutional layers is 1.
[0070] The third block group layer 211f has 6 blocks, and each block is formed by connecting three convolutional layers in series. The convolutional kernel size of the first convolutional layer is 1x1, and the number of channels is 256. The convolutional kernel size of the second convolutional layer is 3x3, and the number of channels is 256. The convolutional kernel size of the third convolutional layer is 1x1, and the number of channels is 1024. The stride of the first convolutional layer is 2, and the strides of the first and second convolutional layers are 1.
[0071] It should be noted that in this embodiment, 4 downsamplings will be performed through the first convolutional layer 211a, the first block layer, the second block group layer 211d, and the third block group layer 211f. Therefore, 4 upsamplings are required. Thus, the number of the upsampling blocks 221 is 4. In other embodiments, the number of the upsampling blocks 221 will vary with the number of downsamplings of the encoder 21, and this is not limited thereto.
[0072] An embodiment of the method for establishing a face texture recognition model of the present invention includes a training data establishment program and a model establishment program.
[0073] Refer to Figure 1 、 3 , for the training data establishment program of this embodiment of the method for establishing a face texture recognition model of the present invention, the steps included in the training data establishment program will be described below.
[0074] In step 31, for each frame of face image, the central processing unit 12 obtains a piece of annotation information including a plurality of annotation points according to the face image.
[0075] It should be noted that in this embodiment, an open-source framework MediaPipe developed by Google Research is used to annotate 468 annotation points on the face image, and each piece of annotation information includes 468 annotation points, but this is not limited thereto.
[0076] In step 32, for each frame of face image, the central processing unit 12 obtains a region of interest related to the position of the specific texture according to a plurality of target annotation points in the annotation information corresponding to the face image.
[0077] Among them, the target annotation points include a first target annotation point (x1, y1) and a second target annotation point (x2, y2). The region of interest is rectangular, and the region of interest has a horizontal axis x range and a vertical axis y range in the face image, as follows:
[0078] The horizontal axis x range:
[0079] from min(point(x1), point(x2)) to min(point(x1), point(x2)) + |point1(x1) – point2(x2)|,
[0080] The range of the vertical axis y:
[0081] from min(point(y1), point(y2)) to min(point(y1), point(y2)) + |point(y1) – point(y2)|)),
[0082] where min(.) is to take the minimum value, and point(.) is to take the coordinate value.
[0083] It should be noted that, in this embodiment, the specific texture is the nasolabial fold, and the target marked points are the 135th marked point and the 236th marked point, but not limited thereto.
[0084] In step 33, for each frame of face image, the central processing unit 12 segments the region of interest from the face image and stores it as a frame of face training image in the storage unit 11.
[0085] Specifically, for each frame of face training image, the central processing unit 12 scales the face training image to a size of 224x384. Then, an expert operates the device 1 to mark the position of the specific texture on a binary image of the same size. The pixels of the binary image that are not the specific texture are represented by 0, and the pixels that are the specific texture are represented by 1. The binary image is a frame of Ground truth image, which is paired with the face training image to form a piece of training data and stored in the storage unit 11, so that the storage unit 11 stores multiple pieces of training data corresponding to the face images respectively.
[0086] It should be noted that, in this embodiment, the central processing unit 12 further pre-processes the face training image and the Ground truth image of the training data to achieve the effect of Data Augmentations. Specifically, the central processing unit 12 normalizes each piece of training data, and then randomly rotates each piece of training data to obtain new training data.
[0087] Refer to Figure 1 、 4 This embodiment of the model establishment procedure of the face texture recognition model establishment method of the present invention will hereinafter describe the steps included in the model establishment procedure.
[0088] In step 41, the graphics processing unit 13 extracts the features of the face training images of the training data by using the feature extraction layer 211, so as to obtain multiple frames of training feature images respectively corresponding to the training data.
[0089] Refer to Figure 5 , step 41 includes steps 411 to 415, and the steps included in step 41 are described below.
[0090] In step 411, the graphics processing unit 13 uses the first convolutional layer 211a to obtain multiple frames of first encoded feature images respectively corresponding to the training data according to the face training images of the training data.
[0091] It should be particularly noted that if the height, width, and number of channels of the face training images of the training data are (H, W, 3), then the height, width, and number of channels of the first encoded feature images are (H / 2, W / 2, 64).
[0092] In step 412, the graphics processing unit 13 uses the pooling layer 211b to reduce the dimension of the first encoded feature image. Among them, the height, width, and number of channels of the first encoded feature image after reducing the dimension are (H / 4, W / 4, 64).
[0093] In step 413, the graphics processing unit 13 uses the first block group layer 211c to obtain multiple frames of second encoded feature images respectively corresponding to the first encoded feature image according to the first encoded feature image. The height, width, and number of channels of the second encoded feature image are (H / 4, W / 4, 256).
[0094] In step 414, the graphics processing unit 13 uses the second block group layer 211d to obtain multiple frames of third encoded feature images respectively corresponding to the second encoded feature image according to the second encoded feature image. The height, width, and number of channels of the third encoded feature image are (H / 8, W / 8, 512).
[0095] In step 415, the graphics processing unit 13 uses the third block group layer 211f to obtain the training feature image according to the third encoded feature image. The height, width, and number of channels of the training feature image are (H / 16, W / 16, 1024).
[0096] In step 42, the graphics processing unit 13 reshapes the training feature image by using the vision transformer 212, so as to obtain multiple frames of reshaped feature images respectively corresponding to the training feature image.
[0097] Refer to Figure 6, step 42 includes steps 421 to 424. The steps included in step 42 are described below.
[0098] In step 421, for each frame of training feature image, the graphics processing unit 13 uses the vision transformer 212 to cut the training feature image into multiple image patches of a fixed size.
[0099] In step 422, for each frame of training feature image, the graphics processing unit 13 uses the vision transformer 212 to linearly map the image patches into multiple vectors respectively corresponding to the image patches.
[0100] In step 423, for each frame of training feature image, the graphics processing unit 13 uses the vision transformer 212 to obtain multiple positional encodings respectively corresponding to the vectors.
[0101] In step 424, for each frame of training feature image, the graphics processing unit 13 uses the vision transformer 212 to obtain a shaped feature image corresponding to the training feature image according to the vectors and the positional encodings.
[0102] In step 43, the graphics processing unit 13 uses the upsampling block 221 of the decoder 22 of the neural network model to obtain multiple texture recognition images respectively corresponding to the shaped feature image and labeled with the specific texture according to the shaped feature image.
[0103] Refer to Figure 7 , step 43 includes steps 431 to 437. The steps included in step 43 are described below.
[0104] In step 431, the graphics processing unit 13 uses one of the attention gates 222 to obtain multiple first attention feature images respectively corresponding to the shaped feature image according to the shaped feature image and the third encoded feature image.
[0105] In step 432, the graphics processing unit 13 uses one of the upsampling blocks 221 to obtain multiple first decoded feature images respectively corresponding to the shaped feature image according to the shaped feature image and the first attention feature image.
[0106] Specifically, the upsampling block 221 doubles the height and width of the shaped feature image while keeping the number of channels unchanged, and then merges the shaped feature image and the first attention feature image one by one to obtain multiple frames of first merged feature images corresponding to the shaped feature image respectively. The height, width, and number of channels of the first merged feature image are (H / 8, W / 8, 1536). The upsampling block 221 then passes the first merged feature image through two convolutional layers to obtain the first decoded feature image. The convolutional kernel size of each convolutional layer is 3X3, the stride is 1, and the number of channels is 512. The height, width, and number of channels of the first decoded feature image are (H / 8, W / 8, 512).
[0107] In step 433, the graphics processing unit 13 uses another one of the attention gates 222 to obtain multiple frames of second attention feature images corresponding to the first decoded feature image according to the first decoded feature image and the second encoded feature image.
[0108] In step 434, the graphics processing unit 13 uses another one of the upsampling blocks 221 to obtain multiple frames of second decoded feature images corresponding to the first decoded feature image according to the first decoded feature image and the second attention feature image.
[0109] Specifically, the upsampling block 221 doubles the height and width of the first decoded feature image while keeping the number of channels unchanged, and then merges the first decoded feature image and the second attention feature image one by one to obtain multiple frames of second merged feature images corresponding to the first decoded feature image respectively. The height, width, and number of channels of the second merged feature image are (H / 4, W / 4, 768). The upsampling block 221 then passes the second merged feature image through two convolutional layers to obtain the second decoded feature image. The convolutional kernel size of each convolutional layer is 3X3, the stride is 1, and the number of channels is 256. The height, width, and number of channels of the second decoded feature image are (H / 4, W / 4, 256).
[0110] In step 435, the graphics processing unit 13 uses another one of the attention gates 222 to obtain multiple frames of third attention feature images corresponding to the second decoded feature image according to the second decoded feature image and the first encoded feature image.
[0111] In step 436, the graphics processing unit 13 uses another one of the upsampling blocks 221 to obtain multiple frames of third decoded feature images corresponding to the second decoded feature image according to the second decoded feature image and the third attention feature image.
[0112] Specifically, the upsampling block 221 doubles the height and width of the second decoded feature image while keeping the number of channels unchanged, and then merges the second decoded feature image and the third attention feature image one by one to obtain multiple frames of third merged feature images respectively corresponding to the first decoded feature image. The height, width, and number of channels of the third merged feature image are (H / 2, W / 2, 320). The upsampling block 221 then passes the third merged feature image through two convolutional layers to obtain the third decoded feature image. The convolutional kernel size of each convolutional layer is 3X3, the stride is 1, and the number of channels is 64. The height, width, and number of channels of the third decoded feature image are (H / 2, W / 2, 64).
[0113] In step 437, the graphics processing unit 13 uses another one of the upsampling blocks 221 to obtain the texture recognition image according to the third decoded feature image.
[0114] Specifically, the upsampling block 221 doubles the height and width of the third decoded feature image while keeping the number of channels unchanged, and then passes the third decoded feature image through two convolutional layers to obtain the texture recognition image. The convolutional kernel size of each convolutional layer is 3X3, the stride is 1, and the number of channels is 2. The height, width, and number of channels of the texture recognition image are (H, W, 2).
[0115] It should be particularly noted that since the pixels of the specific texture to be recognized often only account for about 3% of the pixels of the entire image, most of the traditional neural networks are processing the parts other than the specific texture during the training process. Therefore, in this embodiment, the attention gate 222 is used to enable the neural network to pay more attention to the specific texture part in the image during the training process, so as to enhance the accuracy of the neural network recognition.
[0116] In other embodiments, the attention gate 222 may not be used. With reference to Figure 8 , if the attention gate 222 is not used, step 43 includes steps 431’ to 434’. The following describes the steps included in step 43 when the attention gate 222 is not used.
[0117] In step 431’, the graphics processing unit 13 uses one of the upsampling blocks 221 to obtain multiple frames of first decoded feature images respectively corresponding to the shaping feature image according to the shaping feature image and the third encoded feature image.
[0118] In step 432’, the graphics processing unit 13 uses another one of the upsampling blocks 221 to obtain multiple frames of second decoded feature images respectively corresponding to the first decoded feature image according to the first decoded feature image and the second encoded feature image.
[0119] In step 433’, the graphics processing unit 13 uses another one of the upsampling blocks 221 of the decoder 22 of the neural network model to obtain multiple frames of third decoded feature images respectively corresponding to the second decoded feature image according to the second decoded feature image and the first encoded feature image.
[0120] In step 434’, the graphics processing unit 13 uses another one of the upsampling blocks 221 to obtain the texture recognition image according to the third decoded feature image.
[0121] In step 44, the graphics processing unit 13 uses a loss algorithm to obtain a training loss value according to the texture recognition image and the ground truth image of the training data.
[0122] Specifically, for each frame of texture recognition image, the graphics processing unit 13 converts the texture recognition image into a frame of recognition binary image with a size of (H, W, 1). Among them, for each pixel of the recognition binary image, the value of the pixel is determined according to the values of the pixels at the same pixel position in the two channels of the texture recognition image. That is, the texture recognition image has two channels, namely channel 0 and channel 1. If the value of the pixel in channel 0 is greater than or equal to the value of the pixel in channel 1, the value of the pixel in the recognition binary image is 0; if the value of the pixel in channel 0 is less than the value of the pixel in channel 1, the value of the pixel in the recognition binary image is 1. The pixels with a binary image value of 1 represent the pixels marked with the specific texture, and the pixels with a value of 0 are the pixels not marked with the specific texture.
[0123] The training loss value Dice loss is as follows:
[0124]
[0125] Where X is the set of pixel points of the specific texture marked in the texture recognition image, Y is the set of pixel points of the specific texture marked in the ground truth image of the training data, and ε is a very small positive number to prevent the denominator from being zero.
[0126] In step 45, the graphics processing unit 13 determines whether the training loss value is less than a threshold. When it is determined that the training loss value is not less than the threshold, the process proceeds to step 46; when it is determined that the training loss value is less than the threshold, the process proceeds to step 47.
[0127] In step 46, the graphics processing unit 13 adjusts multiple parameters of the neural network model and repeats steps 41 to 45 using the adjusted neural network model.
[0128] In step 47, the graphics processing unit 13 makes the neural network model into the face texture recognition model.
[0129] In summary, for the method for establishing a face texture recognition model according to the present invention, the graphics processing unit 13 extracts the features of the face training images of the training data by means of the feature extraction layer 211, reshapes the training feature images by means of the vision transformer 212 to enhance the recognition of wider textures, and uses the attention gate 222 to enable the neural network to pay more attention to the specific texture parts in the images during the training process. Without complex parameter adjustment and precise scanning equipment, the object of the present invention can indeed be achieved.
[0130] The above are only the embodiments of the present invention, and the scope of implementation of the present invention cannot be limited thereby. That is, all simple equivalent changes and modifications made according to the claims and the content of the specification of the present invention still fall within the scope of the present invention.
Claims
1. A method for establishing a face texture recognition model, the face texture recognition model being applicable to recognizing a specific texture of a face image to be recognized, and being executed by a device, the device storing multiple pieces of training data, each piece of training data including a face training image and a ground truth image corresponding to the face training image and labeled with the specific texture, characterized in that: The method for establishing the face texture recognition model includes the following steps: (A) Provide a neural network model, the neural network model includes an encoder and a decoder, the encoder includes a feature extraction layer and a vision transformer, and the decoder includes a plurality of upsampling blocks; (B) Use the feature extraction layer of the encoder of the neural network model to extract the features of the face training images of the training data to obtain multiple frames of training feature images respectively corresponding to the training data; (C) Use the vision transformer of the encoder of the neural network model to reshape the training feature images to obtain multiple frames of reshaped feature images respectively corresponding to the training feature images; (D) Use the upsampling blocks of the decoder of the neural network model to obtain multiple frames of texture recognition images respectively corresponding to the reshaped feature images and labeled with the specific texture according to the reshaped feature images; (E) Use a loss algorithm to obtain a training loss value according to the texture recognition images and the ground truth images of the training data; (F) Determine whether the training loss value is less than a threshold; (G) When it is determined that the training loss value is not less than the threshold, adjust multiple parameters of the neural network model, and use the adjusted neural network model to repeat steps (B) to (F); and (H) When it is determined that the training loss value is less than the threshold, make the neural network model become the face texture recognition model.
2. The method for establishing a face texture recognition model according to claim 1, wherein: The device also stores multiple frames of face images respectively corresponding to the training data, and before step (A), it further includes the following steps: (I) For each face image, obtain annotation information including a plurality of annotation points according to the face image; (J) For each face image, obtain a region of interest related to the position where the specific texture is located according to a plurality of target annotation points in the annotation information corresponding to the face image; and (K) For each face image, segment the region of interest from the face image to be used as a face training image.
3. The method for establishing a face texture recognition model according to claim 2, wherein: In step (J), the target annotation points include a first target annotation point (x1, y1) and a second target annotation point (x2, y2), the region of interest is a rectangle, and the horizontal axis x range and vertical axis y range of the region of interest in the face image are as follows: The horizontal axis x range: min(point(x1), point(x2)) to min(point(x1), point(x2)) + |point1(x1) – point 2(x2)|; The vertical axis y range: min(point(y1), point(y2)) to min(point(y1), point(y2)) + |point(y1) – point(y2)|; where min(.) is to take the minimum value, and point(.) is to take the coordinate value.
4. The method for establishing a face texture recognition model according to claim 1, characterized in that: In step (A), the feature extraction layer of the encoder includes a first convolutional layer, a pooling layer, a first block group layer with multiple blocks, a second block group layer with multiple blocks, and a third block group layer with multiple blocks. Each block is formed by cascading multiple convolutional layers. Step (B) includes the following steps: (B-1) Using the first convolutional layer of the feature extraction layer, obtain multiple frames of first encoded feature images respectively corresponding to the training data according to the face training images of the training data; (B-2) Using the pooling layer of the feature extraction layer to reduce the dimension of the first encoded feature images; (B-3) Using the first block group layer of the feature extraction layer, obtain multiple frames of second encoded feature images respectively corresponding to the first encoded feature images according to the first encoded feature images; (B-4) Using the second block group layer of the feature extraction layer, obtain multiple frames of third encoded feature images respectively corresponding to the second encoded feature images according to the second encoded feature images; and (B-5) Using the third block group layer of the feature extraction layer, obtain the training feature images according to the third encoded feature images.
5. The method for establishing a face texture recognition model according to claim 4, characterized in that: Step (D) includes the following steps: (D-1) Using one of the upsampling blocks of the decoder of the neural network model, obtain multiple frames of first decoded feature images respectively corresponding to the shaped feature images according to the shaped feature images and the third encoded feature images; (D-2) Using another one of the upsampling blocks of the decoder of the neural network model, obtain multiple frames of second decoded feature images respectively corresponding to the first decoded feature images according to the first decoded feature images and the second encoded feature images; (D-3) Using another one of the upsampling blocks of the decoder of the neural network model, obtain multiple frames of third decoded feature images respectively corresponding to the second decoded feature images according to the second decoded feature images and the first encoded feature images; and (D-4) Using another one of the upsampling blocks of the decoder of the neural network model, obtain the texture recognition images according to the third decoded feature images.
6. The method for establishing a face texture recognition model according to claim 4, wherein: In step (A), the decoder further includes multiple attention gates respectively corresponding to the upsampling blocks. Step (D) includes the following steps: (D-1) Using one of the attention gates of the decoder of the neural network model, obtain multiple frames of first attention feature images respectively corresponding to the shaped feature images according to the shaped feature images and the third encoded feature images; (D-2) Using one of the upsampling blocks of the decoder of the neural network model, obtain multiple frames of first decoded feature images respectively corresponding to the shaped feature images according to the shaped feature images and the first attention feature images; (D-3) Using another one of the attention gates of the decoder of the neural network model, obtain multiple frames of second attention feature images respectively corresponding to the first decoded feature image according to the first decoded feature image and the second encoded feature image; (D-4) Using another one of the upsampling blocks of the decoder of the neural network model, obtain multiple frames of second decoded feature images respectively corresponding to the first decoded feature image according to the first decoded feature image and the second attention feature image; (D-5) Using another one of the attention gates of the decoder of the neural network model, obtain multiple frames of third attention feature images respectively corresponding to the second decoded feature image according to the second decoded feature image and the first encoded feature image; (D-6) Using another one of the upsampling blocks of the decoder of the neural network model, obtain multiple frames of third decoded feature images respectively corresponding to the second decoded feature image according to the second decoded feature image and the third attention feature image; and (D-7) Using another one of the upsampling blocks of the decoder of the neural network model, obtain the texture recognition image according to the third decoded feature image.
7. The method for establishing a face texture recognition model according to claim 1, wherein: Step (C) includes the following steps: (C-1) For each training feature image, use the vision transformer of the encoder of the neural network model to cut the training feature image into multiple image blocks of a fixed size; (C-2) For each training feature image, use the vision transformer of the encoder of the neural network model to linearly map the image blocks into multiple vectors respectively corresponding to the image blocks; (C-3) For each training feature image, use the vision transformer of the encoder of the neural network model to obtain multiple position encodings respectively corresponding to the vectors; and (C-4) For each training feature image, use the vision transformer of the encoder of the neural network model to obtain a shaped feature image corresponding to the training feature image according to the vectors and the position encodings.
8. The method for establishing a face texture recognition model according to claim 1, wherein: In step (E), the training loss value Dice loss is as follows: where X is the set of pixel points of the texture recognition image annotating the specific texture, and Y is the set of pixel points of the ground truth image of the training data annotating the specific texture.