Feature extraction network training method and apparatus, classification method and apparatus, and electronic device
By adjusting the parameters and class center matrix of the feature extraction network in two stages, the class center offset problem caused by difficult samples is solved, and the accuracy of feature extraction is improved.
Patent Information
- Application Number
- PCT/CN2024/125667
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-10-18
- Publication Date
- 2025-05-08
AI Technical Summary
When using sample sets to train image feature extraction networks, difficult sample images can easily lead to class-center distribution offset, affecting the training effect.
The training method of feature extraction network is adopted in two stages. First, by adjusting the initial class center matrix of multiple class centers, ensuring that the features of the sample images of each category are mapped to the corresponding class center; secondly, determine the main class center and adjust the network parameters until the end condition of the training phase is reached, and the trained feature extraction network is obtained.
The class-center offset is effectively avoided, and the feature extraction performance of the trained feature extraction network for low-quality difficult samples is improved, thereby improving the accuracy of feature extraction.
Smart Images

Figure CN2024125667_08052025_PF_FP_ABST
Abstract
Description
Feature extraction network training method, classification method, device and electronic equipment
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on October 31, 2023, with application number 202311431364.0, and invention name “Training method, classification method, device and electronic device for feature extraction network”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence technology, and more specifically, to a training method, classification method, device and electronic equipment for a feature extraction network. Background Art
[0003] Feature extraction network is a deep learning technique that can extract useful features from input data for effective classification and prediction.
[0004] When using a sample set to train an image feature extraction network, the sample set usually includes high-quality normal sample images and low-quality difficult sample images, such as non-frontal sample images, sample images with low clarity, and sample images with low pixels.
[0005] Technical content
[0006] The embodiments of the present application provide a feature extraction network training method, a classification method, an apparatus, and an electronic device, which can move the features of difficult samples closer to the main class center of the corresponding class to avoid class center shift.
[0007] An embodiment of the present application provides a method for training a feature extraction network, the method comprising: performing the following steps in each iterative cycle of a first training phase: extracting image features of a first sample image using a feature extraction network to be trained; determining a first loss based on the image features of the first sample image, the category label of the first sample image, and initial category center matrices corresponding to multiple category centers of each category in a plurality of categories; adjusting parameters of the feature extraction network to be trained and initial category center matrices corresponding to multiple category centers of each category based on the first loss until a first training phase end condition is met, thereby obtaining a preliminarily trained feature extraction network and first category center matrices corresponding to multiple category centers of each category; performing the following steps in each iterative cycle of a second training phase: extracting image features of a second sample image using the preliminarily trained feature extraction network; determining a second loss based on the image features of the second sample image, the category label of the second sample image, and first category center matrices corresponding to main category centers of each category; the main category center of a category being one of the multiple category centers of that category; adjusting parameters of the preliminarily trained feature extraction network and first category center matrices corresponding to main category centers of each category based on the second loss until a second training phase end condition is met, thereby obtaining a trained feature extraction network, wherein the trained feature extraction network is used for image classification.
[0008] An embodiment of the present application also provides an image classification method, which includes: obtaining an image to be classified; performing feature extraction on the image to be classified using a trained image feature extraction network obtained by the aforementioned feature extraction network training method to obtain target image features; and determining a classification result of the image to be classified based on the target image features.
[0009] An embodiment of the present application also provides a training device for a feature extraction network, which includes a first feature extraction module, a first loss determination module, a first adjustment module, a second feature extraction module, a second loss determination module, and a second adjustment module. The first feature extraction module is used to extract image features of a first sample image through the feature extraction network to be trained; the first loss determination module is used to determine a first loss based on the image features of the first sample image, the category label of the first sample image, and the initial class center matrix corresponding to multiple class centers of each category in multiple categories; the first adjustment module is used to adjust the parameters of the feature extraction network to be trained and the initial class center matrix corresponding to the multiple class centers of each category according to the first loss, to obtain A preliminarily trained feature extraction network and first-class center matrices corresponding to multiple class centers of each category; a second feature extraction module, used to extract image features of a second sample image through the preliminarily trained feature extraction network; a second loss determination module, used to determine a second loss based on the image features of the second sample image, the category label of the second sample image and the first-class center matrix corresponding to the main class center of each category; the main class center of a category is one of the multiple class centers of the category; a second adjustment module, used to adjust the parameters of the preliminarily trained feature extraction network and the first-class center matrix corresponding to the main class center of each category according to the second loss, to obtain a trained feature extraction network, and the trained feature extraction network is used for image classification.
[0010] An embodiment of the present application also provides an image classification device, which includes an image acquisition module, a third feature extraction module and a classification result determination module. The image acquisition module is used to acquire the image to be classified; the third feature extraction module is used to use the trained image feature extraction network obtained by the above-mentioned feature extraction network training device to perform feature extraction on the image to be classified to obtain target image features; the classification result determination module is used to determine the classification result of the image to be classified based on the target image features.
[0011] An embodiment of the present application also provides an electronic device, including a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.
[0012] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores program code, wherein the above method is executed when the program code is executed by a processor.
[0013] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device retrieves the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described method.
[0014] BRIEF DESCRIPTION OF THE DRAWINGS
[0015] FIG1 shows an application scenario diagram of a feature extraction network training method provided in an embodiment of the present application;
[0016] FIG2 is a schematic diagram showing a flow chart of a method for training a feature extraction network according to an embodiment of the present application;
[0017] FIG3 shows a schematic diagram of a method for intercepting a palmprint pixel area according to an embodiment of the present application;
[0018] FIG4 shows a schematic diagram of a network structure of a feature extraction network proposed in an embodiment of the present application;
[0019] FIG5 shows another flow chart of a method for training a feature extraction network according to an embodiment of the present application;
[0020] FIG6 shows a schematic flow chart of an image classification method proposed in an embodiment of the present application;
[0021] FIG7 shows an application flow chart of a feature extraction network proposed in an embodiment of the present application;
[0022] FIG8 shows a connection block diagram of a training device for a feature extraction network proposed in an embodiment of the present application;
[0023] FIG9 shows a connection block diagram of an image classification device proposed in an embodiment of the present application;
[0024] FIG10 shows a structural block diagram of an electronic device for executing the method according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the reference examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.
[0026] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0027] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0028] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0029] It should be noted that the term "plurality" used in this document refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. The character " / " generally indicates an "or" relationship between the associated objects.
[0030] In some methods, when using a sample set to train an image feature extraction network, the sample set often includes both high-quality, normal sample images and low-quality, difficult sample images. When using this sample set to train an image feature extraction network, difficult sample images can easily cause a shift in the distribution of class centers during training, thus affecting the training effect. To address this issue, this application proposes training the feature extraction network in two stages, which are described in detail below.
[0031] The solution of this application mainly uses machine learning to perform image classification.
[0032] In the present application, multiple class centers are set for each of the multiple categories. For each category, a class center under the category is equivalent to a subcategory under the category, that is, in the present application, multiple subcategories are set for each category. Among them, one class center corresponds to an initial class center matrix, and the initial class center matrix is the feature space of the class center. Multiple class centers of each category are pre-constructed, and the initial class center matrix corresponding to each class center is a fully connected matrix. The number of class centers of different categories can be the same or different. Before training the feature extraction network, initial values can be set for the initial class center matrices corresponding to the multiple class centers of each category, and the initial values of the initial class center matrices corresponding to different class centers can be the same or different.
[0033] FIG1 is a schematic diagram of an application scenario according to an embodiment of the present application. As shown in FIG1 , the application scenario includes a terminal device 10 and a server 20 that is communicatively connected to the terminal device 10 via a network.
[0034] The terminal device 10 may be a mobile phone, a computer, an intelligent voice interaction device, an intelligent home appliance, a vehicle-mounted terminal, etc. The terminal device 10 may be provided with a client for displaying data. The network may be a wide area network or a local area network, or a combination of the two.
[0035] Server 20 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0036] If the feature extraction network is trained using the terminal device 10 and the server 20 as shown in Figure 1, the terminal device 10 can upload the first sample image and the second sample image to the server 20. After obtaining the first sample image and the second sample image, the server 20 performs the following steps in each iteration cycle of the first training phase: extracting the image features of the first sample image through the feature extraction network to be trained; determining the first loss based on the image features of the first sample image, the category label of the first sample image and the initial class center matrix corresponding to the multiple class centers of each category; adjusting the parameters of the feature extraction network to be trained and the initial class center matrix corresponding to the multiple class centers of each category according to the first loss until the end condition of the first training phase is met, and the result is obtained. to the preliminarily trained feature extraction network and the first class center matrices corresponding to the multiple class centers of each category; in each iterative cycle of the second training stage, the following steps are performed: extracting image features of the second sample image through the preliminarily trained feature extraction network; determining the second loss according to the image features of the second sample image, the class label of the second sample image and the first class center matrix corresponding to the main class center of each category; the main class center of a category is one of the multiple class centers of the category; according to the second loss, adjusting the parameters of the preliminarily trained feature extraction network and the first class center matrix corresponding to the main class center of each category until the end condition of the second training stage is reached, thereby obtaining the trained feature extraction network, and the trained feature extraction network is used for image classification.
[0037] The present application adopts the above method to achieve the training process of a difficult sample model with a large number of training samples and a large number of low-quality ones. By setting multiple class centers for each category in the multiple categories corresponding to the classification task, in the first training stage of the feature extraction network, the parameters of the feature extraction network to be trained and the initial class center matrices corresponding to the multiple class centers of each category are adjusted according to the first loss determined by the image features of the first sample image, the category label of the first sample image and the initial class center matrices corresponding to the multiple class centers of each category, so as to achieve that for each category, the image features of the sample images belonging to that category (such as normal sample images and difficult sample images) can be mapped to one of the multiple class centers corresponding to the category. In the second training stage of the feature extraction network, after determining the main class center, the parameters of the preliminary trained feature extraction network and the first class center matrix corresponding to the main class center of each class are adjusted according to the second loss determined based on the image features of the second sample image, the category label of the second sample image and the first class center moment corresponding to the main class center of each class, so as to achieve the goal of gradually approaching the main class center for the image features extracted by the preliminary trained feature extraction network for the difficult samples in the second sample image, so as to improve the feature extraction performance of the low-quality difficult sample images in the second sample image while ensuring the feature extraction performance of the preliminary trained feature extraction network for the normal samples in the second sample image, thereby improving the accuracy of the features extracted by the trained feature extraction network.
[0038] After the training of the feature extraction network is completed, the trained feature extraction network can be deployed on the server 20, so that after obtaining the image to be classified, the trained image feature extraction network can be used to extract features of the image to be classified to obtain the target image features; according to the target image features, subsequent image processing operations are performed, such as image classification operations.
[0039] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0040] Please refer to FIG2 , which shows that the present application also provides a method for training a feature extraction network, which can be applied to an electronic device, which can be the terminal device 10 or the server 20 mentioned above. The method includes a first training phase and a second training phase, wherein:
[0041] In each iteration cycle of the first training stage, the following steps S110 to S130 are performed:
[0042] Step S110: extracting image features of the first sample image through the feature extraction network to be trained.
[0043] The first sample image can be obtained by annotating images in an image set, or by obtaining a plurality of pre-stored images with sample category labels from an electronic device or other devices associated with the electronic device as the first sample image. The setting can be made according to actual needs.
[0044] There are multiple first sample images, and each first sample image is marked with a category label. The above-mentioned category label can be set according to the classification task, and the category corresponding to the category label of the first sample image belongs to one of the multiple categories in the above-mentioned classification task. If the classification task is to identify whether the image is qualified, the category label of the sample image is qualified or unqualified. If the classification task is used to classify objects (such as animals and plants) in the image, such as cats, dogs, pigs and other animals, the category label of the sample image can be the specific category to which the object in the image belongs; if the classification task is to identify whether the image is an abnormal image, the category label of the sample image is normal or abnormal. If the classification task is to identify the specific identification information corresponding to a certain image, the category label of the sample image is the identification information of a certain object.
[0045] Taking the classification task of identifying and classifying an object based on its biometric information (e.g., facial image, fingerprint image, palm print image, or iris image), the category label is the identification information of an object. This identification information can be information used to uniquely identify the object, such as the object name or object ID. In this manner, the first sample image can be obtained by extracting a specified area from the initial image. Specifically, key point detection can be performed on the initial image, and based on the key point detection results, a specified area (e.g., a facial pixel area, a fingerprint pixel area, a palm print pixel area, or an iris pixel area) can be extracted from the initial image to serve as the first sample image.
[0046] As shown in FIG3 , exemplarily, when the biometric information is palm print information, the method for obtaining the first sample image may be to perform key point detection on the initial image to obtain key points in the initial image; the key points of the initial image include a first finger gap key point A between the index finger and the middle finger, a second finger gap key point B between the middle finger and the ring finger, and a third finger gap key point C between the ring finger and the little finger; an image coordinate system is established based on the first finger gap key point A, the second finger gap key point B, and the third finger gap key point C in the initial image, and the line between the first finger gap key point A and the third finger gap key point C is the horizontal axis (x-axis) of the image coordinate system. ), a line passing through the second finger joint key point B and perpendicular to the horizontal axis is the vertical axis (y-axis) of the image coordinate system, and the intersection between the horizontal axis and the vertical axis is the origin of the image coordinate system; a point in the initial image that is located on the vertical axis of the image coordinate system and is at a specified distance from the origin of the image coordinate system is used as the palm print center point D in the initial image, and the specified distance can be determined according to the distance between the first finger joint key point A and the third finger joint key point C, and the palm print center point D and the second finger joint key point B are respectively located on both sides of the horizontal axis; a sample image is intercepted from the initial image based on the palm print center point D and the distance between the first finger joint key point A and the third finger joint key point B.
[0047] Specifically, after determining the image coordinate system shown in Figure 3 based on the first, second, and third finger joint key points A, B, and C, the palm print center point D is found at a distance AC from the coordinate origin along the negative y-axis. The distance DE is equal to six-fifths of the distance AC. The distance from point A to point C is multiplied by 3 / 2 to obtain the side length d of the palm print pixel area. The palm print pixel area is extracted as a sample image (i.e., the first sample image) with point D as the center and a square with a side length of d.
[0048] It should be understood that when the biometric features of the object are different, the corresponding methods of intercepting the image of the designated area from the initial image are different, which will not be described in detail here.
[0049] The feature extraction network can be constructed using one or more neural networks. Specifically, the neural network can be any neural network that can extract image features, such as a ResNet residual network, a DenseNet classic network, a VGG convolutional neural network, an AlexNet deep convolutional neural network, a Swin-Transformer network, a MaxViT network, or a LeNet convolutional neural network, and is not specifically limited in this embodiment.
[0050] Step S120: determining a first loss according to the image features of the first sample image, the category label of the first sample image, and the initial class center matrices corresponding to the multiple class centers of each of the multiple categories.
[0051] The above-mentioned step S120 can be: determining the category prediction result of the first sample image based on the similarity between the image features of the first sample image and the initial class center matrices corresponding to multiple class centers of each category, and performing loss calculation based on the category prediction result of the first sample image and the category label of the first sample image to obtain the first loss.
[0052] In this manner, based on the similarities between the image features of the first sample image and the initial class center matrices corresponding to the multiple class centers of each category, a method for determining the class prediction result of the first sample image can be specifically as follows: for each category, the average or maximum value of the similarities between the image features of the first sample image and the initial class center matrices corresponding to the multiple class centers of the category is determined as the reference similarity between the image features of the first sample image and the category; based on the reference similarities between the image features of the first sample image and the category, the probability that the first sample image belongs to each category is determined, wherein the reference similarity between the image features of the first sample image and the category reflects the difference between the image features corresponding to the first sample image and the feature space corresponding to the category; the smaller the difference, the greater the probability that the first sample image is predicted to belong to the category. The probability that the first sample image is predicted to belong to each category is the class prediction result of the first sample image.
[0053] The above-mentioned step S120 can also be: based on the similarity between the image features of the first sample image and the initial class center matrices corresponding to the multiple class centers of each category, at least one reference class center is determined from the multiple class centers corresponding to each category, and the category prediction result of the first sample image is determined based on the similarity between each reference class center of each category and the image features of the first sample image, and the loss is calculated based on the category prediction result of the first sample image and the category label of the first sample image to obtain the first loss.
[0054] The similarity between the image features of the first sample image and the initial class center matrices corresponding to the multiple class centers of each category reflects the difference between the image features corresponding to the first sample image and the feature spaces corresponding to the initial class center matrices of the multiple class centers of each category. The similarity between a reference class center and the image features of the first sample image also reflects the difference between the feature space corresponding to the reference class center and the image features. It should be understood that the smaller the difference, the greater the probability that the predicted class of the first sample image is the class corresponding to the reference class center.
[0055] The similarity between the initial class center matrix corresponding to a class center and the image features of the first sample image can be calculated by calculating the cosine similarity or Euclidean distance between the initial class center matrix corresponding to the class center and the image features of the first sample image.
[0056] When calculating the loss for the category prediction result of the first sample image and the category label of the first sample image, a preset loss function may be used to calculate the loss for the category prediction result of the first sample image and the category label of the first sample image. The preset loss function may be a cross entropy loss function, a mean square error loss function, or a multi-classification cross entropy loss function, and may be set according to actual needs.
[0057] Step S130: According to the first loss, the parameters of the feature extraction network to be trained and the initial class center matrices corresponding to the multiple class centers of each category are adjusted until the end condition of the first training stage is reached, and the preliminarily trained feature extraction network and the first class center matrices corresponding to the multiple class centers of each category are obtained.
[0058] In this embodiment, when the number of iterations in the first training phase reaches a first preset number, the number of cycles corresponding to the iteration cycle reaches a first preset number of cycles, or the first loss is less than a first preset loss threshold, it is considered that the training of the feature extraction network to be trained has reached the end condition of the first training phase. When the training of the feature extraction network to be trained reaches the end condition of the first training phase, the feature extraction network to be trained after the last iterative adjustment is used as the feature extraction network for preliminary training, and the initial class center matrix corresponding to the multiple class centers of each category after the last iterative adjustment is used as the first class center matrix corresponding to the multiple class centers of each category. The first preset number of times, the first preset number of cycles, and the first preset loss threshold can be set according to task requirements and are not specifically limited here.
[0059] It should be noted that multiple first training sets can be set to participate in the training process of steps S110 to S130 as described above. Each first training set includes multiple first sample images. If all first sample images in a first training set participate in training the feature extraction network to be trained once, it is considered that one epoch has been completed. In some embodiments, the first sample images in each first training set can be input into the feature extraction network to be trained in batches for iterative training. The number of batches corresponding to each batch can be set according to actual needs.
[0060] By iteratively adjusting the parameters of the feature extraction network to be trained and the initial class center matrices corresponding to the multiple class centers of each category, the feature extraction accuracy of the feature extraction network to be trained can be improved while constraining the feature space corresponding to the multiple class centers of each category.
[0061] In each iteration cycle of the second training phase, the following steps S140 to S160 are performed:
[0062] Step S140: extracting image features of the second sample image through the preliminarily trained feature extraction network.
[0063] The second sample image should be acquired in the same or similar manner as the first sample image, and the category corresponding to the second sample image should also belong to one of the multiple categories corresponding to the classification task when the initially trained feature extraction network is used in the classification task. Therefore, the process of acquiring and processing the second sample image can refer to the description of the first sample image above and is not further specified here.
[0064] It should be understood that there are multiple second sample images, and the second sample images may be different from the first sample images, or the first sample images may also be used as second sample images to participate in the training of the feature extraction network for preliminary training.
[0065] For a detailed description of extracting the second sample image through the preliminarily trained feature extraction network, please refer to the detailed description of step S110 above, which will not be repeated here.
[0066] Step S150: determining a second loss based on the image features of the second sample image, the category label of the second sample image, and the first category center matrix corresponding to the main category center of each category.
[0067] The main class center of a class is one of the multiple class centers of the class.
[0068] The main class center of a class refers to the class center corresponding to the feature space (class center matrix) where the features of images belonging to the class are mainly distributed.
[0069] In some embodiments, the main class center of a category can be determined based on the number of sample images (e.g., first sample images) attached to multiple class centers of the category. For example, the class center with the largest number of first sample images attached to the multiple class centers of the category is determined as the main class center of the category.
[0070] In other possible implementations, since there are multiple first sample images, when determining the main class center of a category, a reference image set corresponding to each category can be determined, and the reference image set corresponding to a category includes multiple first sample images with the same category label as the category; for each class center of each category, the reference similarity corresponding to the class center is determined based on the similarity between the initial class center matrix of the class center and the image features of each first sample image in the reference image set corresponding to the category; for each category, the class center with the largest reference similarity under the category is determined as the main class center of the category.
[0071] In this embodiment, for each class center of each category, the mean, maximum value or median of the similarity between the initial class center matrix of the class center and the image features of each first sample image in the reference image set corresponding to the category can be determined as the reference similarity corresponding to the class center.
[0072] In some embodiments, there are multiple second sample images, and the above-mentioned step S150 can also be: performing similarity calculation on the image features of each second sample image and the first class center matrix corresponding to the main class center of each category to obtain the second similarity between the image features of each second sample image and the main class center of each category; taking the category to which the main class center having the largest second similarity with the image features of the second sample image among multiple categories belongs as the predicted category of the second sample image; and determining the second loss based on the category label of the second sample image and the predicted category of the second sample image.
[0073] Among them, the method of calculating the second similarity between the image features of each second sample image and the main class center of each category can be similar to the aforementioned calculation process of calculating the similarity between the image features of the first sample image and the initial class center matrix corresponding to each class center, and the process of determining the second loss can also be similar to the aforementioned process of determining the first loss, which will not be described one by one here.
[0074] It should be understood that the above-mentioned method of obtaining the second loss is merely illustrative, and there are other methods for obtaining the second loss, which will not be described in detail in this embodiment.
[0075] Step S160: According to the second loss, adjust the parameters of the initially trained feature extraction network and the first category center matrix corresponding to the main category center of each category until the end condition of the second training stage is met, thereby obtaining the trained feature extraction network.
[0076] In this embodiment, the trained feature extraction network can be used for image classification.
[0077] In this embodiment, when the number of iterations of the initially trained feature extraction network reaches a second preset number, the number of cycles corresponding to the iteration cycle reaches a second preset number of cycles, or the second loss is less than a second preset loss threshold, it can be considered that the training of the initially trained feature extraction network has reached the second training stage termination condition. When the training of the initially trained feature extraction network reaches the second training stage termination condition, the initially trained feature extraction network after the last iteration adjustment is used as the trained feature extraction network. The second preset number of times, the second preset number of cycles, and the second preset loss threshold can be set according to task requirements and are not specifically limited here.
[0078] During the training process of the initially trained feature extraction network, the parameters of the initially trained feature extraction network and the first category center matrix corresponding to the main category center of each category are adjusted according to the second loss, so as to constrain the image features extracted by the initially trained feature extraction network for each second sample image to gradually approach the main category center of the corresponding category during the training process. In other words, even if the second sample image is a difficult sample, such as a non-frontal sample image, a sample image with low clarity, or a sample image with low pixels, the image features extracted by the initially trained feature extraction network for the second sample image as a difficult sample can be constrained to approach the main category center of the corresponding category, so that the trained feature extraction network can not only ensure the accuracy of feature extraction for normal images, but also improve the accuracy of feature extraction for low-quality images, thereby ensuring accurate classification of low-quality or difficult images.
[0079] The training method of the feature extraction network provided in the embodiment of the present application sets multiple class centers for each category, and adjusts the parameters of the feature extraction network to be trained and the initial class center matrix corresponding to the multiple class centers of each category according to the first loss determined based on the image features of the first sample image, the class label of the first sample image, and the initial class center matrix corresponding to the multiple class centers of each category. This can improve the accuracy of feature extraction of the feature extraction network to be trained while constraining the feature space corresponding to the multiple class centers of each category. Subsequently, the parameters of the preliminarily trained feature extraction network and the first class center matrix corresponding to the main class center of each class are adjusted according to the second loss determined by the image features of the second sample image, the category label of the second sample image, and the first class center moment corresponding to the main class center of each class, so that the image features extracted by the preliminarily trained extraction network gradually approach the main class center of each class during the training process. That is, even if the second sample image is a difficult sample, the features of the difficult sample can be moved closer to the main class center of the corresponding class, solving the problem of class center offset caused by the difficult sample. This can ensure that the feature extraction performance of the trained feature extraction network for normal samples in the sample image is improved while improving the feature extraction performance for low-quality difficult samples in the sample image, thereby improving the accuracy of the features extracted by the trained feature extraction network.
[0080] Referring to FIG4 , an embodiment of the present application provides a method for training a feature extraction network, the method comprising:
[0081] Step S210: extracting image features of the first sample image through the feature extraction network to be trained.
[0082] For a detailed description of step S210 , please refer to the detailed description of step S110 above, which will not be repeated here.
[0083] Step S220: For each category, calculate a first similarity between the image feature of the first sample image and the initial class center matrices corresponding to each of the multiple class centers of the category.
[0084] The process of calculating the similarity between image features and the class center matrix can be found in the previous detailed description and will not be repeated here.
[0085] Step S230: determining, among the multiple initial class center matrices of each class, the initial class center matrix having the greatest first similarity with the image features of the first sample image as the first reference class center matrix.
[0086] That is, the first reference class center matrix refers to an initial class center matrix having the greatest first similarity with the image features of the first sample image among the multiple class center matrices of the multiple categories.
[0087] Step S240: Determine a first sub-loss according to the first reference class center matrix and the class label of the first sample image.
[0088] The category to which the class center corresponding to the first reference class center matrix belongs can be used as the predicted category corresponding to the first sample image. The first sub-loss is used to reflect the difference between the predicted category corresponding to the first sample image and the category indicated by the category label of the first sample image. It can be understood that the smaller the difference between the predicted category corresponding to the first sample image and the category indicated by the category label of the first sample image, the higher the accuracy of the image features extracted by the feature extraction network to be trained.
[0089] Step S250: For each category, among the multiple initial class center matrices of the category, the initial class center matrix having the greatest first similarity with the image features of the first sample image is determined as the second reference class center matrix of the category;
[0090] For a category, the second reference class center matrix under the category refers to the initial class center matrix with the maximum first similarity between the image features of the first sample image under the category.
[0091] Step S260: Determine a second sub-loss according to a first similarity between a second reference class center matrix in each of the multiple classes and the image feature of the first sample image.
[0092] Specifically, the second sub-loss may be determined according to the sum or mean of the first similarities between the second reference class center matrix in each of the multiple categories and the image features of the first sample image.
[0093] Step S270: Determine a first loss based on the first sub-loss and the second sub-loss.
[0094] The first loss can be determined based on the first and second sub-losses by performing a weighted summation of the first and second sub-losses, or by determining the sub-loss with the largest loss value between the first and second sub-losses as the first loss. The setting can be adjusted based on actual needs.
[0095] In some embodiments, the above step S270 includes:
[0096] Step S270a: determining a first interval parameter and a first scaling factor based on a first cycle number corresponding to the iteration cycle in which the first sample image participates, wherein the first interval parameter and the first scaling factor are negatively correlated with the first cycle number.
[0097] The first cycle number refers to the cycle number corresponding to the iteration cycle in which the first sample image participates. That is, if the first sample image participates in the kth iteration cycle, the first cycle number corresponding to the iteration cycle in which the first sample image participates is k. It will be understood that different first sample images may participate in different iteration cycles and may also correspond to different first cycle numbers.
[0098] By determining the first interval parameter and the first scaling factor based on the first number of cycles corresponding to the iteration cycle in which the first sample image participates, the parameters can be adaptively set, and the supervision loss is relatively relaxed, ensuring that the feature distribution of normal samples is not affected. At the same time, the situation in which the feature extraction network is easily affected by difficult samples due to unstable gradients in the early stages of training is avoided. Moreover, the first interval parameter and the first scaling factor are negatively correlated with the first number of cycles. Thus, the larger the first number of cycles, the smaller the first interval parameter and the first scaling factor. As the number of iteration cycles increases, the first interval parameter and the first scaling factor are gradually reduced, and the loss requirement can be gradually increased.
[0099] Step S270b: Calculate the additive angular spacing loss based on the first spacing parameter, the first scaling factor, the first sub-loss, and the second sub-loss to obtain a first loss.
[0100] Specifically, the first loss may be calculated using a first additive angular interval loss calculation formula based on the first interval parameter, the first scaling factor, the first sub-loss, and the second sub-loss. The first additive angular interval loss calculation formula is as follows: Among them, lossearlystage is the first loss, cos(θ i,yi ) is the first sub-loss, which is determined based on the category label yi of the first sample image and the first reference class center matrix with the largest first similarity to the image feature of the first sample image among the multiple class center matrices of multiple categories. i,j) is the second sub-loss, which is determined based on the first similarity between the second reference class center matrix under each category in the multiple categories and the image features of the first sample image, s1 is the first interval parameter, and m1 is the first scaling factor. G represents the number of first sample images input during one iterative training process, θ i,yi Represents the angle between the category corresponding to the classification label of the first sample input image and the category corresponding to the reference class center matrix, θ i,j It represents the angle between the category corresponding to the classification label of the first sample input image and the category corresponding to the j-th class center. Specifically, θ i,j =arccos(max f (W jf x)),f=1,2....,F, where x represents the image feature, W jf represents the initial class center matrix of the f-th class center of the j-th class, max f (W jf x) refers to the maximum similarity between the image features and the initial class center matrix, and arccos() is the inverse cosine function.
[0101] In some possible implementations, the process of determining the first interval parameter and the first scaling factor may specifically be: determining the first coefficient based on the first cycle number corresponding to the iteration cycle in which the first sample image participates, the first coefficient being negatively correlated with the first cycle number; determining the first interval adjustment value based on the first coefficient, adding the first interval adjustment value to the first reference interval parameter to obtain the first interval parameter; determining the first scaling adjustment value based on the first coefficient, adding the first scaling adjustment value to the first reference scaling factor to obtain the first scaling factor.
[0102] For example, a first difference value can be obtained by subtracting a first preset number of cycles from a first number of cycles corresponding to an iteration cycle in which the first sample image participates; a ratio of the first difference value to the first preset number of cycles is used as a first coefficient; the first coefficient is multiplied by a first reference interval parameter to obtain a first interval adjustment value, and the first interval adjustment value is added to the first reference interval parameter to obtain a first interval parameter; the first coefficient is multiplied by a specified scaling factor to obtain a first scaling adjustment value, and the first reference scaling factor is added to the first scaling adjustment value to obtain a first scaling factor. The first preset number of cycles can be set as needed, ensuring that the difference between the first preset number of cycles and the first number of cycles is not less than zero. The specific values of the first reference interval parameter, the specified scaling factor, and the first reference scaling factor are not specifically limited here and can be set according to actual needs.
[0103] Step S280: According to the first loss, the parameters of the feature extraction network to be trained and the initial class center matrices corresponding to the multiple class centers of each category are adjusted until the end condition of the first training stage is reached, and the preliminary trained feature extraction network and the first class center matrices corresponding to the multiple class centers of each category are obtained.
[0104] Step S290: extracting image features of the second sample image through the preliminarily trained feature extraction network.
[0105] There are multiple second sample images.
[0106] Step S300: performing similarity calculation on the image features of the second sample image and the first class center matrix corresponding to the main class center of each class to obtain a second similarity between the image features of the second sample image and the main class center of each class.
[0107] Step S310: The category to which the main category center having the second greatest similarity with the image features of the second sample image belongs among the multiple categories is used as the predicted category of the second sample image.
[0108] Step S320: Determine a second loss based on the category label of the second sample image and the predicted category of the second sample image.
[0109] It should be understood that if the category label of the second sample image is consistent with the predicted category of the second sample image, the second loss is smaller; if the category label of the second sample image is inconsistent with the predicted category of the second sample image, the second loss is larger.
[0110] Considering that the parameters of the feature extraction network tend to be stable in the later stage of training, the main goal of the training stage becomes to make the difficult samples far away from the class center corresponding to the normal samples close to the class center of the normal samples. Therefore, it is necessary to use tighter supervision parameters. In the parameter adjustment process, an adaptive parameter setting method can be designed to gradually use a stricter supervision strategy in the later stage of training, while ensuring the recognition effect of normal data, while improving the compatibility of the model with difficult samples. Specifically, in some possible implementation methods, the above step S320 includes:
[0111] Step S320a: determining a second interval parameter and a second scaling factor based on a second cycle number corresponding to the iteration cycle in which the second sample image participates, wherein the second interval parameter and the second scaling factor are negatively correlated with the second cycle number.
[0112] The second cycle number refers to the cycle number corresponding to the iteration cycle in which the second sample image participates. That is, if the second sample image participates in the zth iteration cycle, then the second cycle number corresponding to the iteration cycle in which the second sample image participates is z. It will be understood that different second sample images may participate in different iteration cycles and may also correspond to different second cycle numbers.
[0113] The second interval parameter and the second scaling factor can be determined in the following manner: based on the second cycle number corresponding to the iteration cycle in which the second sample image participates, the second coefficient is determined, and the second coefficient is negatively correlated with the second cycle number; based on the second coefficient, a second interval adjustment value is determined, and the second reference interval parameter is added to the second interval adjustment value to obtain the second interval parameter; based on the second coefficient, a second scaling factor is determined, and the second reference scaling factor is added to the second scaling factor to obtain the second scaling factor.
[0114] Exemplarily, the second preset cycle number is calculated as a difference between the second cycle number corresponding to the iteration cycle in which the second sample image participates, thereby obtaining a second difference value, and the second difference value is compared with the second preset cycle number to obtain a second coefficient; the second coefficient is multiplied by the second reference interval parameter to obtain a second interval adjustment value, and the second reference interval parameter and the second interval adjustment value are added to obtain a second interval parameter; the second coefficient is multiplied by the set scaling factor to obtain a second scaling adjustment value, and the second reference scaling factor is added to the second scaling adjustment value to obtain a second scaling factor. The second preset cycle number can be set as needed, ensuring that the difference between the second preset cycle number and the second cycle number is not less than zero. The specific values of the second reference interval parameter, the set scaling factor, and the second reference scaling factor are not specifically limited here and can be set according to actual needs.
[0115] In some embodiments, to ensure the training effect, the first interval parameter during the training of the feature extraction network to be trained can be constrained to be less than or equal to the second interval parameter during the training of the feature extraction network for preliminary training, and the first scaling factor during the training of the feature extraction network to be trained can be constrained to be less than or equal to the second scaling factor during the training of the feature extraction network for preliminary training.
[0116] Step S320b: Calculate the additive angular separation loss based on the second separation parameter, the second scaling factor, the category label of the second sample image, and the predicted category of the second sample image to obtain a second loss.
[0117] Specifically, the second additive angular interval loss calculation formula can be used to calculate the additive angular interval loss based on the second interval parameter, the second scaling factor, the category label of the second sample image, and the predicted category of the second sample image to obtain the second loss, wherein the second additive angular interval loss calculation formula is as follows: Among them, loss latestage is the second loss, H is the batch size, that is, the number of second sample images input in a batch, cos(θ yi ) is the third sub-loss, which is determined based on the similarity between the image features of the second sample image and the main class center of each category, s2 is the second interval parameter, and m2 is the second scaling factor. yi is the angle between the predicted category of the second sample image and the category corresponding to the category label of the second sample image.
[0118] Step S330: According to the second loss, adjust the parameters of the initially trained feature extraction network and the first category center matrix corresponding to the main category center of each category until the end condition of the second training stage is met, thereby obtaining the trained feature extraction network.
[0119] Referring to FIG. 5 , an embodiment of the present application further provides an image classification method, which can be applied to the above-mentioned electronic device. The method includes:
[0120] Step S410: Obtain an image to be classified.
[0121] The image to be classified may be any image that needs to be classified, and its acquisition process may be similar to the method of acquiring the first sample image and the second sample image in the aforementioned embodiment, which is not specifically limited here.
[0122] It should be understood that to ensure the accuracy of features subsequently extracted from the image to be classified, step S410 may include obtaining an initial image and performing preprocessing on the initial image, such as denoising, enhancement, filtering, and the like. After obtaining the preprocessed initial image, different processing operations may be performed depending on the classification task. For example, if the classification task is object recognition, the region containing the object in the classification task may be cropped, and the object in the classification task may be scaled.
[0123] In some embodiments, the image to be classified is a palm print image, and step S410 includes:
[0124] Step S412: Acquire a hand image.
[0125] Step S414: performing key point detection on the hand image to obtain key points of finger joints in the hand image.
[0126] Step S416: Based on the finger joint key points in the hand image, a palm print pixel area is intercepted from the hand image as a palm print image.
[0127] The specific process of obtaining the palm print image can be found in the previous description, which will not be repeated here.
[0128] Step S420: The trained feature extraction network obtained by using the feature extraction network training method performs feature extraction on the image to be classified to obtain target image features.
[0129] Step S430: Determine the classification result of the image to be classified according to the target image features.
[0130] It should be understood that when the application scenario corresponding to the trained feature extraction network is different, the method of determining the classification result of the image to be classified is different. If the application scenario corresponding to the trained feature extraction network is an identification and authentication scenario, the classification result of the above-mentioned image to be classified is authentication passed or authentication failed; if the application scenario corresponding to the trained feature extraction network is a multi-classification scenario, the classification result of the above-mentioned image to be classified is which category the image to be classified belongs to; if the trained feature extraction network is applied in an image anomaly recognition scenario, the classification result of the above-mentioned image to be classified is whether the image to be classified is normal or abnormal. It should be understood that the application scenario of the above-mentioned trained feature extraction network is merely illustrative, and the classification method determined and the classification results obtained in different application scenarios are different.
[0131] If the trained feature extraction network is applied in an identification and authentication scenario, the above-mentioned step S430 may be that the classification result of the image to be classified includes the authentication result. Specifically, the above-mentioned step S430 may be that the target image feature is matched with the preset reference image feature. If there is a preset reference image feature that matches the target image feature, authentication information including authentication success is generated, or the authentication information associated with the preset reference image feature is used as the authentication information of the image to be classified.
[0132] In this embodiment, the above step S430 may specifically include:
[0133] Step S432: performing similarity calculation on the target image feature and a plurality of reference image features in a preset database to obtain the similarity between the target image feature and each reference image feature.
[0134] Step S434: determining the target reference image feature having the highest similarity to the target image feature based on the similarities between the target image feature and each reference image feature.
[0135] Step S436: using the authentication information associated with the target reference image features as the authentication result of the image to be classified.
[0136] In this way, if the authentication is passed, subsequent operations can be performed, such as sending certain application unlocking data, payment, device unlocking, gate release, smart lock opening, etc.
[0137] In some possible implementations, after using the authentication information associated with the target reference image features as the authentication result of the image to be classified, the method further includes: performing payment processing based on the authentication result of the image to be classified.
[0138] If the trained feature extraction network is applied in a multi-classification task scenario, different categories each have a main class center, and the main class center can be determined according to the aforementioned feature extraction network training method. The above step S430 can be to calculate the similarity between the image features of the image to be classified and the main class center corresponding to each category in the multi-classification task, so as to determine the classification result of the image to be classified according to the similarity between the image features and the main class center of each category.
[0139] It should be understood that the above-mentioned method of determining the classification result of the image to be classified is merely illustrative, and there may be more confirmation methods, which will not be described in detail in this embodiment.
[0140] As shown in Figure 6, this embodiment of the present application provides a method for training a feature extraction network. The trained feature extraction network is used to extract palmprint features of different users. The trained feature extraction network is used to identify and authenticate the palmprint image to be identified, and payment is performed after the identification and authentication are successful. The specific training process and application process are as follows:
[0141] Training phase:
[0142] First, a plurality of initial images including palm images are obtained, each of which corresponds to a piece of label information, and the label information is used to characterize the object to which the palm image belongs.
[0143] For each initial image, the target detection algorithm (such as the yo lov2 detection algorithm) is used to detect the key points in the initial image, and the first finger gap key point A between the index finger and the middle finger, the second finger gap key point B between the middle finger and the ring finger, and the third finger gap key point C between the ring finger and the little finger are obtained; then, an image coordinate system is established based on the first finger gap key point A, the second finger gap key point B, and the third finger gap key point C in the initial image. The line between the first finger gap key point A and the third finger gap key point C is the horizontal axis (x-axis) of the image coordinate system, and the line passing through the second finger gap key point B and perpendicular to the horizontal axis is the image coordinate system. The vertical axis (y-axis) of the coordinate system and the intersection of the horizontal axis and the vertical axis are the origin of the image coordinate system; the point in the initial image that is located on the vertical axis of the image coordinate system and at a specified distance from the origin of the image coordinate system is used as the palm print center point D in the initial image, and the specified distance is determined according to the distance between the first finger joint key point A and the third finger joint key point C. The palm print center point D and the second finger joint key point B are located on both sides of the horizontal axis respectively; a sample image is intercepted from the initial image based on the distance between the palm print center point D and the first finger joint key point A and the third finger joint key point B.
[0144] In this way, a first sample image and a second sample image can be obtained, wherein there are multiple first sample images and multiple second sample images, and the first sample image and the second sample image can be the same or different.
[0145] After obtaining the first sample image and the second sample image, the first sample image and the second sample image may be scaled to the same size to participate in subsequent feature extraction network training.
[0146] When training the feature extraction network, the feature extraction network is applied to a classification task, the number of categories in the classification task is set to E, the initial class center matrices of multiple class centers of each category are set to zero, and the number of multiple class centers corresponding to each category is the same as F. The initial class center matrix corresponding to each class center is a fully connected linear matrix and has the same length as N. The category corresponding to the category label of the sample image should belong to one of the multiple categories in the classification task.
[0147] During the first stage of training, the number of first sample images input for each iteration is G. The image features of the first sample images are extracted using the feature extraction network to be trained. Then, for each category, a first similarity is calculated between the image features of the first sample image and the initial class center matrices corresponding to the multiple class centers of that category. The initial class center matrix with the greatest first similarity to the image features of the first sample image among the multiple initial class center matrices of each category is determined as the first reference class center matrix. A first sub-loss is determined based on the first reference class center matrix and the category label of the first sample image. The initial class center matrix with the greatest first similarity to the image features of the first sample image among the multiple initial class center matrices of each category is determined as the second reference class center matrix for that category. A second sub-loss is determined based on the first similarity between the second reference class center matrix for each category and the image features of the first sample image. Then, a first interval parameter and a first scaling factor are determined based on the first number of iteration cycles corresponding to the first sample image, and the first interval parameter and the first scaling factor are negatively correlated with the first number of cycles. An additive angular interval loss is calculated based on the first interval parameter, the first scaling factor, the first sub-loss, and the second sub-loss to obtain a first loss.
[0148] Specifically, the first loss can be calculated based on the following formula: Among them, lossearlystage is the first loss, cos(θ i,yi ) is the first sub-loss, which is determined based on the category label yi of the first sample image and the first reference class center matrix with the largest first similarity to the image feature of the first sample image among the multiple class center matrices of multiple categories.i,j ) is the second sub-loss, which is determined based on the first similarity between the second reference class center matrix under each category in the multiple categories and the image features of the first sample image, s1 is the first interval parameter, and m1 is the first scaling factor. G represents the number of first sample images input during one iterative training process, θ i,yi Represents the angle between the category corresponding to the classification label of the first sample input image and the category corresponding to the reference class center matrix, θ i,j It represents the angle between the category corresponding to the classification label of the first sample input image and the category corresponding to the j-th class center. Specifically, θ i,j =arccos(max f (W jf x)),f=1,2....,F, where x represents the image feature, W jf The initial class center matrix representing the f-th subclass center of the j-th class.
[0149] In some embodiments, the first training stage is a stage in which the iteration period in which the first sample image participates is less than 10. Since the gradient of model learning is unstable and easily affected by difficult samples, an adaptive parameter setting method can be adopted, and the supervision loss is relatively loose to ensure that the feature distribution of normal samples is not affected. The range of s1 is limited to [24, 48], and the range of m1 is limited to [0.3, 0.5]. The actual first interval parameter and the first scaling factor are calculated as follows: s1 = 24 + (10-epoch1) / 10 * 24; m1 = 0.3 + (10-epoch1) / 10 * 0.2, where epoch1 refers to the number of iteration periods in which the first sample image participates.
[0150] After calculating the first loss, the parameters of the feature extraction network to be trained and the initial class center matrices corresponding to the multiple class centers of each category are adjusted through gradient feedback until the end condition of the first training stage is reached, and the preliminary trained feature extraction network and the first class center matrices corresponding to the multiple class centers of each category are obtained, so that for each category, the image features of the first sample image belonging to the category can be mapped to one of the multiple class centers corresponding to the category.
[0151] During the training process of the second stage, there are multiple second sample images. The image features of the second sample images are extracted through the preliminarily trained feature extraction network, and the similarity between the image features of each second sample image and the first class center matrix corresponding to the main class center of each category is calculated to obtain the second similarity between the image features of each second sample image and the main class center of each category; the category to which the main class center having the largest second similarity with the image features of the second sample image among multiple categories belongs is used as the predicted category of the second sample image; and the second loss is determined based on the category label of the second sample image and the predicted category of the second sample image.
[0152] Specifically, the second loss can be calculated using the following loss function: Among them, loss latestage is the second loss, H is the batch size, that is, the number of second sample images input in a batch, cos(θ yi ) is the third sub-loss, which is determined based on the similarity between the image features of the second sample image and the main class center of each category, s2 is the second interval parameter, and m2 is the second scaling factor. yi is the angle between the predicted category and the category corresponding to the category label.
[0153] In some embodiments, the second training stage refers to a stage in which the iteration period in which the second sample image participates is less than 10. In the second training stage, since the model parameters tend to be stable, the main goal becomes to move the difficult samples that were originally far away from the class center of the normal samples closer to the class center of the normal samples. Therefore, tighter supervision parameters need to be used. An adaptive parameter setting method is also designed to limit the range of s2 to [48, 64] and the range of m2 to [0.5, 0.7]. The actual second interval parameter and the second scaling factor are calculated as follows: s2 = 48 + (10-epch2) / 10*16; m2 = 0.5 + (10-epoch2) / 10*0.2, where epoch2 refers to the number of iteration periods in which the second sample image participates.
[0154] Through the above two stages of training, the feature extraction performance of the trained feature extraction network for normal samples in the sample images can be improved while the feature extraction performance of low-quality difficult sample images in the sample images can be improved, thereby improving the accuracy of the features extracted by the trained target feature extraction network.
[0155] Application stage:
[0156] Referring to FIG7 , when the trained feature extraction network is used for palmprint recognition payment, a hand image of the target object can be captured through the terminal payment device camera; the detection model detects the key points between the target object's fingers; the palm region of interest is extracted based on the hand image and the key point positions; the trained feature extraction network is used to extract the palm print feature encoding vector corresponding to the palm region of interest; and the cosine similarity between the extracted palm print feature encoding vector and the base database features is calculated. The cosine similarity calculation formula is as follows: in, is the palmprint feature encoding vector, For any base database feature, the ID information associated with the base database feature with the highest similarity is taken as the ID information of the target object, and the ID information of the target object is returned to the terminal payment device as the recognition result, so that payment processing can be performed based on the recognition result.
[0157] Compared to facial recognition, palmprint recognition for payment processing is more discreet and therefore more conducive to protecting user privacy. It is also unaffected by factors such as masks, makeup, and sunglasses. Therefore, palmprint recognition technology has broad application prospects in commercial scenarios such as mobile payments and identity verification.
[0158] In order to verify the effectiveness of the feature extraction network trained in this application, a dataset of 1000 identities, each corresponding to 50 normal images and 50 difficult images, was used to verify the effectiveness of the feature extraction network trained using the feature extraction network training method of this application and the feature extraction network trained using the training method of the existing method. The verification results are shown in Table 1.
[0159] According to Table 1, the error recognition rate of the existing method on the difficult data set is significantly higher than that on the normal data set, while the feature extraction network trained using the training method of the present application can perform well on both the normal data set and the difficult data set. It can be seen that the training method of the feature extraction network of the present application can improve the accuracy of the recognition of difficult palmprint data.
[0160] Please refer to Figure 8. Another embodiment of the present application provides a training device 500 for a feature extraction network. The training device 500 for the feature extraction network includes a first feature extraction module 510, a first loss determination module 520, a first adjustment module 530, a second feature extraction module 540, a second loss determination module 550, and a second adjustment module 560. The first feature extraction module 510 is used to extract image features of a first sample image through the feature extraction network to be trained; the first loss determination module 520 is used to determine a first loss based on the image features of the first sample image, the category label of the first sample image, and the initial class center matrix corresponding to multiple class centers of each category in multiple categories; the first adjustment module 530 is used to adjust the parameters of the feature extraction network to be trained and the multiple class centers of each category according to the first loss. The initial class center matrix corresponding to each class center is used to obtain the preliminarily trained feature extraction network and the first class center matrix corresponding to multiple class centers of each category; the second feature extraction module 540 is used to extract the image features of the second sample image through the preliminarily trained feature extraction network; the second loss determination module 550 is used to determine the second loss according to the image features of the second sample image, the category label of the second sample image and the first class center matrix corresponding to the main class center of each category; the main class center of a category is one of the multiple class centers of the category; the second adjustment module 560 is used to adjust the parameters of the preliminarily trained feature extraction network and the first class center matrix corresponding to the main class center of each category according to the second loss to obtain the trained feature extraction network, wherein the trained feature extraction matrix can be used for image classification.
[0161] Please refer to Figure 9. An embodiment of the present application provides an image classification device 600, which includes an image acquisition module 610, a third feature extraction module 620, and a classification result determination module 630. The image acquisition module 610 is used to acquire the image to be classified; the third feature extraction module 620 is used to perform feature extraction on the image to be classified using the trained feature extraction network obtained by the training device of the above-mentioned feature extraction network to obtain target image features; the classification result determination module 630 is used to determine the classification result of the image to be classified based on the target image features.
[0162] Each module in the above-mentioned device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules. It should be noted that the device embodiment in this application corresponds to the aforementioned method embodiment. The specific principles in the device embodiment can be found in the contents of the aforementioned method embodiment, which will not be repeated here.
[0163] An electronic device provided by this application will be described below with reference to FIG10 .
[0164] Please refer to Figure 10. Based on the training method of the feature extraction network provided in the above embodiment, the embodiment of the present application also provides another electronic device 100 including a processor 102 that can execute the above method. The electronic device 100 can be a server or a terminal device, and the terminal device can be a smart phone, tablet computer, computer or portable computer and other devices.
[0165] The electronic device 100 further includes a memory 104 . The memory 104 stores a program capable of executing the contents of the aforementioned embodiments, and the processor 102 can execute the program stored in the memory 104 .
[0166] The processor 102 may include one or more cores for processing data and a message matrix unit. The processor 102 utilizes various interfaces and circuits to connect various components within the electronic device 100. It executes instructions, programs, code sets, or instruction sets stored in the memory 104, and accesses data stored in the memory 104 to perform various functions and process data within the electronic device 100. In some embodiments, the processor 102 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 102 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 102 and may be implemented separately via a communication chip.
[0167] The memory 104 may include a random access memory (RAM) or a read-only memory (ROM). The memory 104 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, and the like. The data storage area may also store data (e.g., training data or images to be classified) acquired by the electronic device 100 during use.
[0168] The electronic device 100 may also include a network module and a screen. The network module is used to receive and send electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, and thus communicate with a communication network or other devices, such as communicating with an audio playback device. The network module may include various existing circuit components for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, user identity modules (SIM) cards, memories, and the like. The network module can communicate with various networks such as the Internet, corporate intranets, wireless networks, or communicate with other devices via wireless networks. The above-mentioned wireless networks may include cellular telephone networks, wireless local area networks, or metropolitan area networks. The screen can display interface content and perform data interaction, such as displaying the molecular property prediction results of the audio to be identified, and recording audio through the screen.
[0169] In some embodiments, electronic device 100 may further include a peripheral interface 106 and at least one peripheral device. Processor 102, memory 104, and peripheral interface 106 may be connected via a bus or signal lines. Each peripheral device may be connected to the peripheral interface via a bus, signal lines, or circuit boards. Specifically, the peripheral devices include, for example, a radio frequency component 108.
[0170] The peripheral interface 106 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 102 and the memory 104. In some embodiments, the processor 102, the memory 104, and the peripheral interface 106 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 102, the memory 104, and the peripheral interface 106 can be implemented on separate chips or circuit boards, which is not limited in this embodiment of the present application.
[0171] The RF component 108 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF component 108 communicates with communication networks and other communication devices via electromagnetic signals. The RF component 108 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the RF component 108 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF component 108 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF component 108 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0172] The embodiment of the present application also provides a structural block diagram of a computer-readable storage medium. The computer-readable storage medium stores program code, which can be called by a processor to execute the method described in the above method embodiment.
[0173] The computer-readable storage medium may be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has storage space for program codes for executing any of the method steps described above. These program codes can be read from or written to one or more computer program products. The program codes can be compressed, for example, in an appropriate form.
[0174] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods described in the various optional implementations described above.
[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for training a feature extraction network, performed by an electronic device, the method comprising: In each iteration of the first training phase, the following steps are performed: Extracting image features of the first sample image by using a feature extraction network to be trained; determining a first loss according to an image feature of the first sample image, a category label of the first sample image, and initial category center matrices corresponding to a plurality of category centers of each category in a plurality of categories; According to the first loss, adjusting the parameters of the feature extraction network to be trained and the initial class center matrices corresponding to the multiple class centers of each category, until the end condition of the first training stage is reached, and obtaining the initially trained feature extraction network and the first class center matrices corresponding to the multiple class centers of each category; In each iteration of the second training phase, the following steps are performed: Extracting image features of the second sample image through a preliminarily trained feature extraction network; Determine a second loss according to the image feature of the second sample image, the category label of the second sample image, and the first category center matrix corresponding to the main category center of each category; the main category center of a category is one of the multiple category centers of the category; According to the second loss, the parameters of the initially trained feature extraction network and the first category center matrix corresponding to the main category center of each category are adjusted until the end condition of the second training stage is reached to obtain a trained feature extraction network, and the trained feature extraction network is used for image classification.
2. The method according to claim 1, wherein: The determining of the first loss according to the image feature of the first sample image, the category label of the first sample image, and the initial category center matrices corresponding to the multiple category centers of each category in the multiple categories includes: For each category, calculating a first similarity between an image feature of a first sample image and an initial class center matrix corresponding to a plurality of class centers of the category; A first loss is determined based on a first similarity between the image feature of the first sample image and initial class center matrices corresponding to the multiple class centers of each category.
3. The method according to claim 2, wherein: The determining of the first loss based on the first similarity between the image feature of the first sample image and the initial class center matrices corresponding to the plurality of class centers of each category comprises: Determine, among the multiple initial class center matrices of each class, an initial class center matrix having the greatest first similarity with the image feature of the first sample image as a first reference class center matrix; Determining a first sub-loss according to the first reference class center matrix and the class label of the first sample image; For each category, determining an initial class center matrix having the greatest first similarity with the image features of the first sample image among a plurality of initial class center matrices of the category as a second reference class center matrix of the category; determining a second sub-loss according to a first similarity between a second reference class center matrix for each of the multiple classes and an image feature of the first sample image; Based on the first sub-loss and the second sub-loss, a first loss is determined.
4. The method according to claim 3, wherein: The determining a first loss based on the first sub-loss and the second sub-loss includes: Determine a first interval parameter and a first scaling factor based on a first cycle number corresponding to an iteration cycle in which the first sample image participates, wherein the first interval parameter and the first scaling factor are negatively correlated with the first cycle number; An additive angular spacing loss is calculated based on the first spacing parameter, the first scaling factor, the first sub-loss, and the second sub-loss to obtain a first loss.
5. The method according to claim 4, wherein: The determining of a first interval parameter and a first scaling factor based on a first cycle number corresponding to an iteration cycle in which the first sample image participates includes: Determine a first coefficient based on a first cycle number corresponding to an iteration cycle in which the first sample image participates, wherein the first coefficient is negatively correlated with the first cycle number; Determine a first interval adjustment value according to the first coefficient, and add the first interval adjustment value to a first reference interval parameter to obtain a first interval parameter; A first scaling adjustment value is determined according to the first coefficient, and the first scaling adjustment value is added to a first reference scaling coefficient to obtain a first scaling coefficient.
6. The method according to any one of claims 1 to 5, wherein: There are multiple first sample images. Before determining the second loss based on the image features of the second sample images, the category labels of the second sample images, and the first category center matrix corresponding to the main category centers of each category, the method further includes: Determine a reference image set corresponding to each category, wherein the reference image set corresponding to a category includes a plurality of first sample images having the same category label as the category; For each category, according to the similarity between the initial class center matrix of each class center under the category and the image features of each first sample image in the reference image set corresponding to the category, determine the reference similarity corresponding to the class center under the category; For each category, the class center with the largest reference similarity under the category is determined as the main class center of the category.
7. The method according to any one of claims 1 to 6, wherein: There are multiple second sample images, and determining the second loss according to the image features of the second sample images, the category labels of the second sample images, and the first category center matrix corresponding to the main category centers of each category includes: Performing similarity calculation on the image features of each of the second sample images and the first class center matrix corresponding to the main class center of each class, to obtain a second similarity between the image features of each of the second sample images and the main class center of each class; taking the category to which the main category center having the greatest second similarity with the image feature of the second sample image belongs among the multiple categories as the predicted category of the second sample image; A second loss is determined based on the category label of the second sample image and the predicted category of the second sample image.
8. The method according to claim 7, wherein: The determining the second loss based on the category label of the second sample image and the predicted category of the second sample image includes: Determine a second interval parameter and a second scaling factor based on a second cycle number corresponding to the iteration cycle in which the second sample image participates, wherein the second interval parameter and the second scaling factor are negatively correlated with the second cycle number; An additive angular interval loss is calculated based on a second interval parameter, a second scaling factor, the category label of the second sample image, and the predicted category of the second sample image to obtain a second loss.
9. The method according to claim 8, wherein: The determining of the second interval parameter and the second scaling factor based on the second cycle number corresponding to the iteration cycle in which the second sample image participates includes: Determine a second coefficient based on a second cycle number corresponding to an iteration cycle in which the second sample image participates, wherein the second coefficient is negatively correlated with the second cycle number; Determine a second interval adjustment value based on the second coefficient, and add a second reference interval parameter to the second interval adjustment value to obtain a second interval parameter; A second scaling adjustment value is determined based on the second coefficient, and a second reference scaling coefficient is added to the second scaling adjustment value to obtain a second scaling coefficient.
10. An image classification method, performed by an electronic device, the method comprising: Get the image to be classified; Using the trained feature extraction network obtained as claimed in any one of claims 1 to 9 to perform feature extraction on the image to be classified to obtain target image features; According to the target image features, a classification result of the image to be classified is determined.
11. The method according to claim 10, wherein: The classification result of the image to be classified includes an authentication result; Determining the classification result of the image to be classified according to the target image feature includes: Calculating the similarity between the target image feature and multiple reference image features in a preset database to obtain the similarity between the target image feature and each reference image feature; Determining, based on the similarities between the target image feature and each reference image feature, a target reference image feature having the highest similarity to the target image feature; The authentication information associated with the target reference image feature is used as the authentication result of the image to be classified.
12. The method according to claim 11, wherein: After taking the authentication information associated with the target reference image feature as the authentication result of the image to be classified, the method further includes: Payment processing is performed based on the authentication result of the image to be classified.
13. The method according to any one of claims 10 to 12, wherein: The image to be classified is a palm print image, and obtaining the image to be classified includes: Acquire a hand image; Performing key point detection on the hand image to obtain finger gap key points in the hand image; Based on the finger joint key points in the hand image, a palm print pixel area is intercepted from the hand image as the palm print image.
14. A training device for a feature extraction network, the device comprising: A first feature extraction module, used for extracting image features of a first sample image through a feature extraction network to be trained; A first loss determination module, configured to determine a first loss according to an image feature of the first sample image, a category label of the first sample image, and initial class center matrices corresponding to a plurality of class centers of each category in a plurality of categories; A first adjustment module is used to adjust the parameters of the feature extraction network to be trained and the initial class center matrix corresponding to the multiple class centers of each category according to the first loss, so as to obtain the initially trained feature extraction network and the first class center matrix corresponding to the multiple class centers of each category; A second feature extraction module, used to extract image features of a second sample image through the preliminarily trained feature extraction network; A second loss determination module is used to determine a second loss according to the image features of the second sample image, the category label of the second sample image, and the first category center matrix corresponding to the main category center of each category; the main category center of a category is one of the multiple category centers of the category; The second adjustment module is used to adjust the parameters of the initially trained feature extraction network and the first category center matrix corresponding to the main category center of each category according to the second loss to obtain a trained feature extraction network, and the trained feature extraction network is used for image classification.
15. An image classification device, comprising: An image acquisition module, used to acquire images to be classified; A third feature extraction module, configured to extract features of the image to be classified using the trained image feature extraction network obtained as claimed in any one of claims 1 to 9, to obtain target image features; The classification result determination module is used to determine the classification result of the image to be classified according to the target image features.
16. An electronic device comprising: one or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method as described in any one of claims 1-9 or 10-13.
17. A computer-readable storage medium storing program code, wherein the program code can be called by a processor to execute the method according to any one of claims 1-9 or 10-13.
18. A computer program product, comprising computer instructions, wherein the computer instructions are stored in a computer-readable storage medium, and when the computer instructions are executed, the method according to any one of claims 1-9 or 10-13 is implemented.
Citation Information
Patent Citations
Face recognition model construction method, face recognition method and related devices
CN111259738A
Video classification method, electronic equipment and storage medium
CN112101091A
Method and device for training image feature extraction model and extracting image features
CN113255694A
Neural network model training method and device
CN114139704A
Feature extraction network training method, feature extraction network classification method, feature extraction network training device, feature extraction network classification device and electronic equipment
CN117152567A
Cited By
Model training method and device
CN120747855A