Target recognition model training method, target recognition method and electronic equipment
By adding training samples and domain losses of low-quality domains to the training of the target recognition model and fusing the feature layers of the two recognition models, the problem of poor recognition effect of target recognition model in complex environments is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111600292.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-24
AI Technical Summary
The existing target recognition model has poor recognition effect in complex environments and is affected by interference factors, resulting in a degradation of recognition performance.
By obtaining the preset recognition model and sample data set, including high-quality and low-quality images, performing the first and second samples, two recognition models are trained, and their feature layers are fused to form a target recognition model.
By adding training samples and domain losses of low-quality domains to the training, the features of low-quality images are brought close to the features of high-quality images, thereby improving the recognition effect and combining the advantages of the two models to improve the recognition accuracy.
Smart Images

Figure CN114299369B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target recognition, and in particular to a training method of a target recognition model, a target recognition method and electronic equipment. Background Art
[0002] Target recognition technology has been widely used. For example, face recognition technology has been widely used in security, finance and other fields. It uses computers to analyze face images to obtain effective recognition information for identity recognition. In order to ensure the recognition effect, the target recognition model currently used is generally trained with high-quality and low-quality images. In actual application scenarios, the environment will be more complex, which will bring many interference factors to target recognition, thereby affecting its recognition performance and resulting in poor recognition effect. Summary of the invention
[0003] In view of this, an embodiment of the present invention provides a training method for a target recognition model, a target recognition method and an electronic device to solve the problem that the existing target recognition model has poor recognition effect in a complex environment.
[0004] According to a first aspect, an embodiment of the present invention provides a method for training a target recognition model, comprising:
[0005] Acquire a preset recognition model and acquire a sample data set, wherein the preset recognition model is obtained by training based on a first quality image, and the sample data set includes a first quality image and a second quality image, wherein the quality of the first quality image is higher than the quality of the second quality image;
[0006] Performing a first sampling on the first quality image and the second quality image in the sample data set to obtain a first sampling data set, and training the preset recognition model based on the first sampling data set to determine a first recognition model;
[0007] Performing a second sampling on the first quality image and the second quality image in the sample data set to obtain a second sampling data set, and training the preset recognition model based on the second sampling data set to determine a second recognition model;
[0008] The first recognition model and the second recognition model are fused at the feature level to determine a target recognition model.
[0009] The training method of the target recognition model provided by the embodiment of the present invention adds training samples of low-quality domains to the training because the samples in complex environments are generally second-quality images, and combines the training method and domain loss to make the features of the second-quality images close to the features of the first-quality images, so that the recognition effect is close to the recognition effect of the first-quality images, and at the same time, the two complementary models are combined to ensure that the recognition effect of the obtained target recognition model for both the first-quality images and the second-quality images is improved; and the training process is based on a preset recognition model, and the two recognition models are trained on this basis, which can improve the recognition effect without affecting the training efficiency.
[0010] In combination with the first aspect, in a first implementation of the first aspect, the first sampling of the first quality image and the second quality image in the sample data set to obtain the first sample data set includes:
[0011] Obtaining the number of input samples of the preset recognition model;
[0012] Determining the number of the first quality images and the second quality images in each training based on the number of the input samples in a ratio of 1:1;
[0013] Based on the determined number of the first quality images and the second quality images, sampling is performed in the sample data set to obtain the first sample data set for each training.
[0014] The training method of the target recognition model provided by an embodiment of the present invention extracts the first quality image and the second quality image in a 1:1 ratio to form a first sampling data set used in each training, and uses the first quality image as the original domain and the second quality image as the target domain, so that the recognition effect of the second quality image is close to that of the first quality image.
[0015] In combination with the first aspect or the first implementation manner of the first aspect, in the second implementation manner of the first aspect, performing a second sampling on the first quality image and the second quality image in the sample data set to obtain a second sample data set includes:
[0016] Obtaining the number of input samples of the preset recognition model and obtaining the ratio of the number of the first quality images to the second quality images in the sample data set;
[0017] Determining the number of the first quality images and the second quality images in each training based on the number of input samples and according to the number ratio;
[0018] Based on the determined number of the first quality images and the second quality images, sampling is performed in the sample data set to obtain the second sample data set for each training.
[0019] The training method of the target recognition model provided by the embodiment of the present invention determines the second sampling data set based on the quantity ratio of the first quality image and the second quality image in the sample data set, and the two quality images can be used to complement each other during subsequent training.
[0020] In combination with the first aspect, in a third implementation manner of the first aspect, the step of fusing the first recognition model with the second recognition model at a feature level to determine a target recognition model includes:
[0021] Connecting the feature layers in the first recognition model and the second recognition model to the normalization layer respectively;
[0022] The two normalization layers are fused to determine the target recognition model.
[0023] The training method of the target recognition model provided by the embodiment of the present invention performs fusion after normalizing connection at the feature layer, so that the obtained target recognition model can combine the advantages of the two models and improve the recognition accuracy of the target recognition model.
[0024] In combination with the first aspect, in a fourth implementation of the first aspect, obtaining a preset recognition model includes:
[0025] Inputting the first quality image into a recognition model to obtain a feature layer output of the recognition model and a classification layer output of the recognition model;
[0026] Corresponding loss function calculations are performed based on the feature layer output and the classification layer output respectively to update the parameters of the recognition model and determine the preset recognition model.
[0027] The target recognition model training method provided by the embodiment of the present invention calculates the loss function at the feature layer and the classification layer respectively to update the parameters of the recognition model, thereby ensuring the accuracy of the trained preset recognition model.
[0028] According to the second aspect, an embodiment of the present invention further provides a target recognition method, including:
[0029] Obtain an image to be recognized;
[0030] Inputting the image to be identified into a target recognition model to determine features to be matched of the image to be identified, wherein the target recognition model is trained according to the first aspect of the present invention or the target recognition model training method described in any embodiment of the first aspect;
[0031] The to-be-matched features are matched with each target feature to determine the target recognition result corresponding to the to-be-recognized image.
[0032] The target recognition method provided by the embodiment of the present invention ensures that the target recognition model obtained by combining two complementary models will improve the recognition effect of the first quality image and the second quality image. Based on this, the target recognition model is used to identify the image to be identified, which can improve the accuracy of the features to be matched, thereby improving the accuracy of the target recognition result.
[0033] According to a third aspect, an embodiment of the present invention provides a training device for a target recognition model, comprising:
[0034] A first acquisition module is used to acquire a preset recognition model and a sample data set, wherein the preset recognition model is obtained by training based on a first quality image, and the sample data set includes a first quality image and a second quality image, wherein the quality of the first quality image is higher than that of the second quality image;
[0035] A first training module, configured to perform a first sampling on the first quality image and the second quality image in the sample data set to obtain a first sampling data set, and train the preset recognition model based on the first sampling data set to determine a first recognition model;
[0036] A second training module, configured to perform a second sampling on the first quality image and the second quality image in the sample data set to obtain a second sampling data set, and train the preset recognition model based on the second sampling data set to determine a second recognition model;
[0037] The fusion module is used to perform feature-layer fusion on the first recognition model and the second recognition model to determine a target recognition model.
[0038] According to a fourth aspect, an embodiment of the present invention further provides a target recognition device, including:
[0039] A second acquisition module is used to acquire an image to be identified;
[0040] a determination module, used for inputting the image to be identified into a target recognition model to determine features to be matched of the image to be identified, wherein the target recognition model is trained according to the first aspect of the present invention or the training method of the target recognition model described in any embodiment of the first aspect;
[0041] The matching module is used to match the to-be-matched features with each target feature to determine the target recognition result corresponding to the to-be-recognized image.
[0042] According to the fifth aspect, an embodiment of the present invention provides an electronic device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the training method of the target recognition model described in the first aspect or any one of the embodiments of the first aspect, or executes the target recognition method described in the second aspect or any one of the embodiments of the second aspect by executing the computer instructions.
[0043] According to the sixth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the training method of the target recognition model described in the first aspect or any one of the embodiments of the first aspect, or execute the target recognition method described in the second aspect or any one of the embodiments of the second aspect.
[0044] It should be noted that the corresponding beneficial effects of the target recognition model training device, target recognition device, electronic device and computer-readable storage medium provided in the embodiments of the present invention can be found in the corresponding description of the target recognition model training method or the target recognition method above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0046] Figure 1 is a flow chart of a method for training a target recognition model according to an embodiment of the present invention;
[0047] Figure 2a is a schematic diagram of the structure of a preset recognition model according to an embodiment of the present invention;
[0048] Figure 2b is a schematic diagram of a training structure of a preset recognition model according to an embodiment of the present invention;
[0049] Figure 3 is a flow chart of a method for training a target recognition model according to an embodiment of the present invention;
[0050] Figure 4 is a schematic diagram of the structure of a target recognition model according to an embodiment of the present invention;
[0051] Figure 5 is a flow chart of a method for training a target recognition model according to an embodiment of the present invention;
[0052] Figure 6 is a structural block diagram of a training device for a target recognition model according to an embodiment of the present invention;
[0053] Figure 7 is a structural block diagram of a target recognition device according to an embodiment of the present invention;
[0054] Figure 8 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0056] The target recognition method provided in the embodiment of the present invention can be applied to face recognition, vehicle recognition, label recognition, etc. There is no limitation on its application scenarios. Specifically, sample data in corresponding scenarios can be collected according to actual needs to train the target recognition model.
[0057] According to an embodiment of the present invention, a training method for a target recognition model or an embodiment of a target recognition method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0058] In this embodiment, a method for training a target recognition model is provided, which can be used in electronic devices such as computers, mobile phones, tablet computers, etc. Figure 1 is a flow chart of a method for training a target recognition model according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0059] S11, obtaining a preset recognition model and obtaining a sample data set.
[0060] The preset recognition model is obtained by training based on a first quality image, the sample data set includes a first quality image and a second quality image, and the quality of the first quality image is higher than that of the second quality image.
[0061] The first quality image can also be called a high-quality image, that is, a high-quality image, which can come from a public data set or an image collected in a controllable environment; the second quality image can also be called a low-quality image, that is, a low-quality image, which can come from an image collected in a complex environment. Of course, the sources of the first quality image and the second quality image are not limited to those described above, and can also be collected in other ways, and no limitation is made here. Among them, the set of first quality images is called a high-quality training sample set, and the set of second quality images is called a low-quality training sample set. The sample data set is a combination of a high-quality training sample set and a low-quality training sample set.
[0062] The preset recognition model can be obtained by training the electronic device using high-quality images in the high-quality training sample set, or can be sent to the electronic device after training by other devices. The method for obtaining the preset recognition model is not limited here, and it is only necessary to ensure that the electronic device can obtain the preset recognition model. The input of the preset recognition model is an image, and the output is an image feature.
[0063] In some optional implementations of this embodiment, obtaining a preset recognition model may include the following steps:
[0064] (1) Inputting the first quality image into the recognition model to obtain the feature layer output of the recognition model and the classification layer output of the recognition model.
[0065] (2) The corresponding loss function is calculated based on the feature layer output and the classification layer output to update the parameters of the recognition model and determine the preset recognition model.
[0066] The recognition model includes a feature extraction network and a classification layer, wherein the feature extraction network can be a residual network or other networks; the classification layer is a fully connected network. The electronic device inputs the first quality image into the recognition model, and obtains the feature layer output of the recognition model and the classification layer output of the recognition model respectively. Then, the loss function is calculated based on the feature layer output and the recognition model output respectively, that is, the domain loss calculation and the classification loss calculation are performed respectively, and the parameters of the recognition model are updated based on the calculation results of the loss function, and finally the preset recognition model is determined.
[0067] like Figure 2a As shown, the recognition model includes an input layer 31 (Input), a convolution layer 32 (Conv), a pooling layer 33 (Pooling), a residual unit 34 (Resblock), a fully connected layer 35 (fc), loss function layers 36, 37, a normalization layer 38, and a connection layer 39. For a description of the structure of each layer, please refer to Table 1 below.
[0068] Table 1 Structure description
[0069] layer describe Input Layer Input image data Convolutional Layer Extract features of the input image Pooling Layer Downsampling Resblock1 x 3 Three consecutive residual units are connected Resblock2 x 4 4 consecutive residual units connected Resblock3 x 6 6 consecutive residual units connected Resblock4 x 3 Three consecutive residual units are connected Fully connected layer Linear Weighted Sum Loss Function Evaluating model accuracy Normalization layer Normalize the feature data Connection Layer Connect the normalized features
[0070] like Figure 2b As shown in the figure, during the training process, the output of the fully connected layer FC1 is the feature layer output, and the output of FC2 is the classification layer output, and the loss functions are calculated for each of them. The output of FC1 is used for domain loss calculation, and the output of FC2 is used for classification loss calculation. For example, the classification loss function is the ArcFace loss function, which is expressed by the following formula:
[0071]
[0072] Where N is the total number of sample images, i is the i-th sample image, and y i is the category label to which the i-th sample image belongs, s is the scaling factor, θ is the angular interval between the weight vector of the recognition model and the feature vector of the training feature of the i-th sample image, t is the angular edge, the training features include the first training feature, 1≤i≤N.
[0073] The domain loss function is the LMMD loss function, which is expressed by the following formula:
[0074]
[0075] Where p and q are the distributions of the original domain and the target domain respectively, C is the number of categories, They are The weight belonging to class c, is the output of the feature layer.
[0076] Of course, the choice of specific loss function is not limited to the above, and other loss functions can be used, and can be set according to actual needs.
[0077] By calculating the loss function at the feature layer and the classification layer respectively to update the parameters of the recognition model, the accuracy of the trained preset recognition model is ensured.
[0078] S12, performing a first sampling on the first quality image and the second quality image in the sample data set to obtain a first sampling data set, and training a preset recognition model based on the first sampling data set to determine a first recognition model.
[0079] The first sampling means that the number of the first quality image and the second quality image input to the preset recognition model each time is the same, and the first quality image is used as the original domain and the second quality image is used as the target domain. Based on this, the electronic device can extract the same number of first quality images and second quality images each time to form a first sampling data set. The first sampling data set is then used to train the preset recognition model, and the parameters of the preset recognition model are updated to determine the first recognition model.
[0080] The first sampling data set includes not only corresponding images, but also image labels and image features. During the training process, the image categories and image features output by the preset recognition model are respectively compared with the image labels and image features corresponding to the first sampling data to perform loss function calculations, thereby updating the parameters of the preset recognition model and finally determining the first recognition model. The training cutoff condition for training the preset recognition model using the first sampling data set can be the number of iterations, or the accuracy of the model, etc., which meets the preset conditions.
[0081] For example, Figure 2b As shown, the electronic device inputs the first sample data set into the input layer 31, and uses the output of the feature layer and the output of the classification layer to calculate the loss function, update the parameters of the preset recognition model, and obtain the first recognition model.
[0082] This step will be described in detail below.
[0083] S13, performing a second sampling on the first quality image and the second quality image in the sample data set to obtain a second sample data set, and training a preset recognition model based on the second sample data set to determine a second recognition model.
[0084] The second sampling is to extract the first quality image and the second quality image according to a certain ratio to form a second acquisition data set. The certain ratio can be set according to actual conditions, or can be determined according to the ratio of all first quality images to all second quality images. For example, if the ratio of all first quality images to all second quality images is 2:1, then the above-mentioned certain ratio is 2:1.
[0085] After determining the second sample data set, the electronic device trains the preset recognition model in a similar manner to the above S12, except that the training uses a different data set. The second recognition model is determined by updating the parameters of the preset recognition model.
[0086] This step will be described in detail below.
[0087] S14, performing feature layer fusion on the first recognition model and the second recognition model to determine a target recognition model.
[0088] The first recognition model and the second recognition model respectively include a feature extraction model and a classification model connected after the feature extraction model. When the electronic device performs feature layer fusion, it respectively intercepts the feature extraction models of the first recognition model and the second recognition model, and connects the feature layers of the two in a preset manner, thereby realizing the fusion of the features extracted by the two feature extraction models to determine the target recognition model. The input of the target recognition model is an image, and the output is the features corresponding to the image.
[0089] The training method for the target recognition model provided in this embodiment adds training samples of low-quality domains to the training because samples in complex environments are generally second-quality images. The training method and domain loss are combined to make the features of the second-quality images close to those of the first-quality images, so that the recognition effect is close to that of the first-quality images. At the same time, two complementary models are combined to ensure that the recognition effect of the target recognition model obtained for both the first-quality images and the second-quality images is improved. In addition, the training process is based on a preset recognition model, and the two recognition models are trained on this basis, which can improve the recognition effect without affecting the training efficiency.
[0090] In this embodiment, a method for training a target recognition model is provided, which can be used in electronic devices such as computers, mobile phones, tablet computers, etc. Figure 3 is a flow chart of a method for training a target recognition model according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:
[0091] S21, obtaining a preset recognition model and obtaining a sample data set.
[0092] The preset recognition model is obtained by training based on a first quality image, the sample data set includes a first quality image and a second quality image, and the quality of the first quality image is higher than that of the second quality image.
[0093] For details, please see Figure 1 S11 of the illustrated embodiment will not be described in detail here.
[0094] S22, performing a first sampling on the first quality image and the second quality image in the sample data set to obtain a first sampling data set, and training a preset recognition model based on the first sampling data set to determine a first recognition model.
[0095] Specifically, the above S22 includes:
[0096] S221, obtaining the number of input samples of the preset recognition model.
[0097] The number of input samples of the preset recognition model represents the input amount input to the preset recognition model during each training, for example, the number of input samples is N. The number of input samples can be obtained when obtaining the preset recognition model, or can be determined by querying the number of nodes in the input layer of the preset recognition model, and so on.
[0098] S222: Determine the number of first quality images and second quality images in each training based on the number of input samples in a 1:1 ratio.
[0099] As described above, if the number of input samples is N, then according to the ratio of 1:1, the number of first quality images is N / 2 and the number of second quality images is N / 2 in each training.
[0100] S223: Based on the determined number of first quality images and second quality images, sampling is performed in the sample data set to obtain a first sample data set for each training.
[0101] After the number of first quality images and the number of second quality images are determined, the electronic device samples data from the sample data set, extracting N / 2 first quality images and N / 2 second quality images each time to form a first sample data set.
[0102] S224: Train a preset recognition model based on the first sampling data set to determine a first recognition model.
[0103] For details about this step, see Figure 1 The relevant description of S22 in the illustrated embodiment will not be repeated here.
[0104] S23, performing a second sampling on the first quality image and the second quality image in the sample data set to obtain a second sample data set, and training a preset recognition model based on the second sample data set to determine a second recognition model.
[0105] Specifically, the above S23 includes:
[0106] S231, obtaining the number of input samples of a preset recognition model and obtaining the ratio of the number of first quality images to the number of second quality images in the sample data set.
[0107] The electronic device counts the number of first quality images and the number of second quality images in the sample data set. Specifically, as described above, the first quality images form a high-quality sample data set, and the second quality images form a low-quality sample data set. The electronic device counts the number of first quality images in the high-quality sample data set and the number of second quality images in the low-quality sample data set, respectively, and calculates the ratio of the two to obtain the ratio of the number of first quality images to second quality images in the sample data set. For example, the calculated ratio is c1:c2.
[0108] S232: Determine the number of first quality images and second quality images in each training according to the number ratio based on the number of input samples.
[0109] After determining that the quantity ratio is c1:c2, the electronic device determines the quantity of the first quality images and the second quality images in each training according to the quantity of the input samples.
[0110] S233: Based on the determined number of first quality images and second quality images, sampling is performed in the sample data set to obtain a second sample data set for each training.
[0111] The electronic device samples corresponding images in the sample data set using the number determined in S232 to obtain a second sample data set for each training.
[0112] S234: Train a preset recognition model based on the second sampling data set to determine a second recognition model.
[0113] For details about this step, see Figure 1 The relevant description of S23 in the illustrated embodiment will not be repeated here.
[0114] S24, performing feature layer fusion on the first recognition model and the second recognition model to determine a target recognition model.
[0115] Specifically, the above S24 includes:
[0116] S241, connecting the feature layers in the first recognition model and the second recognition model to the normalization layer respectively.
[0117] As described above, the first recognition model and the second recognition model include a feature extraction module and a classification module. The electronic device connects the outputs of the feature extraction modules of the first recognition model and the second recognition model, that is, the two feature layers, to a normalization layer respectively.
[0118] S242, the two normalization layers are fused to determine the target recognition model.
[0119] The fusion of the two normalized layers can be performed by splicing or weighting, etc. The specific fusion method is not limited here. After the electronic device fuses the two normalized layers, it can determine the target recognition model.
[0120] like Figure 4As shown, the layers 31-35 on the left are the first recognition model, and the layers 31-35 on the right are the second recognition model. The normalization layer 38 is connected after the fully connected layer 35 of the first recognition model, and the normalization layer 38 is connected after the fully connected layer 35 of the second recognition model. The two normalization layers 38 are fused using the Concat layer 39 to determine the target recognition model shown in 4.
[0121] The training method of the target recognition model provided in this embodiment extracts the first quality image and the second quality image in a 1:1 ratio to form the first sampling data set used in each training, and uses the first quality image as the original domain and the second quality image as the target domain, so that the recognition effect of the second quality image is close to the first quality image. The second sampling data set is determined according to the quantitative ratio of the first quality image to the second quality image in the sample data set, and the two quality images can be used to complement each other during subsequent training. After the normalization connection is performed at the feature layer, fusion is performed so that the obtained target recognition model can combine the advantages of the two models and improve the recognition accuracy of the target recognition model.
[0122] In this embodiment, a method for training a target recognition model is provided, which can be used in electronic devices such as computers, mobile phones, tablet computers, etc. Figure 5 is a flow chart of a method for training a target recognition model according to an embodiment of the present invention. Figure 5 As shown, the process includes the following steps:
[0123] S31, obtaining an image to be recognized.
[0124] The image to be identified may be collected by the electronic device, or obtained by the electronic device from a third-party device, etc. Optionally, the image to be identified may be determined in the following manner:
[0125] (1) Perform object detection on the original image to obtain the object key points in the original image.
[0126] (2) Based on the positions of the key points of the object in the original image, the original image is scaled to a preset size to obtain the image to be identified.
[0127] By scaling each original image to a preset size, the preset recognition model only needs to recognize images of the preset size, thereby improving the recognition accuracy of the preset recognition model.
[0128] Optionally, the original image is an image collected in a monitoring scene. Since the object is constantly moving in the monitoring scene, there are few continuous static images. Therefore, when detecting an object, the electronic device only uses an object detection algorithm to detect the object in the moving area, thereby improving the object detection speed. Among them, the object detection algorithm is a multi-task convolutional neural network (MTCNN) algorithm; or a DenseBox algorithm; or an SSH algorithm, etc. This embodiment does not limit the type of the object detection algorithm.
[0129] S32, inputting the image to be identified into the target recognition model to determine the features to be matched of the image to be identified.
[0130] The target recognition model is obtained by training according to the target recognition model training method described in any of the above embodiments. For the specific structure of the target recognition model, please refer to the detailed description in the above embodiment, which will not be repeated here.
[0131] The electronic device inputs the image to be identified into the target recognition model, and outputs the features to be matched of the image to be identified through the target recognition model.
[0132] S33, matching the feature to be matched with each target feature to determine the target recognition result corresponding to the image to be recognized.
[0133] A feature database may be provided in the electronic device, and the electronic device matches the features to be matched in the feature database to determine the target recognition result corresponding to the image to be recognized. The target recognition result is specifically set according to the application scenario, for example, identity information, vehicle information, etc.
[0134] The target recognition method provided in this embodiment ensures that the target recognition model is obtained by combining two complementary models, thereby improving the recognition effect of the obtained target recognition model on the first quality image and the second quality image. Based on this, the target recognition model is used to identify the image to be identified, which can improve the accuracy of the features to be matched, thereby improving the accuracy of the target recognition result.
[0135] In the present embodiment, a training or target recognition device for a target recognition model is also provided, and the device is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions thereof will not be repeated. As used below, the term "module" may implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and contemplated.
[0136] This embodiment provides a training device for a target recognition model, such as Figure 6 As shown, including:
[0137] A first acquisition module 41 is used to acquire a preset recognition model and a sample data set, wherein the preset recognition model is obtained by training based on a first quality image, and the sample data set includes a first quality image and a second quality image, wherein the quality of the first quality image is higher than that of the second quality image;
[0138] A first training module 42 is used to perform a first sampling on the first quality image and the second quality image in the sample data set to obtain a first sampling data set, and train the preset recognition model based on the first sampling data set to determine a first recognition model;
[0139] A second training module 43 is used to perform a second sampling on the first quality image and the second quality image in the sample data set to obtain a second sampling data set, and train the preset recognition model based on the second sampling data set to determine a second recognition model;
[0140] The fusion module 44 is used to perform feature-layer fusion on the first recognition model and the second recognition model to determine a target recognition model.
[0141] This embodiment provides a target recognition device, such as Figure 7 As shown, including:
[0142] The second acquisition module 51 is used to acquire the image to be recognized;
[0143] A determination module 52, used for inputting the image to be identified into a target recognition model to determine the features to be matched of the image to be identified, wherein the target recognition model is trained according to the target recognition model training method described in any embodiment of the present invention;
[0144] The matching module 53 is used to match the to-be-matched features with each target feature to determine the target recognition result corresponding to the to-be-recognized image.
[0145] The training device of the target recognition model, or the target recognition device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0146] The further functional description of each of the above modules is the same as that of the above corresponding embodiments and will not be repeated here.
[0147] The embodiment of the present invention also provides an electronic device having the above Figure 6The training device of the target recognition model shown in FIG. Figure 7 The target recognition device shown.
[0148] See also Figure 8 , Figure 8 is a schematic diagram of the structure of an electronic device provided by an optional embodiment of the present invention, such as Figure 8 As shown, the electronic device may include: at least one processor 601, such as a CPU (Central Processing Unit), at least one communication interface 603, a memory 604, and at least one communication bus 602. The communication bus 602 is used to realize the connection and communication between these components. The communication interface 603 may include a display screen (Display), a keyboard (Keyboard), and the optional communication interface 603 may also include a standard wired interface and a wireless interface. The memory 604 may be a high-speed RAM memory (Random Access Memory) or a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 604 may optionally be at least one storage device located away from the aforementioned processor 601. The processor 601 may be combined with Figure 6 Or the device described in 7, the application is stored in the memory 604, and the processor 601 calls the program code stored in the memory 604 to execute any of the above method steps.
[0149] The communication bus 602 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The communication bus 602 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0150] Among them, the memory 604 may include a volatile memory (English: volatile memory), such as a random access memory (English: random-access memory, abbreviated: RAM); the memory may also include a non-volatile memory (English: non-volatile memory), such as a flash memory (English: flash memory), a hard disk drive (English: hard disk drive, abbreviated: HDD) or a solid-state drive (English: solid-state drive, abbreviated: SSD); the memory 604 may also include a combination of the above types of memory.
[0151] The processor 601 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and a NP.
[0152] The processor 601 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0153] Optionally, the memory 604 is also used to store program instructions. The processor 601 can call the program instructions to implement the training method of the target recognition model shown in any embodiment of the present application, or the target recognition method shown in any embodiment.
[0154] The embodiment of the present invention also provides a non-transitory computer storage medium, which stores computer executable instructions, and the computer executable instructions can execute the training method of the target recognition model in any of the above method embodiments, or the target recognition method. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FlashMemory), a hard disk (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memory.
[0155] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for training a target recognition model, characterized in that: include: Acquire a preset recognition model and acquire a sample data set, wherein the preset recognition model is obtained by training based on a first quality image, and the sample data set includes a first quality image and a second quality image, wherein the quality of the first quality image is higher than the quality of the second quality image; Performing a first sampling on the first quality image and the second quality image in the sample data set to obtain a first sampling data set, and training the preset recognition model based on the first sampling data set to determine a first recognition model; Performing a second sampling on the first quality image and the second quality image in the sample data set to obtain a second sampling data set, and training the preset recognition model based on the second sampling data set to determine a second recognition model; The first recognition model and the second recognition model are fused at the feature level to determine a target recognition model.
2. The method according to claim 1, characterized in that The first sampling of the first quality image and the second quality image in the sample data set to obtain a first sample data set includes: Obtaining the number of input samples of the preset recognition model; Determining the number of the first quality images and the second quality images in each training based on the number of the input samples in a ratio of 1:1; Based on the determined number of the first quality images and the second quality images, sampling is performed in the sample data set to obtain the first sample data set for each training.
3. The method according to claim 1 or 2, characterized in that: The performing second sampling on the first quality image and the second quality image in the sample data set to obtain a second sample data set includes: Obtaining the number of input samples of the preset recognition model and obtaining the ratio of the number of the first quality images to the second quality images in the sample data set; Determining the number of the first quality images and the second quality images in each training based on the number of input samples and according to the number ratio; Based on the determined number of the first quality images and the second quality images, sampling is performed in the sample data set to obtain the second sample data set for each training.
4. The method according to claim 1, characterized in that: The step of fusing the first recognition model with the second recognition model at the feature level to determine the target recognition model includes: Connecting the feature layers in the first recognition model and the second recognition model to the normalization layer respectively; The two normalization layers are fused to determine the target recognition model.
5. The method according to claim 1, characterized in that: The obtaining of the preset recognition model comprises: Inputting the first quality image into a recognition model to obtain a feature layer output of the recognition model and a classification layer output of the recognition model; Corresponding loss function calculations are performed based on the feature layer output and the classification layer output respectively to update the parameters of the recognition model and determine the preset recognition model.
6. A target recognition method, characterized in that: include: Obtain an image to be recognized; Inputting the image to be identified into a target recognition model to determine features to be matched of the image to be identified, wherein the target recognition model is trained according to the target recognition model training method according to any one of claims 1 to 5; The to-be-matched features are matched with each target feature to determine the target recognition result corresponding to the to-be-recognized image.
7. A training device for a target recognition model, characterized in that: include: A first acquisition module is used to acquire a preset recognition model and a sample data set, wherein the preset recognition model is obtained by training based on a first quality image, and the sample data set includes a first quality image and a second quality image, wherein the quality of the first quality image is higher than that of the second quality image; A first training module, configured to perform a first sampling on the first quality image and the second quality image in the sample data set to obtain a first sampling data set, and train the preset recognition model based on the first sampling data set to determine a first recognition model; A second training module, configured to perform a second sampling on the first quality image and the second quality image in the sample data set to obtain a second sampling data set, and train the preset recognition model based on the second sampling data set to determine a second recognition model; The fusion module is used to perform feature-layer fusion on the first recognition model and the second recognition model to determine a target recognition model.
8. A target recognition device, characterized in that: include: A second acquisition module is used to acquire an image to be identified; A determination module, used for inputting the image to be identified into a target recognition model to determine the features to be matched of the image to be identified, wherein the target recognition model is trained according to the training method of the target recognition model according to any one of claims 1 to 5; The matching module is used to match the to-be-matched features with each target feature to determine the target recognition result corresponding to the to-be-recognized image.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the training method of the target recognition model described in any one of claims 1 to 5, or executes the target recognition method described in claim 6 or 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the training method of the target recognition model described in any one of claims 1 to 5, or to execute the target recognition method described in claim 6 or 7.
Citation Information
Patent Citations
Neural network model migration method and system, electronic device, program and medium
CN108229534A
Image super-resolution reconstruction method based on residual distillation network
CN110111256A