Model training method and device, electronic equipment and storage medium

By performing style conversion of the training samples collected by the first depth camera, samples consistent with the samples collected by the second depth camera are generated, which solves the classification model recognition accuracy problem caused by the difference in image samples between depth cameras, and achieves high recognition accuracy in the second depth camera scene.

CN120147769APending Publication Date: 2025-06-13BEIJING MOMENTA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311707590.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In some scenarios, only image samples from one depth camera may be collected, resulting in poor identification accuracy of classification models in another depth camera scene.

Method used

By obtaining the training samples collected by the first depth camera, performing style conversion, a second training sample that is consistent with the data distribution of the sample collected by the second depth camera, training the benchmark classification model based on the second training sample, and obtaining the target classification model.

Benefits of technology

In the second-depth camera scene, the target classification model generated through style conversion can significantly improve the recognition accuracy, improve the real rate and reduce the false negative rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147769A_ABST
    Figure CN120147769A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device, electronic equipment and a storage medium, and can train a classification model capable of accurately identifying a sample collected by a second depth camera under the condition that the sample collected by the second depth camera is lacked. The model training method comprises the steps that a first training sample is acquired, and the first training sample is acquired by a first depth camera; the first training sample is subjected to style conversion, a second training sample is obtained, the second training sample is consistent with a sample collected by a second depth camera in data distribution, and the first depth camera is different from the second depth camera; and training the reference classification model based on the second training sample to obtain a target classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of deep learning technology, and in particular, to a model training method, apparatus, electronic device, and storage medium.

Background Art

[0002] Generally, depth cameras mainly include TOF cameras and RGBD cameras, and each of the TOF camera and the RGBD camera can collect the same type of image samples. The difference is that there are significant differences between the same type of image samples.

[0003] In some scenarios, it may be possible to collect only the image samples of a certain depth camera mentioned above. Then, the classification model can only be trained based on the above-mentioned image samples. However, when applying the above classification model, in the scene where the image samples may be collected by the other depth camera mentioned above, the classification accuracy of the above classification model may be poor when identifying the image samples collected by the other camera.

Summary of the Invention

[0004] Embodiments of the present application provide a model training method, apparatus, electronic device, and storage medium, which can train a classification model that can accurately identify the samples collected by the second depth camera in the case of a lack of samples collected by the second depth camera.

[0005] In a first aspect, embodiments of the present application provide a model training method, and the method includes:

[0006] Obtain a first training sample, where the first training sample is collected by a first depth camera;

[0007] Perform style conversion on the first training sample to obtain a second training sample, where the data distribution of the second training sample is consistent with that of the samples collected by a second depth camera, and the first depth camera is different from the second depth camera;

[0008] Train a benchmark classification model based on the second training sample to obtain a target classification model.

[0009] In the embodiments of the present application, the first depth camera can be considered as the depth camera that can currently be used to collect training samples. For example, the first depth camera collects the first training sample, and the second depth camera can be considered as the depth camera that can be used to collect samples of the same type as the first depth camera during application. Therefore, by performing style conversion on the first training sample, a second training sample with a data distribution consistent with that of the samples collected by the second depth camera is obtained, that is, the second training sample is considered to be indistinguishable from the samples collected by the second depth camera. Then, the target classification model obtained by training based on the second training sample can be considered to achieve a high recognition effect in the scenario of applying the second depth camera.

[0010] Optionally, performing style conversion on the first training sample to obtain a second training sample, including:

[0011] Performing a preprocessing operation on the first training sample;

[0012] Inputting the first training sample that has undergone the preprocessing operation into a pre-trained style conversion model to obtain the second training sample, where the style conversion model is a generative model based on the cycle-GAN structure.

[0013] In the embodiments of the present application, by using a generative model based on the cycle-GAN structure to perform style conversion on the first training sample, it is possible to obtain a second training sample whose data distribution is relatively consistent with the samples collected by the second depth camera.

[0014] Optionally, the preprocessing operation at least includes an object alignment operation and a central cropping operation.

[0015] In the embodiments of the present application, by performing an object alignment operation and a central cropping operation on the first training sample, while ensuring that the first training sample has the same specifications and contains the same object, background interference can be minimized to the greatest extent, thereby improving the recognition effect of the obtained target classification model.

[0016] Optionally, before training the benchmark classification model based on the second training sample to obtain a target classification model, the method further includes:

[0017] Performing histogram equalization processing on the second training sample to obtain a third training sample;

[0018] Training the benchmark classification model based on the second training sample to obtain a target classification model, including:

[0019] Training the benchmark classification model based on the third training sample to obtain the target classification model.

[0020] In the embodiments of the present application, first, histogram equalization processing is performed on the second training sample obtained by style conversion to obtain a third training sample. Compared with the second training sample, the third training sample has improved image contrast, that is, more detailed features of the objects contained therein can be highlighted from the third training sample. Then, the target classification model trained using the third training sample can have a better recognition effect.

[0021] Optionally, performing histogram equalization processing on the second training sample to obtain a third training sample, including:

[0022] Perform local histogram equalization on the second training sample to obtain the third training sample.

[0023] In the embodiments of the present application, by adopting a specific histogram equalization method, i.e., local histogram equalization, for the second training sample, while enhancing the image contrast of the second training sample, the noise will not be amplified.

[0024] Optionally, before training the benchmark classification model based on the third training sample to obtain the target classification model, the method further includes:

[0025] Perform normalization on the third training sample to obtain a fourth training sample;

[0026] Training the benchmark classification model based on the third training sample to obtain the target classification model includes:

[0027] Training the benchmark classification model based on the fourth training sample to obtain the target classification model.

[0028] In the embodiments of the present application, by performing normalization on the third training sample and then training the benchmark classification model, the training process can be accelerated, and at the same time, the obtained target classification model has a better recognition effect.

[0029] Optionally, before training the benchmark classification model based on the second training sample to obtain the target classification model, the method further includes:

[0030] Obtain a fifth training sample, where the fifth training sample is collected by the second depth camera;

[0031] Training the benchmark classification model based on the second training sample to obtain the target classification model includes:

[0032] Training the benchmark classification model based on the first training sample, the second training sample, and the fifth training sample to obtain the target classification model.

[0033] In the embodiments of the present application, training the benchmark model jointly with the first training sample collected by the first depth camera, the second training sample obtained by style conversion of the first training sample which can be equivalent to the sample collected by the second depth camera, and the fifth training sample collected by the second depth camera can, to a certain extent, improve the recognition effect of the obtained target classification model when applied to the scenario of the second depth camera compared with training the benchmark model only based on the first training sample and the fifth training sample.

[0034] In a second aspect, the embodiments of the present application provide a model training device, and the device includes:

[0035] An acquisition unit for acquiring a first training sample, where the first training sample is acquired by a first depth camera;

[0036] A conversion unit for performing style conversion on the first training sample to obtain a second training sample, where the data distribution of the second training sample is consistent with that of the samples acquired by a second depth camera, and the first depth camera is different from the second depth camera;

[0037] A training unit for training a benchmark classification model based on the second training sample to obtain a target classification model.

[0038] Optionally, the conversion unit includes:

[0039] A preprocessing unit for performing preprocessing operations on the first training sample;

[0040] A style conversion unit for inputting the first training sample that has undergone the preprocessing operation into a pre-trained style conversion model to obtain the second training sample, where the style conversion model is a generative model based on the cycle-GAN structure.

[0041] Optionally, the preprocessing operation includes at least object alignment operation and central cropping operation.

[0042] Optionally, the preprocessing unit is further configured to:

[0043] Perform histogram equalization processing on the second training sample to obtain a third training sample;

[0044] The training unit includes:

[0045] A training subunit for training the benchmark classification model based on the third training sample to obtain the target classification model.

[0046] Optionally, the preprocessing unit is specifically configured to:

[0047] Perform local histogram equalization processing on the second training sample to obtain the third training sample.

[0048] Optionally, the preprocessing unit is further configured to:

[0049] Perform normalization processing on the third training sample to obtain a fourth training sample;

[0050] The training subunit is specifically configured to:

[0051] Train the benchmark classification model based on the fourth training sample to obtain the target classification model.

[0052] Optionally, the obtaining unit is further configured to:

[0053] Obtain a fifth training sample, where the fifth training sample is collected by the second depth camera;

[0054] The training unit is specifically configured to:

[0055] Train the benchmark classification model based on the first training sample, the second training sample, and the fifth training sample to obtain the target classification model.

[0056] In a third aspect, an embodiment of the present application provides an electronic device, where the electronic device includes:

[0057] At least one processor;

[0058] A memory coupled to the processor;

[0059] When the at least one processor executes the computer program stored in the memory, the steps of the method according to any one of the embodiments of the first aspect or the second aspect are implemented.

[0060] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any one of the embodiments of the first aspect are implemented.

[0061] It should be understood that the second to fourth aspects of the embodiments of the present application are consistent with the technical solutions of the first aspect of the embodiments of the present application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar, and will not be repeated.

Description of the Drawings

[0062] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 It is a schematic flowchart of a model training method provided by an embodiment of the present application;

[0064] Figure 2 It is a schematic flowchart of a method for performing style conversion on a first training sample provided by an embodiment of the present application;

[0065] Figure 3 It is a schematic flowchart of a method for obtaining a target classification model provided by an embodiment of the present application;

[0066] Figure 4 Schematic flowchart of a method for performing histogram equalization processing provided by an embodiment of the present application;

[0067] Figure 5 Schematic flowchart of a method for obtaining a target classification model provided by an embodiment of the present application;

[0068] Figure 6 Schematic flowchart of a method for obtaining a target classification model provided by an embodiment of the present application;

[0069] Figure 7 Schematic structural diagram of a model training device provided by an embodiment of the present application;

[0070] Figure 8 Schematic structural diagram of an electronic device provided by an embodiment of the present application.

Specific Embodiments

[0071] To better understand the technical solutions of this specification, the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0072] It should be clear that the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this specification without creative efforts belong to the scope protected by this specification.

[0073] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit this specification. The singular forms of "a", "the" and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0074] Currently, depth cameras mainly include TOF cameras and RGBD cameras. When using TOF cameras and RGBD cameras to collect image samples respectively, although image samples of the same type can be collected, that is, both TOF cameras and RGBD cameras can collect near-infrared image samples and depth image samples, there are significant differences in the same type of image samples collected by TOF cameras and RGBD cameras when presenting the same object.

[0075] Through research by the applicant, it is found that in some scenarios, it may only be possible to collect image samples of one of the above-mentioned depth cameras. Then, the classification model can only be trained based on the above-mentioned image samples. However, when applying the above-mentioned classification model, in the scene where the image samples may be collected by the other depth camera, the classification accuracy of the above-mentioned classification model may be poor when identifying the image samples collected by the other depth camera.

[0076] In view of this, an embodiment of the present application provides a model training method. In this method, by obtaining a first training sample collected by a first depth camera, and then converting the first training sample into a style consistent with the sample collected in the second depth camera scene, a second training sample is obtained. That is, the second training sample can be considered to be collected by the second depth camera in the second depth camera scene. Then, the target classification model trained based on the second training sample can be considered to achieve a high recognition effect in the scene where the second depth camera is applied.

[0077] The following introduces the technical solution provided by the embodiment of the present application with reference to the accompanying drawings. Please refer to Figure 1 , an embodiment of the present application provides a model training method, which is applied to an electronic device. The electronic device can be a notebook, a desktop computer, a server, and a chip module with computing capabilities, and no special limitation is made here. The process of this method is described as follows:

[0078] Step 101: Obtain a first training sample.

[0079] In the embodiment of the present application, the first depth camera can be considered as the camera currently available for collecting image samples, and the first training sample can be considered as the image samples collected by the first depth camera according to actual training requirements. For example, if the purpose of training is to train a live body recognition model (that is, essentially a classification model for distinguishing live bodies from non-live bodies), then the first training sample includes the near-infrared images and depth images corresponding to live bodies, as well as the near-infrared images and depth images corresponding to non-live bodies. Of course, according to other training requirements, the first training sample can only include one of the near-infrared images or depth images, and no special limitation is made here.

[0080] Step 102: Perform style conversion on the first training sample to obtain a second training sample.

[0081] In the embodiment of the present application, the second depth camera can be considered as the camera used to collect image samples in the application scenario, and the main difference between the same type of image samples collected by the first depth camera and the second camera lies in the style difference. Then, after obtaining the first training sample collected by the first depth camera, the first training sample can be subjected to style conversion to obtain a second training sample with the same data distribution as the sample collected by the second depth camera, that is, it is considered that the second training sample and the sample collected by the second depth camera have the same presentation style for the same object.

[0082] It should be understood that the first depth camera and the second depth camera are different. For example, when the first depth camera is an RGBD camera, the second depth camera is a TOF camera; when the first depth camera is a TOF camera, the second depth camera is an RGBD camera.

[0083] Step 103: Train the benchmark classification model based on the second training sample to obtain the target classification model.

[0084] In the embodiments of the present application, after obtaining the second training sample after style conversion, the benchmark classification model can be trained to obtain the target classification model. Then, when applying the target classification model to identify the samples collected by the second depth camera, a relatively accurate identification effect can be achieved.

[0085] For example, the second training sample is the near-infrared image and depth image after style conversion corresponding to a live body, and the near-infrared image and depth image after style conversion corresponding to a non-live body. The benchmark classification model can adopt ResNet, MobileNet, etc., which are not particularly limited here. Then, based on the above second training sample, the above benchmark classification model is trained to obtain a classification model for distinguishing live / non-live bodies.

[0086] Please refer to Figure 2 , a schematic flowchart of a method for performing style conversion on the first training sample provided in the embodiments of the present application. Step 103 can be specifically implemented by executing sub-step 1031 and step 1032:

[0087] Step 1031: Perform a preprocessing operation on the first training sample.

[0088] Step 1032: Input the first training sample after the preprocessing operation into the pre-trained style conversion model to obtain the second training sample.

[0089] In the embodiments of the present application, the preprocessing operation performed on the first training sample may at least include an object alignment operation and a central cropping operation. For example, if the main object included in the first training sample is a human face, then the above object alignment operation can be considered a face alignment operation, and the above central cropping operation is to crop the background area centered on the area where the face is located, so as to minimize background interference while ensuring that the first training sample has the same specifications and contains the same object, thereby improving the identification effect of the obtained target classification model.

[0090] After the above preprocessing, the preprocessed first training sample can be input into the pre-trained style conversion model. This style conversion model can be considered a generation model based on the cycle-GAN structure, that is, using the generation characteristics of the generation model based on the cycle-GAN structure to generate a second training sample with a data distribution relatively consistent with the samples collected by the second depth camera.

[0091] It should be understood that, of course, in addition to performing style conversion on the first training sample based on the generation model of the cycle-GAN structure, other generation models or other methods can also be used for style conversion, and no special limitation is made here.

[0092] Please refer to Figure 3 , which is a schematic flowchart of a method for obtaining a target classification model provided by an embodiment of the present application. Before executing step 103, step 201 can also be executed:

[0093] Step 201: Perform histogram equalization processing on the second training sample to obtain a third training sample.

[0094] Step 103 can be specifically implemented by executing sub-step 202:

[0095] Step 202: Train the benchmark classification model based on the third training sample to obtain a target classification model.

[0096] In the embodiment of the present application, first, histogram equalization processing is performed on the second training sample obtained by style conversion to obtain a third training sample. Compared with the second training sample, the third training sample improves the image contrast, that is, through histogram equalization, more detailed features of the objects included in the third training sample can be highlighted. Then, the target classification model trained using the third training sample can have a better recognition effect.

[0097] Please refer to Figure 4 , which is a schematic flowchart of a method for performing histogram equalization processing provided by an embodiment of the present application. Step 201 can be specifically implemented by sub-step 2011:

[0098] Step 2011: Perform local histogram equalization processing on the second training sample to obtain a third training sample.

[0099] In the embodiment of the present application, by adopting a specific histogram equalization method for the second training sample, that is, local histogram equalization, while improving the image contrast of the second training sample, the noise is not amplified. Then, the target classification model trained using the third training sample can have a better recognition effect.

[0100] Please refer to Figure 5 , which is a schematic flowchart of a method for obtaining a target classification model provided by an embodiment of the present application. Before step 202, step 301 can also be executed:

[0101] Step 301: Perform normalization processing on the third training sample to obtain a fourth training sample.

[0102] Step 202 can be specifically implemented by executing sub-step 302:

[0103] Step 302: Train the benchmark classification model based on the fourth training sample to obtain the target classification model.

[0104] In the embodiments of the present application, by normalizing the third training sample, that is, normalizing the pixel values of the third training sample to [0, 1], and then training the benchmark classification model, the difference between different third training samples can be reduced, the training process can be accelerated, and at the same time, the obtained target classification model has a better recognition effect.

[0105] Please refer to Figure 6 , which is a schematic flowchart of a method for obtaining a target classification model provided by the embodiments of the present application. Before step 103, step 401 may also be executed:

[0106] Step 401: Obtain the fifth training sample, where the fifth training sample is collected by the second depth camera.

[0107] Step 103 may be specifically implemented by executing sub-step 402:

[0108] Step 402: Train the benchmark classification model based on the first training sample, the second training sample, and the fifth training sample to obtain the target classification model.

[0109] In the embodiments of the present application, both the current first depth camera and the second depth camera can be used to collect image samples. However, there may be a situation where the number of image samples collected by the second depth camera is small. For example, the second depth camera has collected the fifth training sample, but the number of the fifth training samples is small. At this time, the second training sample can be regarded as being collected by the second depth camera, so as to increase the number of image samples collected by the second depth camera. On this basis, the first training sample collected by the first depth camera, the second training sample, and the fifth training sample collected by the second depth camera are jointly used to train the benchmark model. Compared with training the benchmark model only based on the first training sample and the fifth training sample, the recognition effect of the obtained target classification model applied in the scenario of the second depth camera can be improved to a certain extent.

[0110] The recognition effects before and after the improvement of the model training method will be described below.

[0111] Taking the classification model with the target classification model being a live body / non - live body as an example, when judging the recognition effect, the following four parameters are introduced: True Positive (TP), True Negative (TN), False Positive (FP), and False Negative (FN). Among them, TP represents that the actual is a live body and the recognition result of the above - mentioned classification model is a live body; TN represents that the actual is a non - live body and the recognition result of the above - mentioned classification model is a non - live body; FP represents that the actual is a non - live body and the recognition result of the above - mentioned classification model is a live body; FN represents that the actual is a live body and the recognition result of the above - mentioned classification model is a non - live body.

[0112] Based on the above four parameters, two main evaluation indicators can be formed: True Positive Rate (TPR) and False Negative Rate (FPR).

[0113] TPR = TP / (TP + FN)

[0114] FPR = FP / (FP + TN)

[0115] In comparison scenario one, the first depth camera is an RGBD camera, and the samples it collects are represented by A; the second depth camera is a TOF camera, and the samples it collects are represented by B. The comparison of the recognition effects before and after the improvement of the model training method is shown in Table 1:

[0116] Table 1

[0117]

[0118]

[0119] Comparing the above - mentioned first group and the second group, it can be seen that when the samples collected in the application scenario are B, compared with directly training using A, for the live - body recognition model obtained by converting the style of A into samples with the same data distribution as B and then training, when applied to B, its TPR is significantly increased from 21.48% to 62.73%.

[0120] In comparison scenario two, the first depth camera is an RGBD camera, and the samples it collects are represented by A; the second depth camera is a TOF camera, and the samples it collects are represented by B. The comparison of the recognition effects before and after the improvement of the model training method is shown in Table 2:

[0121] Table 2

[0122] Serial number Training method FPR TPR 3 AB training, B testing 0.01 97% 4 AB + Training with samples after A style conversion, B testing 0.01 100%

[0123] Comparing the above-mentioned third group with the fourth group, it can be seen that when the collected sample in the application scenario is B, compared with directly using AB for training, after converting A into a sample with the same data distribution as B and then jointly training AB to obtain a live recognition model, when applied to B, its TPR is significantly increased from 97% to 100%.

[0124] Please refer to Figure 7 , based on the same inventive concept, an embodiment of the present application provides a model training device, which includes: an acquisition unit 501, a conversion unit 502, and a training unit 503.

[0125] The acquisition unit 501 is used to acquire a first training sample, where the first training sample is collected by a first depth camera;

[0126] The conversion unit 502 is used to perform style conversion on the first training sample to obtain a second training sample, and the data distribution of the second training sample is the same as that of the sample collected by the second depth camera. The first depth camera is different from the second depth camera;

[0127] The training unit 503 is used to train a benchmark classification model based on the second training sample to obtain a target classification model.

[0128] Optionally, the conversion unit 502 includes:

[0129] A preprocessing unit, which is used to perform preprocessing operations on the first training sample;

[0130] The style conversion unit is used to input the first training sample after the preprocessing operation into a pre-trained style conversion model to obtain a second training sample, where the style conversion model is a generation model based on the cycle-GAN structure.

[0131] Optionally, the preprocessing operation includes at least object alignment operation and central cropping operation.

[0132] Optionally, the preprocessing unit is further used to:

[0133] Perform histogram equalization processing on the second training sample to obtain a third training sample;

[0134] The training unit 503 includes:

[0135] A training subunit, which is used to train a benchmark classification model based on the third training sample to obtain a target classification model.

[0136] Optionally, the preprocessing unit is specifically used to:

[0137] Perform local histogram equalization processing on the second training sample to obtain a third training sample.

[0138] Optionally, the preprocessing unit is further configured to:

[0139] Normalize the third training sample to obtain a fourth training sample;

[0140] The training subunit is specifically configured to:

[0141] Train the benchmark classification model based on the fourth training sample to obtain a target classification model.

[0142] Optionally, the obtaining unit 501 is further configured to:

[0143] Obtain a fifth training sample, where the fifth training sample is collected by the second depth camera;

[0144] The training unit 503 is specifically configured to:

[0145] Train the benchmark classification model based on the first training sample, the second training sample, and the fifth training sample to obtain a target classification model.

[0146] Please refer to Figure 8 , based on the same inventive concept, an embodiment of the present application provides an electronic device 100, which includes at least one processor 601. The processor 601 is configured to execute a computer program stored in a memory to implement the steps of the model training method provided in the embodiment of the present application as Figures 1 - 6 shown.

[0147] Optionally, the processor 601 may specifically be a central processing unit or a specific ASIC, and may be one or more integrated circuits for controlling program execution.

[0148] Optionally, the electronic device 100 may further include a memory 602 coupled to at least one processor 601. The memory 602 may include a ROM, a RAM, and a disk memory. The memory 602 is used to store data required for the operation of the processor 601, that is, instructions that can be executed by at least one processor 601. The at least one processor 601 executes the instructions stored in the memory 602 to execute the method as Figures 1 - 6 shown. Wherein, the number of the memories 602 is one or more.

[0149] Wherein, the entity devices corresponding to the obtaining unit 501, the conversion unit 502, and the training unit 503 may all be the foregoing processor 601. The electronic device 100 may be used to execute the method provided in the embodiment shown in Figures 1 - 6 . Therefore, regarding the functions that can be implemented by each functional module in the electronic device 100, reference may be made to the corresponding descriptions in the embodiment shown in Figures 1 - 6 , which will not be elaborated here.

[0150] The embodiments of the present application further provide a computer storage medium, where the computer storage medium stores computer instructions, and when the computer instructions run on a computer, the computer is caused to execute as Figures 1 - 6 the method described above.

[0151] The foregoing are only the preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification shall be included within the scope of protection of this specification.

Claims

1. A model training method, characterized in that, the method includes: Obtain a first training sample, where the first training sample is collected by a first depth camera; Perform style conversion on the first training sample to obtain a second training sample, and the data distribution of the second training sample is consistent with that of the sample collected by a second depth camera, and the first depth camera is different from the second depth camera; Train a benchmark classification model based on the second training sample to obtain a target classification model.

2. The method according to claim 1, characterized in that, Performing style conversion on the first training sample to obtain a second training sample includes: Performing a preprocessing operation on the first training sample; Input the first training sample after the preprocessing operation into a pre-trained style conversion model to obtain the second training sample, where the style conversion model is a generative model based on the cycle-GAN structure.

3. The method according to claim 2, characterized in that, The preprocessing operation at least includes an object alignment operation and a central cropping operation.

4. The method according to claim 2, characterized in that, Before training the benchmark classification model based on the second training sample to obtain a target classification model, the method further includes: Performing histogram equalization processing on the second training sample to obtain a third training sample; Training the benchmark classification model based on the second training sample to obtain a target classification model includes: Training the benchmark classification model based on the third training sample to obtain the target classification model.

5. The method according to claim 4, characterized in that, Performing histogram equalization processing on the second training sample to obtain a third training sample includes: Performing local histogram equalization processing on the second training sample to obtain the third training sample.

6. The method according to claim 4, characterized in that, Before training the benchmark classification model based on the third training sample to obtain the target classification model, the method further includes: Performing normalization processing on the third training sample to obtain a fourth training sample; Training the benchmark classification model based on the third training sample to obtain the target classification model includes: Training the benchmark classification model based on the fourth training sample to obtain the target classification model.

7. The method according to claim 1, characterized in that, Before training the benchmark classification model based on the second training sample to obtain a target classification model, the method further includes: Obtain a fifth training sample, where the fifth training sample is collected by the second depth camera; Training the benchmark classification model based on the second training sample to obtain a target classification model includes: Training the benchmark classification model based on the first training sample, the second training sample, and the fifth training sample to obtain the target classification model.

8. A model training device, characterized in that, the device includes: An acquisition unit, configured to acquire a first training sample, where the first training sample is acquired by a first depth camera; A conversion unit, configured to perform style conversion on the first training sample to obtain a second training sample, where the data distribution of the second training sample is consistent with that of the samples acquired by a second depth camera, and the first depth camera is different from the second depth camera; A training unit, configured to train a benchmark classification model based on the second training sample to obtain a target classification model.

9. An electronic device, characterized in that, the electronic device includes: at least one processor; a memory coupled to the at least one processor; when the at least one processor executes a computer program stored in the memory, the steps of the method according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the steps of the method according to any one of claims 1-7 are implemented.