Method, apparatus and electronic device for generating feature learning model
By training the target feature encoding model twice, using the features of the first feature and historical feature encoding model, the problem of reoperation when changing the face recognition algorithm is solved, and seamless features are achieved and time and energy savings are achieved.
Patent Information
- Application Number
- CN202111681125.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-30
AI Technical Summary
When face recognition algorithms change in the prior art, the face recognition process needs to be re-operated, resulting in a lot of time and energy consumption.
By extracting the first feature of the first sample data and performing a first training task on the target feature encoding model, then extracting the second feature of the second sample data and obtaining the third feature extracted by the historical feature encoding model, and finally using the second feature and the third feature to perform the second training task on the target feature encoding model until the preset standard is reached.
It is realized that even if the recognition algorithm changes, the characteristics of the sample data it collects are extremely close to the features extracted by the model before the algorithm upgrade, so that there is no need to re-extract the characteristics of the sample data in the base library and build feature indexes, which greatly saves time and energy.
Smart Images

Figure CN114519878B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computer technology, and in particular, to a method, apparatus, and electronic device for generating a feature learning model. Background Art
[0002] The face recognition process generally includes several steps such as face detection, face position calibration, face recognition, feature index generation, and face retrieval. Among them, face detection is to find the area where a face exists in an image; face position calibration is to perform more precise positioning based on the results of face detection and according to the facial features / contour information of the face; face recognition is to map a face image to a high-dimensional feature space through a deep learning algorithm; feature index generation is to generate an index for a large number of face features in the database; face retrieval is to retrieve the database through a face feature index after extracting face features and find face features similar to the currently extracted face features in the database.
[0003] It is found in the actual application process that in the above operation process, when the number of face images in the database is huge, the extraction of face features and the generation of face feature indexes will consume a large amount of time. And when the face recognition algorithm changes, the above face recognition process needs to be operated again completely, which will consume a large amount of time and effort. Summary of the Invention
[0004] The present application provides a method, apparatus, and electronic device for generating a feature learning model to solve the technical problems such as the need to consume a large amount of time and effort for re-operation of the face recognition process when the face feature recognition algorithm changes in the prior art.
[0005] In a first aspect, the present application provides a method for generating a feature learning model, the method comprising:
[0006] Extracting a first feature of first sample data;
[0007] Using the first feature to perform a first training task on a target feature encoding model;
[0008] Extracting a second feature of second sample data;
[0009] Obtaining a third feature corresponding to the second sample data extracted by a historical feature encoding model;
[0010] Using the second feature and the third feature to perform a second training task on the target feature encoding model;
[0011] When the target feature encoding model meets a first preset standard corresponding to the first training task and a second preset standard corresponding to the second training task, determining the target encoding model as a feature learning model.
[0012] In a second aspect, the present application provides an apparatus for generating a feature learning model, the apparatus comprising:
[0013] A feature extraction module, configured to extract first features of first sample data;
[0014] A first training module, configured to perform a first training task on a target feature encoding model;
[0015] The feature extraction module is further configured to extract second features of second sample data;
[0016] An acquisition module, configured to acquire third features corresponding to the second sample data extracted by a historical feature encoding model;
[0017] A second training module, configured to perform a second training task on the target feature encoding model by using the second features and the third features;
[0018] A generation module, configured to determine the target feature encoding model as a feature learning model when the target feature encoding model meets a first preset criterion corresponding to the first training task and meets a second preset criterion corresponding to the second training task.
[0019] In a third aspect, there is provided an electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0020] The memory is configured to store a computer program;
[0021] The processor is configured to implement the steps of the method for generating a feature learning model according to any one of the embodiments of the first aspect when executing the program stored in the memory.
[0022] In a fourth aspect, there is provided a computer-readable storage medium, having stored thereon a computer program, and the computer program, when executed by a processor, implements the steps of the method for generating a feature learning model according to any one of the embodiments of the first aspect.
[0023] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art:
[0024] The method provided by the embodiments of this application extracts the first feature of the first sample data, uses the first feature to perform the first training task on the target feature encoding network model, then extracts the second feature of the second sample data, obtains the third feature extracted by the historical feature encoding network model, and then uses the second feature and the third feature to perform the second training task on the target feature encoding model. The first training task is used to train the high precision of the target feature encoding model. For example, when identifying the same person in different environments, it can be determined that they belong to the same person, that is, it meets the first preset standard. The second training task is used to train the features extracted by the target feature encoding model to be "infinitely close" to the features extracted by the historical feature encoding model, that is, it meets the second preset standard. In this way, even if a certain recognition algorithm changes, such as the algorithm of the target feature encoding model is upgraded, the features of the sample data collected by it can also be ensured to be extremely close to the features extracted by the model before the algorithm upgrade (for example, the historical feature encoding model is the network model before the algorithm upgrade) through a method similar to that introduced in this application. Furthermore, the target feature encoding model can directly use the bottom library, face feature index, etc. before the algorithm upgrade, and no longer needs to re-extract the features of the sample data in the bottom library according to the new algorithm (in order to compare with the features of the test data extracted by the new algorithm later), and there is no need to construct the feature index of the sample data, etc. This greatly saves time and effort. Especially when the quantity of sample data in the bottom library is huge, this effect is particularly obvious. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic flowchart of a method for generating a feature learning model provided by an embodiment of the present invention;
[0026] Figure 2 It is a flowchart of performing the first training task provided by the present invention;
[0027] Figure 3 It is a flowchart of performing the second training task provided by the present invention;
[0028] Figure 4 It is an overall framework diagram of generating a feature learning model provided by the present invention;
[0029] Figure 5 It is a schematic structural diagram of a device for generating a feature learning model provided by an embodiment of the present invention;
[0030] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] To facilitate the understanding of the embodiments of the present invention, the following will further explain with specific embodiments in conjunction with the accompanying drawings. The embodiments do not constitute a limitation to the embodiments of the present invention.
[0033] In response to the technical problems mentioned in the background art, the embodiments of the present application provide a method for generating a feature learning model. Specifically, refer to Figure 1 as shown Figure 1 which is a schematic flowchart of a method for generating a feature learning model provided by the embodiments of the present invention. This method is executed by a target feature encoding model, and the method steps include:
[0034] Step 110, extract the first feature of the first sample data.
[0035] Specifically, in the specific execution process, for the second sample data, it needs to be sample data that meets certain preset standards. For the first sample data, there are not many restrictive requirements. As long as the first sample data and the second sample data are sample data of the same object under different conditions. For example, the first sample data is a person's daily life image, which can include photos of any pose, any expression, such as frontal face, profile face, back face, etc., under any background, for example, photos at any angle, lighting, pose, expression, even in the presence of occlusions. The second sample data, on the other hand, needs to be a photo including relatively standardized and comprehensive user features, such as the user's frontal photo (such as an ID photo). In a specific application scenario, such as identifying the user's identity information through face recognition. Then, the second sample data can be the user's ID photo, and the first sample data can be any of the user's life photos.
[0036] Extract the first feature of the first sample data, such as identifying the face features in the user's life photo. The specific feature extraction process can be implemented through the solutions in the prior art, so no further introduction will be made here.
[0037] Step 120, use the first feature to perform the first training task on the target feature encoding model.
[0038] Specifically, using the first feature, a first training task is performed on the target feature encoding model, that is, the first feature is input into the target feature encoding model to train the target feature encoding model. The main purpose is to train the network model to accurately classify the features of the same user into the same group under various "interference factors" such as different environments, angles, postures, etc. For the features of different users, they are divided into different groups.
[0039] Step 130: Extract the second feature of the second sample data.
[0040] Specifically, the process of feature extraction is the same as the process of extracting the first feature above, and will not be elaborated here.
[0041] Step 140: Obtain the third feature corresponding to the second sample data extracted by the historical feature encoding model.
[0042] Specifically, the process of the historical feature encoding model extracting the third feature of the second sample data can also be achieved by conventional means and will not be elaborated here.
[0043] Step 150: Using the second feature and the third feature, perform a second training task on the target feature encoding model.
[0044] Step 160: When the target feature encoding model meets the first preset standard corresponding to the first training task and the second preset standard corresponding to the second training task, determine the target feature encoding model as the feature learning model.
[0045] Specifically, as introduced in step 120 above, the purpose of performing the first training task on the target feature encoding model using the first feature is to ensure that the target feature encoding model can obtain accurate recognition results when recognizing image data. Even if the image includes various "interference factors" such as different environments, angles, postures, etc., it can still be accurately recognized. That is to say, the first preset standard is that it is hoped that the recognition accuracy of the target feature encoding model for images is higher than the preset accuracy threshold. This includes not only the recognition accuracy for the first sample data, but also the recognition accuracy for any subsequent input images. For example, when different images corresponding to different people are input into the target feature encoding model, the target feature encoding model can recognize these images as the same person.
[0046] While training the target feature encoding model by using the second feature and the third feature simultaneously, it is expected that the second feature extracted by the target feature encoding model can be infinitely close to the third feature extracted by the historical feature encoding model. Usually, the historical feature encoding model is defined as a reference (benchmark) model. The third feature extracted by the historical feature encoding model generally remains unchanged without special circumstances. The second feature is extracted by the target feature encoding model, and its accuracy will be more precise as the target feature encoding model is continuously optimized. In the second training task, it is hoped that the extracted second feature can continuously approach the third feature while ensuring accuracy. That is to say, try to ensure that the features extracted by the target feature encoding model and the historical feature encoding model are consistent in the angular space.
[0047] Through continuous iteration until the target feature encoding model can meet both the first preset standard and the second preset standard.
[0048] In an optional example, the first preset standard corresponding to the first training task is that the accuracy of the target feature encoding model in recognizing images is higher than the preset accuracy threshold, and the second preset standard corresponding to the second training task is that the cosine similarity between the features extracted by the target feature encoding model and the features extracted by the historical feature encoding model is higher than the preset similarity threshold.
[0049] When the target feature encoding model meets the first preset standard and the second network meets the second preset standard, it can be determined that the target feature encoding model is the feature learning model.
[0050] The feature learning model generated in this way can ensure in real time that the model for algorithm upgrade (such as the target feature encoding model in this embodiment) will keep the features extracted consistent with the features extracted by the historical feature encoding model in space regardless of how the algorithm changes.
[0051] When the A feature extracted by the target feature encoding model is extremely close to the B feature extracted by the historical feature encoding model (cosine similarity is higher than the preset similarity threshold), it indicates that the A feature and the B feature belong to the same object. Since then, the image database, feature index, etc. to which the object belongs can be directly used by the network model (such as the target feature encoding model) after the algorithm is continuously upgraded, without the need to create a database corresponding to the network model after the algorithm upgrade, as well as feature indexes, etc., saving a lot of time and manpower.
[0052] Moreover, this method can be applied not only to application scenarios where the algorithm changes (such as algorithm upgrade), but also to application scenarios for cross-model feature comparison.
[0053] For example, when a certain public department needs to search for a specific person, but it is not convenient to publicly disclose the entire person's image information. The feature encoding model corresponding to this public department is used as the historical feature encoding model, and another feature encoding model is selected as the target feature encoding model. After training through the method of this embodiment, it can be ensured that the features extracted by the target feature encoding model are infinitely close to the features extracted by the feature encoding model corresponding to the public department, that is, the effect of recognizing different images of the same person as this person can be achieved. Then, after extracting the image features of the person to be searched in the public department, the image features of multiple people are extracted using the target feature encoding model, and then the image features of different people are compared with the image features of the person to be searched to identify whether it is the person to be searched.
[0054] The method for generating the feature learning model provided by the embodiments of the present application can establish a mapping relationship between two models, that is, ensure that the similarity between different features belonging to the same object extracted by the two models is higher than a preset similarity threshold. Under the condition of ensuring data security and privacy, feature comparison across models can also be achieved.
[0055] Moreover, using the above method, when extracting image features for different objects, the extracted features will be completely different. In this way, feature comparison across models can be completed.
[0056] The idea of the embodiments of the present application is that through the above method, two different versions of models are made to be compatible with respect to the features of the same object. When the features of the same object have established a feature index and an image database under certain conditions, another model can directly use the constructed image database, feature index, etc. through this "mapping relationship" between the two models. Of course, it can also include other content about the same object that has been constructed, etc. Specific details are not limited here.
[0057] The method for generating a feature learning model provided by an embodiment of the present invention extracts the first feature of the first sample data, uses the first feature to perform a first training task on the target feature encoding network model, then extracts the second feature of the second sample data, obtains the third feature extracted by the historical feature encoding network model, and then uses the second feature and the third feature to perform a second training task on the target feature encoding model. The first training task is used to train the high precision of the target feature encoding model. For example, when identifying the same person in different environments, it can be determined that they belong to the same person, that is, it meets the first preset standard. Using the second training task is to train the features extracted by the target feature encoding model to be "infinitely close" to the features extracted by the historical feature encoding model, that is, it meets the second preset standard. In this way, even if a certain recognition algorithm changes, such as the algorithm of the target feature encoding model is upgraded, the features of the sample data collected by it can also be ensured to be extremely close to the features extracted by the model before the algorithm upgrade (for example, the historical feature encoding model is the network model before the algorithm upgrade) through a method similar to that introduced in this application. Furthermore, the target feature encoding model can directly use the bottom library, face feature index, etc. before the algorithm upgrade, and no longer needs to re-extract the features of the sample data in the bottom library according to the new algorithm (in order to compare with the features of the test data extracted by the new algorithm later), and there is no need to construct the feature index of the sample data, etc. This greatly saves time and effort. Especially when the quantity of sample data in the bottom library is huge, this effect is particularly obvious.
[0058] In an optional embodiment, to ensure that the finally obtained feature learning model can meet the first preset standard and the second preset standard at the same time, the first training task and the second training task can be executed simultaneously.
[0059] Further optionally, specifically refer to Figure 2 as shown in Figure 2 which shows the flowchart of executing the first training task.
[0060] Among them, using the first feature to perform the first training task on the target feature encoding model can be achieved in the following way:
[0061] Using the first feature and the pre-obtained first loss function to optimize and train the target feature encoding model until the target feature encoding model reaches the first preset standard corresponding to the first training task.
[0062] Specifically, the first loss function can be a loss function with a classification algorithm function. For example, the ID Loss algorithm (such as the arcface algorithm), or the contrast learning algorithm (such as the Triplet Loss algorithm).
[0063] The ID Loss algorithm maps the face features of different individuals to a high-dimensional angular space through statistical learning methods. Contrastive learning algorithms, such as the Triplet Loss algorithm, map the features of different photos of the same person to adjacent regions and map the face features of different individuals to regions far from each other through contrastive learning methods.
[0064] Regardless of the classification method, the face features of different individuals are mapped to different regions on the unit sphere of the high-dimensional space, while the face features of the same individual are mapped to the same region.
[0065] Figure 2 In [reference], the target feature encoding model extracts the first feature from the first sample data, and then inputs the first feature into the ID Loss algorithm to calculate the difference between the first feature and the target feature (the actual feature of the sample data) in the high-dimensional angular space. Using this difference, the parameters in the target feature encoding model are optimized. The specific implementation can refer to the prior art and will not be elaborated here.
[0066] Figure 3 It shows a flowchart including the process of performing the second training task.
[0067] Specifically refer to Figure 3 As shown, the second sample data is respectively input into the target feature encoding model and the historical feature encoding model, and then the second feature and the third feature can be obtained respectively.
[0068] Then, using the second feature and the third feature, the second training task is performed on the target feature encoding model.
[0069] In an optional implementation, it can be implemented in the following way:
[0070] Using the second feature, the third feature, and the pre-obtained second loss function, the target feature encoding model is optimized and trained until the target feature encoding model reaches the second preset standard corresponding to the second training task.
[0071] Specifically, after obtaining the second feature and the third feature, the second feature and the third feature are simultaneously input into the Cosine Loss function to calculate the difference between the second feature and the third feature.
[0072] Then, using the difference between the second feature and the third feature, the target feature encoding model is iteratively optimized. Until the features extracted by the target feature encoding model and the features of the historical feature encoding model have the smallest difference (that is, ensuring that the cosine similarity between the two is higher than the preset similarity threshold), that is, ensuring that the second feature and the third feature are as consistent as possible in the angular space.
[0073] Through continuous iterative training by the above method until the final generation of the feature learning model ends.
[0074] In the above manner, the two training tasks are trained simultaneously to ensure that the target feature encoding model can process more complex images as much as possible. While still being able to extract image features from complex images, it can also ensure that the extracted image features are infinitely close to the features of the standard images extracted by the historical feature encoding model.
[0075] Figure 4 It is the overall framework diagram of the tasks to be executed by the feature learning model. For details, please refer to Figure 4 as shown. The first training task and the second training task are executed simultaneously. The target feature encoding model in the first training task on the left and the target feature encoding model in the second training task on the right share the same parameters. In practical applications, it is actually the same target feature encoding model. It is only shown in this way here for better illustration. Moreover, Figure 4 as shown. And, Figure 4 the result obtained by the CossineLoss function in will be input into the target feature encoding model to optimize the parameters of the target feature encoding model.
[0076] Therefore, the work performed by the functional modules in this architecture and the specific working principles have been described in detail above, so they will not be elaborated here too much.
[0077] Figure 5 This is a device for generating a feature learning model provided by an embodiment of the present invention. The device includes: a feature extraction module 501, a first training module 502, an acquisition module 503, a second training module 504, and a generation module 505.
[0078] The feature extraction module 501 is used to extract the first feature of the first sample data;
[0079] The first training module 502 is used to perform the first training task on the target feature encoding model;
[0080] The feature extraction module 501 is also used to extract the second feature of the second sample data;
[0081] The acquisition module 503 is used to acquire the third feature corresponding to the second sample data extracted by the historical feature encoding model;
[0082] The second training module 504 is used to perform the second training task on the target feature encoding model by using the second feature and the third feature;
[0083] A generation module 505, configured to determine the target feature encoding model as a feature learning model when the target feature encoding model meets a first preset criterion corresponding to a first training task and a second preset criterion corresponding to a second training task.
[0084] Optionally, the first preset criterion includes: the accuracy of the target feature encoding model in recognizing images is higher than a preset accuracy threshold; the second preset criterion includes: the similarity between the features extracted by the target feature encoding model and the features extracted by the historical feature encoding model is higher than a preset similarity threshold.
[0085] Optionally, the first training module 502 and the second training module 504 execute the training tasks simultaneously.
[0086] Optionally, the second sample data is sample data meeting a preset criterion, and both the first sample data and the second sample data are sample data of the same object under different conditions.
[0087] Optionally, the first feature training module is specifically configured to use the first feature and a pre-obtained first loss function to perform tuning training on the target feature encoding model until the target feature encoding model meets the first preset criterion corresponding to the first training task.
[0088] Optionally, the second feature training module is specifically configured to use the second feature, the third feature, and a pre-obtained second loss function to perform tuning training on the target feature encoding model until the target feature encoding model meets the second preset criterion corresponding to the second training task.
[0089] The functions performed by each component in the generation device of a feature learning model provided in the embodiments of the present invention have been described in detail in any of the above method embodiments, and thus will not be elaborated here.
[0090] An apparatus for generating a feature learning model provided by an embodiment of the present invention extracts a first feature of first sample data, performs a first training task on a target feature encoding network model using the first feature, then extracts a second feature of second sample data, obtains a third feature extracted by a historical feature encoding network model, and then performs a second training task on the target feature encoding model using the second feature and the third feature. The first training task is used to train the high precision of the target feature encoding model. For example, when identifying the same person in different environments, it can be determined that they belong to the same person, that is, it reaches a first preset standard. Using the second training task is to train the features extracted by the target feature encoding model to be "infinitely close" to the features extracted by the historical feature encoding model, that is, it reaches a second preset standard. In this way, even if a certain recognition algorithm changes, such as the algorithm of the target feature encoding model is upgraded, the features of the sample data collected by it can also be ensured to be extremely close to the features extracted by the model before the algorithm upgrade (for example, the historical feature encoding model is the network model before the algorithm upgrade) through a method similar to that introduced in this application. Furthermore, the target feature encoding model can directly use the database, face feature index, etc. before the algorithm upgrade, and no longer needs to re-extract the features of the sample data in the database according to the new algorithm (in order to compare with the features of the test data extracted by the new algorithm later), and there is no need to construct a feature index of the sample data, etc. This greatly saves time and effort. Especially when the quantity of sample data in the database is huge, this effect is particularly obvious.
[0091] As Figure 6 shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete communication with each other through the communication bus 114.
[0092] The memory 113 is used to store a computer program;
[0093] In an embodiment of the present application, when the processor 111 executes the program stored on the memory 113, it implements the method for generating a feature learning model provided by any one of the foregoing method embodiments, including:
[0094] Extract a first feature of first sample data;
[0095] Use the first feature to perform a first training task on the target feature encoding model;
[0096] Extract a second feature of second sample data;
[0097] Obtain a third feature corresponding to the second sample data extracted by the historical feature encoding model;
[0098] Using the second feature and the third feature, perform a second training task on the target feature encoding model;
[0099] When the target feature encoding model meets the first preset standard corresponding to the first training task and the second preset standard corresponding to the second training task, determine the target encoding model as the feature learning model.
[0100] Optionally, the first preset standard includes: the accuracy of the target feature encoding model in recognizing images is higher than a preset accuracy threshold; the second preset standard includes: the similarity between the features extracted by the target feature encoding model and the features extracted by the historical feature encoding model is higher than a preset similarity threshold.
[0101] Optionally, the first training task and the second training task are performed simultaneously.
[0102] Optionally, the second sample data is sample data that meets the preset standard, and both the first sample data and the second sample data are sample data of the same object under different conditions.
[0103] Optionally, use the first feature and a pre-obtained first loss function to fine-tune and train the target feature encoding model until the target feature encoding model meets the first preset standard corresponding to the first training task.
[0104] Optionally, use the second feature, the third feature, and a pre-obtained second loss function to fine-tune and train the target feature encoding model until the target feature encoding model meets the second preset standard corresponding to the second training task.
[0105] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for generating a feature learning model provided in any of the foregoing method embodiments are implemented.
[0106] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0107] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for generating a feature learning model, characterized in that, the method includes: extracting a first feature of the first sample image data; using the first feature to perform a first training task on the target feature encoding model; extracting a second feature of the second sample image data; obtaining a third feature corresponding to the second sample image data extracted by the historical feature encoding model; using the second feature and the third feature to perform a second training task on the target feature encoding model; when the target feature encoding model meets the first preset standard corresponding to the first training task and the second preset standard corresponding to the second training task, determining the target feature encoding model as the feature learning model.
2. The method according to claim 1, characterized in that, the first preset standard includes: the accuracy of the target feature encoding model in recognizing images is higher than a preset accuracy threshold; the second preset standard includes: the similarity between the features extracted by the target feature encoding model and the features extracted by the historical feature encoding model is higher than a preset similarity threshold.
3. The method according to claim 1, characterized in that, the first training task and the second training task are executed simultaneously.
4. The method according to any one of claims 1-3, characterized in that, the second sample image data is sample image data that meets the preset standard, and both the first sample image data and the second sample image data are sample image data of the same object under different conditions.
5. The method according to any one of claims 1-3, characterized in that, the step of using the first feature to perform a first training task on the target feature encoding model specifically includes: using the first feature and a pre-obtained first loss function to perform tuning training on the target feature encoding model until the target feature encoding model meets the first preset standard corresponding to the first training task.
6. The method according to any one of claims 1-3, characterized in that, the step of using the second feature and the third feature to perform a second training task on the target feature encoding model specifically includes: using the second feature, the third feature, and a pre-obtained second loss function to perform tuning training on the target feature encoding model until the target feature encoding model meets the second preset standard corresponding to the second training task.
7. A device for generating a feature learning model, characterized in that, the device includes: a feature extraction module for extracting a first feature of the first sample image data; a first training module for performing a first training task on the target feature encoding model; the feature extraction module is further configured to extract a second feature of the second sample image data; an acquisition module for obtaining a third feature corresponding to the second sample image data extracted by the historical feature encoding model; a second training module for performing a second training task on the target feature encoding model using the second feature and the third feature; A generation module, configured to determine the target feature encoding model as the feature learning model when the target feature encoding model meets a first preset standard corresponding to the first training task and a second preset standard corresponding to the second training task.
8. The apparatus according to claim 7, wherein, the first preset standard includes: the accuracy of the target feature encoding model in recognizing images is higher than a preset accuracy threshold; the second preset standard includes: the similarity between the features extracted by the target feature encoding model and the features extracted by the historical feature encoding model is higher than a preset similarity threshold.
9. An electronic device, wherein, it includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; the memory is used to store a computer program; the processor is configured to implement the steps of the method for generating the feature learning model according to any one of claims 1-6 when executing the program stored on the memory.
10. A computer-readable storage medium, on which a computer program is stored, wherein, the computer program, when executed by a processor, implements the steps of the method for generating the feature learning model according to any one of claims 1-6.
Citation Information
Patent Citations
Business model hyper-parameter configuration determination method and device, equipment and storage medium
CN112488245A