Target recognition method and neural network training method
By using the pre-trained neural network of the first client on the second client to assist in training the neural network of the second client and fusing feature data, the problem of improving the accuracy of deep learning network caused by data silos is solved, and efficient multi-party data utilization and target recognition accuracy are achieved.
Patent Information
- Application Number
- CN202210322086.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-03-29
AI Technical Summary
Due to the existence of data silos, the data that a single enterprise or client can obtain is limited, making it difficult to improve the accuracy limit of deep learning networks. The deployment of existing federated learning methods is difficult, costly and poor flexibility.
By using the neural network pre-trained by the first client on the second client, assisting in training the neural network of the second client, fusing the feature data of the two until the convergence conditions are met, a target neural network with better compatibility is formed.
On the premise of ensuring data security, we can realize the simultaneous utilization of multi-party island data, improve the network's ability to identify multi-party data, and improve the target recognition accuracy.
Smart Images

Figure CN114912572B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly relates to a target recognition method and a training method for a neural network. Background Art
[0002] Data islands refer to data resources accumulated between different enterprises or clients. For purposes such as privacy protection or data security, they are like independent islands and cannot be connected and interacted with each other, and there is a lack of relevance between the data on the islands.
[0003] With the advent of the big data era, data has gradually become a new production factor. Data is the cornerstone of a deep learning network (DNN, Deep Neural Network). However, due to the existence of data islands, the data that a single enterprise or client can obtain is very limited, resulting in difficulty in improving the upper limit of the accuracy of the DNN. Summary of the Invention
[0004] To improve the accuracy of a neural network, embodiments of the present disclosure provide a target recognition method and apparatus, a training method and apparatus for a neural network, an electronic device, and a storage medium.
[0005] In a first aspect, an embodiment of the present disclosure provides a training method for a neural network, which is applied to a second client, and the method includes:
[0006] Obtain a first neural network; the first neural network is pre-trained by a first client using first sample image data, and the first sample image data is first island data that the first client can obtain;
[0007] Input second sample image data into the first neural network to obtain first feature data output by the first neural network; the second sample image data is second island data that the second client can obtain;
[0008] Train a second neural network to be trained according to the second sample image data and the first feature data to obtain a trained second neural network;
[0009] Use the second sample image data to train a target neural network obtained by fusing the trained second neural network and the first neural network until a convergence condition is satisfied.
[0010] In some embodiments, the second sample image data includes multiple target categories, and each target category includes at least one target sample image; the inputting the second sample image data into the first neural network to obtain first feature data output by the first neural network includes:
[0011] Input each target sample image into the first neural network to obtain the first image features corresponding to each target sample image output by the first neural network;
[0012] For each of the target categories, determine the first type of central feature and the class feature range corresponding to the target category according to the first image features of each target sample image in the target category;
[0013] Determine the first type of central feature and the class feature range as the first feature data corresponding to the target category.
[0014] In some embodiments, the determining the first type of central feature and the class feature range corresponding to the target category according to the first image features of each target sample image in the target category includes:
[0015] Determine the first type of central feature corresponding to the target category according to the first image features of each target sample image in the target category;
[0016] Determine the similarity between each first image feature in the target category and the first type of central feature, and determine the class feature range corresponding to the target category according to the maximum and minimum values of the similarity.
[0017] In some embodiments, training the second neural network to be trained according to the second sample image data and the first feature data to obtain the trained second neural network includes:
[0018] Input the second sample image data into the second neural network to be trained to obtain the second feature data output by the second neural network;
[0019] Based on the first difference between the second feature data and the label data included in the second sample image data, and the second difference between the second feature data and the first feature data, adjust the network parameters of the second neural network until the convergence condition is met to obtain the trained second neural network.
[0020] In some embodiments, the second sample image data includes multiple target categories, where each target category includes at least one target sample image; the inputting the second sample image data into the second neural network to be trained to obtain the second feature data output by the second neural network includes:
[0021] Input each target sample image into the second neural network to be trained to obtain the second image features corresponding to each target sample image output by the second neural network;
[0022] For each of the target categories, according to the second image features of each target sample image in the target category, determine the second type of central feature corresponding to the target category;
[0023] Determine the second type of central feature and the second image features corresponding to each target sample image as the second feature data corresponding to the target category.
[0024] In some embodiments, the first feature data of each target category includes the first type of central feature and the class feature range of the target category to which each target sample image belongs; the adjusting the network parameters of the second neural network based on the first difference between the second feature data and the label data included in the second sample image data, and the second difference between the second feature data and the first feature data, includes:
[0025] Determine the first difference according to the second image features corresponding to each target sample image and the label data;
[0026] For each of the target categories, determine the second difference according to the difference between the second type of central feature corresponding to the target category and the first type of central feature, and the difference between the second image features corresponding to each target sample image and the class feature range;
[0027] Adjust the network parameters of the second neural network based on the first difference and the second difference.
[0028] In some embodiments, the training the target neural network obtained by fusing the trained second neural network and the first neural network by using the second sample image data includes:
[0029] Input the second sample image data into the trained second neural network to obtain third feature data output by the second neural network;
[0030] Fuse the first feature data and the third feature data according to the fusion weight determined based on the second sample image data to obtain fused feature data;
[0031] Adjust the network parameters of the target neural network based on the third difference between the fused feature data and the label data included in the second sample image data until the convergence condition is satisfied.
[0032] In some embodiments, the fusing the first feature data and the third feature data according to the fusion weight determined based on the second sample image data to obtain fused feature data includes:
[0033] Input the second sample image data into the attribute network of the target neural network to obtain the attribute information output by the attribute network;
[0034] Determine the first weight of the first feature data and the second weight of the third feature data according to the attribute information;
[0035] Based on the first weight and the second weight, perform fusion processing on the first feature data and the third feature data to obtain the fused feature data.
[0036] In some embodiments, the first neural network includes a first face recognition network, and the first sample image data is first face image data.
[0037] In some embodiments, the second neural network includes a second face recognition network, and the second sample image data is second face image data.
[0038] In a second aspect, an embodiment of the present disclosure provides a target recognition method, including:
[0039] Obtain a to-be-tested image, where the to-be-tested image includes a to-be-tested target;
[0040] Input the to-be-tested image into a pre-trained target recognition network to obtain a recognition result output by the target recognition network; the target recognition network is a target neural network obtained according to the training method of any embodiment of the first aspect.
[0041] In a third aspect, an embodiment of the present disclosure provides a training device for a neural network, which is applied to a second client, and the device includes:
[0042] A network acquisition module, configured to acquire a first neural network; the first neural network is pre-trained by a first client using first sample image data, and the first sample image data is first isolated island data that the first client can acquire;
[0043] A first processing module, configured to input second sample image data into the first neural network to obtain first feature data output by the first neural network; the second sample image data is second isolated island data that the second client can acquire;
[0044] A first training module, configured to train a second neural network to be trained according to the second sample image data and the first feature data to obtain a trained second neural network;
[0045] The second training module is configured to use the second sample image data to train a target neural network obtained by fusing the trained second neural network and the first neural network until a convergence condition is met.
[0046] In some embodiments, the second sample image data includes multiple target categories, where each target category includes at least one target sample image; the first processing module is configured to:
[0047] Input each target sample image into the first neural network to obtain first image features corresponding to each target sample image output by the first neural network;
[0048] For each of the target categories, determine a first class center feature and a class feature range corresponding to the target category according to the first image features of each target sample image in the target category;
[0049] Determine the first class center feature and the class feature range as the first feature data corresponding to the target category.
[0050] In some embodiments, the first processing module is configured to:
[0051] Determine the first class center feature corresponding to the target category according to the first image features of each target sample image in the target category;
[0052] Determine the similarity between each first image feature in the target category and the first class center feature, and determine the class feature range corresponding to the target category according to the maximum and minimum values of the similarity.
[0053] In some embodiments, the first training module is configured to:
[0054] Input the second sample image data into a second neural network to be trained to obtain second feature data output by the second neural network;
[0055] Based on a first difference between the second feature data and label data included in the second sample image data, and a second difference between the second feature data and the first feature data, adjust network parameters of the second neural network until a convergence condition is met to obtain the trained second neural network.
[0056] In some embodiments, the second sample image data includes multiple target categories, where each target category includes at least one target sample image; the first training module is configured to:
[0057] Input each target sample image into the second neural network to be trained, and obtain the second image feature corresponding to each target sample image output by the second neural network;
[0058] For each of the target categories, determine the second type of central feature corresponding to the target category according to the second image features of each target sample image in the target category;
[0059] Determine the second type of central feature and the second image feature corresponding to each target sample image as the second feature data corresponding to the target category.
[0060] In some embodiments, the first feature data of each target category includes the first type of central feature and the class feature range of the target category to which each target sample image belongs; the first training module is configured to:
[0061] Determine the first difference according to the second image feature corresponding to each target sample image and the label data;
[0062] For each of the target categories, determine the second difference according to the difference between the second type of central feature corresponding to the target category and the first type of central feature, and the difference between the second image feature corresponding to each target sample image and the class feature range;
[0063] Adjust the network parameters of the second neural network based on the first difference and the second difference.
[0064] In some embodiments, the second training module is configured to:
[0065] Input the second sample image data into the trained second neural network to obtain the third feature data output by the second neural network;
[0066] Perform a fusion process on the first feature data and the third feature data according to the fusion weight determined based on the second sample image data to obtain the fusion feature data;
[0067] Adjust the network parameters of the target neural network based on the third difference between the fusion feature data and the label data included in the second sample image data until the convergence condition is met.
[0068] In some embodiments, the second training module is configured to:
[0069] Input the second sample image data into the attribute network of the target neural network to obtain the attribute information output by the attribute network;
[0070] Determine a first weight of the first feature data and a second weight of the third feature data according to the attribute information;
[0071] Based on the first weight and the second weight, perform a fusion process on the first feature data and the third feature data to obtain the fused feature data.
[0072] In some embodiments, the first neural network includes a first face recognition network, and the first sample image data is first face image data.
[0073] In some embodiments, the second neural network includes a second face recognition network, and the second sample image data is second face image data.
[0074] In a fourth aspect, an embodiment of the present disclosure provides an object recognition device, including:
[0075] An image acquisition module configured to acquire a to-be-detected image, where the to-be-detected image includes a to-be-detected object;
[0076] A second processing module configured to input the to-be-detected image into a pre-trained object recognition network to obtain a recognition result output by the object recognition network; the object recognition network is an object neural network obtained according to the training method of any one of the first aspects.
[0077] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including:
[0078] A processor; and
[0079] A memory storing computer instructions, and the computer instructions are used to cause the processor to execute the method according to any one of the first aspect or the second aspect.
[0080] In a sixth aspect, an embodiment of the present disclosure provides a storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of the first aspect or the second aspect.
[0081] The neural network training method of the present disclosure embodiment is applied to a second client. The method includes obtaining a first neural network pre-trained by a first client using first sample image data, inputting second sample image data into the first neural network to obtain first feature data output by the first neural network, training a second neural network to be trained according to the second sample image data and the first feature data to obtain a trained second neural network, and training a target neural network formed by fusing the second neural network and the first neural network using the second sample image data until a convergence condition is met. In the embodiment of the present disclosure, while ensuring the security of island data, the simultaneous utilization of multi-party island data is realized. The trained target neural network has better compatibility, improves the network's recognition ability for multi-party data, and greatly improves the target recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0083] Figure 1 is a schematic structural diagram of a neural network training system according to some embodiments of the present disclosure.
[0084] Figure 2 is a flowchart of a neural network training method according to some embodiments of the present disclosure.
[0085] Figure 3 is a flowchart of a neural network training method according to some embodiments of the present disclosure.
[0086] Figure 4 is a flowchart of a neural network training method according to some embodiments of the present disclosure.
[0087] Figure 5 is a flowchart of a neural network training method according to some embodiments of the present disclosure.
[0088] Figure 6 is a flowchart of a neural network training method according to some embodiments of the present disclosure.
[0089] Figure 7 is a flowchart of a neural network training method according to some embodiments of the present disclosure.
[0090] Figure 8 is a flowchart of a neural network training method according to some embodiments of the present disclosure.
[0091] Figure 9 It is a schematic diagram of a method for training a neural network according to some embodiments of the present disclosure.
[0092] Figure 10 It is a flowchart of a method for training a neural network according to some embodiments of the present disclosure.
[0093] Figure 11 It is a flowchart of a method for training a neural network according to some embodiments of the present disclosure.
[0094] Figure 12 It is a schematic diagram of a method for training a neural network according to some embodiments of the present disclosure.
[0095] Figure 13 It is a flowchart of a target recognition method according to some embodiments of the present disclosure.
[0096] Figure 14 It is a block diagram of the structure of a training device for a neural network according to some embodiments of the present disclosure.
[0097] Figure 15 It is a block diagram of the structure of a target recognition device according to some embodiments of the present disclosure.
[0098] Figure 16 It is a block diagram of the structure of an electronic device according to some embodiments of the present disclosure. Detailed Embodiments
[0099] The technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present disclosure belong to the scope of protection of the present disclosure. In addition, the technical features involved in different embodiments of the present disclosure described below can be combined with each other as long as they do not conflict with each other.
[0100] With the advent of the big data era, data has become the most important production factor. Data is the cornerstone of a deep learning network (DNN, Deep Neural Network). If more data can be utilized, there will be more potential to break through the accuracy ceiling of the neural network. However, due to the existence of data islands, the data that can be accumulated and obtained by a single enterprise or client is very limited, so it is difficult to improve the accuracy ceiling of the DNN.
[0101] Taking the face recognition scenario as an example, using more face image data to train a face recognition network is an effective method to improve the accuracy of face recognition. However, due to the high privacy of face data, the data accumulated by different enterprises or clients are stored and maintained independently, and the databases are isolated from each other, forming data islands.
[0102] For example, enterprise A has accumulated a large amount of face data of users wearing masks collected by cameras, while enterprise B has accumulated a large amount of face data of users not wearing masks uploaded by users. If the face data of both enterprise A and enterprise B can be used to train the face recognition network at the same time, the face recognition network can learn the face features of both types of data simultaneously, greatly improving the accuracy of face recognition. However, due to the existence of data islands, the isolated data of enterprise A and enterprise B cannot be obtained simultaneously, resulting in the upper limit of the accuracy of the face recognition network being difficult to further improve.
[0103] In related technologies, the data islands can be broken through the method of Federated Learning to realize the utilization of isolated data. The basic principle of federated learning is to construct an encrypted communication environment between a "server - multiple clients", where each client has isolated data stored and maintained independently. The server sends the DNN to each client, so that each client uses its own isolated data to locally train the DNN, encrypts and sends the trained network parameters and gradients to the server, and then the server performs aggregation processing to obtain the finally trained DNN, so as to achieve the purpose that the isolated data can participate in network training without leaving the local area.
[0104] However, federated learning requires the construction of a huge encrypted communication environment to realize the encrypted communication between the server and each client, with high deployment difficulty, long time cycle, high cost, and at the same time requires the same DNN network structure for each client, greatly restricting the flexibility of isolated data utilization and network training efficiency.
[0105] Based on this, the embodiments of the present disclosure provide a training method and device for a face recognition network, a face recognition method and device, an electronic device, and a storage medium, aiming to break data islands, effectively utilize isolated data from all parties to participate in network training, improve the accuracy of face recognition, and have simple deployment and low cost.
[0106] Figure 1 The structure of the training system of the face recognition network according to the embodiments of the present disclosure is shown. The application environment of the embodiments of the present disclosure will be described below in combination with Figure 1 to illustrate the application environment of the embodiments of the present disclosure.
[0107] As Figure 1As shown, the training system includes a first client 100 and a second client 200. The first client 100 and the second client 200 can establish a wired or wireless communication connection through a network 300.
[0108] In the embodiments of the present disclosure, the first client 100 stores first sample image data, and the second client 200 stores second sample image data. It can be understood that according to different application scenarios, the data types of the first sample image data and the second sample image data can be set accordingly. For example, they can be face data, natural scene data, vehicle data, etc. The present disclosure places no restrictions on this.
[0109] Among them, the first sample image data and the second sample image data are mutually isolated data. That is, the first client 100 cannot obtain the second sample image data, and the second client 200 also cannot obtain the first sample image data.
[0110] In one example, the first client 100 and the second client 200 can be two different enterprises. For example, the first client 100 is enterprise A, which has face data independently accumulated, stored, and maintained, that is, the first sample image data. The second client 200 is enterprise B, which has face data independently accumulated, stored, and maintained, that is, the second sample image data. For the purposes of data value or privacy protection, etc., the data between enterprise A and enterprise B cannot be interconnected. That is, enterprise A cannot obtain the second sample image data, and similarly, enterprise B cannot obtain the first sample image data.
[0111] In another example, the first client 100 and the second client 200 can be different departments of the same enterprise. For example, the first client 100 is department A of enterprise X, which has natural scene data independently accumulated, stored, and maintained, that is, the first sample image data. The second client 200 is department B of enterprise X, which has natural scene data independently accumulated, stored, and maintained, that is, the second sample image data. For the purposes of data value or privacy protection, etc., the data between department A and department B cannot be interconnected. That is, department A cannot obtain the second sample image data, and similarly, department B cannot obtain the first sample image data.
[0112] Of course, it can be understood that the application scenarios of the system in the embodiments of the present disclosure are not limited to the above examples, and can also be applied to any scenario suitable for jointly training a network with two or more isolated data. The present disclosure will not enumerate them here.
[0113] On the basis of what is shown in Figure 1 the embodiments of the present disclosure provide a method for training a neural network. This method can be applied to the second client to enable the first sample image data and the second sample image data to participate in network training simultaneously.
[0114] As Figure 2 shown, in some embodiments, the method for training a neural network according to the examples of the present disclosure includes:
[0115] S210. Obtain a first neural network.
[0116] In the embodiments of the present disclosure, the first face recognition network is pre-trained by the first client using the first sample image data.
[0117] Referring to Figure 1 it can be seen that between the first client 100 and the second client 200, although the isolated data stored in the two cannot be interchanged, a communicable connection is established between the two. Thus, the first client 100 can pre-train the first neural network using the first sample image data that it can obtain, and then send the trained first neural network to the second client 200 through the network 300.
[0118] In the following embodiments of the present disclosure, the process of the first client training the first neural network is specifically described and will not be elaborated here for the time being.
[0119] S220. Input the second sample image data into the first neural network to obtain first feature data output by the first neural network.
[0120] In the embodiments of the present disclosure, after the second client receives the first neural network sent by the first client, it can use the first neural network to assist in training the second neural network, so that the second neural network can also have the ability to recognize the features learned by the first neural network.
[0121] Specifically, the second client can obtain the second sample image data stored by itself. In the embodiments of the present disclosure, first, the second sample image data is input into the received first neural network to obtain first feature data.
[0122] It can be understood that the first neural network is trained using the first sample image data of the first client 100. That is to say, the first neural network has good recognition ability for the features of the first sample image data. In the embodiments of the present disclosure, on the side of the second client 200, the second sample image data is input into the first neural network, so that the recognition features of the first neural network for the second sample image data can be extracted, that is, the first feature data.
[0123] In the embodiments of the present disclosure, it is precisely the first feature data extracted by the first neural network that is used to assist in training the second neural network on the side of the second client 200, so that the second neural network can be compatible with the recognition ability of the first neural network, which will be described in S230 below.
[0124] S230. Train the second neural network to be trained according to the second sample image data and the first feature data, and obtain the trained second neural network.
[0125] Specifically, on the side of the second client, it has the second sample image data stored by itself, and also has the first feature data obtained by performing feature extraction on the second sample image data through S220.
[0126] It can be understood that if only the second sample image data is used to train the second neural network, then the second neural network can only learn the features of the second sample image data, and will not learn the features of the first sample image data. In the embodiments of the present disclosure, the first feature data is fused into the training process of the second neural network, so that the second neural network can be compatible with the feature recognition of the first feature data.
[0127] That is, in the embodiments of the present disclosure, when training the second neural network on the side of the second client 200, the optimization of the objective function mainly includes the following two parts:
[0128] 1) The classification loss of the second sample image data.
[0129] It represents the recognition ability of the second neural network for the second sample image data, that is, the difference between the predicted value and the label value of the classification of the second sample image data.
[0130] 2) The loss between the features extracted by the second neural network and the features extracted by the first neural network.
[0131] It represents the difference between the features extracted by the first neural network and the features extracted by the second neural network for the same second sample image data. It can be understood that the smaller the difference between the two, the stronger the ability of the second neural network to be compatible with the first neural network, and when using the second neural network to recognize a target closer to the first sample image data, the network recognition accuracy is higher.
[0132] Therefore, by constraining the above two loss terms, the second neural network to be trained is trained using the first feature data and the second sample image data, and a second neural network that is compatible with both the first sample image data and the second sample image data is obtained.
[0133] S240. Use the second sample image data to train the target neural network obtained by fusing the trained second neural network and the first neural network until the convergence condition is met.
[0134] Specifically, in the embodiments of the present disclosure, after obtaining the trained second neural network, the second neural network and the first neural network are fused to obtain a target neural network, and then the target neural network is trained again using the second sample image data until the convergence condition is met, and the training is completed to obtain the final target neural network for use in the prediction stage.
[0135] It should be noted that in the embodiments of the present disclosure, the trained second neural network is not directly used as the final target neural network, but the first neural network and the second neural network obtained previously are fused to obtain the target neural network. In this way, when training the target neural network, the features extracted by the network include both the features extracted by the first neural network and the features extracted by the second neural network, so that the target neural network has good robustness for both the first sample image data and the second sample image data.
[0136] The process of training the fusion network will be described in the following embodiments of the present disclosure and will not be elaborated here for the time being.
[0137] In some embodiments, after training the target neural network, the target neural network can be sent to the first client 100. It can be understood that the first client 100 and the second client 200 only transmit the first neural network once at the beginning and the target neural network once at the end, and there is no interaction between their isolated island data. On the premise of ensuring data security, the simultaneous utilization of the isolated island data of both parties is realized.
[0138] As can be seen from the above, in the embodiments of the present disclosure, on the premise of ensuring the security of the isolated island data, the simultaneous utilization of the isolated island data of multiple parties is realized, and the trained target neural network has better compatibility, improves the network's recognition ability for multiple-party data, and greatly improves the target recognition accuracy.
[0139] It should be noted that in the following embodiments of the present disclosure, the neural network training method of the present disclosure will be described by taking the face recognition scenario as an example. That is, in the following embodiments, the first sample image data is the first face image data, the second sample image data is the second face image data, the first neural network can include the first face recognition network, and the second neural network can include the second face recognition network.
[0140] However, it can be understood that the present disclosure is not limited to the face recognition scenario, but can also be any other applicable application scenarios, such as the vehicle recognition, natural scene recognition and other scenarios mentioned above. The present disclosure will not elaborate on this anymore. The following will be described in combination with Figure 2 embodiments.
[0141] In the embodiments of the present disclosure, whether it is the first sample image data or the second sample image data, both include a plurality of sample data, and each sample data includes a target sample image, corresponding label data, and the target category to which it belongs. Taking the face recognition scenario as an example, whether it is the first face data or the second face data, both include a plurality of sample data, and each sample data includes a face sample image, corresponding label data, and a face category. The face category indicates the category to which the face sample image belongs. For example, multiple face sample images of the same person belong to the same face category. The label data represents the ground truth that the corresponding face sample image belongs to a certain face category, and the label data can be obtained by manual annotation.
[0142] In short, both the first face data and the second face data include a plurality of sample data, and these sample data belong to multiple face categories. Among them, each sample data includes a face sample image and the label data corresponding to the face sample image.
[0143] In some embodiments, for the first client 100, the first client 100 can use the first face data stored in itself to train the first face recognition network, so as to obtain the trained first face recognition network. The following is combined with Figure 3 the embodiments for description.
[0144] As Figure 3 shown, in some embodiments, the process of the first client 100 training the first face recognition network includes:
[0145] S310: Input the first sample image data into the first neural network to be trained, and obtain the output result of the first neural network.
[0146] S320: Adjust the network parameters of the first neural network according to the difference between the output result and the label data until the convergence condition is met, and obtain the trained first neural network.
[0147] In the embodiments of the present disclosure, there is no limitation on the specific network structure of the first face recognition network, and any face recognition network suitable for implementation can be adopted. For example, in one example, the first face recognition network can adopt the FaceNet network structure, which is expressed as:
[0148]
[0149] In formula (1), x S represents the face sample image in the first face data, M s represents the first face recognition network, Represents the image features extracted by the first face recognition network. Taking a sample data as an example below, the training process of the first face recognition network will be described.
[0150] Input the face sample image included in the sample data into the first face recognition network to be trained. The first face recognition network performs processes such as convolution, pooling, and classification to obtain the classification result for the face sample image, that is, the output result.
[0151] It can be understood that the output result represents the predicted value of the first face recognition network, while the label data represents the true value of the face sample image. Thus, the difference between the output result and the label data, that is, the loss, is calculated through a pre-constructed objective function. According to this difference, backpropagation is used to optimize and adjust the network parameters of the first face recognition network.
[0152] The above is only described by taking one sample data as an example. For multiple sample data in the first face data, repeat the above process to continuously iterate and optimize the first face recognition network until the convergence condition is met, thereby obtaining the trained first face recognition network.
[0153] After the first client 100 obtains the trained first face recognition network, it can send the first face recognition network to the second client 200 through the network 300. It can be understood that only the first face recognition network is transmitted between the first client 100 and the second client 200, and there is no transmission of any first face data. Therefore, it is ensured that the data on the first client 100 side does not leave the local area, ensuring data security.
[0154] On the side of the second client 200, after receiving the first face recognition network sent by the first client 100, it can use the second face data stored by itself and the first face recognition network to perform compatibility training on the second face recognition network. The following is combined with Figures 4 to 6 Embodiments are described.
[0155] As Figure 4 shown, in some embodiments, in the training method of the present disclosure example, the process of obtaining the first face feature data includes:
[0156] S410. Input each target sample image into the first neural network to obtain the first image feature corresponding to each target sample image output by the first neural network.
[0157] Specifically, in the face recognition scenario, the target sample image, that is, the second face data, includes face sample images in the sample data. Taking one sample data as an example, the face sample image included in the sample data can be input into the first face recognition network, so that the first face recognition network extracts features from the face sample image through, for example, a convolutional layer, and obtains the first image feature corresponding to the face sample image. The specific process can be similar to the aforementioned formula (1), and will not be elaborated here.
[0158] S420. For each target category, according to the first image features of each target sample image in the target category, determine the first type of central feature and the class feature range corresponding to the target category.
[0159] Specifically, in the face recognition scenario, the target category is the face category to which the face sample image belongs. Combining the foregoing, it can be known that the second face data includes multiple face categories. For example, the face sample images belonging to the same person in the second face data can belong to the same face category. That is, each face category includes at least one face sample image.
[0160] Taking any face category as an example, this face category includes N face sample images. Through formula (2) in the aforementioned S410, N first image features corresponding to the N face sample images can be extracted. For this face category, the first type of central feature corresponding to this face category can be calculated according to the first image features corresponding to the N face sample images included, expressed as:
[0161]
[0162] In formula (2), represents the first image feature corresponding to the k-th face sample image in the i-th category of the second face data extracted by using the first face recognition network M s and represents the first type of central feature. It can be understood that the first type of central feature represents the average feature of the face sample images of this face category, and it can reflect the average feature value of this face category.
[0163] For any face category, through the above process in sequence, the first image features corresponding to each face sample image in this face category and the first type of central feature corresponding to this face category can be calculated.
[0164] After obtaining the first type of central feature, the class feature range corresponding to this face category can be calculated according to the first type of central feature. The following will be described in combination with Figure 5 the embodiments.
[0165] S421. Determine the first-class central feature corresponding to the target category according to the first image features of each target sample image in the target category.
[0166] S422. Determine the similarity between each first image feature in the target category and the first-class central feature, and determine the class feature range corresponding to the target category according to the maximum value of the similarity.
[0167] Specifically, the first-class central feature corresponding to each face category can be obtained through the foregoing formula (2), and the first-class central feature represents the average feature of the corresponding face category.
[0168] Taking any face category as an example, this face category includes a total of N face sample images and corresponding first image features, and this face category also includes a corresponding class central feature. In the embodiment of the present disclosure, the similarity between the first image feature of each face sample image and the class central feature can be calculated. This similarity represents the similarity degree between each face sample image and the average feature. In one example, the cosine similarity between each first image feature and the class central feature can be calculated, which is expressed as:
[0169]
[0170] In formula (3), represents the similarity between the k-th first image feature of the i-th class extracted by using the first face recognition network M s and the class central feature, represents the k-th first feature data of the i-th class extracted by using the first face recognition network M s and represents the class central feature of the i-th class.
[0171] After obtaining the similarity between each first image feature and the class central feature, determine the minimum value and the maximum value of the similarity, which are expressed as:
[0172]
[0173]
[0174] In formulas (4) and (5), represents the minimum value of the similarity, represents the maximum value of the similarity, represents each similarity in the i-th class.
[0175] After determining the minimum value of the similarity and the maximum value of the similarity, the maximum value of the similarity can be used as the inner feature boundary of this face category, and the minimum value of the similarity As the outer feature boundary of the face category, the range between the inner and outer feature boundaries is the class feature range corresponding to the face category.
[0176] The above explains one of the face categories. For each face category included in the second face data, the above process is sequentially executed to obtain the class center feature and the class feature range corresponding to each face category.
[0177] S430. Determine the first feature data corresponding to the target category from the first class center feature and the class feature range.
[0178] Specifically, for any face category included in the second face data, the first class center feature of the face category obtained above and the class feature range (S min , S max ) are jointly used as the first feature data of the face category.
[0179] As can be seen from the above, after the second face data of the second client 200 is subjected to feature extraction by the first face recognition network, the first face feature corresponding to each face category is obtained, and the training of the second face recognition network can be assisted according to the first face feature. The following is combined with Figure 6 Embodiments are described.
[0180] As Figure 6 shown, in some embodiments, the process of the second client 200 training the second face recognition network includes:
[0181] S610. Input the second sample image data into the second neural network to be trained, and obtain the second feature data output by the second neural network.
[0182] S620. Based on the first difference between the second feature data and the label data of the second sample image data, and the second difference between the second feature data and the first feature data, adjust the network parameters of the second neural network until the convergence condition is met, and obtain the trained second neural network.
[0183] Specifically, as can be seen from the above, the target items for the compatibility training of the second face recognition network (i.e., the second neural network) mainly include two parts: 1) the classification loss of the second face data; 2) the loss between the features extracted by the second face recognition network and the features extracted by the first face recognition network.
[0184] Thus, in the embodiments of the present disclosure, the second face data (i.e., the second sample image data) can be input into an untrained second face recognition network to obtain the second feature data output by the second face recognition network, that is, the second face feature data. The second face feature data includes the second image features corresponding to each face sample image and the second category center features corresponding to each face category. The following will be described in conjunction with Figure 7 embodiments.
[0185] As Figure 7 shown, in some embodiments, the process of obtaining the second face feature data in the training method of the present disclosure example includes:
[0186] S611. Input each target sample image into the second neural network to be trained, and obtain the second image features corresponding to each target sample image output by the second neural network.
[0187] S612. For each target category, determine the second category center feature corresponding to the target category according to the second image features of each target sample image in the target category.
[0188] S613. Determine the second category center feature and the second image features corresponding to each target sample image as the second feature data corresponding to the target category.
[0189] Generally speaking, the process of calculating the second image features of each face sample image and the second category center features of each face category is similar to the process of calculating the first image features and the first category center features described above. The main difference is that the foregoing first image features and first category center features are obtained based on the first face recognition network, while the second image features and second category center features in this embodiment are obtained based on the second face recognition network.
[0190] Specifically, input each face sample image of the second face data into an untrained second face recognition network. In the embodiments of the present disclosure, any suitable face recognition network can be used for the second face recognition network. The network structure of the second face recognition network can be the same as or different from that of the first face recognition network. The present disclosure does not limit this. For example, in one example, the first face recognition network can also adopt the FaceNet network structure, which is expressed as:
[0191]
[0192] In formula (6), x A represents the face sample image in the second face data, M A represents the second face recognition network, represents the features of the face sample image extracted by the second face recognition network, that is, the second image features described in the present disclosure.
[0193] Taking a sample data of the second face data as an example, the face sample image x included in the sample data A is input into the untrained second face recognition network M A so that the second face recognition network extracts features from the face sample image x based on Equation (6) A to obtain the second image feature corresponding to the face sample image
[0194] Performing the above processing on each sample data of the second face data in sequence, the second image feature corresponding to each face sample image can be obtained.
[0195] For any face category, the second type of central feature corresponding to the face category can be determined according to the respective second image features included in the face category. For example, a certain face category includes N face sample images. For this face category, the calculated second type of central feature is expressed as:
[0196]
[0197] In Equation (7), represents the second image feature corresponding to the k-th face sample image of the i-th category in the second face data extracted by using the second face recognition network M A and represents the second type of central feature.
[0198] For any face category, the second image feature corresponding to each face sample image in the face category and the second type of central feature corresponding to the face category can be calculated in sequence through the above process. The second image feature and the second type of central feature are jointly used as the second feature data of the face category, that is, the second face feature data.
[0199] After obtaining the first face feature data and the second face feature data, the second face recognition network can be supervised and trained. The following is described in combination with Figure 8 embodiments.
[0200] As Figure 8 shown, in some embodiments, the process of adjusting the network parameters of the second face recognition network in the training method of the present disclosure example includes:
[0201] S621. Determine a first difference according to the second image feature and the label data corresponding to each target sample image.
[0202] Specifically, taking a sample data in the second face data as an example, the second image feature corresponding to the face sample image of the sample data can be obtained through the foregoing process. The second face recognition network can predict a classification result, that is, an output result, according to the second image feature.
[0203] It can be understood that the output result represents the predicted value of the second face recognition network, while the label data represents the true value of the face sample image. Thus, the loss between the output result and the label data can be calculated through a pre-constructed loss function, that is, the first difference described in the present disclosure.
[0204] S622. For each target category, determine a second difference according to the difference between the second type of central feature corresponding to the target category and the first type of central feature, and the difference between the second image feature corresponding to each target sample image and the class feature range.
[0205] In the embodiment of the present disclosure, the second difference includes two parts: one is the difference between the second type of central feature and the first type of central feature; the other is a loss term for constraining the second image feature based on the class feature range.
[0206] Specifically, taking any face category in the second face data as an example, through the foregoing Figure 5 embodiment, the first type of central feature corresponding to the face category can be calculated And through the foregoing Figure 7 embodiment, the second type of central feature corresponding to the face category can be calculated Thus, according to the first type of central feature and the second type of central feature the difference between the two can be calculated.
[0207] Meanwhile, for each face category, through the foregoing Figure 5 embodiment, the corresponding class feature range (S min , S max ) can also be obtained. In the embodiment of the present disclosure, the difference between the second image feature of each face sample image and the class feature range is constrained at the same time. For example, in an example, for any sample data, the second image feature corresponding to the face sample image of the sample data and the first type of central feature of the face category to which the face sample image belongs can be calculated for the cosine similarity, and then the cosine similarity is constrained to be less than Thus, the difference between the first type of central feature and the second type of central feature, and the difference between the second image feature and the class feature range together serve as the second difference.
[0208] S623. Based on the first difference and the second difference, adjust the network parameters of the second neural network until the convergence condition is met, and obtain the trained second face recognition network.
[0209] Specifically, combining the above first difference and second difference, optimize and adjust the network parameters of the second face recognition network according to the backpropagation of this difference. For multiple sample data in the second face data, repeat the above process to continuously iterate and optimize the second face recognition network until the convergence condition is met, thereby obtaining the trained second face recognition network.
[0210] It can be understood that in the above process of training the second face recognition network, the first difference represents the constraint on the recognition ability of the second face recognition network for the second face data, and the second difference represents the constraint on the compatibility of the second face recognition network for the first face data. Therefore, based on the above training process, the obtained second face recognition network can have good compatibility and recognition ability for both the first face data and the second face data, improving the face recognition accuracy.
[0211] Figure 9 Shows the schematic diagram of the compatibility training of the second face recognition network M in the training method of the present disclosure. The following will be further described in combination with A For further illustration. Figure 9
[0212] As Figure 9 shown, this training process is carried out on the side of the second client 200. Therefore, based on the foregoing Figure 4 and Figure 5 processes, the first face feature data can be obtained by using the second face data and the first face recognition network M s At the same time, based on the foregoing Figure 6 and Figure 7 processes, the second face feature data can be obtained by using the second face data and the second face recognition network M to be trained A Then, according to the foregoing Figure 8 process, based on the first face feature data, the second face feature data, and the label data in the second face data, the loss including the first difference and the second difference can be calculated by using the pre-constructed loss function. Then, the network parameters of the second face recognition network M A are adjusted by backpropagation according to this loss until the network converges, and the trained second face recognition network M A is obtained.
[0213] As described above, in the embodiments of the present disclosure, the first face recognition network is used to assist in the compatibility training of the second face recognition network, so that the second face recognition network can have good compatibility recognition capabilities for both the first face data and the second face data. At the same time, the network can converge better, improving the face recognition accuracy.
[0214] It can be understood that the trained second face recognition network can be obtained through the above process. In the embodiments of the present disclosure, the second face recognition network is not directly used as the final target face recognition network. Instead, the first face recognition network and the second face recognition network are fused to obtain the target face recognition network, and the second face data is used again to train the target face recognition network to improve the network's compatibility with each isolated island data. The following is an explanation in combination with Figure 10 embodiments.
[0215] As Figure 10 shown, in some embodiments, the training method of the present disclosure example for training the target face recognition network includes:
[0216] S1010: Input the second sample image data into the trained second neural network to obtain the third feature data output by the second neural network.
[0217] S1020: According to the fusion weight determined based on the second sample image data, perform a fusion process on the first feature data and the third feature data to obtain the fusion feature data.
[0218] S1030: Based on the third difference between the fusion feature data and the label data included in the second sample image data, adjust the network parameters of the target neural network until the convergence condition is met.
[0219] In some embodiments, a fusion layer can be added after the feature extraction layers of the first face recognition network and the second face recognition network to perform a fusion process on the extracted features of both.
[0220] Specifically, input the second face data into the trained second face recognition network. Based on the same process as extracting the second face feature data above, the second face recognition network can output the third feature data, that is, the so-called third face feature data. It can be understood that the third face feature data is essentially the same as the second face feature data. Those skilled in the art can understand and fully implement it by referring to the foregoing embodiments, and will not be elaborated herein. At the same time, input the second face data into the first face recognition network. Based on the foregoing embodiment process, the first face recognition network can output the first face feature data.
[0221] When performing fusion processing on the first face feature data and the third face feature data, it is first necessary to determine the fusion weights of the two, and then perform fusion processing on the two data based on the fusion weights.
[0222] In some embodiments, the fusion weights can be determined according to the attribute information of the second face data. The attribute information represents the differences in attributes between the first face data and the second face data. For example, in one example, the first face data mainly includes face data of "children", while the second face data mainly includes face data of "adults", then "age" is the attribute difference between the two isolated data. For another example, in one example, the first face data mainly includes face data of "men", while the second face data mainly includes face data of "women", then "gender" is the attribute difference between the two isolated data.
[0223] Of course, those skilled in the art can understand that the attribute information is not limited to the above examples, and can also be any other attribute information suitable for implementation, as long as it can make the first face data and the second face data have certain differences as a whole. The present disclosure does not limit this.
[0224] In some examples, the target face recognition network may further include an attribute recognition network, so as to extract the attribute information of the face sample image by using the attribute recognition network, so that the target face recognition network can determine the corresponding fusion weights according to the attribute information. This will be described in the following embodiments of the present disclosure and will not be elaborated here for the time being.
[0225] After determining the fusion weights of the first face feature data and the third face feature data, the two can be fused according to the fusion weights to obtain fusion feature data. It can be understood that the fusion feature data simultaneously fuses the feature information of the first face recognition network and the second face recognition network, and thus has a certain representativeness for the features of the two isolated data.
[0226] The target face recognition network can add a classification layer after the fusion layer. The classification layer is, for example, a fully connected layer. The fully connected layer predicts and outputs the output result corresponding to the face sample data according to the input fusion feature data.
[0227] It can be understood that the output result represents the predicted value of the target face recognition network for the face sample image, while the label data represents the true value of the face sample image. Therefore, the third difference, that is, the loss, between the output result and the label data is calculated through a pre-constructed loss function, and the classification layer parameters of the target face recognition network are optimized and adjusted according to the backpropagation of the third difference. For multiple sample data in the second face data, the above process is repeated to continuously iterate and optimize the target face recognition network until the convergence condition is met, so as to obtain the trained target face recognition network.
[0228] As described above, in the embodiments of the present disclosure, the target face recognition network is obtained by fusing the first face recognition network and the second face recognition network, thereby improving the compatibility of the target face recognition network with multi-party isolated data and improving the network accuracy and face recognition accuracy.
[0229] As Figure 11 shown, in some embodiments, the process of fusing the first face feature data and the third face feature data in the training method of the present disclosure example includes:
[0230] S1021. Input the second sample image data into the attribute network of the target neural network to obtain the attribute information output by the attribute network.
[0231] S1022. Determine the first weight of the first feature data and the second weight of the third feature data according to the attribute information.
[0232] S1023. Based on the first weight and the second weight, perform fusion processing on the first feature data and the third feature data to obtain fused feature data.
[0233] Figure 12 shows the schematic diagram of training the target face recognition network in the training method of the present disclosure. The following will be specifically described in conjunction with Figure 12 this.
[0234] As Figure 12 shown, in some embodiments, the target face recognition network includes an attribute network M attr , and the attribute network M attr represents a network for recognizing the attribute information of the second face data. The attribute network M attr can be pre-trained based on the type of attribute information.
[0235] As can be seen from the foregoing, the purpose of introducing the attribute information into the target face recognition network is to better fuse the first face feature data and the second face feature data. Therefore, the type of attribute information can be the attribute information mainly for differentiating the first face data and the second face data.
[0236] For example, in an example, the first face data is mainly the face data of the user without any wear on the face, while the second face data is mainly the face data of the user wearing a mask on the face. Thus, the attribute network M attr can be an attribute network pre-trained for whether the user wears a mask on the face, mainly for extracting the features of the user's facial wear and predicting the corresponding attribute information.
[0237] See Figure 12 shown. In this embodiment, the second face data is input into the attribute network M attrThe attribute information of the second face data can be obtained. The second face data is input into the first face recognition network M s to obtain the first face feature data. The second face data is input into the second face recognition network M A to obtain the second face feature data.
[0238] In some embodiments, when determining the first weight of the first face feature data and the second weight of the second face feature data based on the attribute information, the attribute information can be further smoothed based on a smoothing coefficient.
[0239] For example, in one example, a fully connected layer branch can be added to the second face recognition network M A so that the corresponding smoothing coefficient T can be output according to the second face data, and the attribute information is smoothed using the smoothing coefficient T, expressed as:
[0240]
[0241] In Equation (8), Attr s represents the smoothed attribute information, Attr represents the attribute information output by the attribute network, and T represents the smoothing coefficient output by the second face recognition network.
[0242] After obtaining the smoothed attribute information according to Equation (8), the first weight of the first face feature data and the second weight of the third face feature data can be determined according to the attribute information. Then, based on the first weight and the second weight, linear weighted fusion processing is performed on the first face feature data and the third face feature data to obtain the fusion feature data.
[0243] After obtaining the fusion feature data, the classification layer predicts the corresponding output result according to the fusion feature data, and then, based on the difference between the output result and the label data, backpropagation is used to optimize the parameters of the classification layer of the target face recognition network until the convergence condition is met, and the training of the target face recognition network is completed.
[0244] For the second client 200, after obtaining the trained target face recognition network, the face recognition network can be sent to the first client 100 through the network 300. It can be understood that in the embodiments of the present disclosure, only the first client 100 needs to send the first face recognition network to the second client 200 once, and the second client 200 needs to send the target face recognition network to the first client 100 once. In addition, no other data communication is required, and there is no intercommunication of the isolated island data itself, which protects the security of the isolated island data, and the network architecture is simple, easy to deploy, and has low cost.
[0245] As described above, in the embodiments of the present disclosure, while ensuring the security of island data, the simultaneous utilization of multi-party island data is achieved. The trained target neural network has better compatibility, improves the network's recognition ability for multi-party data, and greatly improves the target recognition accuracy. Moreover, by fusing attribute information to identify and predict target attributes, the compatibility of the network with island data of different attribute information is further enhanced, and the target recognition accuracy is improved.
[0246] Embodiments of the present disclosure provide a target recognition method, which can be applied to an electronic device. The electronic device in the embodiments of the present disclosure can be any suitable device type, such as a mobile terminal, a wearable device, a vehicle-mounted device, a server, a cloud platform, etc., and the present disclosure does not limit this.
[0247] As Figure 13 shown, in some embodiments, the target recognition method exemplified in the present disclosure includes:
[0248] S1310. Obtain a to-be-tested image.
[0249] S1320. Input the to-be-tested image into a pre-trained target recognition network to obtain a recognition result output by the target recognition network.
[0250] Specifically, the target recognition network described in the embodiments of the present disclosure is a target neural network trained according to the training method of any of the foregoing embodiments.
[0251] Taking the face recognition scenario as an example, in face recognition, the to-be-tested image is an image expected to recognize the target face in the image. That is, the to-be-tested image including the to-be-tested face can be input into the target recognition network described in the present disclosure, and thus the recognition result output by the target recognition network can be obtained.
[0252] Of course, it can be understood that the target recognition in the embodiments of the present disclosure is not limited to the face recognition scenario, and can also be any other suitable scenario, such as vehicle recognition, natural scene recognition, etc., and the present disclosure will not elaborate on this.
[0253] As described above, in the embodiments of the present disclosure, since the target recognition network has good compatible recognition ability for multi-party island data, the recognition accuracy of the to-be-tested image is higher, meeting the requirements of high-precision target recognition scenarios.
[0254] Embodiments of the present disclosure provide a training device for a neural network, which can be applied to a second client to enable the first sample image data and the second sample image data to participate in network training simultaneously.
[0255] As Figure 14 shown, in some embodiments, the training device for the neural network exemplified in the present disclosure includes:
[0256] A network acquisition module 10, configured to acquire a first neural network; the first neural network is pre-trained by a first client using first sample image data, and the first sample image data is first isolated island data that the first client can acquire;
[0257] A first processing module 20, configured to input second sample image data into the first neural network to obtain first feature data output by the first neural network; the second sample image data is second isolated island data that the second client can acquire;
[0258] A first training module 30, configured to train a second neural network to be trained according to the second sample image data and the first feature data to obtain a trained second neural network;
[0259] A second training module 40, configured to use the second sample image data to train a target neural network obtained by fusing the trained second neural network and the first neural network until a convergence condition is met.
[0260] As can be seen from the above, in the embodiments of the present disclosure, while ensuring the security of isolated island data, the simultaneous utilization of multi-party isolated island data is achieved. The trained target neural network has better compatibility, improves the network's recognition ability for multi-party data, and greatly improves the target recognition accuracy.
[0261] In some embodiments, the second sample image data includes multiple target categories, where each target category includes at least one target sample image; the first processing module 20 is configured to:
[0262] Input each target sample image into the first neural network to obtain first image features corresponding to each target sample image output by the first neural network;
[0263] For each target category, determine a first category center feature and a category feature range corresponding to the target category according to the first image features of each target sample image in the target category;
[0264] Determine the first category center feature and the category feature range as the first feature data corresponding to the target category.
[0265] In some embodiments, the first processing module 20 is configured to:
[0266] Determine the first category center feature corresponding to the target category according to the first image features of each target sample image in the target category;
[0267] Determine the similarity between each first image feature in the target category and the first category center feature, and determine the category feature range corresponding to the target category according to the maximum and minimum values of the similarity.
[0268] In some embodiments, the first training module 30 is configured to:
[0269] Input the second sample image data into the second neural network to be trained, and obtain the second feature data output by the second neural network;
[0270] Based on the first difference between the second feature data and the label data included in the second sample image data, and the second difference between the second feature data and the first feature data, adjust the network parameters of the second neural network until the convergence condition is met, and obtain the trained second neural network.
[0271] In some embodiments, the second sample image data includes multiple target categories, where each target category includes at least one target sample image; the first training module 30 is configured to:
[0272] Input each target sample image into the second neural network to be trained, and obtain the second image feature corresponding to each target sample image output by the second neural network;
[0273] For each of the target categories, determine the second category center feature corresponding to the target category according to the second image features of each target sample image in the target category;
[0274] Determine the second category center feature and the second image feature corresponding to each target sample image as the second feature data corresponding to the target category.
[0275] In some embodiments, the first feature data of each target category includes the first category center feature and the category feature range of the target category to which each target sample image belongs; the first training module 30 is configured to:
[0276] Determine the first difference according to the second image feature corresponding to each target sample image and the label data;
[0277] For each of the target categories, determine the second difference according to the difference between the second category center feature corresponding to the target category and the first category center feature, and the difference between the second image feature corresponding to each target sample image and the category feature range;
[0278] Based on the first difference and the second difference, adjust the network parameters of the second neural network.
[0279] In some embodiments, the second training module 40 is configured to:
[0280] Input the second sample image data into the trained second neural network to obtain third feature data output by the second neural network;
[0281] Fuse the first feature data and the third feature data according to a fusion weight determined based on the second sample image data to obtain fused feature data;
[0282] Adjust network parameters of the target neural network based on a third difference between the fused feature data and label data included in the second sample image data until a convergence condition is met.
[0283] In some embodiments, the second training module 40 is configured to:
[0284] Input the second sample image data into an attribute network of the target neural network to obtain attribute information output by the attribute network;
[0285] Determine a first weight of the first feature data and a second weight of the third feature data according to the attribute information;
[0286] Fuse the first feature data and the third feature data based on the first weight and the second weight to obtain the fused feature data.
[0287] In some embodiments, the first neural network includes a first face recognition network, and the first sample image data is first face image data.
[0288] In some embodiments, the second neural network includes a second face recognition network, and the second sample image data is second face image data.
[0289] As can be seen from the above, in the embodiments of the present disclosure, while ensuring the security of isolated data, simultaneous utilization of isolated data from multiple parties is achieved. The trained target neural network has better compatibility, improves the network's recognition ability for data from multiple parties, and greatly improves the target recognition accuracy. Moreover, fusing attribute information to identify and predict target attributes further enhances the network's compatibility with isolated data of different attribute information and improves the target recognition accuracy.
[0290] As Figure 15 shown, in some embodiments, the present disclosure provides an example of a target recognition device, including:
[0291] An image acquisition module 50, configured to acquire a to-be-tested image, where the to-be-tested image includes a to-be-tested target;
[0292] A second processing module 60, configured to input the image to be measured into a pre-trained target recognition network, and obtain a recognition result output by the target recognition network; the target recognition network is a target neural network obtained according to the training method described in any implementation manner of the first aspect.
[0293] As can be seen from the above, in the embodiments of the present disclosure, since the target recognition network has good compatibility and recognition capabilities for multi-party island data, the recognition accuracy of the image to be measured is higher, meeting the requirements of high-precision target recognition scenarios.
[0294] In some embodiments, the present disclosure provides an electronic device, including:
[0295] A processor; and
[0296] A memory storing computer instructions for causing the processor to execute the method described in any of the above embodiments.
[0297] In some embodiments, the present disclosure provides a storage medium storing computer instructions for causing a computer to execute the method described in any of the above embodiments.
[0298] Specifically, Figure 16 FIG. shows a schematic structural diagram of an electronic device 600 suitable for implementing the method of the present disclosure. Through Figure 16 the electronic device shown, the corresponding functions of the above-mentioned processor and storage medium can be realized.
[0299] As Figure 16 shown, the electronic device 600 includes a processor 601, which can perform various appropriate actions and processes according to a program stored in the memory 602 or a program loaded from the storage section 608 into the memory 602. In the memory 602, various programs and data required for the operation of the electronic device 600 are also stored. The processor 601 and the memory 602 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0300] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as required so that a computer program read from it can be installed into the storage section 608 as required.
[0301] Specifically, according to an embodiment of the present disclosure, the above method process can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code for performing the above method. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 609, and / or installed from the removable medium 611.
[0302] The flowcharts and block diagrams in the accompanying drawings illustrate the architectures, functions, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0303] Obviously, the above embodiments are only examples for clear illustration and are not limitations on the embodiments. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the embodiments here. And the obvious changes or variations derived therefrom are still within the protection scope of the present disclosure.
Claims
1. A training method for a neural network, characterized in that, Applied to the second client, the method includes: Obtain a first neural network; the first neural network is pre-trained by the first client using first sample image data, and the first sample image data is first island data that the first client can obtain; Input second sample image data into the first neural network to obtain first feature data output by the first neural network; the second sample image data is second island data that the second client can obtain; Train a second neural network to be trained according to the second sample image data and the first feature data to obtain a trained second neural network; Use the second sample image data to train a target neural network obtained by fusing the trained second neural network and the first neural network until a convergence condition is met; The second sample image data includes multiple target categories, where each target category includes at least one target sample image; the inputting the second sample image data into the first neural network to obtain first feature data output by the first neural network includes: Input each target sample image into the first neural network to obtain first image features corresponding to each target sample image output by the first neural network; For each of the target categories, determine a first class center feature and a class feature range corresponding to the target category according to the first image features of each target sample image in the target category; Determine the first class center feature and the class feature range as the first feature data corresponding to the target category.
2. The method according to claim 1, wherein The determining the first class center feature and the class feature range corresponding to the target category according to the first image features of each target sample image in the target category includes: Determine the first class center feature corresponding to the target category according to the first image features of each target sample image in the target category; Determine the similarity between each first image feature in the target category and the first class center feature, and determine the class feature range corresponding to the target category according to the maximum and minimum values of the similarity.
3. The method according to any one of claims 1 to 2, characterized in that, Training the second neural network to be trained according to the second sample image data and the first feature data to obtain a trained second neural network includes: Input the second sample image data into the second neural network to be trained to obtain second feature data output by the second neural network; Based on a first difference between the second feature data and label data included in the second sample image data, and a second difference between the second feature data and the first feature data, adjust network parameters of the second neural network until a convergence condition is met to obtain the trained second neural network.
4. The method according to claim 3, characterized in that, The second sample image data includes multiple target categories, where each target category includes at least one target sample image; the inputting the second sample image data into the second neural network to be trained to obtain second feature data output by the second neural network includes: Input each target sample image into the second neural network to be trained, and obtain the second image features corresponding to each target sample image output by the second neural network; For each of the target categories, determine the second category center feature corresponding to the target category according to the second image features of each target sample image in the target category; Determine the second feature data corresponding to the target category by using the second category center feature and the second image features corresponding to each target sample image.
5. The method according to claim 4, wherein The first feature data of each target category includes the first category center feature and the category feature range of the target category to which each target sample image belongs; the adjusting the network parameters of the second neural network based on the first difference between the second feature data and the label data included in the second sample image data, and the second difference between the second feature data and the first feature data, includes: Determine the first difference according to the second image features corresponding to each target sample image and the label data; For each of the target categories, determine the second difference according to the difference between the second category center feature and the first category center feature corresponding to the target category, and the difference between the second image features corresponding to each target sample image and the category feature range; Adjust the network parameters of the second neural network based on the first difference and the second difference.
6. The method according to claim 1, characterized in that, The training the target neural network obtained by fusing the trained second neural network and the first neural network by using the second sample image data includes: Input the second sample image data into the trained second neural network to obtain the third feature data output by the second neural network; Perform a fusion process on the first feature data and the third feature data according to the fusion weight determined based on the second sample image data to obtain the fusion feature data; Adjust the network parameters of the target neural network based on the third difference between the fusion feature data and the label data included in the second sample image data until the convergence condition is satisfied.
7. The method according to claim 6, wherein The performing a fusion process on the first feature data and the third feature data according to the fusion weight determined based on the second sample image data to obtain the fusion feature data includes: Input the second sample image data into the attribute network of the target neural network to obtain the attribute information output by the attribute network; Determine the first weight value of the first feature data and the second weight value of the third feature data according to the attribute information; Perform a fusion process on the first feature data and the third feature data based on the first weight value and the second weight value to obtain the fusion feature data.
8. The method according to claim 1, wherein The first neural network includes a first face recognition network, and the first sample image data is first face image data; and / or, the second neural network includes a second face recognition network, and the second sample image data is second face image data.
9. A target recognition method, characterized in that, includes: Obtain a to-be-detected image, where the to-be-detected image includes a to-be-detected target; Input the image to be measured into a pre-trained target recognition network to obtain the recognition result output by the target recognition network; the target recognition network is a target neural network obtained according to the training method described in any one of claims 1 to 8.
10. A training device for a neural network, characterized in that, Applied to a second client, the device includes: A network acquisition module configured to acquire a first neural network; the first neural network is pre-trained by a first client using first sample image data, and the first sample image data is first island data that the first client can acquire; A first processing module configured to input second sample image data into the first neural network to obtain first feature data output by the first neural network; the second sample image data is second island data that the second client can acquire; A first training module configured to train a second neural network to be trained according to the second sample image data and the first feature data to obtain a trained second neural network; A second training module configured to use the second sample image data to train a target neural network obtained by fusing the trained second neural network and the first neural network until a convergence condition is met; The second sample image data includes multiple target categories, where each target category includes at least one target sample image; the first processing module is configured to: Input each target sample image into the first neural network to obtain the first image feature corresponding to each target sample image output by the first neural network; For each of the target categories, determine the first class center feature and the class feature range corresponding to the target category according to the first image features of each target sample image in the target category; Determine the first class center feature and the class feature range as the first feature data corresponding to the target category.
11. An object recognition device, characterized in that, Includes: An image acquisition module configured to acquire an image to be measured, where the image to be measured includes a target to be measured; A second processing module configured to input the image to be measured into a pre-trained target recognition network to obtain the recognition result output by the target recognition network; the target recognition network is a target neural network obtained according to the training method described in any one of claims 1 to 8.
12. An electronic device, characterized in that, Includes: A processor; And A memory storing computer instructions for causing the processor to execute the method according to any one of claims 1 to 9.
13. A storage medium, characterized in that, Stores computer instructions for causing a computer to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image recognition method and device, electronic equipment and storage medium
CN110598504A
Joint learning method and system, node and storage medium
CN113191479A