Doorplate identification method and device based on privacy protection split learning

By processing the intermediate representation with batch normalization and binarization layer, and combining cross entropy and distance correlation loss functions, the privacy leakage problem of door sign image data in split learning is solved, and safe training and accurate identification are achieved under the split learning framework.

CN120236271APending Publication Date: 2025-07-01ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510372796.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In split learning, there is a risk of privacy leakage during the transmission and training of the client's door sign image data, especially under the attack of black and white boxes, where the opponent may infer sensitive information through the intermediate representation.

Method used

Using a split learning method based on privacy protection, the batch normalization layer and the binarization layer process the intermediate representation, reduce information redundancy, and introduce cross-entropy loss and distance correlation loss functions, combined with the U-shaped neural network architecture to prevent data leakage.

Benefits of technology

Effectively resist black and white box attacks, prevent the leakage of door sign image data and labels, ensure client privacy and security, while maintaining the accuracy and efficiency of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236271A_ABST
    Figure CN120236271A_ABST
Patent Text Reader

Abstract

The invention discloses a doorplate identification method and device based on privacy protection split learning, and the method comprises the steps: enabling a client to binarize the output of a local model into + 1 and-1, and effectively reducing the reversibility of a client model; in addition, a distance correlation loss function is introduced, and the risk of privacy disclosure is further reduced by reducing the correlation between the intermediate representation and the input doorplate image data. In order to overcome the negative influence of binaryzation on the model performance, the problem of gradient disappearance caused by binaryzation is relieved through a straight-through estimator, and the feature expression ability of intermediate representation is improved through a batch normalization layer. According to the method, the intermediate representation of the client model under the split learning framework is protected, the problem of leakage of the doorplate image data of the client is effectively solved, and black box attack and white box attack aiming at local doorplate image data are successfully defended. Meanwhile, a U-shaped neural network architecture is adopted, so that privacy leakage of the doorplate image tag of the client is prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of privacy image recognition, and in particular, to a house number recognition method based on privacy-preserving split learning. Background Art

[0002] As a key technology in the field of artificial intelligence, deep learning has achieved remarkable results. By constructing deep neural networks, deep learning can extract features from a large amount of raw data and make intelligent decisions, greatly promoting technological progress in multiple fields such as speech recognition and intelligent driving. However, the breakthrough development of deep learning not only depends on powerful computing power and high-quality algorithms, but also relies on a large amount of user data as support. In this process, the acquisition and use of data may disclose sensitive content such as users' personal information and behavior trajectories.

[0003] In the traditional scenario of unmanned vehicle delivery, the vehicle perceives the real-time environment by carrying sensors and cameras, and recognizes information such as house numbers. These data will be transmitted to the server for further storage, analysis, and model training to continuously optimize the precise delivery ability of the unmanned vehicle. Although the above technologies can significantly improve logistics efficiency and accuracy, the subsequent privacy protection issues need special attention. Although the house number itself does not directly involve personal privacy, it is closely related to sensitive information such as residents' addresses and living locations, and may expose users' places of residence or living habits in some cases. If these data are not subject to strict privacy protection measures, they may be misused, resulting in the leakage of personal privacy. Privacy data cannot be directly handed over to a server with sufficient resources for training, and this limitation has given rise to the need for distributed learning.

[0004] As a new paradigm of distributed learning, Split Learning (SL) realizes collaborative training between resource heterogeneous devices by leaving the first few layers of the neural network to be processed by local devices and handing over subsequent deep computations to the server.

[0005] As Figure 1 shown, assuming that the function representation of the complete neural network model is f, in the two-party split learning framework, the model is divided into a client model f C and a server model f S two parts, satisfying During the training process, the client first performs forward propagation based on local house number image data until the split layer stops, and sends the output of this layer (i.e., the intermediate representation) to the server. After receiving the intermediate representation, the server continues to complete the remaining calculations. Finally, the server generates the predicted output of the model and calculates the loss function based on the true labels. During backpropagation, the server first updates its own model parameters, and then returns the gradient of the loss function with respect to the intermediate representation to the client. The client completes the update of the remaining model parameters according to the received gradient information until the entire model converges. In this way, split learning can achieve distributed model training across devices or organizations without directly sharing the original house number image data, only through the exchange of intermediate representations and gradient information. Although in split learning the client only needs to transmit intermediate representation information and does not directly exchange the original data, research shows that there is still a certain risk of privacy leakage in this process. An adversary may infer the client's sensitive house number image data and its related attributes by analyzing these intermediate representations. Therefore, although split learning has certain privacy protection advantages, it still needs to be further optimized in terms of model design and security mechanisms. Summary of the Invention

[0006] The object of the present invention is to propose a house number recognition method and device based on privacy-preserving split learning in view of the deficiencies of the prior art.

[0007] The object of the present invention is achieved by the following technical solutions: A house number recognition method based on privacy-preserving split learning, comprising the following steps:

[0008] S1. Input the house number image data sample into the split learning house number recognition network for forward propagation; the split learning house number recognition network includes a client network and a server network. The input data enters the batch normalization layer and the binarization layer after passing through the client network, and generates a discrete intermediate representation and sends it to the server;

[0009] S2. The server continues to perform forward propagation with the intermediate representation and returns the prediction output by the server network to the client;

[0010] S3. The client calculates a loss function including cross-entropy loss and distance correlation loss through the prediction of the server network;

[0011] S4. Directly send the loss function to the server network for backpropagation. When propagating to the intermediate layer, return the gradient of the loss function with respect to the intermediate representation to the client;

[0012] S5. The client network continues backpropagation through the returned gradient and updates the network parameters;

[0013] S6. Use the trained split learning house number recognition network for specific house number recognition.

[0014] Further, when the house number image data samples are input into the split learning house number recognition network, the data set is divided into several small batches.

[0015] Further, the calculation process of the batch normalization layer specifically includes:

[0016] When normalizing this batch, each input v i ∈V will be transformed into where ∈ is a constant added to the denominator to ensure numerical stability; V = [v1, v2,..., v m T is the intermediate representation output by the client network;

[0017] is the batch mean, is the batch variance.

[0018] Further, the processing method of the binarization layer is specifically:

[0019]

[0020] where v is the input vector, h is the corresponding output vector, and the function is calculated element-wise for the vector.

[0021] Further, the loss function including the cross-entropy loss and the distance correlation loss is specifically:

[0022] where CE(·) is the cross-entropy loss between the predicted output of the server and the true label value Y; DCOR(·) is the distance correlation loss between the input house number image data samples and the client intermediate representation.

[0023] Further, the calculation method of the cross-entropy loss is:

[0024]

[0025] where C is the number of classifications; m is the number of batch house number image samples; Y ic is the sign function, which takes 1 if the true category of sample i in this batch is c, otherwise takes 0; p ic is the predicted probability that sample i belongs to category c.

[0026] Further, the specific calculation steps of the distance correlation loss are as follows:

[0027] ​First, calculate the distance matrices of the batch of house number image samples X and the discrete intermediate representation H respectively:

[0028] A ij = |x i - x j ||, B ij = ||h i - h j ||

[0029] where ||·|| represents the Euclidean distance; X = [x1, x2,... x m T 、H = [h1, h2,... h m T

[0030] Secondly, perform centering processing on the distance matrices to eliminate the offset effect:

[0031]

[0032] Then, according to the centered distance matrices, calculate the distance covariance and distance variance of X and H:

[0033]

[0034] Finally, calculate the distance correlation loss between X and H:

[0035]

[0036] where dCov(X, H) refers to the distance covariance between X and H, and dVar(X) and dVar(H) refer to the distance variances of X and H respectively.

[0037] Furthermore, in S4, the backpropagation method of the intermediate layer is specifically as follows: Given a vector v and the corresponding binarized vector h, the gradient of the loss function with respect to v is approximately calculated in the following way:

[0038]

[0039] where is the loss function, is the derivative of the loss function with respect to the input vector of the binarization layer.

[0040] On the other hand, the specification of the present invention also provides a house number recognition device based on privacy-preserving split learning, including a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, the above-mentioned house number recognition method based on privacy-preserving split learning is implemented.

[0041] ​​On the other hand, the present invention specification also provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the door plate recognition method based on privacy protection split learning is implemented.

[0042] Beneficial effects of the present invention:

[0043] 1. The method of the present invention can ensure that in the house number image recognition system based on split learning, the intermediate representation output by the client model will not leak too much additional information, thereby resisting black box attacks and white box attacks and preventing local house number image data from being leaked to the server.

[0044] 2. The U-shaped neural network architecture adopted by the present invention ensures that the door plate image data label of the client does not leave the local area, preventing the privacy leakage of the label. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 Schematic diagram of split learning algorithm;

[0046] Figure 2 This is a schematic diagram of a black box attack;

[0047] Figure 3 This is a schematic diagram of a white box attack;

[0048] Figure 4 A framework diagram of the doorplate recognition method based on privacy-preserving split learning provided by the present invention;

[0049] Figure 5 A schematic diagram of a neural network model provided by an embodiment of the present invention;

[0050] Figure 6 A visual comparison diagram of the defense strategies provided by an embodiment of the present invention;

[0051] Figure 7 Schematic diagram of a house number recognition device based on privacy-preserving split learning provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.

[0053] In the split learning framework, when each client sends the output of the local model (i.e., the intermediate representation) to the server, due to the high redundancy of the intermediate representation information in floating-point form, it is easy for the server to reverse-engineer the input house number image data, resulting in privacy issues. Therefore, the present invention introduces a privacy protection method for house number recognition, which is also called Binarized Intermediate Representations-based Split Learning, BIRSL. The purpose is to enable resource-constrained clients to cooperate with the server in training while ensuring the privacy of house number image data. The specific content of the embodiments of the present invention is as follows:

[0054] As Figure 4 shown, a house number recognition method based on privacy-preserving split learning mainly consists of the following parts:

[0055] S1. Forward propagation stage of the client model: The client inputs the local house number image data into the backbone network, and then sequentially processes it through the batch normalization layer and the binarization layer, so that the output intermediate representation only contains +1 and -1, thereby reducing the amount of information that the server can obtain and reducing the risk of privacy attacks.

[0056] The processing of the client batch normalization layer is as follows:

[0057] During the training process, the house number image dataset is divided into multiple small batches, and each small batch only contains a part of the samples. The normalization process is based on the mean and variance of this small batch, rather than the statistical information of the global dataset. Although this approach reduces the accuracy of global statistics, it can significantly accelerate the training process and improve the stability of training. In addition, as the training progresses, the mean and variance of the small batches gradually approach the statistical characteristics of the entire dataset, thus ensuring good training results.

[0058] Specifically, given a small batch of house number image data of size m After forward propagation by the client, the intermediate representation V = [v1, v2,..., v m T is obtained, where the batch mean The batch variance When normalizing this batch, each input v i ∈V will be converted to where ∈ is a small constant added to the denominator to ensure numerical stability. If ∈ is not considered, then with zero mean and unit variance can be obtained. In this way, ​After binarization, it can achieve the balance of positive and negative values, effectively avoid certain bits from always maintaining fixed values, prevent the loss of information in this dimension, and thus improve the model's ability to express complex data structures.

[0059] The processing of the client binarization layer is as follows:

[0060] In black-box attacks and white-box attacks, an adversary can use the intermediate representation during the forward propagation of the model to reverse-engineer the client's house number image data. To reduce this reversibility, the present invention introduces a binarization layer and uses the Sign function as the activation function. Specifically, the Sign function can be expressed as:

[0061]

[0062] where v is the input vector, h is the corresponding output vector, and the function calculates element-wise for the vector. The Sign function converts the continuous input into discrete values, namely +1 or -1, retaining only the sign information and discarding the magnitude information. Since the output information is extremely limited, the adversary cannot directly recover the fine-grained features of the input house number image data from the result, significantly reducing the inferable information volume, thereby enhancing privacy protection. In addition, this binarization processing also makes it difficult for subtle changes in the input to be manifested in the output, further protecting the details of the input house number image data.

[0063] S2. Server model forward propagation stage: The server continues the forward propagation based on the received binarized intermediate representation and sends the model prediction output to the client, forming a U-shaped network architecture.

[0064] S3. Client loss function calculation: The client calculates the loss function based on the input house number image data, the intermediate representation, and the server's prediction value, which includes two parts: cross-entropy loss and distance correlation loss.

[0065] The client loss function calculation is as follows:

[0066] Loss function Includes two parts: cross-entropy loss and distance correlation loss, specifically:

[0067]

[0068] where CE(·) is the server's prediction output The cross-entropy loss between the predicted value and the true label value Y; DCOR(·) is the distance correlation loss between the input house number image data sample and the client intermediate representation. Intuitively, the cross-entropy loss improves the classification accuracy by minimizing the difference between the prediction result and the true label; while the distance correlation loss increases the difficulty of reconstructing the house number image data from the intermediate representation by reducing the correlation between the intermediate representation and the input house number image data. Therefore, the combination of the two realizes the collaborative optimization of the classification main task and the privacy protection task. The following is the calculation process of the two losses:

[0069] Cross-entropy loss: Cross-entropy is a metric for measuring the difference between two probability distributions. By introducing the logarithmic function, cross-entropy can effectively avoid the problem of vanishing gradients that may occur at extremely small probability values. Given a multi-class classification problem, the true label of the sample is Y, and the predicted output of the model is Then the cross-entropy loss is defined as:

[0070]

[0071] where C is the number of classes; m is the number of batches of house number image samples; Y ic is the sign function, which takes 1 if the true class of sample i in this batch is c, otherwise takes 0; p ic is the predicted probability that sample i belongs to class c.

[0072] Distance correlation loss: Distance correlation is a statistic used to characterize the dependence or association degree between two random variables. Different from the Pearson correlation coefficient, distance correlation not only considers linear relationships but also can identify non-linear relationships. The value range of distance correlation is between 0 and 1. The closer the value is to 1, the stronger the correlation between the two variables; the closer the value is to 0, the less any association between the two variables. Given a batch of house number image samples Intermediate representation where X = [x1, x2,..., x m T , H = [h1, h2,..., h m T . The specific calculation steps of the distance correlation loss between X and H are as follows:

[0073] First, calculate the distance matrices of X and H respectively:

[0074] A ij = ||x i - x j ||, B ij = ||h​​i -h j ||

[0075] Among them, ||·|| represents the Euclidean distance.

[0076] Secondly, centralize the distance matrix to eliminate the offset effect:

[0077]

[0078] Then, according to the centralized distance matrix, calculate the distance covariance and distance variance between X and H:

[0079]

[0080] Finally, calculate the distance correlation loss between X and H:

[0081]

[0082] Among them, dCov(X, H) refers to the distance covariance between X and H, and dVar(X) and dVar(X) refer to the distance variances of X and H respectively.

[0083] S4. Model backpropagation of the server and the client: The client returns the loss function to the server, and the server first performs backpropagation and updates the parameters.

[0084] S5. The server sends the gradient of the loss function with respect to the intermediate representation to the client, and the client performs local backpropagation and updates the parameters.

[0085] The backpropagation calculation of the binarization layer is as follows:

[0086] Since the Sign function adopted by the binarization layer of the present invention is non-smooth and non-convex, the gradient for all non-zero inputs is zero, which causes the standard backpropagation method to be inapplicable when training a neural network. This phenomenon is called gradient vanishing. To solve this problem, traditional methods use a set of continuous activation functions to approximate the Sign function, making the binarization layer still differentiable. However, this method violates the original intention of reducing the reversibility of the neural network. Therefore, the present invention introduces a straight-through estimator to approximately estimate the gradient. Specifically, given a vector v and the corresponding vector h after binarization, the gradient of the loss function with respect to v can be approximately calculated in the following way:

[0087]

[0088] Among them, is the loss function, is the derivative of the loss function with respect to the input vector of the binarization layer. Intuitively, the straight-through estimator directly bypasses the binarization layer during backpropagation and uses the gradient of the subsequent layer as the gradient of this layer, thus solving the problem of gradient vanishing after introducing the Sign function.

[0089] The model parameters of the client and the server are updated as follows:

[0090]

[0091] where, W S and W C represent the server and client model parameters respectively, η represents the learning rate, represents the loss function.

[0092] S6. Use the trained split learning house number recognition network for specific house number recognition.

[0093] The method of the present invention solves the problem of how to train a complete model for a resource-constrained client while ensuring the privacy of house number image data. A key idea of the present invention is to binarize the forward propagation output (i.e., the intermediate representation) of the client model, making it difficult for the server to reverse the input house number image data. In addition, a distance correlation loss is introduced to further reduce the correlation between the input house number image data and the intermediate representation. Finally, the neural network model is set to a U-shaped architecture, so that the label of the client's house number image data can stay local without leaving, preventing label leakage. To solve the problem of accuracy loss caused by binarization, the present invention introduces a batch normalization layer to maximize the feature expression ability of the intermediate representation.

[0094] 1. Embodiment settings:

[0095] Operating environment: CPU: 2 Intel(R) Xeon(R) Gold 6326 CPUs @ 2.90 GHz; GPU: 1 NVIDIA GeForce RTX 3090 with 24G video memory; Operating system: Ubuntu 18.04.6 LTS; PyTorch version: 1.9.1; Python version: 3.8.13.

[0096] Neural network model: In this embodiment, two convolutional neural network models are adopted: LeNet-5 and VGG-11, and the complete model structures are as Figure 5 shown. Among them, LeNet-5 is composed of a convolutional layer (Conv) and a fully connected layer (FC), and VGG-11 is stacked by modules (Blocks) composed of convolutional layers and pooling layers. In this embodiment, two model splitting methods are designed for different models respectively, as shown in Table 1.

[0097] Table 1 Model splitting mode settings

[0098]

[0099]

[0100] Dataset: In this embodiment, two types of image datasets are used, including the handwritten digit dataset (Modified National Institute of Standards and Technology, MNIST) and the street view house number dataset (Street View House Number, SVHN). The specific composition is shown in Table 2. Among them, MNIST is trained using LeNet-5, and SVHN is trained using VGG-11.

[0101] Table 2 Dataset Information

[0102] Dataset Number of classes Number of training images Number of test images Image size MNIST 10 60000 10000 SVHN 10 73257 26032

[0103] Training hyperparameter settings: The number of training epochs is 30, and the training batch size is 50. For MNIST, the optimizer is Adam, and the learning rate is 0.001; for SVHN, the optimizer is SGD, the learning rate is 0.005, the momentum is 0.9, and the weight decay coefficient is 0.0001. The scaling factor λ of the distance correlation loss is 0.1.

[0104] Privacy attack hyperparameter settings: In the black-box attack, the adversary is assumed to have 10,000 samples in the letters subset of the EMNIST dataset; the inverse network contains 4 deconvolution layers, the optimizer is Adam, the learning rate is 0.0001, the training batch size is 128, and the number of training epochs is 100. In the white-box attack, the optimizer is Adam, the learning rate is 0.001, the batch size is 1, the TV loss scaling factor γ is 0.5, and the number of training epochs is 10,000.

[0105] Specifically, the current privacy attack methods in split learning include black-box attacks and white-box attacks:

[0106] Black-box attack: In the black-box assumption, the adversary cannot obtain the structure or parameters of the client model, can only access the intermediate representation of the split layer, and knows the distribution information of the house number image data. The core goal of the black-box attack is to train an inverse network This network can be regarded as the inverse function of the client model function representation f C When launching an attack, the adversary can input the intermediate representation obtained during training into this inverse network and directly infer the private house number image data. As Figure 2 shown, the process of the black-box attack can be divided into three stages: (1) The adversary uses N known dataset samples Initiate a query on the client model to obtain the intermediate representation (2) Input into the randomly initialized inverse network g to generate a virtual output Then calculate the distance norm between this virtual output and the known sample By continuously updating the parameters of the inverse network, gradually reduce the distance between the virtual output and the known sample ; (3) After the inverse network training is completed, the adversary inputs the intermediate representation f C (x0) corresponding to the target house number image sample x0 into the inverse network, and the reconstruction result of the target sample can be obtained. Therefore, the black-box attack can be transformed into the following optimization problem:

[0107]

[0108] White-box attack: The white-box assumption is a more stringent assumption than the black-box assumption. The adversary can not only access the intermediate representation of the disassembled layers but also obtain the structure and parameters of the client model. The core idea of the white-box attack is to continuously update the virtual sample so that its output under the client model gradually approaches the corresponding output of the target house number image sample, thereby reducing the difference between the two. As Figure 3 shown, the process of the white-box attack can be divided into two stages: (1) Initialize the virtual sample x′, and calculate the virtual output f C (x′) of the client model with x′ as the input; (2) According to the true output f C (x0) of the target house number image sample x0 in the client model, calculate the distance norm between it and f C (x′). Use this distance as the loss function to update x′ until after the iteration is completed, a virtual sample x′* close to the target house number image sample x0 can be obtained. Therefore, the white-box attack can be transformed into the following optimization problem:

[0109]

[0110] where TV(·) represents the Total Variation, which is an additional optimization objective controlled by the parameter γ. The total variation can make the generated image x′ keep piecewise smooth within the region, avoiding local drastic changes, while allowing larger changes along the region boundary, and its degree of change is controlled by the parameter β. The specific definition is as follows:

[0111]

[0112] where x′ i,j is the pixel value of x′ at the position (i, j).

[0113] Evaluation Metrics: In this embodiment, three metrics are used to evaluate the similarity between the image reconstructed by the adversary and the real image, and based on this, the privacy protection ability of the proposed algorithm is judged, including the Structural Similarity Index Measure (SSIM), the Peak Signal-to-Noise Ratio (PSNR), and the Distance Correlation (DCOR). Among them, the value range of SSIM is between [-1, 1]. The closer the value is to 1, the more similar the two images are; the closer the value is to -1, the lower the image similarity. The value range of PSNR is greater than 0. The higher the PSNR, the better the image reconstruction quality. The value range of DCOR is between [0, 1]. The closer the value is to 1, the stronger the dependence relationship between the two variables; the closer the value is to 0, the weaker the dependence relationship between the two variables.

[0114] 2. Steps and Results of the Embodiment:

[0115] The present invention verifies the privacy protection effect and model accuracy of the present invention through numerical experiments, and the numerical experiments are completed by Python. The present invention designs experiments from the following aspects:

[0116] Embodiment 1: In this embodiment, centralized training is used as the baseline, and the difference in prediction accuracy between the BIRSL algorithm and the baseline model in two splitting modes is compared, where BIRSL-1 and BIRSL-2 respectively represent that the BIRSL algorithm is trained in splitting mode 1 and splitting mode 2. As shown in Table 3, there is no obvious decrease in accuracy for the BIRSL algorithm on each dataset, and the prediction accuracy does not exceed 0.6%.

[0117] Although binarization causes the loss of a lot of information in the floating-point representation, the neural network can still effectively learn and extract features from discrete values. From the results, this binarization process does not significantly weaken the expression ability of the neural network. During the training process, the straight-through estimator ensures the continuity of backpropagation, so that the update direction of the model parameter weights can be kept roughly the same as the desired direction. Although this approximate calculation inevitably introduces certain errors, the natural fault tolerance of the neural network can effectively make up for these deficiencies: parameter redundancy allows the network to adjust learning through other paths, and the non-linear structure enhances the robustness of the model to approximate errors.

[0118] Table 3 Comparison of Model Accuracy under Different Training Modes

[0119] Dataset Centralized training BIRSL-1 BIRSL-2 MNIST 0.9904 0.9862 0.9851 SVHN 0.9498 0.9446 0.9455

[0120] Example 2: In this example, the defense performance of the BIRSL algorithm under two privacy attacks, namely black-box attack and white-box attack, was evaluated respectively. The core lies in measuring the effectiveness of the algorithm in privacy protection by calculating the similarity between the image reconstructed by the adversary and the original image. The experiment was conducted in split mode 1, where BIRSL was compared with the split model without defense.

[0121] As shown in Tables 4 and 5, the proposed BIRSL method demonstrated strong defense capabilities in all three metrics of SSIM, PSNR, and DCOR. Especially in the black-box attack scenario, the defense effect of the proposed algorithm was significantly better than that of the other baseline algorithms, showing stronger model robustness. Figure 6 The visual defense effects of the proposed algorithm under the two attacks were further demonstrated. The first row is the original image (Ground Truth, GT), and the second to third rows are the reconstructed images under no defense and the BIRSL algorithm defense respectively. It can be observed from the figure that under the white-box attack, the BIRSL algorithm showed significant defense advantages, and the reconstructed image was almost completely blurred, making it impossible to identify any specific features. Under the black-box attack, for the more complex SVHN dataset, there were still many noise points in the reconstructed image of the BIRSL algorithm, making it difficult for the adversary to extract useful information.

[0122] Table 4 Comparison of privacy protection capabilities of different methods under white-box attack

[0123]

[0124] Table 5 Comparison of privacy protection capabilities of different methods under black-box attack

[0125]

[0126]

[0127] Corresponding to the foregoing embodiment of a house number recognition method based on privacy-preserving split learning, the present invention also provides an embodiment of a house number recognition device based on privacy-preserving split learning.

[0128] See Figure 7 , an embodiment of a house number recognition device based on privacy-preserving split learning provided by an embodiment of the present invention includes a memory and one or more processors. An executable code is stored in the memory, and when the processor executes the executable code, it is used to implement a house number recognition method based on privacy-preserving split learning in the above embodiment.

[0129] An embodiment of the doorplate recognition device based on privacy-preserving split learning provided by the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. At the hardware level, as Figure 7 shown, it is a hardware structure diagram of any device with data processing capabilities where the doorplate recognition device based on privacy-preserving split learning provided by the present invention is located. In addition to Figure 7 the processor, memory, network interface, and non-volatile memory shown, generally, according to the actual functions of any device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.

[0130] The specific implementation processes of the functions and roles of each unit in the above device are specifically described in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.

[0131] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0132] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a method for doorplate recognition based on privacy-preserving split learning in the above embodiment.

[0133] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capability. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.

[0134] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the house number recognition method based on privacy protection split learning.

[0135] Those skilled in the art will readily appreciate other embodiments of the present application after considering the description and practicing the contents disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The description and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the claims.

[0136] It should be understood that the above general description and the detailed description below are only exemplary and explanatory and cannot limit the present application. The present application is not limited to the precise structure described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is limited only by the attached claims.

Claims

1. A doorplate recognition method based on privacy-preserving split learning, characterized in that: The following steps are involved: S1. Input the house number image data sample into the split learning house number recognition network for forward propagation; the split learning house number recognition network includes a client network and a server network. After passing through the client network, the input data enters the batch normalization layer and the binarization layer to generate a discrete intermediate representation and send it to the server; S2, the server continues forward propagation with the intermediate representation and returns the prediction of the server network output to the client; S3, the client calculates a loss function including a cross entropy loss and a distance correlation loss through prediction of the server network; S4, the loss function is sent directly to the server network for back propagation. When it propagates to the intermediate layer, the gradient of the loss function with respect to the intermediate representation is returned to the client; S5, the client network continues to back-propagate through the returned gradient and updates the network parameters; S6. Use the trained split learning house number recognition network to perform specific house number recognition.

2. According to the privacy-preserving split learning-based house number recognition method of claim 1, it is characterized in that: When the house number image data samples are input into the split learning house number recognition network, the data set is divided into several small batches.

3. A doorplate recognition method based on privacy-preserving split learning according to claim 2, characterized in that: The calculation process of the batch normalization layer specifically includes: When normalizing the batch, each input v i ∈V will be converted to where ∈ is a constant added to the denominator to ensure numerical stability; V = [v1, v2, …, v m ] T The intermediate representation of the client network output; is the batch mean, is the batch variance.

4. The doorplate recognition method based on privacy-preserving split learning according to claim 1, characterized in that: The processing method of the binary layer is specifically as follows: Among them, v is the input vector, h is the corresponding output vector, and the function calculates the vector element by element.

5. The doorplate recognition method based on privacy-preserving split learning according to claim 1 is characterized in that: The loss function including the cross entropy loss and the distance correlation loss is specifically: Among them, CE(·) is the predicted output of the server and the cross entropy loss between the true label value Y; DCOR(·) is the distance correlation loss between the input doorplate image data sample and the client intermediate representation, and λ is the scaling factor of the distance correlation loss.

6. A doorplate recognition method based on privacy-preserving split learning according to claim 5, characterized in that: The cross entropy loss is calculated as: Among them, C is the number of categories; m is the number of house number image samples; Y ic is a symbolic function that takes the value 1 if the true category of sample i in the batch is c, otherwise it takes the value 0; pic is the predicted probability that sample i belongs to category c.

7. The doorplate recognition method based on privacy-preserving split learning according to claim 5 is characterized in that: The specific calculation steps of the distance correlation loss are as follows: First, calculate the distance matrix of the batch of house number image samples X and the discrete intermediate representation H respectively: A ij =||x i -x j ||,B ij =h i -h j || where ||·||| represents the Euclidean distance; X = [x1, x2, …, x m ] T , H=[h1,h2,…,h m ] T Secondly, the distance matrix is ​​centralized to eliminate the effect of offset: Then, based on the centralized distance matrix, calculate the distance covariance and distance variance of X and H: Finally, calculate the distance correlation loss between X and H: Where dCov(X, H) refers to the distance covariance between X and H, and dVar(X) and dVar(X) refer to the distance variances of X and H, respectively.

8. The doorplate recognition method based on privacy-preserving split learning according to claim 1 is characterized in that: In S4, the back propagation method of the intermediate layer is specifically as follows: given a vector v and a corresponding binarized vector h, the gradient of the loss function with respect to v is approximately calculated in the following way: in, is the loss function, is the derivative of the loss function with respect to the input vector of the binarization layer.

9. A doorplate recognition device based on privacy-preserving split learning, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, a house number recognition method based on privacy-preserving split learning as described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, a door plate recognition method based on privacy-preserving split learning as described in any one of claims 1 to 8 is implemented.