Multimedia privacy protection-oriented multi-dimensional quantization index evaluation method and device
By constructing a multi-dimensional quantitative indicator system to evaluate the privacy protection effect of multimedia data, the problem of existing technologies being unable to assess the privacy protection needs after multimedia data anonymization is solved, thus achieving scientific and objective privacy protection assessment and data security assurance.
Patent Information
- Application Number
- CN202511134162.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-13
AI Technical Summary
Existing technologies cannot effectively assess whether anonymized multimedia data meets privacy protection requirements, leading to an increased risk of privacy breaches.
By constructing a multi-dimensional quantitative index system, the texture similarity, edge similarity, and semantic similarity between the source image and the encrypted image are evaluated to determine the privacy protection index and judge whether the encrypted image meets the privacy protection requirements.
To scientifically and objectively assess the privacy protection effectiveness of multimedia data, reduce the risk of privacy leaks, and ensure data security.
Smart Images

Figure CN120705916B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information security, in particular to a multi-dimensional quantization index evaluation method and device for multimedia privacy protection. BACKGROUND
[0002] In the digital era, multimedia data (such as images or videos, etc.) is increasing, and multimedia data has become an important carrier of information exchange. However, with the increasing convenience of data sharing and dissemination, the risk of personal privacy leakage also increases. In order to protect individual privacy, various privacy protection technologies have emerged, such as data encryption, anonymous processing, image blurring, data desensitization, etc. Taking data desensitization as an example, data desensitization is a data security technology that aims to transform sensitive information contained in multimedia data through pre-set rules and algorithms to protect private data. The main purpose of data desensitization is to ensure the privacy and security of sensitive information without affecting the value of data use. By performing data desensitization on multimedia data, data security can be protected, and data leakage risk can be effectively reduced.
[0003] However, when the multimedia data is desensitized to obtain the desensitized data, it is impossible to know whether the desensitized data meets the privacy protection requirement. If the desensitized data does not meet the privacy protection requirement, personal privacy will still be leaked when the desensitized data is transmitted, and data security cannot be protected. SUMMARY
[0004] The present application provides a multi-dimensional quantization index evaluation method for multimedia privacy protection, applied to a terminal device to be subjected to multimedia privacy protection, and the method comprises:
[0005] obtaining a source image and an encrypted image; wherein the source image comprises a privacy area and a non-privacy area, and the encrypted image is obtained by performing privacy processing on the privacy area of the source image;
[0006] determining a target texture similarity based on the source image and the encrypted image, the target texture similarity being used to represent the similarity between the texture features of the source image and the texture features of the encrypted image;
[0007] determining a target edge similarity based on the source image and the encrypted image, the target edge similarity being used to represent the similarity between the edge features of the source image and the edge features of the encrypted image;
[0008] determining a target semantic similarity based on the source image and the encrypted image, the target semantic similarity being used to represent the similarity between the semantic features of the source image and the semantic features of the encrypted image;
[0009] determine a privacy protection index based on the target texture similarity, the target edge similarity, and the target semantic similarity; if the privacy protection index satisfies a preset condition, determine that the encrypted image satisfies a privacy protection requirement, otherwise, determine that the encrypted image does not satisfy the privacy protection requirement.
[0010] The application provides a multi-dimensional quantization index evaluation device for multimedia privacy protection, which is applied to a terminal device to be subjected to multimedia privacy protection, and the device comprises:
[0011] An acquisition module is configured to acquire a source image and an encrypted image; the source image comprises a privacy region and a non-privacy region, and the encrypted image is obtained by performing privacy processing on the privacy region of the source image.
[0012] A determination module is configured to determine a target texture similarity based on the source image and the encrypted image, the target texture similarity being used to represent a similarity between texture features of the source image and texture features of the encrypted image; determine a target edge similarity based on the source image and the encrypted image, the target edge similarity being used to represent a similarity between edge features of the source image and edge features of the encrypted image; determine a target semantic similarity based on the source image and the encrypted image, the target semantic similarity being used to represent a similarity between semantic features of the source image and semantic features of the encrypted image; determine a privacy protection index based on the target texture similarity, the target edge similarity, and the target semantic similarity; if the privacy protection index satisfies a preset condition, determine that the encrypted image satisfies a privacy protection requirement, otherwise, determine that the encrypted image does not satisfy the privacy protection requirement.
[0013] The application provides an electronic device, comprising a processor and a machine readable storage medium, the machine readable storage medium storing machine executable instructions capable of being executed by the processor; the processor is configured to execute the machine executable instructions to implement a multi-dimensional quantization index evaluation method for multimedia privacy protection.
[0014] The application provides a computer program product, comprising a computer program, which, when executed by a processor, implements a multi-dimensional quantization index evaluation method for multimedia privacy protection.
[0015] The application provides a machine readable storage medium, which stores machine executable instructions capable of being executed by a processor; wherein the processor is configured to execute the machine executable instructions to implement a multi-dimensional quantization index evaluation method for multimedia privacy protection.
[0016] It can be seen from the above technical solutions that in the embodiments of the present application, the privacy protection index is determined based on the target texture similarity, the target edge similarity and the target semantic similarity between the source image and the encrypted image (such as the desensitized image), and then the encrypted image is evaluated by the privacy protection index to determine whether the encrypted image meets the privacy protection requirement. If the privacy protection index meets the preset condition, it means that the encrypted image meets the privacy protection requirement, so that when the encrypted image is transmitted, the risk of personal privacy leakage is reduced, the data security is protected, and the risk of data leakage is effectively reduced. If the privacy protection index does not meet the preset condition, it means that the encrypted image does not meet the privacy protection requirement, and the encrypted image that does not meet the privacy protection requirement is not transmitted, thereby reducing the risk of personal privacy leakage and protecting data security.
[0017] In the above manner, by constructing multiple visual quantization dimensions (such as texture, edge, semantic, etc.), a multi-dimensional quantization index system is designed, and a visual multi-dimensional quantization index is given, thereby providing a scientific and objective evaluation means for the privacy protection effect in multimedia data, that is, the privacy protection effect in multimedia data can be scientifically and objectively evaluated, an objective evaluation basis is provided for optimization and improvement of privacy protection technologies such as encryption, blurring and desensitization, and the privacy protection effect of the encrypted image can be more accurately evaluated. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 is a flowchart of a multi-dimensional quantization index evaluation method for multimedia privacy protection;
[0019] Figure 2 is a flowchart of a multi-dimensional quantization index evaluation method for multimedia privacy protection;
[0020] Figure 3 is a schematic diagram of a visual multi-dimensional quantization index in an embodiment of the present application;
[0021] Figure 4 is a structural diagram of a multi-dimensional quantization index evaluation device for multimedia privacy protection;
[0022] Figure 5 is a hardware structure diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0023] In the embodiments of the present application, a multi-dimensional quantization index evaluation method for multimedia privacy protection is proposed, which can be applied to a terminal device (such as an Internet of Things device) to be subjected to multimedia privacy protection. As shown in Figure 1 The method can include the following steps:
[0024] In step 101, a source image and an encrypted image are obtained. The source image can include a privacy region and a non-privacy region, and the encrypted image is obtained by performing privacy processing on the privacy region of the source image.
[0025] Step 102, determining a target texture similarity based on the source image and the encrypted image, the target texture similarity being used to represent a similarity degree between a texture feature of the source image and a texture feature of the encrypted image.
[0026] Step 103, determining a target edge similarity based on the source image and the encrypted image, the target edge similarity being used to represent a similarity degree between an edge feature of the source image and an edge feature of the encrypted image.
[0027] Step 104, determining a target semantic similarity based on the source image and the encrypted image, the target semantic similarity being used to represent a similarity degree between a semantic feature of the source image and a semantic feature of the encrypted image.
[0028] Step 105, determining a privacy protection index based on the target texture similarity, the target edge similarity and the target semantic similarity; if the privacy protection index meets a preset condition, determining that the encrypted image meets a privacy protection requirement, and if the privacy protection index does not meet the preset condition, determining that the encrypted image does not meet the privacy protection requirement.
[0029] Illustratively, the target texture similarity is determined based on the source image and the encrypted image, including: determining a first gradient direction and a first gradient amplitude based on a non-private area of the source image, determining a second gradient direction and a second gradient amplitude based on a non-private area of the encrypted image; determining a direction similarity of the source image and the encrypted image based on the first gradient direction and the second gradient direction; determining an amplitude similarity of the source image and the encrypted image based on the first gradient amplitude, the second gradient amplitude, the non-private area of the source image and the non-private area of the encrypted image; determining a first texture similarity of the non-private area based on the direction similarity and the amplitude similarity; determining a third gradient direction based on a private area of the source image, and determining a fourth gradient direction based on a private area of the encrypted image; determining a second texture similarity of the private area based on the third gradient direction, the fourth gradient direction and a total number of pixels in the private area; and determining the target texture similarity based on the first texture similarity and the second texture similarity.
[0030] For example, the determination of the first gradient direction, the first gradient magnitude, the second gradient direction, the second gradient magnitude, the third gradient direction and the fourth gradient direction can include: obtaining a horizontal direction gradient operator and a vertical direction gradient operator, the horizontal direction gradient operator and the vertical direction gradient operator each including K gradient parameters, K being a positive integer; wherein the K gradient parameters are parameters in the target network model and are obtained by optimizing the target network model; determining the horizontal direction gradient of each pixel point in the source image based on the horizontal direction gradient operator and determining the vertical direction gradient of each pixel point in the source image based on the vertical direction gradient operator; determining the first gradient direction and the first gradient magnitude based on the horizontal direction gradient and the vertical direction gradient of each pixel point in the non-private area of the source image; determining the third gradient direction based on the horizontal direction gradient and the vertical direction gradient of each pixel point in the private area of the source image; determining the horizontal direction gradient of each pixel point in the encrypted image based on the horizontal direction gradient operator and determining the vertical direction gradient of each pixel point in the encrypted image based on the vertical direction gradient operator; determining the second gradient direction and the second gradient magnitude based on the horizontal direction gradient and the vertical direction gradient of each pixel point in the non-private area of the encrypted image; and determining the fourth gradient direction based on the horizontal direction gradient and the vertical direction gradient of each pixel point in the private area of the encrypted image.
[0031] For example, the obtaining of the K gradient parameters includes: inputting a sample image into the target network model to output a predicted sensitive probability of each first pixel point in the sample image, a predicted gradient direction of each second pixel point in the private area of the sample image, the predicted sensitive probability indicating a probability that the first pixel point belongs to a sensitive pixel; determining a private sensitive area segmentation loss value based on the predicted sensitive probability of each first pixel point in the sample image and a true sensitive probability of each first pixel point in the sample image; determining a gradient direction consistency loss value based on the predicted gradient direction of each second pixel point in the private area and a true gradient direction of each second pixel point in the private area; determining a target loss value based on the private sensitive area segmentation loss value and the gradient direction consistency loss value, optimizing the target network model based on the target loss value to obtain an optimized model; if the optimized model has not converged, updating the optimized model to the target network model and returning to the operation of inputting the sample image into the target network model; and if the optimized model has converged, obtaining the K gradient parameters from network parameters of the optimized model.
[0032] For example, determining the target edge similarity based on the source image and the encrypted image can include: obtaining a threshold set, the threshold set including a plurality of dynamic thresholds; for each dynamic threshold, determining a first edge binary image corresponding to the source image based on the dynamic threshold, and determining a second edge binary image corresponding to the encrypted image based on the dynamic threshold; a first value in the first edge binary image represents an edge point, and a second value represents a non-edge point; obtaining a common edge image; wherein for each pixel point in the common edge image, if the logical and operation result of the pixel value corresponding to the pixel point in the first edge binary image and the pixel value corresponding to the pixel point in the second edge binary image is the first value, and the absolute value of the difference between the first pixel value corresponding to the pixel point in the source image and the second pixel value corresponding to the pixel point in the encrypted image is not greater than a brightness control threshold, then based on the first pixel value, the second pixel value and the brightness control threshold, the pixel value corresponding to the pixel point in the common edge image is determined; otherwise, the pixel value corresponding to the pixel point in the common edge image is determined as the second value; determining the edge similarity corresponding to the dynamic threshold based on the common edge image and the first edge binary image, and weighting the edge similarity corresponding to each dynamic threshold to obtain the target edge similarity.
[0033] For example, obtaining the threshold set can include but is not limited to: determining the gradient amplitude of the source image based on the horizontal direction gradient and the vertical direction gradient of the source image, and performing a normalization operation on the gradient amplitude of the source image to obtain a first target gradient amplitude; generating a first gradient histogram corresponding to the source image based on the first target gradient amplitude; wherein the first gradient histogram can include the gradient amplitude distribution of each pixel point in the source image. Determine the gradient amplitude of the encrypted image based on the horizontal direction gradient and the vertical direction gradient of the encrypted image, and perform a normalization operation on the gradient amplitude of the encrypted image to obtain a second target gradient amplitude; generate a second gradient histogram corresponding to the encrypted image based on the second target gradient amplitude; wherein the second gradient histogram can include the gradient amplitude distribution of each pixel point in the encrypted image. On this basis, generate a joint gradient histogram based on the first gradient histogram and the second gradient histogram, determine a threshold lower limit value and a threshold upper limit value based on the joint gradient histogram; generate N uniformly distributed dynamic thresholds based on the threshold lower limit value and the threshold upper limit value, the threshold set including N dynamic thresholds; wherein N can be a positive integer greater than 1, N can be pre-configured, or N can be determined based on the effective peak number of the joint gradient histogram.
[0034] For example, determining the target semantic similarity based on the source image and the encrypted image can include: determining the object semantic similarity based on the overlap between the detection box of the source image and the detection box of the encrypted image; the detection box of the source image is the area of each identified object in the source image, and the detection box of the encrypted image is the area of each identified object in the encrypted image; determining the relationship semantic similarity based on the semantic association graph of the source image and the semantic association graph of the encrypted image; the semantic association graph of the source image represents the association relationship of each identified object in the source image, and the semantic association graph of the encrypted image represents the association relationship of each identified object in the encrypted image; determining the scene semantic similarity based on the scene feature vector of the source image and the scene feature vector of the encrypted image; wherein the scene feature vector of the source image is obtained by inputting the source image into a feature extraction network, and the scene feature vector of the encrypted image is obtained by inputting the encrypted image into the feature extraction network; performing weighted operation on the object semantic similarity, the relationship semantic similarity and the scene semantic similarity to obtain a semantic risk value, and determining the target semantic similarity based on the semantic risk value.
[0035] For example, determining the privacy protection index based on the target texture similarity, the target edge similarity and the target semantic similarity can include: determining a first standard deviation based on the target texture similarity, determining a first correlation coefficient based on the target texture similarity and a pre-configured real leakage risk label, and determining a first weighting coefficient corresponding to the target texture similarity based on the first standard deviation and the first correlation coefficient; determining a second standard deviation based on the target edge similarity, determining a second correlation coefficient based on the target edge similarity and the real leakage risk label, and determining a second weighting coefficient corresponding to the target edge similarity based on the second standard deviation and the second correlation coefficient; determining a third standard deviation based on the target semantic similarity, determining a third correlation coefficient based on the target semantic similarity and the real leakage risk label, and determining a third weighting coefficient corresponding to the target semantic similarity based on the third standard deviation and the third correlation coefficient; weighting the target texture similarity, the target edge similarity and the target semantic similarity based on the first weighting coefficient, the second weighting coefficient and the third weighting coefficient to obtain the privacy protection index.
[0036] As can be seen from the above technical solutions, in the embodiments of the present application, the privacy protection index is determined based on the target texture similarity, the target edge similarity and the target semantic similarity between the source image and the encrypted image (such as the desensitized image), and then the encrypted image is evaluated whether it meets the privacy protection requirement through the privacy protection index. If the privacy protection index meets the preset condition, it means that the encrypted image meets the privacy protection requirement, so that when the encrypted image is transmitted, the risk of personal privacy leakage is reduced, the data security is protected, and the risk of data leakage is effectively reduced. If the privacy protection index does not meet the preset condition, it means that the encrypted image does not meet the privacy protection requirement, and the encrypted image that does not meet the privacy protection requirement is not transmitted, thereby reducing the risk of personal privacy leakage and protecting data security.
[0037] In the above manner, by constructing multiple visual quantization dimensions (such as texture, edge, semantic, etc.), a multi-dimensional quantization index system is designed, and a visual multi-dimensional quantization index is given, so as to provide a scientific and objective evaluation means for the privacy protection effect in multimedia data, that is, to scientifically and objectively evaluate the privacy protection effect in multimedia data, to provide an objective evaluation basis for the optimization and improvement of privacy protection technologies such as encryption, blurring, desensitization, and to more accurately evaluate the privacy protection effect of encrypted images.
[0038] The above technical solutions of the embodiments of the present application will be described below in combination with specific application scenarios.
[0039] In order to protect individual privacy, various privacy protection technologies have emerged, such as data encryption, anonymous processing, image blurring, data desensitization, etc. However, when the source image is subjected to privacy protection to obtain an encrypted image, it is impossible to know whether the encrypted image meets the privacy protection requirement. If the encrypted image does not meet the privacy protection requirement, personal privacy will still be leaked when the encrypted image is transmitted, and data security cannot be protected.
[0040] In view of the above finding, a multi-dimensional quantization index evaluation method for multimedia privacy protection is proposed in the embodiments of the present application, which aims to provide a scientific and objective evaluation means for the privacy protection effect in multimedia data (such as images, videos, etc., hereinafter taking images as an example). By constructing multiple visual quantization dimensions (such as texture, edge, semantic, etc.), starting from the visual features of the image, and combining the processing effect of the privacy protection technology on the image, a multi-dimensional quantization index system is designed, which can more accurately evaluate the privacy protection effect of encrypted images, can scientifically and objectively evaluate the privacy protection effect, can perform multi-dimensional quantization analysis on the privacy protection effect, and can provide scientific basis and objective standard for the evaluation and improvement of multimedia privacy protection technology. This method can be applied to high privacy sensitive scenarios such as smart city, medical image, autonomous driving, etc.
[0041] A multi-dimensional quantization index evaluation method for multimedia privacy protection is proposed in the embodiments of the present application, as shown in Figure 2 The method can include the following steps:
[0042] Step 201, acquiring a source image, the source image being an original image that needs to be subjected to privacy protection.
[0043] Exemplarily, a video stream from a camera can be collected, each frame of image in the video stream can be taken as a source image, or part of the images in the video stream can be taken as source images. In addition, a video stream can be obtained from a storage device, each frame of image in the video stream can be taken as a source image, or part of the images in the video stream can be taken as source images. In addition, a real-time streaming can be received, each frame of image in the real-time streaming can be taken as a source image, or part of the images in the real-time streaming can be taken as source images.
[0044] Exemplarily, the source image can be an RGB image, a grayscale image, a video frame sequence, or a multispectral image, and the embodiment does not limit the type of the source image, and the source image can be of any format.
[0045] In step 202, an encrypted image corresponding to the source image is obtained; wherein the source image can include a privacy area and a non-privacy area, and the encrypted image is obtained by performing privacy processing on the privacy area of the source image.
[0046] For example, a data encryption algorithm can be used to perform privacy processing (i.e., encryption processing) on the privacy area of the source image to obtain the encrypted image. An anonymous processing algorithm can be used to perform privacy processing (i.e., anonymous processing) on the privacy area of the source image to obtain the encrypted image. An image blurring algorithm can be used to perform privacy processing (i.e., image blurring processing) on the privacy area of the source image to obtain the encrypted image. A data desensitization algorithm can be used to perform privacy processing (i.e., data desensitization processing) on the privacy area of the source image to obtain the encrypted image.
[0047] Exemplarily, a multi-modal model can be used to identify sensitive content such as faces, license plates, and text, and a data cleaning technique can be used to dynamically screen risk areas to obtain the privacy area of the source image. On this basis, the privacy area of the source image can be processed (e.g., privacy protection processing) to obtain the encrypted image.
[0048] In one possible implementation, the following method can be used to obtain the encrypted image corresponding to the source image. The following method is only an example, and the privacy area of the source image can be processed to obtain the encrypted image.
[0049] After obtaining the source image, the image integrity of the source image can be checked, for example, whether there is file header damage, pixel value overflow, etc. If there is file header damage, pixel value overflow, etc., the process is ended, and the source image is not processed subsequently. Otherwise, the source image is processed subsequently.
[0050] After obtaining the source image, a de-serialization operation can be performed on the source image, that is, if there is metadata in the source image, the metadata in the source image is removed. For example, sensitive fields in the source image such as GPS (Global Positioning System) coordinates, shooting time, device serial number, and other privacy information are removed.
[0051] After obtaining the source image, visual content recognition can be performed on the source image, i.e., a privacy region of the source image is recognized by the multi-modal model, which can include at least one detection box (such as a detection box of a recognized object in the source image, which can be a face object, a license plate object, an ID card object, etc., and no limitation is made on the recognized object, which can be any sensitive object). That is, the multi-modal model detects whether there is a recognized object in the source image and outputs a detection box (such as a rectangular detection box) of the recognized object, which is taken as a privacy region of the source image, so as to obtain the privacy region of the source image.
[0052] For example, the multi-modal model can be a detection model such as a Faster R-CNN detection model, or other models, which are not limited and can detect recognized objects in the source image. The structure and training process of the detection model are not limited in this embodiment. The source image can be input to the detection model, which detects whether there is a recognized object in the source image and outputs a detection box (such as detection box coordinates), a class, and a confidence of the recognized object. The class represents the class of the recognized object corresponding to the detection box, such as a face object, a license plate object, an ID card object, etc. The confidence represents the probability of the detection box corresponding to the recognized object.
[0053] For example, after the detection model outputs the detection box of the recognized object, NMS (non maximum suppression) can be used to eliminate overlapping boxes and retain high-confidence detection boxes.
[0054] For example, before the source image is input to the detection model, the source image can be pre-processed and the pre-processed source image can be input to the detection model. For example, for resolution standardization, the source image can be scaled to a fixed size (such as 512*512), and the source image of the fixed size can be input to the detection model.
[0055] As described above, the privacy region of the source image (i.e., the region where the sensitive object is located) can be obtained, and other regions except the privacy region can be taken as non-privacy regions of the source image. For the privacy region of the source image, privacy processing can be performed on the privacy region, such as Gaussian blurring of the privacy region, and the kernel size of the Gaussian blurring is adaptive to the size of the sensitive object. For example, GAN (Generative Adversarial Networks) can be used to replace the texture of the privacy region. For example, feature replacement can be performed on the privacy region, such as replacing a real face with a virtual face. Of course, the above are only a few examples, and no limitation is made on the privacy processing method, which can achieve privacy protection of the privacy region.
[0056] For the non-private region of the source image, the non-private region can be kept unchanged, or Gaussian noise can be added to the non-private region to destroy the global statistical characteristics of the non-private region, and no restriction is made on the non-private region.
[0057] After the above processing of the source image, the processed source image (i.e., the privacy region is subjected to privacy processing, and the non-private region is kept unchanged or Gaussian noise is added) can be synthesized (i.e., fused) with the source image, and the fused image is taken as the encrypted image corresponding to the source image. Obviously, the source image can include a privacy region and a non-private region, and the encrypted image is obtained by subjecting the privacy region of the source image to privacy processing.
[0058] Step 203, determine the target texture similarity based on the source image and the encrypted image, the target texture similarity being used to represent the similarity between the texture features of the source image and the texture features of the encrypted image. For example, the target texture similarity can be a texture distribution similarity, which can quantify the explicit privacy leakage risk.
[0059] For example, given the source image S and the encrypted image C, the similarity between the texture features of the source image S and the texture features of the encrypted image C can be determined, and then the target texture similarity is obtained, and the determination method of the target texture similarity is not limited, as long as it can represent the similarity between the texture features. In one possible implementation, the target texture similarity can be determined by the following steps:
[0060] Step S11, obtain the horizontal direction gradient operator and the vertical direction gradient operator.
[0061] In one possible implementation, the horizontal direction gradient operator can be configured according to actual needs, and the vertical direction gradient operator can be configured according to actual needs. For example, the horizontal direction gradient operator may be: , and the vertical direction gradient operator may be: . Of course, the above is only an example of the horizontal direction gradient operator and the vertical direction gradient operator, and no limitation is made thereto.
[0062] In a possible implementation, the horizontal direction gradient operator and the vertical direction gradient operator each include K gradient parameters, K is a positive integer, the K gradient parameters are parameters in the target network model, and are obtained by optimizing the target network model. Based on this, the horizontal direction gradient operator and the vertical direction gradient operator can be trainable gradient operators for detecting a privacy-sensitive region, so that the horizontal direction gradient operator and the vertical direction gradient operator automatically learn weights that are more sensitive to privacy features, the horizontal direction gradient operator and the vertical direction gradient operator change with training, can be optimized for specific tasks, dynamically learn an optimal edge detection mode, and the horizontal direction gradient operator and the vertical direction gradient operator continuously learn based on a decrease in a loss function.
[0063] For example, in order to enhance the contribution of the current pixel row and improve the anti-noise performance, the horizontal direction gradient operator and the vertical direction gradient operator can be designed as an antisymmetric structure (such as a symmetric positive-negative weight), to ensure the zero mean of the gradient operator. When the K gradient parameters are three gradient parameters, the three gradient parameters include parameters w 1 、 w 2 、 w 3 For example, the horizontal direction gradient operator can include , and the vertical direction gradient operator can include . When the K gradient parameters are four gradient parameters ( w 1 、 w 2 、 w 3 、 w 4 ), for the horizontal direction gradient operator, the first column is w 1 、 w 2 、 w 3 、 w 4 , the last column is -w 1 、 -w 2 、 -w 3 、 -w 4 , and the remaining two columns are all 0, and for the vertical direction gradient operator, the first row is w 1 、 w 2 、 w 3 、 w 4 , and the last row is-w 1 、 -w 2 、 -w 3 、 -w 4 and the rest two lines are all 0. When the K gradient parameters are of other quantities, the horizontal direction gradient operator and the vertical direction gradient operator can be designed similarly.
[0064] For example, the acquisition process of the K gradient parameters can include but is not limited to:
[0065] The target network model is acquired, which is used to output the predicted sensitive probability of the pixel point and the predicted gradient direction of the pixel point. The network structure of the target network model is not limited, as long as the target network model can output the predicted sensitive probability of the pixel point and the predicted gradient direction of the pixel point.
[0066] In the target network model, a plurality of network parameters can be included, and the K network parameters in all network parameters can be taken as the K gradient parameters (i.e. the gradient parameters to be optimized, such as w 1 、 w 2 、 w 3 As to which network parameters are taken as the gradient parameters, the embodiment is not limited, and the K network parameters can be selected as the K gradient parameters according to actual needs, so that the K gradient parameters can be the parameters in the target network model, and the K gradient parameters are obtained by optimizing the target network model.
[0067] The sample image (which can be a plurality of sample images, and here an example of one sample image) and the label data corresponding to the sample image are acquired. As to the sample image, the sample image can be any image with a private area, which is not limited. As to the label data corresponding to the sample image, the label data can include the real sensitive probability of each first pixel point (each pixel point in the sample image is referred to as a first pixel point) in the sample image, and the real sensitive probability represents the probability that the first pixel point belongs to the sensitive pixel, for example, the real sensitive probability can be 1 or 0, 1 indicates that the first pixel point belongs to the sensitive pixel, i.e. the first pixel point belongs to the private area, and 0 indicates that the first pixel point does not belong to the sensitive pixel, i.e. the first pixel point does not belong to the private area.
[0068] The label data can also include the real gradient direction of each second pixel point (each pixel point in the private area of the sample image is referred to as a second pixel point) in the private area of the sample image, and the real gradient direction is the direction in which the function value of the function increases fastest at the second pixel point, which is not limited.
[0069] The privacy region of the sample image can be a region (such as a rectangular region) in which a sensitive object is located, and the sensitive object can be a face object, a license plate object, an ID card object, etc., without any limitation.
[0070] The label data corresponding to the sample image can be manually annotated by a user or obtained by using an algorithm, and the source of the label data is not limited.
[0071] After obtaining the target network model and the sample image (such as multiple sample images), the sample image can be input into the target network model, and the target network model can output the predicted sensitive probability of each first pixel point in the sample image and the predicted gradient direction of each second pixel point in the privacy region of the sample image.
[0072] For example, since the target network model is used to output the predicted sensitive probability of a pixel point and the predicted gradient direction of the pixel point, after the sample image is input into the target network model, the target network model can process the sample image to obtain the predicted sensitive probability of each first pixel point in the sample image, and the processing process is not limited, and the predicted sensitive probability can be obtained.
[0073] After the sample image is input into the target network model, the target network model can process the sample image to obtain the predicted gradient direction of each second pixel point in the privacy region of the sample image, that is, the target network model first determines the privacy region of the sample image, and then determines the predicted gradient direction of each second pixel point in the privacy region, and the processing process is not limited, and the predicted gradient direction can be obtained.
[0074] For example, the privacy sensitive region segmentation loss value can be determined based on the predicted sensitive probability (i.e., predicted data) of each first pixel point in the sample image and the true sensitive probability (i.e., label data) of each first pixel point in the sample image, and the privacy sensitive region segmentation loss value is used to align the privacy sensitive gradient map output by the target network model with the real labeled privacy region mask. For example, a weighted cross-entropy loss function can be used to determine the privacy sensitive region segmentation loss value to solve the class imbalance problem, and other loss functions can also be used to determine the privacy sensitive region segmentation loss value, without any limitation.
[0075] For example, the privacy sensitive region segmentation loss value can be determined by using the following formula L1 :
[0076]
[0077] In the above formula, N represents the total number of pixel points of the sample image, i represents the first pixel point of the sample image. i represents the weight coefficient, which is used to emphasize the punishment of the privacy area, and can be configured according to actual needs. represents the real sensitive probability of the first pixel point of the sample image, represents the predicted sensitive probability of the first pixel point of the sample image, i The value of the real sensitive probability can be 0 or 1, and the value of the predicted sensitive probability can be between 0 and 1, that is, in the range of 0~1. represents the weight coefficient, which is used to emphasize the punishment of the privacy area, and can be configured according to actual needs. i For example, based on the predicted gradient direction (i.e. predicted data) of each second pixel point in the privacy area of the sample image and the real gradient direction (i.e. label data) of each second pixel point in the privacy area of the sample image, the gradient direction consistency loss value can be determined, which is used to maintain the direction selectivity characteristics of the gradient operator, enhance the robustness, and avoid losing texture details after training. The gradient direction consistency loss value only considers the privacy area of the sample image, and does not consider the non-privacy area of the sample image. For example, a direction difference loss function can be used to determine the gradient direction consistency loss value, and other loss functions can also be used to determine the gradient direction consistency loss value. The loss function is not limited.
[0078] For example, the gradient direction consistency loss value can be determined by the following formula :
[0079] L2
[0080] In the above formula, represents the total number of pixel points of the sample image,
[0081] represents the pixel point set of the privacy area of the sample image, M represents the first pixel point of the privacy area of the sample image. M represents the predicted gradient direction of the first second pixel point in the privacy area, j represents the real gradient direction of the first second pixel point in the privacy area. j j j
[0082] After obtaining the privacy-sensitive region segmentation loss value and the gradient direction consistency loss value, the target loss value can be determined based on the privacy-sensitive region segmentation loss value and the gradient direction consistency loss value, such as a sum value of the privacy-sensitive region segmentation loss value and the gradient direction consistency loss value as the target loss value, or the privacy-sensitive region segmentation loss value and the gradient direction consistency loss value are weighted to obtain the target loss value.
[0083] For example, after obtaining the target loss value, the target network model can be optimized based on the target loss value to obtain an optimized model. For example, the gradient descent algorithm or the like can be used to optimize the target network model, and the optimization method is not limited. The optimization target is to make the target loss value smaller and smaller.
[0084] Then, it can be determined whether the optimized model has converged. For example, if the target loss value is less than a preset threshold, the optimized model has converged, and if the target loss value is not less than the preset threshold, the optimized model has not converged. For example, if the number of iterations reaches a number threshold, the optimized model has converged, and if the number of iterations does not reach the number threshold, the optimized model has not converged. For example, if the iteration time length reaches a time length threshold, the optimized model has converged, and if the iteration time length does not reach the time length threshold, the optimized model has not converged.
[0085] If the optimized model has not converged, the optimized model is updated to the target network model, and the operation of inputting the sample image to the target network model is returned. If the optimized model has converged, K gradient parameters are obtained from the network parameters of the optimized model. For example, the optimized model includes a plurality of network parameters, and K network parameters in all network parameters can be taken as K gradient parameters (such as w 1 、 w 2 、 w 3 )。
[0086] After obtaining the K gradient parameters, the horizontal direction gradient operator and the vertical direction gradient operator can be constructed based on the K gradient parameters. The construction method can refer to the description above, which will not be repeated here.
[0087] Step S12, determining the horizontal direction gradient of each pixel point in the source image based on the horizontal direction gradient operator, and determining the vertical direction gradient of each pixel point in the source image based on the vertical direction gradient operator.
[0088] And, determining the horizontal direction gradient of each pixel point in the encrypted image based on the horizontal direction gradient operator, and determining the vertical direction gradient of each pixel point in the encrypted image based on the vertical direction gradient operator.
[0089] For example, the horizontal gradient can be determined using the following formula: In the above formula, This represents the gradient operator in the horizontal direction. This represents the convolution operator, if Indicates the first element within the source image i If there are 100 pixels, then Indicates the first element within the source image i The horizontal gradient of each pixel, if Indicates the first [image] within the encrypted image i If there are 100 pixels, then Indicates the first [image] within the encrypted image i The horizontal gradient of each pixel.
[0090] For example, the vertical gradient can be determined using the following formula: In the above formula, It can represent the gradient operator in the vertical direction, if Indicates the first element within the source image i If there are 100 pixels, then It can represent the first element within the source image. i The vertical gradient of each pixel, if Indicates the first [image] within the encrypted image i If there are 100 pixels, then It can represent the first part of the encrypted image. i Vertical gradient of each pixel.
[0091] Step S13: Determine the first gradient direction and the first gradient magnitude based on the horizontal and vertical gradients of each pixel in the non-privacy region of the source image; determine the second gradient direction and the second gradient magnitude based on the horizontal and vertical gradients of each pixel in the non-privacy region of the encrypted image.
[0092] For example, for each pixel within a non-privacy region of the source image, a first gradient direction and a first gradient magnitude are determined based on the horizontal and vertical gradients of that pixel. For instance, the first gradient direction of the pixel can be determined using the following formula: The first gradient magnitude of the pixel is determined using the following formula: . The first non-privacy region of the source image represents the... i The first gradient direction of each pixel The first non-privacy region of the source image represents the... i The first gradient magnitude of each pixel. The first non-privacy region of the source image represents the... i Vertical gradient of each pixel The first non-privacy region of the source image represents the...i The horizontal gradient of each pixel.
[0093] For example, in the above formula, if The first non-privacy region of the encrypted image represents the i The vertical gradient of each pixel, and The first non-privacy region of the encrypted image represents the i The horizontal gradient of each pixel, then It can represent the first non-privacy region of an encrypted image. i The second gradient direction of each pixel It can represent the first non-privacy region of an encrypted image. i The second gradient magnitude of each pixel.
[0094] Step S14: Determine the third gradient direction (i.e., the gradient direction of the privacy region) based on the horizontal and vertical gradients of each pixel within the privacy region of the source image; determine the fourth gradient direction based on the horizontal and vertical gradients of each pixel within the privacy region of the encrypted image.
[0095] For example, for each pixel within a privacy region of the source image, the third gradient direction of that pixel is determined based on its horizontal and vertical gradients. For instance, the third gradient direction of the pixel can be determined using the following formula: . The first region within the privacy area of the source image j The third gradient direction of each pixel The first region within the privacy area of the source image j Vertical gradient of each pixel The first region within the privacy area of the source image j The horizontal gradient of each pixel.
[0096] For example, in the above formula, if The first part of the privacy region representing the encrypted image j The vertical gradient of each pixel, and The first part of the privacy region representing the encrypted image j The horizontal gradient of each pixel, then The first part of the privacy region representing the encrypted image j The fourth gradient direction of each pixel.
[0097] Step S15: Determine the directional similarity between the source image and the encrypted image based on the first gradient direction of the non-privacy region of the source image and the second gradient direction of the non-privacy region of the encrypted image.
[0098] For example, directional similarity can be determined based on the first gradient direction of each pixel within the non-privacy region of the source image and the second gradient direction of each pixel within the non-privacy region of the encrypted image. For instance, the first gradient direction of each pixel can be converted into a first unit vector, and the second gradient direction can be converted into a second unit vector. Then, directional similarity can be determined based on the first and second unit vectors of each pixel.
[0099] For example, regarding the first non-privacy area i A number of pixels can be converted into a unit vector using the following formula: , . The first non-privacy region of the source image S represents the first... i The first gradient direction of each pixel The first non-privacy region of the source image S represents the first... i The first unit vector of each pixel. The first non-privacy region of encrypted image C represents the... i The second gradient direction of each pixel The first non-privacy region of encrypted image C represents the... i The second unit vector of each pixel.
[0100] Then, the directional similarity can be determined using the following formula: In the above formula, The directional similarity value represents the directional similarity between the source image S and the encrypted image C. The closer the directional similarity value is to 1, the more consistent the directional distribution between the source image S and the encrypted image C. N represents the total number of pixels in the non-privacy region (the non-privacy regions of the source image S and the encrypted image C are in the same location, and the privacy regions of the source image S and the encrypted image C are in the same location). i Indicates the first [item] within the non-privacy area i 1 pixel i The value range is from 1 to N. The first non-privacy region of the source image S represents the first... i The first unit vector of each pixel. The first non-privacy region of encrypted image C represents the... i The second unit vector of each pixel.
[0101] Step S16: Determine the amplitude similarity between the source image and the encrypted image based on the first gradient magnitude of the non-privacy region of the source image, the second gradient magnitude of the non-privacy region of the encrypted image, the pixel value of each pixel in the non-privacy region of the source image, and the pixel value of each pixel in the non-privacy region of the encrypted image.
[0102] Exemplarily, based on the first gradient amplitude of each pixel in the non-privacy area of the source image, the second gradient amplitude of each pixel in the non-privacy area of the encrypted image, the first mean value (i.e. the average of the pixel values of all pixels in the non-privacy area) and the first standard deviation (i.e. the standard deviation of the pixel values of all pixels in the non-privacy area) of the non-privacy area of the source image, the second mean value and the second standard deviation of the non-privacy area of the encrypted image, the amplitude similarity of the source image and the encrypted image can be determined. For example, the amplitude similarity can be determined by using the following formula: . denotes the amplitude similarity of the source image and the encrypted image, denotes the first mean value, denotes the second mean value, denotes the first standard deviation, denotes the second standard deviation. denotes the average of the first gradient amplitudes of each pixel in the non-privacy area of the source image, denotes the average of the second gradient amplitudes of each pixel in the non-privacy area of the encrypted image. is a constant configured according to actual requirements, which is greater than 0 and is used to prevent the value from being 0, is a constant configured according to actual requirements, which is greater than 0 and is used to prevent the value from being 0.
[0103] Step S17, determining the first texture similarity of the non-privacy area based on the direction similarity and the amplitude similarity.
[0104] Exemplarily, the direction similarity and the amplitude similarity can be weighted to obtain the first texture similarity of the non-privacy area, i.e. the comprehensive texture similarity under the direction similarity and the amplitude similarity. For example, the first texture similarity can be determined by using the following formula : In the above formula, denotes the weighting coefficient of the direction similarity , which can be configured according to actual requirements, denotes the weighting coefficient of the amplitude similarity , which can be configured according to actual requirements.
[0105] Step S18, determining the second texture similarity of the privacy area based on the third gradient direction of the privacy area of the source image, the fourth gradient direction of the privacy area of the encrypted image and the total number of pixels in the privacy area.
[0106] Exemplarily, based on the third gradient direction of each pixel point in the privacy area of the source image, the fourth gradient direction of each pixel point in the privacy area of the encrypted image, and the total number of pixels in the privacy area (the privacy area positions of the source image and the encrypted image are consistent, that is, the total number of pixels in the privacy area is consistent), the second texture similarity of the privacy area can be determined. For example, the second texture similarity can be determined by using the following formula: In the above formula, may represent the second texture similarity of the privacy area, M can represent the total number of pixel points in the privacy area, that is, M can represent the pixel point set of the privacy area, j represents the first j pixel point in the privacy area, j the value range of is 1 to M. may represent the third gradient direction of the first j pixel point in the privacy area of the source image S, may represent the fourth gradient direction of the first j pixel point in the privacy area of the encrypted image C. Based on the above formula, if the privacy area of the encrypted image is completely irrelevant to the privacy area of the source image, then .
[0107] Step S19, determining a target texture similarity based on the first texture similarity and the second texture similarity.
[0108] Exemplarily, the first texture similarity and the second texture similarity can be weighted to obtain the target texture similarity between the source image and the encrypted image. For example, the target texture similarity can be determined by using the following formula : In the above formula, may represent the weighting coefficient of the second texture similarity , which can be configured according to actual requirements, may represent the weighting coefficient of the first texture similarity , which can be configured according to actual scene requirements.
[0109] In summary, in the embodiment, the target texture similarity is determined based on the trainable horizontal direction gradient operator and the vertical direction gradient operator, so as to evaluate the privacy protection effect through the target texture similarity.
[0110] Considering the consistency of the direction and the local texture statistical features, the image after privacy protection should maintain the texture similarity with the source image in the non-privacy protection area, and show significant differences in the privacy area. Based on this, the target texture similarity can be determined based on the first texture similarity of the non-privacy area and the second texture similarity of the privacy area. For the privacy area, the significance of the direction difference in the privacy area is calculated considering the destruction degree, and then the second texture similarity is determined. For the non-privacy area, the direction consistency of the non-privacy area is considered to ensure the naturalness of the texture structure, and the protection effect can be accurately distinguished by using area separation calculation.
[0111] At this point, step 203 is completed, and the target texture similarity between the source image and the encrypted image is obtained.
[0112] Step 204, determining the target edge similarity based on the source image and the encrypted image, the target edge similarity is used to represent the similarity between the edge features of the source image and the edge features of the encrypted image. For example, the target edge similarity can be the edge structure similarity, which can quantify the explicit privacy leakage risk.
[0113] For example, given the source image S and the encrypted image C, the similarity between the edge features of the source image S and the edge features of the encrypted image C can be determined, and then the target edge similarity is obtained. The determination method of the target edge similarity is not limited, as long as it can represent the similarity between the edge features. In one possible implementation, the target edge similarity can be determined by the following steps:
[0114] Step S21, determining the gradient amplitude of the source image based on the horizontal direction gradient and the vertical direction gradient of the source image, and performing normalization operation on the gradient amplitude of the source image to obtain the first target gradient amplitude.
[0115] And, determining the gradient amplitude of the encrypted image based on the horizontal direction gradient and the vertical direction gradient of the encrypted image, and performing normalization operation on the gradient amplitude of the encrypted image to obtain the second target gradient amplitude.
[0116] For example, for each pixel point of the source image, the horizontal direction gradient of the pixel point can be determined based on the horizontal direction gradient operator, and the vertical direction gradient of the pixel point can be determined based on the vertical direction gradient operator. The horizontal direction gradient operator and the vertical direction gradient operator can be configured according to actual needs, and can also include K gradient parameters, which are obtained by optimizing the target network model, see step S11.
[0117] For example, based on the horizontal direction gradient and the vertical direction gradient of the pixel point, the gradient amplitude of the pixel point can be determined, for example, using the following formula to determine the gradient amplitude of the pixel point: , The gradient amplitude of the pixel point can be represented as may represent the horizontal direction gradient of the pixel point, may represent the vertical direction gradient of the pixel point. Then, the gradient amplitude of the pixel point is normalized, such as scaling the gradient amplitude of the pixel point to the pixel range of 0~255, to obtain the first target gradient amplitude of the pixel point. For example, the first target gradient amplitude of the pixel point can be determined by the following formula: may represent the first target gradient amplitude of the pixel point, may represent the minimum value in the gradient amplitudes of all pixel points in the source image, may represent the maximum value in the gradient amplitudes of all pixel points in the source image. k may be configured according to actual requirements, which is a parameter value of the normalization operation, such as scaling the gradient amplitude to the pixel range of 0~255, k may be 255.
[0118] In summary, for each pixel point of the source image, the first target gradient amplitude of the pixel point can be obtained, that is, the first target gradient amplitude of each pixel point of the source image is obtained. Similarly, for each pixel point of the encrypted image, the second target gradient amplitude of the pixel point can be obtained, which can be obtained by replacing the source image with the encrypted image, that is, the second target gradient amplitude of each pixel point of the encrypted image is obtained.
[0119] Step S22, generating a first gradient histogram corresponding to the source image based on the first target gradient amplitude; wherein the first gradient histogram can include the gradient amplitude distribution of each pixel point in the source image.
[0120] And, generating a second gradient histogram corresponding to the encrypted image based on the second target gradient amplitude; wherein the second gradient histogram can include the gradient amplitude distribution of each pixel point in the encrypted image.
[0121] For example, the first gradient histogram can be generated based on the first target gradient amplitude corresponding to each pixel point of the source image , such as generating the first gradient histogram by the following formula: In the above formula, k represents the abscissa of the first gradient histogram, and if the value range of the first target gradient amplitude is 0~255, then k the values of are 0, 1, 2, …, 254, 255 in turn, represents the ordinate of the first gradient histogram, which is used to represent the count value corresponding to the value “ k ”. For example, for each value of “ k ”, each pixel point of the source image (i.e. pixel point i ), if the first target gradient amplitude of the pixel point is equal to k , the count value of is added by 1, if the first target gradient amplitude of the pixel point is not equal to k , the pixel point is ignored. After traversing all the pixel points, the count value of can be obtained.
[0122] For example, for the position with the horizontal coordinate of 0, the count value is counted, and the count value is the number of pixel points with the first target gradient amplitude of 0. For the position with the horizontal coordinate of 1, the count value is counted, and the count value is the number of pixel points with the first target gradient amplitude of 1. For the position with the horizontal coordinate of 255, the count value is counted, and the count value is the number of pixel points with the first target gradient amplitude of 255. Thus, 256 count values can be obtained, and the 256 count values form the first gradient histogram. The first gradient histogram can include 256 coordinate points, and the horizontal coordinates of the 256 coordinate points are 0, 1, 2, …, 254, 255 in turn, and the vertical coordinates of the 256 coordinate points can be the corresponding count values.
[0123] For example, the second gradient histogram can be generated based on the second target gradient amplitude corresponding to each pixel point of the encrypted image. The generation manner of the second gradient histogram is similar to that of the first gradient histogram, and thus will not be repeated here.
[0124] In step S23, a joint gradient histogram is generated based on the first gradient histogram and the second gradient histogram.
[0125] For example, the joint gradient histogram can reflect the common edge features of the source image and the encrypted image, and the encrypted image should retain part of the edge structure, otherwise the practicability will be lost. For example, the joint gradient histogram can be generated by using the following formula: For example, represents the joint gradient histogram, represents the first gradient histogram, represents the second gradient histogram. For the position with the horizontal coordinate of 0 of the joint gradient histogram, the count value is the average value of and , where is the count value corresponding to the horizontal coordinate of 0 of the first gradient histogram, is the count value corresponding to the horizontal coordinate of 0 of the second gradient histogram. For the position with the horizontal coordinate of 255 of the joint gradient histogram, the count value is the average value of and , where It is the count value corresponding to the x-coordinate 255 of the first gradient histogram. It is the count value corresponding to the x-coordinate 255 of the second gradient histogram.
[0126] Step S24: Determine the lower threshold and upper threshold based on the joint gradient histogram.
[0127] For example, the total count value of the joint gradient histogram (i.e., the sum of all count values) can be determined and denoted as . Total ,but Total= *0+ *1+ *2+…+ *254+ *255, This represents the count value corresponding to the position with an x-coordinate of 0 in the joint gradient histogram, and so on. This represents the count value corresponding to the position with an x-coordinate of 255 in the joint gradient histogram.
[0128] For example, a first threshold and a second threshold can be determined based on the total count value, where the first threshold can be less than the second threshold. For instance, the lowest 5% and highest 5% of gradients can be excluded from the valid range to remove potentially noisy or over-encrypted regions, while retaining the middle 90% of the gradient range, which corresponds to significant edges. Thus, the first threshold can be 0.05* Total The second threshold can be 0.95* Total Of course, "0.05" and "0.95" are just examples and can be configured according to actual needs.
[0129] For example, the lower threshold value can be determined based on the first threshold and the second threshold. t min and threshold upper limit t max The lower and upper threshold values form an effective interval, which is then determined by the cumulative distribution of the joint histogram. For example, the lower and upper threshold values can be determined using the following formula: For the above formula, first iterate through... k =0, for *0, if If it is greater than or equal to the first threshold, then t min If the value is 0, then continue iterating. k =1, for *0+ *1, if greater than or equal to the first threshold value, then t min is 1, otherwise, continue to traverse k = 2, is * 0 * 1 * 2, if greater than or equal to the first threshold value, then t min is 2, otherwise, continue to traverse k = 3, and so on, until the first greater than or equal to the first threshold value is found, this corresponding value represents the lower threshold value t min . Similarly, first traverse k = 0, if greater than or equal to the second threshold value, then t max is 0, otherwise, continue to traverse k = 1, if greater than or equal to the second threshold value, then t max is 1, otherwise, continue to traverse k = 2, if greater than or equal to the second threshold value, then t max is 2, otherwise, continue to traverse k = 3, and so on, until the first greater than or equal to the second threshold value is found, this corresponding value represents the upper threshold value t max .
[0130] Step S25, generating N uniformly distributed dynamic thresholds based on the lower threshold value and the upper threshold value, and the N dynamic thresholds form a threshold set, that is, the threshold set can include N dynamic thresholds.
[0131] For example, N can be a positive integer greater than 1. For example, N can be pre-configured, that is, the value of N is configured according to actual needs, or N can also be determined based on the number of effective peaks of the joint gradient histogram. For example, if the number of effective peaks of the joint gradient histogram is M, then N = M + 2. For "2" dynamic thresholds, they can be the lower threshold value and the upper threshold value. For "M" dynamic thresholds, they can be the M dynamic thresholds corresponding to the M effective peaks. How to detect the effective peaks of the joint gradient histogram is not limited in this embodiment, and taking detection of M effective peaks (i.e., significant peaks) as an example.
[0132] For example, based on the lower threshold value tmin and a threshold upper value t max The N uniform distribution dynamic thresholds can be generated by the following formula: In the above formula, TS denotes a threshold set, which includes N dynamic thresholds, denotes the i dynamic threshold, when i is 0, is a threshold lower value t min when i is N-1, is a threshold upper value t max .
[0133] For example, for the edge detection process, a fixed threshold can be configured according to experience, however, if the fixed threshold is set too high, weak edges will be missed, and if the fixed threshold is set too low, noise will be introduced. In view of the above finding, the dynamic threshold is proposed in the embodiment, which can adaptively select a threshold range according to the image content, ensure to cover meaningful edges, and exclude irrelevant areas, that is, to obtain N dynamic thresholds.
[0134] In step S26, for each dynamic threshold in the threshold set, a first edge binary image corresponding to the source image is determined based on the dynamic threshold, and a second edge binary image corresponding to the encrypted image is determined based on the dynamic threshold.
[0135] For example, for each pixel point of the first edge binary image, the pixel point can have a first value or a second value, if the pixel point has the first value (such as 1), it means that the pixel point corresponding to the pixel point in the source image is an edge point, and if the pixel point has the second value (such as 0), it means that the pixel point corresponding to the pixel point in the source image is a non-edge point. In summary, it is necessary to detect whether each pixel point in the source image is an edge point or a non-edge point, and then generate a first edge binary image corresponding to the source image, denoted as For example, the pixel point in the first edge binary image has a first value or a second value, the first value in the first edge binary image can represent an edge point, and the second value can represent a non-edge point.
[0136] As to how to detect whether each pixel point in the source image is an edge point or a non-edge point, an edge detection algorithm can be used, in which the information of the pixel point is compared with the dynamic threshold, and then it is detected whether the pixel point is an edge point, and the detection process is not limited. Obviously, the first edge binary image is related to the dynamic threshold, and different dynamic thresholds will correspond to different first edge binary images. In the above first edge binary image In the formula, t represents a dynamic threshold, i.e. the first edge binary image corresponding to the dynamic threshold t.
[0137] For example, for each pixel point of the second edge binary image, if the pixel point is the first value (e.g. 1), it indicates that the corresponding pixel point in the encrypted image is an edge point, and if the pixel point is the second value (e.g. 0), it indicates that the corresponding pixel point in the encrypted image is a non-edge point. In summary, it is necessary to detect whether each pixel point in the encrypted image is an edge point or a non-edge point, and then generate the second edge binary image corresponding to the encrypted image, denoted as .
[0138] Step S27, obtaining a common edge image based on the first edge binary image and the second edge binary image.
[0139] For example, for each pixel point in the common edge image, if the logical AND operation result of the corresponding pixel value of the pixel point in the first edge binary image and the corresponding pixel value of the pixel point in the second edge binary image is the first value, and the absolute value of the difference between the first pixel value corresponding to the pixel point in the source image and the second pixel value corresponding to the pixel point in the encrypted image is not greater than the brightness control threshold, then the pixel value corresponding to the pixel point in the common edge image is determined based on the first pixel value, the second pixel value and the brightness control threshold; otherwise, the pixel value corresponding to the pixel point in the common edge image is determined as the second value.
[0140] For example, the common edge image can be obtained by using the following formula, denoted as common edge image :
[0141]
[0142] In the above formula, the pixel point in the common edge image is taken as an example, (i,j) i represents the horizontal coordinate of the pixel point, j represents the vertical coordinate of the pixel point. t represents a dynamic threshold, i.e. the first edge binary image, the second edge binary image and the common edge image all correspond to the dynamic threshold t. represents the pixel value corresponding to the pixel point (i,j) in the common edge image. represents the pixel value corresponding to the pixel point (i,j) in the first edge binary image, represents the pixel value corresponding to the pixel point (i,j) in the second edge binary image, represents the logical AND operation, represents the logical AND operation result being the first value (1). pixel point (i,j) a first pixel value corresponding to the pixel point in the source image S, pixel point (i,j) a second pixel value corresponding to the pixel point in the encrypted image C, represents a luminance control threshold.
[0143] In the above formula, if is not 1, and / or, is greater than , it is determined that the pixel point (i,j) the pixel value corresponding to the pixel point in the common edge image is the second value (0).
[0144] In the above formula, , is a threshold value for controlling the luminance change between and . Through experiments, the luminance control threshold may be 20, or other numerical values, which are not limited. As can be seen from the above, through the luminance control threshold , in order to more accurately obtain the common edge set between the source image S and the encrypted image C, not only the edge detection result is considered, but also the luminance difference is considered.
[0145] Step S28, for each dynamic threshold, determining the edge similarity corresponding to the dynamic threshold based on the common edge image corresponding to the dynamic threshold and the first edge binary image corresponding to the dynamic threshold.
[0146] Exemplarily, the edge similarity between the source image S and the encrypted image C can be determined by the following formula: . t represents the dynamic threshold, that is, the edge similarity, the common edge image and the first edge binary image are all corresponding to the dynamic threshold t. represents the edge similarity between the source image S and the encrypted image C. represents the common edge image, represents the first edge binary image.
[0147] represents the norm of the common edge image , for example, represents the sum of all non-zero elements in the common edge image , of course, here is only an example of the norm, which is not limited.
[0148] represents the norm of the first edge binary image , for example, represents the sum of all non-zero elements in the first edge binary image The sum of all non-zero elements, of course, is just an example of a norm here.
[0149] Step S29, the edge similarity corresponding to each dynamic threshold is weighted to obtain a target edge similarity.
[0150] For example, for each dynamic threshold in the threshold set, based on steps S26-S28, the edge similarity corresponding to the dynamic threshold can be obtained. On this basis, the edge similarity corresponding to each dynamic threshold can be weighted to obtain a target edge similarity. For example, the target edge similarity can be determined using the following formula: . denotes the target edge similarity, denotes the weight value of the dynamic threshold t, denotes the edge similarity corresponding to the dynamic threshold t. The dynamic threshold t belongs to the threshold set TS, and the threshold set TS can include N dynamic thresholds.
[0151] For each dynamic threshold in the threshold set TS, a weight value can be assigned to each dynamic threshold. The weight values of different dynamic thresholds can be the same or different. For example, the sum of the weight values of all dynamic thresholds can be 1, that is, represented by the following formula: .
[0152] In summary, in this embodiment, considering that encryption will destroy the image structure, but may excessively destroy the utility, and visual anomalies may attract the attention of attackers. Therefore, by preserving the high edge similarity, the edge position is kept unchanged, but the actual semantics of the edge region is modified, which balances privacy and utility. The ideal edge similarity is sufficient to preserve the structural utility, but the semantic content associated with the edge has been desensitized.
[0153] At this point, step 204 is completed, and the target edge similarity between the source image and the encrypted image is obtained.
[0154] Step 205, determining a target semantic similarity based on the source image and the encrypted image, the target semantic similarity being used to represent the similarity between the semantic features of the source image and the semantic features of the encrypted image. For example, based on the target semantic similarity, privacy protection evaluation and analysis can be realized, the target semantic similarity can represent the semantic correlation between the non-sensitive region (non-private region) and the sensitive region, and can quantify the risk of implicit privacy leakage.
[0155] For example, in the field of privacy protection, semantic risk refers to the possibility that an attacker infers sensitive information by analyzing semantic information (such as objects, scenes, relationships, etc.) in multimedia data (such as images, videos), which can measure the potential threat of implicit privacy leakage in data that can be leaked through advanced semantic understanding.
[0156] For example, given a source image S and an encrypted image C, the similarity between the semantic features of the source image S and the semantic features of the encrypted image C can be determined, thereby obtaining the target semantic similarity. The method for determining this target semantic similarity is not limited, as long as it can characterize the similarity between semantic features. In one possible implementation, the target semantic similarity can be determined using the following steps:
[0157] Step S31: Determine the semantic similarity of objects based on the overlap between the detection boxes of the source image and the detection boxes of the encrypted image. For example, the detection box of the source image can be a region (e.g., a rectangular region) of each identified object (i.e., a sensitive object, such as a face object, a license plate object, an ID card object, etc.) within the source image, and the detection box of the encrypted image can be a region (e.g., a rectangular region) of each identified object within the encrypted image.
[0158] For example, a source image can be input into a detection model (the network structure of this detection model is not limited), the detection model determines the objects to be identified in the source image, and outputs the detection boxes (at least one detection box) containing these objects, denoted as detection box a1, detection box a2, and detection box a3. This process can be referred to step 202. An encrypted image can be input into a detection model, the detection model determines the objects to be identified in the encrypted image, and outputs the detection boxes containing these objects, denoted as detection box b1 and detection box b2.
[0159] For example, object semantic similarity can be determined based on the overlap (IOU) between the detection boxes in the source image and the encryption image. For instance, the greater the overlap between the detection boxes in the source image and the encryption image, the greater the object semantic similarity. For example, object semantic similarity can be determined based on the overlap between detection boxes a1 and b1, a1 and b2, a2 and b1, a2 and b2, a3 and b1, and a3 and b2. Object semantic similarity directly quantifies the risk of semantic leakage by comparing the detection differences between the source image and the encryption image. The risk is assessed by the number of correctly detected objects after encryption.
[0160] For example, the semantic similarity of objects can be determined using the following formula. : . Represents the source image's first... i One detection box, The first part of the encrypted image j One detection box, For object-sensitive weights, multiple object-sensitive weights can be configured according to actual needs. If the source image has 3 detection boxes and the encrypted image has 2 detection boxes, then 6 object-sensitive weights can be obtained.
[0161] Regarding the overlap between detection boxes a1 and b1, The first object-sensitive weight is used to measure the overlap between detection boxes a1 and b2. The second object-sensitive weight is used to measure the overlap between detection boxes a2 and b1. The third object is assigned a sensitive weight, and so on.
[0162] For example, the semantic similarity of objects can be determined based on the overlap between the detection boxes and ground truth boxes in the source image and the overlap between the detection boxes and ground truth boxes in the encrypted image. The greater the overlap between the detection boxes and ground truth boxes in the source image, the greater the semantic similarity of the objects. Similarly, the greater the overlap between the detection boxes and ground truth boxes in the encrypted image, the greater the semantic similarity of the objects.
[0163] For example, ground truth bounding boxes are the actual detection boxes in the source image. If the source image contains three objects (i.e., sensitive objects), then the accurate regions (e.g., rectangular regions) of these three objects need to be obtained. These accurate regions are the three ground truth bounding boxes. Compared to the detection boxes in the source image, which are outputs of the detection model and may contain errors, the ground truth bounding boxes are accurately obtained detection boxes with no or very small errors. These ground truth bounding boxes can be manually labeled by the user or labeled using a certain algorithm; the key is accurate labeling, and there are no restrictions on this.
[0164] For example, based on the detection boxes of the source image, the detection boxes of the encrypted image, and the ground truth bounding boxes, the semantic similarity of objects is determined using the following formula. : . Represents the source image's first... i One detection box, The first part represents the encrypted image. j One detection box, Indicates the first k A real annotation box, You can configure multiple object-sensitive weights according to your actual needs.
[0165] For the ground truth bounding box c1, detection box a1, and detection box b1, The first object-sensitive weight is assigned to the ground truth bounding box c1, the detection box a1, and the detection box b2. The 2nd object-sensitive weight is for the real label box c1, the detection box a2, and the detection box b1, The 3rd object-sensitive weight is for the real label box c1, the detection box a2, and the detection box b2, The 4th object-sensitive weight is for the real label box c1, the detection box a2, and the detection box b2, and so on.
[0166] In step S32, the relationship semantic similarity is determined based on the semantic association graph of the source image and the semantic association graph of the encrypted image. The relationship semantic similarity can be determined by the number of correct relationships retained after encryption to determine the risk.
[0167] For example, the semantic association graph of the source image can represent the association relationship of each identified object in the source image, and the semantic association graph of the encrypted image can represent the association relationship of each identified object in the encrypted image.
[0168] For example, assuming that there are P1 identified objects (i.e., sensitive objects) in the source image, the semantic association graph of the source image can include P1* P1 pixel points. The pixel point in the first row and the first column represents whether the first identified object is associated with the first identified object. For example, if the pixel value of the pixel point is 1, it means that there is an association relationship, and if the pixel value is 0, it means that there is no association relationship. Obviously, this is for the same identified object, so there is an association relationship. The pixel point in the first row and the second column represents whether the first identified object is associated with the second identified object. For example, if a user holds an ID card, the user and the ID card are associated with each other, and if a cat is far away from a vehicle, the vehicle and the cat are not associated with each other. The pixel point in the second row and the first column represents whether the second identified object is associated with the first identified object, and the pixel value of the pixel point in the first row and the second column is the same.
[0169] The pixel point in the first row and the third column represents whether the first identified object is associated with the third identified object, and the pixel point in the third row and the first column represents whether the third identified object is associated with the first identified object. The pixel values of the pixel points in these two positions are the same, and so on.
[0170] Based on the above, the semantic association graph of the source image can be generated based on the source image, and the semantic association graph of the encrypted image can be generated based on the encrypted image. Based on the semantic association graph of the source image and the semantic association graph of the encrypted image, the relationship semantic similarity is determined by using the following formula . The semantic association graph of the source image is represented by The semantic association graph of the encrypted image is represented by The graph edit distance is represented by represents a graph edit distance between a semantic correlation graph of the source image and a semantic correlation graph of the encrypted image. represents a maximum value in and .
[0171] Step S33, determining a scene semantic similarity based on the scene feature vector of the source image and the scene feature vector of the encrypted image. For example, inputting the source image into a feature extraction network (such as a ResNet network) to obtain the scene feature vector of the source image, that is, the feature extraction network performs feature extraction on the source image to obtain the scene feature vector of the source image. Input the encrypted image into the feature extraction network to obtain the scene feature vector of the encrypted image, that is, the feature extraction network performs feature extraction on the encrypted image to obtain the scene feature vector of the encrypted image.
[0172] For example, the scene semantic similarity is determined by the following formula : . represents the scene feature vector of the source image, represents the scene feature vector of the encrypted image, and cos represents the cosine similarity, that is, the cosine similarity between the scene feature vector of the source image and the scene feature vector of the encrypted image.
[0173] Step S34, performing weighted operation on the object semantic similarity, the relationship semantic similarity and the scene semantic similarity to obtain a semantic risk value. For example, the weighting coefficient of the object semantic similarity is greater than the weighting coefficient of the relationship semantic similarity, and the weighting coefficient of the relationship semantic similarity is greater than the weighting coefficient of the scene semantic similarity.
[0174] For example, the semantic risk value is determined by the following formula: , SLS The semantic risk value can be represented as w1 represents the weighting coefficient of the object semantic similarity, which can be configured according to experience, such as 0.5, w2 represents the weighting coefficient of the relationship semantic similarity, which can be configured according to experience, such as 0.3, w3 represents the weighting coefficient of the scene semantic similarity, which can be configured according to experience, such as 0.2. In this embodiment, the weighting coefficients are not limited, w1 may be greater than w2 may be less than w2 , w1 may be greater than w3 may be less than w3 , w2 may be greater than w3 may be less than w3 .
[0175] Step S35, determining a target semantic similarity based on the semantic risk value.
[0176] For example, the target semantic similarity can be determined by the following formula D SR : .
[0177] Thus, the target semantic similarity between the source image and the encrypted image is obtained in step 205.
[0178] In step 206, the privacy protection index is determined based on the target texture similarity, the target edge similarity and the target semantic similarity. For example, the target texture similarity, the target edge similarity and the target semantic similarity can be subjected to weighted operation to obtain the privacy protection index. The first weighting coefficient of the target texture similarity can be greater than the second weighting coefficient of the target edge similarity, or can be less than the second weighting coefficient. The first weighting coefficient can be greater than the third weighting coefficient of the target semantic similarity, or can be less than the third weighting coefficient. The second weighting coefficient can be greater than the third weighting coefficient, or can be less than the third weighting coefficient.
[0179] In one possible implementation, the privacy protection index can be determined by the following steps:
[0180] In step S41, the first standard deviation is determined based on the target texture similarity, the second standard deviation is determined based on the target edge similarity, and the third standard deviation is determined based on the target semantic similarity.
[0181] For example, in the process of determining the visual multi-dimensional quantization index, a plurality of source images (such as each frame image or part of the image in the video stream as a source image) can be obtained. For each source image, the target texture similarity, the target edge similarity and the target semantic similarity corresponding to the source image can be obtained by using the above steps, i.e. a plurality of target texture similarities, a plurality of target edge similarities and a plurality of target semantic similarities are obtained.
[0182] On this basis, the first standard deviation can be determined based on the plurality of target texture similarities, the second standard deviation can be determined based on the plurality of target edge similarities, and the third standard deviation can be determined based on the plurality of target semantic similarities. For example, the standard deviation can be represented by the following formula : , if represents the target texture similarity, then std represents the first standard deviation of all target texture similarities , if represents the target edge similarity, then std represents the second standard deviation of all target edge similarities , if represents the target semantic similarity, then stda third standard deviation representing calculating all target semantic similarities .
[0183] Step S42, determining a first correlation coefficient based on the target texture similarity and a pre-configured real leakage risk label, determining a second correlation coefficient based on the target edge similarity and the real leakage risk label, and determining a third correlation coefficient based on the target semantic similarity and the real leakage risk label.
[0184] For example, the real leakage risk label can be pre-configured, and the real leakage risk label can be configured according to actual scene requirements, such as 0.9, 0.8, 0.5, etc., which is not limited.
[0185] For example, the correlation coefficient can be represented by the following formula : , y The real leakage risk label is represented by r. If The target texture similarity is represented by s, and corr The first correlation coefficient of the target texture similarity and the real leakage risk label is represented by s1. If The target edge similarity is represented by e, and corr The second correlation coefficient of the target edge similarity and the real leakage risk label is represented by e1. If The target semantic similarity is represented by t, and corr The third correlation coefficient of the target semantic similarity and the real leakage risk label is represented by t1. .
[0186] For example, corr The correlation coefficient formula is used to measure the degree of linear correlation between two variables (such as the target texture similarity and the real leakage risk label), and the result is a value between -1 and 1, which can be determined based on the covariance of the target texture similarity and the real leakage risk label, the standard deviation of the target texture similarity (i.e. the first standard deviation of multiple target texture similarities), and the standard deviation of the real leakage risk label.
[0187] Step S43, determining a first weighting coefficient of the target texture similarity based on the first standard deviation and the first correlation coefficient, determining a second weighting coefficient of the target edge similarity based on the second standard deviation and the second correlation coefficient, and determining a third weighting coefficient of the target semantic similarity based on the third standard deviation and the third correlation coefficient.
[0188] For example, the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient can be determined by the following formula: If The first standard deviation of the target texture similarity is represented by s1, The first correlation coefficient represents the similarity of target textures. This represents the first weighting coefficient indicating the similarity of the target textures. If The second standard deviation represents the similarity of the target edges. The second correlation coefficient represents the similarity of the target edges. This represents the second weighting coefficient indicating the similarity of the target edges. If The third standard deviation represents the semantic similarity of the target. The third correlation coefficient, representing the semantic similarity of the targets, The third weighting coefficient represents the semantic similarity of the target.
[0189] In the above formula, k The value of is 1, 2, or 3. If k If the value is 1, then Indicates the first standard deviation. Represents the first correlation coefficient, if k If the value is 2, then Indicates the second standard deviation. This represents the second correlation coefficient, if k If the value is 3, then Indicates the third standard deviation. This represents the third correlation coefficient.
[0190] Step S44: Based on the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient, the target texture similarity, target edge similarity, and target semantic similarity are weighted to obtain the privacy protection index.
[0191] For example, the following formula can be used to determine the privacy protection index (denoted as the index score). ): . Indicates target texture similarity The corresponding first weighting coefficient, Indicates the similarity of target edges The corresponding second weighting coefficient, Represents target semantic similarity The corresponding third weighting coefficient. In summary, this embodiment proposes a dynamic weight fusion method that makes the weights and dimensions positively correlated with the ability to detect privacy leaks, thereby dynamically adapting to different image types.
[0192] This completes step 206, yielding the privacy protection index, or the privacy protection index score.
[0193] Step 207: Determine whether the privacy protection indicator meets the preset conditions. If the privacy protection indicator meets the preset conditions, proceed to step 208; if the privacy protection indicator does not meet the preset conditions, proceed to step 209.
[0194] For example, if a higher privacy protection index indicates better privacy protection, then when the privacy protection index is greater than a preset threshold, the privacy protection index meets the preset condition; when the privacy protection index is less than the preset threshold, the privacy protection index does not meet the preset condition. Alternatively, if a lower privacy protection index indicates better privacy protection, then when the privacy protection index is less than a preset threshold, the privacy protection index meets the preset condition; when the privacy protection index is not less than the preset threshold, the privacy protection index does not meet the preset condition.
[0195] Step 208: Determine if the encrypted image meets privacy protection requirements. In this case, the encrypted image can be transmitted. Since the encrypted image meets privacy protection requirements, data security can be guaranteed.
[0196] Step 209: Determine that the encrypted image does not meet the privacy protection requirements. In this case, do not transmit the encrypted image. Instead, re-process the source object using a different method to obtain the encrypted image, and repeat the above steps.
[0197] In one possible implementation, see Figure 3 The diagram illustrates a multidimensional quantification index for vision. The input layer acquires the source image. The preprocessing layer cleanses the source image, performing tasks such as de-identification and noise injection. The analysis layer performs content recognition on the source image to obtain an encrypted image. For example, it can perform object detection and semantic segmentation to obtain the encrypted image. The verification layer uses a dynamic threshold-based edge detection algorithm to obtain target edge similarity. It can also use a semantic association graph to obtain target semantic similarity. Finally, it uses a texture extraction algorithm based on trainable operators to obtain target texture similarity. The output layer determines a privacy protection index (i.e., a privacy score) based on target edge similarity, target semantic similarity, and target texture similarity. It then analyzes whether the encrypted image meets privacy protection requirements based on this index and provides optimization suggestions when the encrypted image does not meet these requirements.
[0198] From the above technical solutions, in the embodiments of the present application, by constructing multiple visual quantization dimensions (such as texture, edge, semantic, etc. dimensions), a multi-dimensional quantization index system is designed, and a visual multi-dimensional quantization index is given, thereby providing a scientific and objective evaluation means for the privacy protection effect in multimedia data, that is, the privacy protection effect in multimedia data can be scientifically and objectively evaluated, and an objective evaluation basis is provided for optimization and improvement of privacy protection technologies such as encryption, blurring, and desensitization, and the privacy protection effect of encrypted images can be more accurately evaluated. By fusing edge similarity (ES) and texture similarity (TS), the risk of contour and detail leakage is comprehensively covered. Based on the target detection model to identify sensitive areas, the weight distribution of edge and texture is dynamically adjusted. Joint analysis of edge and texture features blocks the double leakage risk of contour and detail. Through the edge detection algorithm based on dynamic threshold optimization to extract image edge information, through the trainable gradient operator to extract texture similarity, through three dimensions such as semantic similarity, combined with adaptive weight distribution and formal security proof, an objective evaluation basis is provided for optimization and improvement of privacy protection technologies such as encryption, blurring, and desensitization. By changing the fixed operator to a trainable operator, image texture information can be more comprehensively extracted according to different scenarios. Through the dynamic threshold edge detection algorithm, the loss of edge information or the increase of noise caused by a single threshold value can be avoided, and the evaluation accuracy of the privacy protection effect is ensured.
[0199] Based on the same application concept as the above method, in the embodiments of the present application, a multi-dimensional quantization index evaluation device for multimedia privacy protection is proposed, which is applied to a terminal device to be subjected to multimedia privacy protection, as shown in Figure 4 The device can include:
[0200] The acquisition module 41 is configured to acquire a source image and an encrypted image; wherein the source image includes a privacy area and a non-privacy area, and the encrypted image is obtained by performing privacy processing on the privacy area of the source image.
[0201] The determining module 42 is configured to determine a target texture similarity based on the source image and the encrypted image, the target texture similarity being used to represent a similarity between texture features of the source image and texture features of the encrypted image; determine a target edge similarity based on the source image and the encrypted image, the target edge similarity being used to represent a similarity between edge features of the source image and edge features of the encrypted image; determine a target semantic similarity based on the source image and the encrypted image, the target semantic similarity being used to represent a similarity between semantic features of the source image and semantic features of the encrypted image; determine a privacy protection index based on the target texture similarity, the target edge similarity and the target semantic similarity; and determine that the encrypted image meets a privacy protection requirement if the privacy protection index meets a preset condition, or determine that the encrypted image does not meet the privacy protection requirement.
[0202] For example, when determining the target texture similarity based on the source image and the encrypted image, the determining module 42 is configured to determine a first gradient direction and a first gradient amplitude based on a non-private area of the source image, determine a second gradient direction and a second gradient amplitude based on a non-private area of the encrypted image, determine a direction similarity between the source image and the encrypted image based on the first gradient direction and the second gradient direction, determine an amplitude similarity between the source image and the encrypted image based on the first gradient amplitude, the second gradient amplitude, the non-private area of the source image and the non-private area of the encrypted image, determine a first texture similarity of the non-private area based on the direction similarity and the amplitude similarity, determine a third gradient direction based on a private area of the source image, determine a fourth gradient direction based on a private area of the encrypted image, determine a second texture similarity of the private area based on the third gradient direction, the fourth gradient direction and a total number of pixels in the private area, and determine the target texture similarity based on the first texture similarity and the second texture similarity.
[0203] For example, when determining the first gradient direction, the first gradient magnitude, the second gradient direction, the second gradient magnitude, the third gradient direction and the fourth gradient direction, the determining module 42 specifically: obtains a horizontal direction gradient operator and a vertical direction gradient operator, the horizontal direction gradient operator and the vertical direction gradient operator each include K gradient parameters, K is a positive integer; wherein the K gradient parameters are parameters in a target network model, and are obtained by optimizing the target network model; determines horizontal direction gradients of each pixel point in the source image based on the horizontal direction gradient operator, and determines vertical direction gradients of each pixel point in the source image based on the vertical direction gradient operator; determines the first gradient direction and the first gradient magnitude based on the horizontal direction gradients and the vertical direction gradients of each pixel point in a non-private area of the source image; determines the third gradient direction based on the horizontal direction gradients and the vertical direction gradients of each pixel point in a private area of the source image; determines horizontal direction gradients of each pixel point in the encrypted image based on the horizontal direction gradient operator, and determines vertical direction gradients of each pixel point in the encrypted image based on the vertical direction gradient operator; determines the second gradient direction and the second gradient magnitude based on the horizontal direction gradients and the vertical direction gradients of each pixel point in a non-private area of the encrypted image; and determines the fourth gradient direction based on the horizontal direction gradients and the vertical direction gradients of each pixel point in a private area of the encrypted image.
[0204] For example, when obtaining the K gradient parameters, the determining module 42 specifically: inputs a sample image into a target network model, and outputs, by the target network model, a predicted sensitive probability of each first pixel point in the sample image, and a predicted gradient direction of each second pixel point in a private area of the sample image, the predicted sensitive probability indicating a probability that the first pixel point belongs to a sensitive pixel; determines a private sensitive area segmentation loss value based on the predicted sensitive probability of each first pixel point in the sample image and a true sensitive probability of each first pixel point in the sample image; determines a gradient direction consistency loss value based on the predicted gradient direction of each second pixel point in the private area and a true gradient direction of each second pixel point in the private area; determines a target loss value based on the private sensitive area segmentation loss value and the gradient direction consistency loss value, optimizes the target network model based on the target loss value to obtain an optimized model; if the optimized model has not converged, updates the optimized model to the target network model, and returns to perform the operation of inputting the sample image into the target network model; and if the optimized model has converged, obtains the K gradient parameters from network parameters of the optimized model.
[0205] For example, the determining module 42 is specifically configured to: obtain a threshold set, the threshold set comprising a plurality of dynamic thresholds; for each dynamic threshold, determine a first edge binary image corresponding to the source image based on the dynamic threshold, and determine a second edge binary image corresponding to the encrypted image based on the dynamic threshold; wherein a first value in the first edge binary image represents an edge point, and a second value represents a non-edge point; obtain a common edge image; wherein for each pixel point in the common edge image, if a logical and operation result of a pixel value corresponding to the pixel point in the first edge binary image and a pixel value corresponding to the pixel point in the second edge binary image is the first value, and an absolute value of a difference between a first pixel value corresponding to the pixel point in the source image and a second pixel value corresponding to the pixel point in the encrypted image is not greater than a brightness control threshold, then determine a pixel value corresponding to the pixel point in the common edge image based on the first pixel value, the second pixel value and the brightness control threshold; otherwise, determine the pixel value corresponding to the pixel point in the common edge image as the second value; determine an edge similarity corresponding to the dynamic threshold based on the common edge image and the first edge binary image, and weight the edge similarity corresponding to each dynamic threshold to obtain the target edge similarity.
[0206] For example, the determining module 42 is specifically configured to: determine a gradient amplitude of the source image based on a horizontal direction gradient and a vertical direction gradient of the source image, and perform a normalization operation on the gradient amplitude of the source image to obtain a first target gradient amplitude; generate a first gradient histogram corresponding to the source image based on the first target gradient amplitude; wherein the first gradient histogram comprises a gradient amplitude distribution of each pixel point in the source image; determine a gradient amplitude of the encrypted image based on a horizontal direction gradient and a vertical direction gradient of the encrypted image, and perform a normalization operation on the gradient amplitude of the encrypted image to obtain a second target gradient amplitude; generate a second gradient histogram corresponding to the encrypted image based on the second target gradient amplitude; generate a joint gradient histogram based on the first gradient histogram and the second gradient histogram, determine a threshold lower limit value and a threshold upper limit value based on the joint gradient histogram; generate N uniformly distributed dynamic thresholds based on the threshold lower limit value and the threshold upper limit value, the threshold set comprising the N dynamic thresholds; wherein N is a positive integer greater than 1, N is pre-configured, or N is determined based on an effective peak number of the joint gradient histogram.
[0207] For example, the determining module 42 determines the target semantic similarity based on the source image and the encrypted image, and specifically is configured to: determine an object semantic similarity based on an overlap degree between a bounding box of the source image and a bounding box of the encrypted image, wherein the bounding box of the source image is a region of each identified object in the source image, and the bounding box of the encrypted image is a region of each identified object in the encrypted image; determine a relationship semantic similarity based on a semantic correlation graph of the source image and a semantic correlation graph of the encrypted image, wherein the semantic correlation graph of the source image represents a correlation relationship of each identified object in the source image, and the semantic correlation graph of the encrypted image represents a correlation relationship of each identified object in the encrypted image; determine a scene semantic similarity based on a scene feature vector of the source image and a scene feature vector of the encrypted image, wherein the scene feature vector of the source image is obtained by inputting the source image into a feature extraction network, and the scene feature vector of the encrypted image is obtained by inputting the encrypted image into the feature extraction network; perform a weighted operation on the object semantic similarity, the relationship semantic similarity, and the scene semantic similarity to obtain a semantic risk value, and determine the target semantic similarity based on the semantic risk value.
[0208] For example, the determining module 42 determines the target semantic similarity based on the source image and the encrypted image, and specifically is configured to: determine an object semantic similarity based on an overlap degree between a bounding box of the source image and a bounding box of the encrypted image, wherein the bounding box of the source image is a region of each identified object in the source image, and the bounding box of the encrypted image is a region of each identified object in the encrypted image; determine a relationship semantic similarity based on a semantic correlation graph of the source image and a semantic correlation graph of the encrypted image, wherein the semantic correlation graph of the source image represents a correlation relationship of each identified object in the source image, and the semantic correlation graph of the encrypted image represents a correlation relationship of each identified object in the encrypted image; determine a scene semantic similarity based on a scene feature vector of the source image and a scene feature vector of the encrypted image, wherein the scene feature vector of the source image is obtained by inputting the source image into a feature extraction network, and the scene feature vector of the encrypted image is obtained by inputting the encrypted image into the feature extraction network; perform a weighted operation on the object semantic similarity, the relationship semantic similarity, and the scene semantic similarity to obtain a semantic risk value, and determine the target semantic similarity based on the semantic risk value.
[0209] Based on the same application concept as the above method, an electronic device is provided in the embodiments of the present application, as shown in Figure 5 The electronic device includes a processor 51 and a machine readable storage medium 52, the machine readable storage medium 52 stores machine executable instructions that can be executed by the processor 51; the processor 51 is configured to execute the machine executable instructions to implement the multi-dimensional quantitative index evaluation method for multimedia privacy protection.
[0210] Based on the same application concept as the above method, the embodiment of the present application also provides a machine readable storage medium, wherein the machine readable storage medium stores a plurality of computer instructions, and the computer instructions can realize the multimedia privacy protection oriented multi-dimensional quantitative index evaluation method when executed by a processor.
[0211] The machine readable storage medium can be any electronic, magnetic, optical, or other physical storage apparatus, and can contain or store information such as executable instructions, data, and the like. For example, the machine readable storage medium can be a RAM (Random Access Memory), a volatile memory, a non-volatile memory, a flash memory, a storage drive (such as a hard disk drive), a solid state drive, any type of storage disc (such as an optical disc, a DVD, and the like), or a similar storage medium, or a combination thereof.
[0212] Based on the same application concept as the above method, the embodiment of the present application also provides a computer program product, wherein the computer program product comprises a computer program, and the computer program realizes the above multimedia privacy protection oriented multi-dimensional quantitative index evaluation method when executed by a processor.
[0213] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can adopt a computer program product in the form of being implemented on one or more computer usable storage media (including but not limited to a disk memory, a CD-ROM, an optical memory, and the like) containing computer usable program codes.
[0214] The above only describes the embodiments of the present application and is not used to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, and the like within the spirit and principle of the present application shall be included in the scope of claims of the present application.
Claims
1. A multi-dimensional quantification index evaluation method for multimedia privacy protection, characterized in that, The method is applied to a terminal device to be subjected to multimedia privacy protection, and comprises the following steps: obtaining a source image and an encrypted image; wherein the source image comprises a privacy area and a non-privacy area, and the encrypted image is obtained by performing privacy processing on the privacy area of the source image; determining a target texture similarity based on the source image and the encrypted image, wherein the target texture similarity is used to represent the similarity between the texture features of the source image and the texture features of the encrypted image; determining a target edge similarity based on the source image and the encrypted image, wherein the target edge similarity is used to represent the similarity between the edge features of the source image and the edge features of the encrypted image; determining a target semantic similarity based on the source image and the encrypted image, wherein the target semantic similarity is used to represent the similarity between the semantic features of the source image and the semantic features of the encrypted image; determining a privacy protection index based on the target texture similarity, the target edge similarity and the target semantic similarity; if the privacy protection index meets a preset condition, it is determined that the encrypted image meets the privacy protection requirement, otherwise, it is determined that the encrypted image does not meet the privacy protection requirement; wherein the determination of the target texture similarity based on the source image and the encrypted image comprises: determining a first gradient direction and a first gradient amplitude based on the non-privacy area of the source image, determining a second gradient direction and a second gradient amplitude based on the non-privacy area of the encrypted image, determining a direction similarity between the source image and the encrypted image based on the first gradient direction and the second gradient direction, determining an amplitude similarity between the source image and the encrypted image based on the first gradient amplitude, the second gradient amplitude, the non-privacy area of the source image and the non-privacy area of the encrypted image, and determining a first texture similarity of the non-privacy area based on the direction similarity and the amplitude similarity; determining a third gradient direction based on the privacy area of the source image, determining a fourth gradient direction based on the privacy area of the encrypted image, and determining a second texture similarity of the privacy area based on the third gradient direction, the fourth gradient direction and the total number of pixels in the privacy area; determining the target texture similarity based on the first texture similarity and the second texture similarity.
2. The method of claim 1, wherein the determination of the first gradient direction, the first gradient amplitude, the second gradient direction, the second gradient amplitude, the third gradient direction and the fourth gradient direction comprises: obtaining a horizontal direction gradient operator and a vertical direction gradient operator, wherein the horizontal direction gradient operator and the vertical direction gradient operator each comprise K gradient parameters, K being a positive integer; wherein the K gradient parameters are parameters in a target network model, and are obtained by optimizing the target network model. determining a horizontal direction gradient of each pixel point in the source image based on the horizontal direction gradient operator, and determining a vertical direction gradient of each pixel point in the source image based on the vertical direction gradient operator; determining the first gradient direction and the first gradient amplitude based on the horizontal direction gradient and the vertical direction gradient of each pixel point in the non-private area of the source image; determining the third gradient direction based on the horizontal direction gradient and the vertical direction gradient of each pixel point in the private area of the source image. determining a horizontal direction gradient of each pixel point in the encrypted image based on the horizontal direction gradient operator, and determining a vertical direction gradient of each pixel point in the encrypted image based on the vertical direction gradient operator; determining the second gradient direction and the second gradient amplitude based on the horizontal direction gradient and the vertical direction gradient of each pixel point in the non-private area of the encrypted image; determining the fourth gradient direction based on the horizontal direction gradient and the vertical direction gradient of each pixel point in the private area of the encrypted image.
3. The method of claim 2, wherein, the K gradient parameters are obtained by: inputting a sample image into a target network model, and outputting, by the target network model, a predicted sensitive probability of each first pixel point in the sample image, and a predicted gradient direction of each second pixel point in a private area of the sample image, the predicted sensitive probability indicating a probability that the first pixel point is a sensitive pixel; determining a private sensitive area segmentation loss value based on the predicted sensitive probability of each first pixel point in the sample image and a true sensitive probability of each first pixel point in the sample image; determining a gradient direction consistency loss value based on the predicted gradient direction of each second pixel point in the private area and a true gradient direction of each second pixel point in the private area; determining a target loss value based on the private sensitive area segmentation loss value and the gradient direction consistency loss value, and optimizing the target network model based on the target loss value to obtain an optimized model; if the optimized model has not converged, updating the optimized model to the target network model, and returning to inputting the sample image into the target network model; and if the optimized model has converged, obtaining the K gradient parameters from network parameters of the optimized model.
4. The method of claim 1, wherein the target edge similarity is determined based on the source image and the encrypted image by: obtaining a threshold set, the threshold set including a plurality of dynamic thresholds; for each dynamic threshold, determining a first edge binary image corresponding to the source image based on the dynamic threshold, and determining a second edge binary image corresponding to the encrypted image based on the dynamic threshold; wherein a first value in the first edge binary image indicates an edge point, and a second value indicates a non-edge point. obtaining a common edge image; wherein, for each pixel point in the common edge image, if a logical and operation result of a pixel value corresponding to the pixel point in the first edge binary image and a pixel value corresponding to the pixel point in the second edge binary image is a first value, and an absolute value of a difference between a first pixel value corresponding to the pixel point in the source image and a second pixel value corresponding to the pixel point in the encrypted image is not greater than a brightness control threshold, then determining a pixel value corresponding to the pixel point in the common edge image based on the first pixel value, the second pixel value and the brightness control threshold; otherwise, determining the pixel value corresponding to the pixel point in the common edge image as a second value; determining an edge similarity corresponding to the dynamic threshold based on the common edge image and the first edge binary image, and weighting the edge similarity corresponding to each dynamic threshold to obtain a target edge similarity.
5. The method of claim 4, wherein, The obtaining of the threshold set comprises: determining a gradient amplitude of the source image based on a horizontal direction gradient and a vertical direction gradient of the source image, and performing a normalization operation on the gradient amplitude of the source image to obtain a first target gradient amplitude; and generating a first gradient histogram corresponding to the source image based on the first target gradient amplitude; wherein, the first gradient histogram comprises a gradient amplitude distribution of each pixel point in the source image; determining a gradient amplitude of the encrypted image based on a horizontal direction gradient and a vertical direction gradient of the encrypted image, and performing a normalization operation on the gradient amplitude of the encrypted image to obtain a second target gradient amplitude; and generating a second gradient histogram corresponding to the encrypted image based on the second target gradient amplitude; generating a joint gradient histogram based on the first gradient histogram and the second gradient histogram, and determining a threshold lower limit value and a threshold upper limit value based on the joint gradient histogram; generating N dynamic thresholds uniformly distributed based on the threshold lower limit value and the threshold upper limit value, the threshold set comprising the N dynamic thresholds; wherein, N is a positive integer greater than 1, N is pre-configured, or N is determined based on an effective peak number of the joint gradient histogram.
6. The method of claim 1, wherein the determining of the target semantic similarity based on the source image and the encrypted image comprises: determining an object semantic similarity based on an overlap degree between a detection frame of the source image and a detection frame of the encrypted image; wherein, the detection frame of the source image is a region of each recognized object in the source image, and the detection frame of the encrypted image is a region of each recognized object in the encrypted image; determining a relationship semantic similarity based on a semantic association graph of the source image and a semantic association graph of the encrypted image; wherein, the semantic association graph of the source image represents an association relationship of each recognized object in the source image, and the semantic association graph of the encrypted image represents an association relationship of each recognized object in the encrypted image. determine a scene semantic similarity based on the scene feature vector of the source image and the scene feature vector of the encrypted image; wherein the scene feature vector of the source image is obtained by inputting the source image into a feature extraction network, and the scene feature vector of the encrypted image is obtained by inputting the encrypted image into the feature extraction network; perform weighted operation on the object semantic similarity, the relationship semantic similarity and the scene semantic similarity to obtain a semantic risk value, and determine a target semantic similarity based on the semantic risk value.
7. The method of claim 1, wherein, The determining of the privacy protection index based on the target texture similarity, the target edge similarity and the target semantic similarity comprises: determining a first standard deviation based on the target texture similarity, determining a first correlation coefficient based on the target texture similarity and a pre-configured real leakage risk label, and determining a first weighting coefficient corresponding to the target texture similarity based on the first standard deviation and the first correlation coefficient; determining a second standard deviation based on the target edge similarity, determining a second correlation coefficient based on the target edge similarity and the real leakage risk label, and determining a second weighting coefficient corresponding to the target edge similarity based on the second standard deviation and the second correlation coefficient; determining a third standard deviation based on the target semantic similarity, determining a third correlation coefficient based on the target semantic similarity and the real leakage risk label, and determining a third weighting coefficient corresponding to the target semantic similarity based on the third standard deviation and the third correlation coefficient; performing weighting on the target texture similarity, the target edge similarity and the target semantic similarity based on the first weighting coefficient, the second weighting coefficient and the third weighting coefficient to obtain the privacy protection index.
8. A multi-dimensional quantification index evaluation device for multimedia privacy protection, characterized by, The device is applied to a terminal device to be subjected to multimedia privacy protection, and the device comprises: an acquisition module configured to acquire a source image and an encrypted image; wherein the source image comprises a privacy region and a non-privacy region, and the encrypted image is obtained by performing privacy processing on the privacy region of the source image; a determination module configured to determine a target texture similarity based on the source image and the encrypted image, wherein the target texture similarity is used to represent a similarity degree between a texture feature of the source image and a texture feature of the encrypted image; determine a target edge similarity based on the source image and the encrypted image, wherein the target edge similarity is used to represent a similarity degree between an edge feature of the source image and an edge feature of the encrypted image; determine a target semantic similarity based on the source image and the encrypted image, wherein the target semantic similarity is used to represent a similarity degree between a semantic feature of the source image and a semantic feature of the encrypted image; determine a privacy protection index based on the target texture similarity, the target edge similarity and the target semantic similarity; and if the privacy protection index satisfies a preset condition, determine that the encrypted image satisfies a privacy protection requirement, otherwise, determine that the encrypted image does not satisfy the privacy protection requirement; wherein when determining the target texture similarity based on the source image and the encrypted image, the determination module is specifically configured to: determining a first gradient direction and a first gradient magnitude based on the non-private region of the source image, determining a second gradient direction and a second gradient magnitude based on the non-private region of the encrypted image; determining a direction similarity of the source image and the encrypted image based on the first gradient direction and the second gradient direction; determining a magnitude similarity of the source image and the encrypted image based on the first gradient magnitude, the second gradient magnitude, the non-private region of the source image and the non-private region of the encrypted image; determining a first texture similarity of the non-private region based on the direction similarity and the magnitude similarity; determining a third gradient direction based on the private region of the source image, determining a fourth gradient direction based on the private region of the encrypted image; determining a second texture similarity of the private region based on the third gradient direction, the fourth gradient direction and the total number of pixels in the private region; determining a target texture similarity based on the first texture similarity and the second texture similarity.
9. An electronic device, comprising: comprising: a processor and a machine readable storage medium storing machine executable instructions executable by the processor; the processor is configured to execute the machine executable instructions to implement the method of any one of claims 1-7.
Citation Information
Patent Citations
Image perception visual safety evaluation method based on edge and texture similarity
CN117336414A