Multidimensional quantitative index evaluation method and device for multimedia privacy protection

By constructing a multi-dimensional quantitative indicator system to evaluate the privacy protection effect of multimedia data, the problem that the existing technology cannot evaluate the privacy protection needs of multimedia data after desensitization is solved, and a scientific and objective privacy protection evaluation and data security assurance are achieved.

CN120705916AActive Publication Date: 2025-09-26HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511134162.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-09-26
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing technologies cannot effectively evaluate whether multimedia data meets privacy protection requirements after desensitization, resulting in an increased risk of privacy leakage.

Method used

By constructing a multidimensional quantitative indicator system, the texture similarity, edge similarity and semantic similarity between the source image and the encrypted image are evaluated, the privacy protection index is determined, and it is judged whether the encrypted image meets the privacy protection requirements.

Benefits of technology

Scientifically and objectively evaluate the privacy protection effect of multimedia data, reduce the risk of privacy leakage, and ensure data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705916A_ABST
    Figure CN120705916A_ABST
Patent Text Reader

Abstract

The invention provides a multi-dimensional quantitative index evaluation method and device for multimedia privacy protection. The method comprises the following steps: acquiring a source image and an encrypted image; determining a target texture similarity based on the source image and the encrypted image; determining target edge similarity based on the source image and the encrypted image; determining a target semantic similarity based on the source image and the encrypted image; determining a privacy protection index based on the target texture similarity, the target edge similarity and the target semantic similarity; and if the privacy protection index satisfies a preset condition, determining that the encrypted image satisfies a privacy protection demand, otherwise, determining that the encrypted image does not satisfy the privacy protection demand. Through the technical scheme of the invention, the personal privacy leakage risk can be reduced, the data security is protected, and the data leakage risk is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information security technology, and in particular to a multi-dimensional quantitative indicator evaluation method and device for multimedia privacy protection. Background Art

[0002] In the digital age, multimedia data (such as images and videos) is increasing in volume, becoming a crucial medium for information exchange. However, as data sharing and dissemination become increasingly convenient, the risk of personal privacy breaches also increases. To protect individual privacy, various privacy protection technologies have emerged, such as data encryption, anonymization, image blurring, and data desensitization. For example, data desensitization is a data security technology that uses pre-defined rules and algorithms to deconstruct sensitive information contained in multimedia data to protect private data. The primary purpose of data desensitization is to ensure the privacy and security of sensitive information without compromising the data's usefulness. Desensitizing multimedia data can protect data security and effectively reduce the risk of data breaches.

[0003] However, when multimedia data is desensitized to obtain desensitized data, it is impossible to determine whether the desensitized data meets the privacy protection requirements. If the desensitized data does not meet the privacy protection requirements, then when the desensitized data is transmitted, personal privacy will still be leaked and data security cannot be protected. Summary of the Invention

[0004] This application provides a multi-dimensional quantitative index evaluation method for multimedia privacy protection, which is applied to a terminal device to be subjected to multimedia privacy protection. The method includes: Acquire a source image and an encrypted image; wherein the source image includes a privacy area and a non-privacy area, and the encrypted image is obtained by performing privacy processing on the privacy area of ​​the source image; Determining a target texture similarity based on the source image and the encrypted image, wherein the target texture similarity is used to represent a degree of similarity between texture features of the source image and texture features of the encrypted image; determining a target edge similarity based on the source image and the encrypted image, wherein the target edge similarity is used to represent a degree of similarity between edge features of the source image and edge features of the encrypted image; determining a target semantic similarity based on the source image and the encrypted image, the target semantic similarity being used to characterize a degree of similarity between semantic features of the source image and semantic features of the encrypted image; A privacy protection index is determined based on the target texture similarity, the target edge similarity, and the target semantic similarity; if the privacy protection index satisfies a preset condition, it is determined that the encrypted image meets the privacy protection requirement; otherwise, it is determined that the encrypted image does not meet the privacy protection requirement.

[0005] The present application provides a multi-dimensional quantitative index evaluation device for multimedia privacy protection, which is applied to a terminal device to be subjected to multimedia privacy protection. The device includes: An acquisition module, configured to acquire a source image and an encrypted image; wherein the source image includes a privacy area and a non-privacy area, and the encrypted image is obtained by performing privacy processing on the privacy area of ​​the source image; A determination module is configured to determine a target texture similarity based on the source image and the encrypted image, the target texture similarity being used to characterize the degree of similarity between texture features of the source image and those of the encrypted image; determine a target edge similarity based on the source image and the encrypted image, the target edge similarity being used to characterize the degree of similarity between edge features of the source image and those of the encrypted image; determine a target semantic similarity based on the source image and the encrypted image, the target semantic similarity being used to characterize the degree of similarity between semantic features of the source image and those of the encrypted image; and determine a privacy protection index based on the target texture similarity, the target edge similarity, and the target semantic similarity; if the privacy protection index satisfies a preset condition, it is determined that the encrypted image meets the privacy protection requirement; otherwise, it is determined that the encrypted image does not meet the privacy protection requirement.

[0006] The present application provides an electronic device, comprising: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement a multidimensional quantitative indicator evaluation method for multimedia privacy protection.

[0007] The present application provides a computer program product, including a computer program, which, when executed by a processor, implements a multi-dimensional quantitative indicator evaluation method for multimedia privacy protection.

[0008] The present application provides a machine-readable storage medium, which stores machine-executable instructions that can be executed by a processor; wherein the processor is used to execute the machine-executable instructions to implement a multi-dimensional quantitative indicator evaluation method for multimedia privacy protection.

[0009] As can be seen from the above technical solution, in the embodiments of the present application, a privacy protection index is determined based on the target texture similarity, target edge similarity, and target semantic similarity between the source image and the encrypted image (e.g., a desensitized image), and then the privacy protection index is used to evaluate whether the encrypted image meets the privacy protection requirements. If the privacy protection index meets the preset conditions, it indicates that the encrypted image meets the privacy protection requirements. In this way, when the encrypted image is transmitted, the risk of personal privacy leakage is reduced, data security is protected, and the risk of data leakage is effectively reduced. If the privacy protection index does not meet the preset conditions, it indicates that the encrypted image does not meet the privacy protection requirements. The encrypted image that does not meet the privacy protection requirements is not transmitted, thereby reducing the risk of personal privacy leakage and protecting data security.

[0010] In the above method, by constructing multiple visual quantitative dimensions (such as texture, edge, semantics, etc.), a multi-dimensional quantitative indicator system is designed, and visual multi-dimensional quantitative indicators are given, thereby providing a scientific and objective evaluation method for the privacy protection effect in multimedia data. That is, it can scientifically and objectively evaluate the privacy protection effect in multimedia data, provide an objective evaluation basis for the optimization and improvement of privacy protection technologies such as encryption, blurring, and desensitization, and can more accurately evaluate the privacy protection effect of encrypted images. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a flow chart of a multi-dimensional quantitative indicator evaluation method for multimedia privacy protection; Figure 2 It is a flow chart of a multi-dimensional quantitative indicator evaluation method for multimedia privacy protection; Figure 3 is a schematic diagram of a visual multidimensional quantitative index in one embodiment of the present application; Figure 4 It is a structural diagram of a multi-dimensional quantitative index evaluation device for multimedia privacy protection; Figure 5 It is a hardware structure diagram of an electronic device in one embodiment of the present application. DETAILED DESCRIPTION

[0012] In the embodiment of the present application, a multi-dimensional quantitative index evaluation method for multimedia privacy protection is proposed. The method can be applied to terminal devices (such as Internet of Things devices, etc.) to be subjected to multimedia privacy protection. Figure 1 FIG. 5 is a flow chart of the method, which may include: Step 101: Acquire a source image and an encrypted image. The source image may include a privacy area and a non-privacy area, and the encrypted image is obtained by performing privacy processing on the privacy area of ​​the source image.

[0013] Step 102: Determine target texture similarity based on the source image and the encrypted image. The target texture similarity is used to represent the degree of similarity between texture features of the source image and texture features of the encrypted image.

[0014] Step 103: Determine target edge similarity based on the source image and the encrypted image. The target edge similarity is used to represent the degree of similarity between edge features of the source image and edge features of the encrypted image.

[0015] Step 104 : determining a target semantic similarity based on the source image and the encrypted image. The target semantic similarity is used to represent the degree of similarity between the semantic features of the source image and the semantic features of the encrypted image.

[0016] Step 105: Determine a privacy protection index based on the target texture similarity, the target edge similarity, and the target semantic similarity; if the privacy protection index meets the preset conditions, determine that the encrypted image meets the privacy protection requirements; if the privacy protection index does not meet the preset conditions, determine that the encrypted image does not meet the privacy protection requirements.

[0017] Exemplarily, determining a target texture similarity based on a source image and an encrypted image includes: determining a first gradient direction and a first gradient magnitude based on a non-privacy area of ​​the source image, and determining a second gradient direction and a second gradient magnitude based on the non-privacy area of ​​the encrypted image; determining a directional similarity between the source image and the encrypted image based on the first gradient direction and the second gradient direction; determining an amplitude similarity between the source image and the encrypted image based on the first gradient magnitude, the second gradient magnitude, the non-privacy area of ​​the source image, and the non-privacy area of ​​the encrypted image; determining a first texture similarity of the non-privacy area based on the directional similarity and the amplitude similarity; determining a third gradient direction based on the privacy area of ​​the source image, and determining a fourth gradient direction based on the privacy area of ​​the encrypted image; determining a second texture similarity of the privacy area based on the third gradient direction, the fourth gradient direction, and the total number of pixels in the privacy area; and determining a target texture similarity based on the first texture similarity and the second texture similarity.

[0018] Exemplarily, the process of determining the first gradient direction, the first gradient magnitude, the second gradient direction, the second gradient magnitude, the third gradient direction, and the fourth gradient direction may include: obtaining a horizontal gradient operator and a vertical gradient operator, wherein the horizontal gradient operator and the vertical gradient operator each include K gradient parameters, where K is a positive integer; wherein the K gradient parameters are parameters within the target network model and are obtained by optimizing the target network model; determining the horizontal gradient of each pixel in the source image based on the horizontal gradient operator, and determining the vertical gradient of each pixel in the source image based on the vertical gradient operator; and determining the vertical gradient of each pixel in the non-privacy area of ​​the source image based on the horizontal gradient operator. A first gradient direction and a first gradient amplitude are determined based on the horizontal gradient and the vertical gradient of the pixel point; a third gradient direction is determined based on the horizontal gradient and the vertical gradient of each pixel point in the privacy area of ​​the source image; the horizontal gradient of each pixel point in the encrypted image is determined based on the horizontal gradient operator, and the vertical gradient of each pixel point in the encrypted image is determined based on the vertical gradient operator; a second gradient direction and a second gradient amplitude are determined based on the horizontal gradient and the vertical gradient of each pixel point in the non-privacy area of ​​the encrypted image; and a fourth gradient direction is determined based on the horizontal gradient and the vertical gradient of each pixel point in the privacy area of ​​the encrypted image.

[0019] Exemplarily, the process of obtaining K gradient parameters specifically includes: inputting a sample image into a target network model, and the target network model outputting a predicted sensitive probability of each first pixel in the sample image and a predicted gradient direction of each second pixel in the privacy area of ​​the sample image, where the predicted sensitive probability represents the probability that the first pixel is a sensitive pixel; determining a privacy-sensitive area segmentation loss value based on the predicted sensitive probability of each first pixel in the sample image and the true sensitive probability of each first pixel in the sample image; determining a gradient direction consistency loss value based on the predicted gradient direction of each second pixel in the privacy area and the true gradient direction of each second pixel in the privacy area; determining a target loss value based on the privacy-sensitive area segmentation loss value and the gradient direction consistency loss value, and optimizing the target network model based on the target loss value to obtain an optimized model; if the optimized model has not converged, updating the optimized model to the target network model, and returning to execute the operation of inputting the sample image into the target network model; if the optimized model has converged, obtaining K gradient parameters from the network parameters of the optimized model.

[0020] Exemplarily, determining a target edge similarity based on a source image and an encrypted image may include: obtaining a threshold set, the threshold set including multiple dynamic thresholds; for each dynamic threshold, determining a first edge binary image corresponding to the source image based on the dynamic threshold, and determining a second edge binary image corresponding to the encrypted image based on the dynamic threshold; a first value in the first edge binary image represents an edge point, and a second value represents a non-edge point; obtaining a common edge image; wherein, for each pixel in the common edge image, if a logical AND operation result of a pixel value corresponding to the pixel in the first edge binary image and a pixel value corresponding to the pixel in the second edge binary image is the first value, and an absolute value of a difference between the first pixel value corresponding to the pixel in the source image and the second pixel value corresponding to the pixel in the encrypted image is not greater than a brightness control threshold, then determining a pixel value corresponding to the pixel in the common edge image based on the first pixel value, the second pixel value, and the brightness control threshold; otherwise, determining a pixel value corresponding to the pixel in the common edge image as the second value; determining an edge similarity corresponding to the dynamic threshold based on the common edge image and the first edge binary image, and weighting the edge similarity corresponding to each dynamic threshold to obtain a target edge similarity.

[0021] Exemplarily, obtaining a threshold set may include, but is not limited to: determining a gradient magnitude of a source image based on the horizontal and vertical gradients of the source image, and normalizing the gradient magnitude of the source image to obtain a first target gradient magnitude; generating a first gradient histogram corresponding to the source image based on the first target gradient magnitude; wherein the first gradient histogram may include the gradient magnitude distribution of each pixel in the source image; determining a gradient magnitude of an encrypted image based on the horizontal and vertical gradients of the encrypted image, and normalizing the gradient magnitude of the encrypted image to obtain a second target gradient magnitude; and generating a second gradient histogram corresponding to the encrypted image based on the second target gradient magnitude; wherein the second gradient histogram may include the gradient magnitude distribution of each pixel in the encrypted image. On this basis, a joint gradient histogram is generated based on the first gradient histogram and the second gradient histogram, and a lower threshold value and an upper threshold value are determined based on the joint gradient histogram; N uniformly distributed dynamic thresholds are generated based on the lower threshold value and the upper threshold value, and the threshold set includes N dynamic thresholds; wherein N can be a positive integer greater than 1, N can be pre-configured, or N can be determined based on the number of valid peaks of the joint gradient histogram.

[0022] Exemplarily, determining the target semantic similarity based on the source image and the encrypted image may include: determining the object semantic similarity based on the overlap between the detection frame of the source image and the detection frame of the encrypted image; the detection frame of the source image is the area of ​​each identified object in the source image, and the detection frame of the encrypted image is the area of ​​each identified object in the encrypted image; determining the relational semantic similarity based on the semantic association graph of the source image and the semantic association graph of the encrypted image; the semantic association graph of the source image represents the association relationship between each identified object in the source image, and the semantic association graph of the encrypted image represents the association relationship between each identified object in the encrypted image; determining the scene semantic similarity based on the scene feature vector of the source image and the scene feature vector of the encrypted image; wherein, the source image is input into a feature extraction network to obtain the scene feature vector of the source image, and the encrypted image is input into a feature extraction network to obtain the scene feature vector of the encrypted image; performing weighted operations on the object semantic similarity, the relational semantic similarity and the scene semantic similarity to obtain a semantic risk value, and determining the target semantic similarity based on the semantic risk value.

[0023] Exemplarily, determining a privacy protection indicator based on target texture similarity, target edge similarity, and target semantic similarity may include: determining a first standard deviation based on target texture similarity, determining a first correlation coefficient based on target texture similarity and a pre-configured true leakage risk label, and determining a first weighting coefficient corresponding to the target texture similarity based on the first standard deviation and the first correlation coefficient; determining a second standard deviation based on target edge similarity, determining a second correlation coefficient based on target edge similarity and the true leakage risk label, and determining a second weighting coefficient corresponding to the target edge similarity based on the second standard deviation and the second correlation coefficient; determining a third standard deviation based on target semantic similarity, determining a third correlation coefficient based on target semantic similarity and the true leakage risk label, and determining a third weighting coefficient corresponding to the target semantic similarity based on the third standard deviation and the third correlation coefficient; and weighting the target texture similarity, target edge similarity, and target semantic similarity based on the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient to obtain a privacy protection indicator.

[0024] As can be seen from the above technical solution, in the embodiments of the present application, a privacy protection index is determined based on the target texture similarity, target edge similarity, and target semantic similarity between the source image and the encrypted image (e.g., a desensitized image), and then the privacy protection index is used to evaluate whether the encrypted image meets the privacy protection requirements. If the privacy protection index meets the preset conditions, it indicates that the encrypted image meets the privacy protection requirements. In this way, when the encrypted image is transmitted, the risk of personal privacy leakage is reduced, data security is protected, and the risk of data leakage is effectively reduced. If the privacy protection index does not meet the preset conditions, it indicates that the encrypted image does not meet the privacy protection requirements. The encrypted image that does not meet the privacy protection requirements is not transmitted, thereby reducing the risk of personal privacy leakage and protecting data security.

[0025] In the above method, by constructing multiple visual quantitative dimensions (such as texture, edge, semantics, etc.), a multi-dimensional quantitative indicator system is designed, and visual multi-dimensional quantitative indicators are given, thereby providing a scientific and objective evaluation method for the privacy protection effect in multimedia data. That is, it can scientifically and objectively evaluate the privacy protection effect in multimedia data, provide an objective evaluation basis for the optimization and improvement of privacy protection technologies such as encryption, blurring, and desensitization, and can more accurately evaluate the privacy protection effect of encrypted images.

[0026] The above technical solutions of the embodiments of the present application are described below in conjunction with specific application scenarios.

[0027] To protect individual privacy, various privacy-preserving technologies have emerged, including data encryption, anonymization, image blurring, and data desensitization. However, when privacy-preserving the source image to generate an encrypted image, it is unclear whether the encrypted image meets privacy protection requirements. If the encrypted image does not meet privacy protection requirements, then transmitting the encrypted image will still result in personal privacy leaks and fail to ensure data security.

[0028] In response to the above findings, the present application proposes a multidimensional quantitative index evaluation method for multimedia privacy protection in an embodiment. This method aims to provide a scientific and objective means of evaluating the privacy protection effectiveness of multimedia data (such as images and videos, with images being used as an example below). By constructing multiple visual quantitative dimensions (such as texture, edges, and semantics), starting from the visual features of the image and combining the processing effect of privacy protection technology on the image, a multidimensional quantitative index system is designed. This method can more accurately evaluate the privacy protection effectiveness of encrypted images, scientifically and objectively evaluate the privacy protection effect, and conduct a multi-dimensional quantitative analysis of the privacy protection effect, providing a scientific basis and objective standards for the evaluation and improvement of multimedia privacy protection technology. This method is applicable to highly privacy-sensitive scenarios such as smart cities, medical imaging, and autonomous driving.

[0029] In the embodiment of this application, a multi-dimensional quantitative index evaluation method for multimedia privacy protection is proposed. Figure 2 FIG. 5 is a flow chart of the method, which may include: Step 201: Acquire a source image, where the source image is an original image that needs to be privacy protected.

[0030] For example, a video stream from a camera can be collected, and each frame of the video stream can be used as a source image, or a portion of the images in the video stream can be used as a source image. Furthermore, a video stream can be obtained from a storage device, and each frame of the video stream can be used as a source image, or a portion of the images in the video stream can be used as a source image. Furthermore, a real-time transport stream can be received, and each frame of the real-time transport stream can be used as a source image, or a portion of the images in the real-time transport stream can be used as a source image.

[0031] Exemplarily, the source image may be an RGB image, a grayscale image, a video frame sequence, or a multispectral image. This embodiment does not limit the type of the source image, and the source image may be in any format.

[0032] Step 202: Obtain an encrypted image corresponding to the source image; wherein the source image may include a privacy area and a non-privacy area, and the encrypted image is obtained by performing privacy processing on the privacy area of ​​the source image.

[0033] For example, a data encryption algorithm can be used to perform privacy processing (i.e., encryption) on the private region of a source image to obtain an encrypted image. An anonymization algorithm can be used to perform privacy processing (i.e., anonymization) on the private region of a source image to obtain an encrypted image. An image blurring algorithm can be used to perform privacy processing (i.e., image blurring) on ​​the private region of a source image to obtain an encrypted image. A data desensitization algorithm can be used to perform privacy processing (i.e., data desensitization) on the private region of a source image to obtain an encrypted image.

[0034] For example, a multimodal model can be used to identify sensitive content, such as faces, license plates, and text, and combined with data cleaning techniques to dynamically screen risk areas to obtain the privacy region of the source image. Based on this, privacy processing (such as privacy protection processing) can be performed on the privacy region of the source image to produce an encrypted image.

[0035] In a possible implementation, the encrypted image corresponding to the source image may be obtained in the following manner. The following manner is merely an example, and the encrypted image may be obtained by performing privacy processing on the privacy region of the source image.

[0036] After obtaining the source image, you can check the integrity of the source image, for example, to see if there is any file header corruption or pixel value overflow. If so, the process ends and the source image is not processed. Otherwise, the source image is processed.

[0037] After obtaining the source image, it can be de-labeled. This means removing metadata from the source image if it contains any. For example, sensitive fields such as GPS (Global Positioning System) coordinates, capture time, device serial number, and other private information can be removed.

[0038] After obtaining the source image, visual content recognition can be performed on the source image. Specifically, the multimodal model can be used to identify the privacy region of the source image. This privacy region can include at least one detection frame (e.g., a detection frame of an identified object within the source image. The identified object can be a face, license plate, ID card, etc., and there are no restrictions on the identified object; it can be any sensitive object). In other words, the multimodal model detects whether the identified object exists within the source image and outputs a detection frame (e.g., a rectangular detection frame) of the identified object. These detection frames of the identified object serve as the privacy region of the source image, thereby obtaining the privacy region of the source image.

[0039] For example, the multimodal model can be a detection model, such as the Faster R-CNN detection model, or any other model. This is not a limitation as long as it can detect the identified object within the source image. This embodiment does not impose any limitations on the structure and training process of the detection model. The source image can be input into the detection model, which detects whether the identified object exists within the source image and outputs the detection frame (e.g., detection frame coordinates), category, and confidence level of the identified object. The category indicates the category of the identified object to which the detection frame corresponds, such as a facial object, license plate object, or ID card object. The confidence level indicates the probability that the detection frame corresponds to the identified object.

[0040] For example, after the detection model outputs the detection box of the identified object, NMS (non maximum suppression) can be used to eliminate overlapping boxes and retain the detection boxes with high confidence.

[0041] For example, before feeding the source image into the detection model, the source image can be preprocessed and then fed into the detection model. For example, to standardize the resolution, the source image can be scaled to a fixed size (e.g., 512*512) and then fed into the detection model with the fixed-size source image.

[0042] In summary, the privacy region of the source image (i.e., the area where the sensitive object is located) can be determined, while the remaining areas outside the privacy region can be considered the non-private region of the source image. Privacy processing can be performed on the privacy region of the source image. For example, a Gaussian blur can be applied to the privacy region, with the kernel size adaptively adjusted to the size of the sensitive object. For example, a GAN (Generative Adversarial Network) can be used to replace the texture of the privacy region. For example, feature replacement can be performed on the privacy region, such as replacing a real face with a virtual one. Of course, the above are just a few examples, and there is no limitation on the privacy processing method. As long as the privacy of the privacy region is protected, it is sufficient.

[0043] For the non-privacy area of ​​the source image, the non-privacy area can be kept unchanged, or Gaussian noise can be added to the non-privacy area to destroy the global statistical characteristics of the non-privacy area. There is no restriction on this non-privacy area.

[0044] After the source image has been processed as described above, the processed source image (i.e., the privacy-sensitive areas have been processed, while the non-privacy areas remain unchanged or have Gaussian noise added) can be synthesized (i.e., fused) with the source image, and the fused image serves as the encrypted image corresponding to the source image. Obviously, the source image may include both privacy-sensitive and non-privacy-sensitive areas, and the encrypted image is obtained by performing privacy processing on the privacy-sensitive areas of the source image.

[0045] Step 203: Determine target texture similarity based on the source image and the encrypted image. The target texture similarity is used to characterize the degree of similarity between the texture features of the source image and the texture features of the encrypted image. For example, the target texture similarity can be texture distribution similarity, which can quantify the risk of explicit privacy leakage.

[0046] For example, given a source image S and an encrypted image C, the similarity between the texture features of the source image S and the texture features of the encrypted image C can be determined, thereby obtaining a target texture similarity. The method for determining the target texture similarity is not limited, as long as it can characterize the similarity between the texture features. In one possible implementation, the target texture similarity can be determined using the following steps: Step S11: Obtain a horizontal gradient operator and a vertical gradient operator.

[0047] In a possible implementation, the horizontal gradient operator can be configured according to actual needs, and the vertical gradient operator can be configured according to actual needs. For example, the horizontal gradient operator It can be: , vertical gradient operator It can be: Of course, the above are just examples of horizontal gradient operators and vertical gradient operators, and there is no limitation to this.

[0048] In one possible implementation, the horizontal gradient operator and the vertical gradient operator each include K gradient parameters, where K is a positive integer. The K gradient parameters are parameters within the target network model and are obtained by optimizing the target network model. Based on this, the horizontal gradient operator and the vertical gradient operator can be trainable gradient operators for detecting privacy-sensitive areas, so that the horizontal gradient operator and the vertical gradient operator automatically learn weights that are more sensitive to privacy features. The horizontal gradient operator and the vertical gradient operator change with training and can be optimized for specific tasks, dynamically learning the optimal edge detection mode. The horizontal gradient operator and the vertical gradient operator continuously learn based on the descent of the loss function.

[0049] For example, in order to enhance the contribution of the current pixel row and improve noise resistance, the horizontal gradient operator and the vertical gradient operator can be designed as an antisymmetric structure (such as symmetry of positive and negative weights) to ensure the zero mean of the gradient operator. When the K gradient parameters are 3 gradient parameters, the 3 gradient parameters include the parameter w 1 、 w 2 、 w 3 For example, the horizontal gradient operator can include , the vertical gradient operator can include . When K gradient parameters are 4 gradient parameters ( w 1 、 w 2 、 w 3 、 w 4 ), for the horizontal gradient operator, the first column is w 1 、 w 2 、 w 3 、 w 4 , the last column is - w 1 、 -w 2 、 -w 3 、 -w 4 , the other two columns are all 0, for the vertical gradient operator, the first line w 1 、 w 2 、 w 3 、 w 4 , the last line is - w 1 、 -w 2 、 -w 3 、 -w 4 , and the other two rows are all 0. When K gradient parameters are other numbers, the horizontal gradient operator and the vertical gradient operator can be designed similarly.

[0050] Exemplarily, the process of obtaining K gradient parameters may include but is not limited to: Obtain a target network model, which is used to output the predicted sensitivity probability of the pixel point and the predicted gradient direction of the pixel point. There is no restriction on the network structure of this target network model, as long as the target network model can output the predicted sensitivity probability of the pixel point and the predicted gradient direction of the pixel point.

[0051] In the target network model, multiple network parameters can be included, and K network parameters among all network parameters can be used as K gradient parameters (i.e., gradient parameters to be optimized, such as w 1 、 w 2 、 w 3 ). Regarding which network parameters are used as gradient parameters, there is no restriction in this embodiment. K network parameters can be selected as K gradient parameters according to actual needs. In this way, the K gradient parameters can be parameters in the target network model, and the K gradient parameters are obtained by optimizing the target network model.

[0052] Obtain a sample image (which can be multiple sample images, one sample image is used as an example here) and label data corresponding to the sample image. For the sample image, the sample image can be any image with a privacy area, and there is no restriction on this. For the label data corresponding to the sample image, the label data can include the true sensitivity probability of each first pixel point in the sample image (each pixel point in the sample image is referred to as a first pixel point). The true sensitivity probability indicates the probability that the first pixel point is a sensitive pixel. For example, the true sensitivity probability can be 1 or 0, 1 indicates that the first pixel point is a sensitive pixel, that is, the first pixel point belongs to the privacy area, and 0 indicates that the first pixel point is not a sensitive pixel, that is, the first pixel point does not belong to the privacy area.

[0053] The label data may also include the true gradient direction of each second pixel point in the privacy area of ​​the sample image (each pixel point in the privacy area of ​​the sample image is referred to as a second pixel point). The true gradient direction is the direction in which the function value at the second pixel point increases fastest, and there is no restriction on this.

[0054] For the privacy area of ​​the sample image, the privacy area can be an area where a sensitive object is located (such as a rectangular area), and the sensitive object can be a facial object, a license plate object, an ID card object, etc., without any restriction.

[0055] The label data corresponding to the sample image can be manually annotated by the user or obtained using a certain algorithm. There is no restriction on the source of the label data.

[0056] After obtaining the target network model and sample images (such as multiple sample images), the sample images can be input into the target network model, and the target network model outputs the predicted sensitivity probability of each first pixel point in the sample image and the predicted gradient direction of each second pixel point in the privacy area of ​​the sample image.

[0057] For example, since the target network model is used to output the predicted sensitivity probability of the pixel and the predicted gradient direction of the pixel, after the sample image is input to the target network model, the target network model can process the sample image based on the sample image to obtain the predicted sensitivity probability of each first pixel in the sample image. There is no restriction on this processing process, as long as the predicted sensitivity probability can be obtained. The predicted sensitivity probability can represent the probability that the first pixel is a sensitive pixel, and the predicted sensitivity probability is a probability value between 0 and 1. For example, when the predicted sensitivity probability is 0.9, it means that the probability that the first pixel is a sensitive pixel is 90%, and when the predicted sensitivity probability is 1, it means that the probability that the first pixel is a sensitive pixel is 100%.

[0058] After the sample image is input into the target network model, the target network model can process the sample image to obtain the predicted gradient direction of each second pixel point in the privacy area of ​​the sample image. That is, the target network model first determines the privacy area of ​​the sample image, and then determines the predicted gradient direction of each second pixel point in the privacy area. There is no restriction on this processing process, as long as the predicted gradient direction can be obtained.

[0059] For example, a privacy-sensitive region segmentation loss value can be determined based on the predicted sensitivity probability (i.e., predicted data) and the actual sensitivity probability (i.e., labeled data) of each first pixel in the sample image. The privacy-sensitive region segmentation loss value is used to align the privacy-sensitive gradient map output by the target network model with the true labeled privacy region mask. For example, a weighted cross-entropy loss function can be used to determine the privacy-sensitive region segmentation loss value to address class imbalance. Other loss functions can also be used to determine the privacy-sensitive region segmentation loss value, and there is no limitation on this.

[0060] For example, the following formula can be used to determine the segmentation loss value of the privacy-sensitive area L1 :

[0061] In the above formula, N Indicates the total number of pixels in the sample image, i The sample image i The first pixel. Represents the weight coefficient, which is used to emphasize the penalty of the privacy area and can be configured according to actual needs. The sample image can be represented byi The true sensitivity probability of the first pixel, The sample image can be represented by i The predicted sensitivity probability of the first pixel point is 0 or 1, and the value of the actual sensitivity probability can be between 0 and 1, that is, in the range of 0~1. Represents the weight coefficient, which is used to emphasize the penalty of the privacy area and can be configured according to actual needs.

[0062] For example, a gradient directional consistency loss value can be determined based on the predicted gradient direction (i.e., predicted data) of each second pixel point within the privacy region of the sample image and the actual gradient direction (i.e., label data) of each second pixel point within the privacy region of the sample image. The gradient directional consistency loss value is used to maintain the directional selectivity of the gradient operator, enhance robustness, and avoid loss of texture details after training. The gradient directional consistency loss value only considers the privacy region of the sample image and does not consider the non-privacy region of the sample image. For example, the gradient directional consistency loss value can be determined using a directional difference loss function, or other loss functions can be used to determine the gradient directional consistency loss value, and there is no restriction on this loss function.

[0063] For example, the gradient direction consistency loss value can be determined using the following formula: L2 :

[0064] In the above formula, M The total number of pixels in the privacy area of ​​the sample image can be represented by M A set of pixels that can represent the privacy area of ​​the sample image. j The first j The second pixel. Can indicate the first j The predicted gradient direction of the second pixel, Can indicate the first j The true gradient direction of the second pixel.

[0065] After obtaining the privacy-sensitive area segmentation loss value and the gradient direction consistency loss value, a target loss value can be determined based on the privacy-sensitive area segmentation loss value and the gradient direction consistency loss value. For example, the sum of the privacy-sensitive area segmentation loss value and the gradient direction consistency loss value is used as the target loss value, or the privacy-sensitive area segmentation loss value and the gradient direction consistency loss value are weighted to obtain the target loss value.

[0066] For example, after obtaining the target loss value, the target network model can be optimized based on the target loss value to obtain an optimized model. For example, the target network model can be optimized using a gradient descent algorithm or other methods. There is no restriction on this optimization method. The optimization goal is to make the target loss value smaller and smaller.

[0067] Then, it can be determined whether the optimized model has converged. For example, if the target loss value is less than a preset threshold, the optimized model has converged; if the target loss value is not less than the preset threshold, the optimized model has not converged. For another example, if the number of iterations reaches a threshold, the optimized model has converged; if the number of iterations does not reach the threshold, the optimized model has not converged. For another example, if the iteration duration reaches a threshold, the optimized model has converged; if the iteration duration does not reach the threshold, the optimized model has not converged.

[0068] If the optimized model has not converged, the optimized model is updated to the target network model, and the operation of inputting the sample image to the target network model is returned. If the optimized model has converged, K gradient parameters are obtained from the network parameters of the optimized model. For example, if the optimized model includes multiple network parameters, K network parameters from all network parameters can be used as K gradient parameters (such as w 1 、 w 2 、 w 3 ).

[0069] After obtaining K gradient parameters, a horizontal gradient operator and a vertical gradient operator can be constructed based on the K gradient parameters. The construction method can be found in the above description and will not be repeated here.

[0070] Step S12: determining the horizontal gradient of each pixel in the source image based on the horizontal gradient operator, and determining the vertical gradient of each pixel in the source image based on the vertical gradient operator.

[0071] Furthermore, the horizontal gradient of each pixel in the encrypted image is determined based on the horizontal gradient operator, and the vertical gradient of each pixel in the encrypted image is determined based on the vertical gradient operator.

[0072] For example, the horizontal gradient can be determined using the following formula: In the above formula, represents the horizontal gradient operator, represents the convolution operator, if Indicates the source image i pixels, then Indicates the source image i The horizontal gradient of pixels, if Indicates the encrypted image i pixels, then Indicates the encrypted image i The horizontal gradient of a pixel.

[0073] For example, the vertical gradient can be determined using the following formula: In the above formula, It can represent the vertical gradient operator, if Indicates the source image i pixels, then It can represent the first i The vertical gradient of a pixel point, if Indicates the encrypted image i pixels, then It can represent the encrypted image i The vertical gradient of each pixel.

[0074] Step S13: Determine a first gradient direction and a first gradient magnitude based on the horizontal gradient and the vertical gradient of each pixel in the non-privacy area of ​​the source image; and determine a second gradient direction and a second gradient magnitude based on the horizontal gradient and the vertical gradient of each pixel in the non-privacy area of ​​the encrypted image.

[0075] Exemplarily, for each pixel in the non-privacy area of ​​the source image, the first gradient direction and the first gradient magnitude of the pixel are determined based on the horizontal gradient and the vertical gradient of the pixel. For example, the first gradient direction of the pixel is determined using the following formula: , use the following formula to determine the first gradient amplitude of the pixel: . Indicates the first i The first gradient direction of the pixel point, Indicates the first i The first gradient amplitude of the pixel point, Indicates the first i The vertical gradient of each pixel, Indicates the first i The horizontal gradient of each pixel.

[0076] For example, in the above formula, if Indicates the first i The vertical gradient of pixels, and Indicates the first i The horizontal gradient of pixels is It can represent the first i The second gradient direction of the pixel point, It can represent the first i The second gradient amplitude of the pixel point.

[0077] Step S14: Determine a third gradient direction (i.e., the gradient direction of the privacy area) based on the horizontal gradient and the vertical gradient of each pixel in the privacy area of ​​the source image; and determine a fourth gradient direction based on the horizontal gradient and the vertical gradient of each pixel in the privacy area of ​​the encrypted image.

[0078] For example, for each pixel in the privacy area of ​​the source image, the third gradient direction of the pixel is determined based on the horizontal gradient and the vertical gradient of the pixel. For example, the third gradient direction of the pixel is determined using the following formula: . Indicates the privacy area of ​​the source image. j The third gradient direction of the pixel point, Indicates the privacy area of ​​the source image. j The vertical gradient of each pixel, Indicates the privacy area of ​​the source image. j The horizontal gradient of each pixel.

[0079] For example, in the above formula, if Indicates the privacy area of ​​the encrypted image. j The vertical gradient of pixels, and Indicates the privacy area of ​​the encrypted image. j The horizontal gradient of pixels is Indicates the privacy area of ​​the encrypted image. j The fourth gradient direction of the pixel.

[0080] Step S15 : determining the directional similarity between the source image and the encrypted image based on the first gradient direction of the non-privacy area of ​​the source image and the second gradient direction of the non-privacy area of ​​the encrypted image.

[0081] For example, directional similarity can be determined based on the first gradient direction of each pixel in the non-privacy area of ​​the source image and the second gradient direction of each pixel in the non-privacy area of ​​the encrypted image. For example, the first gradient direction of each pixel can be converted into a first unit vector of each pixel, and the second gradient direction of each pixel can be converted into a second unit vector of each pixel. Then, directional similarity can be determined based on the first unit vector of each pixel and the second unit vector of each pixel.

[0082] For example, for the first i Pixel points can be converted into unit vectors using the following formula: , . Indicates the first i The first gradient direction of the pixel point, Indicates the first i The first unit vector of the pixel point, Indicates the first i The second gradient direction of the pixel point, Indicates the first i The second unit vector of pixels.

[0083] Then, the direction similarity can be determined using the following formula: In the above formula, Indicates the directional similarity between the source image S and the encrypted image C. The closer the directional similarity value is to 1, the more consistent the directional distribution between the source image S and the encrypted image C. N represents the total number of pixels in the non-privacy area (the non-privacy areas of the source image S and the encrypted image C are in the same position, and the privacy areas of the source image S and the encrypted image C are in the same position). i Indicates the first i pixels, i The value range is from 1 to N. Indicates the first i The first unit vector of the pixel point, Indicates the first i The second unit vector of pixels.

[0084] Step S16: Determine the amplitude similarity between the source image and the encrypted image based on the first gradient amplitude of the non-privacy area of ​​the source image, the second gradient amplitude of the non-privacy area of ​​the encrypted image, the pixel value of each pixel in the non-privacy area of ​​the source image, and the pixel value of each pixel in the non-privacy area of ​​the encrypted image.

[0085] For example, the amplitude similarity between the source image and the encrypted image can be determined based on the first gradient amplitude of each pixel in the non-privacy area of ​​the source image, the second gradient amplitude of each pixel in the non-privacy area of ​​the encrypted image, the first mean (i.e., the average value of the pixel values ​​of all pixels in the non-privacy area) and the first standard deviation (i.e., the standard deviation of the pixel values ​​of all pixels in the non-privacy area) of the source image, and the second mean and second standard deviation of the non-privacy area of ​​the encrypted image. For example, the amplitude similarity can be determined using the following formula: . Indicates the amplitude similarity between the source image and the encrypted image, represents the first mean, represents the second mean, represents the first standard deviation, represents two standard deviations. represents the average value of the first gradient amplitude of each pixel in the non-privacy area of ​​the source image, Represents the average value of the second gradient magnitude of each pixel in the non-privacy area of ​​the encrypted image. It is a constant configured according to actual needs. It can be greater than 0. It is used to prevent the value from being 0. This is a constant configured according to actual needs. It can be greater than 0 to prevent the value from being 0.

[0086] Step S17: determining a first texture similarity of the non-privacy area based on the direction similarity and the amplitude similarity.

[0087] For example, the direction similarity and amplitude similarity can be weighted to obtain the first texture similarity of the non-privacy area, that is, the comprehensive texture similarity under the direction similarity and amplitude similarity. For example, the following formula can be used to determine the first texture similarity : In the above formula, Indicates direction similarity The weighting coefficient can be configured according to actual needs. Indicates amplitude similarity The weighting coefficient can be configured according to the actual scenario requirements.

[0088] Step S18: determining a second texture similarity of the privacy area based on the third gradient direction of the privacy area of ​​the source image, the fourth gradient direction of the privacy area of ​​the encrypted image, and the total number of pixels in the privacy area.

[0089] For example, the second texture similarity of the privacy region can be determined based on the third gradient direction of each pixel in the privacy region of the source image, the fourth gradient direction of each pixel in the privacy region of the encrypted image, and the total number of pixels in the privacy region (the privacy regions of the source image and the encrypted image are located in the same position, i.e., the total number of pixels in the privacy regions is the same). For example, the second texture similarity can be determined using the following formula: In the above formula, It can represent the second texture similarity of the privacy area, and M can represent the total number of pixels in the privacy area, that is, M can represent the pixel set of the privacy area. j Indicates the first j pixels, jThe value range is 1 to M. It can represent the first j The third gradient direction of the pixel point, It can represent the first j Based on the above formula, if the privacy area of ​​the encrypted image is completely unrelated to the privacy area of ​​the source image, then .

[0090] Step S19: Determine a target texture similarity based on the first texture similarity and the second texture similarity.

[0091] For example, the first texture similarity and the second texture similarity can be weighted to obtain the target texture similarity between the source image and the encrypted image. For example, the target texture similarity can be determined using the following formula: : In the above formula, Can represent the second texture similarity The weighting coefficient can be configured according to actual needs. Can represent the first texture similarity The weighting coefficient can be configured according to the actual scenario requirements.

[0092] In summary, in this embodiment, the target texture similarity is determined based on the trainable horizontal gradient operator and the vertical gradient operator, so as to evaluate the privacy protection effect through the target texture similarity.

[0093] Considering the combination of directional consistency and local texture statistics, the privacy-protected image should maintain texture similarity with the source image in non-privacy-protected areas, while showing significant differences in privacy areas. Based on this, the target texture similarity can be determined based on the first texture similarity in the non-privacy area and the second texture similarity in the privacy area. For the privacy area, the degree of damage is taken into account, and the significance of the directional difference within the privacy area is calculated, and then the second texture similarity is determined. For the non-privacy area, the directional consistency of the non-privacy area is considered to ensure the naturalness of the texture structure, and the region separation calculation can accurately distinguish the protection effect.

[0094] At this point, step 203 is completed, and the target texture similarity between the source image and the encrypted image is obtained.

[0095] Step 204: Determine target edge similarity based on the source image and the encrypted image. The target edge similarity is used to characterize the degree of similarity between edge features of the source image and edge features of the encrypted image. For example, the target edge similarity can be edge structure similarity, which can quantify the risk of explicit privacy leakage.

[0096] For example, given a source image S and an encrypted image C, the degree of similarity between the edge features of the source image S and the encrypted image C can be determined, thereby obtaining a target edge similarity. The method for determining the target edge similarity is not limited, as long as it can characterize the degree of similarity between the edge features. In one possible implementation, the target edge similarity can be determined using the following steps: Step S21 : determining the gradient magnitude of the source image based on the horizontal gradient and the vertical gradient of the source image, and performing a normalization operation on the gradient magnitude of the source image to obtain a first target gradient magnitude.

[0097] Furthermore, a gradient magnitude of the encrypted image is determined based on the horizontal gradient and the vertical gradient of the encrypted image, and a normalization operation is performed on the gradient magnitude of the encrypted image to obtain a second target gradient magnitude.

[0098] Exemplarily, for each pixel of the source image, the horizontal gradient of the pixel can be determined based on the horizontal gradient operator, and the vertical gradient of the pixel can be determined based on the vertical gradient operator. The horizontal gradient operator and the vertical gradient operator can be configured according to actual needs, and can also include K gradient parameters, which are obtained by optimizing the target network model, see step S11.

[0099] For example, based on the horizontal gradient and the vertical gradient of the pixel, the gradient amplitude of the pixel can be determined. For example, the gradient amplitude of the pixel can be determined using the following formula: , It can represent the gradient amplitude of the pixel point. It can represent the horizontal gradient of the pixel. It can represent the vertical gradient of the pixel. Then, the gradient amplitude of the pixel is normalized, such as scaling the gradient amplitude of the pixel to the pixel range of 0 to 255, to obtain the first target gradient amplitude of the pixel. For example, the first target gradient amplitude of the pixel can be determined by the following formula: , Indicates the first target gradient amplitude of the pixel point, Represents the minimum value of the gradient magnitude of all pixels in the source image. Indicates the maximum value of the gradient magnitude of all pixels in the source image. k It can be configured according to actual needs. It is the parameter value of the normalization operation. For example, when scaling the gradient amplitude to the pixel range of 0~255, k Can be 255.

[0100] In summary, for each pixel of the source image, we can obtain the first target gradient magnitude of that pixel. Similarly, for each pixel of the encrypted image, we can obtain the second target gradient magnitude of that pixel. Simply replace the source image with the encrypted image to obtain the second target gradient magnitude of each pixel in the encrypted image.

[0101] Step S22: Generate a first gradient histogram corresponding to the source image based on the first target gradient magnitude; wherein the first gradient histogram may include the gradient magnitude distribution of each pixel in the source image.

[0102] And, based on the second target gradient amplitude, a second gradient histogram corresponding to the encrypted image is generated; wherein the second gradient histogram may include the gradient amplitude distribution of each pixel in the encrypted image.

[0103] For example, the first target gradient magnitude corresponding to each pixel point of the source image can be Generate a first gradient histogram, such as by using the following formula: In the above formula, k Represents the horizontal coordinate of the first gradient histogram. If the value range of the first target gradient amplitude is 0~255, then k The values ​​are 0, 1, 2, ..., 254, 255, Represents the vertical coordinate of the first gradient histogram, used to represent the value " k "The corresponding count value. For example, for " k ", traverses each pixel of the source image (i.e. pixel point i ), if the first target gradient amplitude of the pixel is equal to k , then The count value is increased by 1, if the first target gradient amplitude of the pixel point is not equal to k , then ignore the pixel. After traversing all the pixels, we can get The count value of .

[0104] For example, for the position where the horizontal axis is 0, the count value is counted , count value is the number of pixels where the first target gradient amplitude is 0. For the position where the horizontal coordinate is 1, the count value is counted , count value is the number of pixels where the first target gradient amplitude is 1. Similarly, for the position with the horizontal coordinate of 255, the count value is counted , count value is the number of pixels with the first target gradient amplitude of 255. At this point, 256 count values ​​can be obtained, and these 256 count values ​​constitute the first gradient histogram. That is, the first gradient histogram can include 256 coordinate points, and the horizontal coordinates of the 256 coordinate points are 0, 1, 2, ..., 254, 255 in sequence, and the vertical coordinates of the 256 coordinate points can be the corresponding count values.

[0105] Exemplarily, the second gradient histogram may be generated based on the second target gradient magnitude corresponding to each pixel of the encrypted image, in a manner similar to that of generating the first gradient histogram, which will not be repeated here.

[0106] Step S23: Generate a joint gradient histogram based on the first gradient histogram and the second gradient histogram.

[0107] For example, the joint gradient histogram can reflect the common edge features of the source image and the encrypted image, while the encrypted image should retain some edge structures, otherwise it will lose its practicality. For example, the following formula can be used to generate the joint gradient histogram: For example, represents the joint gradient histogram, represents the first gradient histogram, Represents the second gradient histogram. For the position where the horizontal coordinate of the joint gradient histogram is 0, the count value is and The average value of is the count value corresponding to the horizontal coordinate 0 of the first gradient histogram, is the count value corresponding to the horizontal coordinate 0 of the second gradient histogram. Similarly, for the position where the horizontal coordinate of the joint gradient histogram is 255, the count value is and The average value of is the count value corresponding to the abscissa 255 of the first gradient histogram, is the count value corresponding to the abscissa 255 of the second gradient histogram.

[0108] Step S24: determining a lower threshold value and an upper threshold value based on the joint gradient histogram.

[0109] For example, the total count value of the joint gradient histogram (i.e., the sum of all count values) can be determined and recorded as Total ,but Total= *0+ *1+ *2+…+ *254+ *255, Indicates the count value corresponding to the position where the horizontal coordinate of the joint gradient histogram is 0, and so on. Indicates the count value corresponding to the position where the horizontal coordinate of the joint gradient histogram is 255.

[0110] For example, the first threshold and the second threshold can be determined based on the total count value, and the first threshold can be smaller than the second threshold. For example, the lowest 5% of gradients and the highest 5% of gradients can be excluded from the valid interval to remove possible noise or over-densified areas, and the middle 90% of the gradient range is retained. The middle 90% of the gradient range corresponds to a significant edge. In this way, the first threshold can be 0.05* Total , the second threshold can be 0.95* Total Of course, "0.05" and "0.95" are just examples and can be configured according to actual needs.

[0111] Exemplarily, the lower threshold value may be determined based on the first threshold value and the second threshold value. t min and upper threshold value t max , the lower threshold value and the upper threshold value form a valid interval, and the valid interval is determined by the cumulative distribution of the joint histogram. For example, the lower threshold value and the upper threshold value can be determined using the following formula: . For the above formula, first traverse k =0, for *0, if is greater than or equal to the first threshold, then t min Is 0, otherwise, continue traversing k =1, for *0+ *1, if is greater than or equal to the first threshold, then t min If it is 1, otherwise, continue traversing k =2, for *0+ *1+ *2, if is greater than or equal to the first threshold, then t min If it is 2, otherwise, continue traversing k =3, and so on, until the first value greater than or equal to the first threshold is found. ,this The corresponding value represents the lower threshold value t min Similarly, first traverse k =0, if is greater than or equal to the second threshold, then t max Is 0, otherwise, continue traversing k =1, if is greater than or equal to the second threshold, then t max If it is 1, otherwise, continue traversing k =2, if is greater than or equal to the second threshold, then t max If it is 2, otherwise, continue traversing k =3, and so on, until the first value greater than or equal to the second threshold is found. ,this The corresponding value represents the upper threshold value t max .

[0112] Step S25 : Generate N uniformly distributed dynamic thresholds based on the lower threshold value and the upper threshold value, and form a threshold set with these N dynamic thresholds, that is, the threshold set may include N dynamic thresholds.

[0113] Exemplarily, N can be a positive integer greater than 1. For example, N can be pre-configured, that is, the value of N is configured according to actual needs, or N can also be determined based on the number of valid peaks of the joint gradient histogram. For example, if the number of valid peaks of the joint gradient histogram is M, then N=M+2. For "2" dynamic thresholds, it can be a lower threshold value and an upper threshold value. For "M" dynamic thresholds, it can be M dynamic thresholds corresponding to M valid peaks. Regarding how to detect the valid peaks of the joint gradient histogram, there is no limitation in this embodiment, and the detection of M valid peaks (i.e., significant peaks) is taken as an example.

[0114] For example, based on the lower threshold value t min and upper threshold value t max , the following formula can be used to generate N uniformly distributed dynamic thresholds: In the above formula, TS Represents a threshold set, which includes N dynamic thresholds. Indicates the i A dynamic threshold, i When it is 0, is the lower threshold value t min ,exist i When it is N-1, is the upper threshold value t max .

[0115] For example, a fixed threshold can be configured based on experience during edge detection. However, if the fixed threshold is set too high, weak edges will be missed, while if the fixed threshold is set too low, noise will be introduced. In response to the above findings, this embodiment proposes a dynamic threshold, which can adaptively select a threshold range based on image content to ensure coverage of meaningful edges and exclusion of irrelevant areas, thereby obtaining N dynamic thresholds.

[0116] Step S26 : For each dynamic threshold in the threshold set, determine a first edge binary image corresponding to the source image based on the dynamic threshold, and determine a second edge binary image corresponding to the encrypted image based on the dynamic threshold.

[0117] For example, for each pixel point in the first edge binary image, the pixel point can be the first value or the second value. If the pixel point is the first value (such as 1), it means that the pixel point corresponding to the pixel point in the source image is an edge point. If the pixel point is the second value (such as 0), it means that the pixel point corresponding to the pixel point in the source image is a non-edge point. In summary, it is necessary to detect whether each pixel point in the source image is an edge point or a non-edge point, and then generate the first edge binary image corresponding to the source image, which is recorded as For example, a pixel point in the first edge binary image has a first value or a second value. The first value in the first edge binary image may represent an edge point, and the second value may represent a non-edge point.

[0118] Regarding how to detect whether each pixel in the source image is an edge point or a non-edge point, an edge detection algorithm can be used. In the edge detection algorithm, the pixel information is compared with the dynamic threshold, and then the pixel is detected to see whether it is an edge point. There is no restriction on this detection process. Obviously, the first edge binary image is related to the dynamic threshold, and different dynamic thresholds will correspond to different first edge binary images. In the above first edge binary image In the figure, t represents the dynamic threshold, that is, the first edge binary image for the dynamic threshold t.

[0119] For example, for each pixel point of the second edge binary image, the pixel point has the first value or the second value. If the pixel point has the first value (such as 1), it means that the pixel point corresponding to the pixel point in the encrypted image is an edge point. If the pixel point has the second value (such as 0), it means that the pixel point corresponding to the pixel point in the encrypted image is a non-edge point. In summary, it is necessary to detect whether each pixel point in the encrypted image is an edge point or a non-edge point, and then generate the second edge binary image corresponding to the encrypted image, which is recorded as .

[0120] Step S27 : Acquire a common edge image based on the first edge binary image and the second edge binary image.

[0121] Exemplarily, for each pixel point in the common edge image, if the result of a logical AND operation of the pixel value corresponding to the pixel point in the first edge binary image and the pixel value corresponding to the pixel point in the second edge binary image is a first value, and the absolute value of the difference between the first pixel value corresponding to the pixel point in the source image and the second pixel value corresponding to the pixel point in the encrypted image is not greater than the brightness control threshold, then based on the first pixel value, the second pixel value and the brightness control threshold, the pixel value corresponding to the pixel point in the common edge image is determined; otherwise, the pixel value corresponding to the pixel point in the common edge image is determined to be the second value.

[0122] For example, the following formula can be used to obtain the common edge image, which is recorded as the common edge image :

[0123] In the above formula, the pixel points in the common edge image are (i,j) For example, i Represents the horizontal coordinate of the pixel point, j Represents the vertical coordinate of the pixel point. t represents the dynamic threshold, that is, the first edge binary image, the second edge binary image, and the common edge image all correspond to the dynamic threshold t. Represents pixel points (i,j) The corresponding pixel values ​​in the common edge image. Represents pixel points (i,j) The corresponding pixel value in the first edge binary image, Represents pixel points (i,j) The corresponding pixel value in the second edge binary image, Represents the logical AND operation, Indicates that the result of the logical AND operation is the first value (1). Represents pixel points (i,j) The corresponding first pixel value in the source image S, Represents pixel points (i,j) The corresponding second pixel value in the encrypted image C, Indicates the brightness control threshold.

[0124] In the above formula, if is not 1, and / or Greater than , then determine the pixel (i,j) The corresponding pixel value in the common edge image is the second value (0).

[0125] In the above formula, , It is a control and Through experiments, the brightness control threshold It can be 20 or other values, there is no limit on this. ,In order to obtain the common edge set between the source image S and the encrypted image C more accurately, not only the edge detection results but also the brightness difference are considered.

[0126] Step S28 : For each dynamic threshold, determine the edge similarity corresponding to the dynamic threshold based on the common edge image corresponding to the dynamic threshold and the first edge binary image corresponding to the dynamic threshold.

[0127] For example, the edge similarity between the source image S and the encrypted image C can be determined using the following formula: . t represents the dynamic threshold, that is, the edge similarity, the common edge image and the first edge binary image all correspond to the dynamic threshold t. Represents the edge similarity between the source image S and the encrypted image C. represents the common edge image, Represents the first edge binary image.

[0128] Represents the common edge image The norm of, for example, Represents a common edge image The sum of all non-zero elements in . Of course, this is just an example of the norm and there is no restriction on this.

[0129] Represents the binary image for the first edge The norm of, for example, Represents the first edge binary image The sum of all non-zero elements in . Of course, this is just an example of a norm.

[0130] Step S29: Weight the edge similarities corresponding to each dynamic threshold to obtain the target edge similarity.

[0131] For example, for each dynamic threshold in the threshold set, based on steps S26-S28, the edge similarity corresponding to the dynamic threshold can be obtained. On this basis, the edge similarity corresponding to each dynamic threshold can be weighted to obtain the target edge similarity. For example, the target edge similarity can be determined using the following formula: . represents the target edge similarity, represents the weight value of the dynamic threshold t, represents the edge similarity corresponding to the dynamic threshold t. The dynamic threshold t belongs to the threshold set TS, and the threshold set TS may include N dynamic thresholds.

[0132] For each dynamic threshold in the threshold set TS, a weight value can be assigned to each dynamic threshold. The weight values ​​of different dynamic thresholds can be the same or different. For example, the sum of the weight values ​​of all dynamic thresholds can be 1, which is expressed by the following formula: .

[0133] In summary, this embodiment considers that encryption destroys image structure but may overly damage practicality, and visual anomalies may attract the attention of attackers. Therefore, by maintaining high edge similarity, the edge positions remain unchanged, but the actual semantics of the edge region are modified, thus balancing privacy and practicality. Ideal edge similarity is sufficient to preserve structural practicality, but the semantic content associated with the edge is desensitized.

[0134] At this point, step 204 is completed, and the target edge similarity between the source image and the encrypted image is obtained.

[0135] Step 205: Determine target semantic similarity based on the source image and the encrypted image. Target semantic similarity is used to characterize the degree of similarity between the semantic features of the source image and the encrypted image. For example, target semantic similarity can be used to perform privacy protection assessment and analysis. Target semantic similarity can characterize the semantic correlation between non-sensitive areas (non-privacy areas) and sensitive areas, and can quantify the risk of implicit privacy leakage.

[0136] For example, in the field of privacy protection, semantic risk refers to the possibility that attackers can infer sensitive information by analyzing semantic information (such as objects, scenes, relationships, etc.) in multimedia data (such as images and videos). It can measure the potential threats implicit in the data that can leak privacy through high-level semantic understanding.

[0137] For example, given a source image S and an encrypted image C, the similarity between the semantic features of the source image S and the encrypted image C can be determined, thereby obtaining a target semantic similarity. The method for determining the target semantic similarity is not limited, as long as it can represent the similarity between the semantic features. In one possible implementation, the target semantic similarity can be determined using the following steps: Step S31: Determine the semantic similarity of the objects based on the overlap between the detection frames of the source image and the encrypted image. For example, the detection frames of the source image may be regions (e.g., rectangular regions) of each identifiable object (i.e., sensitive objects such as faces, license plates, and ID cards) within the source image, and the detection frames of the encrypted image may be regions (e.g., rectangular regions) of each identifiable object within the encrypted image.

[0138] For example, a source image can be input into a detection model (the network structure of this detection model is not limited), and the detection model determines the objects within the source image and outputs the detection frames (which can be at least one detection frame) where these objects are located, denoted as detection frame a1, detection frame a2, and detection frame a3. This process can be seen in step 202. An encrypted image can be input into the detection model, and the detection model determines the objects within the encrypted image and outputs the detection frames where these objects are located, denoted as detection frame b1 and detection frame b2.

[0139] For example, object semantic similarity can be determined based on the degree of overlap (i.e., IOU) between the detection frames of the source image and the encrypted image. For example, the greater the overlap between the detection frames of the source image and the encrypted image, the greater the object semantic similarity. For example, object semantic similarity can be determined based on the degree of overlap between detection frames a1 and b1, the degree of overlap between detection frames a1 and b2, the degree of overlap between detection frames a2 and b1, the degree of overlap between detection frames a2 and b2, the degree of overlap between detection frames a3 and b1, and the degree of overlap between detection frames a3 and b2. Object semantic similarity directly quantifies the risk of semantic leakage by comparing the detection differences between the source image and the encrypted image. Object semantic similarity determines risk based on the number of correctly detected objects that can still be detected after encryption.

[0140] For example, the following formula is used to determine the semantic similarity of objects : . The source image i detection frames, The encrypted image j detection frames, The object sensitivity weight can be configured as multiple object sensitivity weights according to actual needs. For example, if the source image has three detection boxes and the encrypted image has two detection boxes, six object sensitivity weights can be obtained.

[0141] For the overlap between detection frame a1 and detection frame b1, is the first object sensitive weight, for the overlap between detection frame a1 and detection frame b2, is the second object sensitive weight, for the overlap between detection frame a2 and detection frame b1, is the sensitive weight of the third object, and so on.

[0142] For example, the object semantic similarity can be determined based on the overlap between the detection frame of the source image and the true annotation frame, and the overlap between the detection frame of the encrypted image and the true annotation frame. The greater the overlap between the detection frame of the source image and the true annotation frame, the greater the object semantic similarity. The greater the overlap between the detection frame of the encrypted image and the true annotation frame, the greater the object semantic similarity.

[0143] For example, the ground truth box is the true detection box for the source image. Assuming the source image contains three identifiable objects (i.e., sensitive objects), we need to obtain the accurate regions (such as rectangular regions) of these three identifiable objects. These accurate regions of these three identifiable objects are the three ground truth boxes. Compared with the detection box of the source image, the detection box of the source image is the output of the detection model and may contain errors. The ground truth box is an accurately obtained detection box with no or very small errors. The ground truth box can be manually annotated by the user or annotated using some algorithm. As long as the ground truth box is accurately annotated, there is no restriction on this.

[0144] For example, based on the detection box of the source image, the detection box of the encrypted image, and the real annotation box, the following formula is used to determine the semantic similarity of the object : . The source image i detection frames, The encrypted image j detection frames, Indicates the k A ground-truth annotation box, It is the object sensitivity weight. You can configure multiple object sensitivity weights according to actual needs.

[0145] For the real annotation box c1, detection box a1 and detection box b1, is the first object sensitive weight, for the real annotation box c1, detection box a1 and detection box b2, is the second object sensitive weight, for the real annotation box c1, detection box a2 and detection box b1, is the third object sensitive weight, for the real annotation box c1, detection box a2 and detection box b2, is the sensitive weight of the fourth object, and so on.

[0146] Step S32: Determine the relationship semantic similarity based on the semantic association graph of the source image and the semantic association graph of the encrypted image. The relationship semantic similarity can be used to judge the risk by the number of correct relationships retained after encryption.

[0147] Illustratively, the semantic association graph of the source image may represent the association relationship between each identified object in the source image, and the semantic association graph of the encrypted image may represent the association relationship between each identified object in the encrypted image.

[0148] For example, assuming there are P1 identification objects (i.e., sensitive objects) in the source image, the semantic association map of the source image may include P1*P1 pixels. The pixel in the first row and first column indicates whether the first identification object is associated with the first identification object. For example, a pixel value of 1 indicates that there is an association, and a pixel value of 0 indicates that there is no association. Obviously, this is for the same identification object, so there is an association. The pixel in the first row and second column indicates whether the first identification object is associated with the second identification object. This embodiment does not limit how to determine whether two identification objects are associated. For example, if a user holds an ID card, the user identification object is associated with the ID card identification object. If there is a cat in the distance of a vehicle, the vehicle identification object is not associated with the cat identification object. The pixel in the second row and first column indicates whether the second identification object is associated with the first identification object, and has the same pixel value as the pixel in the first row and second column.

[0149] The pixel point in the 1st row and 3rd column indicates whether there is an association relationship between the 1st identified object and the 3rd identified object, and the pixel point in the 3rd row and 1st column indicates whether there is an association relationship between the 3rd identified object and the 1st identified object. The pixel values ​​of the pixels at these two positions are the same, and so on.

[0150] In summary, the semantic association graph of the source image can be generated based on the source image. Similarly, the semantic association graph of the encrypted image can be generated based on the encrypted image. Based on the semantic association graph of the source image and the semantic association graph of the encrypted image, the following formula is used to determine the semantic similarity of the relationship : . Represents the semantic association graph of the source image, Represents the semantic association graph of the encrypted image, represents the graph edit distance, so that It represents the graph edit distance between the semantic association graph of the source image and the semantic association graph of the encrypted image. Indicates taking and The maximum value in .

[0151] Step S33: Determine scene semantic similarity based on the scene feature vector of the source image and the scene feature vector of the encrypted image. For example, the source image is input into a feature extraction network (e.g., a ResNet network) to obtain the scene feature vector of the source image. Specifically, the feature extraction network extracts features from the source image to obtain the scene feature vector of the source image. The encrypted image is input into a feature extraction network to obtain the scene feature vector of the encrypted image. Specifically, the feature extraction network extracts features from the encrypted image to obtain the scene feature vector of the encrypted image.

[0152] For example, the following formula is used to determine the scene semantic similarity : . represents the scene feature vector of the source image, represents the scene feature vector of the encrypted image, and cos represents the cosine similarity, that is, the cosine similarity between the scene feature vector of the source image and the scene feature vector of the encrypted image.

[0153] Step S34: Perform a weighted calculation on the object semantic similarity, the relationship semantic similarity, and the scene semantic similarity to obtain a semantic risk value. For example, the weighted coefficient of the object semantic similarity is greater than the weighted coefficient of the relationship semantic similarity, and the weighted coefficient of the relationship semantic similarity is greater than the weighted coefficient of the scene semantic similarity.

[0154] For example, the semantic risk value is determined using the following formula: , SLS The semantic risk value can be expressed. w1 The weighted coefficient representing the semantic similarity of the object can be configured based on experience, such as 0.5. w2 The weighted coefficient representing the semantic similarity of the relationship can be configured based on experience, such as 0.3. w3 The weighting coefficient representing the scene semantic similarity can be configured based on experience, such as 0.2. In this embodiment, there is no restriction on this weighting coefficient. w1 Can be greater than w2 , or less than w2 , w1 Can be greater than w3 , or less than w3 , w2 Can be greater than w3 , or less than w3 .

[0155] Step S35: Determine the target semantic similarity based on the semantic risk value.

[0156] For example, the following formula can be used to determine the target semantic similarity D SR : .

[0157] At this point, step 205 is completed, and the target semantic similarity between the source image and the encrypted image is obtained.

[0158] Step 206: Determine a privacy protection indicator based on the target texture similarity, the target edge similarity, and the target semantic similarity. For example, a weighted operation can be performed on the target texture similarity, the target edge similarity, and the target semantic similarity to obtain the privacy protection indicator. The first weighting coefficient of the target texture similarity can be greater than or less than the second weighting coefficient of the target edge similarity. The first weighting coefficient can be greater than or less than the third weighting coefficient of the target semantic similarity. The second weighting coefficient can be greater than or less than the third weighting coefficient.

[0159] In one possible implementation, the privacy protection indicator may be determined using the following steps: Step S41 : determining a first standard deviation based on target texture similarity, determining a second standard deviation based on target edge similarity, and determining a third standard deviation based on target semantic similarity.

[0160] Exemplarily, in the process of determining the visual multidimensional quantitative indicators, multiple source images can be obtained (such as each frame image or part of the image in the video stream as the source image). For each source image, the above steps can be used to obtain the target texture similarity, target edge similarity and target semantic similarity corresponding to the source image, that is, to obtain multiple target texture similarities, multiple target edge similarities and multiple target semantic similarities.

[0161] On this basis, the first standard deviation can be determined based on the texture similarity of multiple targets, the second standard deviation can be determined based on the edge similarity of multiple targets, and the third standard deviation can be determined based on the semantic similarity of multiple targets. For example, the standard deviation can be expressed by the following formula : ,like represents the target texture similarity, then std Indicates calculating the first standard deviation of all target texture similarities ,like represents the target edge similarity, then std Indicates calculating the second standard deviation of all target edge similarities ,like represents the target semantic similarity, then std Indicates calculating the third standard deviation of the semantic similarity of all targets .

[0162] Step S42: determining a first correlation coefficient based on the target texture similarity and a pre-configured true leakage risk label, determining a second correlation coefficient based on the target edge similarity and the true leakage risk label, and determining a third correlation coefficient based on the target semantic similarity and the true leakage risk label.

[0163] Exemplarily, a real leakage risk label may be pre-configured, and the real leakage risk label may be configured according to actual scenario requirements, such as 0.9, 0.8, 0.5, etc., without limitation.

[0164] For example, the correlation coefficient is expressed by the following formula : , y Indicates the real leakage risk label. represents the target texture similarity, then corr Represents the first correlation coefficient between the target texture similarity and the true leakage risk label .like represents the target edge similarity, then corr Represents the second correlation coefficient between the target edge similarity and the true leakage risk label .like represents the target semantic similarity, then corr Represents the third correlation coefficient between the calculated target semantic similarity and the true leakage risk label .

[0165] For example, corr The correlation coefficient formula is used to measure the degree of linear correlation between two variables (such as target texture similarity and true leakage risk label). The result is a value between -1 and 1, which can be determined based on the covariance of target texture similarity and true leakage risk label, the standard deviation of target texture similarity (that is, the first standard deviation of multiple target texture similarities), and the standard deviation of true leakage risk label.

[0166] Step S43: determining a first weighting coefficient for target texture similarity based on the first standard deviation and the first correlation coefficient, determining a second weighting coefficient for target edge similarity based on the second standard deviation and the second correlation coefficient, and determining a third weighting coefficient for target semantic similarity based on the third standard deviation and the third correlation coefficient.

[0167] For example, the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient may be determined by the following formula: .like represents the first standard deviation of the target texture similarity, The first correlation coefficient representing the target texture similarity, Represents the first weighting coefficient of the target texture similarity. Indicates the second standard deviation of target edge similarity, The second correlation coefficient representing the target edge similarity, Represents the second weighting coefficient of target edge similarity. represents the third standard deviation of the target semantic similarity, The third correlation coefficient representing the target semantic similarity, The third weighting coefficient representing the target semantic similarity.

[0168] In the above formula, k The value of is 1, 2, 3, if k The value of is 1, then represents the first standard deviation, represents the first correlation coefficient, if k The value of is 2, then represents the second standard deviation, represents the second correlation coefficient, if k The value of is 3, then represents the third standard deviation, represents the third correlation coefficient.

[0169] Step S44: weighting the target texture similarity, the target edge similarity, and the target semantic similarity based on the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient to obtain a privacy protection index.

[0170] For example, the following formula can be used to determine the privacy protection index (denoted as the index score ): . Indicates target texture similarity The corresponding first weighting coefficient, Represents target edge similarity The corresponding second weighting coefficient, Indicates the target semantic similarity The corresponding third weighting coefficient. In summary, this embodiment proposes a dynamic weight fusion method to make the weight positively correlated with the dimension's ability to discriminate privacy leaks, thereby dynamically adapting to different image types.

[0171] At this point, step 206 is completed, and the privacy protection index, that is, the privacy protection index score value, is obtained.

[0172] Step 207: Determine whether the privacy protection indicator meets the preset conditions. If the privacy protection indicator meets the preset conditions, proceed to step 208; if the privacy protection indicator does not meet the preset conditions, proceed to step 209.

[0173] For example, if a larger privacy protection index indicates better privacy protection, then when the privacy protection index is greater than a preset threshold, the privacy protection index satisfies the preset condition; and when the privacy protection index is not greater than the preset threshold, the privacy protection index does not satisfy the preset condition. Alternatively, if a smaller privacy protection index indicates better privacy protection, then when the privacy protection index is less than the preset threshold, the privacy protection index satisfies the preset condition; and when the privacy protection index is not less than the preset threshold, the privacy protection index does not satisfy the preset condition.

[0174] Step 208: Determine whether the encrypted image meets the privacy protection requirement. In this case, the encrypted image can be transmitted. Since the encrypted image meets the privacy protection requirement, data security can be guaranteed.

[0175] Step 209: Determine that the encrypted image does not meet the privacy protection requirements. In this case, the encrypted image is not transmitted, and a different privacy protection method is used to re-process the source object to obtain an encrypted image, and the above steps are repeated.

[0176] In one possible implementation, see Figure 3 Figure 2 shows a schematic diagram of multi-dimensional visual quantitative indicators. The input layer acquires the source image. The preprocessing layer cleans the source image, such as by de-identifying and injecting noise. The analysis layer identifies the content of the source image to produce an encrypted image. For example, object detection and semantic segmentation can be performed on the source image to produce an encrypted image. The verification layer uses a dynamic threshold-based edge detection algorithm to obtain target edge similarity. It uses a semantic association graph to obtain target semantic similarity. It uses a trainable operator-based texture extraction algorithm to obtain target texture similarity. The output layer determines a privacy protection indicator (i.e., a privacy score) based on target edge similarity, target semantic similarity, and target texture similarity. Based on this privacy protection indicator, the output layer analyzes whether the encrypted image meets privacy requirements and provides optimization suggestions if the encrypted image does not meet privacy requirements.

[0177] As can be seen from the above technical solutions, the embodiments of this application construct multiple visual quantification dimensions (such as texture, edge, and semantic dimensions) to design a multi-dimensional quantitative index system, providing visual multi-dimensional quantitative indicators. This provides a scientific and objective evaluation method for the privacy protection effectiveness of multimedia data. This allows for a scientific and objective evaluation of the privacy protection effectiveness of multimedia data, providing an objective evaluation basis for the optimization and improvement of privacy protection technologies such as encryption, blurring, and desensitization, and more accurately assessing the privacy protection effectiveness of encrypted images. By integrating edge similarity (ES) and texture similarity (TS), the system comprehensively addresses the risk of both contour and detail leakage. Sensitive areas are identified based on an object detection model, and the weight distribution of edges and textures is dynamically adjusted. Joint analysis of edge and texture features mitigates the dual leakage risk of contours and details. Image edge information is extracted using an edge detection algorithm based on dynamic threshold optimization, texture similarity is extracted using a trainable gradient operator, and three dimensions, such as semantic similarity, combined with adaptive weight distribution and formal security proof, provide an objective evaluation basis for the optimization and improvement of privacy protection technologies such as encryption, blurring, and desensitization. By replacing fixed operators with trainable ones, we can more comprehensively extract image texture information based on different scenarios. Using a dynamic threshold edge detection algorithm avoids edge information loss or noise increase caused by a single threshold, ensuring accurate evaluation of privacy protection effectiveness.

[0178] Based on the same application concept as the above method, the embodiment of this application proposes a multi-dimensional quantitative index evaluation device for multimedia privacy protection, which is applied to a terminal device to be subjected to multimedia privacy protection. Figure 4 FIG. 1 is a schematic diagram of the structure of the device, which may include: An acquisition module 41 is configured to acquire a source image and an encrypted image; wherein the source image includes a privacy region and a non-privacy region, and the encrypted image is obtained by performing privacy processing on the privacy region of the source image; A determination module 42 is configured to determine a target texture similarity based on the source image and the encrypted image, the target texture similarity being used to characterize the degree of similarity between texture features of the source image and texture features of the encrypted image; determine a target edge similarity based on the source image and the encrypted image, the target edge similarity being used to characterize the degree of similarity between edge features of the source image and edge features of the encrypted image; determine a target semantic similarity based on the source image and the encrypted image, the target semantic similarity being used to characterize the degree of similarity between semantic features of the source image and semantic features of the encrypted image; and determine a privacy protection index based on the target texture similarity, the target edge similarity, and the target semantic similarity; if the privacy protection index satisfies a preset condition, it is determined that the encrypted image meets the privacy protection requirement; otherwise, it is determined that the encrypted image does not meet the privacy protection requirement.

[0179] Exemplarily, when determining the target texture similarity based on the source image and the encrypted image, the determination module 42 is specifically configured to: determine a first gradient direction and a first gradient magnitude based on the non-privacy area of ​​the source image, and determine a second gradient direction and a second gradient magnitude based on the non-privacy area of ​​the encrypted image; determine a directional similarity between the source image and the encrypted image based on the first gradient direction and the second gradient direction; determine an amplitude similarity between the source image and the encrypted image based on the first gradient magnitude, the second gradient magnitude, the non-privacy area of ​​the source image, and the non-privacy area of ​​the encrypted image; determine a first texture similarity of the non-privacy area based on the directional similarity and the amplitude similarity; determine a third gradient direction based on the privacy area of ​​the source image, and determine a fourth gradient direction based on the privacy area of ​​the encrypted image; determine a second texture similarity of the privacy area based on the third gradient direction, the fourth gradient direction, and the total number of pixels in the privacy area; and determine the target texture similarity based on the first texture similarity and the second texture similarity.

[0180] Exemplarily, when the determination module 42 determines the first gradient direction, the first gradient magnitude, the second gradient direction, the second gradient magnitude, the third gradient direction, and the fourth gradient direction, it is specifically used to: obtain a horizontal gradient operator and a vertical gradient operator, wherein the horizontal gradient operator and the vertical gradient operator each include K gradient parameters, where K is a positive integer; wherein the K gradient parameters are parameters within the target network model and are obtained by optimizing the target network model; determine the horizontal gradient of each pixel point in the source image based on the horizontal gradient operator, and determine the vertical gradient of each pixel point in the source image based on the vertical gradient operator; and determine the non-privacy area of ​​the source image based on the non-privacy area of ​​the source image. The first gradient direction and the first gradient amplitude are determined based on the horizontal gradient and the vertical gradient of each pixel in the domain; the third gradient direction is determined based on the horizontal gradient and the vertical gradient of each pixel in the privacy area of ​​the source image; the horizontal gradient of each pixel in the encrypted image is determined based on the horizontal gradient operator, and the vertical gradient of each pixel in the encrypted image is determined based on the vertical gradient operator; the second gradient direction and the second gradient amplitude are determined based on the horizontal gradient and the vertical gradient of each pixel in the non-privacy area of ​​the encrypted image; and the fourth gradient direction is determined based on the horizontal gradient and the vertical gradient of each pixel in the privacy area of ​​the encrypted image.

[0181] Exemplarily, when the determination module 42 obtains the K gradient parameters, it is specifically configured to: input a sample image into a target network model, and have the target network model output a predicted sensitivity probability for each first pixel in the sample image and a predicted gradient direction for each second pixel in the privacy region of the sample image, where the predicted sensitivity probability represents the probability that the first pixel is a sensitive pixel; determine a privacy-sensitive region segmentation loss value based on the predicted sensitivity probability of each first pixel in the sample image and the true sensitivity probability of each first pixel in the sample image; determine a gradient direction consistency loss value based on the predicted gradient direction of each second pixel in the privacy region and the true gradient direction of each second pixel in the privacy region; determine a target loss value based on the privacy-sensitive region segmentation loss value and the gradient direction consistency loss value, and optimize the target network model based on the target loss value to obtain an optimized model; if the optimized model has not converged, update the optimized model to the target network model, and return to the operation of inputting the sample image into the target network model; if the optimized model has converged, obtain the K gradient parameters from the network parameters of the optimized model.

[0182] Exemplarily, when the determination module 42 determines the target edge similarity based on the source image and the encrypted image, it is specifically used to: obtain a threshold set, the threshold set including multiple dynamic thresholds; for each dynamic threshold, determine a first edge binary image corresponding to the source image based on the dynamic threshold, and determine a second edge binary image corresponding to the encrypted image based on the dynamic threshold; wherein the first value in the first edge binary image represents an edge point, and the second value represents a non-edge point; obtain a common edge image; wherein, for each pixel point in the common edge image, if the pixel value corresponding to the pixel point in the first edge binary image is the same as the pixel value corresponding to the pixel point in the second edge binary image If the result of a logical AND operation of the corresponding pixel values ​​in the common edge image is a first value, and the absolute value of the difference between the first pixel value corresponding to the pixel point in the source image and the second pixel value corresponding to the pixel point in the encrypted image is not greater than the brightness control threshold, then based on the first pixel value, the second pixel value and the brightness control threshold, the pixel value corresponding to the pixel point in the common edge image is determined to be the second value; based on the common edge image and the first edge binary image, the edge similarity corresponding to the dynamic threshold is determined, and the edge similarity corresponding to each dynamic threshold is weighted to obtain the target edge similarity.

[0183] Exemplarily, when the determination module 42 obtains the threshold value set, it is specifically used to: determine the gradient amplitude of the source image based on the horizontal gradient and the vertical gradient of the source image, and perform a normalization operation on the gradient amplitude of the source image to obtain a first target gradient amplitude; generate a first gradient histogram corresponding to the source image based on the first target gradient amplitude; wherein the first gradient histogram includes the gradient amplitude distribution of each pixel in the source image; determine the gradient amplitude of the encrypted image based on the horizontal gradient and the vertical gradient of the encrypted image, and perform a normalization operation on the gradient amplitude of the encrypted image; The encrypted image is normalized by performing a normalization operation on the first target gradient amplitude; a second gradient histogram corresponding to the encrypted image is generated based on the second target gradient amplitude; a joint gradient histogram is generated based on the first gradient histogram and the second gradient histogram, and a lower threshold value and an upper threshold value are determined based on the joint gradient histogram; N uniformly distributed dynamic thresholds are generated based on the lower threshold value and the upper threshold value, and the threshold set includes the N dynamic thresholds; wherein N is a positive integer greater than 1, N is pre-configured, or N is determined based on the number of valid peaks in the joint gradient histogram.

[0184] Exemplarily, when determining the target semantic similarity based on the source image and the encrypted image, the determination module 42 is specifically used to: determine the object semantic similarity based on the overlap between the detection frame of the source image and the detection frame of the encrypted image; wherein the detection frame of the source image is the area of ​​each identified object in the source image, and the detection frame of the encrypted image is the area of ​​each identified object in the encrypted image; determine the relational semantic similarity based on the semantic association graph of the source image and the semantic association graph of the encrypted image; wherein the semantic association graph of the source image represents the association relationship between each identified object in the source image, and the semantic association graph of the encrypted image represents the association relationship between each identified object in the encrypted image; determine the scene semantic similarity based on the scene feature vector of the source image and the scene feature vector of the encrypted image; wherein the source image is input into a feature extraction network to obtain the scene feature vector of the source image, and the encrypted image is input into a feature extraction network to obtain the scene feature vector of the encrypted image; perform a weighted operation on the object semantic similarity, the relational semantic similarity, and the scene semantic similarity to obtain a semantic risk value, and determine the target semantic similarity based on the semantic risk value.

[0185] Exemplarily, when determining the privacy protection index based on the target texture similarity, the target edge similarity and the target semantic similarity, the determination module 42 is specifically used to: determine a first standard deviation based on the target texture similarity, determine a first correlation coefficient based on the target texture similarity and a pre-configured true leakage risk label, and determine a first weighting coefficient corresponding to the target texture similarity based on the first standard deviation and the first correlation coefficient; determine a second standard deviation based on the target edge similarity, determine a second correlation coefficient based on the target edge similarity and the true leakage risk label, and determine a second weighting coefficient corresponding to the target edge similarity based on the second standard deviation and the second correlation coefficient; determine a third standard deviation based on the target semantic similarity, determine a third correlation coefficient based on the target semantic similarity and the true leakage risk label, and determine a third weighting coefficient corresponding to the target semantic similarity based on the third standard deviation and the third correlation coefficient; and weight the target texture similarity, the target edge similarity and the target semantic similarity based on the first weighting coefficient, the second weighting coefficient and the third weighting coefficient to obtain the privacy protection index.

[0186] Based on the same application concept as the above method, an electronic device is proposed in the embodiment of the present application, see Figure 5 As shown, it includes: a processor 51 and a machine-readable storage medium 52, the machine-readable storage medium 52 stores machine-executable instructions that can be executed by the processor 51; the processor 51 is used to execute the machine-executable instructions to implement a multi-dimensional quantitative indicator evaluation method for multimedia privacy protection.

[0187] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, a multi-dimensional quantitative indicator evaluation method for multimedia privacy protection can be implemented.

[0188] The machine-readable storage medium may be any electronic, magnetic, optical, or other physical storage device that may contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.

[0189] Based on the same application concept as the above method, an embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above multi-dimensional quantitative indicator evaluation method for multimedia privacy protection.

[0190] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, the embodiments of the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0191] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A multi-dimensional quantitative index evaluation method for multimedia privacy protection, characterized by: Applied to a terminal device to be subjected to multimedia privacy protection, the method comprises: Acquire a source image and an encrypted image; wherein the source image includes a privacy area and a non-privacy area, and the encrypted image is obtained by performing privacy processing on the privacy area of ​​the source image; Determining a target texture similarity based on the source image and the encrypted image, wherein the target texture similarity is used to represent a degree of similarity between texture features of the source image and texture features of the encrypted image; determining a target edge similarity based on the source image and the encrypted image, wherein the target edge similarity is used to represent a degree of similarity between edge features of the source image and edge features of the encrypted image; determining a target semantic similarity based on the source image and the encrypted image, the target semantic similarity being used to characterize a degree of similarity between semantic features of the source image and semantic features of the encrypted image; A privacy protection index is determined based on the target texture similarity, the target edge similarity, and the target semantic similarity; if the privacy protection index satisfies a preset condition, it is determined that the encrypted image meets the privacy protection requirement; otherwise, it is determined that the encrypted image does not meet the privacy protection requirement.

2. The method according to claim 1, characterized in that The determining of target texture similarity based on the source image and the encrypted image includes: Determining a first gradient direction and a first gradient magnitude based on the non-privacy region of the source image, and determining a second gradient direction and a second gradient magnitude based on the non-privacy region of the encrypted image; determining a directional similarity between the source image and the encrypted image based on the first gradient direction and the second gradient direction; determining an amplitude similarity between the source image and the encrypted image based on the first gradient magnitude, the second gradient magnitude, the non-privacy region of the source image, and the non-privacy region of the encrypted image; and determining a first texture similarity of the non-privacy region based on the directional similarity and the amplitude similarity; determining a third gradient direction based on the privacy area of ​​the source image, and determining a fourth gradient direction based on the privacy area of ​​the encrypted image; and determining a second texture similarity of the privacy area based on the third gradient direction, the fourth gradient direction, and the total number of pixels in the privacy area; A target texture similarity is determined based on the first texture similarity and the second texture similarity.

3. The method according to claim 2, characterized in that The process of determining the first gradient direction, the first gradient magnitude, the second gradient direction, the second gradient magnitude, the third gradient direction, and the fourth gradient direction includes: Obtaining a horizontal gradient operator and a vertical gradient operator, wherein each of the horizontal gradient operator and the vertical gradient operator includes K gradient parameters, where K is a positive integer; wherein the K gradient parameters are parameters within a target network model and are obtained by optimizing the target network model; determining the horizontal gradient of each pixel in the source image based on the horizontal gradient operator, and determining the vertical gradient of each pixel in the source image based on the vertical gradient operator; determining the first gradient direction and the first gradient amplitude based on the horizontal gradient and the vertical gradient of each pixel in the non-privacy area of ​​the source image; and determining the third gradient direction based on the horizontal gradient and the vertical gradient of each pixel in the privacy area of ​​the source image; The horizontal gradient of each pixel in the encrypted image is determined based on the horizontal gradient operator, and the vertical gradient of each pixel in the encrypted image is determined based on the vertical gradient operator; the second gradient direction and the second gradient amplitude are determined based on the horizontal gradient and the vertical gradient of each pixel in the non-privacy area of ​​the encrypted image; and the fourth gradient direction is determined based on the horizontal gradient and the vertical gradient of each pixel in the privacy area of ​​the encrypted image.

4. The method according to claim 3, characterized in that The process of obtaining the K gradient parameters specifically includes: Inputting a sample image into a target network model, the target network model outputs a predicted sensitivity probability for each first pixel in the sample image and a predicted gradient direction for each second pixel in a privacy region of the sample image, where the predicted sensitivity probability represents the probability that the first pixel is a sensitive pixel; determining a privacy-sensitive area segmentation loss value based on a predicted sensitivity probability of each first pixel point in the sample image and a true sensitivity probability of each first pixel point in the sample image; Determining a gradient direction consistency loss value based on a predicted gradient direction of each second pixel point in the privacy area and a true gradient direction of each second pixel point in the privacy area; Determining a target loss value based on the privacy-sensitive area segmentation loss value and the gradient direction consistency loss value, and optimizing a target network model based on the target loss value to obtain an optimized model; If the optimized model has not converged, the optimized model is updated to the target network model, and the operation of inputting the sample image to the target network model is returned to execute; if the optimized model has converged, the K gradient parameters are obtained from the network parameters of the optimized model.

5. The method according to claim 1, wherein The determining of target edge similarity based on the source image and the encrypted image includes: Acquire a threshold set, where the threshold set includes multiple dynamic thresholds; For each dynamic threshold, determining a first edge binary image corresponding to the source image based on the dynamic threshold, and determining a second edge binary image corresponding to the encrypted image based on the dynamic threshold; wherein a first value in the first edge binary image represents an edge point, and a second value represents a non-edge point; Acquire a common edge image; wherein, for each pixel in the common edge image, if a logical AND operation result of a pixel value corresponding to the pixel in the first edge binary image and a pixel value corresponding to the pixel in the second edge binary image is a first value, and an absolute value of a difference between a first pixel value corresponding to the pixel in the source image and a second pixel value corresponding to the pixel in the encrypted image is not greater than a brightness control threshold, then determine the pixel value corresponding to the pixel in the common edge image based on the first pixel value, the second pixel value, and the brightness control threshold; otherwise, determine the pixel value corresponding to the pixel in the common edge image to be a second value; The edge similarity corresponding to the dynamic threshold is determined based on the common edge image and the first edge binary image, and the edge similarity corresponding to each dynamic threshold is weighted to obtain a target edge similarity.

6. The method according to claim 5, characterized in that The acquiring threshold value set includes: Determining a gradient magnitude of the source image based on a horizontal gradient and a vertical gradient of the source image, and performing a normalization operation on the gradient magnitude of the source image to obtain a first target gradient magnitude; generating a first gradient histogram corresponding to the source image based on the first target gradient magnitude; wherein the first gradient histogram includes a gradient magnitude distribution of each pixel in the source image; determining a gradient magnitude of the encrypted image based on a horizontal gradient and a vertical gradient of the encrypted image, performing a normalization operation on the gradient magnitude of the encrypted image to obtain a second target gradient magnitude; and generating a second gradient histogram corresponding to the encrypted image based on the second target gradient magnitude; generating a joint gradient histogram based on the first gradient histogram and the second gradient histogram, and determining a lower threshold value and an upper threshold value based on the joint gradient histogram; N uniformly distributed dynamic thresholds are generated based on the lower threshold value and the upper threshold value, and the threshold set includes the N dynamic thresholds; wherein N is a positive integer greater than 1, N is pre-configured, or N is determined based on the number of valid peaks of the joint gradient histogram.

7. The method according to claim 1, characterized in that The determining of the target semantic similarity based on the source image and the encrypted image includes: Determining object semantic similarity based on a degree of overlap between a detection frame of the source image and a detection frame of the encrypted image; wherein the detection frame of the source image is a region of each identified object within the source image, and the detection frame of the encrypted image is a region of each identified object within the encrypted image; Determining the relational semantic similarity based on the semantic association graph of the source image and the semantic association graph of the encrypted image; wherein the semantic association graph of the source image represents the association relationship between each identified object in the source image, and the semantic association graph of the encrypted image represents the association relationship between each identified object in the encrypted image; Determining scene semantic similarity based on the scene feature vector of the source image and the scene feature vector of the encrypted image; wherein the source image is input into a feature extraction network to obtain the scene feature vector of the source image, and the encrypted image is input into a feature extraction network to obtain the scene feature vector of the encrypted image; A weighted operation is performed on the object semantic similarity, the relationship semantic similarity, and the scene semantic similarity to obtain a semantic risk value, and a target semantic similarity is determined based on the semantic risk value.

8. The method according to claim 1, characterized in that The determining of the privacy protection index based on the target texture similarity, the target edge similarity, and the target semantic similarity includes: determining a first standard deviation based on the target texture similarity, determining a first correlation coefficient based on the target texture similarity and a pre-configured true leakage risk label, and determining a first weighting coefficient corresponding to the target texture similarity based on the first standard deviation and the first correlation coefficient; determining a second standard deviation based on the target edge similarity, determining a second correlation coefficient based on the target edge similarity and the true leakage risk label, and determining a second weighting coefficient corresponding to the target edge similarity based on the second standard deviation and the second correlation coefficient; determining a third standard deviation based on the target semantic similarity, determining a third correlation coefficient based on the target semantic similarity and the true leakage risk label, and determining a third weighting coefficient corresponding to the target semantic similarity based on the third standard deviation and the third correlation coefficient; The target texture similarity, the target edge similarity, and the target semantic similarity are weighted based on the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient to obtain a privacy protection index.

9. A multi-dimensional quantitative index evaluation device for multimedia privacy protection, characterized in that: Applied to a terminal device to be subjected to multimedia privacy protection, the device comprises: An acquisition module, configured to acquire a source image and an encrypted image; wherein the source image includes a privacy area and a non-privacy area, and the encrypted image is obtained by performing privacy processing on the privacy area of ​​the source image; A determination module is configured to determine a target texture similarity based on the source image and the encrypted image, the target texture similarity being used to characterize the degree of similarity between texture features of the source image and those of the encrypted image; determine a target edge similarity based on the source image and the encrypted image, the target edge similarity being used to characterize the degree of similarity between edge features of the source image and those of the encrypted image; determine a target semantic similarity based on the source image and the encrypted image, the target semantic similarity being used to characterize the degree of similarity between semantic features of the source image and those of the encrypted image; and determine a privacy protection index based on the target texture similarity, the target edge similarity, and the target semantic similarity; if the privacy protection index satisfies a preset condition, it is determined that the encrypted image meets the privacy protection requirement; otherwise, it is determined that the encrypted image does not meet the privacy protection requirement.

10. An electronic device, characterized in that: include: a processor and a machine-readable storage medium storing machine-executable instructions capable of being executed by the processor; The processor is configured to execute machine-executable instructions to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Network image retrieval method based on semantic analysis

    CN101751447A

  • Image acquisition method and equipment

    CN113709353A

  • Image perception visual safety evaluation method based on edge and texture similarity

    CN117336414A

  • Image privacy positioning identification method and device based on multi-modal large model

    CN119863691A

  • A new perceptual thresholding for gradient-based local edge detection

    WO1999030270A1

Cited By

  • A method for adaptive encryption and decryption of key regions of images of unmanned aerial vehicles

    CN122601802B