Security analysis methods and devices for protecting the personal privacy of image data

By extracting sensitive metadata and encrypting personalized parameters, combined with multi-resolution perceptual feature similarity mapping, the visual security analysis of image data is optimized, solving the problem of insufficient visual security in existing technologies and achieving more efficient privacy protection and security analysis.

CN120881213BActive Publication Date: 2026-01-06HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511373701.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-01-06
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Existing image and video encryption algorithms mainly focus on cryptographic security, lacking visual security analysis, leading to serious risks of privacy leaks.

Method used

The target scene is determined by extracting sensitive metadata, encryption is performed based on personalized parameters, and multi-resolution perceptual feature similarity mapping and visual security scoring are conducted to optimize security analysis.

Benefits of technology

It achieves personalized privacy protection, improves the accuracy and flexibility of visual security analysis of image data, and enhances the applicability to various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120881213B_ABST
    Figure CN120881213B_ABST
Patent Text Reader

Abstract

The application provides a security analysis method and device for image data individual privacy protection. In one example, the method comprises: performing sensitive metadata extraction on input image data to determine a corresponding target scene; performing encryption processing on the input image according to an individualized parameter corresponding to the target scene; performing down-sampling processing on the input image data and the encrypted image data respectively to obtain input image data and encrypted image data of T different resolutions; determining visual security scores of different resolutions respectively according to the perceptual features of the input image data and the encrypted image data of different resolutions; and determining a visual security evaluation score of the encrypted image data according to the visual security scores of the T different resolutions. The method can improve the accuracy of security analysis of image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of visual security technology for images and videos, and in particular to a security analysis method and device for protecting the individual privacy of image data. Background Technology

[0002] Against the backdrop of accelerated digitalization, the application of image and video data in fields such as medical diagnosis, smart security, and social media has seen explosive growth, and the resulting risks of privacy breaches have become increasingly serious.

[0003] Currently, there are many practical image and video encryption algorithms. However, security analysis of these algorithms mainly focuses on cryptographic security, with little work addressing visual security. Visual security means that the encrypted video content is incomprehensible to the human eye. The higher the visual security provided by the encryption algorithm, the less information an attacker can obtain from the encrypted image, and the more difficult the attack becomes. Summary of the Invention

[0004] In view of this, this application provides a security analysis method and device for protecting the personal privacy of image data.

[0005] Specifically, this application is implemented through the following technical solution:

[0006] According to a first aspect of the embodiments of this application, a security analysis method for protecting the personalized privacy of image data is provided, comprising:

[0007] Sensitive metadata is extracted from the input image data, and the target scene corresponding to the input image data is determined based on the extracted sensitive metadata.

[0008] Based on the target scenario, determine the personalized parameters corresponding to the target scenario;

[0009] The input image is encrypted based on the personalized parameters to obtain encrypted image data;

[0010] The input image data and the encrypted image data are downsampled respectively to obtain T types of input image data with different resolutions and T types of encrypted image data with different resolutions; wherein, T≥2;

[0011] For any resolution of input image data and encrypted image data, based on the perceptual features of the input image data at that resolution and the perceptual features of the encrypted image data at that resolution, a perceptual feature similarity mapping between the input image data at that resolution and the encrypted image data at that resolution is determined, and based on the perceptual feature similarity mapping between the input image data at that resolution and the encrypted image data at that resolution, a visual security score for that resolution is determined.

[0012] Based on the T different resolution visual security scores, the visual security evaluation score of the encrypted image data is determined.

[0013] According to a second aspect of the embodiments of this application, an electronic device is provided, including a processor and a memory, wherein...

[0014] Memory, used to store computer programs;

[0015] The processor, when executing a program stored in memory, implements the method provided in the first aspect.

[0016] According to a third aspect of the embodiments of this application, a computer program product is provided, wherein the computer program product stores a computer program, and the computer program, when executed by a processor, implements the method provided in the first aspect.

[0017] This application's embodiment of the security analysis method for personalized privacy protection of image data involves extracting sensitive metadata from input image data, determining the target scene corresponding to the input image data based on the extracted sensitive metadata, determining personalized parameters corresponding to the target scene, and encrypting the input image based on the determined personalized parameters to obtain encrypted image data. By performing scene recognition based on the extracted sensitive metadata and performing personalized encryption adapted to the scene, personalized privacy protection is achieved, and the scene applicability and flexibility of privacy protection are improved. Downsampling processing is performed on both the input image data and the encrypted image data to obtain various resolutions. The system takes input image data and encrypted image data of various resolutions as input. Based on the perceptual features of the input and encrypted image data at each resolution, it determines the perceptual feature similarity mapping between the input and encrypted image data at each resolution. Based on the perceptual feature similarity mapping between the input and encrypted image data at each resolution, it determines the visual security score for each resolution. Then, based on the visual security scores of various resolutions, it determines the visual security evaluation score for the encrypted image data. By using a perceptual feature-based image visual security evaluation method combined with multi-resolution visual feature analysis, the performance of security analysis is optimized and the accuracy of image data security analysis is improved. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a security analysis method for protecting the individual privacy of image data, as shown in an exemplary embodiment of this application.

[0019] Figure 2 This application provides an exemplary embodiment of a security analysis scheme for protecting the individual privacy of image data, illustrating an overall framework diagram.

[0020] Figure 3 This application provides an exemplary embodiment of a security analysis scheme for protecting the individual privacy of image data, illustrated in the following flowchart.

[0021] Figure 4 This application provides a schematic diagram illustrating the structure of a security analysis device for protecting the individual privacy of image data, as shown in an exemplary embodiment.

[0022] Figure 5 This is a schematic diagram of the hardware structure of an electronic device as illustrated in an exemplary embodiment of this application. Detailed Implementation

[0023] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, and to make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the technical solutions in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.

[0024] Please see Figure 1 This is a flowchart illustrating a security analysis method for protecting the personalized privacy of image data, as provided in an embodiment of this application. Figure 1 As shown, this security analysis method for protecting the individual privacy of image data may include:

[0025] Step S100: Extract sensitive metadata from the input image data and determine the target scene corresponding to the input image data based on the extracted sensitive metadata.

[0026] For example, input image data may include pictures or videos.

[0027] For example, when the input image data is video, for each video frame in the video, the solution provided in the embodiments of this application can be used to achieve secure analysis for personalized privacy protection of image data.

[0028] In this embodiment of the application, in order to improve the flexibility and accuracy of privacy protection and achieve personalized privacy protection, different privacy protection strategies can be selected according to the scenario.

[0029] Accordingly, for input image data, sensitive metadata can be extracted from the input image data, and the corresponding scene (which can be called the target scene) can be determined based on the extracted sensitive metadata.

[0030] Examples of such scenarios include, but are not limited to, personal cloud storage, short video platforms, medical imaging, or industrial inspection.

[0031] Sensitive metadata refers to identifiable information elements embedded in digital images that can be directly or indirectly linked to individuals, devices, or locations. It possesses core characteristics such as identifiability, structure, and contextual relevance.

[0032] For example, the sensitive metadata extraction process may include: binary parsing of input images or videos, EXIF / DICOM extraction, deep feature analysis, and structured sensitive metadata.

[0033] For example, sensitive metadata takes different forms in different scenarios.

[0034] For example, a specific example of sensitive metadata can be as follows:

[0035]

[0036] It should be noted that, in the embodiments of this application, all data involving personal information used in the implementation of the solution are used with the explicit authorization of the relevant personnel and in accordance with the relevant privacy policies and regulations. The relevant processing of the data during use strictly complies with the applicable data protection policies and regulations, and the relevant processing behavior has obtained the unambiguous authorization of the data subject based on the principle of explicit knowledge for the purpose of using the data for the purposes described in this application.

[0037] Step S110: Determine the personalized parameters corresponding to the target scenario based on the target scenario.

[0038] In this embodiment of the application, in order to achieve personalized privacy protection, the corresponding personalized parameters can be determined based on the scene corresponding to the input image data.

[0039] Step S120: Encrypt the input image according to the personalized parameters to obtain encrypted image data.

[0040] For example, selective encryption methods can be used during the encryption process of the input image. For instance, areas with risk values ​​exceeding a preset risk value threshold can be strongly encrypted (e.g., using a high-strength encryption algorithm); while areas with risk values ​​below the preset risk value threshold can retain the original data or undergo mild desensitization processing.

[0041] For example, a “risk value” is an indicator that can represent the sensitivity or risk of privacy breaches of a certain area in an image.

[0042] For example, a higher risk value indicates that the area is more likely to expose users' sensitive information, and therefore requires stronger protection measures.

[0043] For example, the system can identify sensitive areas in an image using an image recognition model to determine the risk value.

[0044] Step S130: Perform downsampling processing on the input image data and the encrypted image data respectively to obtain T types of input image data with different resolutions and T types of encrypted image data with different resolutions; where T≥2.

[0045] In this embodiment of the application, in order to simulate the hierarchical characteristics of the human visual system and optimize the security analysis performance of image data, after obtaining encrypted image data in the manner described above, the input image data and the encrypted image data can be downsampled to obtain T different resolutions of input image data and T different resolutions of encrypted image data, so as to obtain a more accurate security analysis score by performing security analysis on the encrypted image data of different resolutions.

[0046] For example, the above T types of input image data with different resolutions may include original input image data, as well as input image data obtained by downsampling at least one (T-1 types).

[0047] For example, for input image data, the input image data can be sampled T-1 times to obtain T different resolutions of input image data, including the original input image data.

[0048] Similarly, the T types of encrypted image data with different resolutions mentioned above may include the original encrypted image data, as well as at least one (T-1 types) of downsampled encrypted image data.

[0049] Step S140: For any resolution of input image data and encrypted image data, based on the perceptual features of the input image data at that resolution and the perceptual features of the encrypted image data at that resolution, determine the perceptual feature similarity mapping between the input image data at that resolution and the encrypted image data at that resolution, and determine the visual security score for that resolution based on the perceptual feature similarity mapping between the input image data at that resolution and the encrypted image data at that resolution.

[0050] In this embodiment of the application, in order to better analyze the security of encrypted image data, the security of encrypted image data can be analyzed from the perspective of visual security.

[0051] In addition, to optimize the performance of security analysis, image visual security evaluation based on perceptual features can be adopted.

[0052] For example, the security of encrypted image data can be analyzed and evaluated based on changes in perceptual features between input image data and encrypted image data.

[0053] Accordingly, for any resolution of input image data and encrypted image data, the perceptual feature similarity mapping between the input image data and encrypted image data can be determined based on the perceptual features of the input image data and the perceptual features of the encrypted image data at that resolution, and the visual security score for that resolution can be determined based on the perceptual feature similarity mapping between the input image data and encrypted image data at that resolution.

[0054] For example, perceptual features may include, but are not limited to, one or more of structural features, texture features, and naturalness features.

[0055] For example, when the perceived features include one of structural features, texture features, and natural features, the visual security score for that resolution can be determined based on the similarity mapping of the perceived features.

[0056] For example, based on the similarity mapping of the perceptual feature, the similarity score of the perceptual feature is determined, and the similarity score of the perceptual feature is determined as the visual security score for that resolution.

[0057] When perceptual features include multiple features such as structural features, texture features, and natural features, the visual security score for that resolution can be determined based on the similarity mapping of these multiple perceptual features.

[0058] For example, based on the similarity mapping of the multiple perceptual features, the similarity score of each perceptual feature can be determined, and based on the similarity scores of the multiple perceptual features, the visual security score of the resolution can be determined.

[0059] For example, the weighted average of the similarity scores of the various perceptual features can be used to determine the visual security score for that resolution.

[0060] Step S150: Determine the visual security evaluation score of the encrypted image data based on the visual security scores of T different resolutions.

[0061] In this embodiment of the application, after determining the visual security scores for each resolution in the manner described above, the final visual security assessment (VSA) score of the encrypted image data can be determined based on the visual security scores of T different resolutions.

[0062] In one example, determining the visual security evaluation score of encrypted image data based on T different resolutions can include:

[0063] The average of the visual security scores at T different resolutions is used as the visual security evaluation score for the encrypted image data.

[0064] For example, the visual security scores of T different resolutions can be averaged (including arithmetic average or weighted average) to determine the average value of the visual security scores of T different resolutions, and the average value of the visual security scores of T different resolutions can be determined as the visual security score of that resolution.

[0065] For example, the visual security assessment score for encrypted image data can be determined in the following ways:

[0066]

[0067] in, A visual security evaluation score for encrypted image data. Visual safety scores are given for different resolutions.

[0068] It can be seen that, in Figure 1 In the illustrated method, sensitive metadata is extracted from the input image data, and the target scene corresponding to the input image data is determined based on the extracted sensitive metadata. Personalized parameters corresponding to the target scene are then determined, and the input image is encrypted based on these personalized parameters to obtain encrypted image data. By performing scene recognition based on the extracted sensitive metadata and applying personalized encryption to the scene, personalized privacy protection is achieved, and the scene applicability and flexibility of privacy protection are improved. For both the input image data and the encrypted image data, downsampling processing is performed to obtain input image data and encrypted image data at various resolutions. Based on the perceptual features of each resolution of the input image data and encrypted image data, a perceptual feature similarity mapping between the input image data and the encrypted image data is determined. Based on this perceptual feature similarity mapping, a visual security score for each resolution is determined. Furthermore, based on the visual security scores for various resolutions, a visual security evaluation score for the encrypted image data is determined. By combining a perceptual feature-based image visual security evaluation method with multi-resolution visual feature analysis, the performance of security analysis is optimized, and the accuracy of image data security analysis is improved.

[0069] In some embodiments, determining the target scene corresponding to the input image data based on the extracted sensitive metadata includes:

[0070] Based on the extracted sensitive metadata, the scene privacy mapping table is queried to determine the target scene corresponding to the input image data.

[0071] For example, in order to improve the efficiency of scene determination, a scene privacy mapping table can be set in advance, which can record the correspondence between sensitive metadata and scenes.

[0072] In one example, the scene privacy mapping table can be a privacy dimension quantization matrix.

[0073] For example, the privacy dimension quantification matrix is ​​a tool for quantitatively modeling the risks of sensitive metadata through a multi-dimensional assessment system. Risk specifically refers to the likelihood of privacy breaches caused by the malicious use of sensitive metadata and the degree of potential harm, encompassing three core elements: probability of identification (P), intensity of harm (I), and amplification effect (A).

[0074] The specific form of the privacy dimension quantification matrix can be understood through the following example.

[0075]

[0076] In addition, the dynamic weighting mechanism can adjust the weights for different scenarios.

[0077] Based on the obtained sensitive metadata, the sensitive metadata is matched with the privacy dimension quantization matrix to obtain the final scene score, and then the corresponding scene is selected based on the final score.

[0078] In some embodiments, determining the personalized parameters corresponding to the target scenario based on the target scenario includes:

[0079] Based on the target scenario, determine the corresponding personalized parameter configuration strategy; among which, the personalized parameter configuration strategy includes preset mode, semi-automatic mode or expert mode;

[0080] Based on the target personalized parameter configuration strategy, determine the personalized parameters corresponding to the target scenario.

[0081] For example, the specific selection strategies for different modes can be as follows:

[0082] The system can determine whether the input data belongs to a preset standard business scenario. If so, it can be processed using the preset mode. If it belongs to a new or edge scenario, it will further check whether real-time adjustments are needed.

[0083] If real-time adjustments are required, you can choose between expert mode or semi-automatic mode depending on whether the expert is online.

[0084] If no real-time adjustments are required, the default mode will still be used and logs will be recorded. Audit logs will be generated for all decision-making processes, and in case of anomalies, the system will automatically revert to the default mode's basic configuration.

[0085] For example, a specific selection strategy matrix could be as follows:

[0086]

[0087] For example, personalized parameters corresponding to the target scenario can be determined based on the target personalized parameter configuration strategy.

[0088] For example, the categories and functions of personalized parameters can be as follows:

[0089]

[0090] In some embodiments, the perceived features include one or more of structural features, texture features, and naturalness features.

[0091] The aforementioned determination of the perceptual feature similarity mapping between the input image data and the encrypted image data at that resolution, based on the perceptual features of the input image data at that resolution and the perceptual features of the encrypted image data at that resolution, may include:

[0092] When the perceptual features include structural features, the structural features of the input image data at that resolution and the structural features of the encrypted image data at that resolution are determined based on the gradient amplitude and phase consistency features of the image data, respectively; based on the structural features of the input image data at that resolution and the structural features of the encrypted image data at that resolution, the structural feature similarity mapping between the input image data at that resolution and the encrypted image data at that resolution is determined; wherein, the phase consistency features of the image data are determined based on the amplitude and phase at each pixel position in the image data;

[0093] And / or,

[0094] When the perceived features include texture features, the texture features of the input image at that resolution and the texture features of the encrypted image at that resolution are determined based on the directional correlation between each pixel in the image data and its surrounding neighboring pixels. Based on the texture features of the input image data at that resolution and the texture features of the encrypted image data at that resolution, the texture feature similarity mapping between the input image data at that resolution and the encrypted image data at that resolution is determined.

[0095] And / or,

[0096] When the perceived features include natural features, the natural features of the input image data at that resolution and the natural features of the encrypted image data at that resolution are determined by removing the correlation between pixels in the image data. Based on the natural features of the input image data at that resolution and the natural features of the encrypted image data at that resolution, a natural feature similarity mapping between the input image data at that resolution and the encrypted image data at that resolution is determined.

[0097] It should be noted that, in the embodiments of this application, unless otherwise specified, the processing of input image data and encrypted image data refers to the processing of input image data and encrypted image data of the same resolution.

[0098] For example, when the perceptual features include structural features, the gradient magnitude (GM) and phase consistency (PC) of the image data can be combined to extract the structural features of the image data, and the changes in structural features between the input image data and the encrypted image data can be characterized by calculating the structural similarity mapping between the input image data and the encrypted image data.

[0099] For example, the magnitude of an image gradient can be defined as the change in pixel intensity. The gradient magnitude of an image can be represented by a vector that consists of the horizontal and vertical gradients of each pixel in the image, reflecting the maximum intensity of the structural change.

[0100] For example, the gradient magnitude of an image is defined as:

[0101]

[0102] in, For image pixel coordinates in The gradient in the horizontal direction. The gradient is in the vertical direction.

[0103] For example, the gradient in the horizontal direction and the gradient in the vertical direction can be determined respectively in the following ways:

[0104]

[0105] Where * represents the convolution operator, and These are the gradient operators in the horizontal and vertical directions, respectively.

[0106] For example, the edges of an image are essentially abrupt changes in pixel values. In the frequency domain, such abrupt changes require the superposition of multiple sine waves of different frequencies. These sine waves must be phase-aligned at the edge location to form a sharp transition. Therefore, edge regions exhibit phase consistency of multiple frequency components; while the phases of non-edge regions are disordered and cannot form a consistent superposition effect.

[0107] Phase consistency is a structural feature of image data extracted from the frequency domain of image data. It accurately locates structural features in the image by capturing the "synchronization signal" in the frequency domain. In other words, features with similar edges appear more frequently in the same phase.

[0108] In perceiving image structure, the human visual system is more efficient at analyzing in the frequency domain (phase and amplitude) than at processing in the spatial domain alone. Frequency domain methods can more stably capture cross-scale features and are more robust to changes in illumination and noise, while spatial domain processing, although direct and fast, is prone to failure in complex environments. In other words, it can be considered that the human visual system has an advantage in extracting structural information by utilizing the phase and amplitude of each frequency component in the image than in processing solely in the spatial domain.

[0109] Among them, phase consistency is not affected by changes in image brightness, compared to gradient magnitude.

[0110] For example, phase consistency can be determined for a frame of image data in the following way:

[0111]

[0112] in, This is a noise threshold used to further eliminate the effects of noise (because noise in the image can cause random fluctuations in the local phase). Used to filter noise below the noise threshold, reduce meaningless tiny phase alignments, suppress noise response, and improve signal-to-noise ratio. These are positive constants (i.e., constants greater than 0, typically taking extremely small values, much smaller than the sum of amplitudes under normal conditions, such as 10). −6 (), is used to prevent the denominator from being zero, which could lead to numerical instability and maintain stability. These are weighting parameters used to reduce image size. Frequency spread at location has an impact. and They are respectively Amplitude and phase at location, for Phase difference at the location.

[0113] For example, and The formula for calculating can be as follows:

[0114]

[0115]

[0116]

[0117] Wherein, the original image is denoted as The encrypted image is . and They are respectively in scale and direction Even-symmetric and odd-symmetric filters at the location. This represents the average value of the phase. Parameter It is a constant, usually taking the value of 0.65.

[0118] In one example, the structural features of the input image data at this resolution, and the structural features of the encrypted image data at this resolution, determined based on the gradient magnitude and phase consistency characteristics of the image data, may include:

[0119] For any pixel location in the input image data or encrypted image data of this resolution, the larger value between the phase consistency feature of the pixel location and the normalized gradient magnitude is determined as the structural feature of the pixel location; wherein, the normalized gradient magnitude of the pixel location is obtained by normalizing the gradient magnitude of the pixel location using the maximum gradient magnitude in the image data.

[0120] For example, considering that gradient magnitude or phase consistency alone cannot effectively reflect the structural degradation of encrypted image data, structural information of the image data can be extracted by combining gradient magnitude and phase consistency.

[0121] For example, structural information of image data It can be determined in the following ways:

[0122]

[0123] in, This represents the maximum value of the gradient magnitude in the image data. For the normalized gradient magnitude, To determine and The larger value in the range.

[0124] Based on the above formula, for any pixel position in the image The larger value between the normalized gradient magnitude and phase consistency can be determined as the structural feature of the pixel location.

[0125] For example, the structural feature similarity mapping (also known as structural similarity mapping) between input image data and encrypted image data can be determined in the following way:

[0126]

[0127] in, pixel position Structural feature similarity mapping, Pixel positions in the input image data Structural features, To encrypt pixel positions in image data The structural feature is that R is a small positive constant used to avoid the denominator converging to zero.

[0128] For example, when the perceived features include texture features, the texture features of the input image at that resolution and the texture features of the encrypted image at that resolution can be determined based on the directional correlation between each pixel in the image data and its surrounding neighboring pixels.

[0129] In one example, the OSPV (Optimal Sparse Variational Posterior) operator can be used to extract the texture structure of image data and determine the texture feature similarity mapping between the input image data and the encrypted image data.

[0130] For example, for images The gradient direction of each pixel can be calculated to represent the dominant direction of the pixel:

[0131]

[0132] For any pixel x, we can analyze the similarity between pixel x and its surrounding neighboring pixels.

[0133] For example, for any pixel x, its surrounding neighboring pixels may include the top-left neighboring pixel (if it exists), the top neighboring pixel (if it exists), the top-right neighboring pixel (if it exists), the left neighboring pixel (if it exists), the right neighboring pixel (if it exists), the bottom-left neighboring pixel (if it exists), the bottom neighboring pixel (if it exists), and the bottom-right neighboring pixel (if it exists).

[0134] For example, the relationship between pixels Represented using binary pattern:

[0135]

[0136] Here, 1 indicates that the two pixels are basically in the same direction, and 0 indicates that the two pixels have a large difference in direction. The threshold is used to determine the directional relationship between pixels.

[0137] For example, The value of can be π / 15.

[0138] Use the center pixel The relationship between a local region and its neighboring pixels is used to represent the orientation-selective visual pattern of the local region. :

[0139]

[0140] Where n represents the center pixel The size of the neighborhood, i.e., the number of surrounding neighboring pixels, for each element. Represents the center pixel With specific neighboring pixels The binary relation.

[0141] This can be achieved by considering the directional correlation between each pixel and its surrounding neighboring pixels, i.e. to map the size of the image Texture information (i.e., texture features):

[0142]

[0143] in, The pixel position in image I The pixels, where n is the pixel position in image I. The number of neighboring pixels around it.

[0144] For example, the texture feature similarity mapping (which can be called texture similarity mapping) between input image data and encrypted image data can be determined in the following way:

[0145]

[0146] in, pixel position Texture feature similarity mapping, Pixel positions in the input image data Texture features, To encrypt pixel positions in image data Texture features, It is a constant used to prevent the denominator from being 0.

[0147] In one example, where the perceived features include structural features and texture features, the technical solution provided in this application embodiment may further include:

[0148] Based on the structural feature similarity mapping between the input image data and the encrypted image data at this resolution, and the texture feature similarity mapping between the input image data and the encrypted image data at this resolution, the structural and texture similarity mapping between the input image data and the encrypted image data at this resolution is determined; wherein, the structural and texture similarity mapping between the input image data and the encrypted image data at this resolution is used to determine the visual security score at this resolution.

[0149] For example, considering that structural features reflect the overall geometric layout of an image, while texture features describe the distribution characteristics of details, combining structural and texture features can achieve feature complementarity, capturing both macroscopic structure and microscopic details, and avoiding misjudgments caused by a single feature.

[0150] Furthermore, since structural features are sensitive to geometric transformations and texture features are stable to changes in illumination, the combination of structural and texture features can resist various disturbances.

[0151] Furthermore, the combination of structural and textural features is more in line with human visual perception.

[0152] Based on this, structural feature similarity mapping and texture feature similarity mapping can be combined to obtain structural and texture similarity mapping.

[0153] For example, when the perceived features include natural features, the natural features of the image data can be determined by removing the correlation between pixels in the image data.

[0154] For example, image pixels generally have a certain correlation, but this correlation changes when the image is encrypted. When extracting image features, this correlation can mask some detailed information, so it is necessary to remove this correlation to extract the image features.

[0155] For example, research has found that the naturalness of digital images is a feature after removing this correlation, which can be represented by local averaging of the image to remove normalization:

[0156]

[0157] in, It is a constant used to prevent the denominator from being zero; and These represent the values ​​after convolving the image. The standard deviation and local mean of each pixel.

[0158] For example, and It can be determined in the following ways:

[0159]

[0160]

[0161] in, It is a two-dimensional cyclic symmetric Gaussian weighted function. K and L are parameters that define the spatial range of the convolution kernel (Gaussian window). K: controls the radius of the convolution kernel in the horizontal direction (x-axis), that is, it expands K pixels to the left and right from the center pixel. L: controls the radius of the convolution kernel in the vertical direction (y-axis), that is, it expands L pixels upwards and downwards from the center pixel. Represents image I in pixel coordinates The pixel value at that location.

[0162] For example, the natural feature similarity mapping (also known as natural similarity mapping) between input image data and encrypted image data can be determined in the following way:

[0163]

[0164] in, pixel position Natural feature similarity mapping, Pixel positions in the input image data The natural characteristics, To encrypt pixel positions in image data The natural characteristics, It is a constant used to prevent the denominator from being 0.

[0165] In some embodiments, determining the visual security score for a given resolution based on the perceptual feature similarity between the input image data and the encrypted image data at that resolution may include:

[0166] Determine the saliency map of the input image at this resolution;

[0167] Based on the saliency map of the input image at that resolution, and the perceptual feature similarity mapping between the input image data at that resolution and the encrypted image data, the perceptual feature similarity score between the input image data at that resolution and the encrypted image data is determined.

[0168] The visual security score for that resolution is determined based on the perceptual feature similarity score between the input image data and the encrypted image data at that resolution.

[0169] For example, a saliency map can be generated for the original image at each resolution to identify the regions sensitive to human vision. The structural and texture features of the original and encrypted images can be extracted separately, and the similarity mapping between the two can be calculated. The similarity mapping can be weighted and averaged using the saliency map to highlight the differences in important regions, thereby obtaining the perceptual feature similarity score at each resolution. Then, the visual security score at each resolution can be determined based on the perceptual feature similarity score between the input image data and the encrypted image data at each resolution.

[0170] The above method achieves the effect of simulating the human visual system's perception of privacy protection. It can identify the risk of "overall similarity but local leakage" that traditional encryption assessments may overlook, while avoiding the waste of resources caused by over-protecting non-sensitive areas. The final visual security score generated is more in line with actual privacy protection needs.

[0171] Saliency maps are used to represent the most salient and attention-grabbing parts of an image. The purpose of saliency maps is to simulate the areas that the human eye automatically focuses on when viewing an image. The human eye is typically more sensitive to certain elements, such as strong contrasts, vibrant colors, or unique textures. Saliency maps attempt to capture these features from a computer vision perspective.

[0172] For example, the methods for generating the display diagram may include, but are not limited to, the following:

[0173] 1) Based on contrast saliency: This method considers the most salient parts of an image to be those that differ the most from their surroundings. For example, areas that are more vibrant or have strong contrast between light and dark are considered salient areas.

[0174] 2) Based on color and texture saliency: This method identifies salient regions based on color and texture information. For example, certain colors (such as red) or specific textures (such as stripes or spots) tend to attract attention.

[0175] 3) Deep learning-based methods: Convolutional neural networks (CNNs) can be used for training to generate more accurate and complex saliency maps, which can automatically learn and identify salient regions in images.

[0176] 4) Region-based saliency: This method divides the image into several regions and then calculates the saliency of each region. Finally, a saliency map shows the location of these salient regions.

[0177] For example, the saliency map of the input image data can be determined in the following ways. :

[0178]

[0179] in, yes The pixel at position in the The stimulus is delivered to n feature channels, where n is the number of feature channels. It is a normalized function that combines these stimulus maps into a single map, namely the visual saliency map. .

[0180] In one example, the perceptual feature similarity score between the input image data and the encrypted image data can be determined by weighting the perceptual feature similarity mapping between the input image data and the encrypted image data based on the saliency map of the input image data at that resolution.

[0181] Taking perceptual features, including natural features, as an example, the natural feature similarity score between input image data and encrypted image data can be determined in the following way:

[0182]

[0183] in, This is the natural feature similarity score between the input image data and the encrypted image data. pixel position Natural feature similarity mapping.

[0184] In one example, when the perceptual features include multiple features such as structural features, texture features, and naturalness features, the determination of the visual security score for that resolution based on the perceptual feature similarity score between the input image data and the encrypted image data at that resolution may include:

[0185] The weighted average of the similarity scores of various types of perceptual features between the input image data at this resolution and the encrypted image data is determined as the visual security score for this resolution.

[0186] For example, when the perceptual features include multiple types such as structural features, texture features, and natural features, the similarity scores of various types of perceptual features between the input image data and the encrypted image data at that resolution can be determined separately, and the weighted average of the similarity scores of various types of perceptual features between the input image data and the encrypted image data at that resolution can be determined as the visual security score at that resolution.

[0187] In one example, perceptual features include structural features, texture features, and naturalness features.

[0188] The determination of the perceptual feature similarity score between the input image data and the encrypted image data based on the saliency map of the input image at that resolution and the perceptual feature similarity mapping between the input image data and the encrypted image data at that resolution may include:

[0189] Based on the structural feature similarity mapping between the input image data and the encrypted image data at this resolution, and the texture feature similarity mapping between the input image data and the encrypted image data at this resolution, the structural and texture similarity mapping between the input image data and the encrypted image data at this resolution is determined.

[0190] Based on the saliency map of the input image at that resolution, and the structural and texture similarity mapping between the input image data at that resolution and the encrypted image data, a structural and texture similarity score between the input image data at that resolution and the encrypted image data is determined; and,

[0191] Based on the saliency map of the input image at that resolution, and the natural feature similarity mapping between the input image data at that resolution and the encrypted image data, a natural feature similarity score between the input image data at that resolution and the encrypted image data is determined.

[0192] The visual security score for that resolution is determined based on the perceptual feature similarity score between the input image data and the encrypted image data at that resolution, including:

[0193] The visual security score for that resolution is determined by the weighted average of the structural and texture similarity scores between the input image data and the encrypted image data at that resolution, and the naturalness feature similarity scores between the input image data and the encrypted image data at that resolution.

[0194] For example, when the perceived features include structural features, texture features, and natural features, the structural feature similarity mapping and texture feature similarity mapping determined in the manner described above can be combined into a structural and texture similarity mapping.

[0195] On the one hand, the structural and texture similarity score between the input image data and the encrypted image data can be determined based on the saliency map of the input image and the structural and texture similarity mapping between the input image data and the encrypted image data.

[0196] On the other hand, the natural feature similarity score between the input image data and the encrypted image data can be determined based on the saliency map of the input image and the natural feature similarity mapping between the input image data and the encrypted image data.

[0197] Furthermore, for any resolution, the visual security score for that resolution can be determined by the weighted average of the structural and texture similarity scores between the input image data and the encrypted image data at that resolution, and the naturalness feature similarity scores between the input image data and the encrypted image data at that resolution.

[0198] For example, for any given resolution, the visual safety score for that resolution can be determined in the following way. :

[0199]

[0200] in, This represents the structural and texture similarity score between the input image data and the encrypted image data at this resolution. Score the natural feature similarity between the input image data and the encrypted image data at this resolution. These are weight parameters used to adjust... and Relative importance.

[0201] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, the technical solutions provided in the embodiments of this application are described below with reference to specific examples.

[0202] In this embodiment, a security analysis scheme for personalized privacy protection of image data is designed, taking into account the problem that current image or video privacy protection technologies typically adopt a one-size-fits-all approach and the importance of visual security.

[0203] Security analysis solutions for personalized privacy protection of image data can include: pre-analyzing the privacy concerns emphasized in different application scenarios, such as personal cloud storage and short video platforms, and adjusting the parameter weights in the security assessment accordingly. By constructing a full-process system of scene awareness, dynamic parameter tuning, and quality compensation, accurate and efficient image or video privacy protection can be achieved.

[0204] For example, sensitive information can be accurately located by determining the data type, and the corresponding protection strength can be selected based on the characteristics of the scenario. A tiered protection mechanism is employed during processing, prioritizing high-risk areas, reducing unnecessary computational consumption, improving processing speed, and significantly minimizing the impact on the original image quality. This approach can quickly meet compliance requirements and flexibly adapt to special needs, significantly improving processing efficiency and data usability while ensuring data security.

[0205] In the specific security assessment of each scenario, the image visual security assessment based on perceptual features is mainly adopted.

[0206] This scheme can use the changes in the structure, texture, and naturalness features of images at different resolutions to calculate the visual security of encrypted images.

[0207] For example, multi-resolution models can be established to simulate the hierarchical features of the human visual system, and structural and texture features of the original and encrypted images can be extracted. The naturalness of the images can then be represented by a local averaging and normalization operation. Furthermore, by averaging the visual security scores across all multi-resolution scales, a final visual security score can be calculated, enabling personalized image data security analysis.

[0208] For example, in this embodiment, in order to achieve secure analysis for the protection of personalized privacy in image data, the system may include seven modules:

[0209] The first to the third modules achieve personalized analysis by building a complete process system of scene perception, dynamic parameter tuning, and quality compensation.

[0210] For example, the first module is a scene perception and feature extraction module; the second module is a personalized parameter configuration module; and the third module is a dynamic privacy processing engine.

[0211] The fourth to seventh modules implement specific security analysis.

[0212] For example, the fourth module is a downsampling module, which uses multi-resolution to represent the hierarchical features of the human visual system; the fifth module is a perceptual feature extraction module, which is used to extract the structural features, texture features and naturalness features of the image; the sixth module is a weighted scoring module, which is used to weight the above feature data to obtain a visual safety score; and the seventh module is a comprehensive scoring module, which is used to average the visual safety scores at different resolutions to obtain the final visual safety evaluation score.

[0213] The above seven modules enable security analysis aimed at protecting the personalized privacy of image data.

[0214] An exemplary overall framework diagram of a security analysis scheme for protecting the personalized privacy of image data can be found in [reference needed]. Figure 2 .

[0215] For example, such as Figure 3 As shown, in this embodiment, a security analysis scheme for protecting the personalized privacy of image data may include the following steps:

[0216] S1: Implement scene awareness and feature extraction.

[0217] S1.1: The input data can be classified and sensitive metadata extracted using a multimodal scene recognition engine.

[0218] Examples of such scenarios may include, but are not limited to, personal cloud storage or short video platforms.

[0219] S1.2: Establish a scenario privacy mapping table and design a privacy dimension quantification matrix.

[0220] S2: Build a personalized parameter configuration system.

[0221] S2.1: The three-level parameter configuration system dynamically loads processing strategies based on the scenario type.

[0222] For example, the three-level parameter configuration mode may include a preset mode (calling pre-trained industry templates), a semi-automatic mode (based on contrastive learning-based scene matching combined with automatically generated parameter suggestions), and an expert mode (providing real-time visual feedback).

[0223] For example, the system can determine which scene an image belongs to through the scene perception and feature extraction module, then automatically match the corresponding privacy processing strategy according to the scene type and processing strategy mapping table, and dynamically load the corresponding parameters according to the selected configuration mode (preset, semi-automatic, expert).

[0224] For example, if a configuration mode selection instruction is detected, such as a configuration mode selection instruction manually entered by the user, the configuration mode can be determined based on the configuration mode selection instruction; if no configuration mode selection instruction is detected, the configuration mode can be automatically selected by the system.

[0225] S2.2: Achieve cross-scenario knowledge transfer through a federated learning framework, update the global model every 24 hours under differential privacy protection, and continuously optimize parameter recommendation accuracy.

[0226] For example, a federated learning framework can be introduced to achieve cross-scenario knowledge transfer and parameter recommendation optimization.

[0227] Specifically, users process and fine-tune their image data and models locally on their devices. The system does not collect the original images but only retains the parameter updates of the local model. Every 24 hours, these local updates are uploaded to a central server after adding differential privacy noise, thus protecting user privacy from being reverse-engineered. The central server aggregates model updates from different terminals to generate a new global model that incorporates privacy processing experience from various scenarios. Subsequently, the global model is synchronously distributed to all terminals to optimize the parameter recommendation effect in the next round. In this way, the system can continuously improve its adaptability to diverse scenarios while protecting individual data security, achieving more intelligent and personalized privacy processing strategy recommendations.

[0228] S3: Employs a selective processing pipeline through a dynamic privacy engine, combined with a quality compensation algorithm.

[0229] S3.1: In the core processing phase, the dynamic privacy engine employs a selective processing pipeline.

[0230] For example, areas with risk values ​​exceeding a threshold can be strongly encrypted or obscured, while other areas retain the original data or are only slightly anonymized.

[0231] Among them, "risk value" is an indicator that can represent the sensitivity or exposure risk of a certain area in an image in terms of privacy leakage.

[0232] For example, a higher risk value indicates that the area is more likely to expose users' sensitive information, and therefore requires stronger protection measures.

[0233] Combining several deployment models significantly reduces processing latency, while parallelizing sensitive area detection and processing improves throughput.

[0234] For example, by deploying lightweight, optimized local models to replace cloud computing, and combining model compression and inference acceleration technologies, the system's response time and latency in processing images can be significantly reduced, enabling faster privacy detection and protection.

[0235] S3.2: Employ a quality compensation algorithm to perform super-resolution reconstruction on non-sensitive areas and feather the edges of sensitive areas.

[0236] For example, since privacy processing can reduce image quality or even create a sense of abruptness, quality compensation algorithms can restore the visual integrity and naturalness of an image as much as possible while ensuring privacy protection.

[0237] For example, assuming a selfie contains a face (sensitive area) and a background landscape (non-sensitive area), quality compensation can be achieved in the following ways:

[0238] Step 1: Classify the detection area.

[0239] A privacy detection model is used to mark faces as sensitive areas. Other background elements such as the sky, buildings, and clothing are considered non-sensitive areas.

[0240] Step 2: Sensitive area processing (blurring + feathering edges).

[0241] Apply Gaussian blur to the face area (to protect privacy). To avoid abrupt blur boundaries, feather the edges of the face, i.e., gradually change the transparency of the boundaries to achieve a smooth transition.

[0242] Step 3: Super-resolution reconstruction of non-sensitive areas.

[0243] For unblurred background areas, a super-resolution reconstruction model is used to enhance the image to a higher resolution. Even if the original image is heavily compressed, more details can be recovered in non-sensitive areas.

[0244] S4: Simulates the hierarchical features of the human visual system through multi-resolution representation.

[0245] For example, the hierarchical features of the human visual system can be simulated by establishing multi-resolution representations.

[0246] For example, for a frame I, it can be downsampled T-1 times to obtain T frames of images at different resolutions: ;in, This is the initial image.

[0247] S5. Extract the structural, textural, and natural features of the image.

[0248] S5.1 Combine gradient magnitude and phase consistency to extract structural features of the image; represent the changes in structural features by calculating the structural similarity mapping between the original image (i.e., the image before encryption) and the selectively encrypted image (i.e., the image after encryption).

[0249] For example, the magnitude of an image gradient can be defined as the change in pixel intensity. The gradient magnitude of an image can be represented by a vector that consists of the horizontal and vertical gradients of each pixel in the image, reflecting the maximum intensity of the structural change.

[0250] For example, the gradient magnitude of an image is defined as:

[0251]

[0252] in, Let I be the pixel coordinates in image I. The gradient in the horizontal direction. The gradient is in the vertical direction.

[0253] For example, the gradient in the horizontal direction and the gradient in the vertical direction can be determined respectively in the following ways:

[0254]

[0255] Where * represents the convolution operator, and These are the gradient operators in the horizontal and vertical directions, respectively.

[0256] For example, phase consistency is a structural feature of image data extracted from the frequency domain of the image data, meaning that features with similar edges appear more frequently in the same phase.

[0257] It is assumed that the human visual system is better at extracting structural information by utilizing the phase and amplitude of various frequency components in an image than by processing it spatially.

[0258] Among them, phase consistency is not affected by changes in image brightness, compared to gradient magnitude.

[0259] For example, phase consistency can be determined for a frame of image data in the following way:

[0260]

[0261] in, This is the noise threshold, used to further eliminate the influence of noise. These are positive numbers (usually taking very small values) used to maintain stability. These are weighting parameters used to reduce image size. Frequency spread at location has an impact. and They are respectively Amplitude and phase at location, for Phase difference at the location.

[0262] Considering that gradient magnitude or phase consistency alone cannot effectively reflect the structural degradation of encrypted image data, structural information of the image data can be extracted by combining gradient magnitude and phase consistency.

[0263] For example, structural information of image data It can be determined in the following ways:

[0264]

[0265] in, This represents the maximum value of the gradient magnitude in the image data. For the normalized gradient magnitude, To determine and The larger value in the range.

[0266] Based on the above formula, for any pixel position in the image The larger value between the normalized gradient magnitude and phase consistency can be determined as the structural feature of the pixel location.

[0267] For example, the structural feature similarity mapping (also known as structural similarity mapping) between input image data and encrypted image data can be determined in the following way:

[0268]

[0269] in, pixel position Structural feature similarity mapping, Pixel positions in the input image data Structural features, To encrypt pixel positions in image data The structural feature is that R is a constant used to prevent the denominator from being 0.

[0270] S5.2. Use the OSVP operator to extract the texture structure of the image and calculate the texture similarity mapping between the original image and the selectively encrypted image.

[0271] For example, for images The gradient direction of each pixel can be calculated to represent the dominant direction of the pixel:

[0272]

[0273] For any pixel x, we can analyze the similarity between pixel x and its surrounding neighboring pixels.

[0274] For example, the relationship between pixels Represented using binary pattern:

[0275]

[0276] Here, 1 indicates that the two pixels are basically in the same direction, and 0 indicates that the two pixels have a large difference in direction. The threshold is used to determine the directional relationship between pixels.

[0277] For example, The value of can be π / 15.

[0278] Use the center pixel The relationship between a local region and its neighboring pixels is used to represent the orientation-selective visual pattern of the local region. :

[0279]

[0280] Where n represents the size of the neighborhood of the center pixel x, that is, the number of surrounding neighboring pixels.

[0281] This can be achieved by considering the directional correlation between each pixel and its surrounding neighboring pixels, i.e. to map the size of the image Texture information (i.e., texture features):

[0282]

[0283] For example, the texture feature similarity mapping (which can be called texture similarity mapping) between input image data and encrypted image data can be determined in the following way:

[0284]

[0285] in, pixel position Texture feature similarity mapping, Pixel positions in the input image data Texture features, To encrypt pixel positions in image data Texture features, It is a constant used to prevent the denominator from being 0.

[0286] For example, structural feature similarity mapping can be used. Texture feature similarity mapping By combining these, we obtain a structure and texture similarity mapping:

[0287]

[0288] S5.3: Extract the natural features of the image and represent the changes in natural features by calculating the natural similarity mapping between the original image and the selectively encrypted image.

[0289] For example, image pixels generally have a certain correlation, but this correlation changes when the image is encrypted. When extracting image features, this correlation can mask some detailed information, so it is necessary to remove this correlation to extract the image features.

[0290] For example, research has found that the naturalness of digital images is a feature after removing this correlation, which can be represented by local averaging of the image to remove normalization:

[0291]

[0292] in, It is a constant used to prevent the denominator from being zero; and These represent the values ​​after convolving the image. The standard deviation and local mean of each pixel.

[0293] For example, and It can be determined in the following ways:

[0294]

[0295]

[0296] in, It is a two-dimensional cyclic symmetric Gaussian weighted function. K and L are parameters that define the spatial range of the convolution kernel (Gaussian window). K: controls the radius of the convolution kernel in the horizontal direction (x-axis), that is, it expands K pixels to the left and right from the center pixel. L: controls the radius of the convolution kernel in the vertical direction (y-axis), that is, it expands L pixels upwards and downwards from the center pixel. Represents image I in pixel coordinates The pixel value at that location.

[0297] For example, the natural feature similarity mapping (also known as natural similarity mapping) between input image data and encrypted image data can be determined in the following way:

[0298]

[0299] in, pixel position Natural feature similarity mapping, Pixel positions in the input image data The natural characteristics, To encrypt pixel positions in image data The natural characteristics, It is a constant used to prevent the denominator from being 0.

[0300] S6. Use the saliency map of the original image to weight the feature similarity mapping to obtain visual security scores at different resolutions.

[0301] For example, since for any resolution, the result obtained in the manner described above... and It has the same resolution as the image data, and since the number used to assess the visual security of an image is generally a single number, it can be... and The mapping is represented by two scores (which can be called structural and texture similarity scores and naturalness similarity scores, respectively), to characterize the two similarity mappings.

[0302] For example, the saliency map of the original image can be compared with... and By combining these, we obtain the structural and texture similarity score and the naturalness similarity score:

[0303]

[0304]

[0305] in, The structural and texture similarity score between the input image data and the encrypted image data. pixel position Structure and texture similarity mapping. This is the natural feature similarity score between the input image data and the encrypted image data. pixel position Natural feature similarity mapping.

[0306] For any given resolution, the visual security score for that resolution can be determined by the weighted average of the structural and texture similarity scores between the input image data and the encrypted image data at that resolution, and the naturalness feature similarity scores between the input image data and the encrypted image data at that resolution.

[0307] For example, for any given resolution, the visual safety score for that resolution can be determined in the following way. :

[0308]

[0309] in, This represents the structural and texture similarity score between the input image data and the encrypted image data at this resolution. Score the natural feature similarity between the input image data and the encrypted image data at this resolution. These are weight parameters used to adjust... and Relative importance.

[0310] S7. The final visual security score is calculated by averaging the visual security scores across all multi-resolution scales.

[0311] For example, for each pair of original image O and encrypted image E, at different resolutions Each of these can yield a visual safety score (VS). Therefore, T visual safety scores can be obtained. .

[0312] The final Visual Security Assessment (VSA) score is obtained by averaging the visual security scores of all image pairs at different resolutions.

[0313]

[0314] As can be seen, in this embodiment, by adopting a personalized image data security analysis scheme, more accurate security analysis of image data can be achieved in different scenarios.

[0315] Secondly, by adopting an image visual security evaluation method based on perceptual features, the similarity of three perceptual features—image structure, texture, and naturalness—is calculated to obtain visual security scores at different resolutions. Compared with classic evaluation schemes, this method has very good performance.

[0316] Finally, by constructing a dual innovation mechanism of scenario customization and visual perception quantification, three major improvements have been achieved in the field of privacy protection:

[0317] First, it automatically identifies sensitive areas for different scenarios and uses multi-resolution visual feature analysis to simulate the perceptual characteristics of the human eye. By quantifying changes in features such as structure and texture, it evaluates the encryption effect, ensuring high-strength protection of critical information while preserving the original details of non-sensitive areas to the maximum extent. This solves the dilemma of traditional methods where "over-desensitization leads to data invalidation" or "insufficient protection leads to leakage."

[0318] Secondly, by dynamically adjusting parameters based on scene characteristics and in conjunction with a hierarchical protection mechanism, complex calculations are only performed on high-risk areas, which saves a great deal of computing resources compared to global processing, greatly improving video processing speed and significantly reducing hardware costs.

[0319] Third, regarding compliance adaptation and personalized expansion, it has built-in compliance templates for multiple industries, which can generate compliant and de-identified data with one click. It also supports custom parameter weights (the dynamic weight mechanism can adjust the weights for different scenarios) to meet the needs of special scenarios.

[0320] The system includes pre-set data anonymization guidelines (templates) for multiple industries. Each template contains the types of information that must be protected, recommended processing methods, and compliance guidelines. Users simply need to select the desired industry type in the interface or API, and the system will automatically load the corresponding parameters. The system will then automatically identify the sensitive elements in the image and perform anonymization according to industry standards.

[0321] This evaluation system breaks away from the limitations of traditional fixed thresholds. Through multi-scale feature fusion scoring, it enables the acquisition of security solutions tailored to the data characteristics of different scenarios. By transforming the laws of human visual perception into a computable visual security scoring model, it breaks through the one-size-fits-all approach and provides a quantitative tool and implementation path for dynamically balancing the privacy and practical value of image data.

[0322] The method provided in this application has been described above. The apparatus provided in this application is described below:

[0323] Please see Figure 4 This is a schematic diagram of the structure of a security analysis device for protecting the personalized privacy of image data, as provided in an embodiment of this application. Figure 4 As shown, the security analysis device for protecting the individual privacy of image data may include:

[0324] The first determining unit is used to extract sensitive metadata from the input image data and determine the target scene corresponding to the input image data based on the extracted sensitive metadata.

[0325] The second determining unit is used to determine personalized parameters corresponding to the target scene based on the target scene;

[0326] An encryption unit is used to encrypt the input image according to the personalized parameters to obtain encrypted image data;

[0327] A sampling unit is used to perform downsampling processing on the input image data and the encrypted image data respectively to obtain T kinds of input image data with different resolutions and T kinds of encrypted image data with different resolutions; wherein, T≥2;

[0328] The third determining unit is used to determine the perceptual feature similarity mapping between the input image data and the encrypted image data at any resolution, based on the perceptual features of the input image data at that resolution and the perceptual features of the encrypted image data at that resolution, and to determine the visual security score at that resolution based on the perceptual feature similarity mapping between the input image data and the encrypted image data at that resolution.

[0329] The third determining unit is further configured to determine the visual security evaluation score of the encrypted image data based on the T different resolution visual security scores.

[0330] For example, the specific implementation process of the security analysis scheme for image data personal privacy protection in each unit of the above-mentioned security analysis device can be found in the relevant description in the above embodiments, and will not be repeated here in the embodiments of this application.

[0331] This application also provides an electronic device, including a processor and a memory, wherein the memory is used to store computer programs; and the processor is used to execute the program stored in the memory to implement the security analysis method for image data personalized privacy protection described above.

[0332] Please see Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device may include a processor 501 and a memory 502 storing machine-executable instructions. The processor 501 and the memory 502 can communicate via a system bus 503. Furthermore, by reading and executing the machine-executable instructions corresponding to the security analysis logic for image data personal privacy protection stored in the memory 502, the processor 501 can execute the security analysis method for image data personal privacy protection described above.

[0333] The memory 502 mentioned in this document can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For example, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0334] In some embodiments, a machine-readable storage medium, such as Figure 5 The memory 502 in the memory, which is a machine-readable storage medium, stores machine-executable instructions. When these machine-executable instructions are executed by a processor, they implement the security analysis method for protecting the personalized privacy of image data described above. For example, the machine-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0335] This application also provides a computer program product that stores a computer program, and when a processor executes the computer program, it causes the processor to execute the security analysis method for protecting the individual privacy of image data described above.

Claims

1. A security analysis method for image data individual privacy protection, characterized in that, The method comprises the following steps: Sensitive metadata extraction is performed on input image data, and a target scene corresponding to the input image data is determined according to the extracted sensitive metadata; The input image data comprises a picture or a video; An individualized parameter corresponding to the target scene is determined according to the target scene; An encryption process is performed on the input image according to the individualized parameter, and encrypted image data is obtained; Down-sampling processes are respectively performed on the input image data and the encrypted image data, and T types of input image data with different resolutions and T types of encrypted image data with different resolutions are obtained; wherein T≥2; For input image data and encrypted image data of any resolution, a perceptual feature similarity mapping between the input image data and the encrypted image data of the resolution is determined according to the perceptual feature of the input image data of the resolution and the perceptual feature of the encrypted image data of the resolution, and a visual security score of the resolution is determined according to the perceptual feature similarity mapping between the input image data and the encrypted image data of the resolution; A visual security evaluation score of the encrypted image data is determined according to the T types of visual security scores of different resolutions. The perceptual feature similarity mapping between the input image data and the encrypted image data of the resolution is determined according to the perceptual feature of the input image data of the resolution and the perceptual feature of the encrypted image data of the resolution, which comprises the following steps: The structural features of the input image data of the resolution and the structural features of the encrypted image data of the resolution are respectively determined according to the gradient amplitude and phase consistency features of the image data, and the structural feature similarity mapping between the input image data and the encrypted image data of the resolution is determined according to the structural features of the input image data of the resolution and the structural features of the encrypted image data of the resolution; wherein the phase consistency features of the image data are determined according to the amplitude and phase of each pixel position in the image data; The texture features of the input image data of the resolution and the texture features of the encrypted image data of the resolution are respectively determined according to the direction correlation of each pixel and the surrounding adjacent pixels in the image data, and the texture feature similarity mapping between the input image data and the encrypted image data of the resolution is determined according to the texture features of the input image data of the resolution and the texture features of the encrypted image data of the resolution; The naturalness features of the input image data of the resolution and the naturalness features of the encrypted image data of the resolution are respectively determined by removing the correlation between the pixels in the image data, and the naturalness feature similarity mapping between the input image data and the encrypted image data of the resolution is determined according to the naturalness features of the input image data of the resolution and the naturalness features of the encrypted image data of the resolution.

2. The method of claim 1, wherein, The target scene corresponding to the input image data is determined according to the extracted sensitive metadata, which comprises the following steps: The target scene corresponding to the input image data is determined by querying a scene privacy mapping table according to the extracted sensitive metadata. The scene privacy mapping table comprises a privacy dimension quantization matrix, and the privacy dimension quantization matrix is used for quantitatively modeling the risk of sensitive metadata by a multi-dimension evaluation system.

3. The method of claim 1, wherein, The individualized parameter corresponding to the target scene is determined according to the target scene, and the individualized parameter corresponding to the target scene comprises: According to the target scene, a target individualized parameter configuration strategy is determined, wherein the individualized parameter configuration strategy comprises a preset mode, a semi-automatic mode or an expert mode. According to the target individualized parameter configuration strategy, the individualized parameter corresponding to the target scene is determined.

4. The method of claim 1, wherein, The gradient amplitude and the phase consistency feature of the image data are used to determine the structure feature of the input image data of the resolution and the structure feature of the encrypted image data of the resolution, respectively, and the method comprises the following steps: For any pixel position in the input image data or the encrypted image data of the resolution, the larger value between the phase consistency feature of the pixel position and the normalized gradient amplitude is determined as the structure feature of the pixel position, wherein the normalized gradient amplitude of the pixel position is obtained by normalizing the gradient amplitude of the pixel position by the maximum gradient amplitude in the image data.

5. The method of claim 1, wherein, In the case that the perception feature comprises the structure feature and the texture feature, the method further comprises: According to the structure feature similarity mapping between the input image data and the encrypted image data of the resolution and the texture feature similarity mapping between the input image data and the encrypted image data of the resolution, the structure and texture similarity mapping between the input image data and the encrypted image data of the resolution is determined, wherein the structure and texture similarity mapping between the input image data and the encrypted image data of the resolution is used to determine the visual security score of the resolution.

6. The method of claim 1, wherein, The visual security score of the resolution is determined according to the perception feature similarity between the input image data and the encrypted image data of the resolution, and the method comprises the following steps: The saliency map of the input image of the resolution is determined. According to the saliency map of the input image of the resolution and the perception feature similarity mapping between the input image data and the encrypted image data of the resolution, the perception feature similarity score between the input image data and the encrypted image data of the resolution is determined. According to the perception feature similarity score between the input image data and the encrypted image data of the resolution, the visual security score of the resolution is determined.

7. The method of claim 6, wherein, In the case that the perception feature comprises multiple types of structure feature, texture feature and naturalness feature, the visual security score of the resolution is determined according to the perception feature similarity score between the input image data and the encrypted image data of the resolution, and the method comprises the following steps: The weighted average value of the multiple different types of perception feature similarity scores between the input image data and the encrypted image data of the resolution is determined as the visual security score of the resolution.

8. The method of claim 6, wherein, The perception feature comprises the structure feature, the texture feature and the naturalness feature. The saliency map of the input image at the resolution, and the perceptual feature similarity mapping between the input image data at the resolution and the encrypted image data, determine a perceptual feature similarity score between the input image data at the resolution and the encrypted image data, comprising: The structural feature similarity mapping between the input image data at the resolution and the encrypted image data, and the texture feature similarity mapping between the input image data at the resolution and the encrypted image data, determine a structural and texture similarity mapping between the input image data at the resolution and the encrypted image data; The saliency map of the input image at the resolution, and the structural and texture similarity mapping between the input image data at the resolution and the encrypted image data, determine a structural and texture similarity score between the input image data at the resolution and the encrypted image data; and, The saliency map of the input image at the resolution, and the natural feature similarity mapping between the input image data at the resolution and the encrypted image data, determine a natural feature similarity score between the input image data at the resolution and the encrypted image data; The perceptual feature similarity score between the input image data at the resolution and the encrypted image data, determine a visual security score at the resolution, comprising: The weighted average of the structural and texture similarity score between the input image data at the resolution and the encrypted image data, and the natural feature similarity score between the input image data at the resolution and the encrypted image data, is determined as the visual security score at the resolution.

9. The method of claim 1, wherein, The visual security scores at the T different resolutions, determine a visual security evaluation score of the encrypted image data, comprising: The average of the visual security scores at the T different resolutions, is determined as the visual security evaluation score of the encrypted image data.

10. An electronic device, comprising: The computer program product comprises a processor and a memory, wherein, The memory is used to store a computer program; The processor is used to execute the program stored on the memory, and realize the method of any one of claims 1-9.

11. A computer program product, characterised in that, The computer program product stores a computer program, and the computer program is executed by the processor to realize the method of any one of claims 1-9. The computer program product stores a computer program, and the computer program is executed by the processor to realize the method of any one of claims 1-9.

Citation Information

Patent Citations

  • Image privacy protection method and system

    CN108197453A

  • Perception visual safety assessment method and system

    CN112465028A

  • Hash-based semi-parameter perception encryption visual security analysis method and system

    CN114240827A

  • Visual security evaluation method for selectively encrypted image

    CN116740388A

Cited By

  • Secure transmission method and system for intelligent network connection vehicle data and electronic equipment

    CN121966994A