A dynamic data desensitization method based on multi-dimensional perturbation mapping
Patent Information
- Application Number
- CN202611034675.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-29
AI Technical Summary
该类方式虽然能够隐藏原始数据,但在多终端访问环境下,仍容易通过跨设备截图比对、展示日志拼接、字段顺序映射及页面结构关联等方式推断敏感字段之间的真实对应关系
本发明,通过获取访问员工编号、访问设备编号以及页面字段布局关系,构建设备侧脱敏转子、字段邻接集合和字段暴露位序图,并结合页面中的字段位置关系、字段类型关联关系以及业务关联关系识别结构暴露位置,实现了敏感字段在不同员工、不同设备及不同页面场景下的差异化动态脱敏处理,能够有效降低传统固定脱敏方式中因跨设备截图比对、展示日志拼接和字段位置反推导致的敏感信息关联泄露风险,提高数据展示过程中的结构隐匿能力。
Smart Images

Figure CN122839428A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data security processing technology, and in particular to a dynamic data desensitization method based on multidimensional perturbation mapping. Background Technology
[0002] With the continuous improvement of enterprise digital transformation and data asset management systems, employee information, equipment information, business records, and log data are increasingly used in operations management, business collaboration, remote work, and data analysis. To prevent the leakage of sensitive data during display, sharing, and access, it is typically necessary to anonymize sensitive fields such as employee IDs, equipment IDs, customer information, and business identifiers. Existing data anonymization technologies are widely used in enterprise management platforms, cloud service platforms, industrial internet systems, and multi-terminal collaborative systems to meet data security and privacy protection requirements.
[0003] Currently, most common data anonymization methods employ fixed masks, uniform replacements, or static hash mappings to generate consistent anonymization results for the same sensitive field across different employees, devices, and page scenarios. While these methods can hide the original data, in multi-terminal access environments, the true correspondence between sensitive fields can still be inferred through cross-device screenshot comparisons, log splicing, field order mapping, and page structure associations. Furthermore, existing solutions typically only anonymize the field content itself, lacking joint analysis of field adjacency relationships, page layout structure, and business relationships. This makes it difficult to effectively block the path of data restoration based on structural exposure locations, resulting in a high risk of data leakage in complex business scenarios. Summary of the Invention
[0004] This invention provides a dynamic data desensitization method based on multidimensional perturbation mapping. By constructing a device-side desensitization rotor, a field exposure sequence diagram, and segment hierarchical perturbation rules, it performs differentiated dynamic perturbation processing on different segments in the sensitive field to be displayed. Combined with non-real association perturbation tags, it achieves cross-device structural interference, so that the same sensitive field generates different desensitization display results under different employees, different devices, and different page scenarios. This effectively blocks the leakage of sensitive information caused by cross-device association inference, screenshot comparison, and log splicing, while taking into account both business readability and data security.
[0005] A dynamic data anonymization method based on multidimensional perturbation mapping includes the following steps: S1, obtain the access employee ID, access device ID, sensitive field to be displayed, and the relationship between adjacent fields of the sensitive field to be displayed in the current data access request; perform Hash fusion on the access employee ID and access device ID to generate a device-side desensitization rotor bound to the current access employee and the current access device; and identify the structural exposure position of the sensitive field to be displayed in the page that can be compared across devices based on the relationship between adjacent fields, and generate a field exposure position chart. S2, based on the device-side desensitization rotor, perform perturbation mapping on the field exposure sequence diagram, determine the anchored segments that need to retain business readability and the free segments that need to block cross-device association in the sensitive fields to be displayed, and assign perturbation paths different from other devices to the free segments according to the device-side desensitization rotor, and generate segment layer perturbation rules exclusive to the current employee under the current device; S3, perform dynamic desensitization processing on the sensitive fields to be displayed according to the segment layering perturbation rules, so that the anchored segment maintains the local consistency required for business identification, so that the free segment presents different desensitization forms on different devices, and insert non-real association perturbation marks controlled by the device-side desensitization rotor between adjacent fields, generating dynamic desensitization display results for blocking cross-device screenshot comparison, display log splicing, and field position reverse inference.
[0006] Optionally, S1 includes: S11, receive the current data access request, extract the accessing employee number from the login session, permission token or access log, extract the accessing device number from the terminal registration information, device fingerprint or client certificate, and extract the sensitive fields to be displayed from the page rendering task. At the same time, read the page field layout configuration, determine the adjacent fields before and after the sensitive fields to be displayed, the fields in the same column, the fields in the same group, and the related fields in the same business card, and form a field adjacency set. S12, perform format unification, invalid character removal and encoding standardization on the access employee number and access device number respectively, and concatenate the standardized access employee number and standardized access device number with the page identifier, input the Hash function to obtain the device-side desensitization rotor, so that the desensitization rotor is simultaneously constrained by employee identity, device identity and page scenario; S13, based on the field adjacency set, calculates the display distance, arrangement order, field type similarity and business association strength of the sensitive field to be displayed and its adjacent fields on the page, identifies the structural exposure location by comparison of screenshots from different devices, splicing of display logs or reverse inference of field position, and uses the sensitive field to be displayed and its adjacent fields as nodes, and uses the adjacency, same column, same group or business association relationship between fields as edges to generate a field exposure position graph.
[0007] Optionally, S12 includes: S121, perform format unification, invalid character removal and encoding standardization processing on the access employee number and access device number respectively, and convert employee numbers and device numbers from different sources and in different formats into standard numbers; S122, the standardized access employee number, the standardized access device number and the current page identifier are concatenated in a preset order, and the field type identifier and the desensitization rule version number are introduced to form the rotor input string; S123, input the rotor input string into the Hash function to obtain the original Hash digest, and extract or map the rotor seed, disturbance sequence number and fragment selection factor from the original Hash digest to form the device-side desensitized rotor.
[0008] Optionally, S13 includes: S131, based on the field adjacency set, obtain the coordinate position, arrangement number, field type and business attribute of the sensitive field to be displayed and each adjacent field in the page layout, calculate the display distance, arrangement order difference, field type similarity and business association strength between the sensitive field to be displayed and the adjacent fields respectively, and form a quantitative result of field relationship; S132, Based on the field relationship quantification results, calculate the structural exposure value of each adjacent field relative to the sensitive field to be displayed. When the structural exposure value reaches the exposure judgment threshold, determine the positional relationship between the corresponding fields as the structural exposure position. S133, take the sensitive fields to be displayed and the adjacent fields that meet the structural exposure judgment conditions as field nodes, and establish field relationship edges according to whether there are adjacency, same column, same group or business association relationships between fields. At the same time, use the structural exposure value as the weight of the field relationship edge to generate a field exposure position graph.
[0009] Optionally, S2 includes: S21. Based on the field exposure position sequence diagram, obtain the structural exposure value and field relationship edge weight corresponding to the sensitive field to be displayed, divide the sensitive field to be displayed according to character position, delimiter position or business code segment, generate multiple candidate field fragments, and calculate the fragment exposure degree of each candidate field fragment. S22. Based on the fragment exposure, field business readability requirements, and device-side desensitization rotor, each candidate field fragment is hierarchically judged. Fragments used to maintain business identification, page verification, or manual retrieval are identified as anchored fragments, and fragments that correspond to high-weight relationship edges in the field exposure sequence diagram and maintain the same form across different devices, leading to structural association, are identified as free fragments. S23, based on the disturbance sequence number and fragment selection factor in the desensitized rotor on the equipment side, assign at least one disturbance path among character replacement, position rotation, local rearrangement, virtual character insertion, or length preservation disturbance to each free fragment in the free fragment set, and combine the anchored fragment set, the free fragment set, and their corresponding disturbance paths to generate a fragment layering disturbance rule exclusive to the current employee under the current equipment.
[0010] Optionally, S22 includes: S221. Based on the readability requirements of the field business, the role of each candidate field fragment in business identification, page verification and manual retrieval is scored to obtain the fragment business readability score. S222, Based on the desensitized rotor on the equipment side, perform equipment-differentiated selection of candidate field segments, generate segment free selection values, determine the disturbed candidate field segments under the current employee and current equipment combination through the segment free selection values, and generate rotor selection results; S223, the fragment exposure, fragment service readability score and rotor selection result are jointly judged. When the service readability score of the candidate field fragment reaches the readability threshold and the fragment exposure is lower than the exposure threshold, it is determined as an anchored fragment. When the fragment exposure of the candidate field fragment reaches the exposure threshold, or is selected by the device-side desensitized rotor as a fragment to be detached, it is determined as a detached fragment, and anchored fragments and detached fragments are given priority to not overlap.
[0011] Optionally, S23 includes: S231, Construct a perturbation path library, encode character replacement, position rotation, local rearrangement, virtual character insertion and length preservation perturbation into different path numbers, and set corresponding applicable conditions for each type of perturbation path; S232, for each free fragment in the free fragment set, read the disturbance sequence number and fragment selection factor in the desensitized rotor on the equipment side, and calculate the corresponding disturbance path index by combining the free fragment sequence number, fragment length and fragment exposure. Select at least one disturbance path from the disturbance path library by the disturbance path index so that the free fragments corresponding to different employees or different equipment obtain different disturbance paths. S233 marks the anchored fragment set as reserved, binds each free fragment in the free fragment set to its corresponding perturbation path, perturbation intensity and output position, and generates fragment hierarchical perturbation rules according to the original field fragment order.
[0012] Optionally, S3 includes: S31, read the segment layer perturbation rules, identify the anchored segment set, and perform in-situ preservation, local truncation preservation, or fixed mask preservation processing on the anchored segments; S32, for the set of free fragments, according to the perturbation path, perturbation intensity and output position configured for each free fragment in the fragment layer perturbation rules, performs character replacement, position rotation, local rearrangement, virtual character insertion or length preservation perturbation processing respectively, so that the free fragments are driven by the desensitization rotor on different device sides to generate different desensitization forms on different access devices; S33. Based on the field relationship edge between the sensitive field to be displayed and adjacent fields in the field exposure position diagram, generate non-real association disturbance marks using the device-side desensitization rotor, and insert the non-real association disturbance marks into the display interval, field suffix or page hidden audit area between the sensitive field to be displayed and adjacent fields. Finally, synthesize the dynamic desensitization display result according to the output position in the fragment layering disturbance rule.
[0013] Optionally, S32 includes: S321, for each free segment in the free segment set, read the corresponding disturbance path, disturbance intensity and output position in the segment layer disturbance rule, and determine the disturbance control parameters of the free segment in combination with the desensitized rotor on the equipment side; S322, perform corresponding perturbation processing on the free fragment according to the perturbation path. When the perturbation path is character replacement, map the fragment characters bit by bit according to the desensitization rotor on the device side. When the perturbation path is position rotation, change the character order inside the fragment according to the perturbation sequence number. When the perturbation path is local rearrangement, rearrange the characters at the selected position in the fragment. When the perturbation path is virtual character insertion, insert a non-real character at the specified position in the fragment. When the perturbation path is length-preserving perturbation, replace the fragment with a perturbation character sequence with the same length as the original fragment. S323 rearranges the output of each free segment after perturbation according to the output position in the segment layer perturbation rule, and records the mapping relationship between the original segment number and the output position after perturbation, so that the same free segment forms different display forms and different display positions under the action of desensitizing rotors on different device sides.
[0014] Optionally, S33 includes: S331, extract the field relationship edges between the sensitive field to be displayed and the adjacent fields from the field exposure position graph, and read the edge weight, adjacent field identifier and relationship type corresponding to each field relationship edge to generate a relationship edge identifier sequence as the basis for generating non-real association disturbance markers; S332, based on the relation edge identifier sequence, combines the device-side desensitized rotor, the sensitive field identifier to be displayed, the adjacent field identifier, the edge weight and the relation type into a hash to generate a non-real association disturbance mark that is bound to the current employee, the current device and the current field relation edge, so that the same field relation forms different marks on different devices; S333: Determine the insertion position of the non-real association perturbation mark based on the edge weight and fragment layering perturbation rules, insert it into the display interval between the sensitive field to be displayed and the adjacent field, the field suffix, or the hidden audit area of the page, and synthesize the perturbation results of the anchored fragment and the free fragment according to the output position in the fragment layering perturbation rules to generate dynamic desensitization display results.
[0015] The beneficial effects of this invention are: This invention, by acquiring the accessing employee ID, accessing device ID, and page field layout relationships, constructs a device-side desensitization rotor, a field adjacency set, and a field exposure sequence diagram. It then identifies the structural exposure location by combining the field position relationships, field type association relationships, and business association relationships within the page. This enables differentiated and dynamic desensitization processing of sensitive fields across different employees, devices, and page scenarios. It effectively reduces the risk of sensitive information leakage caused by cross-device screenshot comparisons, log splicing, and field position deduction in traditional fixed desensitization methods, and improves the structural concealment capability during data display.
[0016] This invention divides sensitive fields to be displayed into anchored segments and free segments. Combining segment exposure, business readability scores, and the device-side desensitization rotor's disturbance control capabilities, the free segments are subjected to character replacement, positional rotation, partial rearrangement, virtual character insertion, and length-preserving disturbance processing. This allows the same sensitive field to form different desensitization forms on different devices, while ensuring that the anchored segments still have the ability to identify business data, verify pages, and perform manual retrieval. This balances data security and business availability, and improves the adaptability of dynamic desensitization results in complex business scenarios.
[0017] This invention generates non-real association perturbation markers based on field relationship edges in the field exposure sequence diagram, and inserts these non-real association perturbation markers into display intervals, field suffixes, or hidden audit areas on pages. Simultaneously, it dynamically synthesizes the desensitization results by combining fragment layering perturbation rules, so that the display results on different devices form differentiated interference at the structural association level. This blocks attackers from using field correspondences to perform association inference and reverse reconstruction, further enhancing the anti-association analysis capability and security audit capability in the data desensitization process. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1This is a schematic diagram of the desensitization method according to an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the construction of segment layer perturbation rules in an embodiment of the present invention. Detailed Implementation
[0020] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. Those skilled in the art may employ other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.
[0021] like Figures 1-2 As shown, a dynamic data anonymization method based on multidimensional perturbation mapping includes the following steps: S1. Obtain the access employee ID, access device ID, sensitive fields to be displayed, and the relationship between adjacent fields of the sensitive fields to be displayed in the current data access request. Perform Hash fusion on the access employee ID and access device ID to generate a device-side desensitization rotor bound to the current access employee and the current access device. Based on the relationship between adjacent fields, identify the structural exposure position of the sensitive fields to be displayed in the page that can be compared across devices and generate a field exposure position chart. S2, based on the device-side desensitization rotor, perform perturbation mapping on the field exposure sequence diagram, determine the anchor segments that need to retain business readability in the sensitive fields to be displayed and the free segments that need to block cross-device association, and assign perturbation paths to free segments that are different from other devices according to the device-side desensitization rotor, and generate segment layer perturbation rules exclusive to the current employee under the current device. S3 performs dynamic desensitization processing on sensitive fields to be displayed according to the segment layering perturbation rules, so that the anchored segment maintains the local consistency required for business identification, and the free segment presents different desensitization forms on different devices. It also inserts non-real association perturbation marks controlled by the device-side desensitization rotor between adjacent fields, and generates dynamic desensitization display results to block cross-device screenshot comparison, display log splicing, and field position inference.
[0022] S1 includes: S11: Receive the current data access request, extract the accessing employee ID from the login session, permission token, or access log, extract the accessing device ID from the terminal registration information, device fingerprint, or client certificate, and extract the sensitive fields to be displayed from the page rendering task. Simultaneously, read the page field layout configuration to determine the adjacent fields before and after the sensitive fields to be displayed, fields in the same column, fields in the same group, and related fields within the same business card, forming a field adjacency set, specifically including: (1) Receive the current data access request, parse the request header, login session, permission token and access log, extract the accessing employee number, parse the terminal registration information, device fingerprint and client certificate, extract the accessing device number, parse the page rendering task, extract the sensitive fields to be displayed and their field identifiers, and generate an access tuple, represented as: ; in, To access tuples, To access the employee ID, For accessing device number, For sensitive fields to be displayed, This serves as the identifier for the current page. (2) Read the page field layout configuration corresponding to the current page identifier, obtain the display coordinates, field width, field height, column number, field group number and business card number of each field in the page, and represent each field as a layout feature vector to characterize the positional relationship and structural belonging relationship of the field in the page, as follows: ; in, For the first The layout feature vector of each field. , The first The horizontal and vertical coordinates of each field on the page. , The first The display width and display height of each field. For the first The column number of each field. For the first The field group number to which each field belongs. For the first Each field corresponds to a business card number. For the first The field type of each field; (3) Taking the sensitive field to be displayed as the central field, compare the layout feature vectors of other fields on the page with the sensitive field to be displayed, and calculate the field distance, column consistency, field group consistency, business card consistency and field type correlation respectively. Based on the above results, filter out adjacent fields, fields in the same column, fields in the same group and related fields in the same business card to form a field adjacency set, which is represented as: ; in, For the field adjacency set, For the first page Candidate fields, For sensitive fields to be displayed, The page display distance between the candidate field and the sensitive field to be displayed. For the preset distance threshold, , These are the column numbers of the candidate field and the sensitive field to be displayed, respectively. , These are the field group numbers for the candidate fields and the sensitive fields to be displayed, respectively. , These are the business card numbers for the candidate fields and the sensitive fields to be displayed, respectively. To determine the correlation between candidate field types and the sensitive field types to be displayed. Associat thresholds with field types; ; ; in, , These represent the average and standard deviation of the display distances between all fields on the current page. This is the distance adjustment coefficient. , These represent the average and standard deviation of the type correlation of fields on the current page, respectively. This is the correlation adjustment coefficient; S12, perform format unification, invalid character removal and encoding standardization on the access employee number and access device number respectively, and concatenate the standardized access employee number and standardized access device number with the page identifier, input the Hash function to obtain the device-side desensitization rotor, so that the desensitization rotor is simultaneously constrained by employee identity, device identity and page scenario; S13, based on the field adjacency set, calculates the display distance, arrangement order, field type similarity and business association strength of the sensitive field to be displayed and its adjacent fields on the page, identifies the structural exposure location by comparison of screenshots from different devices, splicing of display logs or reverse inference of field position, and uses the sensitive field to be displayed and its adjacent fields as nodes, and uses the adjacency, same column, same group or business association relationship between fields as edges to generate a field exposure position graph.
[0023] S12 includes: S121, perform format standardization, invalid character removal, and encoding standardization on the accessing employee ID and accessing device ID respectively, converting employee IDs and device IDs from different sources and in different formats into standard IDs, avoiding the generation of different desensitized rotors for the same employee or the same device due to differences in capitalization, spaces, separators, or encoding, as shown below: ; ; in, This is the standardized visiting employee ID. This refers to the standardized access device number. Original visiting employee ID The original access device number, For the purpose of standardizing the format, Invalid character removal processing For encoding standardization processing; S122, the standardized access employee ID, the standardized access device ID, and the current page identifier are concatenated in a preset order, and field type identifiers and desensitization rule version numbers are introduced to form a rotor input string. This ensures that the generated device-side desensitization rotor is not only affected by employee and device identities, but also constrained by page scenarios, field types, and desensitization rule versions, as shown below: ; in, For rotor input string, This serves as the identifier for the current page. This is the version number of the desensitization rule; S123, input the rotor input string into a hash function to obtain the original hash digest, and extract or map the rotor seed, disturbance sequence number, and segment selection factor from the original hash digest. Combine these to form the device-side desensitized rotor, which is used to determine the disturbance direction, disturbance intensity, and display transformation path of sensitive fields, as shown below: ; ; ; ; ; in, This is the original hash digest. For Hash functions, For the desensitized rotor on the equipment side, For rotor seeds, The disturbance number is... Select factors for the fragment. To extract from the original hash digest The Ranked first Extracting a summary fragment. To convert the summary fragment into numerical values, For modulo operation, This represents the number of disturbance paths. The number of shards that can be used for sensitive fields.
[0024] S13 includes: S131, based on the field adjacency set, obtain the coordinate position, arrangement number, field type, and business attribute of the sensitive field to be displayed and each adjacent field in the page layout. Calculate the display distance, arrangement order difference, field type similarity, and business association strength between the sensitive field to be displayed and its adjacent fields to form a quantitative result of the field relationship, expressed as follows: ; ; ; ; ; ; ; ; ; ; ; in, For the first Quantification results of the field relationships between adjacent fields and the sensitive fields to be displayed To display distance, Due to differences in arrangement order, For field type similarity, For the strength of business association, , For the first Page coordinates of adjacent fields, , The page coordinates for the sensitive fields to be displayed. For the first The sequence number of adjacent fields on the page. This refers to the sequence number of the sensitive fields to be displayed on the page. For field type label similarity, Similarity is based on field length features. For field format feature similarity, , , These are the weight coefficients corresponding to the field type label, field length feature, and field format feature, respectively. For the first The field type of each adjacent field. For the field types of the sensitive fields to be displayed, The basic similarity coefficient for heterogeneous fields. For the first The field length of each adjacent field. The length of the sensitive field to be displayed. This is the difference in field length. This represents the number of identical positions of the two fields in a character composition pattern. This represents the total number of positions in the character composition pattern. , , These are the weighting coefficients for page co-occurrence, API calls, and business logic flow, respectively. For page co-occurrence relevance, For the relevance of API calls, For the relevance of business flow, For the first The number of times each adjacent field and the sensitive field to be displayed appear together on the same page. For the first The number of times each adjacent field appears on the page. The number of times the sensitive field to be displayed appears on the page. For the first The number of times that an adjacent field and a sensitive field to be displayed are simultaneously called by the same API. This represents the total number of API calls. For the first The number of times that an adjacent field and the sensitive field to be displayed participate in the same business process node. This represents the total number of business process iterations. S132, based on the field relationship quantification results, calculate the structural exposure value of each adjacent field relative to the sensitive field to be displayed. When the structural exposure value reaches the exposure judgment threshold, the positional relationship between the corresponding fields is determined as the structural exposure position, which is used to characterize the existence of a structural correspondence that can be compared by screenshots from different devices, spliced display logs, or deduced from the field position. It is expressed as: ; ; in, For the first The structure exposure value corresponding to each adjacent field , , , These are the weighting coefficients for display distance, sorting order, field type, and business association, respectively. The results of the structural exposure location determination. Indicates the first The structural exposure locations are formed between each adjacent field and the sensitive field to be displayed. This indicates that no structural exposure location has been formed. The exposure threshold; ; in, , These represent the mean and standard deviation of the relationship structure exposed values for all fields on the current page, respectively. Exposure adjustment coefficient; S133: The sensitive fields to be displayed and adjacent fields that meet the structural exposure criteria are taken as field nodes. Field relationship edges are established based on whether there are adjacencies, same columns, same groups, or business associations between fields. At the same time, the structural exposure value is used as the weight of the field relationship edge to generate a field exposure order graph, represented as follows: ; ; ; ; in, Expose bit order graphs for fields. For a collection of field nodes, For the set of field relationship edges, For the set of edge weights of field relationships, For sensitive fields to be displayed, For the first Adjacent fields, This indicates the sensitive fields to be displayed and the first... There are adjacency, same column, same group, or business relationship between adjacent fields. For sensitive fields to be displayed and the first Edge weights between adjacent fields.
[0025] S2 includes: S21, based on the field exposure positional graph, obtain the structural exposure value and field relationship edge weight corresponding to the sensitive field to be displayed. Segment the sensitive field to be displayed according to character position, delimiter position, or business code segment to generate multiple candidate field fragments. Calculate the fragment exposure degree of each candidate field fragment to determine whether the fragment has a structural association risk that could be compared across devices or inferred from the field position. Specifically, this includes: (1) Extract the nodes corresponding to the sensitive fields to be displayed from the field exposure order graph, as well as the field relationship edges between the nodes and adjacent fields. Obtain the structural exposure value and edge weight of each field relationship edge, and calculate the comprehensive exposure weight of the sensitive fields to be displayed, expressed as: ; in, The overall exposure weight of the sensitive fields to be displayed. Expose the weights of field relationship edges in the ordinal graph for each field. The number of adjacent fields that have a relationship edge with the sensitive field to be displayed; (2) Based on the character position, delimiter position, and business coding segment rules of the sensitive field to be displayed, the sensitive field to be displayed is segmented. The character position is used to divide the field into the first segment, the middle segment, and the last segment. The delimiter position is used to identify the coding units separated by horizontal lines, slashes, or spaces. The business coding segment rules are used to identify field segments with business meaning, such as department code, employee code, equipment code, and serial number. A set of candidate field segments is generated, represented as follows: ; ; in, For the first Candidate field fragments, The number of candidate field fragments, For the first The starting position of each candidate field fragment within the sensitive field to be displayed. For the first The end position of each candidate field fragment in the sensitive field to be displayed. To extract the first from the sensitive fields to be displayed Ranked first Bit-formed field fragments; (3) For each candidate field fragment, calculate the fragment exposure degree by combining its character length ratio, location sensitivity, business code segment sensitivity, and the comprehensive exposure weight of the sensitive field to be displayed. The higher the fragment exposure degree, the higher the degree to which the fragment forms a structural association in cross-device screenshot comparison, display log splicing, or field position reverse inference, as expressed as: ; in, For the first Fragment exposure of each candidate field fragment. For the first The character length of each candidate field fragment. The total character length of the sensitive field to be displayed. For the first Position sensitivity of candidate field fragments For the first Sensitivity of business code segments for each candidate field fragment , , These are the weighting coefficients for length percentage, position sensitivity, and business code segment sensitivity, respectively. ; ; ; ; ; ; ; ; in, Sensitivity to boundary location, Sensitivity to center position For order stability sensitivity, , , These are the weighting coefficients for the boundary position, the center position, and the order stability, respectively. For the first The starting position of each candidate field fragment. For the first The end position of each candidate field segment. The total length of the sensitive fields to be displayed. This represents the minimum distance between a candidate field fragment and the beginning or end of the field. For the first The center position of each candidate field fragment The center position of the sensitive field to be displayed. For the first The number of times each candidate field fragment maintains a fixed order in the historical page layout. To count the total number of times a page layout is historical, For encoding type sensitivity, For encoding frequency sensitivity, For the stability of the coding structure, , , These are the weighting coefficients for encoding type, encoding frequency, and encoding structure stability, respectively. For the first Sensitive weights for the encoding type of each candidate field fragment. To preset the maximum encoding sensitivity weight, For the first The number of unique values of each candidate field fragment in historical data. This represents the total number of historical data records. For the first The number of times a candidate field fragment maintains the same encoding structure across different pages, systems, or devices. This refers to the total number of pages or system versions counted. S22. Based on the fragment exposure, field business readability requirements, and device-side desensitization rotor, each candidate field fragment is hierarchically judged. Fragments used to maintain business identification, page verification, or manual retrieval are identified as anchored fragments, and fragments that correspond to high-weight relationship edges in the field exposure sequence diagram and maintain the same form across different devices, leading to structural association, are identified as free fragments. S23, based on the disturbance sequence number and fragment selection factor in the desensitized rotor on the equipment side, assign at least one disturbance path among character replacement, position rotation, local rearrangement, virtual character insertion, or length preservation disturbance to each free fragment in the free fragment set, and combine the anchored fragment set, the free fragment set, and their corresponding disturbance paths to generate a fragment layering disturbance rule exclusive to the current employee under the current equipment.
[0026] S22 includes: S221, based on the field business readability requirements, the role of each candidate field fragment in business identification, page verification, and manual retrieval is scored to obtain the fragment business readability score, expressed as: ; in, For the first Business readability score for each candidate field fragment Contribution to business identification To verify the contribution of the page, Contribution to manual retrieval , , These are the weighting coefficients for business identification, page verification, and manual retrieval, respectively. ; ; ; ; in, As a unique contribution value, Differentiate contribution values by business category. Contribute value to business continuity , , These are the weighting coefficients for uniqueness, business category differentiation, and business continuity, respectively. For the first The number of unique values corresponding to each candidate field fragment This represents the total number of historical samples. For including the first The number of records where each candidate field fragment uniquely corresponds to a business category. This represents the total number of business categories. For the first The number of times each candidate field fragment remains consistent across consecutive nodes in the business process. This represents the total number of nodes in the business process. ; ; ; ; in, For location visibility, For visual stability, Page reference frequency, , , These are the weighting coefficients for uniqueness, business category differentiation, and business continuity, respectively. For the first Center position of each candidate field fragment The total length of the sensitive fields to be displayed. For the first The horizontal offset of each candidate field fragment across different pages. For the first The vertical offset of each candidate field fragment across different pages. Page width For page height, For including the first The number of pages for each candidate field fragment. This represents the total number of pages in the system. ; ; ; ; in, To retrieve the hit frequency, To improve retrieval differentiation, For manual input of reuse rate, , , These are the weighting coefficients for retrieval hit frequency, retrieval discrimination, and manual input reuse rate, respectively. To use the first The number of times each candidate field fragment is retrieved and successfully matched. This represents the total number of searches. To use the first The number of results returned when candidate field fragments are used as search criteria. This represents the maximum number of results the system is allowed to return. For the first The number of times each candidate field fragment was entered repeatedly. Enter the total number of times manually; S222, Based on the desensitized rotor on the equipment side, perform equipment-differentiated selection on candidate field segments, generate segment free selection values, determine the disturbed candidate field segments under the current employee and current equipment combination using the segment free selection values, and generate rotor selection results, represented as: ; ; in, For the first detached selection values for each candidate field fragment For the rotor seed in the desensitized rotor on the equipment side, For the segment selection factor in the desensitized rotor on the equipment side, This refers to the disturbance sequence number in the desensitized rotor on the equipment side. For numerical conversion functions, For modulo operation, Choose the modulus radix for the segment. To select a threshold for free, For rotor selection results, Indicates the first The candidate word fragments were selected by the device-side desensitization rotor as fragments to be detached; ; in, , These represent the mean and standard deviation of the free selection values for all candidate field segments, respectively. This is the free regulation coefficient; S223, the fragment exposure degree, fragment service readability score, and rotor selection result are jointly determined. When the service readability score of a candidate field fragment reaches the readability threshold and the fragment exposure degree is lower than the exposure threshold, it is identified as an anchored fragment. When the fragment exposure degree of a candidate field fragment reaches the exposure threshold, or is selected as a fragment to be detached by the device-side desensitized rotor, it is identified as a detached fragment. Priority is given to ensuring that anchored fragments and detached fragments do not overlap, as shown below: ; ; ; in, For anchoring fragment set, For a collection of free fragments, For the first Candidate field fragments, Readability threshold For the exposure threshold, This indicates that the set of anchored segments and the set of free segments do not overlap; ; ; in, , These represent the mean and standard deviation of the business readability scores for all candidate field fragments, respectively. This is a readability adjustment factor. , These represent the mean and standard deviation of the exposure of all candidate field fragments, respectively. This is the exposure adjustment coefficient.
[0027] S23 includes: S231, Construct a perturbation path library, encoding character replacement, positional rotation, local rearrangement, virtual character insertion, and length preservation perturbations as different path numbers, and setting corresponding applicable conditions for each type of perturbation path, represented as: ; in, For the perturbation path library, Replace the path with characters. For position rotation load path, For local path rearrangement, Insert a path for virtual characters. Preserve the perturbation path for the length; S232, for each free fragment in the free fragment set, read the disturbance sequence number and fragment selection factor in the desensitized rotor on the equipment side, and calculate the corresponding disturbance path index by combining the free fragment sequence number, fragment length, and fragment exposure. Then, select at least one disturbance path from the disturbance path library based on the disturbance path index, so that free fragments corresponding to different employees or different equipment obtain different disturbance paths, as shown below: ; ; in, For the first The perturbation path index corresponding to each free segment For modulo operation, This refers to the disturbance sequence number in the desensitized rotor on the equipment side. For the segment selection factor in the desensitized rotor on the equipment side, The sequence number of the free fragment. The number of paths in the perturbation path library. For the first The perturbation path corresponding to each free segment The disturbance path library is numbered as follows The disturbance path; S233 marks the anchored fragment set as reserved, binds each free fragment in the free fragment set to its corresponding perturbation path, perturbation strength, and output position, and generates fragment hierarchical perturbation rules according to the original field fragment order. This ensures that the desensitization process for the current employee on the current device retains necessary business readability while performing differentiated perturbations on highly exposed fragments, as shown below: ; ; ; in, For segment layer perturbation rules, For the first Candidate field fragments, For anchoring fragment set, For a collection of free fragments, For preservation purposes, For the first The perturbation path corresponding to each free segment For the first The intensity of the disturbance of a free fragment, For the first The perturbation output position of a free segment To normalize the exposure of fragments, This represents the number of candidate field fragments.
[0028] S3 includes: S31 reads the segment layer perturbation rules, identifies the anchored segment set, and performs in-situ preservation, partial truncation preservation, or fixed mask preservation processing on the anchored segments to ensure that they can still support business identification, manual verification, and retrieval location on the current page, while avoiding the complete exposure of the original sensitive fields, as shown below: ; in, For the first The desensitization result after the anchoring fragments are preserved. For the first The original character content of each anchored segment, For the first The retention processing type corresponding to each anchored segment For from the first The anchored segment retains the previous one The reserved substring formed by the bit characters For the first In each anchor segment, the remaining characters, excluding the reserved characters, undergo fixed mask replacement processing. For a length of All characters are replaced using a fixed mask. S32, for the set of free fragments, according to the perturbation path, perturbation intensity and output position configured for each free fragment in the fragment layer perturbation rules, performs character replacement, position rotation, local rearrangement, virtual character insertion or length preservation perturbation processing respectively, so that the free fragments are driven by the desensitization rotor on different device sides to generate different desensitization forms on different access devices; S33. Based on the field relationship edges between the sensitive field to be displayed and adjacent fields in the field exposure sequence diagram, generate non-real association disturbance marks using the device-side desensitization rotor, and insert the non-real association disturbance marks into the display interval, field suffix, or page hidden audit area between the sensitive field to be displayed and adjacent fields. Finally, synthesize the dynamic desensitization display result according to the output position in the fragment layering disturbance rule.
[0029] S32 includes: S321, for each free segment in the free segment set, read the corresponding disturbance path, disturbance intensity, and output position from the segment layer disturbance rule, and determine the disturbance control parameters for that free segment in conjunction with the desensitized rotor on the equipment side, expressed as: ; in, For the first Disturbance control parameters for each free segment, For the first A free fragment For the first The perturbation path corresponding to each free segment For the first The perturbation intensity corresponding to each free fragment For the first The output position of each free fragment For the desensitized rotor on the equipment side; S322, perform corresponding perturbation processing on the free fragment according to the perturbation path. When the perturbation path is character replacement, map the fragment characters bit by bit according to the device-side desensitization rotor. When the perturbation path is position rotation, change the character order within the fragment according to the perturbation sequence number. When the perturbation path is local rearrangement, rearrange the characters at the selected position within the fragment. When the perturbation path is virtual character insertion, insert a non-real character at the specified position of the fragment. When the perturbation path is length-preserving perturbation, replace the fragment with a perturbation character sequence with the same length as the original fragment, as shown below: ; in, For the first The result after perturbing a free fragment For character replacement functions, For order rotation functions, It is a local rearrangement function. Functions for inserting virtual characters The length-preserving perturbation function, This refers to the disturbance sequence number in the desensitized rotor on the equipment side. For the first The length of each free segment; S323, the perturbation results of each free segment are rearranged and output according to the output position in the segment layer perturbation rule, and the mapping relationship between the original segment number and the perturbation output position is recorded, so that the same free segment forms different display forms and different display positions under the action of the desensitization rotor on different device sides, as shown in: ; ; in, For the set of outputs perturbed by free fragments, For a collection of free fragments, This represents the location mapping relationship of the free fragments.
[0030] S33 includes: S331, extract the field relationship edges between the sensitive field to be displayed and its adjacent fields from the field exposure position graph, and read the edge weight, adjacent field identifier, and relationship type corresponding to each field relationship edge to generate a relationship edge identifier sequence, which serves as the basis for generating non-real association perturbation markers, represented as follows: ; in, Identify the sequence of relation edges. For sensitive fields to be displayed, For the first Adjacent fields, For sensitive fields to be displayed and the first Edge weights between adjacent fields For field relationship types, Expose the set of field relationship edges in the bit order graph for each field; S332, based on the relation edge identifier sequence, combines the device-side desensitized rotor, the identifier of the sensitive field to be displayed, the adjacent field identifier, the edge weight, and the relation type into a hash to generate a non-real association disturbance mark bound to the current employee, the current device, and the current field relation edge. This causes the same field relation to form different marks on different devices, as shown below: ; in, For the first Non-real association perturbation markers corresponding to the field relationship edges. For Hash functions, For the desensitized rotor on the equipment side, The length of the perturbation mark. To extract the first part of the hash result Bit; S333, based on edge weights and fragment layering perturbation rules, determines the insertion position of the non-real association perturbation marker, inserts it into the display interval between the sensitive field to be displayed and adjacent fields, field suffixes, or the hidden audit area of the page, and synthesizes the perturbation results of anchored fragments and free fragments according to the output positions in the fragment layering perturbation rules to generate a dynamic desensitization display result, represented as: ; ; in, For the first The insertion position of a non-realistic correlation perturbation marker. To achieve a high exposure edge weight threshold, To set a low exposure edge weight threshold, To display the results of dynamic desensitization, For the result composition function, To preserve the processed result set for the anchored fragment, This is the set of results after perturbation processing of free fragments. This is a set of non-realistic correlation perturbation markers and their insertion positions. For segment layer perturbation rules; ; ; in, , These represent the average and standard deviation of the weights of all field relationship edges in the field exposure ordinal graph. For high exposure adjustment coefficient, The low exposure adjustment coefficient.
[0031] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0032] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A dynamic data desensitization method based on multidimensional perturbation mapping, characterized in that, Includes the following steps: S1, obtain the access employee ID, access device ID, sensitive field to be displayed, and the relationship between adjacent fields of the sensitive field to be displayed in the current data access request; perform Hash fusion on the access employee ID and access device ID to generate a device-side desensitization rotor bound to the current access employee and the current access device; and identify the structural exposure position of the sensitive field to be displayed in the page that can be compared across devices based on the relationship between adjacent fields, and generate a field exposure position chart. S2, based on the device-side desensitization rotor, perform perturbation mapping on the field exposure sequence diagram, determine the anchored segments that need to retain business readability and the free segments that need to block cross-device association in the sensitive fields to be displayed, and assign perturbation paths different from other devices to the free segments according to the device-side desensitization rotor, and generate segment layer perturbation rules exclusive to the current employee under the current device; S3, perform dynamic desensitization processing on the sensitive fields to be displayed according to the segment layering perturbation rules, so that the anchored segment maintains the local consistency required for business identification, so that the free segment presents different desensitization forms on different devices, and insert non-real association perturbation marks controlled by the device-side desensitization rotor between adjacent fields, generating dynamic desensitization display results for blocking cross-device screenshot comparison, display log splicing, and field position reverse inference.
2. The dynamic data desensitization method based on multidimensional perturbation mapping according to claim 1, characterized in that, S1 includes: S11, receive the current data access request, extract the accessing employee number from the login session, permission token or access log, extract the accessing device number from the terminal registration information, device fingerprint or client certificate, and extract the sensitive fields to be displayed from the page rendering task. At the same time, read the page field layout configuration, determine the adjacent fields before and after the sensitive fields to be displayed, the fields in the same column, the fields in the same group, and the related fields in the same business card, and form a field adjacency set. S12, perform format unification, invalid character removal and encoding standardization on the access employee number and access device number respectively, and concatenate the standardized access employee number and standardized access device number with the page identifier, input the Hash function to obtain the device-side desensitization rotor, so that the desensitization rotor is simultaneously constrained by employee identity, device identity and page scenario; S13, based on the field adjacency set, calculates the display distance, arrangement order, field type similarity and business association strength of the sensitive field to be displayed and its adjacent fields on the page, identifies the structural exposure location by comparison of screenshots from different devices, splicing of display logs or reverse inference of field position, and uses the sensitive field to be displayed and its adjacent fields as nodes, and uses the adjacency, same column, same group or business association relationship between fields as edges to generate a field exposure position graph.
3. The dynamic data desensitization method based on multidimensional perturbation mapping according to claim 2, characterized in that, S12 includes: S121, perform format unification, invalid character removal and encoding standardization processing on the access employee number and access device number respectively, and convert employee numbers and device numbers from different sources and in different formats into standard numbers; S122, the standardized access employee number, the standardized access device number and the current page identifier are concatenated in a preset order, and the field type identifier and the desensitization rule version number are introduced to form the rotor input string; S123, input the rotor input string into the Hash function to obtain the original Hash digest, and extract or map the rotor seed, disturbance sequence number and fragment selection factor from the original Hash digest to form the device-side desensitized rotor.
4. The dynamic data desensitization method based on multidimensional perturbation mapping according to claim 3, characterized in that, S13 includes: S131, based on the field adjacency set, obtain the coordinate position, arrangement number, field type and business attribute of the sensitive field to be displayed and each adjacent field in the page layout, calculate the display distance, arrangement order difference, field type similarity and business association strength between the sensitive field to be displayed and the adjacent fields respectively, and form a quantitative result of field relationship; S132, Based on the field relationship quantification results, calculate the structural exposure value of each adjacent field relative to the sensitive field to be displayed. When the structural exposure value reaches the exposure judgment threshold, determine the positional relationship between the corresponding fields as the structural exposure position. S133, take the sensitive fields to be displayed and the adjacent fields that meet the structural exposure judgment conditions as field nodes, and establish field relationship edges according to whether there are adjacency, same column, same group or business association relationships between fields. At the same time, use the structural exposure value as the weight of the field relationship edge to generate a field exposure position graph.
5. The dynamic data desensitization method based on multidimensional perturbation mapping according to claim 1, characterized in that, S2 includes: S21. Based on the field exposure position sequence diagram, obtain the structural exposure value and field relationship edge weight corresponding to the sensitive field to be displayed, divide the sensitive field to be displayed according to character position, delimiter position or business code segment, generate multiple candidate field fragments, and calculate the fragment exposure degree of each candidate field fragment. S22. Based on the fragment exposure, field business readability requirements, and device-side desensitization rotor, each candidate field fragment is hierarchically judged. Fragments used to maintain business identification, page verification, or manual retrieval are identified as anchored fragments, and fragments that correspond to high-weight relationship edges in the field exposure sequence diagram and maintain the same form across different devices, leading to structural association, are identified as free fragments. S23. Based on the disturbance sequence number and fragment selection factor in the desensitized rotor on the equipment side, assign at least one disturbance path among character replacement, position rotation, local rearrangement, virtual character insertion, or length preservation disturbance to each free fragment in the free fragment set. Combine the anchored fragment set, the free fragment set, and their corresponding disturbance paths to generate a fragment layering disturbance rule exclusive to the current employee under the current equipment.
6. The dynamic data desensitization method based on multidimensional perturbation mapping according to claim 5, characterized in that, S22 includes: S221. Based on the readability requirements of the field business, the role of each candidate field fragment in business identification, page verification and manual retrieval is scored to obtain the fragment business readability score. S222, Based on the desensitized rotor on the equipment side, perform equipment-differentiated selection of candidate field segments, generate segment free selection values, determine the disturbed candidate field segments under the current employee and current equipment combination through the segment free selection values, and generate rotor selection results; S223, the fragment exposure, fragment service readability score and rotor selection result are jointly judged. When the service readability score of the candidate field fragment reaches the readability threshold and the fragment exposure is lower than the exposure threshold, it is determined as an anchored fragment. When the fragment exposure of the candidate field fragment reaches the exposure threshold, or is selected by the device-side desensitized rotor as a fragment to be detached, it is determined as a detached fragment, and anchored fragments and detached fragments are given priority to not overlap.
7. The dynamic data desensitization method based on multidimensional perturbation mapping according to claim 6, characterized in that, S23 includes: S231, Construct a perturbation path library, encode character replacement, position rotation, local rearrangement, virtual character insertion and length preservation perturbation into different path numbers, and set corresponding applicable conditions for each type of perturbation path; S232, for each free fragment in the free fragment set, read the disturbance sequence number and fragment selection factor in the desensitized rotor on the equipment side, and calculate the corresponding disturbance path index by combining the free fragment sequence number, fragment length and fragment exposure. Select at least one disturbance path from the disturbance path library by the disturbance path index so that the free fragments corresponding to different employees or different equipment obtain different disturbance paths. S233 marks the anchored fragment set as reserved, binds each free fragment in the free fragment set to its corresponding perturbation path, perturbation intensity and output position, and generates fragment hierarchical perturbation rules according to the original field fragment order.
8. The dynamic data desensitization method based on multidimensional perturbation mapping according to claim 1, characterized in that, S3 includes: S31, read the segment layer perturbation rules, identify the anchored segment set, and perform in-situ preservation, local truncation preservation, or fixed mask preservation processing on the anchored segments; S32, for the set of free fragments, according to the perturbation path, perturbation intensity and output position configured for each free fragment in the fragment layer perturbation rules, performs character replacement, position rotation, local rearrangement, virtual character insertion or length preservation perturbation processing respectively, so that the free fragments are driven by the desensitization rotor on different device sides to generate different desensitization forms on different access devices; S33. Based on the field relationship edge between the sensitive field to be displayed and adjacent fields in the field exposure position diagram, generate non-real association disturbance marks using the device-side desensitization rotor, and insert the non-real association disturbance marks into the display interval, field suffix or page hidden audit area between the sensitive field to be displayed and adjacent fields. Finally, synthesize the dynamic desensitization display result according to the output position in the fragment layering disturbance rule.
9. A dynamic data desensitization method based on multidimensional perturbation mapping according to claim 8, characterized in that, S32 includes: S321, for each free segment in the free segment set, read the corresponding disturbance path, disturbance intensity and output position in the segment layer disturbance rule, and determine the disturbance control parameters of the free segment in combination with the desensitized rotor on the equipment side; S322, perform corresponding perturbation processing on the free fragment according to the perturbation path. When the perturbation path is character replacement, map the fragment characters bit by bit according to the desensitization rotor on the device side. When the perturbation path is position rotation, change the character order inside the fragment according to the perturbation sequence number. When the perturbation path is local rearrangement, rearrange the characters at the selected position in the fragment. When the perturbation path is virtual character insertion, insert a non-real character at the specified position in the fragment. When the perturbation path is length-preserving perturbation, replace the fragment with a perturbation character sequence with the same length as the original fragment. S323 rearranges the output of each free segment after perturbation according to the output position in the segment layer perturbation rule, and records the mapping relationship between the original segment number and the output position after perturbation, so that the same free segment forms different display forms and different display positions under the action of desensitizing rotors on different device sides.
10. A dynamic data desensitization method based on multidimensional perturbation mapping according to claim 9, characterized in that, S33 includes: S331, extract the field relationship edges between the sensitive field to be displayed and the adjacent fields from the field exposure position graph, and read the edge weight, adjacent field identifier and relationship type corresponding to each field relationship edge to generate a relationship edge identifier sequence as the basis for generating non-real association disturbance markers; S332, based on the relation edge identifier sequence, combines the device-side desensitized rotor, the sensitive field identifier to be displayed, the adjacent field identifier, the edge weight and the relation type into a hash to generate a non-real association disturbance mark that is bound to the current employee, the current device and the current field relation edge, so that the same field relation forms different marks on different devices; S333: Determine the insertion position of the non-real association perturbation mark based on the edge weight and fragment layering perturbation rules, insert it into the display interval between the sensitive field to be displayed and the adjacent field, the field suffix, or the hidden audit area of the page, and synthesize the perturbation results of the anchored fragment and the free fragment according to the output position in the fragment layering perturbation rules to generate dynamic desensitization display results.