Data anonymization adjustment method and system based on dynamic association risk analysis

By obtaining the temporal, spatial and semantic correlation characteristics in cross-domain scenarios and dynamically calling anonymization strategy, the sensitive information leakage problem caused by cross-domain data association risks is solved, and the balance between data privacy and availability is achieved, and different business scenarios and security needs are adapted to different business scenarios and security needs.

CN120449213AActive Publication Date: 2025-08-08BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510947293.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-08-08
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

Existing data anonymization technology cannot effectively identify the associated risks between cross-domain data, resulting in the leakage of sensitive information and lack of dynamic adaptability, unable to cope with the expansion of new data sets and external knowledge bases, resulting in an imbalance in privacy and data availability.

Method used

By obtaining domain features with temporal correlation, spatial correlation and semantic correlation in cross-domain scenarios, using time window overlap frequency, external map knowledge base and BERT model, weighted fusion is performed, and anonymization strategy is called dynamically to protect data privacy.

Benefits of technology

It realizes dynamic perception and adaptive balance of cross-domain data association risks and anonymization levels, effectively protects data privacy, reduces leakage risks, and maintains the analysis value of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449213A_ABST
    Figure CN120449213A_ABST
Patent Text Reader

Abstract

The invention discloses a data anonymization adjustment method and system based on dynamic association risk analysis. The method comprises the following steps: acquiring all original data associated with two types of systems forming a cross-domain scene in a preset period; extracting a first domain feature and a second domain feature with time correlation and space correlation from all the original data; obtaining the time matching rate of the two types of features based on the time window overlapping frequency; acquiring spatial relevance of the two types of features based on an external map knowledge base; obtaining semantic relevance of the two types of features; carrying out weighted fusion on the three types of relevance to obtain a relevance edge weight, and determining a path risk level formed by the two types of features according to the weight; and dynamically calling a corresponding anonymization strategy according to the path risk level, and executing the strategy on the corresponding data. According to the method, the problems of leakage risk and the like caused by cross-domain association of the weakly anonymized data set can be effectively solved, and the balance of dynamic risk perception and anonymization level self-adaption is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data privacy protection, and in particular to a data anonymization adjustment method and system based on dynamic correlation risk analysis. Background Art

[0002] In today's data processing and privacy protection fields, traditional data anonymization technologies have significant limitations. Traditional methods, such as k-anonymity and differential privacy, are designed to focus solely on single datasets and fail to account for the risks of cross-domain data correlation. For example, transaction times and medical records may have temporal and spatial matching, but these traditional methods are unable to identify such potential correlation risks, making it difficult to effectively prevent the leakage of sensitive information through cross-domain data correlation.

[0003] At the same time, existing solutions lack dynamic adaptability. Faced with new public datasets and the expansion of external knowledge bases, these solutions lack corresponding response mechanisms and are unable to promptly address dynamically changing risks. For example, the issue of geolocation accuracy in a specific case study clearly demonstrates the lag of current data privacy protection solutions in the face of dynamic risks, and their inability to quickly adjust to ensure data security when risks change.

[0004] In addition, existing technologies also face the problem of imbalance between utility and privacy; although excessive anonymization processing protects privacy to a certain extent, it undermines the availability of data, for example, it may cause distortion in medical time series analysis; and if weak anonymization processing is used, it cannot effectively resist association reasoning attacks, making it difficult to ensure the privacy of data. Summary of the Invention

[0005] In view of this, the embodiments of the present disclosure provide a data anonymization adjustment method and system based on dynamic association risk analysis, which can solve the problems existing in the prior art such as the risk of leakage of weakly anonymized data sets caused by cross-domain association.

[0006] In a first aspect, an embodiment of the present disclosure provides a data anonymization adjustment method based on dynamic correlation risk analysis, comprising: Obtain all raw data associated with the first type system and the second type system within a preset period; wherein the first type system and the second type system are two types of systems constituting a cross-domain scenario; Extracting two types of cross-domain scene information from all the original data, the two types of cross-domain scene information including first domain features and second domain features having temporal correlation and spatial correlation; Obtaining a time matching rate between the first domain feature and the second domain feature based on a time window overlapping frequency; Acquire the spatial correlation between the first domain feature and the second domain feature based on an external map knowledge base; Obtaining semantic relevance between texts corresponding to the first domain feature and the second domain feature; Performing weighted fusion on the time matching rate, the spatial correlation, and the semantic correlation to obtain an association edge weight between the first domain feature and the second domain feature; Determining a path risk level formed by the first domain feature and the second domain feature according to the associated edge weight; The corresponding anonymization strategy is dynamically called according to the path risk level, and the anonymization strategy is executed on the data corresponding to the first domain feature and the second domain feature.

[0007] In a second aspect, the embodiments of the present disclosure further provide a data anonymization adjustment system based on dynamic correlation risk analysis, including: A raw data acquisition module, configured to acquire all raw data associated with the first and second types of systems within a preset period; wherein the first and second types of systems are two types of systems constituting a cross-domain scenario; A domain feature acquisition module, configured to extract two types of cross-domain scene information from all the raw data, wherein the two types of cross-domain scene information include a first domain feature and a second domain feature having temporal correlation and spatial correlation; a time matching rate acquisition module, configured to obtain a time matching rate between the first domain feature and the second domain feature based on a time window overlapping frequency; A spatial correlation acquisition module, configured to acquire the spatial correlation between the first domain feature and the second domain feature based on an external map knowledge base; A semantic relevance acquisition module, configured to acquire the semantic relevance of texts corresponding to the first domain feature and the second domain feature; a fusion module, configured to perform weighted fusion on the time matching rate, the spatial correlation, and the semantic correlation to obtain an association edge weight between the first domain feature and the second domain feature; a path risk level acquisition module, configured to determine a path risk level formed by the first domain feature and the second domain feature according to the associated edge weight; A dynamic anonymization module is used to dynamically call a corresponding anonymization strategy according to the path risk level, and execute the anonymization strategy on the data corresponding to the first domain feature and the second domain feature.

[0008] In a third aspect, the embodiments of the present disclosure further provide a computer device that adopts the following technical solution: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any of the above-mentioned data anonymization adjustment methods based on dynamic correlation risk analysis.

[0009] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium storing computer instructions for enabling a computer to execute any of the above-mentioned data anonymization adjustment methods based on dynamic correlation risk analysis.

[0010] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, comprising a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.

[0011] The data anonymization adjustment method based on dynamic correlation risk analysis disclosed in the present application obtains all raw data of the first type of system and the second type of system that are associated within a preset period constituting a cross-domain scenario; extracts two types of cross-domain scenario information from all the raw data, and the two types of cross-domain scenario information include first domain features and second domain features with time correlation and spatial correlation; obtains the time matching rate of the first domain features and the second domain features based on the overlapping frequency of the time window; obtains the spatial correlation of the first domain features and the second domain features based on an external map knowledge base; obtains the semantic correlation of the text corresponding to the first domain features and the second domain features, and comprehensively considers the correlations in multiple aspects such as time, space and semantics, so as to more accurately discover the data correlation patterns and potential risks between cross-domain systems, and provide data analysis and decision-making. Provide stronger support; perform weighted fusion of time matching rate, spatial correlation, and semantic correlation to obtain the association edge weight of the first domain feature and the second domain feature; determine the path risk level formed by the first domain feature and the second domain feature based on the association edge weight; dynamically call the corresponding anonymization strategy according to the path risk level, and execute the anonymization strategy on the data corresponding to the first domain feature and the second domain feature. It can adopt corresponding anonymization strategies for data of different risk levels, effectively protect the privacy and sensitive information of related data, and reduce the risk of data leakage. The use of dynamic anonymization strategies can effectively avoid the impact of excessive anonymization on data availability, so that the data can still be used for valuable analysis and research. This method has strong adaptability and flexibility and can cope with different business scenarios and security requirements.

[0012] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the following specifically cites preferred embodiments and describes them in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0014] Figure 1 A flowchart of a data anonymization adjustment method based on dynamic correlation risk analysis provided in an embodiment of the present disclosure.

[0015] Figure 2 A flowchart of a method for obtaining a time matching rate between a first domain feature and a second domain feature provided in an embodiment of the present disclosure.

[0016] Figure 3 A flowchart of a method for acquiring spatial correlation between first domain features and second domain features provided in an embodiment of the present disclosure.

[0017] Figure 4 A flowchart of a method for obtaining semantic relevance between texts corresponding to a first domain feature and a second domain feature provided in an embodiment of the present disclosure.

[0018] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0020] Reference Figure 1 In a first aspect, the present application discloses a data anonymization adjustment method based on dynamic correlation risk analysis, which is used for anonymization adjustment of cross-domain risk data. The method specifically includes: S100, obtaining all raw data associated between the first type system and the second type system within a preset period; wherein the first type system and the second type system are two types of systems constituting a cross-domain scenario.

[0021] Wherein, all original data include structured data and / or unstructured data.

[0022] Obtaining comprehensive raw data is the foundation for subsequent analysis, ensuring its completeness and accuracy. Only with sufficient relevant data can we more accurately identify data correlation patterns and potential risks between two systems.

[0023] In this example, the first type of system is a medical system, and the second type of system is a financial system. Assuming a preset period of one month (from June 1, 2025, to June 30, 2025), the system interface and data extraction tools can be used to extract the raw data associated with each other from the target medical system and the target financial system during this period. For example, when a patient pays for medical treatment at a medical institution using a financial payment method, a related record is left in both systems. These records can be fully extracted, including the patient's basic information, visit time, department, payment amount, payment time, and so on.

[0024] S200 , extracting two types of cross-domain scene information from all original data, the two types of cross-domain scene information including first domain features and second domain features having temporal correlation and spatial correlation.

[0025] Extracting domain features with temporal and spatial correlation can reduce data redundancy, focus on information relevant to the analysis, and provide more targeted data for subsequent correlation analysis.

[0026] The first domain feature is preferably medical information, and the second domain feature is preferably financial transaction information. Furthermore, the extracted raw data can be filtered to select medical information related to the first domain feature, such as the date of the visit, the location of the visit, and the disease diagnosis, and to select financial transaction information related to the second domain feature, such as the transaction date, the location of the transaction, and the transaction amount. For example, if a patient visited Hospital A on June 10 and paid at the same time at the hospital's cashier using a bank card, these medical information and financial transaction information can be extracted separately.

[0027] S300 , obtaining a time matching rate between the first domain feature and the second domain feature based on a time window overlapping frequency.

[0028] The temporal matching rate quantifies the temporal correlation between features in two domains. A high temporal matching rate may indicate a strong correlation between the two domains, helping to identify potential abnormal correlations or patterns, providing important evidence for risk assessment.

[0029] S400: Acquire spatial correlation between first domain features and second domain features based on an external map knowledge base.

[0030] Spatial correlation can further verify the relationship between the characteristics of two fields. In actual situations, medical treatment and financial transactions usually occur in the same place or a nearby place. If the spatial correlation does not conform to common sense, there may be data anomalies or potential risks.

[0031] S500: Obtaining semantic relevance between texts corresponding to the first domain feature and the second domain feature.

[0032] Semantic relevance can reveal the intrinsic connection between the features of two domains at the text level. By analyzing semantics, some hidden related information can be discovered, thereby improving the accuracy of risk assessment.

[0033] S600: Perform weighted fusion on the time matching rate, spatial correlation, and semantic correlation to obtain the correlation edge weight between the first domain feature and the second domain feature.

[0034] Weighted fusion can comprehensively consider the impact of multiple factors on the degree of correlation between the features of two fields, and obtain a more comprehensive and accurate correlation edge weight; by adjusting the weight, the importance of different factors can be highlighted according to actual needs.

[0035] S700: Determine the path risk level formed by the first domain feature and the second domain feature according to the associated edge weight.

[0036] Determining the path risk level can classify and manage different associated paths, making it easier to take targeted measures.

[0037] S800: Dynamically call the corresponding anonymization strategy according to the path risk level, and execute the anonymization strategy on the data corresponding to the first domain feature and the second domain feature.

[0038] Dynamically calling anonymization policies can take appropriate protective measures based on different risk levels, ensuring data security while preserving data availability to a certain extent. This avoids the risk of data losing its analytical value due to over-anonymization, or data leakage due to insufficient anonymization.

[0039] The data anonymization adjustment method based on dynamic correlation risk analysis disclosed in the present application obtains all raw data of the first type of system and the second type of system that are associated within a preset period constituting a cross-domain scenario; extracts two types of cross-domain scenario information from all the raw data, and the two types of cross-domain scenario information include first domain features and second domain features with time correlation and spatial correlation; obtains the time matching rate of the first domain features and the second domain features based on the overlapping frequency of the time window; obtains the spatial correlation of the first domain features and the second domain features based on an external map knowledge base; obtains the semantic correlation of the text corresponding to the first domain features and the second domain features, and comprehensively considers the correlations in multiple aspects such as time, space and semantics, so as to more accurately discover the data correlation patterns and potential risks between cross-domain systems, and provide data analysis and decision-making. Provide stronger support; perform weighted fusion of time matching rate, spatial correlation, and semantic correlation to obtain the association edge weight of the first domain feature and the second domain feature; determine the path risk level formed by the first domain feature and the second domain feature based on the association edge weight; dynamically call the corresponding anonymization strategy according to the path risk level, and execute the anonymization strategy on the data corresponding to the first domain feature and the second domain feature. It can adopt corresponding anonymization strategies for data of different risk levels, effectively protect the privacy and sensitive information of related data, and reduce the risk of data leakage. The use of dynamic anonymization strategies can effectively avoid the impact of excessive anonymization on data availability, so that the data can still be used for valuable analysis and research. This method has strong adaptability and flexibility and can cope with different business scenarios and security requirements.

[0040] Reference Figure 2 The method of S300, “obtaining the time matching ratio between the first domain feature and the second domain feature based on the time window overlapping frequency”, i.e., the method of obtaining the time matching ratio between the first domain feature and the second domain feature, specifically includes: S310: Determine the total time from the earliest time to the latest time included in all the acquired original data, and record it as the total time window.

[0041] Determining the total time window provides a unified time range for subsequent calculations, allowing us to perform feature matching analysis within a clear time period and avoiding the uncertainty of the data range.

[0042] Suppose we have raw data from two domains (e.g., healthcare and finance). In the healthcare system, the earliest patient visit is recorded at 08:00 on March 1, 2025, and the latest is recorded at 23:59 on March 31, 2025. In the finance system, the earliest transaction is also recorded at 09:00 on March 1, 2025 (assuming that the financial system opens at 9:00), and the latest is also recorded at 17:00 on March 31 (the end of the business day). Taking into account the data from both systems, we determine the total time range to be from 08:00 on March 1, 2025, to 23:59 on March 31, 2025. The total duration is 30 days, 15 hours, and 59 minutes, which converts to approximately 735.98 hours, giving an approximate total time window of 736 hours.

[0043] S320: Determine a preset sub-time window according to the total time window.

[0044] Specifically, the preset sub-time windows divide the total time window into smaller time periods, which makes it easier for us to check the time matching between features in a local range, thereby improving the matching accuracy and efficiency.

[0045] Preferably, the preset sub-time window can be a ±1-hour window. This means that for any time point, we use it as the center and extend it forward and backward by 1 hour each to form a 2-hour sub-time window. For example, if an event occurs at 12:00 on January 5, 2025, then the corresponding preset sub-time window is from 11:00 to 13:00 on January 5, 2025.

[0046] In this embodiment, the preset sub-time windows and the total time window are measured in hours, and the ratio of the preset sub-time windows to the total time windows is preferably in the range of 0.001-0.1. In the above example, the total time window is 736 hours, and the preset sub-time window is 2 hours. The ratio is 2 / 736 ≈ 0.0027, which is within a reasonable range. This ratio ensures that possible time correlations can be detected without making the sub-windows too large, which may lead to inaccurate matching.

[0047] S330, obtaining the occurrence time difference between the first domain feature and the second domain feature within the same preset sub-time window. When the occurrence time difference is not greater than the preset duration, it is recorded as a match between the first domain feature and the second domain feature.

[0048] By setting a preset time length to determine whether features match, we can more accurately capture meaningful temporal associations between features, filter out events that are in the same sub-time window but have too long a time interval and may not have actual association, and improve the effectiveness of matching.

[0049] Specifically, suppose a patient in the medical system is rushed to the emergency room due to a sudden heart attack at 11:15 AM on March 22, 2025. Meanwhile, the stock price of a medical-related company in the financial system experiences unusual fluctuations at 11:30 AM on March 22, 2025. These two events fall within the same predefined sub-time window (10:15 AM to 12:15 PM). The time difference between their occurrences is 15 minutes. Since 15 minutes is not greater than 30 minutes, this is considered a match between the medical and financial system features.

[0050] The preset duration should be set based on the characteristics of the two cross-domain systems. In the healthcare-finance scenario, from a medical perspective, a patient's sudden illness could quickly impact the stock performance of related medical companies. Considering the financial market's response speed to events and the time it takes for medical events to spread to the financial market, the preset duration can be set between 0.5 and 1 hour.

[0051] S340, obtaining all the times that the first domain feature and the second domain feature generate matches within the total time window, and recording them as target matching times.

[0052] The number of target matches directly reflects the frequency of temporal matching between the features of the two systems. It is an important basis for calculating the temporal matching rate and provides specific data support for the subsequent analysis of the temporal correlation between the two systems.

[0053] For example, within the time window from 08:00 on March 1, 2025, to 23:59 on March 31, 2025, we iterate over all medical system characteristics (such as patient visits) and financial system characteristics (such as stock fluctuations and financial transactions) and apply the matching rules above. Statistics show that there were 30 matches between medical and financial system characteristics during this 736-hour period, so the target number of matches is 30.

[0054] S350: Obtain a time matching rate between the first domain feature and the second domain feature according to the target matching times and the total time window.

[0055] Among them, the time matching rate is : ,in, is the number of target matches, is the total time window.

[0056] The temporal matching ratio quantifies the temporal correlation between the characteristics of the healthcare and financial systems. This metric allows us to assess whether there is a significant temporal connection between the two systems, providing a reference for decision-making in healthcare investment and financial risk assessment.

[0057] The method for obtaining the temporal matching ratio between the first domain features and the second domain features disclosed in S310-S350 can effectively analyze the temporal correlation between the features of the medical system and the financial system, two seemingly unrelated domains, and discover potential cross-domain relationships, such as whether the outbreak of certain diseases will cause changes in the financial status of related medical companies. Quantifying the temporal matching ratio can provide valuable information for medical industry investors, financial institutions, and medical policymakers. Investors can adjust their investment strategies based on the temporal matching ratio, financial institutions can assess the risks of medical-related financial products, and medical policymakers can predict the impact of medical events on financial markets. This analytical method has certain flexibility and scalability, and can be applied to different time frames and data sizes. The preset sub-time windows and preset durations can also be adjusted as needed to accommodate different analytical needs.

[0058] Reference Figure 3 The method of S400 “obtaining the spatial correlation between the first domain feature and the second domain feature based on the external map knowledge base”, i.e., the method of obtaining the spatial correlation between the first domain feature and the second domain feature, specifically includes: S410: Call an external map knowledge base to determine the latitude and longitude information of the actual occurrence location corresponding to the first domain feature, and record it as the first location information.

[0059] S420: Call an external map knowledge base to determine the latitude and longitude information of the actual occurrence location corresponding to the second domain feature, and record it as the second location information.

[0060] Converting specific address information into longitude and latitude information provides a unified and standard data format for the subsequent accurate calculation of geographic distances. Longitude and latitude are a globally accepted geographic coordinate system that can accurately locate any point on the earth, avoiding calculation errors caused by non-standard address expressions or differences in address systems in different regions.

[0061] S430 : Obtaining a geographical distance between the first domain feature and the second domain feature according to the Haversine formula, the first location information, and the second location information.

[0062] The geographical distance is : ; .

[0063] in, is the latitude of the first location, is the latitude of the second location, is the latitude difference between the first and second locations (in radians), is the difference in longitude between the first and second locations (in radians), is the radius of the Earth (usually taken as 6371 kilometers (or 3958.8 miles).

[0064] The Haversine formula takes into account the spherical nature of the Earth and can relatively accurately calculate the actual distance between any two points on the Earth. In geospatial analysis, accurate distance calculation is the basis for determining the spatial relationship between locations and provides reliable data support for subsequent assessments of spatial relevance.

[0065] S440: Obtain spatial correlation between the first domain feature and the second domain feature according to the geographical distance.

[0066] Among them, the spatial correlation is : ; is the geographical distance, is the attenuation coefficient, which can be flexibly set according to actual needs.

[0067] Quantifying the spatial correlation between the characteristics of the medical and financial systems through geographic distance transforms abstract spatial relationships into concrete indicators. This helps intuitively compare the relationships between different medical and financial locations and uncover the spatial interactions between the two systems. For practitioners in the medical and financial industries, this quantitative spatial correlation analysis can provide a strong basis for their decision-making. For example, when seeking financial support, hospitals can prioritize financial institutions with close proximity and strong connections, and financial institutions can also focus on nearby medical institutions when investing in the medical industry.

[0068] The method for obtaining spatial correlations between first-domain features and second-domain features disclosed in S410-S440 enables spatial correlation analysis between features from two distinct domains: the medical and financial systems. This breaks down barriers between these two sectors. By analyzing the spatial relationships between event locations within the two systems, potential cross-domain connections can be identified, such as whether geographic proximity promotes interaction between medical innovation and financial investment. This provides new insights for further research into the medical-financial ecosystem. Accurate spatial correlation analysis results can provide precise decision-making support for medical companies, financial institutions, and relevant regulatory authorities. Medical companies can rationally plan financing channels and partners based on their spatial correlation with financial institutions. Financial institutions can assess investment risks and potential returns based on proximity and spatial correlation with medical sites, optimizing investment strategies. Regulators can formulate more effective policies and regulatory measures based on the spatial distribution and correlation between the two systems. This data-driven analysis method, based on geographic data, reveals the spatial connections between the medical and financial systems. This analysis helps promote the coordinated spatial development of the two industries, promote the rational allocation of medical and financial resources, and improve the quality of medical services and the efficiency of financial services for society as a whole.

[0069] Reference Figure 4 The method of S500 “obtaining the semantic relevance of the text corresponding to the first domain feature and the second domain feature”, i.e., the method of obtaining the semantic relevance of the text corresponding to the first domain feature and the second domain feature, specifically includes: S510: Obtain a first associated text of a first domain feature in a first type of system.

[0070] By clearly obtaining the text content related to medical information, a specific data basis is provided for subsequent semantic analysis. Complete medical information contains rich medical semantics, which can comprehensively reflect the relevant circumstances of the medical event and help to accurately analyze the semantic characteristics of the medical field.

[0071] Specifically, in a medical system, the first domain feature is a medical visit, which may include the patient's basic information, symptom description, diagnosis, treatment plan, etc. For example, a medical visit record such as "Patient Zhang San, male, 55 years old, was admitted to the hospital with cough and fever for three days, diagnosed with upper respiratory tract infection, and given antibiotics" is the first associated text.

[0072] S520: Use the BERT model to extract a semantic vector of the text field in the first associated text, and record it as a first semantic vector.

[0073] The first associated text, "Patient Zhang San, male, 55 years old, was admitted to the hospital with a cough and fever for three days, diagnosed with an upper respiratory tract infection, and given antibiotics," is input into the BERT model. The BERT model processes the text, converting each word or phrase in the text into a corresponding semantic vector. Ultimately, it extracts the semantic vector for the entire text field. This vector, represented as a multidimensional array, contains the semantic features of the medical information and is recorded as the first semantic vector.

[0074] S530: Obtain a second associated text of the second domain feature in the second type of system.

[0075] Obtaining complete text related to transaction information provides a specific semantic description of transaction events in the financial system. These texts contain key information about the transaction, providing data support for analyzing semantic features in the financial field and facilitating semantic association analysis with medical consultation information.

[0076] Specifically, in the target financial system, the second domain feature is a transaction. This transaction information may include the transaction time, amount, recipient, and purpose. For example, a transaction such as "On June 24, 2025, Zhang San paid 500 yuan for medical expenses to a certain hospital" constitutes the second associated text.

[0077] S540: Use the BERT model to extract a semantic vector of the text field in the second associated text, and record it as a second semantic vector.

[0078] The second associated text "On June 24, 2025, Zhang San paid 500 yuan in medical expenses to a hospital" is input into the BERT model. The BERT model processes the text, converts the words and phrases in it into semantic vectors, and extracts the semantic vector of the entire text field, which is recorded as the second semantic vector.

[0079] S550: Calculate the vector similarity between the first semantic vector and the second semantic vector using a cosine similarity formula.

[0080] Among them, the vector similarity is : A is the first semantic vector, is the second semantic vector, is the modulus of the first semantic vector, is the modulus of the second semantic vector.

[0081] The first semantic vector and the second semantic vector have been obtained. The cosine similarity formula is used to calculate the vector similarity between them. Cosine similarity measures their similarity by calculating the cosine value of the angle between two vectors. The value range is between -1 and 1. The closer the value is to 1, the more similar the two vectors are.

[0082] S560: Obtain semantic relevance of texts corresponding to the first domain feature and the second domain feature based on the vector similarity.

[0083] The semantic relevance is : , is the vector similarity.

[0084] Vector similarity is used to quantify the semantic relevance of the text corresponding to the first-domain features and the second-domain features, converting abstract semantic relationships into specific indicators. This helps to quickly screen out information pairs with strong semantic relevance from massive amounts of medical and transaction information, discover potential connections between the medical and financial systems, and provide a basis for further data analysis and decision-making.

[0085] The method for obtaining semantic correlations between text corresponding to first-domain features and second-domain features, disclosed in S510-S560, enables semantic correlation analysis of different types of information in medical and financial systems, breaking down domain boundaries. By analyzing the semantic relationships between medical visit information and transaction information, potential connections between medical behaviors and financial transactions can be discovered, such as the correlation between patient payment patterns and disease types. Accurate semantic correlation analysis helps extract valuable information from massive amounts of medical and financial data. For example, it can uncover payment patterns for treatment of certain diseases or the financial transaction habits of patients with specific conditions. This information can provide deep insights for healthcare providers, financial institutions, and regulators, enabling them to make more informed decisions. For the healthcare industry, understanding the semantic correlations between medical visit information and transaction information can help hospitals optimize billing strategies and manage medical insurance reimbursements. For the financial industry, it can assess the risks and potential of healthcare-related transactions and develop financial products more suitable for healthcare scenarios. Furthermore, regulators can use this correlation information to strengthen oversight of the medical and financial markets, safeguarding patient rights and market stability.

[0086] The method of S600 of "performing weighted fusion of the temporal matching rate, spatial correlation, and semantic correlation to obtain the correlation edge weight between the first domain feature and the second domain feature" specifically includes: S610: Determine key risk factors in the data scenarios based on the data scenarios corresponding to the first and second types of systems. S620, dynamically determining a time matching weight, a spatial association weight, and a semantic association weight based on key risk factors; S630 , performing weighted summation of the time matching rate, the spatial correlation, and the semantic correlation according to the time matching weight, the spatial correlation weight, and the semantic correlation weight, to obtain the correlation edge weight between the first domain feature and the second domain feature.

[0087] The associated edge weight is : ;in, , is the time matching weight, is the spatial association weight, is the semantic association weight, For time correlation, For spatial correlation, For semantic relevance.

[0088] Furthermore, for medical data scenarios, the time matching between consultation time and medication records is the key leakage path, so time correlation is more important in data anonymization in this scenario. For social media data scenarios, semantic analysis of text content (such as "cancer hospital check-in") is the main source of risk, so semantic correlation is more important in data anonymization in this scenario. For financial transaction scenarios, the proximity of transaction locations to sensitive places (such as hospitals) is the core risk, so spatial correlation is more important in data anonymization in this scenario. Based on the core risk factors in each scenario, the time matching weight ( ), spatial association weight ( ) and semantic association weight ( ) are assigned different values, as follows: 1) Medical data scenario: Set time matching weight = 0.5, spatial association weight = 0.2, semantic association weight = 0.3. 2) Social Media Data Scenario: Setting Time Matching Weight = 0.2, spatial association weight = 0.3, semantic association weight = 0.5. 3) Financial transaction scenario: Setting time matching weight = 0.3, spatial association weight = 0.5, semantic association weight = 0.2.

[0089] Regarding the method of S700 “determining the path risk level formed by the first domain feature and the second domain feature according to the associated edge weight”, in a first embodiment, the method specifically includes: A100, obtaining a third domain feature in a tripartite system that is associated with a first domain feature and a second domain feature in a location.

[0090] Among them, the three-party system is a social media system, the first field feature is medical information, the second field feature is financial transaction information, and the third field feature is social media related information. Furthermore, the third field feature is preferably location-related text information associated with the location corresponding to the financial transaction information.

[0091] Introducing the third domain characteristics of social media systems can enrich the information dimension. The information in social media often contains real-time feedback from the public and potential risk signals. By obtaining social media information that is location-related to the first and second domain characteristics, potential factors that may affect medical treatment and financial transactions can be discovered, providing a more comprehensive basis for subsequent risk assessment.

[0092] A200,constructs a directed weighted graph based on all first domain features, second domain features, and third domain features.,The edges in the directed weighted graph are the,association edge weights between two adjacent features, and the nodes are,each feature.

[0093] Directed weighted graphs can intuitively represent the relationships and strengths of associations between features in different domains. By using edge weights, the degree of association between different features can be quantified, providing a clear structure and data foundation for subsequent calculations of path risk. They visualize complex information relationships and facilitate analysis and understanding of the interactions between features in different domains.

[0094] A300, based on the directed weighted graph, determines each association path including the first domain feature, the second domain feature, and the third domain feature, and obtains the total association weight of each association path.

[0095] In the constructed directed weighted graph, find the connection path that includes medical information, financial transaction information, and social media related information. For example, consider a connection path: medical information → financial transaction information → social media related information. To calculate the total connection weight of this path, simply multiply the weights of the edges along the path. Assuming that in the above path, the edge weight from medical information to financial transaction information is 0.8, and the edge weight from financial transaction information to social media related information is 0.7, then the total connection weight of this connection path is 0.8 × 0.7 = 0.56.

[0096] Determining the association path and the total association weight can quantify and integrate the association relationship between features in different fields. By calculating the total association weight, the comprehensive association degree between different features on a path can be measured, which helps to screen out paths with stronger associations and provide key quantitative indicators for subsequent risk assessment.

[0097] A400 determines the path sensitivity corresponding to each associated path according to the path sensitivity formula.

[0098] Among them, The path sensitivity corresponding to the associated path is : ,in, For the The node sensitivity corresponding to the associated path, For the The total association weight corresponding to the association paths.

[0099] Path sensitivity comprehensively considers the importance of nodes and the association strength of paths. Through node sensitivity, the impact of certain key information nodes on path risk can be highlighted; through the total association weight, the degree of association between different features on the path can be reflected. Path sensitivity can more accurately assess the degree of risk contained in an associated path, providing a more reasonable basis for determining the path risk level.

[0100] A500 determines the path risk level corresponding to the path sensitivity based on the preset sensitivity range.

[0101] The preset sensitivity range may include: when the path sensitivity is 0-0.7, determining that the path risk level corresponding to the path sensitivity is low risk; when the path sensitivity is greater than 0.7, determining that the path risk level corresponding to the path sensitivity is high risk.

[0102] Converting path sensitivity into specific risk levels makes risk assessment results more intuitive and easier to understand. Different risk levels can provide clear references for decision makers, allowing them to take appropriate measures based on the risk levels, such as focusing on monitoring and intervening in high-risk paths and conducting routine management of low-risk paths.

[0103] This solution integrates information from three fields: healthcare, finance, and social media. It can comprehensively assess the path risks formed by the characteristics of different fields. By considering the correlations between information from different fields, it can discover potential risks that are difficult to detect in a single field, providing an effective method for cross-domain risk management. Through the calculation of directed weighted graphs and path sensitivity, complex information correlations and risk levels are quantified, making risk assessment more objective and accurate. At the same time, directed weighted graphs visualize information relationships, making it easier for managers to intuitively understand the correlations between different information and the distribution of risks. Based on the path risk level, decision makers can develop targeted risk management strategies. For high-risk paths, timely measures can be taken to control and prevent risks. For low-risk paths, resources can be reasonably allocated for routine management. This helps improve the efficiency and effectiveness of risk management and ensure the stable operation of the system.

[0104] Regarding the method of S700 “determining the path risk level formed by the first domain feature and the second domain feature according to the associated edge weight”, in a second embodiment, the method specifically includes: B100, according to the path sensitivity formula, determine the path sensitivity corresponding to each pair of the first domain feature and the second domain feature.

[0105] Among them, The path sensitivity corresponding to the first domain feature and the second domain feature is : ,in, For the The sensitivity of nodes corresponding to the first domain features and the second domain features, For the The associated edge weights corresponding to the first domain features and the second domain features.

[0106] B200 determines the path risk level corresponding to the path sensitivity based on the preset sensitivity range.

[0107] This solution provides a concise and efficient method to evaluate the path risk between the characteristics of the first field and the characteristics of the second field. Through simple formula calculation and risk level classification, it can quickly process a large number of feature combinations, saving evaluation time and cost. This embodiment can perform correlation analysis on the characteristics of different fields (such as medical information and financial transaction information) to discover the potential risk relationship between different fields, which helps to break down the barriers between fields and realize cross-field risk management and decision support. The path risk level provides decision makers with a clear decision-making basis. Whether it is the medical industry, the financial industry or the regulatory authorities, they can take corresponding measures according to the risk level to optimize resource allocation, reduce potential risks, and ensure the safe and stable operation of the system.

[0108] The method of S800, "dynamically calling a corresponding anonymization strategy based on the path risk level, and executing the anonymization strategy on the data corresponding to the first domain feature and the second domain feature," specifically includes: A100, when the path risk level is high risk, the first anonymization strategy is called.

[0109] When the path risk level is high, it means that the data may cause serious harm to patients or related entities after being leaked. Calling the first anonymization strategy of strong anonymization can protect the privacy and security of the data to the greatest extent and reduce the risk of data leakage.

[0110] A200 , respectively obtain a first associated text of the first domain feature in the first type of system and a second associated text of the second domain feature in the second type of system.

[0111] The associated text contains more detailed information related to the first-domain features and the second-domain features. Obtaining these associated texts can provide a more comprehensive understanding of the context and background of the data, which helps to accurately extract key fields in the subsequent process and avoid missing important information that may need to be anonymized, thereby improving the anonymization effect.

[0112] A300 extracts key fields from the first associated text and the second associated text respectively.

[0113] Extracting key fields can focus on the most critical and sensitive parts of the data, avoiding unnecessary processing of the entire associated text. This can improve the efficiency of anonymization while ensuring that only information that truly needs to be protected is anonymized, reducing the impact on data availability.

[0114] A400 performs anonymization adjustment on all key fields based on the first anonymization strategy.

[0115] By anonymizing key fields, sensitive information in the data can be effectively protected and privacy risks caused by data leakage can be reduced. In high-risk situations, this strong anonymization process can ensure that even if the data is illegally obtained, it is difficult for attackers to identify individual identities from the anonymized data, thereby protecting the rights and interests of data subjects.

[0116] The methods disclosed in A100-A400 dynamically call anonymization strategies based on the path risk level, and can flexibly adjust the strength of anonymization based on the actual risk situation. A strong anonymization strategy is adopted in high-risk situations, and other anonymization strategies are adopted in low-risk situations, achieving dynamic anonymization adjustment while achieving a balance between privacy protection and data availability. By obtaining associated text and extracting key fields, sensitive information in the data can be comprehensively identified and protected. Not only are the first-domain features and second-domain features themselves anonymized, but related detailed information is also processed, improving the integrity of data protection. In today's strict privacy regulatory environment, this solution helps enterprises and institutions meet regulatory requirements and avoid legal risks due to data leaks. By effectively anonymizing sensitive data, the privacy rights and interests of data subjects can be protected and user trust in data processing can be enhanced.

[0117] Furthermore, the method of A400 "adjusting the anonymization of all key fields based on the first anonymization strategy" specifically includes: A410, determines the actual type of the key field.

[0118] Accurately determining the actual type of key fields is the basis for subsequent targeted anonymization processing. Different types of fields have different characteristics and sensitivities. Only by clarifying the field type can the appropriate anonymization method be adopted to ensure the effectiveness and rationality of the anonymization processing.

[0119] A420, when the actual type is a time type, generalize the precise time information corresponding to the key field to a broad time interval.

[0120] In this embodiment, a generalized time processing method is adopted, that is, the precise time information is generalized into a broader time interval. For example, a time accurate to the minute such as "2023-05-10 14:30" is converted into "May 2023". In this way, the accuracy of the time information is reduced, making it difficult for attackers to associate and identify personal information based on the precise time.

[0121] Time information often serves as a crucial clue for linking and identifying personal information. By generalizing precise time information to a broad time range, the accuracy of time information is reduced, making it difficult for attackers to use precise time to link other data and identify individuals. At the same time, the approximate time range is preserved to a certain extent, preventing the data's temporal attributes from being completely lost, thus ensuring the data's usability in certain scenarios.

[0122] A430, when the actual type is a location type, generalizes the precise geographic information corresponding to the key field into a broad geographic range, and the GPS accuracy of the broad geographic range is lower than the GPS accuracy of the precise geographic information.

[0123] In this embodiment, a method of blurring the location is used, specifically to reduce the accuracy of the GPS coordinates from 3 decimal places to 1 place, which will expand the positioning error from approximately 10 meters to 1 kilometer. In this way, although the approximate geographic location information is still retained, the specific location of the individual cannot be accurately determined, thereby protecting the individual's location privacy.

[0124] Precise geographic information can directly locate an individual's specific location, which can easily lead to the leakage of personal location privacy. Generalizing precise geographic information into a broad geographic range, while retaining the approximate geographic location information, increases the positioning error range, making it impossible for attackers to accurately determine the individual's specific location, effectively protecting the individual's location privacy; moreover, a broad geographic range can still provide a certain geographic location reference, meeting the needs of some applications that do not require high location accuracy.

[0125] A440, when the actual type is a numeric type, adds Laplace noise to the numeric field corresponding to the key field.

[0126] Numerical key fields may contain sensitive information, such as financial transaction amounts and personal health indicators. Adding Laplace noise can slightly perturb the values without changing the overall distribution of the data. This makes it difficult for attackers to accurately infer the original sensitive information from the perturbed values, even if the data is compromised, thus protecting data privacy. Furthermore, the noised values still retain statistical significance and will not significantly impact numerically-based analyses and calculations.

[0127] The method disclosed in A410-A440 adopts different anonymization methods for different types of key fields, and can perform targeted processing according to the characteristics and sensitivity of the fields, thereby achieving more refined privacy protection. Compared with the unified anonymization method, this method can better protect privacy while minimizing the impact on data availability. By anonymizing various sensitive key fields such as time, place, and numerical values, the accuracy and sensitivity of information that can be used to identify personal identity in the data are comprehensively reduced, greatly enhancing data security and reducing the risk of data leakage. Under the requirements of strict privacy regulations, this solution helps enterprises and institutions meet compliance needs. At the same time, on the basis of protecting privacy, the availability of data is retained as much as possible, so that the data can still be used to a certain extent for business activities such as analysis and statistics after anonymization, thus achieving a balance between privacy protection and business needs.

[0128] Furthermore, for the method S800 of "dynamically calling the corresponding anonymization strategy according to the path risk level, and executing the anonymization strategy on the data corresponding to the first domain feature and the second domain feature", it specifically includes: when the path risk level is low risk, the original data can be retained, or the second anonymization strategy can be called for light processing.

[0129] Among them, "calling the second anonymization strategy for light processing" includes: for example, in the age field, directly converting the specific age value into an age segment, such as converting the specific age into a range such as "20-30 years old". This can protect personal privacy to a certain extent, while retaining the approximate information of the data, meeting some analysis needs that do not require high data accuracy.

[0130] In the medical system, patient information includes information about a patient named "Li Si" undergoing appendectomy surgery at "XX Hospital" on June 20, 2025, along with detailed symptoms and diagnosis. In the financial system, transaction information indicates that Li Si paid 5,000 yuan for surgery to "XX Hospital" on June 20, 2025. A risk assessment of the link between these two sets of information revealed that due to the hospital's high patient volume and the relatively weak correlation between the transaction records, this path was classified as low risk. Accurately determining the path risk level is fundamental to selecting an appropriate anonymization strategy. Only with a clear risk level can appropriate measures be implemented to avoid over- or under-protection. In both medical and financial data scenarios, different risk levels correspond to different privacy protection requirements. Properly determining risk levels helps balance privacy protection and data utilization.

[0131] For example, a medical research team is studying the effectiveness and costs of appendicitis surgery over a specific time period. They need accurate patient and transaction data to analyze surgical costs, recovery outcomes, and other factors. Given the extremely high data accuracy requirements and the low risk assessment, they can retain the original data. For example, in the medical system, the data includes "Li Si," "June 20, 2025," "XX Hospital," and "appendicitis surgery," while in the financial system, the data includes "Li Si," "June 20, 2025," "XX Hospital," and "5,000 yuan." Preserving the original data satisfies business scenarios requiring high data accuracy and ensures the smooth operation of the business. In fields like medical research, accurate data is crucial for drawing accurate research conclusions, and retaining the original data provides a reliable basis for research.

[0132] Specific examples of invoking the second anonymization strategy for light processing include: Time information processing can include generalizing the date "June 20, 2025" in medical and transaction information to "June 2025." This reduces the accuracy of the time, but still preserves the approximate time range. Patient identity processing can include replacing the patient name "Li Si" with "Patient L," concealing the patient's true identity while still distinguishing between different patients. Transaction amount processing can include converting a transaction amount of "5,000 yuan" into a range of "4,000-6,000 yuan." While obscuring the specific amount, it preserves the approximate range. Location information processing can include generalizing "XX Hospital" to "a certain hospital in this city," reducing the accuracy of the location information. Invoking the second anonymization strategy for light processing can protect individual privacy to a certain extent while preserving the general information of the data. It is suitable for analysis scenarios where data accuracy is not critical. In some macro-level medical data statistics and analysis, lightly anonymized data is sufficient to meet analytical requirements and effectively reduce the risk of patient privacy leakage.

[0133] In this embodiment, dynamic processing methods are selected based on the path risk level, enabling flexible response to diverse risk scenarios. In scenarios involving medical and financial data, different combinations of medical and transaction information may present different risks. In low-risk scenarios, the original data can be retained to meet high-precision requirements such as scientific research, or lightly anonymized to enhance privacy protection, achieving a dynamic balance between privacy protection and business needs. Even with lightly anonymized data, the data's general information can be retained, allowing it to be used for analytical and statistical work requiring less precision. For example, this could include analyzing the medical revenue scale of different hospitals or disease visit trends within a specific time period. This approach improves data usability while protecting privacy, avoiding situations where data becomes unusable due to over-protection. In the medical and financial sectors, data privacy protection is subject to strict regulatory oversight. This solution helps enterprises and institutions meet compliance requirements. By appropriately processing low-risk data, it reduces the legal risks associated with data leaks, safeguards patients' privacy rights, and enhances user trust in data processing by medical and financial institutions.

[0134] In this embodiment, the first type of system and the second type of system can be a medical system and a financial system respectively, or can be a financial system and a social media system respectively, or can be other cross-domain related scenario systems such as a medical system and a social media system.

[0135] Example 1: Data anonymization scenario in a cross-domain medical-financial scenario.

[0136] The scenario is as follows: A hospital needs to anonymize patients' medical records for disease research. At the same time, it needs to prevent attackers from re-identifying patients by linking financial transaction data. Medical records and financial transaction data may be correlated to a certain extent. If not handled properly, this correlation may expose patients' private information.

[0137] 1) Input data: Medical data includes the patient's visit time (accurate to the minute), the GPS coordinates of the visit location (accurate to three decimal places), and the diagnosis (lung cancer). This information is sensitive enough to identify the patient. Financial data includes the transaction time (accurate to the minute), the GPS coordinates of the transaction location (accurate to three decimal places), and the transaction amount (2,000 yuan). This financial transaction data may be correlated with medical data, for example, the patient may make related purchases nearby after the visit.

[0138] 2) Processing: Identifying associated risks. Specifically, the path "visit time → transaction location" was discovered. Analysis revealed a strong association between the visit time and the transaction location, with an edge weight of 0.955 and a sensitivity of 2.02. If the path risk level formed by the first and second domain characteristics is determined to be high, the visit time can be generalized to "May 2023" and the GPS downgraded to "31.2°N, 121.4°E." This data anonymization adjustment reduced the attack success rate from 72% using traditional methods to 2%, effectively defending against attacks that exploit data association for identity recognition and significantly improving data security and privacy.

[0139] Example 2: Finance-Social Cross-Domain Association Scenario: A bank needs to anonymize credit card transaction data to support consumption trend analysis and prevent users from being re-identified by associating with social media check-in data.

[0140] 1) Input data: Financial data includes: transaction time 2023-08-15 19:30, transaction location 31.2304°N, 121.4737°E (café), and transaction amount 58 RMB. Social data includes: check-in time 2023-08-15 19:45, check-in location 31.2305°N, 121.4736°E, and text: "Coffee time after weekend overtime #workplacelife."

[0141] 2) After identifying the associated risks, the path is obtained: transaction location → check-in location → text content. The corresponding edge weight is 0.942 and the sensitivity is 2.02.

[0142] If the path formed by the first, second, and third domain features is determined to be high risk, the transaction location can be downgraded to "31.2°N, 121.4°E" and the text replaced with "drinks consumed during break time." This data anonymization adjustment reduces the attack success rate from 22% using traditional methods to 0.33%.

[0143] In the existing technology, traditional data anonymization technologies (such as k-anonymity and differential privacy) focus only on a single dataset and are unable to identify correlation risks between cross-domain data. The data anonymization adjustment method based on dynamic correlation risk analysis disclosed in this application obtains raw data from first- and second-category systems, extracts first-domain features and second-domain features with temporal and spatial correlations, and calculates correlation edge weights based on a comprehensive consideration of temporal matching rate, spatial correlation, and semantic correlation to determine the path risk level. This enables the solution to accurately identify potential correlation risks between cross-domain data, such as the possible spatiotemporal matching relationship between medical treatment records and financial transaction records, effectively preventing the leakage of sensitive information through cross-domain data correlation.

[0144] Existing solutions lack a responsiveness mechanism to dynamic changes such as new public datasets and the expansion of external knowledge bases. The data anonymization adjustment method based on dynamic correlation risk analysis disclosed in this application is based on dynamic correlation risk analysis. It can calculate correlation edge weights and path risk levels based on real-time data conditions and external knowledge bases (such as external map knowledge bases). When new public datasets or external knowledge base expansions appear, the solution can promptly reassess the correlation risks between data and dynamically adjust the anonymization strategy, avoiding dynamic risks such as the geographic location accuracy issue in the specific case, ensuring that adjustments can be made quickly to protect data security when risks change.

[0145] Existing technologies suffer from the problems of excessive anonymization damaging data availability, while weak anonymization fails to guarantee privacy. This application discloses a data anonymization adjustment method based on dynamic correlation risk analysis. This method dynamically calls the corresponding anonymization strategy based on the path risk level. For low-risk paths, a mild anonymization process can be used, which protects a certain degree of privacy while retaining data availability and avoiding data distortion caused by excessive anonymization. For high-risk paths, a stronger anonymization strategy is used to effectively resist correlation inference attacks and ensure data privacy, thus achieving a balance between utility and privacy.

[0146] The data anonymization adjustment method based on dynamic correlation risk analysis disclosed in this application can effectively prevent sensitive information from being leaked through cross-domain data correlation by accurately identifying correlation risks between cross-domain data and dynamically adjusting the anonymization strategy. This solution can greatly improve the security and privacy protection of data. Whether it is medical data, financial data, or data in other fields, it can be more reliably protected in cross-domain scenarios, reducing the risks and losses caused by data leakage. The method can adjust the anonymization strategy in a timely manner according to the dynamically changing risk situation. It has strong dynamic adaptability and flexibility. When faced with new public data sets, external knowledge base expansion, and other dynamic risks, it can respond quickly to ensure that the data is always in a secure state. This makes the solution more practical and stable in complex and changing real-world environments. Under the premise of ensuring data privacy, by reasonably adjusting the anonymization strategy, the damage to data availability caused by excessive anonymization is avoided. A certain degree of data availability is retained, so that the data can still be used for various analysis and business applications, such as medical time series analysis and financial risk assessment. This helps to fully realize the business value of data and promote data-driven decision-making and innovation. This solution is applicable to two types of systems that constitute cross-domain scenarios and has broad versatility. Whether it is the system association between different industries or the association between different types of systems in the same industry, this solution can be used to adjust data anonymization. At the same time, the framework and method of the solution are scalable and can be further improved and optimized according to actual needs, such as adding more association factors or adjusting the weight calculation method.

[0147] In a second aspect, the present application discloses a data anonymization adjustment system based on dynamic correlation risk analysis, which is used to execute the data anonymization adjustment method based on dynamic correlation risk analysis disclosed in the first aspect of the present application. The system specifically includes: The raw data acquisition module is used to acquire all raw data associated with the first type system and the second type system within a preset period; wherein the first type system and the second type system are two types of systems constituting a cross-domain scenario; A domain feature acquisition module is used to extract two types of cross-domain scene information from all raw data. The two types of cross-domain scene information include first domain features and second domain features with temporal correlation and spatial correlation; A time matching rate acquisition module, configured to obtain a time matching rate between the first domain feature and the second domain feature based on a time window overlapping frequency; A spatial correlation acquisition module, configured to acquire the spatial correlation between the first domain feature and the second domain feature based on an external map knowledge base; A semantic relevance acquisition module, used to acquire the semantic relevance of the text corresponding to the first domain feature and the second domain feature; A fusion module is used to perform weighted fusion of the temporal matching rate, spatial correlation, and semantic correlation to obtain the correlation edge weight between the first domain feature and the second domain feature; A path risk level acquisition module is used to determine the path risk level formed by the first domain feature and the second domain feature according to the associated edge weight; The dynamic anonymization module is used to dynamically call the corresponding anonymization strategy according to the path risk level, and execute the anonymization strategy on the data corresponding to the first domain feature and the second domain feature.

[0148] A computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc.

[0149] The processor can be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and can control other components in the computer device to perform desired functions. In one embodiment of the present disclosure, the processor is configured to execute the computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the data anonymization adjustment method based on dynamic correlation risk analysis described in various embodiments of the present disclosure.

[0150] Those skilled in the art should understand that in order to solve the technical problem of how to obtain a good user experience, this embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the scope of protection of this disclosure.

[0151] like Figure 5 The present invention provides a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Figure 5 The computer device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0152] like Figure 5As shown, a computer device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) or programs loaded from a storage device into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0153] Typically, the following devices can be connected to the I / O interface: input devices such as sensors or visual information acquisition devices; output devices such as display screens; storage devices such as tapes and hard disks; and communication devices. The communication device can allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or by wire to exchange data. Figure 5 A computer device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0154] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the data anonymization adjustment method based on dynamic correlation risk analysis of the embodiment of the present disclosure are performed.

[0155] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.

[0156] According to an embodiment of the present disclosure, a computer-readable storage medium stores non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the data anonymization adjustment method based on dynamic correlation risk analysis described in each embodiment of the present disclosure are performed.

[0157] The above-mentioned computer-readable storage media include, but are not limited to, optical storage media (e.g., CD-ROMs and DVDs), magneto-optical storage media (e.g., MOs), magnetic storage media (e.g., magnetic tapes or mobile hard disks), media with built-in rewritable non-volatile memory (e.g., memory cards), and media with built-in ROM (e.g., ROM cartridges).

[0158] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.

[0159] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0160] In the present disclosure, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "including," "comprising," "having," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0161] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.

[0162] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.

[0163] Various changes, substitutions, and modifications may be made to the technology described herein without departing from the teachings defined by the appended claims. Moreover, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of things, means, methods, and actions described above. Currently existing or later developed processes, machines, manufactures, compositions of things, means, methods, or actions that perform substantially the same function or achieve substantially the same results as the corresponding aspects described herein may be utilized. Accordingly, the appended claims include within their scope such processes, machines, manufactures, compositions of things, means, methods, or actions.

[0164] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0165] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A data anonymization adjustment method based on dynamic correlation risk analysis, characterized in that: include: Obtain all raw data associated with the first type system and the second type system within a preset period; wherein the first type system and the second type system are two types of systems constituting a cross-domain scenario; Extracting two types of cross-domain scene information from all the original data, the two types of cross-domain scene information including first domain features and second domain features having temporal correlation and spatial correlation; Obtaining a time matching rate between the first domain feature and the second domain feature based on a time window overlapping frequency; Acquire the spatial correlation between the first domain feature and the second domain feature based on an external map knowledge base; Obtaining semantic relevance between texts corresponding to the first domain feature and the second domain feature; Performing weighted fusion on the time matching rate, the spatial correlation, and the semantic correlation to obtain an association edge weight between the first domain feature and the second domain feature; Determining a path risk level formed by the first domain feature and the second domain feature according to the associated edge weight; The corresponding anonymization strategy is dynamically called according to the path risk level, and the anonymization strategy is executed on the data corresponding to the first domain feature and the second domain feature.

2. The data anonymization adjustment method based on dynamic correlation risk analysis according to claim 1 is characterized in that: The obtaining of a time matching rate between the first domain feature and the second domain feature based on the time window overlapping frequency includes: Determine the total time from the earliest time to the latest time included in all the acquired raw data, and record it as the total time window; Determining a preset sub-time window according to the total time window; Obtaining the occurrence time difference between the first domain feature and the second domain feature within the same preset sub-time window; when the occurrence time difference is not greater than a preset time length, it is recorded as a match between the first domain feature and the second domain feature; Obtain all the times that the first domain feature and the second domain feature generate matches within the total time window, and record them as target matching times; Obtaining a time matching rate between the first domain feature and the second domain feature according to the target matching times and the total time window; The time matching rate is : ,in, is the number of target matches, is the total time window.

3. The data anonymization adjustment method based on dynamic correlation risk analysis according to claim 2 is characterized in that: The acquiring the spatial correlation between the first domain feature and the second domain feature based on the external map knowledge base includes: Calling an external map knowledge base to determine the latitude and longitude information of the actual occurrence location corresponding to the first domain feature, and recording it as the first location information; Calling an external map knowledge base to determine the latitude and longitude information of the actual location corresponding to the second domain feature, and recording it as the second location information; Obtaining a geographical distance between the first domain feature and the second domain feature according to the Haversine formula, the first location information, and the second location information; Obtaining the spatial correlation between the first domain feature and the second domain feature according to the geographical distance; The spatial correlation is : ;in, is the geographical distance, is the attenuation coefficient.

4. The data anonymization adjustment method based on dynamic correlation risk analysis according to claim 3 is characterized in that: The obtaining of the semantic relevance of the text corresponding to the first domain feature and the second domain feature includes: Obtaining a first associated text of the first domain feature in the first type of system; Extracting a semantic vector of the text field in the first associated text using the BERT model, recorded as a first semantic vector; Obtaining a second associated text of the second domain feature in the second type of system; Extracting a semantic vector of the text field in the second associated text using the BERT model, recorded as a second semantic vector; Calculating the vector similarity between the first semantic vector and the second semantic vector using a cosine similarity formula; Obtaining semantic relevance of texts corresponding to the first domain feature and the second domain feature according to the vector similarity; The semantic relevance is : , is the vector similarity.

5. The data anonymization adjustment method based on dynamic correlation risk analysis according to claim 4 is characterized in that: The weighted fusion of the time matching rate, the spatial correlation, and the semantic correlation to obtain the correlation edge weight between the first domain feature and the second domain feature includes: Determine key risk factors in the data scenarios according to the data scenarios corresponding to the first and second types of systems; Dynamically determine the time matching weight, spatial association weight, and semantic association weight based on the key risk factors; Performing weighted summation of the time matching rate, the spatial correlation, and the semantic correlation according to the time matching weight, the spatial correlation weight, and the semantic correlation weight to obtain an association edge weight between the first domain feature and the second domain feature; The associated edge weight is : ; in, , is the time matching weight, is the spatial association weight, is the semantic association weight, For time correlation, For spatial correlation, For semantic relevance.

6. The data anonymization adjustment method based on dynamic correlation risk analysis according to claim 5 is characterized in that: Determining a path risk level formed by the first domain feature and the second domain feature according to the associated edge weight includes: Acquire a third domain feature in a tripartite system that has a location association with the first domain feature and the second domain feature; Based on all the first domain features, the second domain features, and the third domain features, a directed weighted graph is constructed, where an edge in the directed weighted graph represents an associated edge weight between two adjacent features, and a node represents each feature; Based on the directed weighted graph, determining each association path including the first domain feature, the second domain feature, and the third domain feature, and obtaining a total association weight of each association path; Determining the path sensitivity corresponding to each associated path according to a path sensitivity formula; No. The path sensitivity corresponding to the associated path is : ,in, For the The node sensitivity corresponding to the associated path, For the The total association weight corresponding to the association paths; According to the preset sensitivity range, a path risk level corresponding to the path sensitivity is determined.

7. The data anonymization adjustment method based on dynamic correlation risk analysis according to claim 5 is characterized in that: Determining a path risk level formed by the first domain feature and the second domain feature according to the associated edge weight includes: Determine the path sensitivity corresponding to each pair of the first domain feature and the second domain feature according to the path sensitivity formula; No. The path sensitivity corresponding to the first domain feature and the second domain feature is : ,in, For the The node sensitivity corresponding to the first domain feature and the second domain feature, For the The weight of the associated edges corresponding to the first domain feature and the second domain feature; According to the preset sensitivity range, a path risk level corresponding to the path sensitivity is determined.

8. The data anonymization adjustment method based on dynamic correlation risk analysis according to claim 1 is characterized in that: The dynamically calling a corresponding anonymization strategy according to the path risk level and executing the anonymization strategy on data corresponding to the first domain feature and the second domain feature includes: When the path risk level is high risk, calling a first anonymization strategy; Respectively obtaining a first associated text of the first domain feature in the first type of system and a second associated text of the second domain feature in the second type of system; extracting key fields from the first associated text and the second associated text respectively; Anonymization adjustment is performed on all the key fields based on the first anonymization strategy.

9. The data anonymization adjustment method based on dynamic correlation risk analysis according to claim 8, characterized in that: The anonymizing adjustment of all the key fields based on the first anonymization strategy includes: Determine the actual type of the key field; When the actual type is a time type, generalizing the precise time information corresponding to the key field into a broad time interval; When the actual type is a location type, generalizing the precise geographic information corresponding to the key field into a broad geographic interval, wherein the GPS accuracy of the broad geographic interval is lower than the GPS accuracy of the precise geographic information; When the actual type is a numerical type, Laplace noise is added to the numerical field corresponding to the key field.

10. A data anonymization adjustment system based on dynamic correlation risk analysis, characterized in that: include: A raw data acquisition module, configured to acquire all raw data associated with the first and second types of systems within a preset period; wherein the first and second types of systems are two types of systems constituting a cross-domain scenario; A domain feature acquisition module, configured to extract two types of cross-domain scene information from all the raw data, wherein the two types of cross-domain scene information include a first domain feature and a second domain feature having temporal correlation and spatial correlation; a time matching rate acquisition module, configured to obtain a time matching rate between the first domain feature and the second domain feature based on a time window overlapping frequency; A spatial correlation acquisition module, configured to acquire the spatial correlation between the first domain feature and the second domain feature based on an external map knowledge base; A semantic relevance acquisition module, configured to acquire the semantic relevance of texts corresponding to the first domain feature and the second domain feature; a fusion module, configured to perform weighted fusion on the time matching rate, the spatial correlation, and the semantic correlation to obtain an association edge weight between the first domain feature and the second domain feature; a path risk level acquisition module, configured to determine a path risk level formed by the first domain feature and the second domain feature according to the associated edge weight; A dynamic anonymization module is used to dynamically call a corresponding anonymization strategy according to the path risk level, and execute the anonymization strategy on the data corresponding to the first domain feature and the second domain feature.

11. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data anonymization adjustment method based on dynamic correlation risk analysis described in any one of claims 1-9.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the data anonymization adjustment method based on dynamic correlation risk analysis according to any one of claims 1 to 9.

13. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Private data identification and desensitization method, system and device and storage medium

    CN116049877A

  • Cross-unit data management method

    CN119989418A

  • Cross-industry data security sharing method and system based on data desensitization and medium

    CN120223391A

  • Providing data of a motor vehicle

    US20220070148A1

Cited By

  • Data cross-platform secure transmission method and system based on industrial internet

    CN121727870A