Computer data processing system and method based on artificial intelligence

By performing data unitization and selective privacy processing locally on the data holder, generating private data units carrying metadata, and performing aggregation processing through an aggregation coordinator, the problems of data privacy security and utility retention in a distributed computing environment are solved, and consistency of privacy policies across links and efficient data processing are achieved.

CN120654267AInactive Publication Date: 2025-09-16广州新华学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510803981.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In a distributed computing environment, how to efficiently preprocess and aggregate sensitive data while ensuring data privacy and security, and ensure that the preprocessing results are connected with the privacy requirements of subsequent data usage links, so as to achieve consistency and transferability of privacy policies across links.

Method used

Data unitization and selective privacy processing are performed locally on the data holder to generate private data units carrying metadata, and then aggregated through the aggregation coordinator. The aggregation coordinator does not access the original sensitive data and only relies on metadata for calculations. At the same time, version compatibility and privacy protection mechanisms are introduced.

Benefits of technology

It effectively prevents individual information leakage in a distributed environment, maximizes data utility, ensures consistency of privacy policies across links, and improves data processing efficiency and availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654267A_ABST
    Figure CN120654267A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer data processing, in particular to a computer data processing system and method based on artificial intelligence, and the method comprises the following steps: carrying out the data unitization processing of the original sensitive data of a data holder at the local of each data holder according to a first preset rule, and generating a data unit; according to a second preset rule, selective privacy processing is executed on the data unit, a privacy data unit carrying metadata is generated, and the metadata comprises associated information used for follow-up aggregation and privacy processing information representing a privacy processing mode; receiving, by an aggregation coordinator, a privatized data unit and its metadata from one or more of the data holders; through the cooperation, the effects that original data are not out of the local in a distributed environment, individual information leakage is effectively prevented, data utility is reserved to the maximum extent, and cross-link privacy strategy consistency is ensured are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer data processing, and in particular to a computer data processing system and method based on artificial intelligence. Background Art

[0002] In the scenario of computer-based processing of dispersed sensitive data, there is a challenge in ensuring the privacy and security of data throughout its entire lifecycle, including transmission, preprocessing, aggregation, distributed training, model updates, model inference and result output, storage, and destruction, while maintaining data availability, processing efficiency, and model performance. Currently, in a distributed computing environment, for sensitive data dispersedly stored in multiple data holders, how to design a preprocessing and aggregation coordination mechanism, while ensuring that the original data does not go beyond the local control scope of each data holder, can effectively prevent the direct leakage of individual original sensitive information, maximize the data utility for subsequent distributed joint analysis or statistical modeling, and ensure that the preprocessing results under this mechanism can be connected with the privacy requirements of subsequent data use links (such as model updates and result release) to achieve the overall privacy protection goal throughout the data processing process, is a technical problem that needs to be solved urgently.

[0003] Specifically, such scenarios usually have the following specific characteristics: physical isolation and logical association of data sources coexist, and the sensitive data of each data holder is physically dispersed and not directly shared, but business needs require logical association analysis or aggregation statistics of these dispersed data; mandatory localization and privacy pre-positioning of preprocessing operations. Due to data sensitivity and the constraint of "staying local", core privacy-enhancing preprocessing operations must be completed in the local environment before the data leaves the holder's control domain; the indirectness of the aggregation process and the dependence of the result utility. Data aggregation cannot be directly based on the original data, but relies on the intermediate form of data after local preprocessing; the effectiveness of the aggregation results is highly dependent on the degree of retention and conversion method of the original data information by the local preprocessing method; the consistency and transferability requirements of privacy policies across links. The privacy protection measures adopted in local preprocessing and their strength need to be perceived and compatible with subsequent distributed computing or data publishing links to ensure that the privacy protection effect is continuously effective and non-conflicting during the data flow process.

[0004] Therefore, how to perform privacy-enhancing preprocessing and aggregation of sensitive data efficiently and with minimal information loss in a distributed environment, while ensuring that the preprocessing results can be connected with the privacy requirements of subsequent data usage links, so as to achieve consistency and transferability of privacy policies across links, and ultimately achieve the overall privacy protection goal throughout the data processing process, is an important challenge currently faced.

[0005] In view of the above problems, the existing technology is in urgent need of improvement. Summary of the Invention

[0006] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a computer data processing system and method based on artificial intelligence.

[0007] In a first aspect, the present invention provides a computer data processing method based on artificial intelligence, applied to a distributed computing environment, the method comprising the following steps: Performing data unitization processing on the original sensitive data of each data holder locally according to a first preset rule to generate data units; performing selective privacy processing on the data unit according to a second preset rule to generate a privacy-protected data unit carrying metadata, wherein the metadata includes associated information for subsequent aggregation and privacy-protected information representing a privacy-protected manner, wherein the privacy-protected information is used to identify a degree of privacy protection; receiving, through the aggregation coordinator, the privacy-enhanced data units and metadata thereof from one or more data holders; The aggregation coordinator performs aggregation processing on the received privacy-enhanced data units according to the third preset rule and the associated information in the metadata to generate an aggregation result.

[0008] Data unitization refers to breaking down the original sensitive data of the data holder into smaller, easier-to-manage and process data units based on preset rules. This can be achieved by dividing the data based on data structure, semantic content or business logic.

[0009] Selective privacy processing refers to the application of one or more privacy protection technologies to data units according to preset rules to reduce the privacy risks of data units. It can be achieved by using technologies such as differential privacy, homomorphic encryption, secure multi-party computing, data desensitization, data generalization, and data suppression. Its main purpose is to maximize the availability of data while protecting privacy.

[0010] A privacy-enhanced data unit with metadata refers to a data unit that has undergone selective privacy processing and has metadata attached that describes its own characteristics and processing process. The metadata contains associated information for subsequent aggregation.

[0011] An aggregation coordinator is a logical or physical entity that is responsible for receiving privacy-protected data units from multiple data holders and performing aggregation operations. It can be an independent server, a computing cluster, or a node in a distributed system. Its main purpose is to centrally process decentralized privacy-protected data and generate aggregation results.

[0012] Preset rules refer to a predefined set of logic, parameters or algorithms used to guide data unitization, privacy processing and aggregation processing. They can exist in the form of configuration files, policy scripts or model parameters. Their main purpose is to ensure the controllability, consistency and standardization of the data processing process.

[0013] The aggregation coordinator does not access the original sensitive data when performing aggregation processing, which means that the aggregation coordinator only performs calculations based on the received privacy data units and the metadata they carry, and does not directly contact or process the original sensitive data stored locally by the data holder. This is mainly to ensure that the original data does not leave the local area, and fundamentally avoid the leakage of original sensitive information during transmission or aggregation.

[0014] In a second aspect, there is provided an artificial intelligence-based computer data processing system, applied in a distributed computing environment, characterized in that the system comprises: a primary processing module, performing data unitization processing on the original sensitive data of each data holder locally according to a first preset rule to generate data units; a secondary processing module, configured to perform selective privacy processing on the data unit according to a second preset rule, generating a privacy-protected data unit carrying metadata, wherein the metadata includes associated information for subsequent aggregation and privacy-protected information representing a privacy-protected manner, wherein the privacy-protected information is used to identify a degree of privacy protection; A receiving and coordinating module receives, through an aggregation coordinator, the privacy-enhanced data units and metadata thereof from one or more data holders; The aggregation processing module performs aggregation processing on the received privacy-enhanced data units according to the third preset rule and the associated information in the metadata through the aggregation coordinator to generate an aggregation result.

[0015] Compared with the prior art, the present invention has the following beneficial effects: By performing fine-grained data unitization and controllable selective privacy processing locally on the data holder's side, and appending metadata containing associated information and privacy processing information to the processed data, the aggregation coordinator can rely entirely on these metadata to accurately aggregate the privacy-protected data units without accessing the original sensitive data. At the same time, the privacy processing information in the metadata supports the subsequent data usage links to perceive and connect the privacy protection strategies of local applications, achieving the effect of keeping the original data from leaving the local area, effectively preventing individual information leakage, maximizing data utility, and ensuring consistency of privacy policies across links in a distributed environment. This solves the problem in existing technologies that it is impossible to efficiently perform distributed data processing while ensuring data privacy, and has the advantage of improving data processing efficiency and data availability while ensuring data privacy. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 Flowchart of the present invention.

[0017] Figure 2 It is a structural diagram of the present invention.

[0018] In the figure: 101, primary processing module; 102, secondary processing module; 103, receiving coordination module; 104, aggregation processing module. DETAILED DESCRIPTION

[0019] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limiting the present invention.

[0020] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0021] In the traditional existing scenario of computer-based processing of dispersed stored sensitive data, we are faced with the challenge of how to ensure the privacy security of data in all links of the entire life cycle, such as transmission, preprocessing, aggregation, distributed training, model update, model reasoning and result output, storage and destruction, while maintaining data availability, processing efficiency and model performance. At present, in a distributed computing environment, for sensitive data dispersedly stored in multiple data holders, how to design a preprocessing and aggregation coordination mechanism under the premise of ensuring that the original data does not go beyond the scope of their local control, which can effectively prevent the direct leakage of individual original sensitive information, and maximize the retention of data utility for subsequent distributed joint analysis or statistical modeling, and ensure that the preprocessing results under this mechanism can be connected with the privacy requirements of subsequent data use links, so as to achieve the overall privacy protection goal throughout the data processing process, is a technical problem that needs to be solved urgently. In this regard, this application proposes a computer data processing method based on artificial intelligence.

[0022] like Figure 1 The computer data processing method based on artificial intelligence is applied to a distributed computing environment, and the method includes the following steps: Performing data unitization processing on the original sensitive data of the data holder locally at each data holder according to a first preset rule to generate data units; Performing selective privacy processing on the data unit according to a second preset rule to generate a privacy-protected data unit carrying metadata, the metadata including associated information for subsequent aggregation and privacy processing information representing a privacy processing method, the privacy processing information being used to identify a degree of privacy protection; Receiving, through an aggregation coordinator, a privacy-preserving data unit and its metadata from one or more data holders; The aggregation coordinator performs aggregation processing on the received privacy-enhanced data units according to the third preset rule and the associated information in the metadata to generate an aggregation result, wherein the aggregation coordinator does not access the original sensitive data when performing the aggregation processing.

[0023] Among them, data unitization processing refers to decomposing the original sensitive data of the data holder into smaller, easier to manage and process data units according to preset rules. It can be achieved by dividing based on data structure, semantic content or business logic.

[0024] Selective privacy processing refers to the application of one or more privacy protection technologies to data units according to preset rules to reduce the privacy risks of data units. It can be achieved by using technologies such as differential privacy, homomorphic encryption, secure multi-party computing, data desensitization, data generalization, and data suppression. Its main purpose is to maximize the availability of data while protecting privacy.

[0025] A privacy-enhanced data unit with metadata refers to a data unit that has undergone selective privacy processing and has metadata attached that describes its own characteristics and processing process. The metadata contains associated information for subsequent aggregation.

[0026] An aggregation coordinator is a logical or physical entity that is responsible for receiving privacy-protected data units from multiple data holders and performing aggregation operations. It can be an independent server, a computing cluster, or a node in a distributed system. Its main purpose is to centrally process decentralized privacy-protected data and generate aggregation results.

[0027] Preset rules refer to a predefined set of logic, parameters or algorithms used to guide data unitization, privacy processing and aggregation processing. They can exist in the form of configuration files, policy scripts or model parameters. Their main purpose is to ensure the controllability, consistency and standardization of the data processing process.

[0028] The aggregation coordinator does not access the original sensitive data when performing aggregation processing, which means that the aggregation coordinator only performs calculations based on the received privacy data units and the metadata they carry, and does not directly contact or process the original sensitive data stored locally by the data holder. This is mainly to ensure that the original data does not leave the local area, and fundamentally avoid the leakage of original sensitive information during transmission or aggregation.

[0029] Specifically, the overall working principle of this solution is as follows: First, each data holder locally performs data unitization based on a first preset rule, breaking the original sensitive data into multiple data units. Next, these data units are selectively privacy-enhanced based on a second preset rule, applying different privacy protection techniques to different data units or privacy requirements, generating private data units. During this process, metadata is appended to each private data unit, recording contextual information for subsequent aggregation and identifying the privacy-enhancing approach employed. Subsequently, an aggregation coordinator receives private data units and their corresponding metadata from one or more data holders. Finally, the aggregation coordinator aggregates these private data units based on a third preset rule and the contextual information contained in the metadata of the received private data units, generating the final aggregation result. Throughout the aggregation process, the aggregation coordinator strictly adheres to the principle of not accessing the original sensitive data and performs calculations based solely on the private data and metadata. In this way, data holders locally protect sensitive information, while the aggregation coordinator performs global aggregation analysis based on this protected data and auxiliary information, achieving a balance between privacy protection and data utility.

[0030] A specific embodiment of this application constructs and implements a full-lifecycle sensitive data privacy protection solution covering key stages such as data preprocessing and transmission, distributed computing training, model updates, model inference, and output. Specifically, during the data preprocessing and transmission phase, each data holder first locally unitizes the original sensitive data. Then, selective privacy processing (e.g., combining encryption and desensitization techniques) is performed according to a second pre-set rule to generate private data units with detailed metadata (including contextual information, privacy processing method, protection level, and rule version identifier), thereby ensuring that the original data remains local and is securely transmitted. When preparing data for the distributed computing training phase, an aggregation coordinator receives and aggregates these private data units based on the metadata without accessing the original data. Simultaneously, it utilizes version compatibility checks, privacy impact assessments, and a dynamic privacy-utility balance adjustment mechanism (including complex alternative adjustment operation selection and chain effect strength analysis) to generate aggregated results that are safe for subsequent distributed training (e.g., federated learning). During the model update phase, the newly added data also follows the local preprocessing and centralized aggregation process described in this application, ensuring privacy during the update process. During the model inference and output phases, a policy basis is continuously provided for input data protection and output result desensitization. Ultimately, the localized, refined processing, metadata-driven aggregation and coordination, and consistent rule management and dynamic balancing mechanisms provided by this invention provide the core technical support and specific operational procedures for building a full-lifecycle integrated privacy protection framework, ensuring the coordination and consistency of privacy measures across all stages.

[0031] The present application further proposes that the steps of receiving, by the aggregation coordinator, the privacy-preserving data unit and its metadata from one or more data holders include: Setting version identifiers for the first preset rule and the second preset rule and writing metadata; Before performing the aggregation process, the aggregation coordinator obtains a version identifier of the first preset rule and a version identifier of the second preset rule; If the aggregation coordinator detects an inconsistency based on the obtained version identifier of the first preset rule or the obtained version identifier of the second preset rule; The aggregation coordinator processes the privacy-enhanced data units with inconsistent versions according to preset compatibility processing rules to generate compatible privacy-enhanced data units for subsequent aggregation processing.

[0032] Among them, the version identifier refers to a mark used to uniquely identify a specific version of the preset rule, which can be implemented by a digital sequence, hash value or timestamp. The compatibility processing rule refers to a predetermined logic or algorithm set used to process the privacy data units generated due to inconsistent preset rule versions, which can be implemented by data conversion functions, data standardization methods or data filtering strategies. A compatible privacy data unit refers to a data unit that can be correctly aggregated with other privacy data units after being processed by the compatibility processing rules. It can be a data unit obtained by adjusting, converting or reformatting the original privacy data unit.

[0033] This solution addresses the issue of inaccurate aggregation results caused by different data holders using different versions of preset rules for data processing in a distributed computing environment. By proposing a version compatibility mechanism, we address the issue of inaccurate aggregation results in distributed computing environments. First, by assigning version identifiers to the first and second preset rules, we implement version management for data processing rules. This allows each private data unit to clearly identify the rule version based on which it was generated. Second, the metadata of the private data unit generated locally by the data holder includes the version identifiers of the first and second preset rules used to generate the private data unit. This design enables the aggregation coordinator to obtain the rule version information used for each data unit, providing a basis for subsequent version compatibility testing. Then, before performing aggregation processing, the aggregation coordinator obtains the version identifiers of the first and second preset rules from the metadata of the received private data unit. Through centralized version information collection, the aggregation coordinator maintains a comprehensive understanding of the rule versions used by the participating data units. Finally, if the aggregation coordinator detects version inconsistencies, it processes the private data unit with the inconsistency according to the preset compatibility processing rules to generate a compatible private data unit for subsequent aggregation processing. This step is the core of this solution. Through compatibility processing, data units generated under different rule versions can be processed uniformly, ensuring the accuracy and reliability of the aggregated results. Building on the foundational distributed privacy-preserving data processing framework, this solution introduces rule version identification and compatibility processing mechanisms to address data compatibility issues arising from differences in rule version evolution or deployment in distributed environments. This combination makes the distributed processing framework more robust and usable in practical applications, capable of handling more complex real-world scenarios, ensuring the accuracy of aggregated results, and thus improving the reliability of the entire data processing process.

[0034] In some of the above-mentioned embodiments of the present application, a scheme is proposed for setting version identifiers for the first preset rule and the second preset rule, and performing compatibility processing when the version identifiers are inconsistent. This scheme can enable the aggregation coordinator to identify the difference in rule versions based on which the data units generated by different data holders are based, and perform preliminary processing when inconsistencies are detected. This can improve the flexibility and compatibility of the system in processing data generated from different versions of rules. However, in its implementation process, the puzzle component information is only generated based on the two images, and there is a lack of pixel-level processing of the images themselves. This may cause the puzzle components to lack fineness and diversity in visual presentation, and fail to fully utilize the detailed features of the two images, making the final puzzle effect less than ideal. How to ensure that the privacy-improved data units after compatibility processing still meet the minimum privacy protection requirements and avoid privacy leaks due to compatibility processing is a problem that needs to be solved.

[0035] This application further proposes that the aggregation coordinator processes the privacy-preserving data units with inconsistent versions according to a preset compatibility processing rule to generate compatible privacy-preserving data units, including the following steps: For each private data unit to be processed that has inconsistent versions, determine one or more initial privacy-preserving attribute values ​​for the private data unit based on the privacy processing information in its metadata that characterizes the privacy processing method adopted; Evaluating, based on predetermined impact assessment parameters related to the privacy-preserving attributes in the pre-set compatibility processing rules, the expected change in one or more initial privacy-preserving attribute values ​​when the compatibility processing rules are applied to the privacy-preserving data unit; Based on the initial privacy-preserving attribute value and the expected change, calculate one or more expected final privacy-preserving attribute values ​​of the compatible privacy-preserving data unit that will be formed after the privacy-preserving data unit is processed by the compatibility processing rule; Compare the expected final privacy protection attribute value with the preset minimum privacy protection requirement; If the minimum privacy protection requirements are met, the aggregation coordinator processes the privacy-protected data units with inconsistent versions according to the preset compatibility processing rules to generate compatible privacy-protected data units and records the expected final privacy-protected attribute value or the change information related to the initial privacy-protected attribute value; If the minimum privacy protection requirements are not met, the aggregation coordinator will execute preset adjustment measures, which include adjusting the compatibility processing rules or excluding the privacy-enhanced data units with inconsistent versions from the current processing.

[0036] This solution addresses the issue of inconsistencies in the privacy protection level of private data units, which may result in a reduction in privacy protection during compatibility processing. It proposes a mechanism that balances privacy protection during compatibility processing. First, for each private data unit to be processed, its initial privacy-preserving attribute value is determined using the privacy-preserving information in its metadata. This lays the foundation for evaluating the impact of the compatibility processing. Then, using the impact assessment parameters associated with the privacy-preserving attributes in the pre-set compatibility processing rules, the expected change in the initial privacy-preserving attribute value due to the compatibility processing is evaluated, thereby predicting the privacy protection level after the compatibility processing. By combining the initial privacy-preserving attribute value and the expected change, the expected final privacy-preserving attribute value of the compatible private data unit is calculated, enabling a quantitative assessment of the impact of the compatibility processing on privacy protection. By comparing the expected final privacy-preserving attribute value with the pre-set minimum privacy protection requirement, it is determined whether the compatibility processing will reduce the privacy protection level below an acceptable range. If the minimum privacy protection requirement is met, the compatibility processing is performed, and the expected final privacy-preserving attribute value or related change information is recorded in the metadata, allowing subsequent steps to understand the privacy protection level of the data. If the minimum privacy protection requirements are not met, adjustment measures are taken, including adjusting the compatibility processing rules or excluding the data unit, so as to avoid privacy leakage caused by compatibility processing. Through the above steps, this solution can ensure data compatibility while ensuring that the degree of privacy protection of the data is not lower than the preset minimum requirements, thereby achieving a balance between compatibility processing and privacy protection. Combined with the data unitization and privacy processing of the original sensitive data according to the rules at the data holder's local location, and the aggregation processing of the received privacy-protected data units according to the rules by the aggregation coordinator, this solution further enhances the robustness and security of the data processing process in a distributed computing environment without accessing the original sensitive data. In particular, when processing data from different rule versions, it can effectively deal with potential privacy leakage risks and ensure the consistency and effectiveness of privacy protection in the entire data processing chain.

[0037] This application further proposes that when, after adjusting the compatibility processing rules, the expected final privacy protection attribute value still does not meet the minimum privacy protection requirement, and if excluding the private data units with inconsistent versions from the current processing will cause the data utility of the aggregation result to fall below a preset utility threshold, the aggregation coordinator executes the preset adjustment measures, including the following steps: Obtaining a target privacy protection attribute reference value corresponding to the minimum privacy protection requirement and a target data utility reference value corresponding to a preset utility threshold; For privacy-enhanced data units with inconsistent versions, select an adjustment operation from a set of preset alternative adjustment operations, each of which corresponds to a different expected privacy protection result and expected data utility result; The selected adjustment operation is determined based on the degree of conformity between the expected final privacy-preserving attribute value of the privacy-preserving data unit after applying the selected adjustment operation and the target privacy-preserving attribute reference value, as well as the degree of conformity between the expected data utility of the processed data on the aggregation result and the target data utility reference value, so as to achieve a preset balance condition. Apply the determined adjustment operation to the privacy-enhanced data unit with the inconsistent version.

[0038] The minimum privacy protection requirement refers to the minimum security level or standard that a system or application must meet to protect individual privacy. It can be expressed using thresholds of privacy metrics such as the ε value in differential privacy, the k value in k-anonymity, and the l value in l-diversity. The preset utility threshold refers to the minimum data quality or amount of information that must be retained in the aggregation results to meet subsequent data usage requirements. It can be expressed using thresholds of data utility metrics such as data accuracy, model performance indicators (such as precision and recall), and statistical bias. The target privacy protection attribute reference value refers to a specific value or standard that directly corresponds to the minimum privacy protection requirement and is used to quantitatively evaluate the privacy protection effect after the adjustment operation. It can use the same measurement units and representation as the minimum privacy protection requirement. The target data utility reference value refers to a specific value or standard that directly corresponds to the preset utility threshold and is used to quantitatively evaluate the degree of data utility preservation after the adjustment operation. It can use the same measurement units and representation as the preset utility threshold. The preset alternative adjustment operations refer to a series of optional processing methods pre-set by the system for private data units with inconsistent versions. For example, they may include perturbations of different intensities, generalizations of different granularities, deletion of some features, or conversion to specific formats. Each operation aims to balance privacy protection and data utility in different ways. The expected privacy protection result refers to the privacy protection attribute value of the data unit predicted by the evaluation model or calculation method after applying a certain alternative adjustment operation to the private data unit. It can be measured by indicators such as differential privacy budget and anonymity. The expected data utility result refers to the contribution or impact of the data unit on the data utility of the final aggregation result predicted by the simulation aggregation or evaluation model after applying a certain alternative adjustment operation to the private data unit. It can be measured by indicators such as the amount of information loss and the impact on downstream task performance. The degree of compliance refers to the degree to which the expected final privacy-preserving attribute value meets or approaches the target privacy-preserving attribute reference value, as well as the expected data utility of the processed data on the aggregation result and the target data utility reference value, meeting or approaching the preset target. This can be calculated or determined using distance metrics, proportions, hierarchical classifications, or by satisfying specific inequality conditions. The preset balance condition refers to a decision rule or function used to comprehensively consider the degree of privacy-preserving compliance and the degree of data utility compliance and to select the optimal adjustment action accordingly. This can be constructed using weighted summation, multi-objective optimization functions, decision trees, or machine learning-based models.

[0039] This method is triggered when the aggregation coordinator processes privacy-enhanced data units with inconsistent versions. This occurs when previous attempts to adjust compatibility processing rules have failed to ensure that the expected final privacy-protection attribute values ​​meet the minimum requirements, and when an assessment reveals that simply excluding these data units would cause the data utility of the final aggregation result to fall below a preset threshold. At this point, the system initiates a refined adjustment process. First, the system obtains a target privacy-protection attribute reference value corresponding to the minimum privacy protection requirement, as well as a target data utility reference value corresponding to a preset utility threshold. These reference values ​​define the boundaries of the privacy protection and data utility targets that the system must strive to achieve under the current predicament. Next, for these data units with inconsistent versions, the system selects from a set of pre-defined candidate adjustment actions. These candidate actions are different processing strategies designed for these data units, each of which has been pre-evaluated based on the potential privacy protection and data utility outcomes it might bring. The system does not make random selections, but rather determines the optimal action based on a preset balance condition. This balance condition comprehensively considers two key factors: first, the degree of conformity between the expected final privacy-preserving attribute value of the privacy-preserving data unit and the target privacy-preserving attribute reference value after applying a certain adjustment operation; and second, the degree of conformity between the expected data utility of the data processed by the operation and the target data utility reference value for the final aggregation result. By simultaneously evaluating the degree of conformity of these two dimensions and weighing them under the balance condition, the system can identify the adjustment operation that can maximize privacy protection to approach the target while minimizing data utility loss to approach the target. Once the appropriate adjustment operation is determined, the aggregation coordinator applies it to the privacy-preserving data unit with inconsistent versions. This process allows the system to avoid simplistic and crude handling methods when faced with conflicts between privacy protection and data utility, and instead adopt a quantitative, balanced, and optimized strategy, thereby maximizing data availability and ensuring the quality of subsequent aggregation processing while meeting basic privacy requirements. This mechanism, which quantifies goals, evaluates multiple options, and makes trade-off decisions based on dual compliance, enables the system to find a better balance when dealing with complex conflicts caused by inconsistent data versions. It effectively solves the technical difficulties in specific scenarios where privacy protection is insufficient and data exclusion leads to low utility.

[0040] This application further proposes that when the compliance assessment includes at least two privacy protection attribute dimensions and at least two data utility dimensions, and there is a preset mutual influence relationship between the dimensions, the steps of specifically determining the selected adjustment operation include: Obtaining a preset set of influencing parameters, adjusting a calculation method for calculating the degree of compliance with the privacy protection attribute and the degree of compliance with the data utility in a preset balance condition, and obtaining an adjusted balance condition and an adjusted degree of compliance calculation method; For each candidate adjustment operation, based on its direct expected changes to each privacy-preserving attribute dimension and each data utility dimension, and combined with the obtained set of impact parameters, determine the comprehensive expected results of the candidate adjustment operation on each dimension. The comprehensive expected results include the expected final privacy-preserving attribute value after applying the candidate adjustment operation and the expected data utility of the processed data on the aggregation result; Based on the degree of conformity between the expected final privacy protection attribute value in the comprehensive expected result and the target privacy protection attribute reference value, as well as the degree of conformity between the expected data utility in the comprehensive expected result and the target data utility reference value, the adjusted conformity calculation method is applied, and the calculated conformity is applied to the adjusted balance condition to determine the selected adjustment operation.

[0041] This solution introduces a more sophisticated adjustment operation selection mechanism for handling private data units with inconsistent versions, when simple adjustment or exclusion strategies are insufficient to simultaneously meet privacy protection and data utility requirements, and when the evaluation involves multiple, interdependent dimensions. First, the system obtains a preset set of impact parameters that capture the complex interdependencies between different privacy and utility dimensions. For example, increasing the ε value of differential privacy may directly reduce privacy protection, but may also indirectly improve certain data utility metrics by affecting data availability. Conversely, increasing the degree of data desensitization may directly improve privacy protection but also negatively impact data utility, and this impact may propagate across different utility dimensions. Based on this understanding of these interplays, the system adjusts the balance conditions and compliance calculation methods used to evaluate candidate adjustment operations based on the obtained set of impact parameters. This adjustment can manifest itself as modifying the weights of each dimension in the balance condition or adjusting the sensitivity of each dimension in the compliance calculation function. For example, if a privacy dimension has a significant negative impact on a key utility dimension, its weight in the balancing condition can be reduced, or a more relaxed standard can be used when calculating its compliance, thereby giving greater consideration to utility in the trade-off. Next, for each feasible alternative adjustment action, the system not only considers its direct expected changes to each dimension but also uses a set of impact parameters to calculate how these direct changes propagate through mutual influence relationships, thereby determining the combined expected outcome of the action on all relevant dimensions. This combined expected outcome more comprehensively reflects the overall state after applying the action. Finally, the system applies an adjusted compliance calculation method to evaluate the degree of compliance of each alternative action's combined expected outcome with the target reference value and substitutes these compliance degrees into the adjusted balancing condition for calculation. By comparing the evaluation results of different alternative actions under the adjusted balancing condition, the system can select the adjustment action that best balances privacy protection and data utility goals in the context of the interplay of multiple dimensions. This approach makes the selection of adjustment operations more intelligent and optimized by explicitly modeling and utilizing the mutual influence relationships between dimensions, overcoming the suboptimal results that may be caused by simple strategies and ensuring that the aggregation results in complex scenarios can better meet the overall requirements.

[0042] This application further proposes obtaining a preset set of influencing parameters, adjusting a calculation method for calculating the degree of compliance with the privacy protection attribute and the degree of compliance with the data utility in a preset balance condition, and obtaining the adjusted balance condition and the adjusted degree of compliance calculation method, including the following steps: When the compliance assessment includes at least two privacy protection attribute dimensions and at least two data utility dimensions, and there is a preset mutual influence relationship between the dimensions, obtaining a preset set of influence parameters, and identifying key influence dimensions in the privacy protection attribute dimensions and the data utility dimensions based on the obtained set of influence parameters; Based on the information related to the identified key impact dimensions in the obtained impact parameter set, the weight parameters corresponding to the identified key impact dimensions in the preset balance conditions are adjusted to form an adjusted balance condition, and the sensitivity parameters corresponding to the identified key impact dimensions in the function used to calculate the degree of compliance with the privacy protection attributes and the degree of compliance with the data utility are adjusted to form an adjusted degree of compliance calculation method.

[0043] This solution achieves a balance between privacy protection and data utility by employing a strategy that identifies and adjusts key influencing dimensions. First, by analyzing the acquired set of influencing parameters, the system identifies which dimensions, within the privacy protection attribute dimension and the data utility dimension, have an impact on other dimensions or the overall balance state. These dimensions are identified as key influencing dimensions. Identifying these key dimensions allows subsequent adjustments to focus on factors that influence the overall outcome. Based on this, the solution adjusts the weight parameters corresponding to these key influencing dimensions in the preset balancing conditions based on information related to these key influencing dimensions in the influencing parameter set. By adjusting the weights of these key dimensions, the balancing conditions reflect the impact of changes in these dimensions on the overall balance, thereby guiding the selected adjustment operations toward optimizing the key dimensions. Furthermore, based on information related to these key influencing dimensions in the influencing parameter set, the solution adjusts the sensitivity parameters corresponding to these key influencing dimensions in the function used to calculate the degree of compliance. Adjusting the sensitivity parameters ensures that the degree of compliance calculation captures changes in the key dimensions, allowing evaluation of the performance of different adjustment operations on these key dimensions. This collaborative approach of identifying key dimensions and adjusting their weights and sensitivities can overcome the bias that may arise from adjusting all dimensions, achieve a balance between privacy protection and data utility, and improve data processing performance.

[0044] This application further proposes a calculation method for calculating the degree of compliance with privacy protection attributes and data utility in a preset balance condition based on a set of influencing parameters. Specifically, this adjustment method can simply assign weights or adjust sensitivities based on the attributes or direct influence strength of each dimension itself. For example, a static weight is preset for each dimension, or only the direct influence strength between dimensions is considered to adjust the calculation method. This can reflect the importance of dimensions to a certain extent. However, in its implementation process, only the attributes or direct influence of the dimensions themselves are considered, and the complex interactions and indirect influences between dimensions are not fully captured, especially the chain effects propagated through multi-step paths. This may lead to an inaccurate assessment of the true importance of the dimensions, and an inability to effectively identify those key dimensions that have a profound impact on the overall balance through chain reactions, making it difficult to achieve a more accurate and dynamic balance between privacy protection and data utility. In order to solve the above problems, this application proposes a method for identifying key influencing dimensions, including: Construct the influence path between the privacy protection attribute dimension and the data utility dimension; For each privacy protection attribute dimension and data utility dimension, evaluate its chain effect on other dimensions along the impact path; Based on the evaluation results of chain effects, the key impact dimensions in the privacy protection attribute dimension and data utility dimension are determined.

[0045] Specifically, first, we construct the impact pathways between privacy-preserving attribute dimensions and data utility dimensions. This is equivalent to building a model that clearly depicts how the dimensions are interconnected and influence each other. For example, reducing the strength of a privacy-preserving dimension may directly affect the data utility dimension while also indirectly affecting other privacy-preserving dimensions. This pathway provides a foundational structure for subsequent analysis. Second, for each privacy-preserving attribute dimension and data utility dimension, we assess the cascading impacts on other dimensions along the impact pathway. This involves considering not only the direct impact of one dimension on another but also how this impact propagates through intermediate dimensions, indirectly impacting even more distant dimensions. Assessing cascading impacts provides a more comprehensive measure of the importance and influence of a dimension within the entire dimensional network. Finally, based on the results of the cascading impact assessment, we identify the key influencing dimensions within the privacy-preserving attribute and data utility dimensions. By quantifying and comparing the strength of the cascading impacts across dimensions, we can identify those dimensions with larger and broader impacts on other dimensions. These dimensions are core elements in the overall balance between privacy protection and data utility. After identifying these key dimensions, the weights in the balance conditions and the sensitivity parameters in the compliance calculation method can be adjusted more specifically.

[0046] This application further proposes an evaluation result based on chain effects. The steps for determining the key impact dimensions in the privacy protection attribute dimension and the data utility dimension include: Quantify the chain effect strength of each privacy protection attribute dimension and data utility dimension along the impact path. The chain effect strength is based on the evaluation results of the chain effect. Compare the intensity of knock-on effects with pre-set criticality thresholds; If the chain effect strength exceeds the criticality threshold, the dimension is identified as a critical impact dimension.

[0047] Based on the constructed dimensional impact paths and the assessed chain effects, this solution quantifies the chain effect strength of each dimension and compares it against a pre-set standard to identify key dimensions with significant impact on the entire dimensional network. Quantifying chain effect strength converts complex interactions between dimensions into comparable numerical values, which better reflects the true influence of dimensions than simply considering direct effects or simple sums. By setting thresholds, the stringency of screening for key dimensions can be flexibly controlled according to actual needs. The importance of identified key dimensions is reflected in subsequent balance condition adjustments and compliance calculations, for example, by assigning higher weights or adjusting sensitivity parameters, making the final balance decisions and compliance assessments more accurate and effective. This solution leverages the results of the impact paths and chain effect assessments, further quantifying and filtering them to extract key information. This key information is then used to adjust parameters in subsequent calculations, allowing for a more precise consideration of the true importance of each dimension when balancing privacy protection and data utility. Ultimately, the optimal adjustment is selected to generate compatible privacy-compliant data units that meet requirements and are integrated into the overall data processing workflow. This process of progressive development, information refinement and utilization reflects the systematicness and effectiveness of the plan. This application further proposes to quantify the chain effect strength generated by each privacy protection attribute dimension and data utility dimension along the impact path. The chain effect strength is based on the evaluation results of the chain effect, including the following steps: Obtain a set of privacy-preserving attribute dimensions and data utility dimensions, denoted as dimension set V, where V contains n dimensions v_1, v_2, ..., v_n; Based on the evaluation results of the chain effect, construct an nxn direct impact matrix A, where the matrix element A[i][j] represents the direct impact intensity of dimension v_i on dimension v_j, and A[i][j] is greater than or equal to 0; Set an attenuation factor α, where 0<α<1. The attenuation factor α represents the proportion of the attenuation of each propagation step. Calculate the inverse matrix of the matrix (I-αA), where I is the nxn identity matrix. Calculate the total influence matrix T, where T = (I - αA)^(-1) - I; For each dimension v_i in the dimension set V, calculate its chain influence intensity CI_i, where CI_i is equal to the sum of all elements in the i-th row of the total influence matrix T, but excluding the element T[i][i], that is, CI_i = sum_{j=1, j!=i}^n T[i][j]; The calculated chain influence intensity CI_i is used as the chain influence intensity of dimension vi_i.

[0048] This scheme constructs a direct impact matrix between dimensions, introduces an attenuation factor to simulate the attenuation of the impact during the propagation process, and uses matrix inversion operations to calculate the total impact matrix, thereby quantifying the chain impact intensity of each dimension on other dimensions.

[0049] For example, consider a scenario involving three dimensions: differential privacy ε, k-anonymity k, and model accuracy. Through expert evaluation or data analysis, construct the direct influence matrix A. For example, A = [[0, 0.2, 0.1], [0.3, 0, 0.05], [0.1, 0.15, 0]]. Set the decay factor α to 0.5. Calculate (I - αA) = [[1, -0.1, -0.05], [-0.15, 1, -0.025], [-0.05, -0.075, 1]]. Calculate the inverse matrix of (I - αA). Then calculate the total influence matrix T = (I - αA)^-1 - I. For example, the calculation results are T = [[0.016, 0.109, 0.056], [0.165,0.013, 0.061], [0.061, 0.117, 0.008]] (example values, not exact calculation results). Calculate the chain effect strength of each dimension: CI_1 = T[1][2] + T[1][3] = 0.109 + 0.056 = 0.165; CI_2 = T[2][1] + T[2][3] = 0.165 + 0.061 = 0.226; CI_3 = T[3][1] + T[3][2] = 0.061 +0.117 = 0.178. These calculated CI values ​​are used as the chain effect strength of each dimension.

[0050] like Figure 2 The computer data processing system based on artificial intelligence is applied to a distributed computing environment, and the system includes: A primary processing module 101 performs data unitization processing on the original sensitive data of each data holder locally according to a first preset rule to generate data units; Secondary processing module 102 performs selective privacy processing on the data unit according to a second preset rule to generate a privacy-protected data unit carrying metadata. The metadata includes associated information for subsequent aggregation and privacy processing information representing the privacy processing method. The privacy processing information is used to identify the degree of privacy protection. A receiving and coordinating module 103 receives, through an aggregation coordinator, privacy-enhanced data units and metadata from one or more data holders; The aggregation processing module 104 performs aggregation processing on the received privacy-enhanced data units according to the third preset rule and the associated information in the metadata through the aggregation coordinator to generate an aggregation result.

[0051] Specifically, the primary processing module 101 and the secondary processing module 102 are deployed in the local computing environment of each data holder and are responsible for pre-processing the original sensitive data. The primary processing module 101 performs data unitization processing on the original sensitive data according to the first preset rule, breaking the original data into data units of smaller granularity. Then, the secondary processing module 102 performs selective privacy processing on these data units according to the second preset rule, applying different privacy protection technologies according to different data unit characteristics or privacy requirements, thereby protecting data privacy while maximizing the availability of data. During this process, the secondary processing module 102 also generates corresponding metadata for each processed data unit. These metadata contain associated information for subsequent aggregation and privacy processing information that characterizes the privacy processing method adopted. This information is crucial for the aggregation coordination module to correctly understand and process the privacy data unit. After processing is completed, the secondary processing module 102 sends the generated privacy data unit and its metadata to the receiving coordination module 103. The receiving coordination module 103 receives the privacy-protected data units and their metadata from one or more data holders, and the aggregation processing module 104 performs aggregation processing on the received privacy-protected data units according to the third preset rule and the associated information in the metadata. Importantly, when performing aggregation processing, the aggregation processing module 104 only operates based on the privacy-protected data units and their metadata, and does not directly access or process the original sensitive data, which further ensures the security of the original data from the system architecture level. In this way, the system separates the local pre-processing of sensitive data from the central or distributed aggregation coordination function, allowing data holders to contribute data value without exposing the original data, while the aggregation processing module 104 can perform effective calculations on the received privacy-protected data, thereby achieving a balance between privacy protection and data utility in a distributed environment.

[0052] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and description only describe the principles of the present invention. Various changes and improvements are possible without departing from the spirit and scope of the present invention, and such changes and improvements fall within the scope of the invention as claimed.

Claims

1. A computer data processing method based on artificial intelligence, applied to a distributed computing environment, characterized in that: The method comprises the following steps: Performing data unitization processing on the original sensitive data of each data holder locally according to a first preset rule to generate data units; performing selective privacy processing on the data unit according to a second preset rule to generate a privacy-protected data unit carrying metadata, wherein the metadata includes associated information for subsequent aggregation and privacy-protected information representing a privacy-protected manner, wherein the privacy-protected information is used to identify a degree of privacy protection; receiving, through the aggregation coordinator, the privacy-enhanced data units and metadata thereof from one or more data holders; The aggregation coordinator performs aggregation processing on the received privacy-enhanced data units according to the third preset rule and the associated information in the metadata to generate an aggregation result.

2. The computer data processing method based on artificial intelligence according to claim 1, characterized in that: The step of receiving, by the aggregation coordinator, the privacy-preserving data unit and its metadata from one or more data holders comprises: Setting version identifiers for the first preset rule and the second preset rule and writing metadata; The aggregation coordinator obtains a version identifier of the first preset rule and a version identifier of the second preset rule before performing the aggregation process; If the aggregation coordinator detects an inconsistency based on the obtained version identifier of the first preset rule or the obtained version identifier of the second preset rule; The aggregation coordinator processes the privacy-enhanced data unit with inconsistent versions according to a preset compatibility processing rule to generate a compatible privacy-enhanced data unit.

3. The computer data processing method based on artificial intelligence according to claim 2, characterized in that: The aggregation coordinator processes the privacy-enhanced data unit with inconsistent versions according to a preset compatibility processing rule to generate a compatible privacy-enhanced data unit, including the following steps: Determining one or more initial privacy protection attribute values ​​of the private data unit based on the privacy processing information in the metadata; evaluating, based on predetermined impact assessment parameters related to the privacy-preserving attributes in a preset compatibility processing rule, an expected change in the one or more initial privacy-preserving attribute values ​​when the compatibility processing rule is applied to the privacy-preserving data unit; Based on the initial privacy protection attribute value and the expected change amount, calculate one or more expected final privacy protection attribute values ​​of the compatible privacy protection data unit that will be formed after the privacy-protected data unit is processed by the compatibility processing rule; Comparing the expected final privacy protection attribute value with a preset minimum privacy protection requirement; If the minimum privacy protection requirement is met, the aggregation coordinator processes the privacy-protected data unit with inconsistent versions according to the preset compatibility processing rules to generate the compatible privacy-protected data unit, and records the expected final privacy-protected attribute value or change information related to the initial privacy-protected attribute value; If the minimum privacy protection requirement is not met, the aggregation coordinator executes a preset adjustment measure, which includes adjusting the compatibility processing rule or excluding the privacy-enhanced data unit with inconsistent versions from the current processing.

4. The computer data processing method based on artificial intelligence according to claim 3, characterized in that: When, after adjusting the compatibility processing rules, the expected final privacy protection attribute value still does not meet the minimum privacy protection requirement, and if excluding the private data unit with inconsistent versions from the current processing will cause the data utility of the aggregation result to be lower than a preset utility threshold, the aggregation coordinator executes the preset adjustment measures, including the following steps: Obtaining a target privacy protection attribute reference value and a target data utility reference value; Selecting an adjustment operation from a set of preset candidate adjustment operations, each of the candidate adjustment operations corresponding to a different expected privacy protection result and expected data utility result; The selected adjustment operation is determined based on the degree of conformity between the expected final privacy-preserving attribute value of the privacy-preserving data unit after applying the selected adjustment operation and the target privacy-preserving attribute reference value, and the degree of conformity between the expected data utility of the processed data on the aggregation result and the target data utility reference value, so as to achieve a preset balance condition. Apply the determined adjustment operation to the privacy-enhanced data unit with the inconsistent version.

5. The computer data processing method based on artificial intelligence according to claim 4, characterized in that: When the compliance evaluation includes at least two privacy protection attribute dimensions and at least two data utility dimensions, and there is a preset mutual influence relationship between the dimensions, the step of specifically determining the selected adjustment operation includes: Obtaining a preset set of influencing parameters, adjusting a calculation method for calculating the degree of compliance with the privacy protection attribute and the degree of compliance with the data utility in the preset balance condition, and obtaining an adjusted balance condition and an adjusted degree of compliance calculation method; For each candidate adjustment operation, based on the direct expected changes to each privacy-preserving attribute dimension and each data utility dimension, and in combination with the acquired set of influencing parameters, determine the comprehensive expected results of the candidate adjustment operation on each dimension. The comprehensive expected results include the expected final privacy-preserving attribute value after applying the candidate adjustment operation and the expected data utility of the processed data on the aggregation result. Based on the degree of conformity between the expected final privacy protection attribute value in the comprehensive expected result and the target privacy protection attribute reference value, as well as the degree of conformity between the expected data utility in the comprehensive expected result and the target data utility reference value, the adjusted conformity calculation method is applied for calculation, and the calculated conformity is applied to the adjusted balance condition to determine the selected adjustment operation.

6. The computer data processing method based on artificial intelligence according to claim 5, characterized in that: The steps of obtaining a preset set of influencing parameters, adjusting a calculation method for calculating the degree of compliance with the privacy protection attribute and the degree of compliance with the data utility in the preset balance condition, and obtaining the adjusted balance condition and the adjusted degree of compliance calculation method include: Identifying key influencing dimensions in the privacy protection attribute dimension and the data utility dimension based on the acquired influencing parameter set; Based on the information related to the identified key impact dimension in the obtained impact parameter set, the weight parameter corresponding to the identified key impact dimension in the preset balance condition is adjusted to form an adjusted balance condition, and the sensitivity parameter corresponding to the identified key impact dimension in the function used to calculate the compliance degree of the privacy protection attribute and the compliance degree of the data utility is adjusted to form an adjusted compliance calculation method.

7. The computer data processing method based on artificial intelligence according to claim 6, characterized in that: The step of identifying key influencing dimensions in the privacy protection attribute dimension and the data utility dimension based on the acquired influencing parameter set includes: Constructing an influence path between the privacy protection attribute dimension and the data utility dimension; For each of the privacy protection attribute dimensions and the data utility dimension, evaluating the chain effects on other dimensions along the impact path; Based on the evaluation results of the chain effects, key impact dimensions in the privacy protection attribute dimension and the data utility dimension are determined.

8. The computer data processing method based on artificial intelligence according to claim 7, characterized in that: The step of determining the key impact dimension in the privacy protection attribute dimension and the data utility dimension based on the evaluation result of the chain effect includes: quantifying the chain impact strength generated by each of the privacy protection attribute dimensions and the data utility dimension along the impact path, wherein the chain impact strength is based on an evaluation result of the chain impact; Comparing the chain reaction intensity with a preset criticality threshold; If the chain impact strength exceeds the criticality threshold, the dimension is determined as the critical impact dimension.

9. The computer data processing method based on artificial intelligence according to claim 8, characterized in that: The step of quantifying the chain impact strength generated by each of the privacy protection attribute dimensions and the data utility dimension along the impact path, wherein the chain impact strength is based on the evaluation result of the chain impact, comprises: Obtain a set of the privacy protection attribute dimension and the data utility dimension, denoted as dimension set V, where V contains n dimensions v_1, v_2, ..., v_n; Based on the evaluation results of the chain effect, construct an nxn direct impact matrix A, where the matrix elements A[i][j] represent the direct impact strength of dimension v_i on dimension v_j, and A[i][j] is greater than or equal to 0; Set an attenuation factor α, where 0<α<1, and the attenuation factor α represents the proportion of the attenuation of each propagation step, and calculate the inverse matrix of the matrix (I-αA), where I is the nxn identity matrix; Calculate the total influence matrix T, where T = (I - αA)^(-1) - I; For each dimension v_i in the dimension set V, calculate its chain influence intensity CI_i, where CI_i is equal to the sum of all elements in the i-th row of the total influence matrix T, but excluding the element T[i][i], that is, CI_i = sum_{j=1, j!=i}^n T[i][j]; The calculated chain influence intensity CI_i is used as the chain influence intensity of the dimension vi_i.

10. An artificial intelligence-based computer data processing system, using an artificial intelligence-based computer data processing method according to any one of claims 1 to 9, characterized in that: The system includes: a primary processing module, performing data unitization processing on the original sensitive data of each data holder locally according to a first preset rule to generate data units; a secondary processing module, configured to perform selective privacy processing on the data unit according to a second preset rule, generating a privacy-protected data unit carrying metadata, wherein the metadata includes associated information for subsequent aggregation and privacy-protected information representing a privacy-protected manner, wherein the privacy-protected information is used to identify a degree of privacy protection; A receiving and coordinating module receives, through an aggregation coordinator, the privacy-enhanced data units and metadata thereof from one or more data holders; The aggregation processing module performs aggregation processing on the received privacy-enhanced data units according to the third preset rule and the associated information in the metadata through the aggregation coordinator to generate an aggregation result.