Cross-mechanism privacy protection data system and method based on deep learning

By using a deep learning-based cross-institutional privacy-preserving data system, the problems of low data transmission efficiency and insufficient privacy protection in cross-institutional data collaboration are solved, achieving secure and efficient data transmission and accurate data classification.

CN120956482AInactive Publication Date: 2025-11-14GUIZHOU JINYIJING INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511145817.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies for cross-organizational data collaboration suffer from problems such as large data transmission volumes, high bandwidth requirements, heavy computational pressure, and susceptibility to processing delays or system crashes, while also failing to provide effective data privacy protection.

Method used

A deep learning-based cross-institutional privacy-preserving data system is adopted, including modules for data acquisition, classification, compression and encryption, dynamic privacy budget allocation, and cross-institutional data transmission. The data classification module determines the security level of the data, a multi-objective optimization model selects an appropriate compression and encryption combination scheme, and an intelligent compression and encryption module performs encryption and compression. An adversarial transfer model is used for feature-aligned data transmission.

Benefits of technology

It enables secure and efficient data transmission across institutions, improves the accuracy of data classification and privacy protection, reduces the consumption of computing resources, and enhances data transmission efficiency and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956482A_ABST
    Figure CN120956482A_ABST
Patent Text Reader

Abstract

The invention relates to the general technical field of artificial intelligence, in particular to a cross-mechanism privacy protection data system and method based on deep learning, and the system comprises a data collection module which collects mechanism data and characterization data; the data grading module is used for grading the mechanism data; the data scheme selection module is used for selecting a data compression and encryption combination scheme; the dynamic privacy budget distribution module is used for calculating a budget distribution proportion; the intelligent compression and encryption module is used for encrypting and compressing the mechanism data; and the cross-mechanism data transmission module is used for transmitting the mechanism data according to the feature alignment data. According to the method, deep learning is carried out on the mechanism data, and data grading and compression encryption scheme selection are carried out on the mechanism data, so that the security requirement of cross-mechanism privacy protection data is met, meanwhile, a data entity with a smaller size is generated, and the calculation pressure of data cross-mechanism is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of general artificial intelligence technology, and in particular to a cross-institutional privacy-preserving data system and method based on deep learning. Background Technology

[0002] Cross-organizational data collaboration is an innovative model that enables multiple independent organizations to securely share data resources and conduct joint analysis without disclosing the original data. Cross-organizational data collaboration typically involves the transmission and interaction of large amounts of data, and the data processing methods used in existing technologies during communication lead to high communication costs. On the one hand, to protect data privacy, transmitted data is often fully encrypted and undergoes complex privacy processing, significantly increasing data volume. On the other hand, in collaborative scenarios involving multiple parties, frequent information exchange and verification processes further exacerbate the communication burden. Traditional cross-organizational data processing systems typically centralize data for analysis and processing at a central node. This centralized architecture faces enormous computational pressure when dealing with large-scale cross-organizational data, easily leading to processing delays or even system crashes. Moreover, the data transmission to the central node not only faces the risk of privacy leaks but also consumes significant time and network resources.

[0003] Chinese Patent Publication No. CN111783803B discloses an image processing method and apparatus for achieving privacy protection, belonging to the field of computer technology. The technical solution of this invention is: to provide an image processing method for achieving privacy protection, employing a technical concept of desensitizing and compressing privacy images at the terminal, and performing image recognition on the server based on the desensitized and compressed data. The method extracts the frequency domain features of the privacy image at the terminal, and preprocesses and filters the frequency domain feature maps to adaptively obtain feature maps more conducive to image recognition. Compared with simple image compression, this improves the compression rate during image transmission, effectively protects data privacy, reduces data processing volume, and thus improves the effectiveness of image processing. However, this solution is not suitable for cross-institutional data processing systems and cannot achieve cross-institutional data collaboration for data transmission and interaction. Summary of the Invention

[0004] To address this, the present invention provides a cross-institutional privacy-preserving data system and method based on deep learning, which overcomes the problems of large data volume, high bandwidth requirements, heavy communication burden, high computational pressure, and easy processing delays or even system crashes when dealing with large-scale cross-institutional data in the prior art.

[0005] To achieve the above objectives, in one aspect, the present invention provides a cross-institutional privacy-preserving data system based on deep learning, comprising: The data acquisition module is used to acquire institutional data and characterization data; The data grading module is used to analyze the data grade of the organization's data, calibrate and adjust the data grade of the organization's data based on the number of abnormal accesses in the characterization data, and adjust the adjustment process of the data grade of the organization's data based on the metadata perfection index calculated from the characterization data. The data scheme selection module is used to analyze the data compression and encryption combination scheme according to the data level, select the data compression and encryption combination scheme using a multi-objective optimization model based on the comprehensive evaluation function value calculated from the characterization data, send the data compression and encryption combination scheme to the intelligent compression and encryption module, modify the data compression and encryption combination scheme based on the data compression ratio calculated from the characterization data, and correct the modification process of the data compression and encryption combination scheme based on the CPU utilization in the characterization data. The dynamic privacy budget allocation module is used to calculate the budget allocation ratio according to the utility quantification method, obtain the budget allocation ratio, and send the budget allocation ratio to the intelligent compression and encryption module. It is also used to calibrate the budget allocation ratio according to the data encryption time in the characterization data, and to adjust the calibration process of the budget allocation ratio according to the data usage scenario security requirement index. The intelligent compression and encryption module is used to encrypt and compress the organization's data according to the data compression and encryption combination scheme and the budget allocation ratio to obtain encrypted and compressed data. The cross-agency data transmission module is used to output feature-aligned data based on the encrypted compressed data using an adversarial migration model, and then transmit the feature-aligned data.

[0006] Further, the data classification module acquires data sensitivity and data usage index, and calculates the data security requirement index Sq based on data sensitivity Sm, data usage index Sy, first weighting coefficient α1, and second weighting coefficient α2, setting Sq = α1 × Sm + α2 × Sy to obtain the data security requirement index Sq. The data security requirement index Sq is then compared with a preset data security requirement index range Sq0. Based on the comparison result, the data classification is determined, and the data level of the organization's data is output based on the determination result. When Sq < Sq0min, the data classification module determines that the data classification is public level data and outputs the public level data as the data level of the organization's data; When Sq0min≤Sq≤Sq0max, the data classification module determines the data classification status as restricted data and outputs the restricted data as the data level of the organization data. When 1 > Sq > Sq0max, the data classification module determines that the data classification is sensitive data and outputs the sensitive data as the data level of the organization data. When Sq=1, the data classification module determines that the data is classified as top secret and outputs the top secret data as the data level of the organization's data.

[0007] Furthermore, the data grading module compares the number of abnormal data accesses Sf in the characterizing data with a preset number of abnormal data accesses Sf0, judges the abnormal data access situation based on the comparison result, and calibrates the data grade of the institutional data based on the judgment result, wherein: When Sf≤Sf0, the data classification module determines the abnormal data access situation as a normal situation and does not calibrate the data level; When Sf > Sf0, the data classification module determines the abnormal data access as an abnormal situation, calibrates the data level, and calibrates the preset data security requirement index Sq0 by using the abnormal access coefficient α. Let e ​​be the base of the natural logarithm. The calibrated preset data security requirement index interval Sq0` is obtained. Sq0` is set to α × Sq0. The preset data security requirement index interval Sq0 is replaced with the calibrated preset data security requirement index interval Sq0`. The data security requirement index Sq is then re-compared with the calibrated preset data security requirement index interval Sq0`. The data classification module also adjusts the data level of the institutional data based on abnormal data access, wherein: When the abnormal data access situation is normal, the preset data security requirement index Sq0 is adjusted by the cold start coefficient I. Sq01 is set to I×Sq0 to obtain the adjusted preset data security requirement index range Sq01. The preset data security requirement index range Sq0 is replaced with the adjusted preset data security requirement index range Sq01, and the data security requirement index Sq is compared with the adjusted preset data security requirement index range Sq01 again. When the abnormal data access situation is an abnormal situation, the preset data security requirement index Sq0` is adjusted by the cold start coefficient I, and Sq02 is set to I×Sq0` to obtain the adjusted preset data security requirement index range Sq02. The preset data security requirement index range Sq0 is replaced with the adjusted preset data security requirement index range Sq02, and the data security requirement index Sq is compared with the adjusted preset data security requirement index range Sq02 again.

[0008] Furthermore, the data grading module calculates the metadata completeness index Fp based on the metadata integrity index Sw, metadata consistency index Ss, metadata standardization index Sk, metadata update timeliness index Sj, and metadata access convenience index Sp in the data. Let Fp = A3×Sw + A4×Ss + A5×Sk + A6×Sj + A7×Sp, to obtain the metadata completeness index Fp. The metadata completeness index Fp is then compared with a preset metadata completeness index Fp0. Based on the comparison result, the metadata completeness is judged, and the data level adjustment process for the institutional data is adjusted according to the judgment result. Wherein: When Fp > Fp0, the data grading module determines that the metadata completeness is normal and does not adjust the data grading process. When Fp≤Fp0, the data grading module determines that the metadata completeness is abnormal and adjusts the data grading process. The cold start coefficient l is adjusted by using the preset metadata completeness index Fp0 and the metadata completeness index Fp. The cold start coefficient l` is set to (Fp0 / Fp)×l to obtain the adjusted cold start coefficient l`. The cold start coefficient l is replaced with the adjusted cold start coefficient l`, and the preset data security requirement index range is readjusted according to the adjusted cold start coefficient l`. The data grading is then re-analyzed according to the adjusted preset data security requirement index range.

[0009] Furthermore, the data scheme selection module analyzes the data compression and encryption combination scheme of the institutional data according to the data level to obtain the data compression and encryption combination scheme of the institutional data, wherein: When the data is classified as top secret, the data scheme selection module uses the Zstandard compression algorithm to compress the data of the organization to obtain compressed data, and uses the AES-256-GCM encryption algorithm to encrypt the compressed data to obtain encrypted top secret data. When the data level is sensitive data, the data scheme selection module uses the Brotli compression algorithm to compress the data to obtain compressed data, and uses the ChaCha20-Poly1305 encryption algorithm to encrypt the compressed data to obtain encrypted sensitive data. When the data level is restricted, the data scheme selection module uses the LZ4 compression algorithm to compress the data to obtain compressed data, and uses the AES-128-CBC encryption algorithm to encrypt the compressed data to obtain encrypted restricted data. When the data level is public, the data scheme selection module uses the gzip compression algorithm to compress the data to obtain compressed public-level data. The data scheme selection module constructs a multi-objective optimization model according to a multi-objective optimization model construction method, which includes: Step B01: Calculate the comprehensive evaluation function value F based on security S, computational resource consumption C, transmission efficiency T, eighth weight coefficient ws, ninth weight coefficient wc, and tenth weight coefficient wt. Set F = ws × S + wc × (1 - C) + wt × T to obtain the comprehensive evaluation function value. Step B02: Set constraints to optimize the comprehensive evaluation function value to obtain a multi-objective optimization model; The data scheme selection module inputs security, computational resource consumption, transmission efficiency, eighth weight coefficient, ninth weight coefficient, and tenth weight coefficient into a multi-objective optimization model, outputs a comprehensive evaluation function value, and compares the comprehensive evaluation function value F with a preset comprehensive evaluation function F0. Based on the comparison result, the module determines the state of the comprehensive evaluation function value and selects a data compression and encryption combination scheme for the institutional data based on the determination result. When F≥F0, the data scheme selection module determines that the comprehensive evaluation function value is in a normal state, selects the data compression and encryption combination scheme, and sends the data compression and encryption combination scheme to the intelligent compression and encryption module. When F < F0, the data scheme selection module determines that the comprehensive evaluation function value is in an abnormal state and does not select the data compression and encryption combination scheme.

[0010] Further, the data scheme selection module calculates the data compression ratio Ys based on the original data size Dd and the compressed data size Yd in the characterization data, sets Ys = Dd / Yd, obtains the data compression ratio Ys, compares the data compression ratio Ys with the preset data compression ratio Ys0, judges the data compression ratio based on the comparison result, and corrects the data compression and encryption combination scheme based on the judgment result, wherein: When Ys≥Ys0, the data scheme selection module determines that the data compression ratio is normal and does not modify the data compression and encryption combination scheme. When Ys < Ys0, the data scheme selection module determines that the data compression ratio is abnormal, corrects the data compression and encryption combination scheme, calibrates the eighth weight coefficient ws using the compression coefficient K, and sets... Let e ​​be the base of the natural logarithm. The calibrated eighth weight coefficient ws` is obtained. Set ws` = K × ws, wc remains unchanged, and wt` = (1 - K) × wt. The calibrated comprehensive evaluation function value F1 is then obtained. The calibrated comprehensive evaluation function value F1 is then compared again with the preset comprehensive evaluation function F0. The data scheme selection module compares the CPU utilization rate Lc in the characterizing data with the preset CPU utilization rate Lc0, judges the CPU utilization based on the comparison result, and corrects the data compression and encryption combination scheme modification process and the data level adjustment process based on the judgment result, wherein: When Lc≤Lc0, the data scheme selection module determines that the CPU utilization is normal and does not correct the process of data compression and encryption combination scheme modification and data level adjustment. When Lc > Lc0, the data scheme selection module determines that the CPU utilization is abnormal, corrects the process of modifying the data compression and encryption combination scheme and adjusting the data level, and corrects the preset compression ratio Ys0 by using the coefficient g. Let e ​​be the base of the natural logarithm. The corrected preset data compression ratio Ys0` is obtained. Ys0` is set to Ys0 / (0.47+g). The preset data compression ratio Ys0 is replaced with the corrected preset data compression ratio Ys0`. The data compression ratio Ys is then compared again with the corrected preset data compression ratio Ys0`. The data scheme selection module corrects the ninth weight coefficient wc using coefficient g, setting wc` = wc × g, to obtain the corrected ninth weight coefficient wc`. The data scheme selection module corrects the preset metadata perfection index Fp0 using coefficient g, setting Fp0` = g × Fp0, to obtain the corrected preset metadata perfection index Fp0`. The data classification adjustment process is corrected based on the corrected preset metadata perfection index.

[0011] Furthermore, the dynamic privacy budget allocation module calculates the budget allocation ratio for the institutional data using a utility quantification method, wherein the utility quantification method includes: Step C01, based on the compression ratio Ys, the maximum compression ratio Ysmax, and the initial privacy budget. Encryption strength value Perform calculations and set Obtain the privacy consumption value ; Step C02, based on the encryption key length K, the maximum security key length Kmax, the minimum security key length Kmin, and the total privacy budget. Privacy consumption value Perform calculations and set To obtain the encryption strength value ; Step C03, set the privacy consumption value Encryption strength value Initial privacy budget Total privacy budget The compression utility coefficient g1, the encryption utility coefficient g2, and the Lagrange multiplier W are used to calculate the budget allocation ratio using the Lagrange function, thus obtaining the budget allocation ratio. The dynamic privacy budget allocation module sends the budget allocation ratio to the intelligent compression and encryption module. The dynamic privacy budget allocation module compares the data encryption time Th with the preset data encryption time Th0, judges the data encryption time based on the comparison result, and adjusts the budget allocation ratio based on the judgment result. Wherein: When Th≤Th0, the dynamic privacy budget allocation module determines that the data encryption time consumption is normal and does not adjust the budget allocation ratio; When Th > Th0, the dynamic privacy budget allocation module determines that the data encryption time consumption is an abnormal situation and adjusts the budget allocation ratio. This adjustment is made using a time consumption coefficient h. The adjusted encryption strength value is obtained. The budget allocation ratio is optimized based on the adjusted encryption strength value. The dynamic privacy budget allocation module compares the data usage scenario security requirement index As with the preset data usage scenario security requirement index As0, judges the security status of the data usage scenario based on the comparison result, and calibrates the adjustment process of the budget allocation ratio, the data level, and the correction process of the data compression and encryption combination scheme based on the judgment result. When As≤As0, the dynamic privacy budget allocation module determines that the data usage scenario security is normal and does not calibrate the process of adjusting the budget allocation ratio, data level, and data compression and encryption combination scheme correction. When As > As0, the dynamic privacy budget allocation module determines that the data usage scenario security situation is abnormal, calibrates the adjustment process of the budget allocation ratio, the data level, and the data compression and encryption combination scheme correction process, and calibrates the preset data encryption time Th0 through the security factor v. Let e ​​be the base of the natural logarithm. The calibrated preset data encryption time Th0` is obtained. Th0` is set to v × Th0. The preset data encryption time Th0 is replaced with the calibrated preset data encryption time Th0`. The data encryption time Th is then compared again with the calibrated preset data encryption time Th0`. The dynamic privacy budget allocation module calibrates the preset data security requirement index interval Sq0 using the security coefficient v, obtaining the calibrated preset data security requirement index interval Sq0`. Sq0` is set to v × Sq0. The preset data security requirement index interval Sq0 is replaced with the calibrated preset data security requirement index interval Sq0`. A preset data security requirement index range Sq0` is established, and the data security requirement index Sq is re-compared with the calibrated preset data security requirement index range Sq0`. The dynamic privacy budget allocation module calibrates the preset data compression ratio Ys0 through the security coefficient v to obtain the calibrated preset data compression ratio Ys0`. Ys0` is set to Ys0 / (0.05+v). When Ys0` is greater than 1, it is set to equal 1. The preset data compression ratio Ys0 is replaced with the calibrated preset data compression ratio Ys0`, and the data compression ratio Ys is re-compared with the calibrated preset data compression ratio Ys0`.

[0012] Furthermore, the intelligent compression and encryption module encrypts and compresses the organization data according to the data compression and encryption combination scheme and the budget allocation ratio, obtains the data compression and encryption combination scheme and encryption strength value according to the data level, calculates the privacy consumption value according to the budget allocation ratio, and compresses and encrypts the organization data using the encryption strength value, the privacy consumption value, and the data compression and encryption combination scheme to obtain encrypted compressed data.

[0013] Furthermore, the cross-organizational data transmission module constructs the adversarial transfer model according to the adversarial transfer model construction method to obtain the adversarial transfer model, inputs the encrypted compressed data into the adversarial transfer model, outputs feature-aligned data, and transmits the feature-aligned data.

[0014] On the other hand, the present invention also provides a method for a cross-institutional privacy-preserving data system based on deep learning, comprising: Step S1: Collect institutional data and characterization data; Step S2 involves analyzing the data level of the organization data, calibrating and adjusting the data level of the organization data based on the number of abnormal accesses in the characterization data, and adjusting the adjustment process of the data level of the organization data based on the metadata perfection index calculated from the characterization data. Step S3: Analyze the data compression and encryption combination scheme according to the data level, select the data compression and encryption combination scheme using a multi-objective optimization model based on the comprehensive evaluation function value calculated from the characterization data, and send the data compression and encryption combination scheme to the intelligent compression and encryption module. Also, correct the data compression and encryption combination scheme based on the data compression ratio calculated from the characterization data, and correct the correction process of the data compression and encryption combination scheme based on the CPU utilization in the characterization data. Step S4: Calculate the budget allocation ratio according to the utility quantification method to obtain the budget allocation ratio, and send the budget allocation ratio to the intelligent compression and encryption module. Also, calibrate the budget allocation ratio according to the data encryption time in the characterization data, and adjust the calibration process of the budget allocation ratio according to the data usage scenario security requirement index. Step S5: Encrypt and compress the organization data according to the data compression and encryption combination scheme and the budget allocation ratio to obtain encrypted and compressed data; Step S6: Using the adversarial migration model, output the feature alignment data based on the encrypted compressed data, and then transmit the feature alignment data.

[0015] Compared with existing technologies, the beneficial effects of this invention are as follows: the system collects institutional data and characterization data through a data acquisition module, enabling subsequent encryption and compression of the data for secure and efficient transmission; the data classification module determines the data classification status, categorizing institutional data into different security levels, making data transmission between different institutions more secure, ensuring the privacy of institutional data transmission, and improving the security of institutional data; the data classification process is calibrated by the number of abnormal data accesses, improving the accuracy of data security requirement judgments, thereby enhancing the accuracy of data classification and achieving more precise data classification. Abnormal access scenarios adjust the preset data security requirement index range to further improve the accuracy of data classification judgment. The data classification module, by judging the completeness of metadata, adjusts the impact of the cold start coefficient on the preset data security requirement index range, further improving the accuracy of data security requirement index judgment. The intelligent compression and encryption module maximizes comprehensive performance by constructing a multi-objective optimization model, improving the accuracy of intelligent compression and encryption module scheme selection. The intelligent compression and encryption module, by judging the data compression ratio, corrects the judgment process of the comprehensive evaluation function value, increases the security weight ratio, thereby improving the accuracy of judging the comprehensive evaluation function value. The intelligent compression and encryption module assesses CPU utilization and corrects the assessment of data compression ratio, improving the accuracy of data compression ratio judgment, reducing the impact of computing resource weights on the comprehensive evaluation function value, and improving the accuracy of metadata completeness assessment. The dynamic privacy budget allocation module improves the security of organizational data and the accuracy of intelligent compression and encryption by calculating budget allocation ratios. This module assesses data encryption time consumption and adjusts the budget allocation ratio, increasing the compression budget allocation and decreasing the encryption budget allocation. Finally, the dynamic privacy budget allocation module assesses the security of data usage scenarios and adjusts the budget allocation accordingly. The system calibrates the data encryption time, preset data security requirement index range, and preset data compression ratio to optimize the accuracy of data encryption time assessment, reduce the impact of the data usage scenario security requirement index on data level assessment, and improve the accuracy of data compression ratio assessment. The cross-institutional data transmission module uses adversarial transfer learning to align institutional data features, improving data transmission efficiency. It also assesses the number of effective feature extraction samples, calibrates the data usage scenario security requirement index and data security requirement index, and optimizes the accuracy of data encryption time assessment. This improves the data security requirement index's assessment of data level and enhances the security of institutional data. Attached Figure Description

[0016] Figure 1This is a schematic diagram of the structure of the deep learning-based cross-institutional privacy-preserving data system in this embodiment; Figure 2 This is a flowchart illustrating the cross-institutional privacy protection data method based on deep learning in this embodiment. Detailed Implementation

[0017] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0018] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0019] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0020] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0021] Please see Figure 1 The diagram shown is a structural schematic of a cross-institutional privacy-preserving data system based on deep learning, as described in this embodiment. The system includes: The data acquisition module is used to acquire institutional data and characterization data; The data grading module is used to analyze the data grading of the organization's data, calibrate and adjust the data grading of the organization's data based on the number of abnormal accesses in the characterization data, and adjust the adjustment process of the data grading of the organization's data based on the metadata perfection index calculated from the characterization data. The data grading module is connected to the data acquisition module. The data scheme selection module is used to analyze the data compression and encryption combination scheme according to the data level, select the data compression and encryption combination scheme using a multi-objective optimization model based on the comprehensive evaluation function value calculated from the characterization data, send the data compression and encryption combination scheme to the intelligent compression and encryption module, correct the data compression and encryption combination scheme based on the data compression ratio calculated from the characterization data, and correct the correction process of the data compression and encryption combination scheme based on the CPU utilization in the characterization data. The data scheme selection module is connected to the data classification module. The dynamic privacy budget allocation module is used to calculate the budget allocation ratio according to the utility quantification method, obtain the budget allocation ratio, and send the budget allocation ratio to the intelligent compression and encryption module. It is also used to calibrate the budget allocation ratio according to the data encryption time in the characterization data, and to adjust the calibration process of the budget allocation ratio according to the data usage scenario security requirement index. The dynamic privacy budget allocation module is connected to the data scheme selection module. The intelligent compression and encryption module is used to encrypt and compress the organization's data according to the data compression and encryption combination scheme and the budget allocation ratio to obtain encrypted and compressed data. The intelligent compression and encryption module is connected to the data scheme selection module and the dynamic privacy budget allocation module. A cross-agency data transmission module is used to output feature-aligned data based on the encrypted compressed data using an adversarial migration model, and to transmit the feature-aligned data. The cross-agency data transmission module is connected to the intelligent compression and encryption module.

[0022] Specifically, the system is installed in a cross-institutional equipment management terminal. Through the collection of cross-institutional data, it classifies the collected institutional data into hierarchical categories, achieving privacy-level management during institutional data transmission to avoid privacy leakage risks. Intelligent compression and encryption reduce the computational burden across institutions, improving the transmission efficiency of privacy-protected data. Specifically, the system uses a data acquisition module to collect institutional data and representative data, achieving secure and efficient data transmission and improving the efficiency of institutional data transmission. The data classification module determines the data classification status, dividing institutional data into different security levels, making the transmission of institutional data between different institutions more secure, enhancing the privacy of institutional data transmission, and improving the security of institutional data. The system aims to improve the accuracy of data security requirement assessment by calibrating the data level judgment process based on the number of abnormal data accesses. This, in turn, enhances the accuracy of data grading. The data grading module adjusts the preset data security requirement index range based on abnormal access patterns, further improving the accuracy of data grading. It also adjusts the impact of the cold start coefficient on the preset data security requirement index range by assessing the completeness of metadata, further improving the accuracy of data security requirement index judgment. The intelligent compression and encryption module improves the accuracy of intelligent compression and encryption scheme selection by constructing a multi-objective optimization model. Finally, the intelligent compression and encryption module compresses data... The system compares the situation and corrects the judgment process of the comprehensive evaluation function value to improve the accuracy of the judgment. The intelligent compression and encryption module judges the CPU utilization and corrects the judgment of the data compression ratio to improve the accuracy of the data compression ratio judgment, reduce the influence of computing resource weight on the comprehensive evaluation function value, and improve the accuracy of the judgment of metadata completeness. The dynamic privacy budget allocation module improves the security of organizational data and the accuracy of intelligent compression and encryption by calculating the budget allocation ratio. The dynamic privacy budget allocation module judges the data encryption time and adjusts the budget allocation ratio, so that the adjusted compression budget allocation increases and the encryption budget allocation decreases. The privacy budget allocation module assesses the security of data usage scenarios, calibrates preset data encryption time, preset data security requirement index ranges, and preset data compression ratios, optimizes the accuracy of data encryption time assessment, reduces the impact of the data usage scenario security requirement index on data level judgment, and improves the accuracy of data compression ratio judgment. The cross-institutional data transmission module uses adversarial transfer learning to align institutional data features, improving data transmission efficiency. It assesses the number of effective feature extraction samples, calibrates the data usage scenario security requirement index and the data security requirement index, optimizes the accuracy of data encryption time assessment, and thus improves the data security requirement index's impact on data level judgment.Improve the security of institutional data.

[0023] Specifically, the data acquisition module collects institutional data and characterization data. Institutional data refers to data used by various institutions for cross-institutional transmission. Characterization data includes metadata integrity index, metadata consistency index, metadata standardization index, metadata update timeliness index, metadata access convenience index, data anomaly access frequency, security, computing resource consumption, transmission efficiency, original data size, compressed data size, CPU utilization, and data encryption time. Metadata is information used to describe the attributes of institutional data. The metadata integrity index refers to the completeness of the metadata of each institution. This embodiment does not limit the numerical form of the metadata integrity index; those skilled in the art can set it according to actual needs. For example, the metadata integrity index can be set to a value between 0 and 1, with the closer the value is to 1, the more complete the metadata. The metadata consistency index refers to the degree of matching between metadata and real data, where real data refers to data actually generated and existing in each organization. This embodiment does not limit the numerical form of the metadata consistency index; those skilled in the art can set it according to actual needs. For example, the metadata consistency index can be set to a value between 0 and 1, with the closer the value is to 1, the more complete the metadata. The metadata standardization index refers to the degree of consistency of metadata across different organizations. This embodiment does not limit the numerical form of the metadata standardization index; those skilled in the art can set it according to actual needs. The metadata standardization index can be set between 0 and 1, with a value closer to 1 indicating more complete metadata. The metadata update timeliness index refers to the timeliness of metadata updates. This embodiment does not limit the numerical form of the metadata update timeliness index; those skilled in the art can set it according to actual needs. For example, the metadata update timeliness index can be set between 0 and 1, with a value closer to 1 indicating more complete metadata. The metadata acquisition convenience index refers to the ease and convenience of acquiring metadata. This embodiment does not limit the numerical form of the metadata acquisition convenience index; those skilled in the art can set it according to actual needs. For example, the metadata acquisition convenience index can be set between 0 and 1. The closer the value is to 1, the more complete the metadata is. The data acquisition module collects metadata integrity index, metadata consistency index, metadata standardization index, metadata update timeliness index, and metadata access convenience index using the Alation commercial metadata management tool. The Alation commercial metadata management tool is a tool for collecting metadata. The number of abnormal data accesses refers to the number of times the organization's data resources have been subjected to abnormal access behavior. The data acquisition module collects the number of abnormal data accesses through log monitoring. The log monitoring data refers to the data obtained by each organization through daily monitoring. The security refers to the degree of security of the organization's data. This embodiment does not limit the calculation method of security.Those skilled in the art can freely choose according to actual needs, such as mapping security data to 0-1. The security data refers to the data acquired for security indicators. The computational resource consumption refers to the degree of computational resource consumption of the organization's data. This embodiment does not limit the calculation method of computational resource consumption; those skilled in the art can freely choose according to actual needs, such as normalization. The transmission efficiency refers to the efficiency of data transmission between organizations. This embodiment does not limit the calculation method of transmission efficiency; those skilled in the art can freely choose according to actual needs, such as setting T = transmitted data volume / transmission time. The transmitted data volume refers to the amount of data transmitted by the organization, the transmission time refers to the time consumed by the organization's data before and after transmission, the original data size refers to the data size before compression, and the compressed data size refers to the data size after compression. The data acquisition module collects the original data size and compressed data size through system data file attributes. The system data file attributes refer to the information contained in the file of the system where the organization's data resides. The CPU utilization rate refers to the degree to which the organization's data is utilized by the central processing unit (CPU). The unit, or Central Processing Unit, is used to execute computer commands. The data acquisition module obtains CPU utilization via an API (Application Programming Interface). The data encryption time refers to the time consumed in encrypting organizational data. The data acquisition module collects this encryption time using a system encryption timer, a system tool for tracking data encryption time.

[0024] Specifically, the data acquisition module collects institutional data and characterization data to achieve secure and efficient data transmission, thereby improving the efficiency of institutional data transmission.

[0025] Specifically, the data classification module acquires data sensitivity and data usage index, and calculates the data security requirement index Sq based on data sensitivity Sm, data usage index Sy, first weighting coefficient α1, and second weighting coefficient α2. Sq is set as α1×Sm+α2×Sy to obtain the data security requirement index Sq. The data security requirement index Sq is compared with a preset data security requirement index range Sq0. Based on the comparison result, the data classification is determined, and the data level of the organization's data is output based on the determination result. When Sq < Sq0min, the data classification module determines that the data classification is public level data and outputs the public level data as the data level of the organization's data; When Sq0min≤Sq≤Sq0max, the data classification module determines the data classification status as restricted data and outputs the restricted data as the data level of the organization data. When 1 > Sq > Sq0max, the data classification module determines that the data classification is sensitive data and outputs the sensitive data as the data level of the organization data. When Sq=1, the data classification module determines that the data is classified as top secret and outputs the top secret data as the data level of the organization's data.

[0026] Specifically, the data sensitivity refers to the sensitivity level of the organization's data. This embodiment does not limit the method of obtaining data sensitivity data; those skilled in the art can set it according to actual needs. For example, data sensitivity can be obtained through a sensitive data scanning tool, and the data sensitivity value can be set between 0 and 1. The data usage index refers to a value that maps the usage of the organization's data to 0, 0.5, and 1, where 0 is basic data, 0.5 is business data, and 1 is analytical data. Basic data refers to the value of the organization's data having no usage, business data refers to the value of the organization's data having business uses, and analytical data refers to the value of the organization's data having analytical uses. This embodiment does not limit the method of judging no usage, business uses, and analytical uses; those skilled in the art can set it according to actual needs. For example, no usage, business uses, and analytical uses can be judged through business continuity analysis. The data security requirement index refers to the specific numerical manifestation of the organization's data security status. The first weighting coefficient refers to the weighting coefficient used to calculate the data security requirement index, and the second weighting coefficient refers to the weighting coefficient used to calculate the data security requirement index. Regarding the weighting coefficients used to calculate the data security requirement index, this embodiment does not limit the first weighting coefficient α1 and the second weighting coefficient α2. Those skilled in the art can set them according to actual needs. For example, the first weighting coefficient α1 can be set to 0.7 and the second weighting coefficient α1 can be set to 0.3. The preset data security requirement index range refers to the preset value for judging the data security requirement index. This embodiment does not limit the specific value of the preset data security requirement index range. Those skilled in the art can set it according to actual conditions, as long as it meets the requirements for judging the data classification. For example, the preset data security requirement index range Sq0 can be set to: 0.68≤Sq0≤0.76. Sq0min refers to the minimum value of the preset data security requirement index range Sq0, and Sq0max refers to the maximum value of the preset data security requirement index range Sq0. The data classification refers to the data security level of the organization judged based on the data security requirement index and the preset data security requirement index range. The data classification includes public data, restricted data, sensitive data, and top-secret data.

[0027] Specifically, the data classification module obtains the data security requirement index, judges the data classification status, and divides the institutional data into different security levels, making the transmission of institutional data between different institutions more secure, achieving the privacy of institutional data transmission, and improving the security of institutional data.

[0028] Specifically, the data grading module compares the number of abnormal data accesses Sf in the characterizing data with a preset number of abnormal data accesses Sf0, judges the abnormal data access situation based on the comparison result, and calibrates the data grade of the institutional data based on the judgment result, wherein: When Sf≤Sf0, the data classification module determines the abnormal data access situation as a normal situation and does not calibrate the data level; When Sf > Sf0, the data classification module determines the abnormal data access as an abnormal situation, calibrates the data level, and calibrates the preset data security requirement index Sq0 by using the abnormal access coefficient α. e is the base of the natural logarithm. The preset data security requirement index interval Sq0` after calibration is obtained. Sq0` is set to α×Sq0. The preset data security requirement index interval Sq0 is replaced with the preset data security requirement index interval Sq0` after calibration. The data security requirement index Sq is then compared with the preset data security requirement index interval Sq0` after calibration.

[0029] Specifically, the preset number of abnormal data accesses refers to a preset value for judging the number of abnormal data accesses. This embodiment does not limit the specific value of the preset number of abnormal data accesses. Those skilled in the art can set it according to the actual situation, as long as it meets the requirements for judging abnormal data access. For example, the preset number of abnormal data accesses Sf0 can be set to: 5 times / day ≤ Sf0 ≤ 7 times / day. The abnormal data access situation refers to whether the number of abnormal data accesses is normal based on the difference between the number of abnormal data accesses and the preset number of abnormal data accesses. The abnormal data access situation includes normal data access situation and abnormal data access situation. The abnormal access coefficient α is a coefficient used to calibrate the preset data security requirement index range.

[0030] Specifically, the data grading module acquires the number of abnormal data accesses through log monitoring and judges the abnormal data access situation. When the abnormal data access situation is normal, the number of abnormal data accesses has little impact on the data grading judgment result, and no calibration is performed on the data grading judgment process. When the abnormal data access situation is abnormal, the number of abnormal data accesses has a significant impact on the data grading judgment result, and the data grading judgment process is calibrated. By setting an abnormal access coefficient that decreases from 0.94 to infinitely close to 0.83 as the number of abnormal data accesses increases, the data security requirement index is calibrated. This makes the preset data security requirement index range after calibration decrease as the number of abnormal data accesses increases, thereby reducing the impact of the number of abnormal data accesses on the data grading judgment, improving the accuracy of the data security requirement judgment, and thus improving the accuracy of the data grading judgment, achieving more accurate data grading.

[0031] Specifically, the data classification module also adjusts the data classification level of the organization's data based on abnormal data access, wherein: When the abnormal data access situation is normal, the preset data security requirement index Sq0 is adjusted by the cold start coefficient I. Sq01 is set to I×Sq0 to obtain the adjusted preset data security requirement index range Sq01. The preset data security requirement index range Sq0 is replaced with the adjusted preset data security requirement index range Sq01, and the data security requirement index Sq is compared with the adjusted preset data security requirement index range Sq01 again. When the abnormal data access situation is an abnormal situation, the preset data security requirement index Sq0` is adjusted by the cold start coefficient I, and Sq02 is set to I×Sq0` to obtain the adjusted preset data security requirement index range Sq02. The preset data security requirement index range Sq0 is replaced with the adjusted preset data security requirement index range Sq02, and the data security requirement index Sq is compared with the adjusted preset data security requirement index range Sq02 again.

[0032] Specifically, the cold start coefficient I refers to a coefficient used to adjust the preset data security requirement index range. This embodiment does not limit the method of obtaining the cold start coefficient. Those skilled in the art can set it according to the actual situation. For example, the cold start coefficient I can be set as: 0.67≤l≤0.89.

[0033] Specifically, the data classification module adjusts the preset data security requirement index range through a cold start coefficient, thereby reducing the adjusted preset data security requirement index range, improving the accuracy of data classification judgment, enhancing the security level of institutional data, and further improving the accuracy of data classification judgment.

[0034] Specifically, the data grading module calculates the metadata completeness index Fp based on the metadata integrity index Sw, metadata consistency index Ss, metadata standardization index Sk, metadata update timeliness index Sj, and metadata access convenience index Sp in the data. Let Fp = A3×Sw + A4×Ss + A5×Sk + A6×Sj + A7×Sp, to obtain the metadata completeness index Fp. The metadata completeness index Fp is then compared with a preset metadata completeness index Fp0. Based on the comparison result, the metadata completeness is judged, and the data level adjustment process for the institutional data is adjusted according to the judgment result. Wherein: When Fp > Fp0, the data grading module determines that the metadata completeness is normal and does not adjust the data grading process. When Fp≤Fp0, the data grading module determines that the metadata completeness is abnormal and adjusts the data grading process. The cold start coefficient l is adjusted by using the preset metadata completeness index Fp0 and the metadata completeness index Fp. The cold start coefficient l` is set to (Fp0 / Fp)×l to obtain the adjusted cold start coefficient l`. The cold start coefficient l is replaced with the adjusted cold start coefficient l`, and the preset data security requirement index range is readjusted according to the adjusted cold start coefficient l`. The data grading is then re-analyzed according to the adjusted preset data security requirement index range.

[0035] Specifically, the metadata completeness index refers to the specific numerical representation of the metadata completeness of each organization. The third weighting coefficient A3, the fourth weighting coefficient A4, the fifth weighting coefficient A5, the sixth weighting coefficient A6, and the seventh weighting coefficient A7 are all weighting coefficients used to calculate the metadata completeness index. This embodiment does not limit the acquisition method of the third weighting coefficient A3, the fourth weighting coefficient A4, the fifth weighting coefficient A5, the sixth weighting coefficient A6, and the seventh weighting coefficient A7. Those skilled in the art can set them according to actual needs, such as setting A3=A4=A5=A6=A7=0.2. The preset metadata completeness index refers to the preset value for judging the metadata completeness index. This embodiment does not limit the specific value of the preset metadata completeness index. Those skilled in the art can set it according to actual conditions, as long as it meets the requirements for judging the metadata completeness. For example, the preset metadata completeness index Fp0 can be set to: 30%≤Fp0≤70%. The metadata completeness status refers to whether the metadata completeness index is normal based on the metadata completeness index and the preset metadata completeness index. The metadata completeness status includes normal and abnormal metadata completeness.

[0036] Specifically, the data grading module calculates the metadata completeness index to determine the metadata completeness. When the metadata completeness is normal, the metadata completeness index has little impact on the cold start coefficient, and no adjustment is made to the cold start coefficient. When the metadata completeness is abnormal, the metadata completeness index has a significant impact on the cold start coefficient, and the cold start coefficient is adjusted. By adjusting the cold start coefficient through the metadata completeness index, the adjusted cold start coefficient increases as the metadata completeness index increases. This adjusts the impact of the cold start coefficient on the preset data security requirement index range, further improving the accuracy of the data security requirement index judgment and thus further improving the accuracy of the data grading judgment.

[0037] Specifically, the data scheme selection module analyzes the data compression and encryption combination scheme of the institutional data according to the data level to obtain the data compression and encryption combination scheme of the institutional data, wherein: When the data is classified as top secret, the data scheme selection module uses the Zstandard compression algorithm to compress the data of the organization to obtain compressed data, and uses the AES-256-GCM encryption algorithm to encrypt the compressed data to obtain encrypted top secret data. When the data level is sensitive data, the data scheme selection module uses the Brotli compression algorithm to compress the data to obtain compressed data, and uses the ChaCha20-Poly1305 encryption algorithm to encrypt the compressed data to obtain encrypted sensitive data. When the data level is restricted, the data scheme selection module uses the LZ4 compression algorithm to compress the data to obtain compressed data, and uses the AES-128-CBC encryption algorithm to encrypt the compressed data to obtain encrypted restricted data. When the data level is public, the data scheme selection module uses the gzip compression algorithm to compress the data to obtain compressed public-level data. The data scheme selection module constructs a multi-objective optimization model according to a multi-objective optimization model construction method, which includes: Step B01: Calculate the comprehensive evaluation function value F based on security S, computational resource consumption C, transmission efficiency T, eighth weight coefficient ws, ninth weight coefficient wc, and tenth weight coefficient wt. Set F = ws × S + wc × (1 - C) + wt × T to obtain the comprehensive evaluation function value. Step B02: Set constraints to optimize the comprehensive evaluation function value to obtain a multi-objective optimization model; The data scheme selection module inputs security, computational resource consumption, transmission efficiency, eighth weight coefficient, ninth weight coefficient, and tenth weight coefficient into a multi-objective optimization model, outputs a comprehensive evaluation function value, and compares the comprehensive evaluation function value F with a preset comprehensive evaluation function F0. Based on the comparison result, the module determines the state of the comprehensive evaluation function value and selects a data compression and encryption combination scheme for the institutional data based on the determination result. When F≥F0, the data scheme selection module determines that the comprehensive evaluation function value is in a normal state, selects the data compression and encryption combination scheme, and sends the data compression and encryption combination scheme to the intelligent compression and encryption module. When F < F0, the data scheme selection module determines that the comprehensive evaluation function value is in an abnormal state and does not select the data compression and encryption combination scheme.

[0038] Specifically, the data compression and encryption combination scheme refers to a scheme for compressing and encrypting organizational data according to data levels. The Zstandard compression algorithm is an algorithm for compressing organizational data. The AES-256-GCM encryption algorithm, short for Advanced Encryption Standard with 256-bit key in Galois / Counter Mode, is an encryption algorithm based on a 256-bit key advanced encryption standard in counter mode. Counter mode refers to a mode that uses a counter to calculate the key. The Brotli compression algorithm is a high-efficiency data compression format compression algorithm. High-efficiency data compression format refers to a format with higher compression efficiency than general data compression. The ChaCha20-Poly1305 encryption algorithm, short for ChaCha20-Poly1305 Authenticated Encryption with Associated Data (AEAD) Cipher, is an encryption algorithm for the ChaCha20-Poly1305 authentication encryption scheme with associated data. Associated data refers to different data with interconnected properties. The authentication encryption scheme refers to a scheme that adds authentication functionality to the encryption scheme. The LZ4 compression algorithm is short for Lempel–Ziv. 4. The Chinese name is "Extremely Fast Compression Algorithm," where "extremely fast" refers to its extremely high compression speed. The full name of the AES-128-CBC encryption algorithm is Advanced Encryption Standard with 128-bit key in Cipher Block Chaining mode. The Cipher Block Chaining mode refers to a mode for encrypting data longer than a single block. "Longer than a single block" means data blocks longer than a single block, and "block data" refers to data of a fixed length. The full name of the gzip compression algorithm is GNU... ZIP refers to a compression algorithm in GNU compression. The multi-objective optimization model refers to a computational model that takes security, computational resource consumption, transmission efficiency, and the eighth, ninth, and tenth weighting coefficients as inputs and outputs a comprehensive evaluation function value. The eighth, ninth, and tenth weighting coefficients are coefficients used to calculate the comprehensive evaluation function value F. This embodiment does not limit the method of obtaining the eighth weighting coefficient ws, the ninth weighting coefficient wc, and the tenth weighting coefficient wt. Those skilled in the art can freely choose according to actual needs, such as setting top-secret data: ws=0.6, wc=0.2, wt=0.2; sensitive data: ws=0.4, wc=0.3, wt=0.3; and restricted data: ws=0.2, wc=0.3, wt=0.5. The constraint conditions refer to additional conditions for optimizing the comprehensive evaluation function value. This embodiment does not limit the setting method of the constraint conditions. Those skilled in the art can freely choose according to actual needs. For example, the value of the security S can be set to satisfy: S≥0.8. Under the condition that the value of the security S is satisfied, the comprehensive evaluation function value is calculated to obtain the optimized comprehensive evaluation function value. Alternatively, the value of the computing resource consumption C can be set to satisfy: C≤90%. Under the condition that the value of the computing resource consumption is satisfied, the comprehensive evaluation function value is calculated to obtain the optimized comprehensive evaluation function value. The comprehensive evaluation function value refers to... The institutional data is evaluated based on a comprehensive assessment. The preset comprehensive assessment function value refers to a pre-defined value used to judge the comprehensive assessment function value. This embodiment does not limit the specific value of the preset comprehensive assessment function value; those skilled in the art can set it according to actual conditions, as long as it meets the judgment requirements of the comprehensive assessment function value state. For example, the preset comprehensive assessment function value F0 can be set to: 0.89≤F0≤0.95. The comprehensive assessment function value state refers to the state of the comprehensive assessment function value, judged based on its comparison with the preset comprehensive assessment function value, indicating whether it is normal. The comprehensive assessment function value state includes normal and abnormal states.

[0039] Specifically, the data scheme selection module constructs a multi-objective optimization model to maximize overall performance while meeting safety requirements, thereby improving the accuracy of the comprehensive evaluation function value judgment and enhancing the accuracy of the data scheme selection module in selecting schemes.

[0040] Specifically, the data scheme selection module calculates the data compression ratio Ys based on the original data size Dd and the compressed data size Yd in the characterization data, setting Ys = Dd / Yd to obtain the data compression ratio Ys. The data compression ratio Ys is then compared with a preset data compression ratio Ys0. Based on the comparison result, the data compression ratio is judged, and the data compression and encryption combination scheme is modified according to the judgment result. When Ys≥Ys0, the data scheme selection module determines that the data compression ratio is normal and does not modify the data compression and encryption combination scheme. When Ys < Ys0, the data scheme selection module determines that the data compression ratio is abnormal, corrects the data compression and encryption combination scheme, calibrates the eighth weight coefficient ws using the compression coefficient K, and sets... e is the base of the natural logarithm. The eighth weight coefficient ws` after calibration is obtained. ws` is set to K×ws, wc remains unchanged, wt`=(1-K)×wt. The comprehensive evaluation function value F1 after calibration is obtained. The comprehensive evaluation function value F1 after calibration is compared with the preset comprehensive evaluation function F0.

[0041] Specifically, the data compression ratio refers to the ratio of the data size before and after compression. The preset data compression ratio is a preset value for judging the data compression ratio. This embodiment does not limit the specific value of the preset data compression ratio. Those skilled in the art can set it according to the actual situation, as long as it meets the requirements for judging the data compression ratio. For example, the preset data compression ratio Ys0 can be set as 95% ≤ Ys0 ≤ 99%. The data compression ratio condition refers to whether the data compression ratio is normal based on the data compression ratio and the preset data compression ratio. The data compression ratio condition includes normal and abnormal conditions.

[0042] Specifically, the data scheme selection module calculates the data compression ratio and judges its status. When the data compression ratio is normal, it has little impact on the judgment result of the comprehensive evaluation function value, and no calibration is performed on the judgment process. When the data compression ratio is abnormal, it has a significant impact on the judgment result of the comprehensive evaluation function value, and the judgment process is calibrated. The security weight is calibrated by setting a compression coefficient that increases from 1.13 to infinitely close to 1.42 as the data compression ratio decreases, thereby calibrating the comprehensive evaluation function value. The calibrated security weight value increases as the data compression ratio decreases, increasing the proportion of security weight and thus improving the accuracy of the judgment of the comprehensive evaluation function value.

[0043] Specifically, the data scheme selection module compares the CPU utilization rate Lc in the characterizing data with the preset CPU utilization rate Lc0, judges the CPU utilization based on the comparison result, and corrects the data compression and encryption combination scheme modification process and the data level adjustment process based on the judgment result, wherein: When Lc≤Lc0, the data scheme selection module determines that the CPU utilization is normal and does not correct the process of data compression and encryption combination scheme modification and data level adjustment. When Lc > Lc0, the data scheme selection module determines that the CPU utilization is abnormal, corrects the process of modifying the data compression and encryption combination scheme and adjusting the data level, and corrects the preset compression ratio Ys0 by using the coefficient g. Let e ​​be the base of the natural logarithm. The corrected preset data compression ratio Ys0` is obtained. Ys0` is set to Ys0 / (0.47+g). The preset data compression ratio Ys0 is replaced with the corrected preset data compression ratio Ys0`. The data compression ratio Ys is then compared again with the corrected preset data compression ratio Ys0`. The data scheme selection module corrects the ninth weight coefficient wc using coefficient g, setting wc` = wc × g, to obtain the corrected ninth weight coefficient wc`. The data scheme selection module corrects the preset metadata perfection index Fp0 using coefficient g, setting Fp0` = g × Fp0, to obtain the corrected preset metadata perfection index Fp0`. The data classification adjustment process is corrected based on the corrected preset metadata perfection index.

[0044] Specifically, the preset CPU utilization rate is a preset value for judging CPU utilization. This embodiment does not limit the specific value of the preset CPU utilization rate. Those skilled in the art can set it according to the actual situation, as long as it meets the requirements for judging CPU utilization. For example, the preset CPU utilization rate Lc0 can be set to 70% ≤ Lc0 ≤ 80%. The CPU utilization status refers to whether the CPU utilization is normal based on the CPU utilization rate and the preset CPU utilization rate. The CPU utilization status includes normal status and abnormal status.

[0045] Specifically, the data scheme selection module judges the CPU utilization. When the CPU utilization is normal, it does not correct the judgment of the data compression ratio, nor does it correct the computing resource weight and the preset metadata perfection index. When the CPU utilization is abnormal, it corrects the judgment of the data compression ratio, the computing resource weight, and the preset metadata perfection index. By setting a utilization coefficient that decreases from 0.52 to infinitely close to 0.32 as the CPU utilization increases, the preset data compression ratio is corrected, so that the corrected preset data compression ratio increases with the increase of CPU utilization, thereby improving the accuracy of the judgment of the data compression ratio. By setting the utilization coefficient, the computing resource weight is adjusted, so that the corrected computing resource weight decreases with the increase of CPU utilization, thereby reducing the influence of the computing resource weight on the comprehensive evaluation function value. By setting the utilization coefficient, the preset metadata perfection index is corrected, so that the corrected preset metadata perfection index decreases with the increase of CPU utilization, thereby improving the accuracy of the judgment of the metadata perfection status.

[0046] Specifically, the dynamic privacy budget allocation module calculates the budget allocation ratio for the organization's data using a utility quantification method, which includes: Step C01, based on the compression ratio Ys, the maximum compression ratio Ysmax, and the initial privacy budget. Encryption strength value Perform calculations and set Obtain the privacy consumption value ; Step C02, based on the encryption key length K, the maximum security key length Kmax, the minimum security key length Kmin, and the total privacy budget. Privacy consumption value Perform calculations and set To obtain the encryption strength value ; Step C03, set the privacy consumption value Encryption strength value Initial privacy budget Total privacy budget The compression utility coefficient g1, the encryption utility coefficient g2, and the Lagrange multiplier W are used to calculate the budget allocation ratio using the Lagrange function, thus obtaining the budget allocation ratio. The dynamic privacy budget allocation module sends the budget allocation ratio to the intelligent compression and encryption module.

[0047] Specifically, the utility quantification method refers to a method for calculating the budget allocation ratio. The maximum compression ratio refers to the maximum value of the compression ratio before and after data compression. This embodiment does not limit the specific value of the maximum compression ratio; those skilled in the art can obtain it according to actual conditions, such as by statistically analyzing the ratio of data size before and after compression. The initial privacy budget refers to the initial value of the data privacy budget allocation. This embodiment does not limit the specific value of the initial privacy budget; those skilled in the art can obtain it through a differential privacy platform. The differential privacy platform refers to a platform for data analysis under the premise of protecting the institution's data privacy. The privacy consumption value refers to the risk level of the institution's data privacy leakage. The encryption key length refers to the number of bits used to encrypt the institution's data. The maximum security key length refers to the maximum number of bits in the security key. The minimum security key length refers to the minimum number of bits in the security key. This embodiment does not limit the method of obtaining the encryption key length, maximum security key length, and minimum security key length; those skilled in the art can obtain them according to actual conditions, such as by using the key length calculator built into the system. The total privacy budget refers to the sum of the budgets allocated for data compression and encryption. This embodiment does not limit the method of obtaining the total privacy budget. The encryption strength value refers to the degree of encryption of the organization's data. The compression utility coefficient is a coefficient calculated using the Lagrange multiplier method. This embodiment does not limit the specific value of the compression utility coefficient, as long as it meets the calculation requirements for the budget allocation ratio. For example, the compression utility coefficient g1 can be set to: 0≤g1≤1. The encryption utility coefficient is a coefficient calculated using the Lagrange multiplier method. This embodiment does not limit the specific value of the encryption utility coefficient, as long as it meets the calculation requirements for the budget allocation ratio. For example, the encryption utility coefficient can be set to: 0≤g1≤1. The coefficient g2 is defined as: 0 ≤ g2 ≤ 1. The Lagrange multiplier is a coefficient used to calculate the Lagrange function. This embodiment does not limit the specific value of the Lagrange multiplier, as long as it meets the calculation requirements of the budget allocation ratio. For example, the Lagrange multiplier W can be set as: -5 ≤ W ≤ 15. The Lagrange function refers to the function used to calculate the budget allocation ratio. The budget allocation ratio refers to the vertical ratio obtained according to the utility quantification method. This embodiment does not limit the method of sending the budget allocation ratio. Those skilled in the art can obtain it themselves according to the actual situation and send it through the institutional data system platform.

[0048] Specifically, the dynamic privacy budget allocation module constructs a budget allocation ratio by considering privacy risks and encryption strength, enabling dynamic allocation of the budget ratio in the compression and encryption stages, thereby improving the security of institutional data and the accuracy of intelligent compression and encryption.

[0049] Specifically, the dynamic privacy budget allocation module compares the data encryption time Th with the preset data encryption time Th0, judges the data encryption time based on the comparison result, and adjusts the budget allocation ratio according to the judgment result, wherein: When Th≤Th0, the dynamic privacy budget allocation module determines that the data encryption time consumption is normal and does not adjust the budget allocation ratio; When Th > Th0, the dynamic privacy budget allocation module determines that the data encryption time consumption is an abnormal situation and adjusts the budget allocation ratio. This adjustment is made using a time consumption coefficient h. The adjusted encryption strength value is obtained. The budget allocation ratio is optimized based on the adjusted encryption strength value.

[0050] Specifically, the preset data encryption time refers to a preset value for judging the data encryption time. This embodiment does not limit the specific value of the preset data encryption time. Those skilled in the art can set it according to the actual situation, as long as it meets the requirements for judging the data encryption time. For example, the preset data encryption time Th0 can be set to: 10MB / s≤Th0≤15MB / s. The data encryption time status refers to whether the data encryption time is normal based on the judgment between the data encryption time and the preset data encryption time. The data encryption time status includes normal status and abnormal status.

[0051] Specifically, the dynamic privacy budget allocation module judges the data encryption time. When the data encryption time is normal, the data encryption time has little impact on the budget allocation and no adjustment is made to the budget allocation. When the data encryption time is abnormal, the data encryption time has a significant impact on the budget allocation and the budget allocation ratio is adjusted, so that the adjusted compression budget allocation increases and the encryption budget allocation decreases.

[0052] Specifically, the dynamic privacy budget allocation module compares the data usage scenario security requirement index As with the preset data usage scenario security requirement index As0, judges the security status of the data usage scenario based on the comparison result, and calibrates the adjustment process of the budget allocation ratio, the data level, and the correction process of the data compression and encryption combination scheme based on the judgment result, wherein: When As≤As0, the dynamic privacy budget allocation module determines that the data usage scenario security is normal and does not calibrate the process of adjusting the budget allocation ratio, data level, and data compression and encryption combination scheme correction. When As > As0, the dynamic privacy budget allocation module determines that the data usage scenario security situation is abnormal, calibrates the adjustment process of the budget allocation ratio, the data level, and the data compression and encryption combination scheme correction process, and calibrates the preset data encryption time Th0 through the security factor v. Let e ​​be the base of the natural logarithm. The calibrated preset data encryption time Th0` is obtained. Th0` is set to v × Th0. The preset data encryption time Th0 is replaced with the calibrated preset data encryption time Th0`. The data encryption time Th is then compared again with the calibrated preset data encryption time Th0`. The dynamic privacy budget allocation module calibrates the preset data security requirement index interval Sq0 using the security coefficient v, obtaining the calibrated preset data security requirement index interval Sq0`. Sq0` is set to v × Sq0. The preset data security requirement index interval Sq0 is replaced with the calibrated preset data security requirement index interval Sq0`. A preset data security requirement index range Sq0` is established, and the data security requirement index Sq is re-compared with the calibrated preset data security requirement index range Sq0`. The dynamic privacy budget allocation module calibrates the preset data compression ratio Ys0 through the security coefficient v to obtain the calibrated preset data compression ratio Ys0`. Ys0` is set to Ys0 / (0.05+v). When Ys0` is greater than 1, it is set to equal 1. The preset data compression ratio Ys0 is replaced with the calibrated preset data compression ratio Ys0`, and the data compression ratio Ys is re-compared with the calibrated preset data compression ratio Ys0`.

[0053] Specifically, the data usage scenario security requirement index refers to the specific numerical representation of the security status of an organization's data usage scenario. This embodiment does not limit the method of obtaining the data usage scenario security requirement index As. Those skilled in the art can set it according to actual needs, such as obtaining it through expert experience, setting 0≤As≤1. The preset data usage scenario security requirement index is a preset value of the data usage scenario security requirement index. This embodiment does not limit the specific value of the preset data usage scenario security requirement index. Those skilled in the art can set it according to actual conditions, as long as it meets the requirements for judging the security status of the data usage scenario. For example, the preset data usage scenario security requirement index As0 can be set to: 0.56≤As0≤0.72. The data usage scenario security status refers to whether the data usage scenario security requirement index is normal, judged based on the data usage scenario security requirement index and the preset data usage scenario security requirement index. The data usage scenario security status includes the data usage scenario security status being normal and the data usage scenario security status being abnormal.

[0054] Specifically, the dynamic privacy budget allocation module assesses the security status of data usage scenarios. When the data usage scenario is under normal security conditions, it does not calibrate the preset data encryption time, preset data security requirement index range, and preset data compression ratio. When the data usage scenario is under abnormal security conditions, it calibrates these parameters. The preset data encryption time is calibrated by setting a security coefficient that decreases from 0.95 to approximately 0.77 as the data usage scenario security requirement index increases. This ensures that the calibrated preset data encryption time decreases as the data usage scenario security requirement index increases, thereby optimizing the accuracy of the data encryption time assessment. The preset data security requirement index range is also calibrated by setting a security coefficient, ensuring that the calibrated range decreases as the index increases, thus reducing the impact of the data usage scenario security requirement index on data level assessment. Finally, the preset data compression ratio is calibrated by setting a utilization coefficient, ensuring that the calibrated ratio increases as the index increases, thereby improving the accuracy of the data compression ratio assessment.

[0055] Specifically, the intelligent compression and encryption module encrypts and compresses the organization's data according to the data compression and encryption combination scheme and the budget allocation ratio. It obtains the data compression and encryption combination scheme and the encryption strength value according to the data level, calculates the privacy consumption value according to the budget allocation ratio, and compresses and encrypts the organization's data using the encryption strength value, the privacy consumption value, and the data compression and encryption combination scheme to obtain encrypted and compressed data.

[0056] Specifically, the encrypted compressed data refers to the data obtained by encrypting and compressing the organization's data according to the data compression and encryption combination scheme and the budget allocation ratio. This embodiment does not limit the specific encryption and compression method; those skilled in the art can choose according to the actual situation. For example, when the organization's data level is top secret, the encryption strength value can be set. satisfy: The privacy consumption value is calculated based on the budget allocation ratio. The organization's data is compressed and encrypted based on the encryption strength value, the privacy consumption value, and the data compression and encryption combination scheme of top-secret data.

[0057] Specifically, the intelligent compression and encryption module encrypts and compresses the organization's data according to the data compression and encryption combination scheme and the budget allocation ratio, thereby achieving the purpose of compressing and encrypting the organization's data, improving the security and privacy of the organization's data, and reducing the data transmission volume of the organization.

[0058] Specifically, the cross-organizational data transmission module constructs the adversarial transfer model according to the adversarial transfer model construction method to obtain the adversarial transfer model, inputs the encrypted compressed data into the adversarial transfer model, outputs feature-aligned data, and transmits the feature-aligned data.

[0059] Specifically, the adversarial transfer model refers to a generative adversarial network model that takes the encrypted compressed data as input and outputs feature-aligned data. In this embodiment, the cross-institutional data transmission module constructs the adversarial transfer model using an adversarial transfer model construction method, which includes: The adversarial analysis dataset is divided into three parts: 70% for adversarial analysis training, 15% for adversarial analysis validation, and 15% for adversarial analysis testing. The adversarial analysis training set is input into a generative adversarial network (GAN) model for training. The adversarial analysis validation set is input into the trained GAN model for iterative hyperparameter optimization. The adversarial analysis testing set is input into the optimized GAN model for adversarial analysis testing, yielding the test results. The total number of samples in the adversarial analysis testing set is set to h0, the number of correctly executed adversarial analysis test samples is set to h, and the adversarial analysis test accuracy is set to H, where H = h / h0. The adversarial analysis test accuracy H is compared with the preset adversarial analysis test accuracy H0. Based on the comparison results, the training performance of the iteratively optimized GAN model is judged, and the judgment result is output. When H≥H0, the cross-organizational data transmission module determines that the training of the iteratively optimized generative adversarial network model has reached the target, outputs the iteratively optimized generative adversarial network model as an adversarial analysis model, and provides the adversarial analysis model to the cross-organizational data transmission module. When H < H0, the cross-institutional data transmission module determines that the trained Generative Adversarial Network (GAN) model after iterative optimization is substandard. It then updates the adversarial analysis dataset to obtain an updated dataset. Based on this updated dataset, the GAN model is trained, its hyperparameters iteratively optimized, and analyzed and tested until the training meets the standards. The adversarial analysis dataset refers to the data set used to construct the adversarial analysis model. It includes historically acquired encrypted and compressed data and corresponding feature-aligned data. The adversarial analysis training set refers to the dataset within the adversarial analysis dataset used for training the GAN model. The GAN model refers to the basic architecture used to construct the adversarial analysis model. The adversarial analysis validation set refers to the dataset within the adversarial analysis dataset used to validate the training results of the GAN model. The adversarial analysis test set refers to the dataset within the adversarial analysis dataset used for testing the GAN model. The adversarial analysis test accuracy is the ratio of the number of correct adversarial analysis test samples to the total number of samples in the adversarial analysis test set. The preset adversarial analysis test accuracy is the target value for the GAN model used in this iterative optimization. The preset values ​​for judging the training compliance status are not limited in this embodiment. Those skilled in the art can set them according to actual needs, as long as they meet the requirements for judging the training compliance status of the iteratively optimized generative adversarial network model. For example, the preset adversarial analysis test accuracy H0 can be set to: 90%≤H0≤95%. The training compliance status of the iteratively optimized generative adversarial network model refers to the accuracy status of the trained generative adversarial network model. The training compliance status of the iteratively optimized generative adversarial network model includes compliance and non-compliance. The feature alignment data refers to the data after matching different data structures of the encrypted compressed data using the adversarial transfer model. The different data structures refer to the different data structures of the encrypted compressed data from different institutions when transmitting across institutions. This embodiment does not limit the transmission method of the feature alignment data. Those skilled in the art can set it according to actual needs, as long as it meets the transmission requirements of the feature alignment data. It can be transmitted through a differential privacy model. The differential privacy model refers to a method of transmitting the encrypted compressed data after noise interference. The noise interference refers to introducing controllable random variables into the transmission of the encrypted compressed data.

[0060] Specifically, the cross-organizational data transmission module, by constructing an adversarial migration model, makes it easier to transmit organizational data, thereby improving the efficiency of organizational data transmission.

[0061] Specifically, please refer to Figure 2As shown, this is a flowchart illustrating the cross-institutional privacy-preserving data method based on deep learning in this embodiment. The method includes: Step S1: Collect institutional data and characterization data; Step S2 involves analyzing the data level of the organization data, calibrating and adjusting the data level of the organization data based on the number of abnormal accesses in the characterization data, and adjusting the adjustment process of the data level of the organization data based on the metadata perfection index calculated from the characterization data. Step S3: Analyze the data compression and encryption combination scheme according to the data level, select the data compression and encryption combination scheme using a multi-objective optimization model based on the comprehensive evaluation function value calculated from the characterization data, and send the data compression and encryption combination scheme to the intelligent compression and encryption module. Also, correct the data compression and encryption combination scheme based on the data compression ratio calculated from the characterization data, and correct the correction process of the data compression and encryption combination scheme based on the CPU utilization in the characterization data. Step S4: Calculate the budget allocation ratio according to the utility quantification method to obtain the budget allocation ratio, and send the budget allocation ratio to the intelligent compression and encryption module. Also, calibrate the budget allocation ratio according to the data encryption time in the characterization data, and adjust the calibration process of the budget allocation ratio according to the data usage scenario security requirement index. Step S5: Encrypt and compress the organization data according to the data compression and encryption combination scheme and the budget allocation ratio to obtain encrypted and compressed data; Step S6: Using the adversarial migration model, output the feature alignment data based on the encrypted compressed data, and then transmit the feature alignment data.

[0062] Specifically, the method classifies the institutional data to obtain institutional data levels, selects the optimal compression and encryption scheme by constructing a multi-objective optimization model, obtains the privacy budget value by allocating the budget to the compression and encryption stages, performs deep learning on cross-institutional privacy protection data, classifies the cross-institutional privacy protection data and selects the optimal compression and encryption scheme, thereby meeting the security requirements of cross-institutional privacy protection data while generating smaller data entities and reducing the computational pressure of cross-institutional data.

[0063] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A cross-institutional privacy-preserving data system based on deep learning, characterized in that, include: The data acquisition module is used to acquire institutional data and characterization data; The data grading module is used to analyze the data grade of the organization's data, calibrate and adjust the data grade of the organization's data based on the number of abnormal accesses in the characterization data, and adjust the adjustment process of the data grade of the organization's data based on the metadata perfection index calculated from the characterization data. The data scheme selection module is used to analyze the data compression and encryption combination scheme according to the data level, select the data compression and encryption combination scheme using a multi-objective optimization model based on the comprehensive evaluation function value calculated from the characterization data, send the data compression and encryption combination scheme to the intelligent compression and encryption module, modify the data compression and encryption combination scheme based on the data compression ratio calculated from the characterization data, and correct the modification process of the data compression and encryption combination scheme based on the CPU utilization in the characterization data. The dynamic privacy budget allocation module is used to calculate the budget allocation ratio according to the utility quantification method, obtain the budget allocation ratio, and send the budget allocation ratio to the intelligent compression and encryption module. It is also used to calibrate the budget allocation ratio according to the data encryption time in the characterization data, and to adjust the calibration process of the budget allocation ratio according to the data usage scenario security requirement index. The intelligent compression and encryption module is used to encrypt and compress the organization's data according to the data compression and encryption combination scheme and the budget allocation ratio to obtain encrypted and compressed data. The cross-agency data transmission module is used to output feature-aligned data based on the encrypted compressed data using an adversarial migration model, and then transmit the feature-aligned data.

2. The cross-institutional privacy-preserving data system based on deep learning according to claim 1, characterized in that, The data grading module acquires data sensitivity and data usage index, and calculates the data security requirement index Sq based on data sensitivity Sm, data usage index Sy, first weighting coefficient α1, and second weighting coefficient α2. Sq is set as α1×Sm+α2×Sy to obtain the data security requirement index Sq. The data security requirement index Sq is compared with a preset data security requirement index range Sq0. Based on the comparison result, the data grading status is judged, and the data level of the organization's data is output according to the judgment result. When Sq < Sq0min, the data classification module determines that the data classification is public level data and outputs the public level data as the data level of the organization's data; When Sq0min≤Sq≤Sq0max, the data classification module determines the data classification status as restricted data and outputs the restricted data as the data level of the organization data. When 1 > Sq > Sq0max, the data classification module determines that the data classification is sensitive data and outputs the sensitive data as the data level of the organization data. When Sq=1, the data classification module determines that the data is classified as top secret and outputs the top secret data as the data level of the organization.

3. The cross-institutional privacy-preserving data system based on deep learning according to claim 2, characterized in that, The data grading module compares the number of abnormal data accesses Sf in the characterizing data with a preset number of abnormal data accesses Sf0, judges the abnormal data access situation based on the comparison result, and calibrates the data grade of the institutional data based on the judgment result, wherein: When Sf≤Sf0, the data classification module determines the abnormal data access situation as a normal situation and does not calibrate the data level; When Sf > Sf0, the data classification module determines the abnormal data access as an abnormal situation, calibrates the data level, and calibrates the preset data security requirement index Sq0 by using the abnormal access coefficient α. Let e ​​be the base of the natural logarithm. The calibrated preset data security requirement index interval Sq0` is obtained. Sq0` is set to α × Sq0. The preset data security requirement index interval Sq0 is replaced with the calibrated preset data security requirement index interval Sq0`. The data security requirement index Sq is then re-compared with the calibrated preset data security requirement index interval Sq0`. The data classification module also adjusts the data level of the institutional data based on abnormal data access, wherein: When the abnormal data access situation is normal, the preset data security requirement index Sq0 is adjusted by the cold start coefficient I. Sq01 is set to I×Sq0 to obtain the adjusted preset data security requirement index range Sq01. The preset data security requirement index range Sq0 is replaced with the adjusted preset data security requirement index range Sq01, and the data security requirement index Sq is compared with the adjusted preset data security requirement index range Sq01 again. When the abnormal data access situation is an abnormal situation, the preset data security requirement index Sq0` is adjusted by the cold start coefficient I, and Sq02 is set to I×Sq0` to obtain the adjusted preset data security requirement index range Sq02. The preset data security requirement index range Sq0 is replaced with the adjusted preset data security requirement index range Sq02, and the data security requirement index Sq is compared with the adjusted preset data security requirement index range Sq02 again.

4. The cross-institutional privacy-preserving data system based on deep learning according to claim 3, characterized in that, The data grading module calculates the metadata completeness index Fp based on the metadata integrity index Sw, metadata consistency index Ss, metadata standardization index Sk, metadata update timeliness index Sj, and metadata access convenience index Sp in the data. Let Fp = A3×Sw + A4×Ss + A5×Sk + A6×Sj + A7×Sp, then the metadata completeness index Fp is obtained. The metadata completeness index Fp is compared with a preset metadata completeness index Fp0. Based on the comparison result, the metadata completeness is judged, and the data level adjustment process of the institutional data is adjusted according to the judgment result, wherein: When Fp > Fp0, the data grading module determines that the metadata completeness is normal and does not adjust the data grading process. When Fp≤Fp0, the data grading module determines that the metadata completeness is abnormal and adjusts the data grading process. The cold start coefficient l is adjusted by using the preset metadata completeness index Fp0 and the metadata completeness index Fp. The cold start coefficient l` is set to (Fp0 / Fp)×l to obtain the adjusted cold start coefficient l`. The cold start coefficient l is replaced with the adjusted cold start coefficient l`, and the preset data security requirement index range is readjusted according to the adjusted cold start coefficient l`. The data grading is then re-analyzed according to the adjusted preset data security requirement index range.

5. The cross-institutional privacy-preserving data system based on deep learning according to claim 4, characterized in that, The data scheme selection module analyzes the data compression and encryption combination schemes of the institutional data according to the data level, and obtains the data compression and encryption combination scheme of the institutional data, wherein: When the data is classified as top secret, the data scheme selection module uses the Zstandard compression algorithm to compress the data of the organization to obtain compressed data, and uses the AES-256-GCM encryption algorithm to encrypt the compressed data to obtain encrypted top secret data. When the data level is sensitive data, the data scheme selection module uses the Brotli compression algorithm to compress the data to obtain compressed data, and uses the ChaCha20-Poly1305 encryption algorithm to encrypt the compressed data to obtain encrypted sensitive data. When the data level is restricted, the data scheme selection module uses the LZ4 compression algorithm to compress the data to obtain compressed data, and uses the AES-128-CBC encryption algorithm to encrypt the compressed data to obtain encrypted restricted data. When the data level is public, the data scheme selection module uses the gzip compression algorithm to compress the data to obtain compressed public-level data. The data scheme selection module constructs a multi-objective optimization model according to a multi-objective optimization model construction method, which includes: Step B01: Calculate the comprehensive evaluation function value F based on security S, computational resource consumption C, transmission efficiency T, eighth weight coefficient ws, ninth weight coefficient wc, and tenth weight coefficient wt. Set F = ws × S + wc × (1 - C) + wt × T to obtain the comprehensive evaluation function value. Step B02: Set constraints to optimize the comprehensive evaluation function value to obtain a multi-objective optimization model; The data scheme selection module inputs security, computational resource consumption, transmission efficiency, eighth weight coefficient, ninth weight coefficient, and tenth weight coefficient into a multi-objective optimization model, outputs a comprehensive evaluation function value, and compares the comprehensive evaluation function value F with a preset comprehensive evaluation function F0. Based on the comparison result, the module determines the state of the comprehensive evaluation function value and selects a data compression and encryption combination scheme for the institutional data based on the determination result. When F≥F0, the data scheme selection module determines that the comprehensive evaluation function value is in a normal state, selects the data compression and encryption combination scheme, and sends the data compression and encryption combination scheme to the intelligent compression and encryption module. When F < F0, the data scheme selection module determines that the comprehensive evaluation function value is in an abnormal state and does not select the data compression and encryption combination scheme.

6. The cross-institutional privacy-preserving data system based on deep learning according to claim 5, characterized in that, The data scheme selection module calculates the data compression ratio Ys based on the original data size Dd and the compressed data size Yd in the characterization data, setting Ys = Dd / Yd to obtain the data compression ratio Ys. The data compression ratio Ys is then compared with a preset data compression ratio Ys0. Based on the comparison result, the data compression ratio is judged, and the data compression and encryption combination scheme is modified according to the judgment result. When Ys≥Ys0, the data scheme selection module determines that the data compression ratio is normal and does not modify the data compression and encryption combination scheme. When Ys < Ys0, the data scheme selection module determines that the data compression ratio is abnormal, corrects the data compression and encryption combination scheme, calibrates the eighth weight coefficient ws using the compression coefficient K, and sets... e is the base of the natural logarithm. The calibrated eighth weight coefficient ws` is obtained. ws` is set to K × ws, wc remains unchanged, and wt` = (1 - K) × wt. The calibrated comprehensive evaluation function value F1 is obtained. The calibrated comprehensive evaluation function value F1 is then compared again with the preset comprehensive evaluation function F0. The data scheme selection module compares the CPU utilization rate Lc in the data with the preset CPU utilization rate Lc0. Based on the comparison result, the CPU utilization is judged. Based on the judgment result, the process of correcting the data compression and encryption combination scheme and the adjustment process of data level adjustment are corrected. Where: When Lc≤Lc0, the data scheme selection module determines that the CPU utilization is normal and does not correct the process of data compression and encryption combination scheme modification and data level adjustment. When Lc > Lc0, the data scheme selection module determines that the CPU utilization is abnormal, corrects the process of modifying the data compression and encryption combination scheme and adjusting the data level, and corrects the preset compression ratio Ys0 by using the coefficient g. Let e ​​be the base of the natural logarithm. The corrected preset data compression ratio Ys0` is obtained. Ys0` is set to Ys0 / (0.47+g). The preset data compression ratio Ys0 is replaced with the corrected preset data compression ratio Ys0`. The data compression ratio Ys is then compared again with the corrected preset data compression ratio Ys0`. The data scheme selection module corrects the ninth weight coefficient wc using coefficient g, setting wc` = wc × g, to obtain the corrected ninth weight coefficient wc`. The data scheme selection module corrects the preset metadata perfection index Fp0 using coefficient g, setting Fp0` = g × Fp0, to obtain the corrected preset metadata perfection index Fp0`. The data classification adjustment process is corrected based on the corrected preset metadata perfection index.

7. The cross-institutional privacy-preserving data system based on deep learning according to claim 6, characterized in that, The dynamic privacy budget allocation module calculates the budget allocation ratio for the organization's data using a utility quantification method, which includes: Step C01: Adjust the encryption strength value based on the compression ratio Ys, the maximum compression ratio Ysmax, and the initial privacy budget. Line calculation, setting Obtain the privacy consumption value ; Step C02, based on the encryption key length K, the maximum security key length Kmax, the minimum security key length Kmin, and the total privacy budget. Privacy consumption value Perform calculations and set To obtain the encryption strength value ; Step C03, set the privacy consumption value Encryption strength value Initial privacy budget Total privacy budget The compression utility coefficient g1, the encryption utility coefficient g2, and the Lagrange multiplier W are used to calculate the budget allocation ratio using the Lagrange function, thus obtaining the budget allocation ratio. The dynamic privacy budget allocation module sends the budget allocation ratio to the intelligent compression and encryption module. The dynamic privacy budget allocation module compares the data encryption time Th with the preset data encryption time Th0, judges the data encryption time based on the comparison result, and adjusts the budget allocation ratio based on the judgment result. Wherein: When Th≤Th0, the dynamic privacy budget allocation module determines that the data encryption time consumption is normal and does not adjust the budget allocation ratio; When Th > Th0, the dynamic privacy budget allocation module determines that the data encryption time consumption is an abnormal situation and adjusts the budget allocation ratio. This adjustment is made using a time consumption coefficient h. The adjusted encryption strength value is obtained. The budget allocation ratio is optimized based on the adjusted encryption strength value. The dynamic privacy budget allocation module compares the data usage scenario security requirement index As with the preset data usage scenario security requirement index As0, judges the security status of the data usage scenario based on the comparison result, and calibrates the adjustment process of the budget allocation ratio, the data level, and the correction process of the data compression and encryption combination scheme based on the judgment result. When As≤As0, the dynamic privacy budget allocation module determines that the data usage scenario security is normal and does not calibrate the process of adjusting the budget allocation ratio, data level, and data compression and encryption combination scheme correction. When As > As0, the dynamic privacy budget allocation module determines that the data usage scenario security situation is abnormal, calibrates the adjustment process of the budget allocation ratio, the data level, and the data compression and encryption combination scheme correction process, and calibrates the preset data encryption time Th0 through the security factor v. Let e ​​be the base of the natural logarithm. The calibrated preset data encryption time Th0` is obtained. Th0` is set to v × Th0. The preset data encryption time Th0 is replaced with the calibrated preset data encryption time Th0`. The data encryption time Th is then compared again with the calibrated preset data encryption time Th0`. The dynamic privacy budget allocation module calibrates the preset data security requirement index interval Sq0 using the security coefficient v, obtaining the calibrated preset data security requirement index interval Sq0`. Sq0` is set to v × Sq0. The preset data security requirement index interval Sq0 is replaced with the calibrated preset data security requirement index interval Sq0`. A preset data security requirement index range Sq0` is established, and the data security requirement index Sq is re-compared with the calibrated preset data security requirement index range Sq0`. The dynamic privacy budget allocation module calibrates the preset data compression ratio Ys0 through the security coefficient v to obtain the calibrated preset data compression ratio Ys0`. Ys0` is set to Ys0 / (0.05+v). When Ys0` is greater than 1, it is set to equal 1. The preset data compression ratio Ys0 is replaced with the calibrated preset data compression ratio Ys0`, and the data compression ratio Ys is re-compared with the calibrated preset data compression ratio Ys0`.

8. The cross-institutional privacy-preserving data system based on deep learning according to claim 7, characterized in that, The intelligent compression and encryption module encrypts and compresses the organization's data according to the data compression and encryption combination scheme and the budget allocation ratio. It obtains the data compression and encryption combination scheme and the encryption strength value according to the data level, and calculates the privacy consumption value according to the budget allocation ratio. The module then compresses and encrypts the organization's data using the encryption strength value, the privacy consumption value, and the data compression and encryption combination scheme to obtain encrypted and compressed data.

9. The cross-institutional privacy-preserving data system based on deep learning according to claim 8, characterized in that, The cross-organizational data transmission module constructs the adversarial migration model according to the adversarial migration model construction method, obtains the adversarial migration model, inputs the encrypted compressed data into the adversarial migration model, outputs feature-aligned data, and transmits the feature-aligned data.

10. A method applied to a deep learning-based cross-agency privacy-preserving data system as described in any one of claims 1-9, characterized in that, The method includes: Step S1: Collect institutional data and characterization data; Step S2 involves analyzing the data level of the organization data, calibrating and adjusting the data level of the organization data based on the number of abnormal accesses in the characterization data, and adjusting the adjustment process of the data level of the organization data based on the metadata perfection index calculated from the characterization data. Step S3: Analyze the data compression and encryption combination scheme according to the data level, select the data compression and encryption combination scheme using a multi-objective optimization model based on the comprehensive evaluation function value calculated from the characterization data, and send the data compression and encryption combination scheme to the intelligent compression and encryption module. Also, correct the data compression and encryption combination scheme based on the data compression ratio calculated from the characterization data, and correct the correction process of the data compression and encryption combination scheme based on the CPU utilization in the characterization data. Step S4: Calculate the budget allocation ratio according to the utility quantification method to obtain the budget allocation ratio, and send the budget allocation ratio to the intelligent compression and encryption module. Also, calibrate the budget allocation ratio according to the data encryption time in the characterization data, and adjust the calibration process of the budget allocation ratio according to the data usage scenario security requirement index. Step S5: Encrypt and compress the organization data according to the data compression and encryption combination scheme and the budget allocation ratio to obtain encrypted and compressed data; Step S6: Using the adversarial migration model, output the feature alignment data based on the encrypted compressed data, and then transmit the feature alignment data.

Citation Information

Patent Citations

  • Image processing method and apparatus for achieving privacy protection

    CN111783803B