High-value data asset operation method of domestic alternative equipment in credential and credential field, electronic equipment and storage medium

By adopting domestically produced alternative equipment in the field of information technology innovation, and combining large-model retrieval data tagging, rule chain analysis, and a privacy computing layered isolation architecture, the problems of low efficiency in large-model retrieval and data privacy security in the field of information technology innovation have been solved, realizing the safe and efficient operation of high-value data assets and enhancing the competitiveness of the information technology innovation industry.

CN121458445APending Publication Date: 2026-02-03CHINA TELECOM DIGITAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511344011.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

In the current field of information technology innovation, large-scale model retrieval is inefficient, retrieval results are unlabeled, the coverage and accuracy of data retrieval question sets are insufficient, data privacy and security are difficult to guarantee, and the valuation of data assets is unscientific, which affects the effective operation of high-value data assets.

Method used

By adopting domestically produced alternative equipment, adding search data tags through large-scale model search conditions, and combining application relationships and custom rule events for filtering, verification, and analysis, a rule chain and result set are established. Data extraction is performed using the Python scripting language, and a domestically developed and secure privacy computing layered isolation architecture is adopted to ensure data privacy. Data value assessment and iterative screening are performed based on Pareto optimal solutions.

Benefits of technology

It significantly improves the efficiency of large model retrieval and the coverage and accuracy of data retrieval question sets, safeguards data privacy and security, ensures the high value and scientific operation of data assets, and enhances the core competitiveness in the field of information technology innovation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121458445A_ABST
    Figure CN121458445A_ABST
Patent Text Reader

Abstract

The invention provides a method for operating high-value data assets of domestic replacement equipment in the credential and credential field, electronic equipment and a storage medium, and relates to the field of credential and credential data asset security. The method comprises the following steps of: retrieving and collecting data by adopting a large model, and establishing a rule chain containing marks and a result set through an application relationship and a user-defined rule; extracting data by using a Python script in an aggregation data wide table form, and finishing data extraction analysis in combination with a large model; the data privacy is guaranteed by adopting a credit-creative domestic privacy calculation layered isolation architecture; according to the six value evaluation indexes, credible data are converted into multiple sets through a Pareto optimal solution, and data meeting all the indexes are screened out through multiple times of overlapping iteration and converted into high-value assets. According to the method, the data retrieval efficiency and security are improved, and scientific operation of high-value data assets is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data asset security in the field of domestic IT innovation, and in particular to a method, electronic device and storage medium for operating high-value data assets in the field of domestic alternative equipment. Background Technology

[0003] The development of the information technology application innovation (ITAI) industry still faces numerous technical challenges. On the hardware side, the performance of general-purpose CPUs is constrained by physical laws; integrated circuit density and clock frequency are approaching their limits, and multi-core architectures are limited by on-chip network communication efficiency. On the software side, the complexity of application algorithms is constantly increasing, such as machine learning and graph computing, which involve complex distributed data processing and place higher demands on node computing performance. Simultaneously, in the data convergence process of the existing ITAI environment, there are security bottlenecks when running large-scale application model queries. The large models themselves also suffer from insufficient security, unlabeled search results, and limited coverage of the query set. These problems seriously affect the effective operation of high-value data assets in the ITAI field. Summary of the Invention

[0004] Purpose of the invention: To propose a method, electronic device, and storage medium for the operation of high-value data assets in the field of information technology innovation using domestically produced alternative equipment. The aim is to solve the technical problems existing in the data asset operation process in the current information technology innovation field, such as low efficiency of large model retrieval, unlabeled retrieval results, insufficient coverage and accuracy of data query sets, difficulty in ensuring data privacy and security, and unscientific data asset value assessment, so as to achieve safe and efficient operation of high-value data assets in the information technology innovation field.

[0005] In a first aspect, this invention proposes a method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment, comprising the following steps:

[0006] S1. In the data center service network environment, a large model is used to search the range of collected data. By combining application relationships and custom rule events with the data center service data wide table, the data results of the search conditions are filtered, verified and analyzed to establish rule chains and result sets. The search conditions include search data tags.

[0007] S2. The result set of the rule chain is formed into a data set in the form of an aggregated data wide table, and the relevant data in the database is completely and comprehensively extracted using the Python scripting language;

[0008] S3. Privacy technology is used to protect the data privacy and security of the data retrieval process for unlabeled data in the rule chain result set. The privacy technology adopts a domestically developed and secure privacy computing layered isolation architecture.

[0009] S4. Trusted and secure data protected by privacy technology is transformed into N data sets based on six value assessment indicators: data integrity, data standardization, data accuracy, data timeliness, data richness, and data coverage. The Pareto optimal solution is used to store overlapping data between the sets into new data sets. After multiple overlapping iterations, the data that finally meets all N indicators is transformed into data assets.

[0010] In a further embodiment of the first aspect, step S1, which establishes the rule chain and result set, specifically includes:

[0011] Before training the large model, the search criteria are analyzed and labeled.

[0012] By combining application relationships and custom rule events with a wide table of data center service data, the data results of retrieval condition queries are filtered, verified, and analyzed to establish a rule chain.

[0013] The rule chain is adjusted and reconfigured according to the order of the indicators, and the marked query results are output to form a result set;

[0014] When the search condition rule chain is used as a parameter in the large model to retrieve data, the large model filters the data by matching the rule chain result set and puts the data that does not match into the rule chain result set.

[0015] In a further embodiment of the first aspect, after the complete and comprehensive extraction of relevant data from the database using the Python scripting language in step S2, a data analysis process is also included:

[0016] The question content is fed into the large model as a parameter, the results are obtained and stored in the results table;

[0017] Perform a join operation on the results table and the profile wide table to achieve distribution analysis and index analysis of the data retrieval process, and obtain the profile tag value for each question set;

[0018] The number of questions is calculated by aggregating different profile tag values ​​to complete the data analysis.

[0019] During the data acquisition and analysis process, the data matching the profile label features are combined with the Python language and a large model. The input table structure, dimensions, dimension values ​​and commonly used indicator data cover various combinations of indicators, dimensions and dimension values ​​under various conditions.

[0020] In a further embodiment of the first aspect, the domestically developed and secure privacy computing layered isolation architecture in step S3 includes a trusted execution environment layer and a privacy computing layer, wherein:

[0021] The Trusted Execution Environment (TEE) layer utilizes Zhaoxin's ZX-TCT trusted computing technology to construct a trust chain based on the Zhaoxin CPU during system startup and runtime, passing the trust relationship up from the Zhaoxin CPU's embedded instructions to the OS layer; it extends the Linux operating system, enables the IMA integrity measurement function, and extends the trust chain to the privacy computing framework layer; all integrity measurement values ​​are stored in the domestically produced TPM2.0 hardware device; and a trusted execution environment is constructed through remote authentication technology.

[0022] In the privacy computation layer, all execution modules of the platform, user-defined operators, and user data are maintained in a trust chain. The platform completes privacy computation through cryptographic techniques such as homomorphic encryption, differential privacy, and unintended transmission, as well as blockchain technologies such as smart contracts and consortium blockchains.

[0023] In a further embodiment of the first aspect, the specific method for data processing using the Pareto optimal solution in step S4 is as follows:

[0024] Construct an objective function f, mapping the n-dimensional decision space Ω to a k-dimensional objective space. The objective function includes six optimization objectives: data integrity, data normalization, data accuracy, data timeliness, data richness, and data coverage. The decision space covers the entire variable space. The expression of the objective function is:

[0025] minf(x)=(f1(x),f2(x),…,f6(x)),x∈Ω

[0026] Among them, f1(x) corresponds to the data integrity index, f2(x) corresponds to the data standardization index, f3(x) corresponds to the data accuracy index, f4(x) corresponds to the data timeliness index, f5(x) corresponds to the data richness index, and f6(x) corresponds to the data coverage index.

[0027] Based on the Pareto dominance rule, the decision vectors are compared. If the decision vector x is not greater than and is less than the decision vector y on any target and at least on one target, then x dominates y.

[0028] Decision vectors that are not dominated by any vector in the decision space are selected to form a Pareto optimal solution set. Based on this solution set, the trusted and secure data is transformed into 6 data sets.

[0029] In a further embodiment of the first aspect, the rule chain result set is stored in the data table demo_userprofilecrowd, which includes fields id, crowdname, calculate_label, and userprofile_date. The calculate_label field records the list of profile labels to be analyzed in the form of a JSON array. The profile analysis results are stored in the data table demo_userprofile_crowd_overview, which includes fields id, crowdid, labelname, overview_result, and userprofile_date. The overview_result field stores the profile analysis results associated with the Socket label, and the userprofile_date field records the date of the data used for the profile analysis.

[0030] In a further embodiment of the first aspect, in the trusted execution environment layer, remote authentication technology generates a proof report through a remote authentication service, and the management end verifies the proof report through PrivateCA and Trust Verifier to ensure the security of the trusted execution environment.

[0031] In a further embodiment of the first aspect, during multiple overlapping iterations of the data set, the data in the new data set is re-verified according to the six value assessment indicators after each iteration. If there is any data that does not meet the indicators, it is removed until all the data in the new data set meets the six indicators.

[0032] In a second aspect, the present invention provides an electronic device comprising: a processor and a memory storing computer program instructions; wherein the processor, when executing the computer program instructions, implements the method described in the first aspect for operating high-value data assets of domestically produced alternative equipment in the field of information technology innovation.

[0033] A third aspect of the present invention provides a computer-readable storage medium storing at least one executable instruction, which, when executed on an electronic device, causes the electronic device to perform the method described in the first aspect for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment.

[0034] Compared with the prior art, the present invention has at least the following beneficial effects:

[0035] (1) In the process of large model retrieval, the present invention adds retrieval data tags to the retrieval conditions, which effectively makes up for the lack of tags in the existing large model retrieval data results. At the same time, by combining application relationships, custom rule events and wide tables of data center service data for filtering and verification analysis, the efficiency of large model retrieval is significantly improved.

[0036] (2) By forming a data set from the result set of the rule chain into an aggregated wide table, and using the Python scripting language to completely extract the database data, a complete input is provided for the generation of the data question set. During the data analysis process, the combination of Python language and large model input of various data information covers various indicators, dimensions and combinations of dimension values, which greatly improves the coverage, accuracy and practicality of the data question set.

[0037] (3) Adopting a domestically developed and secure privacy computing layered isolation architecture, data privacy security is protected from both the trusted execution environment layer and the privacy computing layer. This not only meets the localization requirements in the field of domestic IT innovation, but also effectively reduces the risk of data leakage and ensures the security of data during the data retrieval process through a variety of encryption technologies and blockchain technology.

[0038] (4) Based on six indicators of data integrity, standardization, accuracy, timeliness, richness and coverage, and combined with Pareto optimal solution, trusted and secure data are processed and iteratively screened to ensure that the final transformed data assets have high value. This provides an effective method for the scientific operation of data assets in the field of information technology innovation and helps to enhance the core competitiveness of enterprises in the field of information technology innovation. Attached Figure Description

[0039] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0040] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention can be practiced without one or more of these details. In other instances, certain technical features well-known in the art have not been described in order to avoid obscuring the invention.

[0041] This embodiment discloses a method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment. For details, please refer to [link to specific steps]. Figure 1 :

[0042] 1. Large-scale model retrieval and rule chain / result set establishment

[0043] In a data center service network environment, a large model is used to search the collected data range. Unlike existing large model retrieval methods, this invention adds search data tags to the search conditions to compensate for the lack of tags in the data results of large model retrieval. By combining application relationships and custom rule events with a wide table of data center service data, the data results queried based on the search conditions are filtered, verified, and analyzed, thereby establishing rule chains and result sets.

[0044] The specific process is as follows: Before the large model training begins, the search conditions are analyzed and marked; then, with the help of application relationships and custom rule events, and combined with the wide table of data center service data, the data results of the search condition query are filtered, verified and analyzed to construct a rule chain; the rule chain is adjusted and reconfigured in terms of indicator order, and the marked query results are output to form a result set; when the search condition rule chain is put into the large model as a parameter for data retrieval, the large model will match the retrieved data with the rule chain result set for filtering, and the data that does not match will be put into the rule chain result set.

[0045] 2. Data extraction and analysis

[0046] The result set of the rule chain is organized into a data set in the form of an aggregated wide table. The relevant data in the database is extracted completely and comprehensively using the Python scripting language, providing complete input for the subsequent generation of the data query set.

[0047] After data extraction, data retrieval analysis is performed: the question content is fed into the large model as parameters, the results are obtained and stored in the results table; the results table and the profile wide table are joined to perform distribution analysis and indicator analysis of the data retrieval process, thereby obtaining the profile tag value for each question set; the number of questions corresponding to different profile tag values ​​is aggregated and calculated to complete the data retrieval analysis. During the data retrieval analysis, data matching profile tag features is combined with the large model using Python language, inputting table structure, dimensions, dimension values, and commonly used indicator data to cover various combinations of indicators, dimensions, and dimension values, thereby improving the coverage, accuracy, and practicality of the retrieved question sets.

[0048] The rule chain result set is stored in the data table demo_userprofilecrowd, which contains fields id, crowdname, calculate_label, and userprofile_date. The calculate_label field records the list of profile labels to be analyzed via a JSON array. The profile analysis results are stored in the data table demo_userprofile_crowd_overview, which includes fields id, crowdid, labelname, overview_result, and userprofile_date. The overview_result field stores the profile analysis results associated with the Socket label, and the userprofile_date field records the date of the data used for the profile analysis.

[0049] 3. Data privacy and security protection

[0050] For unlabeled data in the rule chain result set, a domestically developed and secure privacy-preserving computation layered isolation architecture is adopted to ensure data privacy and security during the data retrieval process. This architecture includes a trusted execution environment layer and a privacy computation layer.

[0051] The Trusted Execution Environment (TEE) layer utilizes Zhaoxin's ZX-TCT trusted computing technology to construct a trust chain based on the Zhaoxin CPU during system startup and runtime, passing the trust relationship upwards from the Zhaoxin CPU's embedded instructions to the OS layer. It extends the Linux operating system by enabling the IMA integrity measurement function, extending the trust chain to the privacy computing framework layer. All integrity measurement values ​​are stored within domestically produced TPM2.0 hardware devices, ensuring their trustworthiness and immutability. The TEE is constructed through remote authentication technology, which generates a proof report via a remote authentication service. The management end verifies the proof report using PrivateCA and a TrustVerifier, further ensuring the security of the TEE.

[0052] In the privacy computing layer, all execution modules of the platform, user-defined operators, and user data are maintained in a trust chain to ensure their integrity and immutability. The platform uses cryptographic techniques such as homomorphic encryption, differential privacy, and unintended transmission, as well as blockchain technologies such as smart contracts and consortium blockchains, to complete privacy computing and realize the mining of data value.

[0053] 4. Data asset transformation

[0054] Trustworthy and secure data protected by privacy technology is transformed into six data sets based on six value assessment indicators: data integrity, data standardization, data accuracy, data timeliness, data richness, and data coverage. The Pareto optimal solution is used to store overlapping data from these sets into new data sets. After multiple overlapping iterations, the data that finally meets all six indicators is transformed into data assets.

[0055] The specific method for data processing using Pareto optimal solutions is as follows: An objective function f is constructed, mapping the n-dimensional decision space Ω to a k-dimensional objective space. The objective function contains the six optimization objectives mentioned above, and the decision space covers the entire variable space. The objective function expression is minf(x)=(f1(x),f2(x),…,f6(x)), x∈Ω, where f1(x) corresponds to the data integrity index, f2(x) corresponds to the data standardization index, f3(x) corresponds to the data accuracy index, f4(x) corresponds to the data timeliness index, f5(x) corresponds to the data richness index, and f6(x) corresponds to the data coverage index. Based on the Pareto dominance rule, decision vectors are compared. If decision vector x is not greater than and at least less than decision vector y on any objective, then x dominates y. Decision vectors not dominated by any vector in the decision space are selected to form a Pareto optimal solution set. Based on this solution set, the reliable and secure data is transformed into six data sets.

[0056] During multiple overlapping iterations of the dataset, the data in the new dataset is re-verified according to the six value assessment indicators after each iteration. If any data does not meet the indicators, it is removed until all the data in the new dataset meets the six indicators.

[0057] The technical solution of the present invention will be further described in detail below with reference to specific embodiments.

[0058] I. Environmental Preparation

[0059] A data center service network environment that meets the requirements of domestic IT innovation was built. The hardware uses domestically produced servers based on Zhaoxin CPUs and is equipped with domestically produced TPM2.0 hardware devices. On the software side, a Linux operating system was installed, and the IMA integrity measurement function was enabled. Components related to the trusted execution environment based on Zhaoxin ZX-TCT trusted computing technology were deployed, as well as a privacy computing platform that includes cryptographic technologies such as homomorphic encryption, differential privacy, and unintended transmission, and blockchain technologies such as smart contracts and consortium blockchains. The database uses a domestically produced MySQL database to store data such as rule chain result sets and profile analysis results.

[0060] II. Large-scale model retrieval and rule chain / result set establishment

[0061] 1. Select a large model suitable for data retrieval. Before training the large model, analyze the retrieval conditions (such as data type, data generation time, data source, etc.) based on the characteristics of the data in the data center and the retrieval requirements. Add a unique retrieval data tag to each retrieval condition, such as "Data type_Business system A_2024" and "Data generation time_202405_Business system B".

[0062] 2. Analyze the application relationships in the data center services, such as the data interaction relationships between business systems and the data processing flow relationships. At the same time, formulate custom rule events based on actual business needs, such as "only retain valid business data generated in the last 3 months" and "remove data that does not conform to the standard format".

[0063] 3. Construct a wide table for data center services. This wide table integrates relevant data from various business systems in the data center services, including basic data information (data identifier, data type, generation time, source system, etc.) and business attribute information (business number, business status, associated users, etc.).

[0064] 4. Input the analyzed and labeled search criteria into the large model. The large model, based on application relationships and custom rule events, combined with the data center service data wide table, filters, verifies, and analyzes the data results queried based on the search criteria. For example, for the search criteria labeled "Data Type_Business System A_2024", the large model filters out data related to Business System A based on application relationships, and then, based on the custom rule event "Only retain valid business data generated in the last 3 months", removes data generated more than 3 months ago or invalid data, thereby establishing a rule chain.

[0065] 5. Adjust and reconfigure the order of indicators in the established rule chain. For example, move the data validity verification indicator to the front of the rule chain to prioritize the removal of invalid data and improve the efficiency of subsequent processing. After configuration, output the marked query results to form a result set.

[0066] 6. Use the search condition rule chain as a parameter to search the data again in the large model. When the large model searches the data, it filters the result set of the rule chain and puts the data that does not match successfully (such as data that does not meet any rule in the rule chain) into the rule chain result set to improve the result set data.

[0067] III. Data Extraction and Data Analysis

[0068] 1. The rule chain result set is organized into a data set in the form of an aggregated wide table. This aggregated wide table contains complete information about all data in the rule chain result set and is organized in a format that facilitates subsequent processing.

[0069] 2. Write a Python script to connect to a domestic MySQL database using a Python database connection library (such as pymysql). Use the script to extract complete and comprehensive data related to the aggregated wide table in the database. The extracted data includes basic data information, business attribute information, and tag information.

[0070] 3. Design the data retrieval queries, such as "Query the number of data for different business statuses in business system A in the past month" or "Statistics on the number of business transactions for each associated user in business system B". These queries are used as parameters and fed into the large model. The large model processes each query and generates results. These results are stored in a result table, which includes fields such as query identifier, query content, large model output results, and generation time.

[0071] 4. Construct a wide profile table. This wide table is based on business data to build user profiles, business profiles and other related profile information. For example, the user profile wide table includes user identifier, basic user information (name, age, gender, region, etc.), and user behavior information (business processing records, data access records, etc.). The business profile wide table includes business identifier, business type, business scale, business cycle and other information.

[0072] 5. Use SQL statements to perform join operations on the result table and the profile wide table. For example, join the result table for "the number of data for different business statuses of business system A in the past month" with the business profile information of business system A in the business profile wide table to obtain the profile tag value corresponding to each query set, such as the profile tag value for business status "completed" and the profile tag value for business status "processing".

[0073] 6. Based on different profile tag values, use aggregation functions (such as the count function) to calculate the corresponding number of questions. For example, count the number of questions with a business status of "completed" and the number of questions with a business status of "processing", and complete the data collection and analysis process.

[0074] 7. During the data retrieval and analysis process, data matching profile tag features (such as data matching the "age 25-35 years old" tag feature in user profiles and data matching the "large business scale" tag feature in business profiles) are combined with a large model using Python. The large model is then fed with table structures (such as the field structure of the result table and the profile wide table), dimensions (such as the time dimension, business status dimension, and user age dimension), dimension values ​​(such as "last month" and "last 3 months" in the time dimension, and "completed" and "processing" in the business status dimension), and commonly used indicator data (such as the number of data, the number of business transactions, and the amount). This allows the large model to cover various combinations of indicators, dimensions, and dimension values ​​under different circumstances, further improving the coverage, accuracy, and practicality of the data retrieval question set.

[0075] Meanwhile, the rule chain result set is stored in the `demo_userprofilecrowd` table of a domestic MySQL database. The `id` field in this table is a unique identifier, the `crowdname` field records the result set name, the `calculate_label` field records the list of profile labels to be analyzed (e.g., ["Business Status", "User Age", "Business Scale"]) via a JSON array, and the `userprofile_date` field records the data processing date. The profile analysis results are stored in the `demo_userprofile_crowd_overview` table. The `id` field is a unique identifier, the `crowdid` field is associated with the `id` field of the `demo_userprofilecrowd` table, the `labelname` field records the profile label name (e.g., "Business Status" and "User Age"), the `overview_result` field stores the profile analysis results associated with the associated Socket labels (e.g., the number of users with a "Completed" business status and the number of users aged 25-35), and the `userprofile_date` field records the date of the data used for the profile analysis.

[0076] IV. Data Privacy and Security Protection

[0077] 1. Construction of Trusted Execution Environment Layer

[0078] By utilizing Zhaoxin's ZX-TCT trusted computing technology, when a domestic server based on Zhaoxin CPU starts up, it sequentially performs trust measurements on the Bootloader, OS kernel, and OS applications, starting from the Zhaoxin CPU's embedded instructions, to build a trust chain based on the Zhaoxin CPU during system startup, and passes the trust relationship from the Zhaoxin CPU's embedded instructions up to the OS layer.

[0079] Extend the Linux operating system by integrating the IMA integrity measurement module into the operating system kernel, enabling the IMA integrity measurement function to perform integrity measurements on applications, configuration files, etc. in the operating system, and further extending the trust chain from the OS layer to the privacy computing framework layer.

[0080] All integrity metrics (including metrics during system startup and IMA integrity metrics) are stored in the domestically produced TPM2.0 hardware device. The TPM2.0 device uses hardware encryption to ensure the reliability and immutability of the metrics.

[0081] Deploy a remote authentication service, which is based on Zhaoxin ZX-TCT trusted computing technology and TPM2.0 devices, to generate a certificate report containing system integrity information. The management end authenticates the remote authentication service using a certificate generated by PrivateCA (Private Certificate Authority), and at the same time, uses a TrustVerifier to compare the integrity metric in the certificate report with the preset "GoodKnownWhitelistTable" to verify the integrity and trustworthiness of the system, thereby building a controllable trusted execution environment and realizing trusted computing for upper-layer privacy computing applications.

[0082] 2. Privacy Computation Layer Operation

[0083] All execution modules of the privacy computing platform (such as data encryption modules, data decryption modules, model training modules, etc.), user-defined operators based on business needs (such as custom data cleaning operators, data statistical calculation operators, etc.), and user's original data are included in the protection scope of the trust chain. The integrity of these modules, operators, and data is monitored and protected in real time through the IMA integrity measurement function and TPM2.0 devices to ensure their integrity and immutability.

[0084] When privacy-preserving computations are required, appropriate technologies should be selected based on the data processing needs: For scenarios requiring data-sharing computations but where the original data cannot be leaked, homomorphic encryption is used to encrypt the data before sending it to participating nodes. These nodes then perform computations directly on the encrypted data, and the final result is obtained after decryption. For scenarios requiring data privacy protection while allowing for a certain amount of data error, differential privacy technology is used to add appropriate noise to the dataset, preventing attackers from inferring specific individual data from the dataset. For scenarios requiring data exchange between two or more parties without leaking their respective data, inadvertent transmission technology is used to achieve secure data exchange.

[0085] Simultaneously, a privacy computing consortium is constructed using consortium blockchain technology, bringing together participating nodes (such as different business departments or enterprises) to the consortium blockchain. Smart contracts define the rights and obligations of each party, data computation rules, and result sharing methods. During the privacy computing process, all computational operations and data flow information are recorded on the consortium blockchain. The decentralized and immutable nature of the consortium blockchain ensures the transparency and traceability of the computation process, further guaranteeing data privacy and security, and the credibility of the computation results.

[0086] By operating the privacy computing layer, while avoiding the risk of data leakage, it assists multiple parties in outputting models suitable for their own businesses. This not only reduces the additional costs caused by violating data-related laws and regulations and data leakage, but also reduces the operating costs of marketing activities and other businesses through accurate models.

[0087] V. Data Asset Transformation

[0088] 1. Determine the valuation indicators and calculation methods

[0089] (1) Data integrity: The proportion of missing data in a statistical dataset. The lower the proportion of missing data, the higher the data integrity. The calculation formula is: Data integrity = (1 - number of missing data entries / total number of data entries) × 100%.

[0090] (2) Data Standardization: Evaluated from two aspects: data format and logical rationality. Regarding data format, the percentage of data that conforms to the statistical format standard; regarding logical rationality, the percentage of data whose statistical data item values ​​and the logical relationships between data items are reasonable. Data Standardization = (Number of data items conforming to the format standard + Number of data items with reasonable logic) / (2 × Total number of data items) × 100%.

[0091] (3) Data accuracy: First, verify the reliability of the data source. For data from reliable sources, calculate the proportion of error-free and abnormal data. Data accuracy = Number of error-free and abnormal data entries / Number of data entries from reliable sources × 100%. Abnormal data is identified through data analysis and business scenario knowledge, such as data that significantly deviates from a normal distribution or data that does not conform to business logic.

[0092] (4) Data Timeliness: Set a time decay factor according to the data type. For example, the time decay factor for user consumption data is set to 0.1 per month. Data Timeliness = Initial Data Value × (1 - Time Decay Factor × Number of Months for Data Storage). The initial data value is set to a value between 0 and 1 according to the importance of the data.

[0093] (5) Data richness: The proportion of valuable information items in the statistical data assets out of the total number of information items. It also considers the correlation and distinctiveness between data items, and reduces information for redundant data items. Data richness = (Number of valuable information items - Number of redundant information items) / Total number of information items × 100%.

[0094] (6) Data Coverage: The proportion of sample data covered by statistical data assets. Data Coverage = Number of covered sample data entries / Total number of sample data entries × 100%.

[0095] 2. Application of Pareto optimal solution

[0096] Construct an objective function f(x) = (f1(x), f2(x), ..., f6(x)), x ∈ Ω, where x represents each data point in the trusted and secure dataset, Ω represents the trusted and secure dataset (decision space), f1(x) corresponds to the calculated value of the data integrity index, f2(x) corresponds to the calculated value of the data standardization index, f3(x) corresponds to the calculated value of the data accuracy index, f4(x) corresponds to the calculated value of the data timeliness index, f5(x) corresponds to the calculated value of the data richness index, and f6(x) corresponds to the calculated value of the data coverage index. The objective function is minf(x), that is, to seek the data with the optimal calculated values ​​of each index (data integrity, standardization, accuracy, timeliness, richness, and coverage as high as possible).

[0097] The Pareto dominance rule is used to compare each data point (decision vector) in the dataset. For example, for data x and data y, if f1(x)≥f1(y), f2(x)≥f2(y), f3(x)≥f3(y), f4(x)≥f4(y), f5(x)≥f5(y), f6(x)≥f6(y), and at least one index satisfies f1(x)≥f2(y), then f1(x)≥f2(y), f3(x)≥f3(y), f4(x)≥f4(y), f5(x)≥f5(y), f6(x)≥f6(y), and at least one index satisfies f1(x)≥f2(y). i (x)>f i If (y)(i=1,2,…,6), then x dominates y.

[0098] Data that is not dominated by any other data is selected to form a Pareto optimal solution set. Based on this solution set, the trustworthy and secure data are transformed into six data sets according to six indicators: the optimal set of data integrity, the optimal set of data standardization, the optimal set of data accuracy, the optimal set of data timeliness, the optimal set of data richness, and the optimal set of data coverage.

[0099] 3. Overlapping Iteration of Data Sets and Transformation of Data Assets

[0100] Identify overlapping data among the six datasets and store this overlapping data in a new dataset. Then, validate the data in the new dataset again against the six value assessment indicators. If any data fails to meet the preset standards (e.g., data completeness below 90%, data accuracy below 85%), remove it.

[0101] Repeat the overlapping data screening and verification process described above. After multiple overlapping iterations, the data in the new data set meets all the preset standards of the six indicators. At this point, the data in the new data set is transformed into high-value data assets for use in business decision-making, model training, service optimization, and other scenarios in the field of information technology innovation.

[0102] By implementing this embodiment, many technical problems in the operation of data assets in the existing information technology innovation field can be effectively solved, enabling the safe and efficient operation of high-value data assets and promoting the further development of the information technology innovation industry.

[0103] The logical ideas behind the methods disclosed in the above embodiments can be implemented in whole or in part through software, hardware, firmware or other arbitrary combinations.

[0104] When implemented in hardware, a hardware architecture can be constructed that employs domestically developed, secure privacy-preserving computation layers for isolation, reducing computational complexity and ensuring the security of privacy-preserving data computation. This hardware architecture consists of two parts: a trusted execution environment layer and a privacy-preserving computation layer.

[0105] Trusted Execution Environment Layer: Utilizing Zhaoxin ZX-TCT trusted computing technology, a trust chain based on the Zhaoxin CPU is constructed during system startup and runtime, passing the trust relationship upwards from the Zhaoxin CPU's embedded instructions all the way to the OS layer. Simultaneously, by extending the Linux operating system and enabling the IMA integrity measurement function, the trust chain is further extended to the application layer, i.e., the privacy computing framework layer. Furthermore, all integrity measurement values ​​are stored within domestically produced TPM2.0 hardware devices to ensure their trustworthiness and immutability. Finally, through remote authentication technology, a more controllable trusted execution environment is constructed, enabling trusted computing for upper-layer applications.

[0106] Privacy Computation Layer: All execution modules of the platform, user-defined operators, and user data are maintained in a trust chain to ensure their integrity and immutability. The platform utilizes cryptographic techniques such as homomorphic encryption, differential privacy, and unintended transmission, as well as blockchain technologies such as smart contracts and consortium blockchains, to perform privacy computation and realize the mining of data value.

[0107] The practical application of privacy computing based on domestically produced chips proves that domestically produced chips can perfectly meet the hardware requirements of privacy computing, validating the feasibility of the privacy computing industry breaking away from its dependence on foreign chips. It also demonstrates that, while avoiding the risk of data leakage, privacy computing can assist multiple parties in developing models suitable for their own businesses. This not only helps clients reduce additional costs incurred due to violations of data-related laws and regulations and data breaches, but also reduces the operational costs of marketing campaigns through accurate models.

[0108] When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions or computer programs.

[0109] When computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions according to the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives (SSDs).

[0110] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification. Although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the invention itself. Various changes in form and detail may be made without departing from the spirit and scope of the invention as defined in the appended claims.

Claims

1. A method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment, characterized in that, Includes the following steps: S1. In the data center service network environment, a large model is used to search the range of collected data. By combining application relationships and custom rule events with the data center service data wide table, the data results of the search conditions are filtered, verified and analyzed to establish rule chains and result sets. The search conditions include search data tags. S2. The result set of the rule chain is formed into a data set in the form of an aggregated data wide table, and the relevant data in the database is completely and comprehensively extracted using the Python scripting language; S3. Privacy technology is used to protect the data privacy and security of the data retrieval process for unlabeled data in the rule chain result set. The privacy technology adopts a domestically developed and secure privacy computing layered isolation architecture. S4. Trusted and secure data protected by privacy technology is transformed into N data sets based on six value assessment indicators: data integrity, data standardization, data accuracy, data timeliness, data richness, and data coverage. The Pareto optimal solution is used to store overlapping data between the sets into new data sets. After multiple overlapping iterations, the data that finally meets all N indicators is transformed into data assets.

2. The method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment as described in claim 1, characterized in that, Step S1 involves establishing the rule chain and result set, specifically including: Before training the large model, the search criteria are analyzed and labeled. By combining application relationships and custom rule events with a wide table of data center service data, the data results of retrieval condition queries are filtered, verified, and analyzed to establish a rule chain. The rule chain is adjusted and reconfigured to change the order of the indicators, and the marked query results are output to form a result set. When the search condition rule chain is used as a parameter in the large model to retrieve data, the large model filters the data by matching the rule chain result set and puts the data that does not match into the rule chain result set.

3. The method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment as described in claim 1, characterized in that, After extracting the relevant data from the database completely and comprehensively using the Python scripting language as described in step S2, the process also includes data analysis: The question content is fed into the large model as a parameter, the results are obtained and stored in the results table; Perform a join operation on the results table and the profile wide table to achieve distribution analysis and index analysis of the data retrieval process, and obtain the profile tag value for each question set; The number of questions is calculated by aggregating different profile tag values ​​to complete the data analysis. During the data acquisition and analysis process, the data matching the profile label features are combined with the Python language and a large model. The input table structure, dimensions, dimension values ​​and commonly used indicator data are used to cover various combinations of indicators, dimensions and dimension values.

4. The method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment as described in claim 1, characterized in that, Step S3 of the domestically developed secure privacy computing layered isolation architecture includes a trusted execution environment layer and a privacy computing layer, wherein: The Trusted Execution Environment (TEE) layer utilizes Zhaoxin's ZX-TCT trusted computing technology to construct a trust chain based on the Zhaoxin CPU during system startup and runtime, passing the trust relationship up from the Zhaoxin CPU's embedded instructions to the OS layer; it extends the Linux operating system, enables the IMA integrity measurement function, and extends the trust chain to the privacy computing framework layer; all integrity measurement values ​​are stored in the domestically produced TPM2.0 hardware device; and a trusted execution environment is constructed through remote authentication technology. In the privacy computation layer, all execution modules of the platform, user-defined operators, and user data are maintained in a trust chain. The platform completes privacy computation through cryptographic techniques such as homomorphic encryption, differential privacy, and unintended transmission, as well as blockchain technologies such as smart contracts and consortium blockchains.

5. The method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment according to claim 1, characterized in that, The specific method for data processing using the Pareto optimal solution in step S4 is as follows: Construct an objective function f, mapping the n-dimensional decision space Ω to a k-dimensional objective space. The objective function includes six optimization objectives: data integrity, data normalization, data accuracy, data timeliness, data richness, and data coverage. The decision space covers the entire variable space. The expression of the objective function is: minf(x)=(f1(x),f2(x),…,f6(x)),x∈Ω Among them, f1(x) corresponds to the data integrity index, f2(x) corresponds to the data standardization index, f3(x) corresponds to the data accuracy index, f4(x) corresponds to the data timeliness index, f5(x) corresponds to the data richness index, and f6(x) corresponds to the data coverage index. Based on the Pareto dominance rule, the decision vectors are compared. If the decision vector x is not greater than and is less than the decision vector y on any target and at least on one target, then x dominates y. Decision vectors that are not dominated by any vector in the decision space are selected to form a Pareto optimal solution set. Based on this solution set, the trusted and secure data is transformed into 6 data sets.

6. The method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment as described in claim 3, characterized in that, The rule chain result set is stored in the data table demo_userprofilecrowd, which includes fields id, crowdname, calculate_label, and userprofile_date. The calculate_label field records the list of profile labels to be analyzed in JSON array format. The profile analysis results are stored in the data table demo_userprofile_crowd_overview, which includes fields id, crowdid, labelname, overview_result, and userprofile_date. The overview_result field stores the profile analysis results associated with the Socket label, and the userprofile_date field records the date of the data used for profile analysis.

7. The method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment as described in claim 4, characterized in that, In the Trusted Execution Environment (TEE) layer, remote authentication technology generates a proof report through a remote authentication service. The management end verifies the proof report through PrivateCA and Trust Verifier to ensure the security of the TEE.

8. The method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment as described in claim 5, characterized in that, During multiple overlapping iterations of the dataset, the data in the new dataset is re-verified according to the six value assessment indicators after each iteration. If any data does not meet the indicators, it is removed until all the data in the new dataset meets the six indicators.

9. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; the processor executes the computer program instructions to implement a method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one executable instruction, which, when executed on an electronic device, causes the electronic device to perform a method for operating high-value data assets in the field of information technology innovation using domestically produced alternative equipment as described in any one of claims 1 to 8.