Classification method, system and equipment for outputting uplink transmission data and medium

By acquiring, preprocessing, clustering, and storing data on the blockchain, the problem of blockchain data silos and secure sharing is solved, data quality and utilization are improved, data security and compliance are ensured, and cross-chain data analysis and application are realized.

CN120670520APending Publication Date: 2025-09-19YUNNAN POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510529359.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Blockchain data has problems such as data silos, uneven data quality, low data utilization, high security risks and increased regulatory compliance requirements, making data sharing across departments, regions and businesses difficult to achieve.

Method used

By obtaining raw data from the data source, performing preprocessing and diversified proofreading, formulating data classification rules, using clustering to identify features, generating a classification structure containing classification labels and timestamps, and using blockchain technology for on-chain storage and verification to generate a unique index.

Benefits of technology

It improves data quality and utilization, ensures data security and compliance, realizes cross-chain data analysis and application, and improves the efficiency and credibility of data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670520A_ABST
    Figure CN120670520A_ABST
Patent Text Reader

Abstract

The invention discloses a classification method, system and device for outputting uplink transmission data and a medium, and belongs to the technical field of data classification and block chain data, and the method comprises the steps: obtaining original data in an uplink transmission data list from a data source, carrying out the preprocessing and diversified proofreading of the original data, and obtaining a data list; comprising abnormal value removal, missing value supplementation and error value correction; formulating a data classification rule according to business requirements, and identifying data classification features by adopting a clustering mode; based on the classification features, data classification calibration is carried out, and a classification structure containing classification tags, original data features and timestamps is generated; and performing uplink verification and classified uplink storage on the classified calibration data, and generating a unique index. Through automatic data collection, a multi-dimensional classification rule, a high-precision classification algorithm and block chain cochain verification, the data processing efficiency and accuracy are improved, and the authenticity, integrity and traceability of classified data are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data classification and blockchain data technology, and in particular to a classification method, system, device and medium for outputting on-chain transmission data. Background Art

[0002] Driven by the goals of carbon peak and carbon neutrality, data transmission and utilization require cross-departmental, cross-regional, and cross-business data sharing. With the rapid development of blockchain technology, various blockchain platforms and applications have emerged, generating massive amounts of on-chain data. This data holds enormous value, but also faces issues such as data silos, uneven data quality, and difficulty in effective data utilization. To better unlock the value of blockchain data, it is crucial to effectively classify and manage it. The difficulties and pain points of blockchain data classification are primarily reflected in the following aspects:

[0003] The volume and diversity of data are exploding: Blockchains store a wide range of data, including transaction records, smart contract code, on-chain asset information, and user behavior data. This volume is growing exponentially. Traditional data management methods struggle to cope with this vast and complex data.

[0004] The data value is huge, but the utilization rate is low: Blockchain data contains valuable information such as transaction behavior, user portraits, and market trends. However, due to the lack of effective classification and governance, the value of this data is difficult to be fully explored and utilized.

[0005] The phenomenon of data silos is serious: the data formats between different blockchain platforms are not unified, and there is a lack of interoperability, resulting in serious data silos and making cross-chain data analysis and application difficult.

[0006] The need for data security and privacy protection is growing: Blockchain data is open, transparent, and tamper-proof, but it also faces security risks such as data leakage and privacy infringement. Data needs to be classified and graded, and appropriate security measures must be taken.

[0007] Increasing regulatory compliance requirements: As the application of blockchain technology continues to expand, governments around the world are increasingly strengthening their oversight of blockchain data. Data classification is necessary to meet regulatory compliance requirements in different countries and regions.

[0008] In summary, as the open sharing of on-chain data places increasing demands on security classification and assurance technologies, there is an urgent need to address the issue of secure data sharing across departments, regions, and businesses. This requires breakthroughs in key data classification technologies, breaking down data silos, improving data quality, unlocking data value, ensuring data security, and meeting regulatory compliance requirements. This will promote the healthy development and application of blockchain technology, provide effective support for data security governance, and enhance the security management and control of on-chain data. Summary of the Invention

[0009] In view of the above-mentioned problems, the present invention is proposed.

[0010] Therefore, the technical problem solved by the present invention is: how to break through the key technology of data classification, break data silos, improve data quality, mine data value, and ensure data security.

[0011] To solve the above technical problems, the present invention provides the following technical solutions: a classification method for outputting on-chain transmission data, which comprises the following steps: obtaining the original data in the on-chain transmission data list from the data source, preprocessing and diversified proofreading the original data, including outlier removal, missing value supplementation and error value correction; formulating data classification rules according to business needs, and using clustering to identify data classification features; based on the classification features, performing data classification calibration, and generating a classification structure including classification labels, original data features and timestamps; performing on-chain verification and classified on-chain storage of the classification calibration data, and generating a unique index.

[0012] As a preferred solution of the classification method for outputting chain-transmitted data described in the present invention, the method includes: obtaining the original data in the chain-transmitted data list from the data source, including determining the data source according to the data generation situation and data classification and processing requirements, clarifying the data platform to be collected, collecting data information including transaction record data, sensor parameter data and market environment data, and collecting data sources from different data platforms; according to the determined data source, identifying the corresponding data platform and database information, selecting the required data interface or data form as the data source to obtain the required data; performing data preprocessing on the original data extracted from the data source, and the data preprocessing content includes data cleaning, format conversion, and data normalization; and performing diversified proofreading on the preprocessed data.

[0013] As a preferred solution of the classification method for output uplink transmission data described in the present invention, wherein: the data classification rules are formulated according to business needs, and the data classification characteristics are identified by clustering, including conducting business needs analysis, clarifying the data classification scenario based on the name, source and scale of the collected data, organizing the required data types, and clarifying the goals and standards of data classification; conducting data attribute analysis, clarifying the attributes and characteristics of the analyzed data based on the obtained data classification information, and determining the dimensions of data classification; formulating data classification rules based on business needs analysis and data attribute analysis; and identifying data classification characteristics based on the formed data classification rules.

[0014] As a preferred solution of the classification method for output uplink transmission data described in the present invention, the data classification calibration is performed based on the classification features to generate a classification structure including classification labels, original data features and timestamps, including selecting an algorithm for classifying data features and selecting a suitable classification algorithm based on the business needs obtained through analysis; using the determined algorithm, comparatively analyzing the identified classification features to identify differences and similarities between the data, and determining the classification basis of the data based on the distance by quantitatively calculating the distance between the features, which can be expressed as:

[0015]

[0016] Where D is the total number of features of the data, For data x i At the value of d features, is the data μ k The similarity is normalized at the value of d features to obtain the corresponding quantitative difference values ​​between different features.

[0017] As a preferred solution of a classification method for outputting uplink transmission data described in the present invention, wherein: the data classification calibration is performed based on the classification features to generate a classification structure including classification labels, original data features and timestamps, and also includes data classification according to the corresponding quantitative difference values ​​between different features. The classification method adopts line classification, surface classification or hybrid classification method to classify and calibrate the data according to business needs. The line classification method classifies the data according to a single feature, the surface classification method classifies the data according to multiple features, and the hybrid classification method combines the line classification method and the surface classification method to perform multi-level classification on the data, and finally forms a classification result; according to the determined classification result, a corresponding classification label is generated for each data, and the classification label includes the category information of the data, the original data features and the timestamp. The classification label, the original data features and the timestamp information are integrated to construct a classification structure for the classified data.

[0018] As a preferred solution of the classification method for outputting data for chain transmission described in the present invention, the chain verification and classified chain storage of the classified and calibrated data include: storing the processed data on the chain, storing the classified and calibrated data on the chain in different categories according to the data classification structure, and the chain storage is achieved through the distributed ledger technology of the blockchain; comparing and verifying the data stored on the chain, verifying the stored data using the consensus mechanism of the blockchain, calculating a hash value for each data block, and storing the hash value on the blockchain, and verifying whether the hash value of the data is consistent through the consensus mechanism of the blockchain network to ensure that the data has not been tampered with.

[0019] As a preferred solution of the classification method for outputting data for transmission on the chain as described in the present invention, the generation of a unique index includes generating a unique index for the data that has passed the on-chain verification, generating a unique index for each classified data based on the data storage characteristics and data structure, and generating a unique index based on the classification label, timestamp, and hash value information of the data. The unique index is generated based on the classification label, timestamp, and hash value information of the data, and the classified data stored on the chain can be quickly queried and retrieved through the unique index.

[0020] Another object of the present invention is to provide a classification system for outputting uplink transmission data.

[0021] In order to solve the above technical problems, the present invention provides the following technical solutions: a classification system for outputting on-chain transmission data, comprising: a data acquisition module, a data classification feature recognition module, a classification structure generation module and an on-chain storage module; the data acquisition module is used to obtain the original data in the on-chain transmission data list from the data source, and pre-process and multi-dimensionally proofread the original data, including outlier removal, missing value supplementation and error value correction; the data classification feature recognition module is used to formulate data classification rules according to business needs, and identify data classification features by clustering; the classification structure generation module is used to perform data classification calibration based on classification features, and generate a classification structure including classification labels, original data features and timestamps; the on-chain storage module is used to perform on-chain verification and classified on-chain storage of the classification calibration data, and generate a unique index.

[0022] The present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, the steps of the classification method for outputting uplink transmission data are implemented.

[0023] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the method for classifying output uplink transmission data are implemented.

[0024] The present invention achieves beneficial effects by automating data collection through open-source or independently developed tools, reducing manual intervention and improving the efficiency and real-time nature of data collection. Furthermore, through multi-source data comparison and automated verification, data accuracy and integrity are ensured, reducing data errors and biases. The present invention also utilizes advanced algorithms and tools to remove outliers, fill in missing values, and correct erroneous values, improving data quality and laying a solid foundation for subsequent analysis.

[0025] This invention develops multi-dimensional classification rules (e.g., by attribute, source, sensitivity, etc.) based on business needs and data characteristics, improving classification flexibility and accuracy. By analyzing the most sensitive characteristics of the data (e.g., privacy importance), classification rules are dynamically adjusted to ensure effective protection of sensitive data. Furthermore, the invention's classification rules are closely aligned with business needs, ensuring that data classification directly supports business decisions and application scenarios.

[0026] This invention utilizes advanced classification algorithms (line classification, surface classification, and hybrid classification) to improve the accuracy and efficiency of data classification and adapt to the needs of large-scale data processing. It also generates dynamic classification labels based on data features and timestamps, ensuring that the classification results reflect the latest data status. The classification labels, raw data features, and timestamps are integrated into structured data for easy subsequent storage and query, improving the standardization of data management.

[0027] This invention uses blockchain technology to store data on-chain, ensuring data immutability and traceability, and enhancing data security and credibility. Blockchain consensus mechanisms (such as Proof-of-Work and Proof-of-Stake) are also used to automatically verify data, ensuring its authenticity and integrity while reducing manual intervention. Finally, a unique index is generated for the classified data, facilitating rapid query and retrieval, improving data management efficiency and user experience. The transparency of blockchain technology allows for auditability of the data storage and verification process, enhancing data credibility and compliance. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0029] Figure 1 An overall flow chart of a classification method for output uplink transmission data provided by one embodiment of the present invention.

[0030] Figure 2 A diagram of a header table establishment process for a classification method for outputting uplink transmission data provided by one embodiment of the present invention.

[0031] Figure 3 This is a diagram illustrating an example of the FP-Tree establishment process for a classification method for outputting uplink transmission data provided by one embodiment of the present invention.

[0032] Figure 4 A system solution module diagram of a classification system for outputting uplink transmission data provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0033] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0034] Example 1, with reference to Figure 1 , which is the first embodiment of the present invention, provides a classification method for outputting on-chain transmission data, including: obtaining the original data in the on-chain transmission data list from the data source, preprocessing and diversified proofreading the original data, including outlier removal, missing value supplementation and error value correction; formulating data classification rules according to business needs, and using clustering to identify data classification features; based on the classification features, performing data classification calibration, and generating a classification structure including classification labels, original data features and timestamps; performing on-chain verification and classified on-chain storage of the classification calibration data, and generating a unique index.

[0035] S1. Obtain the original data in the chain transmission data list from the data source, and preprocess and multi-faceted proofread the original data, including outlier removal, missing value addition, and error value correction.

[0036] Core raw data from the on-chain data list is collected from the data source. Data processing methods, including pre-processing and multi-faceted proofreading, are used to ensure data accuracy and completeness. Using open-source data tools or proprietary tools, combined with pre-research and background information on the on-chain data, automated collection of a range of on-chain data required during blockchain transactions, including transaction records, sensor parameters, and market environment data, is achieved. Data is then processed to remove outliers, interpolate missing values, and correct erroneous values ​​to ensure accuracy and rationality in subsequent classification steps.

[0037] S101. Determine the data source based on data generation and data classification and processing requirements, identify the data platform to be collected, and collect data information including transaction records, sensor parameter data, and market environment data. The data sources collected come from different systems or platforms, such as exchanges, IoT devices, and market data providers.

[0038] S102. Based on the sources determined in S101, identify the corresponding data platform and database information, select the required data interface or data form as the data source to obtain the required data. Use open source data tools or self-developed tools to automatically extract data from various data sources regularly or in real time;

[0039] S103: Preprocess the raw data automatically extracted from the data source in S102. Data preprocessing includes data cleaning, format conversion, and data normalization. This step includes the following operations: outlier removal, which uses statistical methods or machine learning algorithms to identify and remove outliers in the data; missing value processing, which fills in missing data using methods such as the mean, median, and interpolation; and error correction, which corrects erroneous values ​​in the data through data validation rules or manual review.

[0040] S104: Perform multiple proofreading on the data pre-processed in step S103 to ensure data accuracy and completeness. Proofreading methods include comparison with other data sources, manual review, or the use of automated verification tools. Proofreading information includes data item completeness, data information integrity, and data timestamp accuracy.

[0041] Traditional data collection usually relies on manual or semi-automated methods, with a single data source and lack of real-time performance. Data preprocessing usually requires manual intervention and is inefficient. Existing data cleaning technologies often rely on simple rules or scripts, which are difficult to handle complex outliers, missing values, and erroneous values, and lack automation capabilities. The present invention uses open source tools or self-developed tools to achieve automated data collection, reduce manual intervention, and improve the efficiency and real-time performance of data collection. At the same time, it can ensure the accuracy and completeness of data and reduce data errors and deviations through multi-source data comparison and automated verification. The present invention also uses advanced algorithms and tools to remove outliers, fill in missing values, and correct erroneous values ​​in data, thereby improving data quality and laying a solid foundation for subsequent analysis.

[0042] S2. Develop data classification rules based on business needs and use clustering to identify data classification features.

[0043] According to business needs, basic rules for classifying data transmitted on-chain are formulated. Based on the data's corresponding attributes or characteristics, attribute clustering or characteristic clustering is used to determine the corresponding classification features of the data, facilitating subsequent data classification based on these features. The method for formulating data classification rules is based on the data classification dimensions and the breadth of data sources. Based on the data's name, description, and source, categories such as transaction data, user data, transmission parameters, and blockchain information data are separated. Furthermore, based on the most sensitive features contained in different types of data, such as the most sensitive feature of privacy information is privacy importance, data classification features are then separated.

[0044] S201. Conduct a business needs analysis. Based on the name, source, and size of the collected data, analyze the business needs for data classification and clarify the goals and standards for data classification. For example, determine whether classification is necessary based on data type, data source, data sensitivity, etc. Once the business needs are determined, formulate guidance documents and information for rule development to facilitate subsequent data feature determination.

[0045] S202: Conduct data attribute analysis. Based on the rule-making guidance documents and information obtained in S201, clarify the attributes or characteristics of the data and determine the dimensions for data classification. For example, transaction data may include attributes such as transaction amount, transaction time, and transaction parties; user data may include attributes such as user ID, user behavior, and user preferences. Organize and compile a list of data attributes or characteristics corresponding to the data for use in classification rule development.

[0046] S203. According to step S201 and step S202, basic rules for data classification are formulated. For example, transaction data is classified according to attributes such as transaction type, transaction amount, and transaction time; user data is classified according to attributes such as user ID, user behavior, and user preference; sensor parameter data is classified according to attributes such as sensor type, parameter value, and timestamp; market environment data is classified according to attributes such as market type, market trend, and market fluctuation. Basic data classification rule information is formed for subsequent identification and classification of sensitive data features.

[0047] S204: Based on the basic data classification rule information generated in step S203, identify sensitive features of the data. Further refine the classification rules based on the sensitive features of different types of data. For example, private data can be classified based on privacy importance, and sensitive data can be classified based on the security level of the data. Appropriate feature attributes for different data types are determined for subsequent classification.

[0048] Traditional data classification usually relies on manual experience or simple rules, with a single classification dimension, which makes it difficult to cope with complex data characteristics and business needs. In addition, the existing technology has limited ability to identify sensitive features, usually relying on static rules, and it is difficult to dynamically adapt to changes in data. The present invention formulates multi-dimensional classification rules (such as by attributes, sources, sensitivity, etc.) based on business needs and data characteristics to improve the flexibility and accuracy of classification. By analyzing the most sensitive features of the data (such as privacy importance), the classification rules are dynamically adjusted to ensure that sensitive data is effectively protected. At the same time, the classification rules of the present invention are closely integrated with business needs to ensure that data classification can directly support business decisions and application scenarios.

[0049] S3. Based on the classification features, data classification calibration is performed to generate a classification structure including classification labels, original data features and timestamps.

[0050] Using algorithms to compare and analyze the differences and similarities between the classification features of the on-chain data in S2, a combination of line classification, surface classification, and hybrid classification is used to classify the data according to business classification requirements, generate corresponding classification labels, and identify the data category. The classification labels, original data features, and timestamps are integrated to construct a reference basis for classification data and a data classification structure.

[0051] S301: Select an algorithm for data feature classification. Based on the business requirements and data characteristics analyzed in step S2, select an appropriate classification algorithm. Based on the efficiency, accuracy, and classification capabilities required by the classification business scenario, perform decision analysis and select common classification algorithms, including decision trees, support vector machines, and K-means clustering. Select an appropriate algorithm for subsequent classification calibration.

[0052] S302. Using the algorithm determined in S301, perform a comparative analysis of the classification features determined in S2 to identify differences and similarities between the data. By quantitatively calculating the similarity or distance between the features, the data classification basis is determined based on the magnitude of the similarity or the distance. The similarity is between 0 and 1, with the closer to 1 the greater the similarity, the more likely the data corresponding to the two types of features are to be classified together. The distance is between 0 and 1, with the closer to 0 the distance, the more likely the data corresponding to the two types of features are to be classified together. Through analysis, the corresponding quantitative difference values ​​between different features are obtained.

[0053] S303. Classify the data based on the numerical values ​​corresponding to the quantified features determined in S302. The classification method can be based on business needs, using line classification, surface classification, or a hybrid classification method to classify and calibrate the data. The three methods include: line classification, which classifies data based on a single feature, such as transaction amount. Surface classification, which classifies data based on multiple features, such as transaction amount and transaction time. Hybrid classification: Combining line classification and surface classification to perform multi-level classification on the data, ultimately generating a classification result.

[0054] S304: Generate a corresponding classification label for each data item based on the classification results determined in S303. The classification label includes the data category information, original data features, and timestamp. The classification label, original data features, and timestamp are integrated to construct a reference basis and data structure for the classified data, facilitating subsequent data storage and query.

[0055] Traditional data classification usually relies on simple clustering algorithms or manual calibration, which has low classification accuracy and is difficult to process large-scale data. Classification labels are usually statically generated and lack adaptability to dynamic changes in data. The present invention uses advanced classification algorithms (line classification, surface classification, and hybrid classification) to improve the accuracy and efficiency of data classification and adapt to the needs of large-scale data processing. At the same time, dynamic classification labels are generated based on data features and timestamps to ensure that the classification results can reflect the latest status of the data. The classification labels, original data features, and timestamps are integrated into structured data to facilitate subsequent storage and query, thereby improving the standardization of data management.

[0056] S4. Verify and store the classified and calibrated data on the chain in different categories, and generate a unique index.

[0057] The classified and calibrated data is then stored and verified on-chain. The classified and calibrated data is connected to its data structure and categorized for on-chain storage. Blockchain consensus mechanisms and other mechanisms are then used to verify the stored data on-chain, ensuring its authenticity and integrity. Finally, a unique index is generated for the classified data based on its storage characteristics and structure to facilitate subsequent query and retrieval.

[0058] S401: Store the processed data on-chain. The categorized and labeled data is stored on-chain according to its data structure. This storage can be achieved through blockchain's distributed ledger technology, ensuring data immutability and traceability.

[0059] S402: Compare and verify the data stored on the blockchain. This process uses the blockchain's consensus mechanism (e.g., PoW, PoS, etc.) to verify the stored data and ensure its authenticity and integrity. The specific steps include data hash calculation: calculating a hash value for each data block and storing the hash value on the blockchain; and consensus verification: verifying the consistency of the data's hash value through the blockchain network's consensus mechanism to ensure the data has not been tampered with.

[0060] S403. Generate a unique index for the data that has been verified on the blockchain. Based on the data storage characteristics and structure, a unique index is generated for each categorized data. This unique index is generated based on information such as the data's category label, timestamp, and hash value, facilitating subsequent queries and retrieval. Using this unique index, users can quickly query and retrieve categorized data stored on the blockchain. This query and retrieval can be performed using a blockchain explorer, API, or custom query tools.

[0061] Traditional data storage usually relies on centralized databases, which are subject to the risk of data tampering and loss, and lack transparency and traceability. Data verification usually relies on manual review or simple verification rules, making it difficult to ensure the authenticity and integrity of the data. The present invention uses blockchain technology to achieve data on-chain storage, ensuring the immutability and traceability of data, and enhancing the security and credibility of the data. At the same time, blockchain consensus mechanisms (such as PoW and PoS) are used to automatically verify the data to ensure the authenticity and integrity of the data and reduce manual intervention. Finally, a unique index is generated for the classified data to facilitate quick query and retrieval, improving the efficiency of data management and user experience. The transparency of blockchain technology makes the data storage and verification process auditable, enhancing the credibility and compliance of the data.

[0062] Example 2, reference Figure 2 and Figure 3 , which is the second embodiment of the present invention, provides a classification method for output uplink transmission data. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.

[0063] During the data preparation phase, electric carbon data classification requires two types of data: 1. An electric carbon sensitive data lookup table, which contains the names of sensitive fields in electric carbon data (i.e., a rule table); 2. An electric carbon data table for data classification, which contains multiple sections of each data field. These sections represent specific objects to be classified, such as transmission lines and substations. Each section contains a variable amount of specific data for various fields in the electric carbon system, such as active power, reactive power, node voltage amplitude and phase angle, node marginal price, and maximum and minimum active power output values ​​(i.e., a data table). This example requires data classification based on sensitivity.

[0064] The specific method of establishing the sensitive data dictionary is: enumerate all the objects to be classified in the electric carbon data table and use them as indexes in the sensitive data dictionary; traverse the sensitive data index, for example: search for all data fields under the index name (such as a line, substation, etc.) in the electric carbon data table; open a corresponding project set for the index in the sensitive data dictionary; based on the electric carbon sensitive data lookup table, judge whether all data fields under the index are sensitive. If sensitive, add the field name to the project set of the current index; after the traversal is completed, the construction of the sensitive data dictionary is completed.

[0065] In this case, the FP-Tree (Frequent Pattern Tree) based on the sensitive data dictionary is used to mine frequent item sets of sensitive data, thereby providing a basis for classification. The establishment of the FP-Tree first depends on the establishment of the header table. The specific approach is: scan all the item sets in the sensitive data dictionary to obtain the count of all frequent 1-item sets; delete items with support lower than the threshold, put the frequent 1-item sets into the header table, and sort them in descending order of support; scan all the item sets in the sensitive data dictionary, and remove all non-frequent 1-item sets in the item sets under each index, and then sort them in descending order of support. The support refers to the probability of the item appearing in all index item sets, which is calculated by the ratio of the number of times the item appears to the number of index sets. The establishment of the header table can be used Figure 2 Specific example description.

[0066] As can be seen in the figure, each row of the data is an index set, and the items in the header table are sorted by their frequency of occurrence in each index set. In this example, the support threshold is set to 20%, so items below this threshold are removed from the original dataset.

[0067] Secondly, it is necessary to establish an FP-Tree for frequent item set mining based on the item header table. Initially, there is no data in the FP-Tree, so when establishing the FP-Tree, we need to read the sorted item sets in the order of the index in the dictionary and insert them into the FP-Tree. For each item set, the items need to be inserted into the FP-Tree in the order of the sorting of the item header table: the nodes with the highest sorting are ancestor nodes, and the nodes with the lowest sorting are descendant nodes. If there is a common ancestor, the count of the corresponding common ancestor node is increased by 1. After the insertion, if a new node appears, the node corresponding to the item header table will be linked to the new node through the node linked list. Until all the item sets are inserted into the FP-Tree, the establishment of the FP-Tree is completed, and the classification of sensitive data is basically achieved. The specific approach is as follows: Figure 3 shown.

[0068] Depend on Figure 3 It can be seen that the FP-Tree establishment process is based on Figure 2 Taking the data in the six-item header table as an example, the specific process of inserting the items in the first four index sets into the FP-Tree is demonstrated.

[0069] After obtaining the FP-Tree and header table, this example uses the FP-growth algorithm to mine frequent itemsets. First, traverse the items at the bottom of the header table. For each item in the header table corresponding to the FP-Tree, find its corresponding conditional pattern base. The conditional pattern is the FP subtree corresponding to the leaf node of the node to be mined. After obtaining the FP subtree, set the count of each node in the subtree to the count of the leaf node, and delete nodes with counts below the support. At this point, recursive mining can be performed based on the conditional pattern to obtain frequent itemsets. At this point, we can extract the sensitive features of the objects to be classified based on the sensitive data dictionary and the sensitive data frequent itemsets. The specific method is as follows: first, traverse each index in the sensitive data dictionary and establish a feature set dictionary corresponding to each index: the index in the feature set dictionary is still the object to be classified, and the content is the feature vector corresponding to the index. The dimension of the vector is the same as the number of frequent itemsets, and each field is a 0, 1 variable that represents whether each frequent item set is included in the index item set of the sensitive data dictionary; during the traversal, for the current index, first take out the corresponding item set in the sensitive data dictionary, and then traverse the frequent item sets to determine whether each frequent item set is included in the index item set. If included, the corresponding feature variable takes the value of 1, otherwise it is 0; after the traversal is completed, the sensitive data classification set of electric carbon data can be obtained.

[0070] Example 3 is the third embodiment of the present invention, which differs from the first two embodiments in that:

[0071] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0072] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0073] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0074] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0075] Example 4, with reference to Figure 4 , which is the fourth embodiment of the present invention, provides a classification system for outputting uplink transmission data, including a data acquisition module, a data classification feature recognition module, a classification structure generation module and an uplink storage module.

[0076] The data acquisition module is used to obtain the original data in the chain transmission data list from the data source, and pre-process and multi-dimensionally proofread the original data, including outlier removal, missing value addition and error value correction.

[0077] The data classification feature identification module is used to formulate data classification rules according to business needs and identify data classification features using clustering.

[0078] The classification structure generation module is used to perform data classification calibration based on classification features and generate a classification structure containing classification labels, original data features and timestamps.

[0079] The on-chain storage module is used to verify the classification and calibration data on the chain, store them on the chain in different categories, and generate a unique index.

[0080] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A classification method for outputting uplink transmission data, characterized by: include, Obtain the raw data in the chain transmission data list from the data source, and perform preprocessing and diversified proofreading on the raw data, including outlier removal, missing value addition, and error value correction; Formulate data classification rules based on business needs and use clustering to identify data classification features; Based on the classification features, data classification calibration is performed to generate a classification structure containing classification labels, original data features and timestamps; The classified and calibrated data is verified and stored on the chain in different categories, and a unique index is generated.

2. The method for classifying output uplink transmission data according to claim 1, wherein: The method of obtaining the original data in the uplink transmission data list from the data source includes: Determine the data source based on data generation and data classification and processing requirements, clarify the data platform to be collected, and collect data information including transaction record data, sensor parameter data, and market environment data. The data sources are collected from different data platforms; Based on the determined data source, identify the corresponding data platform and database information, and select the required data interface or data form as the data source to obtain the required data; Perform data preprocessing on the raw data extracted from the data source. Data preprocessing includes data cleaning, format conversion, and data normalization. Perform diversified proofreading on the preprocessed data.

3. The method for classifying output uplink transmission data according to claim 2, wherein: The data classification rules are formulated according to business needs, and the data classification features are identified by clustering, including: Conduct business needs analysis, clarify data classification scenarios based on the name, source, and scale of the collected data, organize the required data types, and clarify the goals and standards for data classification; Conduct data attribute analysis, and based on the obtained data classification information, clearly analyze the attributes and characteristics of the data and determine the dimensions of data classification; Formulate data classification rules based on business needs analysis and data attribute analysis; According to the formed data classification rules, data classification features are identified.

4. The method for classifying output uplink transmission data according to claim 3, wherein: The data classification is calibrated based on the classification features to generate a classification structure containing classification labels, original data features and timestamps, including: Select an algorithm for data feature classification and choose an appropriate classification algorithm based on the business needs obtained through analysis; Using a certain algorithm, we can compare and analyze the identified classification features, identify the differences and similarities between the data, and quantify the distance between the features to determine the classification basis of the data based on the distance, which can be expressed as: Where D is the total number of features of the data, For data x i At the value of d features, is the data μ k The similarity is normalized at the value of d features to obtain the corresponding quantitative difference values ​​between different features.

5. The method for classifying output uplink transmission data according to claim 4, wherein: The data classification is calibrated based on the classification features to generate a classification structure including classification labels, original data features and timestamps, and also includes: Data is classified based on the corresponding quantitative difference values ​​between different features. The classification method is based on business needs, and line classification, surface classification, or hybrid classification is used to classify and calibrate the data. Line classification classifies data according to a single feature, surface classification classifies data according to multiple features, and hybrid classification combines line classification and surface classification to perform multi-level classification on the data, ultimately forming a classification result. According to the determined classification results, a corresponding classification label is generated for each data. The classification label contains the category information of the data, the original data characteristics and the timestamp. The classification label, the original data characteristics and the timestamp information are integrated to construct the classification structure of the classified data.

6. The method for classifying output uplink transmission data according to claim 4, wherein: The verification and classified storage of the classified calibration data on the chain includes: The processed data is stored on the chain, and the classified and calibrated data is stored on the chain in different categories according to the data classification structure. The chain storage is achieved through the distributed ledger technology of the blockchain; The data stored on the chain is compared and verified, and the stored data is verified using the consensus mechanism of the blockchain. The hash value is calculated for each data block and stored on the blockchain. The consensus mechanism of the blockchain network is used to verify whether the hash value of the data is consistent to ensure that the data has not been tampered with.

7. The method for classifying output uplink transmission data according to claim 4, wherein: As stated and generate a unique index, including, For data that has passed the chain verification, a unique index is generated for the data. According to the data storage characteristics and data structure, a unique index is generated for each classified data. The unique index is generated based on the classification label, timestamp, and hash value information of the data. Through the unique index, the classified data stored on the chain can be quickly queried and retrieved.

8. A classification system for output uplink transmission data, applying a classification method for output uplink transmission data according to any one of claims 1 to 7, characterized in that: include: Data acquisition module, data classification feature recognition module, classification structure generation module and chain storage module; The data acquisition module is used to obtain the original data in the chain transmission data list from the data source, and perform preprocessing and diversified proofreading on the original data, including outlier removal, missing value addition and error value correction; The data classification feature identification module is used to formulate data classification rules according to business needs and identify data classification features using a clustering method; The classification structure generation module is used to perform data classification calibration based on classification features and generate a classification structure including classification labels, original data features and timestamps; The on-chain storage module is used to perform on-chain verification and classified on-chain storage of the classified calibration data, and generate a unique index.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of a classification method for output uplink transmission data according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a method for classifying output uplink transmission data according to any one of claims 1 to 7 are implemented.