Classification and grading method of traceability sensitive data
By constructing hierarchical indicators based on enterprise business relevance and blockchain storage, the problems of vague concepts and weak business targeting in traditional data classification and grading schemes have been solved, and the accurate classification and secure sharing of traceability sensitive data have been achieved.
Patent Information
- Application Number
- CN202411772354.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2044-12-04
AI Technical Summary
Traditional data classification and grading schemes are vague in concept and lack business relevance, making it difficult to effectively manage the sharing and protection of sensitive corporate data.
By constructing a hierarchical index based on the relevance of enterprise business and its weights, and combining social network analysis and blockchain technology, the source-tracing sensitive data is classified into primary and secondary categories, and the initial sensitivity classification results are optimized and stored on the blockchain to ensure data security.
It enables precise classification and grading of sensitive data for traceability, ensuring the credibility and security of data sharing. It solves the problems of vague concepts and weak business relevance in traditional solutions, and improves the effectiveness of data sharing and privacy protection capabilities.
Smart Images

Figure CN119692855B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data classification and grading, and specifically to a method, system, storage medium, and electronic device for classifying and grading sensitive data for tracing its origin. Background Technology
[0002] The advent of the big data era has made data a crucial asset for enterprises. When sharing data externally, companies must determine which data can be shared openly and which data must not be disclosed. Therefore, to ensure enterprise data security and regulate the management of enterprise privacy data, the classification and grading of sensitive data has become a research hotspot.
[0003] In related technologies, traditional classification and grading schemes generally classify data first based on the business characteristics of the industry, and then grade the data according to the importance of the country, society, and enterprises. For example, the paper (Zhang Xiaoyi, Dai Yicong. Water Conservancy Data Classification and Grading and Security Protection Technology [J]. Yangtze River, 2023, 54(S2):232-237.DOI:10.16232 / j.cnki.1001-4179.2023.S2.053) proposes a data security protection system adapted to the characteristics of water conservancy business.
[0004] However, this method has a rather vague classification concept and is not very business-specific. Summary of the Invention
[0005] (a) Technical problems to be solved
[0006] To address the shortcomings of existing technologies, this invention provides a method, system, storage medium, and electronic device for classifying and grading sensitive data for tracing its origins, solving the technical problems of traditional classification and grading schemes having vague concepts and weak business relevance.
[0007] (II) Technical Solution
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] A method for classifying and grading sensitive data for tracing its origin includes:
[0010] All source-tracing sensitive data are classified into primary and secondary categories in sequence;
[0011] Construct a tiered index; wherein the tiered index includes at least one or a combination of the following: business relevance of the enterprise and its weight, and the impact of data sensitivity, data importance, and the degree of impact of data loss or leakage;
[0012] Based on hierarchical indicators other than the degree of relevance to enterprise business and its weight, a data classification model is constructed to perform initial sensitivity classification of traceability sensitive data under each first-level subclass.
[0013] Based on the enterprise business relevance and its weight, the initial sensitivity classification results are optimized, and the final classification and grading results of each traceability sensitive data are obtained.
[0014] Preferably, the step of sequentially classifying all source-tracing sensitive data into primary and secondary categories includes:
[0015] The source-tracing sensitive data is classified into primary categories according to attribute characteristics, and then classified into secondary categories based on the data content, ownership, and access permissions of the source-tracing sensitive data under each primary subcategory.
[0016] Preferably, the process of constructing the enterprise business relevance includes:
[0017] Acquire and preprocess transaction data between various enterprises in the upstream and downstream of the large-scale manufacturing industry chain; wherein, the data indicators of the transaction data include enterprise ID, transaction ID, order quantity, unit price, and invoice date;
[0018] Based on the preprocessed transaction data, calculate the R, F, and M values using enterprise ID as the unit;
[0019] Using the Social Network Analysis (SNA) method, a social network model of enterprise transaction relationships is constructed to calculate the business relevance of each enterprise; including:
[0020] The freshness Fre and contractual relationship Con between any two enterprises are statistically analyzed. Combined with the corresponding R, F, and M values, the sum is used to measure the strength of the transaction relationship between any two enterprises. This process is repeated for all enterprises to obtain the adjacency matrix of the transaction relationship strength between enterprises.
[0021] The freshness Fre is used to characterize whether there is a transaction relationship between the two companies within the benchmark date. If there is, the value is 1, otherwise the value is 0. The contractual relationship Con is used to characterize whether the transaction volume between the two companies within the benchmark date reaches the set benchmark value. If it does, the value is 1, otherwise the value is 0.
[0022] The adjacency matrix is transformed into a business transaction relationship social network model using the uCinet visualization tool. In this model, nodes represent different businesses, lines between nodes represent transaction relationships between two businesses, and the data above the lines represents the business relevance between nodes.
[0023] Based on the aforementioned enterprise transaction relationship social network model, the degree centrality of the node corresponding to each enterprise is calculated and used as the enterprise's business relevance.
[0024] Preferably, the preprocessing of the transaction data includes:
[0025] The consumption amount is represented by the variable amount, where amount = quantity * unit price; and quantity represents the number of orders and unit price represents the unit price.
[0026] Extract the purchase date and purchase time by invoice date to distinguish transactions created by the company at different times;
[0027] Delete invalid data.
[0028] Preferably, the calculation of R, F, and M values based on preprocessed transaction data, using enterprise ID as the unit, includes:
[0029] Select a company ID and calculate the company's R value based on the interval between the company's most recent purchase date and the base date;
[0030] The F-value of a company is obtained by counting the number of transactions made by the company within a unit of time after the baseline date; different transaction IDs between the company and the same trading partner within the same unit of time are considered as one transaction.
[0031] The company's M-value is obtained by calculating the amount of its purchases after the base date.
[0032] Preferably, the optimization of the initial sensitivity classification result based on the enterprise's business relevance and its weight is expressed as:
[0033] S i ′ j =S ij +C o w
[0034] Among them, S ij Let set A be ij initial sensitivity, A ij C represents the set of tags for source-tracing sensitive data with first-level subclass j and sensitivity level i; o represents the degree of business relevance of the enterprise; w represents the weight of the business relevance of the enterprise.
[0035] Preferably, the classification and grading method further includes:
[0036] All traceability-sensitive data and their final classification and grading results are stored on the blockchain.
[0037] A classification and grading system for tracing sensitive data includes:
[0038] The classification module is used to classify all traceability-sensitive data into primary and secondary categories in sequence.
[0039] A construction module is used to construct hierarchical indicators; wherein, the hierarchical indicators include at least one or a combination of several of the following: enterprise business relevance and its weight, data sensitivity, data importance, and the impact of data loss or leakage;
[0040] The grading module is used to build a data classification model based on grading indicators other than the enterprise business relevance and its weight, so as to perform initial sensitivity grading of the traceability sensitive data under each first-level subclass;
[0041] And to optimize the initial sensitivity classification results based on the enterprise's business relevance and its weight, and to obtain the final classification and grading results for each piece of traceability sensitive data.
[0042] A storage medium storing a computer program for classifying and grading source-tracing sensitive data, wherein the computer program causes a computer to execute the source-tracing sensitive data classification and grading method as described above.
[0043] An electronic device, comprising:
[0044] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including methods for performing the classification and grading of source-tracing sensitive data as described above.
[0045] (III) Beneficial Effects
[0046] This invention provides a method, system, storage medium, and electronic device for classifying and grading sensitive data for tracing its origin. Compared with existing technologies, it has the following advantages:
[0047] In this invention, all traceability-sensitive data are sequentially classified into primary and secondary categories; hierarchical indicators are constructed; based on these hierarchical indicators, excluding enterprise business relevance and its weight, a data classification model is built to initially classify the traceability-sensitive data under each primary subcategory; based on enterprise business relevance and its weight, the initial sensitivity classification results are optimized to obtain the final classification and grading result for each piece of traceability-sensitive data. Building upon traditional solutions and considering the characteristics of traceability business, this invention uses the business relevance between enterprises as an indicator for sensitive data classification, providing a new approach to classifying and grading enterprise traceability-sensitive data. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 A block diagram illustrating a method for classifying and grading sensitive data for tracing the source, provided in an embodiment of the present invention;
[0050] Figure 2 A block diagram illustrating another method for classifying and grading sensitive data for tracing the source, provided in an embodiment of the present invention;
[0051] Figure 3 A schematic diagram of a data classification tree for tracing sensitive data provided in an embodiment of the present invention;
[0052] Figure 4 A schematic diagram of a social network model of enterprise transaction relationships provided in an embodiment of the present invention;
[0053] Figure 5 This is an example diagram of a source-tracing sensitive data classification scheme provided in an embodiment of the present invention. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0055] This application provides a method, system, storage medium, and electronic device for classifying and grading sensitive data for tracing its origins, which solves the technical problems of traditional classification and grading schemes having vague concepts and weak business relevance.
[0056] The technical solution in this application is to solve the above-mentioned technical problems, and the general idea is as follows:
[0057] First, it should be noted that the term "traceability" refers to the process of recording and tracking information at each stage of a product's production, processing, transportation, and sales to achieve traceability and management throughout its entire lifecycle. Because traceability-sensitive data involves numerous stakeholders and stages, when assigning responsibility based on this data, it's crucial to consider the data's sensitivity level to determine which data can be shared externally and which cannot be publicly disclosed.
[0058] This invention primarily categorizes and classifies sensitive data involved in traceability operations to facilitate on-demand sharing of product data during the traceability process. Building upon traditional solutions and considering the characteristics of traceability operations, this invention uses the business relationships between enterprises (traceability entities) as an indicator for classifying sensitive data.
[0059] Furthermore, most traditional sensitive data classification and grading schemes adopt a centralized data storage scheme, that is, the data is stored uniformly in a central database. When the database is maliciously attacked, the stored data will be lost or leaked, causing irreversible losses.
[0060] In response, this invention also utilizes blockchain technology as a storage medium to store all traceability-sensitive data and its final classification and grading results, thereby eliminating security risks in data storage and transmission of sensitive data classification and grading schemes.
[0061] In summary, the embodiments of the present invention solve the problems of vague concepts, weak business targeting, and weak system security in traditional classification and grading schemes, and realize the classification, grading and sharing of traceability sensitive data.
[0062] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0063] Example 1:
[0064] like Figure 1 As shown, this embodiment of the invention provides a method for classifying and grading sensitive data for tracing its origin, including:
[0065] S1. All traceability sensitive data are classified into primary and secondary categories in sequence;
[0066] S2. Construct hierarchical indicators; wherein, the hierarchical indicators include at least one or a combination of several of the following: enterprise business relevance and its weight, data sensitivity, data importance, and the impact of data loss or leakage.
[0067] S3. Based on hierarchical indicators other than the enterprise business relevance and its weight, construct a data classification model to perform initial sensitivity classification of the traceability sensitive data under each first-level subclass.
[0068] S4. Based on the enterprise business relevance and its weight, optimize the initial sensitivity classification result and obtain the final classification and grading result of each traceability sensitive data.
[0069] Based on traditional solutions, this invention, combined with the characteristics of traceability business, uses the degree of business correlation between enterprises as an indicator for classifying sensitive data, providing a new approach for classifying and grading sensitive traceability data for enterprises.
[0070] In an optional embodiment, as Figure 2 shown, the classification and grading method provided by the embodiments of the present invention further includes:
[0071] S5. Store all traceability-sensitive data and their final classification and grading results on the blockchain.
[0072] The embodiments of the present invention use the blockchain as the underlying storage tool to store traceability-sensitive data, ensuring the credibility during data sharing and preventing the traceability process from being untrustworthy due to malicious tampering of data.
[0073] Next, each step of the above solution will be introduced in detail:
[0074] In step S1, all traceability-sensitive data is classified at the first level and then at the second level in sequence.
[0075] Referring to the opinions given in the "Guidelines for the Classification and Grading of Industrial Data (Trial)", classify the large-scale manufacturing industry data according to factors such as manufacturing industry requirements, characteristics, business needs, data sources, and uses, and classify the data in the industrial and information fields into three levels: general data, important data, and core data according to the degree of harm caused by data tampering, destruction, leakage, or illegal acquisition and illegal use to national security, public interests, or the legitimate rights and interests of individuals and organizations. After determining the data classification target, determine the data classification system according to the manufacturing industry characteristics, and then divide and classify the traceability-sensitive data according to the classification system and rules.
[0076] In this step of classification, first classify the traceability-sensitive data according to attribute characteristics at the first level, and then classify at the second level based on the data content, ownership, and access rights of the traceability-sensitive data under each first-level subclass, so as to group the same type of data into one category.
[0077] Exemplarily, data classification can use the form of a classification tree to subdivide data types layer by layer, forming a classification scheme with clear classification details. The form of the classification tree is as Figure 3 shown.
[0078] In step S2, construct grading indicators.
[0079] After data classification, this step continues to grade the data for sensitivity through the following four indicators to restrict the access rights of other users:
[0080] 1) Data sensitivity S, which divides the data into three levels: first level, second level, and third level. The higher the level, the higher the data sensitivity level.
[0081] Exemplarily, assume 0 < S < 10. When 0 < S ≤ S l the data sensitivity level is the first level. When S l < S ≤ S mThe data sensitivity level is level 2, S>S m The data sensitivity level is Level 3. For example, Level 3 sensitive data includes personal information and financial data.
[0082] 2) The importance of the system where the data resides.
[0083] If the system containing the data is critical to the business, the data level is high; conversely, if a system outage has little impact on the business, the data level is low.
[0084] 3) The extent of the impact of data loss or leakage.
[0085] Data loss or leakage that causes significant damage and impact is classified as high-level, while data with minor impact is classified as medium-low-level.
[0086] 4) Business relevance of the enterprise C o With weight w.
[0087] Enterprise relevance represents the business relationship between two enterprises, and its construction process is described later. Higher business relevance indicates closer transactions between the enterprises, involves more sensitive data related to both enterprises, and a higher degree of data sharing between them; that is, the sensitivity level of this type of data relative to the two enterprises is lower. Slight differences in the relevance between different enterprises will result in different sensitivity levels for the same category of data when accessed by different enterprises. Sensitive data can be classified using different methods. In this embodiment of the invention, sensitive data is classified based on the enterprise's business characteristics, and then graded according to different categories of data based on business relevance. Enterprises can assign weights to business relevance according to their needs; higher weight values indicate that the enterprise considers more business relevance, and enterprises with higher relevance values can access a higher level of sensitive data.
[0088] The large-scale manufacturing supply chain involves numerous upstream and downstream enterprises with complex and ever-changing business relationships. Some traceability-sensitive data may involve multiple companies, but when sharing such data, it is often difficult to measure the correlation between the data and the companies. Therefore, the RFM model is an important tool for analyzing customer value. It assesses customer value by analyzing three dimensions: the time since the last purchase (Recency), the frequency of purchases within a certain period (Frequency), and the monetary value of purchases within a certain period (Monetary). This paper describes the business relationships between manufacturing companies using the RFM model and constructs a business relationship model using the VSM model.
[0089] To elaborate further, regarding the business relevance C of the enterprise oThis invention employs the RFM method to calculate business relevance indicators and utilizes the SNA method to abstract and construct a network diagram of large-scale manufacturing enterprises and their interrelationships, thereby deriving the business relevance between enterprises. Specific steps include:
[0090] S100: Acquire and preprocess transaction data between various enterprises in the upstream and downstream of the large-scale manufacturing industry chain.
[0091] Large-scale manufacturing enterprises have huge order volumes, but the transactions between enterprises are relatively fixed. The following data indicators are selected here: enterprise ID (user ID), transaction ID (trans ID), quantity, unit price, and invoice date.
[0092] When calculating the three metrics of the RFM model, the data must first be preprocessed:
[0093] First, the consumption amount is represented by the variable amount, where amount = quantity * unit price; and quantity represents the number of orders and unit price represents the unit price.
[0094] Secondly, the purchase date and purchase time can be extracted by the invoice date to distinguish transactions created by the company at different times;
[0095] Finally, delete invalid data. That is, delete data that is missing information (such as missing user ID, or invalid data with zero or negative order quantity or amount), and keep only valid data.
[0096] S200. Based on the preprocessed transaction data, calculate the R, F, and M values using the enterprise ID as the unit; including:
[0097] R-value: Select a company ID and calculate the company's R-value based on the interval between the company's most recent purchase date and the base date.
[0098] For example, assuming the base date is October 9, 2024, and the company's most recent purchase date is October 9, 2024, then the corresponding R value is 0 (R∈[0,12], in months).
[0099] F-value: The F-value is the number of transactions a company makes per unit of time after the baseline date. count trans ID This represents the sum of the number of different transaction IDs within n months after the base date. It is important to note that different transaction IDs between the same company and the same trading partner within the same time period are considered as one transaction.
[0100] M-value: The amount of money the company spends on products after the statistical base date. The company's M-value is obtained (M>0).
[0101] S300 uses the Social Network Analysis (SNA) method to construct a social network model of enterprise transaction relationships and calculate the degree of business relevance for each enterprise.
[0102] The process of tracing information sharing involves a large amount of data, and each piece of information plays a different role in data classification. Considering the overall data when calculating the correlation between nodes would significantly increase the computational difficulty. By selecting data with high node similarity for sensitivity level classification, the computational load can be reduced while optimizing the classification effect. Social Network Analysis (SNA) was proposed by British anthropologist Radcliffe-Brown in his analysis of social structures. Considering that the relationships between upstream and downstream enterprises in a large-scale manufacturing supply chain can form a corresponding network model, this embodiment of the invention uses the SNA method to abstract and construct a network graph of large-scale manufacturing enterprises and their interrelationships, which is beneficial for identifying nodes with different transaction correlations.
[0103] Specifically, step S300 includes:
[0104] S301. Calculate the freshness Fre and contractual relationship Con between any two enterprises, and sum the corresponding R, F, and M values to measure the strength of the transaction relationship between any two enterprises. Iterate through all enterprises to obtain the adjacency matrix of the transaction relationship strength between enterprises.
[0105] The freshness Fre is used to characterize whether there is a transaction relationship between the two companies within the benchmark date. If there is, the value is 1, otherwise the value is 0. The contractual relationship Con is used to characterize whether the transaction volume between the two companies within the benchmark date reaches the set benchmark value. If it does, the value is 1, otherwise the value is 0.
[0106] For example, the degree of relationships between enterprises is summarized based on the above five types of relationships and transformed into an adjacency matrix as shown in Table 1:
[0107] Table 1
[0108]
[0109]
[0110] In this table, A1-A7 represent 7 suppliers, B1-B8 represent 8 manufacturers, C1-C8 represent 8 transporters, and D1-D8 represent 8 distributors. The adjacency matrix (1-modular matrix) is formed by scoring the five indicators of R, F, M, freshness Fre, and contractual relationship Con among the companies. That is, the values in the matrix are derived from R+F+M+Fre+Con, where 0 indicates that there is no transaction relationship between the two companies during this period, and the larger the value, the closer the transaction is.
[0111] S302. Using the uCinet visualization tool, the adjacency matrix is transformed into the following form: Figure 4 The diagram shows a social network model of corporate transaction relationships.
[0112] like Figure 4 As shown, nodes in the business transaction relationship social network model represent different enterprises, connections between nodes represent transaction relationships between two enterprises, and data above the connections represent the business relevance between nodes.
[0113] Furthermore, to make the transaction relationships between nodes clearer, nodes with a correlation score greater than or equal to 3.0 are selected in the diagram, and the nodes with the highest correlation scores are highlighted in bold. It is evident from the diagram that A6, B2, B5, B3, and C8 are at the center of the transaction relationship network, with relatively close relationships to the other companies in the network; while A1, A2, D1, D5, and D6 are at the periphery of the transaction relationship network, with relatively loose relationships to the other nodes.
[0114] S303. Based on the enterprise transaction relationship social network model, calculate the degree centrality of the node corresponding to each enterprise, and use it as the enterprise business correlation degree of that enterprise.
[0115] Centrality measures whether a node is at the center or the edge of a network. There are three metrics for measuring centrality: degree centrality, center centrality, and proximity centrality. This invention measures the closeness between nodes, specifically using degree centrality to calculate the degree.
[0116] Continuing with the example above, Figure 4 The social network model of enterprise transaction relationships in the model can qualitatively show the influence of each node. The degree centrality results show that the degree centrality of A6, B2, B5, B3, and C8 are 91, 83, 85, 80, and 79, respectively, while the degree centrality of A1, A2, D1, D5, and D6 are 50, 53, 62, 58, and 57, respectively.
[0117] This completes the process of building the enterprise's business relevance.
[0118] In summary, indicators 1), 2), and 3) above refer to common data classification and grading methods, while indicator 4) is innovatively introduced in this embodiment of the invention. It considers the business relationship between the two enterprises when sharing traceability data, avoiding the inability to accurately locate the sharing subject and scope when simply using a common classification and grading scheme. See steps S3 and S4 for details:
[0119] In step S3, a data classification model is constructed based on hierarchical indicators other than the enterprise business relevance and its weight, so as to perform initial sensitivity classification of the traceability sensitive data under each first-level subclass.
[0120] Considering the business characteristics of manufacturing enterprises, traceability data can be categorized into R&D, production, operation and maintenance, management, and external data, and further subdivided under these five primary subcategories. Figure 5 The data is shown as a second-level subclass. Using qualitative indicators 1), 2), and 3), the categorized data can be assigned to corresponding sensitivity levels. Enterprises can then determine the sharing conditions for each level of data based on their respective management needs.
[0121] Figure 5 The source-tracing sensitive data classification and grading scheme divides data into primary subcategories based on business dimensions, and further subdivides the data within each major category at a finer granular level. To adapt to the scalability of this scheme under different classification methods, [the following is omitted as it is not relevant to the main text]. Figure 5 The specific classification results are abstracted into the data classification model shown in Table 2.
[0122] Table 2
[0123]
[0124] In this table, T class -j indicates the first-level subclass of source-sensitive data, which is the coarsest-grained classification; T sense -i indicates the sensitivity level of the data being traced. For example, here the data sensitivity level is divided into three levels, with different disclosure methods and management models set for data of different sensitivity levels; A ij Let A represent the set of source-tracing sensitive data tags with first-level subclass j and sensitivity level i, where A ij ={a1,a2,...,a n}, a k This represents the secondary feature label of subclass j and sensitivity level i.
[0125] In step S4, based on the enterprise business relevance and its weight, the initial sensitivity classification result is optimized, and the final classification and classification result of each traceability sensitive data is obtained.
[0126] Understandably, based on the obtained correlation analysis results, we can see the degree of connection between various enterprises. In the actual process of sharing traceability sensitive data, traceability sensitive data is often only of reference value to enterprises with which they have transactions. Therefore, considering the relationship between enterprises can accurately locate the scope of data sharing and prevent data from being used maliciously.
[0127] Determining the relationships between entities provides a new approach to classifying and grading sensitive data for tracing its origins. Limiting the scope of data sharing based on the strength of relationships between enterprises can more effectively ensure data security and privacy, while maximizing data sharing.
[0128] Accordingly, this step, based on the above 4) business relevance of enterprises, optimizes the initial sensitivity classification results by adjusting the proportion of self-adjusting indicators with weighting coefficients; expressed as:
[0129] S i ′ j =S ij +C o w
[0130] Among them, S ij Let set A be ij initial sensitivity, A ij C represents the set of tags for source-tracing sensitive data with first-level subclass j and sensitivity level i; o represents the degree of business relevance of the enterprise; w represents the weight of the degree of business relevance of the enterprise.
[0131] Finally, A was reclassified based on the optimized sensitivity. ij The sensitivity level of internal data can be used to classify and grade sensitive data based on business relevance.
[0132] In step S5, all traceability-sensitive data and their final classification and grading results are stored on the blockchain.
[0133] Thus, this embodiment of the invention completes the entire process of classifying and grading sensitive data for tracing the source.
[0134] Example 2:
[0135] This invention provides a classification and grading system for tracing sensitive data, comprising:
[0136] The classification module is used to classify all traceability-sensitive data into primary and secondary categories in sequence.
[0137] A construction module is used to construct hierarchical indicators; wherein, the hierarchical indicators include at least one or a combination of several of the following: enterprise business relevance and its weight, data sensitivity, data importance, and the impact of data loss or leakage;
[0138] The grading module is used to build a data classification model based on grading indicators other than the enterprise business relevance and its weight, so as to perform initial sensitivity grading of the traceability sensitive data under each first-level subclass;
[0139] And to optimize the initial sensitivity classification results based on the enterprise's business relevance and its weight, and to obtain the final classification and grading results for each piece of traceability sensitive data.
[0140] Example 3:
[0141] This invention provides a storage medium storing a computer program for classifying and grading sensitive data for tracing its origin, wherein the computer program causes a computer to execute the classification and grading method for sensitive data for tracing its origin as described above.
[0142] Example 4:
[0143] This invention provides an electronic device, comprising:
[0144] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including methods for performing the classification and grading of source-tracing sensitive data as described above.
[0145] It is understood that the classification and grading system, storage medium and electronic device for traceability sensitive data provided in the embodiments of the present invention correspond to the classification and grading method for traceability sensitive data provided in the embodiments of the present invention. The explanation, examples and beneficial effects of the relevant contents can be referred to the corresponding parts of the classification and grading method, and will not be repeated here.
[0146] In summary, compared with existing technologies, it has the following beneficial effects:
[0147] 1. The embodiments of the present invention are aimed at sharing traceability-sensitive data in large-scale manufacturing industries, and provide ideas for classifying and grading traceability-sensitive data of manufacturing enterprises.
[0148] 2. The embodiments of the present invention take into account the business relationships between enterprises when classifying and grading sensitive data for tracing the source, which solves the problem of weak data sharing caused by the lack of specificity of traditional classification and grading schemes, and resolves the contradiction between data sharing and privacy protection.
[0149] 3. The embodiments of the present invention use blockchain as the underlying storage tool to store traceability sensitive data, which ensures the credibility of data sharing and prevents data from being maliciously tampered with, thus making the traceability process unreliable.
[0150] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0151] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method of classifying a hierarchy of sensitive data for tracing, characterized by, The application comprises the following steps: Classifying all traceability sensitive data in turn in a first level and a second level; Building a hierarchical index, wherein the hierarchical index at least comprises enterprise business correlation degree and its weight, and a combination of one or any of data sensitivity, data importance, and influence degree of data loss or leakage; Building a data classification model based on the hierarchical index except for the enterprise business correlation degree and its weight, to classify the traceability sensitive data under each first level subcategory in an initial sensitivity level; Optimizing the initial sensitivity classification result based on the enterprise business correlation degree and its weight, to obtain a final classification result of each traceability sensitive data; The classification of all traceability sensitive data in turn in a first level and a second level comprises: Classifying the traceability sensitive data in a first level according to attribute characteristics, and classifying the traceability sensitive data in a second level based on data content, ownership, and access permission of the traceability sensitive data under each first level subcategory; The process of building the enterprise business correlation degree comprises: Obtaining and preprocessing transaction data between upstream and downstream enterprises in a large-scale manufacturing industry chain, wherein the data index of the transaction data comprises enterprise ID, transaction ID, order quantity, unit price, and invoice date; Calculating R, F, and M values based on the preprocessed transaction data in units of enterprise ID; Building an enterprise transaction relationship social network model by using a social network analysis (SNA) method, and calculating the enterprise business correlation degree of each enterprise, which comprises: Statistically analyzing the freshness (Fre) and contract relationship (Con) between two enterprises, combining the corresponding R, F, and M values, summing up, and measuring the transaction relationship strength between any two enterprises, and traversing all enterprises to obtain an adjacency matrix of the transaction relationship strength; The freshness (Fre) is used to represent whether there is a transaction relationship between two enterprises within a benchmark date, and is valued at 1 if there is a transaction relationship, otherwise is valued at 0; the contract relationship (Con) is used to represent whether the transaction volume between two enterprises within a benchmark date reaches a set benchmark value, and is valued at 1 if it reaches the benchmark value, otherwise is valued at 0; Using a ucinet visualization tool to convert the adjacency matrix into an enterprise transaction relationship social network model, wherein the nodes on the enterprise transaction relationship social network model represent different enterprises, the connection between the nodes represents the existence of a transaction relationship between two enterprises, and the data on the connection represents the business correlation degree between the nodes; Based on the enterprise transaction relationship social network model, the degree centrality result of the node corresponding to each enterprise is calculated and used as the enterprise business correlation degree of the enterprise.
2. The classification hierarchy method of claim 1, wherein, The preprocessing process of the transaction data comprises: The consumption amount is represented by a variable amount, amount = quantity * unit price; wherein quantity represents the order quantity, and unit price represents the unit price; The purchase date and purchase time are extracted from the invoice date to distinguish the transactions created by the enterprise at different time points; Invalid data is deleted.
3. The classification hierarchy method of claim 2, wherein, The calculation of R, F, and M values based on the preprocessed transaction data in units of enterprise ID comprises: Selecting the enterprise ID, calculating the R value of the enterprise based on the interval time between the latest purchase date of the enterprise and the benchmark date; Statistics of the number of transactions of the enterprise in a unit time after the benchmark date, obtaining the F value of the enterprise; wherein the different transaction IDs between the enterprise and the same transaction object in the same unit time are regarded as a transaction; Statistics of the consumption amount of the products purchased by the enterprise after the benchmark date, obtaining the M value of the enterprise.
4. The classification hierarchy method of claim 1, wherein, The initial sensitivity classification result is optimized based on the business correlation degree of the enterprise and the weight thereof; and is represented as: S' ij = S ij + C o w Wherein, S ij represents the sensitivity of the set A ij The initial sensitivity of A ij represents the label set of the traceability sensitive data with the first sub-class j and the sensitivity level i; C o represents the enterprise business correlation degree; w represents the enterprise business correlation degree weight.
5. The classification grading method according to any one of claims 1 to 4, characterized in that, Further comprising: Storing all the traceability sensitive data and the final classification and grading results thereof on the blockchain.
6. A classification hierarchy system for tracing sensitive data, characterized by, A computer program product for performing the classification and grading method according to claim 1, comprising: A classification module for sequentially performing primary classification and secondary classification on all the traceability sensitive data; A construction module for constructing grading indicators; wherein the grading indicators at least include the business correlation degree and the weight thereof, and one or any combination of the data sensitivity, the data importance, and the influence degree of data loss or leakage; A grading module for constructing a data classification model based on the grading indicators other than the business correlation degree and the weight thereof, to perform initial sensitivity grading on the traceability sensitive data under each primary sub-class; And for optimizing the initial sensitivity grading result based on the business correlation degree and the weight thereof, to obtain the final classification and grading result of each piece of traceability sensitive data.
7. A storage medium, characterized by A computer program product for performing the classification and grading method according to claim 1, comprising:
8. An electronic device, comprising: One or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include a program for performing the classification and grading method of the traceability sensitive data according to any one of claims 1-5. Comprise: One or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include a program for performing the classification and grading method of the traceability sensitive data according to any one of claims 1-5.
Citation Information
Patent Citations
Business data flow security risk analysis method and system, storage medium and terminal
CN116506217A
Historical town network construction and classification evaluation method combining historical and modern tour data
CN116882828A