A hash and feature selection based hardware and software device fingerprint fine-grained generation method

By using hashing and feature selection methods, representative and stable features are selected, and combined with Bayesian networks to generate fingerprints of power equipment, the problems of duplicate and easily counterfeited identification codes are solved, and efficient and secure equipment identification and management are achieved.

CN119537891BActive Publication Date: 2025-11-18BEIJING UNIV OF POSTS & TELECOMM +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411419780.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-12
Publication Date
2025-11-18
Estimated Expiration
2044-10-12

AI Technical Summary

Technical Problem

Traditional methods of identifying electrical equipment are prone to duplicate identification codes. Since the meaning of the codes is publicly available, they are easily counterfeited, leading to system intrusion and losses.

Method used

A hash-based and feature-selection method is adopted to select representative and stable features through information entropy and autocorrelation coefficient to generate device fingerprints, and then use hash codes and Bayesian networks for fine-grained identification and differentiation.

Benefits of technology

It improves the accuracy and security of device identification, reduces computing and storage costs, enhances the system's adaptability and fraud prevention capabilities, and ensures the uniqueness and reliability of device identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119537891B_ABST
    Figure CN119537891B_ABST
Patent Text Reader

Abstract

The application discloses a hardware and software device fingerprint fine-grained generation method based on hash and feature selection and belongs to the technical field of power network hardware and software device fingerprint identification. Based on information entropy and autocorrelation coefficient, features used for generating device fingerprints are screened out from a large number of features, representative and stable features are extracted from the device fingerprints, the selected feature parameters are converted into hash codes, the hash codes are used for device fingerprint generation and matching steps, and through the fine-grained device fingerprint generation step based on feature selection, hash codes and a Bayesian network, different devices can be accurately identified and distinguished in a hierarchical manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of fingerprint generation technology for power network hardware and software devices, and relates to a fine-grained method for generating fingerprints of hardware and software devices based on hashing and feature selection. Background Technology

[0002] The Power Internet is a critical network infrastructure that meets the development needs of intelligent power grid hardware and software, characterized by low latency, high reliability, and wide coverage. It represents an emerging business model and application paradigm formed by the deep integration of next-generation information and communication technologies with the power industry. By building platforms based on technologies such as the Internet of Things (IoT), cloud computing, and big data, the Power Internet enables interconnection and interoperability of power grid equipment, data sharing, intelligent operation and maintenance, and smart electricity consumption, providing comprehensive information support and intelligent services for the power industry. The Power Internet identification system is a crucial component of its network architecture, serving as the central hub supporting interconnection and interoperability. The core of this system lies in the generation of fingerprints for hardware and software devices within the power environment. Device fingerprints are unique identifiers that can identify physical assets such as machines and products, as well as virtual resources such as systems and software. Traditional identification methods are prone to duplicate identifiers, and the meanings of these codes are publicly available, making them easy to guess and counterfeit, potentially leading to serious system intrusions and even significant losses. Summary of the Invention

[0003] This invention addresses the problems of existing technologies by providing a fine-grained method for generating hardware and software device fingerprints based on hashing and feature selection.

[0004] A fine-grained method for generating hardware and software device fingerprints based on hashing and feature selection includes the following steps: based on information entropy and autocorrelation coefficient, features for generating device fingerprints are selected from numerous features; representative and stable features are extracted from the device fingerprints; the selected feature parameters are converted into hash codes for use in device fingerprint generation and matching steps; through the fine-grained device fingerprint generation steps based on feature selection, hash codes, and Bayesian networks, different devices are accurately identified and distinguished hierarchically.

[0005] A method for generating fine-grained fingerprints of hardware and software devices based on hashing and feature selection further includes the following steps:

[0006] Step S1: When the system receives a fingerprint generation request, it calls the data interface or the backend database according to the request type.

[0007] Step S2: Extract different feature data for different modules, perform algorithm analysis on the extracted feature data, filter redundant and invalid attribute features, and segment coarse-grained attributes.

[0008] Step S3: Perform hash operations on the extracted attribute features to generate a fine-grained fingerprint of the device.

[0009] Step S4: Transmit the fingerprint to a Bayesian network constructed from historical data for classification and feed the results back to the server.

[0010] Step S5: Merge the fingerprints of different modules of the device in the order of chip, module and operating system to generate device fingerprint and add corresponding verification code.

[0011] The advantages of this invention are: it combines feature selection and hashing steps to achieve more efficient and accurate device identification and classification.

[0012] Improved accuracy: This method improves the accuracy of device identification by selecting the most representative and distinguishable features from the device fingerprint through a feature selection step.

[0013] Reduced computational costs: By utilizing a hashing step, selected features can be converted into compact hash codes, thereby reducing storage and computational costs. Compared to traditional device fingerprint generation methods, this approach is more efficient in maintaining device databases and performing real-time identification.

[0014] Enhanced security: Fine-grained device fingerprint generation provides better protection against fraud and unauthorized access. More accurate device identification helps the system better recognize trusted devices, thus enhancing security.

[0015] High adaptability: Due to the use of feature selection and hashing techniques, this method is highly adaptable to different environments and scenarios. It can be effectively applied and deployed on mobile devices, in network communications, and in IoT environments.

[0016] Reducing the false recognition rate: Through carefully designed feature selection and hashing techniques, the false recognition rate can be effectively reduced, improving the stability and reliability of the system. This means the system can more reliably identify devices, reducing the negative impact of misjudgments.

[0017] The proposed feature-based fingerprint generation scheme can effectively extract features from different types of devices. The features of each type of device are representative. In addition, the diversity of hash functions makes the generated fingerprints more random, making it difficult to find patterns and thus difficult to counterfeit.

[0018] This invention can integrate the inherent attributes and behavioral characteristics of a device to generate a device identifier, which makes it easier to manage and update device fingerprints and authenticate device access.

[0019] The device fingerprint generation scheme uses a hash function, which can prevent duplicate device fingerprints from appearing in a system and ensure the uniqueness of device identifiers.

[0020] Recording the attribute elements of the generated fingerprint facilitates subsequent monitoring and analysis of device behavior, enabling timely updates to the device fingerprint.

[0021] By inferring the relationships between attributes based on changes in the original fingerprint data, a Bayesian network can be constructed to classify devices after fingerprint generation and distinguish whether the device is joining the network for the first time.

[0022] This invention proposes a fine-grained method for generating hardware and software device fingerprints based on hashing and feature selection, which achieves efficient and accurate device identification through feature selection and hashing techniques.

[0023] First, the feature selection step uses information entropy and autocorrelation coefficients to filter out the most representative and discriminative features from the device fingerprint. This process effectively filters out redundant and invalid attributes, ensuring that the selected features accurately reflect the uniqueness of the device and improving the accuracy of device identification. This method allows for the selection of the few most useful features for device identification from a large pool of data, reducing the complexity and computational load of data processing and improving the overall performance of the system.

[0024] Secondly, converting selected features into compact hash codes significantly reduces storage and computation costs. Compared to traditional device fingerprinting methods, this approach is more efficient in maintaining device databases and performing real-time identification. The use of hash codes not only makes the data more compact, effectively reducing storage space usage, but also accelerates data processing and transmission. Simultaneously, the generated fine-grained device fingerprints better prevent fraudulent activities and unauthorized access, enhancing system security. Fine-grained device fingerprints allow for precision down to every detail of the device, preventing any spoofing or imitation and ensuring the reliability and security of device identification.

[0025] Finally, classifying device fingerprints using Bayesian networks accurately determines the device category and the uniqueness of the fingerprint, further improving the intelligence and automation of device management. Bayesian networks utilize historical data and empirical knowledge to build a powerful classification model that can quickly and accurately classify devices after fingerprint generation and determine whether the device is joining the network for the first time. Furthermore, recording the attribute elements of the generated fingerprint facilitates subsequent monitoring and analysis of device behavior, allowing for timely updates to the device fingerprint and ensuring the accuracy and timeliness of device identification. This continuous monitoring and dynamic updating mechanism enables the system to respond quickly to changes, improving the flexibility and adaptability of device management.

[0026] This fine-grained method for generating hardware and software device fingerprints based on hashing and feature selection not only solves the problems of duplicate identifiers and susceptibility to counterfeiting in traditional device fingerprint generation methods, but also significantly improves the accuracy and security of device identification. Its superior performance in adaptability, efficiency, and security makes this method promising for broad applications, effectively applicable to power network hardware and software scenarios, providing a comprehensive and reliable device identification and management solution. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. As shown in the figures:

[0028] Figure 1 This is a flowchart of the fingerprint generation method for power grid hardware and software devices according to the present invention.

[0029] Figure 2 This is a flowchart illustrating the feature selection algorithm in the fingerprint generation method for power grid hardware and software devices of the present invention.

[0030] Figure 3 This is a flowchart illustrating the fingerprint Bayesian network application of the power grid hardware and software devices of the present invention.

[0031] Figure 4 This is a parameter relationship diagram for the fingerprint Bayesian network application of the power grid hardware and software devices of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Example 1: As Figure 1 As shown, a fine-grained method for generating software and hardware device fingerprints based on hashing and feature selection can generate uniquely encoded fingerprint identifiers for power equipment. Furthermore, it can integrate features from different levels of modules within the power grid software and hardware equipment to generate a device fingerprint with encrypted encoding meaning, thereby improving the security of the power grid software and hardware equipment system.

[0034] Includes the following steps:

[0035] Step S1: When the system receives a fingerprint generation request, it calls the data interface or the backend database according to the request type.

[0036] Step S2: Extract different feature data for different modules, perform algorithm analysis on the extracted feature data, filter redundant and invalid attribute features, and segment coarse-grained attributes.

[0037] Step S3: Perform hash operations on the extracted attribute features to generate a fine-grained fingerprint of the device.

[0038] Step S4: Transmit the fingerprint to a Bayesian network constructed from historical data for classification and feed the results back to the server.

[0039] Step S5: Merge the fingerprints of different modules of the device in the order of chip, module and operating system to generate device fingerprint and add corresponding verification code.

[0040] like Figure 2 As shown, the data preprocessing steps include: providing the original dataset to Onehot encoded categorical data, imputing missing values, and collecting features to filter the dataset.

[0041] The feature selection process includes: feature discrimination and stability assessment, feature importance scoring, and differential selection.

[0042] The device fingerprint generation process includes: the filtered feature step, and the device fingerprint generation + check code generation step.

[0043] like Figure 3 As shown, the initial steps involve extracting features from power equipment modules and sending them to the server. Fine-grained device fingerprints are then generated for each module. These fingerprints are compared with an existing fingerprint database to generate feature vectors, which are then input into a Bayesian network for classification to distinguish between equipment categories and whether the equipment is new or old. The server then provides the results.

[0044] like Figure 4 As shown, the chip connects to the module, size, supplier, function and power consumption, the module connection address, communication password, port number and communication baud rate, the operating system connection system version, network traffic and kernel parameters, and the above connections are judged before output.

[0045] Example 2: Figure 1 , Figure 2 , Figure 3 and Figure 4 As shown, a method for generating hardware and software device fingerprints based on hashing and feature selection includes the following steps:

[0046] Step S1: When the system receives a fingerprint generation request, it calls the data interface or the backend database according to the request type.

[0047] Step S2: Extract different feature data for different modules, perform algorithm analysis on the extracted feature data, filter redundant and invalid attribute features, and segment coarse-grained attributes.

[0048] Step S3: Perform hash operations on the extracted attribute features to generate a fine-grained fingerprint of the device.

[0049] Step S4: Transmit the fingerprint to a Bayesian network constructed from historical data for classification and feed the results back to the server.

[0050] Step S5: Merge the fingerprints of different modules of the device in the order of chip, module and operating system to generate device fingerprint and add corresponding verification code.

[0051] In step S1, when an encoding request is received, data is retrieved from the database or API through the corresponding data encoding interface.

[0052] In step S2, the coarse-grained attributes of different level modules of the segmentation device are processed by pruning and segmenting the complex attribute features to achieve fine-grained features.

[0053] In step S3, the data in different modules of the device are hashed to generate a fine-grained fingerprint of the device.

[0054] In step S2, the method for selecting the features used to generate the device fingerprint is as follows: First, the acquired data is preprocessed to supplement the missing values ​​of some features and the categorical data is encoded; then, the discriminative power and stability of the features are calculated and weighted and sorted to determine the features used to generate the device fingerprint.

[0055] The method for calculating the discriminative power of features is to evaluate the amount of information in each feature in the dataset using information entropy. The higher the information entropy, the more information the feature contains, indicating that the data is more uncertain and difficult to distinguish; conversely, the lower the information entropy, the more certain and easily distinguishable the data for that feature is, and it does not have high discriminative power.

[0056] The calculation method is as follows Where X represents the characteristics of hardware and software devices in the power environment, x i These are the feature parameters; the final numerical values ​​are used to identify the discriminative power of the features.

[0057] The stability of features is assessed using the autocorrelation coefficient (ACF) to evaluate the stability of each feature field in the dataset. A high stability means that its value changes little at different times, more accurately reflecting the inherent characteristics of the device and thus improving the uniqueness and robustness of the device fingerprint. When the absolute value of the ACF coefficient is greater than 0.8, it indicates strong correlation in the data, potentially with some volatility; when the absolute value of the ACF coefficient is less than 0.2, it indicates almost no correlation in the data, and the changes tend to be stable. The calculation method is as follows:

[0058]

[0059] The final numerical value is used to indicate the stability of the feature.

[0060] Let X t X is a characteristic of hardware and software devices in the power environment. t-k Features after translation The method for feature selection, which combines the discriminative power and stability of a set of feature parameters, is as follows: Stability and discriminative power are often contradictory in feature selection. Some features with high stability may contribute little to the discriminative power between different classes, while some features with high discriminative power may be unstable and susceptible to noise and interference. This method aims to balance stability and discriminative power, selecting the optimal subset of features, as expressed by the formula: feature w =weight·H w -(1-weight)·|ACF w Where weight represents the weights for discrimination and stability, and generally, the weight for discrimination is considered greater than the weight for stability; H w and ACF w These are the information entropy and ACF coefficient of the w-th feature, respectively. The features with the highest final scores will be selected as fingerprint generation elements.

[0061] The final determined feature method for generating device fingerprints is differential filtering. Specifically, to avoid loss of generality, the attribute features and behavioral features are first arranged from largest to smallest. The difference k of the i-th feature... i The calculation formula is: k i =feature i+1 -feature i Among them, feature i+1 The feature representing the (i+1)th feature. iLet represent the i-th feature. The case with the highest score means that the feature score drops sharply, which is the end position for selecting the feature.

[0062] In this way, the most important features can be selected, while those features that have less impact on device fingerprinting can be ignored.

[0063] In step S3, the hash algorithm for generating the device fingerprint includes SHA-1, SHA-256, SHA-384, and SHA-512 algorithms.

[0064] In step S4, the Bayesian network structure is first constructed using historical data and empirical knowledge.

[0065] In step S4, the training data is preprocessed and the parameters of the Bayesian network are learned.

[0066] In step S4, the fingerprint is transmitted to a Bayesian network model trained based on historical data, and the device category and whether the device fingerprint is generated for the first time are determined. The determination result is transmitted to the server.

[0067] In step S5, after combining the fingerprints of different modules of the device, a check code is added as a suffix. The check code algorithm is either a parity check algorithm, a Hamming check algorithm, or a cyclic redundancy check algorithm.

[0068] Generate a checksum using a predefined checksum algorithm.

[0069] This invention simulates and generates corresponding simulation data for electricity meters and terminals in a real-world power grid scenario. To accurately monitor and analyze the power grid's operational status, based on the real-world scenario and to select the features to be extracted from the modules, the following attributes were selected as the filtering criteria: communication protocol, communication method, module ID, manufacturer, wiring method, rated current, and rated voltage.

[0070] First, the data in the dataset is preprocessed to supplement missing values ​​of some features and encode categorical data.

[0071] To reduce the impact of encoding on the algorithm, the encoded data is normalized to the [0,1] interval to facilitate subsequent data stability analysis.

[0072] In data preprocessing, the information entropy algorithm is used to measure the discriminative power of the data. Let X represent the characteristics of hardware and software devices in the power environment, x... i The data, using the information entropy algorithm as a feature parameter, has the following discriminative measure:

[0073]

[0074] Algorithm: Information Entropy Algorithm. Used to evaluate the discriminative power of data. Based on the frequency distribution of the data, it determines the diversity and entropy of the data by calculating the probability distribution of different values ​​in the dataset, thereby calculating the uncertainty and disorder of the attribute data.

[0075] Input: D: A dataset containing n objects, each with m-dimensional attribute parameters.

[0076] Output: Discrimination index.

[0077] method:

[0078] By analyzing the data of n objects in each of the m-dimensional features using the information entropy algorithm, the final results are sorted in descending order.

[0079] Considering the ease of expressing this metric in reality and its rationality in subsequent algorithms, it was ultimately decided to store the obtained discrimination evaluation metrics in descending order. By sorting the evaluation metrics in descending order, it is convenient to identify and extract the data or features with the highest discrimination, which facilitates subsequent algorithms and analysis.

[0080] The results are shown in Table 1 below:

[0081] Table 1

[0082]

[0083]

[0084] The stability of the data is measured using the autocorrelation coefficient (ACF) algorithm. Let X... t X is a characteristic of hardware and software devices in the power environment. t-k Features after translation Given a set of characteristic parameters, the stability of the data using the autocorrelation coefficient (ACF) algorithm is as follows:

[0085]

[0086] Input: D: A dataset containing n objects, each object having m-dimensional attribute parameters, and k is the data shift amount.

[0087] Output: Stability evaluation metrics.

[0088] method:

[0089] 1) Set the corresponding k value according to the amount of data in the dataset.

[0090] 2) By analyzing the data of n objects in each of the m-dimensional features using the information entropy algorithm, the results are finally sorted in descending order.

[0091] By comprehensively comparing simulated data, 3 was selected as the data offset for k, and the resulting discrimination evaluation indicators were stored in ascending order. Storing the indicators in ascending order allows for the rapid identification of parameters with good stability as references.

[0092] The results are shown in Table 2 below:

[0093] Table 2

[0094]

[0095]

[0096] Based on the calculated discrimination and stability, the two evaluation criteria are weighted and fused for judgment. The specific formula is shown below:

[0097] feature w =weight·H w -(1-weight)·|ACF w |

[0098] Where weight is the weight, H w and ACF w These are the information entropy and ACF coefficient of the w-th feature, respectively.

[0099] By comprehensively comparing simulated data, the data's discriminative power is a more important factor to consider when generating device fingerprints, so the weight is set to 0.6; at the same time, the calculated comprehensive evaluation score will be stored in descending order.

[0100] The specific results are shown in Table 3 below:

[0101] Table 3

[0102]

[0103]

[0104] The parameters used to generate fingerprints are finally extracted through a differential algorithm, and the selected attributes are stored in the database when the parameters are generated, which facilitates the subsequent monitoring of the device itself and its modules.

[0105] By using fine-grained features selected based on device hierarchy through differential mapping, the topology of a Bayesian network is constructed, such as... Figure 4 As shown.

[0106] The model is then trained using the training data to learn its parameters, thereby determining the conditional probability distribution of each node.

[0107] Considering that the collected training dataset may not be complete and cannot obtain accurate prior probabilities, the maximum likelihood estimation method is used for parameter learning.

[0108] The feature vector of the subsequent fingerprint change and the device fingerprint vector are respectively input into two trained Bayesian network models for testing. Finally, the device type and whether the device is a new device are determined by using a pre-set threshold. The server then provides feedback on the judgment result.

[0109] After determining the device category and successfully generating fingerprints for each module, the fine-grained fingerprints of each module are merged in the order of chip, onboard module, and operating system, generating the device fingerprint and appending the corresponding verification code. This fingerprint is then stored in the database corresponding to the device category.

[0110] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for fine-grained generation of hardware and software device fingerprints based on hashing and feature selection, characterized in that, Based on information entropy and autocorrelation coefficient, features for generating device fingerprints are selected from numerous features. Representative and stable features are extracted from the device fingerprints. The selected feature parameters are converted into hash codes for device fingerprint generation and matching steps. Through fine-grained device fingerprint generation steps based on feature selection, hash codes and Bayesian networks, different devices are accurately identified and distinguished hierarchically. The method for selecting features for generating device fingerprints is as follows: First, the acquired data is preprocessed to supplement missing values ​​of some features, and categorical data is encoded; then, the discriminative power and stability of the features are calculated and weighted to determine the features used to generate device fingerprints. The method for calculating the discriminative power of features is as follows: Information entropy is used to evaluate the amount of information contained in each feature in the dataset. The higher the information entropy, the more information the feature contains, indicating that the data is more uncertain and difficult to distinguish; conversely, the lower the information entropy, the more certain and easily distinguishable the data for that feature, but it does not have high discriminative power. The calculation method is as follows: Where X represents the characteristics of hardware and software devices in the power environment, x i These are the feature parameters; the final numerical values ​​are used to identify the discriminative power of the features. The stability of features is assessed using the Autocorrelation Coefficient (ACF) to evaluate the stability of each feature field in the dataset. A feature field with high stability exhibits smaller value changes at different times, more accurately reflecting the inherent characteristics of the device and thus improving the uniqueness and robustness of the device fingerprint. When the absolute value of the ACF coefficient is greater than 0.8, it indicates strong correlation in the data, potentially with some volatility; when the absolute value of the ACF coefficient is less than 0.2, it indicates almost no correlation in the data, and the changes tend to be stable. The calculation method is as follows: The final numerical value is used to identify the stability of the feature, X. t X is a characteristic of hardware and software devices in the power environment. t-k Features after translation Let k be the average of a set of feature parameters, k be the data shift amount, and k be the dataset of N objects. The method for feature selection based on a combination of discriminative power and stability is as follows: Stability and discriminative power are often contradictory in feature selection. Some features with high stability may contribute little to the discriminative power between different categories, while some features with high discriminative power may be unstable and easily affected by noise and interference. A balance is struck between stability and discriminative power to select the optimal subset of features. The formula is: feature w =weight·H w -(1-weight)·|ACF w Where weight represents the weights for discrimination and stability, with discrimination having a greater weight than stability; H w and ACF w These are the information entropy and ACF coefficient of the w-th feature, respectively. Finally, features with high scores are selected as fingerprint generation elements. The final determined feature method for generating device fingerprints is as follows: Differential filtering. To avoid loss of generality, the attribute features and behavioral features are first arranged from largest to smallest, and the difference k of the i-th feature is... i The calculation formula is: k i =feature i+1 -feature i Among them, feature i+1 The feature representing the (i+1)th feature. i This represents the case where the score of the i-th feature is the highest, which means that the feature score drops sharply. This means selecting the end position of the feature. In this way, we can filter out the most important features and ignore those features that have little impact on the device fingerprint. The fingerprint is transmitted to a Bayesian network model trained on historical data to determine the device category and whether the device's fingerprint is generated for the first time. The determination result is then transmitted to the server.

2. The method for fine-grained generation of hardware and software device fingerprints based on hashing and feature selection according to claim 1, characterized in that, It also includes the following steps: Step S1: When the system receives a fingerprint generation request, it calls the data interface or the backend database according to the request type; Step S2: Extract different feature data for different modules, perform algorithm analysis on the extracted feature data, filter redundant and invalid attribute features, and segment coarse-grained attributes; Step S3: Perform hash operations on the extracted attribute features to generate a fine-grained fingerprint of the device; Step S4: Transmit the fingerprint to a Bayesian network constructed from historical data for classification and feed the results back to the server; Step S5: Merge the fingerprints of different modules of the device in the order of chip, module and operating system to generate device fingerprint and add corresponding verification code.

3. The method for fine-grained generation of hardware and software device fingerprints based on hashing and feature selection according to claim 2, characterized in that, Step S1 also includes the following steps: upon receiving an encoding request, data is retrieved from the database or API interface through the corresponding data encoding interface. Step S2 also includes the following steps: segmenting the coarse-grained attributes of different level modules of the segmentation device, and pruning and segmenting the complex attribute features to achieve fine-grained feature processing. Step S3 also includes the following step: hashing the data from different modules of the device to generate a fine-grained fingerprint of the device. Step S4 also includes the following steps: first, the Bayesian network structure is constructed using historical data and empirical knowledge.

4. The method for fine-grained generation of hardware and software device fingerprints based on hashing and feature selection according to claim 2, characterized in that, Step S3 also includes the following steps: generating a device fingerprint using a hash algorithm, including SHA-1, SHA-256, SHA-384, and SHA-512 algorithms.

5. The method for fine-grained generation of hardware and software device fingerprints based on hashing and feature selection according to claim 2, characterized in that, Step S4 also includes the following steps: preprocessing the training data and learning the parameters of the Bayesian network.

6. The method for fine-grained generation of hardware and software device fingerprints based on hashing and feature selection according to claim 2, characterized in that, Step S5 also includes the following steps: after combining the fingerprints of different modules of the device, a check code is added as a suffix. The check code algorithm is a parity check algorithm, a Hamming check algorithm, or a cyclic redundancy check algorithm.

7. The method for fine-grained generation of hardware and software device fingerprints based on hashing and feature selection according to claim 1, characterized in that, In data preprocessing, the information entropy algorithm is used to measure the discriminative power of the data. Let X be a feature of the hardware and software devices in the power environment, x i The data, using the information entropy algorithm as a feature parameter, has the following discriminative measure: Algorithm: The information entropy algorithm is used to evaluate the discriminative power of data. Based on the frequency distribution of the data, it determines the diversity and entropy of the data by calculating the probability distribution of different values ​​in the dataset, thereby calculating the uncertainty and disorder of the attribute data. Input: D: A dataset containing n objects, each object having m-dimensional attribute parameters. Output: Discrimination index, method: By analyzing the data of n objects for each feature in m dimensions using the information entropy algorithm, the results are finally sorted in descending order. Considering the ease of expressing this indicator in reality and the rationality of its use in subsequent algorithms, it was ultimately decided to store the obtained discrimination evaluation indicators in descending order. By sorting the evaluation indicators in descending order, the data or features with the highest discrimination can be identified and extracted.

8. The method for fine-grained generation of hardware and software device fingerprints based on hashing and feature selection according to claim 1, characterized in that, In data preprocessing, the stability of the data is measured using the autocorrelation coefficient (ACF) algorithm, where X is the autocorrelation coefficient. t X is a characteristic of hardware and software devices in the power environment. t-k Features after translation Given a set of characteristic parameters, the stability of the data using the autocorrelation coefficient (ACF) algorithm is measured as follows: Input: D: A dataset containing N objects, each object having m-dimensional attribute parameters, and k being the data shift amount. Output: Stability evaluation metrics method: 1) Set a corresponding k value based on the amount of data in the dataset. 2) By analyzing the data of N objects for each feature in m dimensions using the information entropy algorithm, the results are finally sorted in descending order. By comprehensively comparing simulated data, 3 was selected as the k-data offset, and the resulting discrimination evaluation index was stored in ascending order.

Citation Information

Patent Citations

  • Equipment fingerprint generation method and equipment

    CN116074051A

  • Unsupervised optimal anomaly detection model selection method and system for power data

    CN116167004A