Method and system for generating fingerprint of IoT device
By collecting and processing protocol slogan data of IoT devices, using mask language model and clustering technology to generate stable device fingerprints, the problems of low detection efficiency and unstable matching of IoT devices are solved, and efficient identification and automatic update of device fingerprints are achieved.
Patent Information
- Application Number
- CN202510846977.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The existing IoT device detection technology is low in efficiency and poor in real time, and cannot recognize silent devices. The fingerprint matching is unstable, and the manual maintenance cost is high.
The protocol slogan data is collected through the network scanning tool, and the word participle is used to decompose it into a byte-level subword data set. The pre-trained mask language model learns the semantics and subword context relationships, optimizes the embedding stability, generates slogan-level embeddings, and uses the supervised comparison loss fine-tuning method to generate regular expression device fingerprints and updates the fingerprint library.
It realizes automatic fingerprint generation of active detection equipment, identifying silent equipment, stabilizing fingerprint matching, and reducing manual maintenance costs.
Smart Images

Figure CN120353881B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet of Things security technology, and in particular to a method and system for generating fingerprints for Internet of Things devices. Background Art
[0002] With the rapid development of the Internet of Things (IoT), smart devices have become widely used in smart homes, industrial control, healthcare, and other fields. However, the proliferation of IoT devices also poses serious security risks. Due to default configuration vulnerabilities and unauthorized access, a large number of IoT devices have become entry points for network attacks. Therefore, covert and efficient IoT device detection technology has become a key component of network security protection, asset management, and vulnerability troubleshooting. By identifying the types of IoT devices on the network and their service status, data support can be provided for risk warning, access control, and threat tracing.
[0003] In the existing technology, mainstream IoT device detection technologies mainly rely on two methods: active scanning and passive traffic analysis. Among them, passive traffic analysis technology extracts device traffic characteristics and then uses machine learning, deep learning and other technologies to identify device types. However, this method relies on the active communication behavior of the device, has low detection efficiency and poor real-time performance, and cannot identify silent devices. Active scanning technology sends standardized detection requests to the target network and identifies devices based on the returned information. It mainly relies on manually generated regular expression fingerprint libraries and identifies device types by matching protocol banner data. However, since protocol banners often contain dynamic content such as timestamps, random IDs or minor version numbers, fingerprint matching will be unstable, and manual maintenance of the fingerprint database is very difficult. Experts need to update it to adapt to new devices. With the continuous emergence of new devices and software, a lot of human resources and time must be invested.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention
[0005] In response to the above problems, the present application provides a method and system for generating fingerprints of IoT devices, which can actively detect and optimize technology, realize automatic generation of fingerprints of active detection devices and reduce the manual maintenance cost of device fingerprints.
[0006] To achieve the purpose of this application, this application provides the following technical solutions:
[0007] In a first aspect, the present application provides a method for generating fingerprints of IoT devices, comprising:
[0008] Collect the labeled protocol slogan data of IoT devices through network scanning tools, and decompose the protocol slogan data into byte-level subword datasets through a word segmenter;
[0009] A masked language model is pre-trained on a byte-level subword dataset to learn the relationship between word semantics and subword context. A supervised contrastive loss fine-tuning method is used to optimize the embedding's stability to dynamic content, generating slogan-level embeddings suitable for downstream device fingerprinting tasks.
[0010] Perform dimensionality reduction clustering on slogan-level embeddings to obtain device clusters, extract common substrings from each device cluster, and convert them into regular expression device fingerprints;
[0011] Compare the regular expression device fingerprint with the existing fingerprint database and update the fingerprint database.
[0012] In one possible implementation, the step of collecting the protocol slogan data with labels of IoT devices through a network scanning tool and decomposing the protocol slogan data into byte-level subword datasets through a word segmenter includes:
[0013] Use a port scanning tool to scan the available ports of the IP address, combine it with a protocol scanning tool to capture the protocol banner data, and use a network mapping tool to obtain the labels corresponding to the protocol banner data to obtain the protocol banner data with labels;
[0014] Preprocess the collected slogan data to remove special symbols and HTML structure symbols, and convert uppercase characters to lowercase;
[0015] Train a byte-level tokenizer to decompose labeled protocol slogan data into byte-level subwords.
[0016] In a possible implementation, the step of training a byte-level word segmenter to decompose the protocol slogan data with labels into byte-level subwords includes:
[0017] Decompose each slogan into a sequence of single characters and map them into a sequence of ASCII bytes;
[0018] Initialize the vocabulary containing all single characters and count the frequencies of characters and adjacent byte pairs;
[0019] Merge the most frequent byte pairs in the current dataset into new subwords and update the slogan dataset;
[0020] The merging process is repeated until the vocabulary size reaches a first preset threshold or the number of merging times reaches a second preset threshold.
[0021] In one possible implementation, the steps of pre-training a masked language model using a byte-level subword dataset to learn the relationship between word semantics and subword context, and fine-tuning the model using a supervised contrastive loss method to optimize the embedding's stability to dynamic content, thereby generating a slogan-level embedding suitable for downstream device fingerprinting tasks, include:
[0022] Pre-train the masked language model on a byte-level subword dataset to learn the semantic and contextual relationships of subwords;
[0023] Conduct multiple detections on the same IP address and port at different times, extract protocol banner data and record them as banner pairs, filter the dynamic content in the banner pairs using regular expression rules, and generate static banner data;
[0024] The static slogan pairs generated from the same slogan pair are used as positive sample pairs, and slogans are extracted from different slogan pairs to generate negative sample pairs. The cosine similarity is used for pre-screening to construct positive and negative sample pairs.
[0025] A pre-trained masked language model is used to generate subword embeddings for each slogan. The embeddings of all subwords in the sequence are averaged to generate slogan-level embeddings suitable for downstream device fingerprint recognition tasks. A first loss function is constructed to use supervised contrastive loss to narrow the cosine similarity of positive sample embeddings and push the similarity of negative sample embeddings further away.
[0026] In a possible implementation, the first loss function is:
[0027] ;
[0028] in, is the first loss function, For slogans The embedding vector of For slogans The positive sample embedding of For slogans The embedding vector of the slogan Negative sample embedding, Calculate the cosine similarity, is the number of positive sample pairs, is the temperature parameter.
[0029] In one possible implementation, the masked language model is pre-trained using a cross-entropy loss function; the cross-entropy loss function is:
[0030] ;
[0031] in, is the cross entropy loss function, M represents the set of masked subwords, is the true value of subword i, is the probability predicted by the model based on the context.
[0032] In one possible implementation, the steps of performing dimensionality reduction clustering on slogan-level embeddings to obtain device clusters, extracting common substrings from each device cluster, and converting them into regular expression device fingerprints include:
[0033] We perform linear dimensionality reduction on slogan-level embeddings through principal component analysis, and compare and verify this with unified manifold approximation and projective nonlinear dimensionality reduction.
[0034] Hierarchical density clustering is used to cluster the reduced slogan-level embeddings to generate device clusters.
[0035] Counting the metadata tag distribution of the slogans within each cluster and calculating a tag consistency score; if the consistency score is higher than a fourth preset threshold, assigning the most common tag category to the cluster as the device information tag of the cluster;
[0036] Samples are selected from each cluster to extract common substrings, generating device fingerprints in the form of regular expressions. Combined with the device information labels of the cluster, structured fingerprint rules are generated.
[0037] In a possible implementation, after the step of selecting samples from each cluster to extract common substrings and generate a device fingerprint in the form of a regular expression, the method further includes:
[0038] Apply the regular expression to all slogans in the cluster and calculate the coverage. If the coverage is less than the fifth preset threshold, reselect samples or adjust the regular expression rules until the coverage is greater than or equal to the fifth preset threshold.
[0039] In one possible implementation, the step of comparing the regular expression device fingerprint with an existing fingerprint database and updating the fingerprint database includes:
[0040] The newly generated regular expression fingerprint is matched one by one with the rules in the original fingerprint database for semantic equivalence;
[0041] The fingerprint that matches successfully inherits the corresponding device information in the database;
[0042] The fingerprint rules that are not successfully matched are used as new rules to update the fingerprint rule base.
[0043] In a second aspect, the present application further provides an IoT device fingerprint generation system for executing the above-mentioned IoT device fingerprint generation method, the system comprising:
[0044] The data collection module is used to collect the protocol slogan data with labels of IoT devices through network scanning tools, and decompose the protocol slogan data into byte-level subword datasets through a word segmenter;
[0045] The embedding generation module is used to pre-train a masked language model using a byte-level subword dataset to learn the relationship between word semantics and subword context. It also uses a supervised contrastive loss fine-tuning method to optimize the embedding's stability to dynamic content, generating slogan-level embeddings suitable for downstream device fingerprinting tasks.
[0046] The fingerprint generation module is used to perform dimensionality reduction clustering on slogan-level embeddings to obtain device clusters, extract common substrings from each device cluster, and convert them into regular expression device fingerprints;
[0047] The fingerprint library update module is used to compare the regular expression device fingerprint with the existing fingerprint database and update the fingerprint library.
[0048] The technical solution provided by this application may have the following beneficial effects:
[0049] The method and system for generating fingerprints for IoT devices provided in this application can obtain protocol slogan data through active detection, use supervised contrast loss to fine-tune the masked language model to generate stable semantic embedding to cope with dynamic content, and combine dimensionality reduction and HDBSCAN clustering to automatically generate and update device fingerprints in the form of regular expressions, thereby realizing automatic generation of IoT device fingerprints based on active detection, thereby being able to identify silent devices, stabilize fingerprint matching, and reduce the cost of manual fingerprint maintenance.
[0050] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not limit the present invention. Obviously, the drawings described below are only some embodiments of the present disclosure. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0052] Figure 1 A schematic diagram of a flow chart of a method for generating fingerprints for an IoT device provided in an embodiment of the present application;
[0053] Figure 2 A flowchart of step S100 of a method for generating a fingerprint for an IoT device provided in an embodiment of the present application;
[0054] Figure 3A flowchart of step S200 of a method for generating a fingerprint for an IoT device provided in an embodiment of the present application;
[0055] Figure 4 A flowchart of step S300 of a method for generating a fingerprint for an IoT device provided in an embodiment of the present application;
[0056] Figure 5 A schematic diagram of the process of step S400 of a method for generating a fingerprint for an IoT device provided in an embodiment of the present application;
[0057] Figure 6 A schematic diagram of the structure of a fingerprint generation system for an IoT device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0058] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0059] This example embodiment first provides a method for generating fingerprints for IoT devices with dynamic interference power management under conditions of sparse feedback and missing observations. Figure 1 As shown in , the method for generating fingerprints of IoT devices may include the following steps:
[0060] Step S100: collecting protocol slogan data with labels of IoT devices through a network scanning tool, and decomposing the protocol slogan data into byte-level subword data sets through a word segmenter.
[0061] Step S200: Pre-training a masked language model with a byte-level sub-word dataset to learn the relationship between word semantics and sub-word contexts, and optimizing the embedding stability to dynamic content through supervised contrast loss fine-tuning to generate slogan-level embeddings suitable for downstream task device fingerprint recognition.
[0062] Step S300: performing dimensionality reduction clustering on the slogan-level embedding to obtain device clusters, and extracting common substrings from each device cluster and converting them into regular expression device fingerprints.
[0063] Step S400: Compare the regular expression device fingerprint with the existing fingerprint database and update the fingerprint database.
[0064] Through the above-mentioned IoT device fingerprint generation method, it is possible to distinguish the dynamic and static parts in the slogan, propose a fine-tuning method based on supervised contrast loss, optimize the stability of embedding, enhance the robustness to dynamic content, and automatically generate fingerprint rules through dimensionality reduction and clustering, supplement and update the fingerprint rule library, and adapt to new devices and software, so as to achieve the acquisition of protocol slogan data through active detection, supervise the contrast loss fine-tuning mask language model to generate stable semantic embedding to cope with dynamic content, and combine dimensionality reduction and HDBSCAN clustering to automatically generate and update device fingerprints in the form of regular expressions, realize the identification of silent devices, and improve the stability of fingerprint matching.
[0065] Below, we will refer to Figures 2 to 5 Each step of the above-mentioned method for generating fingerprints of IoT devices in this example implementation is described in more detail.
[0066] In step S100, protocol slogan data with tags of IoT devices are collected through a network scanning tool, and the protocol slogan data are decomposed into byte-level subword data sets through a word segmenter.
[0067] It is understandable that the protocol banner data includes header information of HTTP, FTP, and SSH protocols; and the tag includes device type, device manufacturer, and model data.
[0068] In a possible implementation, step S100 may further include the following sub-steps:
[0069] In step S110, a port scanning tool is used to scan available ports of the IP address, and a protocol scanning tool is used to capture protocol banner data. Then, a network mapping tool is used to obtain labels corresponding to the protocol banner data, thereby obtaining protocol banner data with labels.
[0070] It should be noted that data collection is divided into two parts. The first is the collection of protocol banner data. Zmap is used as a port scanning tool to scan the common ports of the target IP address in the network space, including ports 80, 21 and 22, to obtain a list of open ports; secondly, the protocol scanning tool ZGrab is used to capture the banner data of protocols such as HTTP, FTP, SSH, etc. The banner data includes protocol header information, such as the Server field of HTTP and the welcome message of FTP; finally, device metadata is obtained from the existing network mapping engine to provide preliminary labels for some protocol banners, including device type, manufacturer, and model labels.
[0071] In step S120, the collected slogan data is pre-processed to remove special symbols and HTML structural symbols, and convert uppercase characters into lowercase characters.
[0072] It is understandable that the collected slogan data removes special symbols and line breaks, retains letters, numbers and necessary separators, and converts uppercase characters in the slogans to lowercase to ensure consistent data format.
[0073] In step S130 , a byte-level word segmenter is trained to decompose the protocol slogan data with labels into byte-level subwords.
[0074] It can be understood that the word segmenter training part uses the BPE word segmenter to segment the slogan and generate a byte-level subword sequence.
[0075] Furthermore, step 130 may further include the following sub-steps:
[0076] In step S131, each slogan is decomposed into a single character sequence and mapped into an ASCII byte sequence.
[0077] In step S132, a vocabulary containing all single characters is initialized, and the frequencies of characters and adjacent byte pairs are counted.
[0078] It is understandable that after initialization, the vocabulary is a single character (such as the ASCII character set), and each character in the slogan dataset exists independently.
[0079] In step S133, the most frequent byte pairs in the current data set are merged into new subwords, and the slogan data set is updated.
[0080] It is understandable that the frequency of occurrence of all adjacent character pairs is counted, and the most frequent pairs (such as "ab") are selected and merged into new subwords (such as " <ab>"), update the data set ("ab" in the original slogan is replaced by " <ab>").
[0081] In step S134 , the merging process is repeated until the vocabulary size reaches a first preset threshold or the number of merging times reaches a second preset threshold.
[0082] In step S200, a masked language model is pre-trained using a byte-level sub-word dataset to learn the relationship between word semantics and sub-word contexts. The stability of the embedding to dynamic content is optimized through supervised contrastive loss fine-tuning, generating slogan-level embeddings suitable for fingerprint recognition of downstream task devices.
[0083] It can be understood that the semantic embedding of the protocol slogans is generated using a pre-trained masked language model to capture the contextual relationship of sub-words, and fine-tuned through supervised contrastive loss to optimize the stability of the embedding to dynamic content, generating high-quality embedding representations suitable for downstream task device fingerprint recognition.
[0084] In a possible implementation, step S200 may further include the following sub-steps:
[0085] In step S210 , the masked language model is pre-trained on the byte-level sub-word dataset to learn the semantics and contextual relationships of the sub-words.
[0086] It should be noted that the pre-trained BERT model is used as the masked language model and fine-tuned on the slogan dataset after word segmentation; the fine-tuning dataset is obtained by converting the protocol slogan dataset into a subword sequence using the trained BPE word segmenter, and the maximum subword sequence length of each slogan is restricted; a part of the subwords in each slogan subword sequence is randomly selected for masking operation, and the selected subwords are replaced with a special tag [MASK]; with the goal of predicting masked subwords, the model's semantic understanding of the protocol slogans is optimized.
[0087] Optionally, the model is trained using the cross entropy loss function:
[0088] ;
[0089] in, is the cross entropy loss function, M represents the set of masked subwords, is the true value of subword i, is the probability predicted by the model based on the context.
[0090] In step S220, multiple detections are performed on the same IP address and port at different times, protocol banner data are extracted and recorded as banner pairs, and dynamic content in the banner pairs is filtered using regular expression rules to generate static banner data.
[0091] It is understandable that each pair of banners reflects the response of the same device at different times, and dynamic content may cause differences in banners; the regular expressions used include the timestamp regular expression [0-9]{4}-[0-9]{2}-[0-9]{2}, the random ID regular expression [0-9a-f]{8}, and the dynamic version number regular expression [0-9]+\.[0-9]+\.[0-9]+; dynamic content includes timestamp and random ID.
[0092] In step S230, static slogan pairs generated from the same slogan pair are used as positive sample pairs, slogans are extracted from different slogan pairs to generate negative sample pairs, and cosine similarity is used for pre-screening to construct positive and negative sample pairs.
[0093] It is understandable that static slogans from the same slogan pair are positive samples, and slogans randomly selected from slogans of different devices are negative samples. To ensure the semantic differences of negative samples, pre-calculated cosine similarity is used to filter negative samples, and slogans with similarity below a third preset threshold are retained.
[0094] In step S240, a subword embedding of each slogan is generated through a pre-trained masked language model, and the embeddings of all subwords in the sequence are averaged to generate a slogan-level embedding suitable for downstream task device fingerprint recognition. A first loss function is constructed to use supervised contrast loss to bring the cosine similarity of the positive sample embedding closer and push the similarity of the negative sample embedding further away.
[0095] It should be noted that the subword embedding of each slogan is generated through the fine-tuned BERT model, and the embeddings of all subwords in the sequence are averaged (mean pooling) to generate slogan-level embeddings, reducing the local impact of dynamic content on the embedding.
[0096] Furthermore, the first loss function is:
[0097] ;
[0098] in, is the first loss function, For slogans The embedding vector of For slogans The positive sample embedding of For slogans The embedding vector of the slogan Negative sample embedding, Calculate the cosine similarity, is the number of positive sample pairs, is the temperature parameter.
[0099] It can be understood that the first loss function is used to bring the positive samples closer and push the negative samples further away. In the case of bringing the positive samples closer, the slogan embeddings of the same device at different times after mean pooling are made highly similar in the semantic space, thereby strengthening the consistency of the static core semantics and ignoring the interference of dynamic content; in the case of pushing the negative samples further away, the slogan embeddings of different devices with large semantic differences after cosine filtering are significantly separated in the semantic space, thereby enhancing the semantic recognition between devices.
[0100] In step S300 , dimensionality reduction clustering is performed on the slogan-level embedding to obtain device clusters, and common substrings are extracted from each device cluster and converted into regular expression device fingerprints.
[0101] It should be noted that the dimensionality reduction strategy adopts a combination of principal component analysis (PCA) and unified manifold approximation and projection (UMAP) to ensure the retention of semantic information; the clustering adopts the HDBSCAN algorithm to cluster the embeddings after dimensionality reduction, generate device clusters, and mine groups of slogans with similar semantics.
[0102] In one possible implementation, step S300 may include the following sub-steps:
[0103] In step S310, linear dimensionality reduction is performed on the slogan-level embedding through principal component analysis, and a comparison and verification is performed by combining unified manifold approximation with projected nonlinear dimensionality reduction.
[0104] It should be noted that PCA linear dimensionality reduction: By calculating the covariance matrix of the embedding matrix, the eigenvector with the largest eigenvalue is extracted to construct a low-dimensional projection matrix. The embedding vector is projected into a low-dimensional space through matrix multiplication to generate a reduced-dimensional embedding representation. UMAP nonlinear dimensionality reduction verification: To verify the semantic preservation effect of PCA dimensionality reduction, UMAP is used for nonlinear dimensionality reduction comparison; based on manifold learning, a topological representation of high-dimensional data is constructed, and the distance relationship between data points is retained by optimizing the cross-entropy loss of low-dimensional embeddings; the number of neighbors is set to control the degree of preservation of local structure, and the minimum distance is set to avoid too dense embedding points. The cosine distance is used as a metric to adapt to the semantic characteristics of the embedding vector.
[0105] In step S320, hierarchical density clustering is used to cluster the reduced-dimensional slogan-level embeddings to generate device clusters.
[0106] It should be noted that the clustering uses the HDBSCAN algorithm to cluster the embeddings after dimensionality reduction, generate device clusters, and mine groups of slogans with similar semantics.
[0107] In step S330, the metadata tag distribution of the slogans in each cluster is counted and the tag consistency score is calculated; if the consistency score is higher than the fourth preset threshold, the most common tag category is assigned to the cluster as the device information tag of the cluster.
[0108] It should be noted that for each cluster, the distribution of metadata tags (device type, manufacturer, model) of the slogans in the cluster is counted and the tag consistency score is calculated; if the consistency score is higher than the threshold, the most common tag is assigned to the cluster as the device information tag of the cluster.
[0109] In step S340 , samples are selected from each cluster to extract common substrings, and a device fingerprint in the form of a regular expression is generated. The device information tags of the cluster are combined to generate a structured fingerprint rule.
[0110] It can be understood that several slogan samples with the most similar semantics to the cluster center are selected from each cluster as representative slogans; first, the mean vector of all embedded vectors in the cluster is calculated as the cluster center to represent the semantic characteristics of the cluster, and the cosine similarity of each slogan in the cluster with the cluster center is calculated. After sorting, several slogans with the highest similarity are selected to ensure that the samples are representative; secondly, multiple sequence alignment is performed on the selected slogan samples to extract common substrings, and a multiple sequence alignment algorithm based on dynamic programming is used to perform character-level alignment on the slogans to identify common substrings; finally, according to the semantic structure of the slogan, a regular expression is designed to cover the changing parts such as version number and secondary fields, and the device fingerprint in the form of a regular expression is obtained.
[0111] Structured fingerprint rules are complete fingerprints containing device information and are used for library management. Device fingerprints in the form of regular expressions are the core matching component for real-time identification. Combining the two, structured rules address the need for dynamic fingerprint library updates, while regular expressions address the need for stable and efficient matching. This layered decoupling of storage and execution adapts to the diversity of IoT devices while ensuring high recognition performance.
[0112] The optional, structured fingerprint rule format is as follows:
[0113] .
[0114] Optionally, after step S340, the method further includes:
[0115] In step S341, the regular expression is applied to all slogans in the cluster and the coverage is calculated. If the coverage is less than the fifth preset threshold, the samples are reselected or the regular expression rules are adjusted until the coverage is greater than or equal to the fifth preset threshold.
[0116] It is understandable that the regular expression is applied to all slogans in the cluster, and the proportion of slogans that are successfully matched is counted to ensure that the coverage reaches a high level. If the coverage is insufficient, the samples are reselected or the regular expression rules are adjusted until the requirements are met.
[0117] In step S400, the regular expression device fingerprint is compared with the existing fingerprint database to update the fingerprint database.
[0118] It should be noted that the comparison between regular expression device fingerprints and existing fingerprint databases uses semantic equivalence for matching.
[0119] Furthermore, step S400 may include the following sub-steps:
[0120] In step S410, the newly generated regular expression fingerprint is matched one by one with the rules in the original fingerprint database for semantic equivalence.
[0121] It is understandable that for each new regular fingerprint, it is compared one by one with the rules in the library: it is converted into a finite automaton and checked whether the languages are equivalent.
[0122] In step S420, the fingerprint that has been successfully matched inherits the corresponding device information in the database.
[0123] It is understandable that equivalence inherits device information, such as manufacturer and model, without the need for new rules, avoiding duplicate storage and keeping the library concise.
[0124] In step S430, the fingerprint rules that have not been successfully matched are used as new rules to update the fingerprint rule library.
[0125] It is understandable that if the new regular expression is not semantically equivalent to all the rules in the library, such as matching the sensor slogan of VendorY and there is no corresponding rule in the library, it will be added to the library as a new rule and bound to its device information. The device information is automatically obtained through the cluster annotation in step 300.
[0126] Furthermore, in this exemplary embodiment, a system for generating fingerprints of IoT devices is provided, which is used to execute the above-mentioned method for generating fingerprints of IoT devices. Figure 6 As shown in , the system may include: a data collection module, an embedding generation module, a fingerprint generation module and a fingerprint library update module.
[0127] The data collection module is used to collect the protocol slogan data with labels of IoT devices through network scanning tools, and decompose the protocol slogan data into byte-level subword datasets through a word segmenter.
[0128] The embedding generation module is used to pre-train a masked language model using a byte-level subword dataset to learn the relationship between word semantics and subword context. It also uses supervised contrastive loss fine-tuning to optimize the embedding's stability to dynamic content and generate slogan-level embeddings suitable for downstream device fingerprint recognition tasks.
[0129] The fingerprint generation module is used to perform dimensionality reduction clustering on slogan-level embeddings to obtain device clusters, extract common substrings from each device cluster, and convert them into regular expression device fingerprints.
[0130] The fingerprint library update module is used to compare the regular expression device fingerprint with the existing fingerprint database and update the fingerprint library.
[0131] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
[0132] The above embodiments are intended only to illustrate the technical solutions of the present application and are not intended to limit them. The present application is not limited to the precise structures described above and illustrated in the accompanying drawings, and it cannot be assumed that the specific implementation of the present application is limited to these descriptions. For those skilled in the art of the present application, any changes and modifications made without departing from the concept of the present application should be deemed to fall within the scope of protection of the present application.< / ab> < / ab>
Claims
1. A method for generating fingerprints of IoT devices, characterized in that: include: Collect the labeled protocol slogan data of IoT devices through network scanning tools, and decompose the protocol slogan data into byte-level subword datasets through a word segmenter; By pre-training a masked language model on a byte-level subword dataset, we learn the relationship between word semantics and subword context. We also use a supervised contrastive loss fine-tuning method to optimize the embedding's stability to dynamic content, generating slogan-level embeddings suitable for downstream device fingerprinting tasks. This includes: The masked language model is pre-trained on a byte-level subword dataset to learn the semantics and contextual relationships of subwords. The same IP address and port are probed multiple times at different times, and protocol slogan data is extracted and recorded as slogan pairs. The dynamic content in the slogan pairs is filtered using regular expression rules to generate static slogan data. The static slogan pairs generated by the same slogan pair are used as positive sample pairs, and slogans are extracted from different slogan pairs to generate negative sample pairs. Cosine similarity is used for pre-screening to construct positive and negative sample pairs. The subword embedding of each slogan is generated using the pre-trained masked language model. The embeddings of all subwords in the sequence are averaged to generate slogan-level embeddings suitable for downstream device fingerprint recognition tasks. The first loss function is constructed to use supervised contrast loss to narrow the cosine similarity of the positive sample embeddings and push the similarity of the negative sample embeddings away. Perform dimensionality reduction clustering on slogan-level embeddings to obtain device clusters. Then extract common substrings from each device cluster and convert them into regular expression device fingerprints, including: Linear dimensionality reduction is performed on slogan-level embeddings through principal component analysis, and then compared and verified by combining unified manifold approximation with projected nonlinear dimensionality reduction. Hierarchical density clustering is used to cluster the reduced slogan-level embeddings to generate device clusters. The metadata label distribution of slogans within each cluster is counted, and the label consistency score is calculated. If the consistency score is higher than a fourth preset threshold, the most common label classification is assigned to the cluster as the cluster's device information label. Samples are selected from each cluster to extract common substrings, generating device fingerprints in the form of regular expressions. Combined with the cluster's device information labels, structured fingerprint rules are generated. Compare the regular expression device fingerprint with the existing fingerprint database and update the fingerprint database.
2. The method for generating fingerprints of IoT devices according to claim 1, characterized in that: The steps of collecting the protocol slogan data with labels of IoT devices through a network scanning tool and decomposing the protocol slogan data into byte-level subword data sets through a word segmenter include: Use a port scanning tool to scan the available ports of the IP address, combine it with a protocol scanning tool to capture the protocol banner data, and use a network mapping tool to obtain the labels corresponding to the protocol banner data to obtain the protocol banner data with labels; Preprocess the collected slogan data to remove special symbols and HTML structure symbols, and convert uppercase characters to lowercase; Train a byte-level tokenizer to decompose labeled protocol slogan data into byte-level subwords.
3. The method for generating fingerprints of IoT devices according to claim 2, characterized in that: The step of training a byte-level word segmenter to decompose the protocol slogan data with labels into byte-level subwords includes: Decompose each slogan into a sequence of single characters and map them into a sequence of ASCII bytes; Initialize the vocabulary containing all single characters and count the frequencies of characters and adjacent byte pairs; Merge the most frequent byte pairs in the current dataset into new subwords and update the slogan dataset; The merging process is repeated until the vocabulary size reaches a first preset threshold or the number of merging times reaches a second preset threshold.
4. The method for generating fingerprints of IoT devices according to claim 1, wherein: The first loss function is: ; in, is the first loss function, For slogans The embedding vector of For slogans The positive sample embedding of For slogans The embedding vector of the slogan Negative sample embedding, Calculate the cosine similarity, is the number of positive sample pairs, is the temperature parameter.
5. The method for generating fingerprints of IoT devices according to claim 1, wherein: The masked language model is pre-trained using a cross-entropy loss function; the cross-entropy loss function is: ; in, is the cross entropy loss function, M represents the set of masked subwords, is the true value of subword i, is the probability predicted by the model based on the context.
6. The method for generating fingerprints of IoT devices according to claim 1, wherein: After the step of selecting samples from each cluster to extract common substrings and generate a device fingerprint in the form of a regular expression, the method further includes: Apply the regular expression to all slogans in the cluster and calculate the coverage. If the coverage is less than the fifth preset threshold, reselect samples or adjust the regular expression rules until the coverage is greater than or equal to the fifth preset threshold.
7. The method for generating fingerprints of IoT devices according to claim 1, characterized in that: The step of comparing the regular expression device fingerprint with the existing fingerprint database and updating the fingerprint database includes: The newly generated regular expression fingerprint is matched one by one with the rules in the original fingerprint database for semantic equivalence; The fingerprint that matches successfully inherits the corresponding device information in the database; The fingerprint rules that are not successfully matched are used as new rules to update the fingerprint rule base.
8. A fingerprint generation system for an Internet of Things device, characterized in that: The system is used to execute the method for generating fingerprints of an IoT device according to any one of claims 1 to 7, and the system includes: The data collection module is used to collect the protocol slogan data with labels of IoT devices through network scanning tools, and decompose the protocol slogan data into byte-level subword datasets through a word segmenter; The embedding generation module is used to pre-train a masked language model using a byte-level subword dataset to learn the relationship between word semantics and subword context. It also uses a supervised contrastive loss fine-tuning method to optimize the embedding's stability to dynamic content, generating slogan-level embeddings suitable for downstream device fingerprinting tasks. This module includes: The masked language model is pre-trained on a byte-level subword dataset to learn the semantics and contextual relationships of subwords. The same IP address and port are probed multiple times at different times, and protocol slogan data is extracted and recorded as slogan pairs. The dynamic content in the slogan pairs is filtered using regular expression rules to generate static slogan data. The static slogan pairs generated by the same slogan pair are used as positive sample pairs, and slogans are extracted from different slogan pairs to generate negative sample pairs. Cosine similarity is used for pre-screening to construct positive and negative sample pairs. The subword embedding of each slogan is generated using the pre-trained masked language model. The embeddings of all subwords in the sequence are averaged to generate slogan-level embeddings suitable for downstream device fingerprint recognition tasks. The first loss function is constructed to use supervised contrast loss to narrow the cosine similarity of the positive sample embeddings and push the similarity of the negative sample embeddings away. The fingerprint generation module is used to perform dimensionality reduction clustering on slogan-level embeddings to obtain device clusters, extract common substrings from each device cluster, and convert them into regular expression device fingerprints. This module includes: Linear dimensionality reduction is performed on slogan-level embeddings through principal component analysis, and then compared and verified by combining unified manifold approximation with projected nonlinear dimensionality reduction. Hierarchical density clustering is used to cluster the reduced slogan-level embeddings to generate device clusters. The metadata label distribution of slogans within each cluster is counted, and the label consistency score is calculated. If the consistency score is higher than a fourth preset threshold, the most common label classification is assigned to the cluster as the cluster's device information label. Samples are selected from each cluster to extract common substrings, generating device fingerprints in the form of regular expressions. Combined with the cluster's device information labels, structured fingerprint rules are generated. The fingerprint library update module is used to compare the regular expression device fingerprint with the existing fingerprint database and update the fingerprint library.
Citation Information
Patent Citations
Fine-grained Internet of Things equipment automatic identification method based on firmware simulation
CN117171417A
Networking device identification method and system based on incremental learning and word embedding model
CN118467984A