Threat intelligence knowledge sharing method for malicious traffic logs
By adopting decentralized storage platform IPFS and smart contract technology on the cyber threat intelligence platform, the challenges of privacy protection and automated selection of model training are solved, and efficient and secure threat intelligence data sharing and model training are achieved.
Patent Information
- Application Number
- CN202510229285.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
AI Technical Summary
Existing cyberthreat intelligence platforms have challenges in privacy protection and automation selection of model training, resulting in data privacy breaches and inefficient model training.
The decentralized storage platform IPFS is used to obtain threat intelligence data, generate pending training data sets through data cleaning and standardization, and select appropriate machine learning models for adaptive iterative training through automatic selection mechanisms, and ensure the privacy and integrity of the data through smart contracts.
It improves the efficiency and effectiveness of model training, ensures the privacy and integrity of data, avoids the risk of leakage of sensitive data, and traces and verifies the ownership and usage records of the model through decentralized storage and smart contracts.
Smart Images

Figure CN120074928A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly to a method for sharing threat intelligence knowledge for malicious traffic logs. Background Art
[0002] A network threat intelligence platform is a system for collecting, analyzing, sharing, and utilizing network attack and defense data. Its main functions include collection and collation of threat data, identification and analysis of attack patterns, and sharing and application of intelligence. To improve the accuracy and timeliness of threat response, network threat intelligence platforms have begun to widely apply machine learning technology to discover potential attack behaviors and abnormal patterns by analyzing massive amounts of network traffic, logs, event data, etc. With the continuous upgrading of network attack means, traditional rule-based protection methods can no longer meet the needs of modern network security, and more and more enterprises and organizations are turning to data-driven threat detection and response systems. In these platforms, blockchain technology, as a decentralized distributed ledger technology, has gradually been applied to improve the security and credibility of data. Through technologies such as smart contracts, blockchain can automatically achieve authorized sharing of intelligence, security verification, and record updating. In this way, data providers can share threat intelligence in a decentralized and encrypted manner without worrying about data being tampered with or leaked.
[0003] However, the current network threat intelligence platforms still face significant challenges in terms of privacy protection. Especially in the field of network threat intelligence, data usually contains sensitive information such as network traffic, log records, IP addresses, attack behaviors, etc. These data often involve the core assets of enterprises and user privacy. If these data are directly stored in the blockchain in plain text form, sensitive security information may be exposed, resulting in serious privacy leakage and potential security risks. Therefore, how to ensure data availability while preventing public access to data content has become a major problem in the application of blockchain technology in threat intelligence platforms.
[0004] Secondly, there are still certain limitations in the model training and automated selection of data in existing threat intelligence platforms. Currently, many existing threat intelligence platforms rely on manual intervention to complete model training and selection. This approach will greatly increase the workload of users and is easily affected by subjective human factors, resulting in inefficiency and instability in the training process. In threat intelligence data, the types of data sources and data formats are diverse, including various types of data such as network traffic, log files, and sensor data. These data have great differences in quality, dimension, and features. How to select the most suitable machine learning model for training has become a complex and cumbersome task.
[0005] Therefore, there is an urgent need to provide a solution to solve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a threat intelligence knowledge sharing method for malicious traffic logs, which improves the limitation of relying on manual intervention in model selection.
[0007] A threat intelligence knowledge sharing method for malicious traffic logs provided by the present invention adopts the following technical solutions:
[0008] Obtain a threat intelligence data set from the decentralized storage platform IPFS, perform data cleaning and data type standardization on the threat intelligence data set to obtain a training data set to be processed, and perform preprocessing on the training data set to be processed to obtain a model training data set and a model test data set;
[0009] Select a training model based on an automatic selection mechanism, perform model adaptive iterative training to obtain the final output file of the model, obtain the specific type of the model based on the final output file of the model, and perform model evaluation based on the specific type of the model;
[0010] Upload the evaluated model to the IPFS network to obtain a unique model hash value, generate a pair of public and private keys, use the private key to digitally sign the model hash value based on the elliptic curve digital signature algorithm, verify the validity of the signature based on the public key, and store the generated model hash value, digital signature, and public key in the decentralized storage platform IPFS.
[0011] Optionally, during the process of obtaining the threat intelligence data set from the decentralized storage platform IPFS, it includes:
[0012] After confirming that the threat intelligence data is correct, download the threat intelligence data based on the provided IPFS hash value and save it as a specified file name to the user's local storage to obtain the threat intelligence data set. When the threat intelligence data cannot be downloaded normally or the hash value does not match, abort the download task and record the error information.
[0013] Optionally, during the process of performing data cleaning and data type standardization on the threat intelligence data set to obtain a training data set to be processed, it includes:
[0014] Select a reading method based on the extension name of the threat intelligence data file. For nested structure data, extract the internal nested information based on the multi-level expansion method. For list type fields, uniformly process them as strings;
[0015] Perform type detection and conversion on each column in the file data frame, identify IP address fields based on regular expressions and convert them to integer numerical representations; uniformly convert numerical data to integer processing based on missing value filling; standardize boolean values to 0 and 1 representations; determine the processing method for string columns based on whether they contain numerical values, convert numerical strings to integers, and encode non-numerical strings as numerical values based on a hash function;
[0016] Delete duplicate data files, and save the data files after data cleaning and data type standardization to the specified file path to obtain the training data set to be processed.
[0017] Optionally, in the process of preprocessing the training data set to be processed to obtain a model training data set, it includes:
[0018] Analyze the characteristics of the training data set to be processed to obtain the data set scale, the number of features, and the class distribution, perform standardization processing on numerical features, perform encoding conversion on categorical features to meet the model requirements, and divide the data set into a model training data set and a model test data set based on the task requirements.
[0019] Optionally, in the process of executing the selection of a training model based on an automatic selection mechanism, it includes:
[0020] Judge whether there are labels in the target column of the model training data set. If there are no labels, select a clustering model. If there are labels, judge the type and its distribution of the labels. When the labels in the target column are of numerical type and the number of unique values in the model training data set is greater than 10, select a regression model. Otherwise, select a classification model.
[0021] Optionally, in the process of selecting a clustering model when there are no labels, it includes: when the number of samples in the model training data set is less than 1000, select the K-Means model. Otherwise, select the DBSCAN model.
[0022] Optionally, in the process of selecting a regression model, it includes: when the number of samples in the model training data set is less than 1000, select the linear regression model. When the number of features is greater than 10, select the random forest regression model. Otherwise, select the SVR model.
[0023] Optionally, in the process of selecting a classification model, it includes: when the number of samples in the model training data set is less than 1000, select the logistic regression model. When the number of categorical features is more than the number of numerical features, select the random forest model. Otherwise, select the SVC model.
[0024] Optionally, in the process of evaluating the model based on the specific type of the model, it includes:
[0025] Obtain the feature data and target column data of the model test data set, divide the data proportionally, use the loaded model to predict the test set to obtain the prediction results, compare the prediction results with the true values of the test data set, and generate the performance evaluation indicators of the model.
[0026] Optionally, in the process of generating the performance evaluation indicators of the model, it includes:
[0027] For the regression model, the performance evaluation indicators include calculating the mean squared error, root mean squared error, mean absolute error, and R 2 score;
[0028] For the classification model, the performance evaluation indicators include calculating the accuracy rate, precision rate, recall rate, and F1 score;
[0029] For the clustering model, the performance evaluation indicator is to calculate the silhouette coefficient to evaluate the distribution of data points in different clusters and obtain the evaluation of the clustering quality of the model for the data.
[0030] A threat intelligence knowledge sharing method for malicious traffic logs provided by the present invention has the beneficial effects as follows:
[0031] 1. Aiming at the diversity and complexity of threat intelligence data sources, the present invention can flexibly select the optimal machine learning model according to the data characteristics, avoiding the limitations of relying on manual intervention in model selection in traditional methods, thereby improving the efficiency and effect of model training;
[0032] 2. The present invention uses the decentralized storage platform IPFS to upload the trained model and its evaluation results to the chain, ensures the integrity and privacy of the data when accessed by authorized parties through smart contracts, and innovatively proposes the concept of "model as intelligence", which does not directly expose the training data itself. Instead, it provides threat intelligence services for data users through model parameters and inference results, effectively avoiding the risk of leakage of sensitive data;
[0033] 3. The upload, sharing, and verification of each model and data are processed transparently and cannot be tampered with. Through decentralized storage and smart contracts, the ownership and usage records of the model can be traced and verified, ensuring the rights and interests of all participating parties and the security of the data. Description of the Drawings
[0034] Figure 1 It is a flowchart of the threat intelligence knowledge sharing method for malicious traffic logs provided by the embodiments of the present invention;
[0035] Figure 2 It is a flowchart of model selection provided by the embodiments of the present invention. Detailed Embodiments
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art in the field to which the present invention pertains. The words such as "including" used herein are intended to mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects.
[0037] An embodiment of the present invention provides a threat intelligence knowledge sharing method for malicious traffic logs. Refer to Figure 1 , including:
[0038] S1. Obtain a threat intelligence data set from the decentralized storage platform IPFS, perform data cleaning and data type standardization on the threat intelligence data set to obtain a training data set to be processed, and perform preprocessing on the training data set to be processed to obtain a model training data set and a model test data set;
[0039] S2. Select a training model based on an automatic selection mechanism, perform model adaptive iterative training to obtain the final output file of the model, obtain the specific type of the model based on the final output file of the model, and perform model evaluation based on the specific type of the model;
[0040] S3. Upload the evaluated model to the IPFS network to obtain a unique model hash value, generate a pair of public and private keys, use the private key to digitally sign the model hash value based on the elliptic curve digital signature algorithm, verify the validity of the signature based on the public key, and store the generated model hash value, digital signature, and public key in the decentralized storage platform IPFS.
[0041] In some embodiments, during the execution of step S1, it includes:
[0042] S1.1. Obtain a threat intelligence data set from the decentralized storage platform IPFS;
[0043] S1.2. Perform data cleaning and data type standardization on the threat intelligence data set to obtain a training data set to be processed;
[0044] S1.3. Perform preprocessing on the training data set to be processed to obtain a model training data set and a model test data set.
[0045] Specifically, in the process of executing step S1.1 to obtain the threat intelligence data set based on the decentralized storage platform IPFS, it includes: after confirming that the threat intelligence data is correct, downloading the threat intelligence data based on the provided IPFS hash value and saving it as a specified file name to the user's local storage to obtain the threat intelligence data set. When the threat intelligence data cannot be downloaded normally or the hash value does not match, the download task is aborted and the error information is recorded.
[0046] Specifically, in the process of executing step S1.2 to perform data cleaning and data type standardization on the threat intelligence data set to obtain the training data set to be processed, it includes: selecting the reading method based on the extension of the threat intelligence data file. For nested structure data, extracting the internal nested information based on the multi-layer expansion method. For list type fields, uniformly processing them as strings.
[0047] Furthermore, perform type detection and conversion on each column in the file data frame. Identify the IP address field based on regular expressions and convert it to an integer numerical representation; uniformly convert numerical data to integer processing based on missing value filling; standardize boolean values to 0 and 1 representations; determine the processing method of the string column based on whether it contains numerical values. Convert numerical strings to integers, and encode non-numerical strings as numerical values based on the hash function.
[0048] Furthermore, delete the duplicate data files and save the data files after data cleaning and data type standardization to the specified file path to obtain the training data set to be processed.
[0049] Specifically, in the process of executing step S1.3 to preprocess the training data set to be processed to obtain the model training data set, it includes: analyzing the characteristics of the training data set to be processed to obtain the data set scale, the number of features, and the category distribution, performing standardization processing on numerical features, performing encoding conversion on categorical features to meet the model requirements, and dividing the data set into a model training data set and a model test data set based on the task requirements.
[0050] In some embodiments, in the process of executing step S2, it includes:
[0051] S2.1. Select a training model based on the automatic selection mechanism;
[0052] S2.2. Perform model adaptive iterative training to obtain the final output file of the model;
[0053] S2.3. Perform model evaluation based on the specific type of the model.
[0054] Specifically, in the process of executing step S2.1 to select a training model based on the automatic selection mechanism, refer to Figure 2, including: determining whether there are labels in the target column of the model training dataset. If there are no labels, a clustering model is selected; if there are labels, the type and distribution of the labels are determined. When the labels in the target column are of numerical type and the number of unique values in the model training dataset is greater than 10, a regression model is selected; otherwise, a classification model is selected.
[0055] Furthermore, in the process of selecting a clustering model when there are no labels, it includes: when the number of samples in the model training dataset is less than 1000, the K-Means model is selected; otherwise, the DBSCAN model is selected.
[0056] Furthermore, in the process of selecting a regression model, it includes: when the number of samples in the model training dataset is less than 1000, the linear regression model is selected; when the number of features is greater than 10, the random forest regression model is selected; in other cases, the SVR model is selected.
[0057] Furthermore, in the process of selecting a classification model, it includes: when the number of samples in the model training dataset is less than 1000, the logistic regression model is selected; when the number of classification features is more than the number of numerical features, the random forest model is selected; in other cases, the SVC model is selected.
[0058] Specifically, when performing step S2.2 to obtain the final output file of the model through model adaptive iterative training, it includes: performing model adaptive iterative training, and after the training ends, generating the final output file of the model, including the structure, parameters of the model, and the intermediate result files generated during the training.
[0059] Furthermore, read the previously trained model file according to the provided model path, and obtain the specific type of the model through the records in the model information.
[0060] Specifically, when performing step S2.3 to evaluate the model based on the specific type of the model, it includes: obtaining the feature data and target column data of the model test dataset, dividing the data proportionally, using the loaded model to predict the test set to obtain the prediction results, comparing the prediction results with the true values of the test dataset, and generating the performance evaluation indicators of the model.
[0061] Actually, in order to ensure the adaptability of the evaluation data, the data will be preprocessed before evaluation to ensure that all features are numerical, and encoding conversion will be performed on the necessary non-numerical columns.
[0062] Furthermore, in the process of generating the performance evaluation indicators of the model, it includes: for the regression model, the performance evaluation indicators include calculating the mean squared error, root mean squared error, mean absolute error, and R 2 score; for the classification model, the performance evaluation indicators include calculating the accuracy, precision, recall, and F1 score.
[0063] Furthermore, the evaluation of the clustering model focuses on the quality of data grouping. Since there is no explicit target column in the clustering task, for the clustering model, the performance evaluation metric is to calculate the silhouette coefficient to evaluate the distribution of data points in different clusters, thereby obtaining an evaluation of the clustering quality of the model for the data. When the clustering model belongs to the K-Means type, the inertia metric will also be calculated to reflect the average distance from the data points to the cluster centers.
[0064] Actually, by calculating the mean squared error, root mean squared error, mean absolute error, and R 2 score, the error magnitude and goodness of fit of the regression model in continuous numerical prediction tasks can be quantified; by calculating accuracy, precision, recall, and F1 score, the prediction effect of the classification model in the classification task can be reflected, especially in multi-class classification tasks, it can provide detailed performance analysis.
[0065] In some embodiments, during the execution of step S3, it includes:
[0066] S3.1. Upload the evaluated model to the IPFS network to obtain a unique model hash value;
[0067] S3.2. Generate a pair of public and private keys based on the owner of the model;
[0068] S3.3. Store the generated model hash value, digital signature, and public key in the decentralized storage platform IPFS.
[0069] Specifically, during the execution of step S3.1, it includes: after the model training and evaluation are completed, obtain the locally stored model file and model-related parameter files and upload them to the IPFS network, and obtain a unique model hash value returned by the system. The hash value is used to locate and access the model file.
[0070] Furthermore, after the blockchain upload of the model data is completed, the system will update the status record of the model, which includes marking the model as having been uploaded to the chain, and at the same time recording the unique identifier (model ID) of the model in the blockchain. In addition, the system will regularly monitor the status of the data on the chain to ensure that the model data can be verified and accessed at any time.
[0071] Specifically, during the execution of step S3.2, it includes: generating a pair of public and private keys based on the owner of the model, using the elliptic curve digital signature algorithm to sign the calculated model hash value with the private key, and using the public key to verify the validity of the signature.
[0072] Actually, through this process, a digital signature is generated. The signature is the encrypted result of the model hash, proving that the model was created by the owner of the private key and that the model content has not been modified.
[0073] Further, perform step S3.3 to store the generated model hash value, digital signature, and public key in the decentralized storage platform IPFS.
[0074] In fact, when the model is uploaded and recorded on the blockchain, any third party can query the model's hash value, signature, and public key, and verify whether the signature of the model is valid through the public key to confirm the identity of the model's owner.
[0075] Although the embodiments of the present invention have been described in detail above, it is obvious to those skilled in the art that various modifications and changes can be made to these embodiments. However, it should be understood that such modifications and changes fall within the scope and spirit of the present invention described in the claims. Moreover, the present invention described herein can have other embodiments and can be implemented or realized in various ways.
Claims
1. A threat intelligence knowledge sharing method for malicious traffic logs, characterized in that: The following steps are involved: A threat intelligence data set is obtained from the decentralized storage platform IPFS, data cleaning and data type standardization are performed on the threat intelligence data set to obtain a training data set to be processed, and the training data set to be processed is preprocessed to obtain a model training data set and a model test data set; Select a training model based on the automatic selection mechanism, perform model adaptive iterative training to obtain the final output file of the model, obtain the specific type of the model based on the final output file of the model, and perform model evaluation based on the specific type of the model; Upload the evaluated model to the IPFS network to obtain a unique model hash value, generate a pair of public and private keys, use the private key to digitally sign the model hash value based on the elliptic curve digital signature algorithm, verify the validity of the signature based on the public key, and store the generated model hash value, digital signature and public key in the decentralized storage platform IPFS.
2. According to claim 1, a threat intelligence knowledge sharing method for malicious traffic logs is characterized in that: The process of obtaining threat intelligence datasets based on the decentralized storage platform IPFS includes: After confirming that the threat intelligence data is correct, the threat intelligence data is downloaded based on the provided IPFS hash value and saved as the specified file name to the user's local storage to obtain the threat intelligence data set. When the threat intelligence data cannot be downloaded normally or the hash value does not match, the download task is terminated and the error information is recorded.
3. A method for sharing threat intelligence knowledge for malicious traffic logs according to claim 1, characterized in that: The process of cleaning the threat intelligence data set and standardizing the data type to obtain the training data set to be processed includes: The reading method is selected based on the extension of the threat intelligence data file. For nested structured data, the internal nested information is extracted based on the multi-layer expansion method. For list type fields, they are uniformly processed as strings. Perform type detection and conversion on each column in the file data frame. Identify IP address fields based on regular expressions and convert them to integer values. Convert numeric data to integers based on missing value filling. Standardize Boolean values to 0 and 1. Determine the processing method of string columns based on whether they contain numeric values. Convert numeric strings to integers and encode non-numeric strings to numeric values based on hash functions. Delete duplicate data files, and save the data files after data cleaning and data type standardization to the specified file path to obtain the training data set to be processed.
4. A method for sharing threat intelligence knowledge for malicious traffic logs according to claim 1, characterized in that: The process of preprocessing the training data set to be processed to obtain the model training data set includes: Analyze the characteristics of the training data set to be processed to obtain the data set size, number of features and category distribution, standardize the numerical features, encode the category features to adapt to the model requirements, and divide the data set into model training data set and model test data set based on task requirements.
5. A method for sharing threat intelligence knowledge for malicious traffic logs according to claim 4, characterized in that: The process of selecting a training model based on the automatic selection mechanism includes: Determine whether the target column of the model training data set has a label. If there is no label, select the clustering model. If there is a label, determine the type of label and its distribution. When the label of the target column is a numeric type and there are more than 10 unique values in the model training data set, select the regression model, otherwise select the classification model.
6. A method for sharing threat intelligence knowledge for malicious traffic logs according to claim 5, characterized in that: The process of selecting a clustering model without labels includes: when the number of samples in the model training data set is less than 1000, the K-Means model is selected, otherwise the DBSCAN model is selected.
7. A method for sharing threat intelligence knowledge for malicious traffic logs according to claim 5, characterized in that: The process of selecting a regression model includes: When the number of samples in the model training data set is less than 1000, the linear regression model is selected. When the number of features is greater than 10, the random forest regression model is selected. In other cases, the SVR model is selected.
8. A method for sharing threat intelligence knowledge for malicious traffic logs according to claim 5, characterized in that: The process of selecting a classification model includes: When the number of samples in the model training data set is less than 1000, the logistic regression model is selected. When the number of categorical features is greater than the number of numerical features, the random forest model is selected. In other cases, the SVC model is selected.
9. A method for sharing threat intelligence knowledge for malicious traffic logs according to claim 1, characterized in that: The process of model evaluation based on the specific type of model includes: Obtain the feature data and target column data of the model test data set, divide the data in proportion, use the loaded model to predict the test set to obtain the prediction results, compare the prediction results with the true values of the test data set, and generate the performance evaluation indicators of the model.
10. A method for sharing threat intelligence knowledge for malicious traffic logs according to claim 1, characterized in that: The process of generating performance evaluation indicators for the model includes: For regression models, performance evaluation metrics include calculating mean square error, root mean square error, mean absolute error, and R 2 Fraction; For classification models, performance evaluation metrics include calculating accuracy, precision, recall, and F1 score; For clustering models, the performance evaluation indicator is to calculate the silhouette coefficient to evaluate the distribution of data points in different clusters and obtain the evaluation of the model's clustering quality for the data.
Citation Information
Cited By
High-quality data set quality evaluation method based on data and model collaboration
CN121808314A