Data security information automatic classification and danger level automatic determination method and system based on AI algorithm

Through the automatic classification of data security intelligence and automatic determination of risk levels based on AI algorithms, the problems of low efficiency and insufficient accuracy of manual analysis in existing technologies have been solved, the automatic classification of intelligence data and accurate determination of risk levels have been realized, and the efficiency of data security monitoring and early warning has been improved.

CN120744693APending Publication Date: 2025-10-03HENAN POLICE ACAD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510628048.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing data security intelligence classification and risk assessment technologies rely on manual analysis, which has the characteristics of low efficiency, insufficient assessment accuracy, long response time, and difficulty in coping with complex data security environments.

Method used

An AI-based method for automatic classification of data security intelligence and automated risk level determination is adopted. Intelligence data is collected through web crawlers, OCR image recognition, and natural language processing technologies. The BERT-BiLSTM-CRF model is used to extract key features. Combined with machine learning algorithms and dynamic risk determination models, the automatic classification of intelligence data and accurate determination of risk levels are achieved.

Benefits of technology

It has improved the automation and accuracy of data security intelligence processing, significantly improved the efficiency of data security monitoring and early warning, and reduced manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744693A_ABST
    Figure CN120744693A_ABST
Patent Text Reader

Abstract

The invention discloses a data security intelligence automatic classification and danger level automatic determination method and system based on an AI algorithm. The method comprises the following steps: automatically collecting intelligence data from various data leakage intelligence sources; key features are extracted by adopting an entity recognition model based on a BERT-BiLSTM-CRF fusion architecture, and joint characterization and feature extraction are performed on multi-modal data by adopting a multi-modal feature fusion technology; constructing a data security information automatic classification model by using a machine learning algorithm; on the basis of the classified intelligence data, features related to risks are extracted, a dynamic risk judgment model is utilized, the threat possibility and the influence degree are evaluated in combination with an OWASP risk matrix, and automatic judgment of the danger level is achieved; and carrying out continuous learning and model optimization. According to the method, the automation degree, accuracy and timeliness of data security information processing can be improved, manual intervention is reduced, and the efficiency of data security monitoring and early warning is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data security technology, and in particular to a method and system for automatic classification of data security intelligence and automatic determination of risk levels based on an AI algorithm. Background Art

[0002] With the rapid development of the digital economy, data security issues are becoming increasingly prominent, with threats such as data leaks, cyberattacks, and malware emerging one after another. Existing data security intelligence classification and risk assessment technologies typically rely on manual analysis, resulting in low efficiency, insufficient assessment accuracy, and long response times, making them inadequate for today's complex data security environment. Therefore, there is an urgent need for a more efficient and accurate automated method for processing data security intelligence. Summary of the Invention

[0003] In response to the problems that existing data security intelligence classification and risk assessment technologies usually rely on manual analysis, have low efficiency, insufficient assessment accuracy, long response time, and are difficult to cope with the current complex data security environment, this invention proposes a method and system for automatic classification of data security intelligence and automatic assessment of risk levels based on AI algorithms, so as to improve the automation, accuracy and timeliness of data security intelligence processing, reduce manual intervention, and significantly improve the efficiency of data security monitoring and early warning.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions:

[0005] On one hand, the present invention proposes a method for automatic classification of data security intelligence and automatic determination of risk levels based on an AI algorithm, comprising:

[0006] Step 1: Automatically collect intelligence data from various data leakage intelligence sources through web crawlers, interface calls, OCR image recognition technology, and natural language processing technology;

[0007] Step 2: The collected intelligence data is cleaned and standardized, and an entity recognition model based on the BERT-BiLSTM-CRF fusion architecture is used to extract key features, including event type, data type involved, data sensitivity, and leak scale. Multimodal feature fusion technology is also used to jointly represent and extract features from multimodal data including text and images.

[0008] Step 3: Use a machine learning algorithm to build a data security intelligence automatic classification model. The data security intelligence automatic classification model adopts a hierarchical classification architecture and classifies intelligence data through the trained data security intelligence automatic classification model.

[0009] Step 4: Based on the classified intelligence data, extract risk-related features, use a dynamic risk assessment model, and combine it with the OWASP risk matrix to assess the threat likelihood and impact, thereby achieving automated determination of the risk level. The dynamic risk assessment model uses a time decay factor to dynamically adjust the risk weight.

[0010] Furthermore, during the intelligence data collection process, data sources are prioritized according to their update frequency, sensitivity, and content quality, and data sources are collected based on the priority ranking.

[0011] Furthermore, during the intelligence data collection process, an anti-crawler mechanism of dynamic UA camouflage and request frequency control is adopted.

[0012] Furthermore, the machine learning algorithms include random forest, support vector machine and deep neural network.

[0013] Furthermore, adversarial training is added when training the data security intelligence automatic classification model.

[0014] Furthermore, the step 4 also includes: constructing a dynamic risk determination model through a fuzzy reasoning algorithm.

[0015] Furthermore, the assessment of threat likelihood and impact using the OWASP risk matrix includes:

[0016] The fuzzy output is converted into a risk score of 0-100, which is divided into extremely high, high, medium, and low risk levels according to the preset risk score threshold. 0-30 points represent low risk level, 31-60 points represent medium risk level, 61-80 points represent high risk level, and 81-100 points represent extremely high risk level.

[0017] Furthermore, after step 4, the method further comprises:

[0018] Step 5: Through the real-time feedback mechanism and historical data, the XGBoost+Transformer hybrid architecture is used to dynamically adjust the weights of the classification and risk assessment models.

[0019] Another aspect of the present invention provides a data security intelligence automatic classification and risk level automatic determination system based on AI algorithm, comprising:

[0020] The intelligence data automatic collection module is used to automatically collect intelligence data from various data leakage intelligence sources through web crawlers, interface calls, OCR image recognition technology and natural language processing technology;

[0021] The data preprocessing and feature extraction module is used to clean and standardize the collected intelligence data. It uses an entity recognition model based on the BERT-BiLSTM-CRF fusion architecture to extract key features, including event type, data type involved, data sensitivity, and leak scale. It also uses multimodal feature fusion technology to jointly represent and extract features from multimodal data including text and images.

[0022] An AI classification model building module is used to build a data security intelligence automatic classification model using a machine learning algorithm. The data security intelligence automatic classification model adopts a hierarchical classification architecture and classifies intelligence data through the trained data security intelligence automatic classification model;

[0023] The automated risk level determination module is used to extract risk-related features based on classified intelligence data, use a dynamic risk determination model, and combine it with the OWASP risk matrix to assess the threat likelihood and impact, thereby achieving automated risk level determination. The dynamic risk determination model uses a time decay factor to dynamically adjust the risk weight.

[0024] Furthermore, the system also includes:

[0025] The continuous learning and model optimization module is used to dynamically adjust the weights of the classification and risk assessment models through real-time feedback mechanisms and historical data, using the XGBoost+Transformer hybrid architecture.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] This invention uses AI algorithms to achieve automatic classification of data security intelligence and accurate determination of risk levels, effectively reducing the cost of manual operations, improving the efficiency and accuracy of intelligence analysis, and significantly enhancing data security risk identification and emergency response capabilities. It is suitable for complex and changeable data security monitoring and management scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a basic flow chart of a method for automatic classification of data security intelligence and automatic determination of risk levels based on an AI algorithm in an embodiment of the present invention;

[0029] Figure 2 This is a basic flow chart of another method for automatic classification of data security intelligence and automatic determination of risk levels based on an AI algorithm according to an embodiment of the present invention;

[0030] Figure 3 This is a schematic diagram of the architecture of a system for automatic classification of data security intelligence and automatic determination of risk levels based on an AI algorithm according to an embodiment of the present invention;

[0031] Figure 4 This is a schematic diagram of the architecture of another system for automatic classification of data security intelligence and automatic determination of risk levels based on AI algorithms in an embodiment of the present invention. DETAILED DESCRIPTION

[0032] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments:

[0033] like Figure 1 As shown, a method for automatic classification of data security intelligence and automatic determination of risk levels based on AI algorithms includes:

[0034] S101: Automatic intelligence data collection. Through web crawlers, API calls, OCR image recognition technology, and natural language processing (NLP) technology, intelligence data is automatically collected from various data leak intelligence sources. Anti-crawler mechanisms such as dynamic UA camouflage and request frequency control are used during the collection process to ensure the stability and effectiveness of the collection.

[0035] S102: Data Preprocessing and Feature Extraction. Collected intelligence data is cleaned and standardized. An entity recognition model based on the BERT-BiLSTM-CRF fusion architecture is used to extract key features, such as event type, data type involved, data sensitivity, and breach scale. Multimodal feature fusion technology is also used to jointly represent and extract features from multimodal data, including text and images.

[0036] S103: AI Classification Model Construction. Utilize machine learning algorithms (random forest, support vector machine, or deep neural network) to construct an automatic data security intelligence classification model. Specifically, a hierarchical classification architecture is proposed, with a cascade model that performs coarse-grained classification followed by fine-grained classification. Adversarial training is also added to enhance model robustness. Intelligence data is then classified using the trained automatic data security intelligence classification model.

[0037] S104: Automated Risk Level Determination. Based on the characteristics of classified intelligence data, a dynamic risk assessment model is used, combined with the OWASP risk matrix, to perform quantitative assessments using a fuzzy inference algorithm to achieve automated and accurate risk level determination. The dynamic risk assessment model uses a time-attenuation factor to dynamically adjust risk weights to accurately reflect the current risk level in real time.

[0038] Through the above technical solutions, the automation, accuracy and timeliness of data security intelligence processing can be improved, manual intervention can be reduced, and the efficiency of data security monitoring and early warning can be significantly improved.

[0039] Furthermore, S101 specifically includes:

[0040] S101.1: Data Source Identification and Classification

[0041] Identify target data sources, such as public web pages (e.g., news sites, forums, social media), API interfaces (e.g., Twitter API, GitHub API), image data (e.g., screenshots of leaked data, images in PDF documents), and unstructured text (e.g., forum posts, chat logs). Then, categorize these data sources and prioritize them based on update frequency and sensitivity, prioritizing the acquisition of important sources.

[0042] To ensure efficient collection, data sources can be prioritized based on their update frequency, sensitivity, and content quality. Specifically, a weighted scoring system is used to calculate the priority:

[0043] P=w1*F+w2*S+w3*Q

[0044] Where F is the update frequency, S is the sensitivity, Q is the content quality, and w1, w2, and w3 are weights.

[0045] As an implementable embodiment, P = 0.4*F + 0.3*S + 0.3*Q. Assuming F = 0.8 (high frequency update), S = 0.9 (high sensitivity), Q = 0.7 (good quality), then P = 0.4*0.8 + 0.3*0.9 + 0.3*0.7 = 0.77.

[0046] S101.2: Web crawler configuration and execution

[0047] Public web data sources identified by S101.1 are processed by a web crawler. The crawler is configured using the Scrapy framework, with dynamic User-Agent spoofing implemented using the fake-useragent library. The DOWNLOAD_DELAY parameter is set to control request frequency, and proxy pools and cookie management are used to mitigate anti-crawler mechanisms. For dynamically loaded web pages, Selenium is used in conjunction with Headless Chrome to retrieve the rendered HTML content.

[0048] Specifically, by disguising the User-Agent to simulate different devices and browsers, crawler detection can be circumvented. As an implementation method, this can be achieved in the following way.

[0049] from fake_useragent import UserAgent

[0050] ua=UserAgent()

[0051] headers={'User-Agent':ua.random}

[0052] Specifically, setting a download delay and controlling the request frequency can avoid triggering the anti-crawler mechanism. As an implementable method, this can be achieved in the following way.

[0053] #Scrapy settings.py

[0054] DOWNLOAD_DELAY=2#Basic delay 2 seconds

[0055] RANDOMIZE_DOWNLOAD_DELAY=True #Randomize the delay

[0056] Specifically, a proxy pool is used to switch IP addresses to cope with the anti-crawler mechanism. As an implementable method, it can be implemented in the following way.

[0057]

[0058] S101.3: API interface call

[0059] Connect the API data source identified in S101.1 to S101.3. Obtain the API key, configure the request parameters, retrieve data through HTTP requests (GET or POST), and parse the returned JSON or XML data to extract intelligence information.

[0060] S101.4: Image Data Processing and OCR Recognition

[0061] In parallel with S101.2 and S101.3, the image data in S101.1 is processed. First, image preprocessing operations (such as grayscale, denoising, and binarization) are performed, and then Tesseract-OCR is used to extract text information from the image.

[0062] Image preprocessing operations include grayscale conversion, denoising, and binarization to improve OCR accuracy. The code example is as follows.

[0063] import cv2

[0064] image=cv2.imread('image.png')

[0065] gray=cv2.cvtColor(image,cv2.COLOR_BGR2GRAY)

[0066] denoised=cv2.fastNlMeansDenoising(gray)

[0067] _,binary=cv2.threshold(denoised,127,255,cv2.THRESH_BINARY)

[0068] Use Tesseract-OCR to extract text information from images. The code example is as follows.

[0069] import pytesseract

[0070] text=pytesseract.image_to_string(binary,lang='eng')

[0071] S101.5: Preliminary Applications of Natural Language Processing (NLP)

[0072] The text data collected by S101.2, S101.3 and S101.4 are aggregated and preliminarily cleaned (HTML tags, URLs, and special characters are removed). Then, word segmentation and entity pre-identification are performed to prepare data for subsequent feature extraction.

[0073] As an implementation method, text cleaning is performed by removing irrelevant content such as HTML tags and URLs. The code example is as follows.

[0074]

[0075] As an implementation method, Jieba is used for Chinese and NLTK is used for English when performing word segmentation. The code example is as follows.

[0076] import jieba

[0077] text="Data Leakage Incident"

[0078] words = jieba.cut(text)

[0079] print(list(words))#['data','leak','event']

[0080] S101.6: Data Storage and Management

[0081] Store the data processed by S101.5 in a database (such as MongoDB or MySQL), record metadata such as collection time and source, and ensure data traceability and management.

[0082] As an implementation, MongoDB is used to store unstructured data, and MySQL is used to store structured metadata. The code example is as follows.

[0083]

[0084] Furthermore, S102 specifically includes:

[0085] S102.1: Data Cleansing and Standardization

[0086] The collected data was obtained from the S101 database. The text data was cleaned to remove HTML tags, URLs, special characters, and emoticons, and the encoding format and uppercase and lowercase characters were standardized. The image data was also grayscaled, denoised, and binarized. Finally, the time format and numerical data were standardized.

[0087] Specifically, data cleaning is performed in the following manner:

[0088] Deduplication: Use hashing or pandas’ drop_duplicates() to remove duplicate records;

[0089] Missing value processing: For data with missing key fields, fill in default values ​​(such as "unknown") or delete records;

[0090] Standardized formats: standardized date formats (e.g., YYYY-MM-DD) and units (e.g., GB is standardized to bytes).

[0091] Specifically, the standardization process is performed in the following manner:

[0092] Text normalization: convert the text to lowercase and encode it to UTF-8;

[0093] Numerical normalization: normalize the numerical features (x = (x-min) / (max-min)) or standardize (z = (x-mean) / std).

[0094] S102.2: Entity Recognition Model Construction and Training

[0095] The cleaned text data from S102.1 was used to build an entity recognition model. Using the BERT-BiLSTM-CRF architecture, we annotated the text with entities in BIO format using annotation tools (such as BRAT) to prepare training data. The model was then trained using the AdamW optimizer and the CRF loss function to extract key features (such as event type and data sensitivity).

[0096] Entity Recognition Based on BERT-BiLSTM-CRF:

[0097] Specifically, in the entity recognition model, context-related word vectors are generated through the pre-trained BERT model; bidirectional LSTM (BiLSTM) is used to capture the long-distance dependencies of the sequence. As an implementable method, the hidden layer size is set to 256; and CRF is used to optimize entity boundaries to ensure labeling consistency.

[0098] The entity recognition model training process is as follows:

[0099] 1. Prepare annotated dataset (such as BIO format);

[0100] 2. Fine-tune the model, using the Adam optimizer with a learning rate of 2e -5 ;

[0101] 3. Verify model performance and adjust hyperparameters (such as the number of LSTM layers).

[0102] The code example is as follows:

[0103]

[0104]

[0105] S102.3: Multimodal Feature Fusion

[0106] After S102.1 processing, the text and image data are separately extracted for features: BERT is used to generate feature vectors for text, and ResNet-50 is used to extract features for images. The text and image features are then fused through concatenation or an attention mechanism to construct a multimodal feature vector.

[0107] As an implementable method, ResNet-50 is used to extract image features. The code example is as follows.

[0108] import torchvision.models as models

[0109] resnet=models.resnet50(pretrained=True)

[0110] image_features=resnet(image_tensor)

[0111] As an implementable method, BERT is used to extract text features. The code example is as follows.

[0112] from transformers import BertTokenizer,BertModel

[0113] tokenizer=BertTokenizer.from_pretrained('bert-base-uncased')

[0114] model=BertModel.from_pretrained('bert-base-uncased')

[0115] inputs=tokenizer(text,return_tensors='pt')

[0116] text_features=model(**inputs).pooler_output

[0117] S102.4: Feature Storage and Management

[0118] The entity information extracted in S102.2 and the feature vectors generated in S102.3 are stored in the database to prepare for subsequent classification and analysis. Specifically, the text length and keyword frequency (e.g., "leak," "password") are extracted, and feature importance analysis is performed using random forests to retain the top-N features.

[0119] Furthermore, S103 specifically includes:

[0120] S103.1: Data preparation and feature selection

[0121] Obtain preprocessed and feature-extracted data from the S102 database and prepare a dataset for training the classification model. Split the dataset into training, validation, and test sets in a 7:2:1 ratio. Ensure that the data has classification labels (e.g., "cyberattack," "data breach") and select features relevant to the classification task, such as text features, image features, and entity recognition results. Use the chi-squared test to assess feature importance.

[0122] S103.2: Classification Model Architecture Design

[0123] The data from S103.1 is used to design a hierarchical classification architecture. First, perform coarse-grained classification (such as event type), and then perform fine-grained classification (such as event subtype) under each coarse-grained category. Select algorithms such as random forest, support vector machine, or deep neural network to build the model. Among them, random forest is suitable for high-dimensional features, and parameter settings such as n_estimators = 100 (indicating the use of 100 decision trees) and max_depth = 10 (limiting the maximum depth of each tree to 10). Support vector machine (SVM) is suitable for small samples, and the kernel function is radial basis function RBF. Deep neural network can specifically choose Transformer or CNN to process complex patterns.

[0124] S103.3: Model Training and Optimization

[0125] The model designed in S103.2 was applied to training, and parameters were optimized using cross-validation and grid search. Adversarial training was also implemented to improve model robustness by adding adversarial examples.

[0126] S103.4: Model Evaluation and Adjustment

[0127] Evaluate the model trained with S103.3 using the test set, analyzing metrics such as precision, recall, and F1-score. Adjust the model architecture or parameters based on the results, and retrain if necessary. Specifically, use grid search to adjust the learning rate, hidden layer size, and other parameters. Evaluate the model using 5-fold cross-validation.

[0128] Furthermore, S104 specifically includes:

[0129] S104.1: Risk Feature Extraction

[0130] Obtain classified intelligence data from S103's database and extract risk-related features, such as event type, data type, data sensitivity, and breach scale. Event type, such as "data breach" or "DDoS attack." Data type, such as PII (personally identifiable information) or financial data. Data sensitivity can be scored based on confidentiality (1-5). Breach scale, such as the number of affected users or the amount of data.

[0131] S104.2: Construction of dynamic risk assessment model

[0132] The risk characteristics of S104.1 are input into the dynamic risk assessment model, and the threat possibility and impact are assessed in combination with the OWASP risk matrix. Specifically, a dynamic risk assessment model is constructed through a fuzzy inference algorithm, and fuzzy rules are defined (such as "if the leak is large and the data sensitivity is high, the risk is high"). Fuzzy membership functions are used to quantify the characteristics and output risk scores, specifically converting the fuzzy output into a risk score of 0-100. A time decay factor is introduced to dynamically adjust the risk weight. Specifically, the risk decreases over time, and an exponential decay function w(t) = e is used. -λt , where the decay rate λ is adjusted according to business needs (such as 0.1), and t represents the number of days after the event occurs.

[0133] S104.3: Risk Level Determination

[0134] The S104.2 model is used to assign a risk score to each piece of intelligence data, and the risk levels are classified into extremely high, high, medium, and low according to the preset risk score thresholds. Specifically, the risk score is calculated as follows:

[0135] Score = w1*feature1+w2*feature2+...

[0136] As an implementable method, 0-30 points represent a low risk level, 31-60 points represent a medium risk level, 61-80 points represent a high risk level, and 81-100 points represent an extremely high risk level.

[0137] S104.4: Result storage and feedback

[0138] The judgment results of S104.3 are stored in the database and notified to the security team through a real-time feedback mechanism for timely response.

[0139] Furthermore, if Figure 2 As shown, the present invention also proposes another method for automatic classification of data security intelligence and automatic determination of risk levels based on AI algorithms, including:

[0140] S201: Automatic intelligence data collection. Through web crawlers, API calls, OCR image recognition technology, and natural language processing (NLP) technology, intelligence data is automatically collected from various data leak intelligence sources. Anti-crawler mechanisms such as dynamic UA camouflage and request frequency control are used during the collection process to ensure the stability and effectiveness of the collection.

[0141] S202: Data Preprocessing and Feature Extraction. Collected intelligence data is cleaned and standardized. An entity recognition model based on the BERT-BiLSTM-CRF fusion architecture is used to extract key features, such as event type, data type involved, data sensitivity, and breach scale. Multimodal feature fusion technology is also used to jointly represent and extract features from multimodal data, including text and images.

[0142] S203: AI Classification Model Construction. Utilize machine learning algorithms (random forest, support vector machine, or deep neural network) to construct a data security intelligence automatic classification model. Specifically, a hierarchical classification architecture is proposed, with a cascade model that performs coarse-grained classification followed by fine-grained classification. Adversarial training is also added to enhance model robustness. Intelligence data is then classified using the trained data security intelligence automatic classification model.

[0143] S204: Automated Risk Level Determination. Based on the classified intelligence data features, a dynamic risk assessment model is used, combined with the OWASP risk matrix, to perform a quantitative assessment using a fuzzy inference algorithm to achieve automated and accurate risk level determination. The dynamic risk assessment model uses a time-attenuation factor to dynamically adjust risk weights to accurately reflect the current risk level in real time.

[0144] S205: Continuous Learning and Model Optimization. Leveraging real-time feedback and historical data, we employ a dynamic weighted multi-model ensemble approach (XGBoost + Transformer hybrid architecture) to optimize classification models and risk assessment models to adapt to the ever-changing data security environment. We also dynamically adjust the weights in the risk assessment model.

[0145] This invention uses AI algorithms to achieve automatic classification of data security intelligence and accurate determination of risk levels, effectively reducing the cost of manual operations, improving the efficiency and accuracy of intelligence analysis, and significantly enhancing data security risk identification and emergency response capabilities.

[0146] Furthermore, S205 specifically includes:

[0147] S205.1: Establishment of real-time feedback mechanism

[0148] Collect feedback from the security team on S204 intelligence data and risk assessments, and establish a real-time feedback channel.

[0149] S205.2: Historical Data Analysis

[0150] Combine feedback from S205.1 with historical data to regularly analyze cases where model prediction errors or risk assessments are inaccurate.

[0151] S205.3: Model Optimization and Update

[0152] The analysis results of S205.2 are used to optimize the model, and the XGBoost+Transformer hybrid architecture is adopted to dynamically adjust the weights of the data security intelligence automatic classification model and the dynamic risk judgment model.

[0153] S205.4: Continuous Training and Deployment

[0154] Regularly train the S205.3 optimized model using the latest data and feedback information, update parameters, and deploy it to the production environment.

[0155] Based on the above embodiments, Figure 3 As shown, the present invention also proposes a data security intelligence automatic classification and risk level automatic determination system based on AI algorithm, including:

[0156] The intelligence data automatic collection module is used to automatically collect intelligence data from various data leakage intelligence sources through web crawlers, interface calls, OCR image recognition technology and natural language processing technology;

[0157] The data preprocessing and feature extraction module is used to clean and standardize the collected intelligence data. It uses an entity recognition model based on the BERT-BiLSTM-CRF fusion architecture to extract key features, including event type, data type involved, data sensitivity, and leak scale. It also uses multimodal feature fusion technology to jointly represent and extract features from multimodal data including text and images.

[0158] An AI classification model building module is used to build a data security intelligence automatic classification model using a machine learning algorithm. The data security intelligence automatic classification model adopts a hierarchical classification architecture and classifies intelligence data through the trained data security intelligence automatic classification model;

[0159] The automated risk level determination module is used to extract risk-related features based on classified intelligence data, use a dynamic risk determination model, and combine it with the OWASP risk matrix to assess the threat likelihood and impact, thereby achieving automated risk level determination. The dynamic risk determination model uses a time decay factor to dynamically adjust the risk weight.

[0160] Furthermore, if Figure 4 As shown, the present invention also proposes another data security intelligence automatic classification and risk level automatic determination system based on AI algorithm. Different from the above data security intelligence automatic classification and risk level automatic determination system, the system also includes:

[0161] The continuous learning and model optimization module is used to dynamically adjust the weights of the classification and risk assessment models through real-time feedback mechanisms and historical data, using the XGBoost+Transformer hybrid architecture.

[0162] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for automatic classification of data security intelligence and automatic determination of risk levels based on AI algorithms, characterized in that: include: Step 1: Automatically collect intelligence data from various data leakage intelligence sources through web crawlers, interface calls, OCR image recognition technology, and natural language processing technology; Step 2: The collected intelligence data is cleaned and standardized, and an entity recognition model based on the BERT-BiLSTM-CRF fusion architecture is used to extract key features, including event type, data type involved, data sensitivity, and leak scale. Multimodal feature fusion technology is also used to jointly represent and extract features from multimodal data including text and images. Step 3: Use a machine learning algorithm to build a data security intelligence automatic classification model. The data security intelligence automatic classification model adopts a hierarchical classification architecture and classifies intelligence data through the trained data security intelligence automatic classification model. Step 4: Based on the classified intelligence data, extract risk-related features, use a dynamic risk assessment model, and combine it with the OWASP risk matrix to assess the threat likelihood and impact, thereby achieving automated determination of the risk level. The dynamic risk assessment model uses a time decay factor to dynamically adjust the risk weight.

2. The method for automatic classification of data security intelligence and automatic determination of risk level based on AI algorithm according to claim 1 is characterized in that: During the intelligence data collection process, data sources are prioritized according to their update frequency, sensitivity, and content quality, and data sources are collected based on the priority ranking.

3. The method for automatic classification of data security intelligence and automatic determination of risk level based on AI algorithm according to claim 1 is characterized in that: During the intelligence data collection process, an anti-crawler mechanism of dynamic UA camouflage and request frequency control is adopted.

4. The method for automatic classification of data security intelligence and automatic determination of risk level based on AI algorithm according to claim 1 is characterized in that: The machine learning algorithms include random forest, support vector machine and deep neural network.

5. The method for automatic classification of data security intelligence and automatic determination of risk level based on AI algorithm according to claim 1 is characterized in that: When training the data security intelligence automatic classification model, adversarial training is added.

6. The method for automatic classification of data security intelligence and automatic determination of risk level based on AI algorithm according to claim 1 is characterized in that: The step 4 also includes: constructing a dynamic risk determination model through a fuzzy reasoning algorithm.

7. The method for automatic classification of data security intelligence and automatic determination of risk level based on AI algorithm according to claim 1 is characterized in that: The assessment of threat likelihood and impact using the OWASP risk matrix includes: The fuzzy output is converted into a risk score of 0-100, which is divided into extremely high, high, medium, and low risk levels according to the preset risk score threshold. 0-30 points represent low risk level, 31-60 points represent medium risk level, 61-80 points represent high risk level, and 81-100 points represent extremely high risk level.

8. The method for automatic classification of data security intelligence and automatic determination of risk level based on AI algorithm according to claim 1 is characterized in that: After step 4, the method further comprises: Step 5: Through the real-time feedback mechanism and historical data, the XGBoost+Transformer hybrid architecture is used to dynamically adjust the weights of the classification and risk assessment models.

9. A data security intelligence automatic classification and risk level automatic determination system based on AI algorithm, characterized by: include: The intelligence data automatic collection module is used to automatically collect intelligence data from various data leakage intelligence sources through web crawlers, interface calls, OCR image recognition technology and natural language processing technology; The data preprocessing and feature extraction module is used to clean and standardize the collected intelligence data. It uses an entity recognition model based on the BERT-BiLSTM-CRF fusion architecture to extract key features, including event type, data type involved, data sensitivity, and leak scale. It also uses multimodal feature fusion technology to jointly represent and extract features from multimodal data including text and images. An AI classification model building module is used to build a data security intelligence automatic classification model using a machine learning algorithm. The data security intelligence automatic classification model adopts a hierarchical classification architecture and classifies intelligence data through the trained data security intelligence automatic classification model; The automated risk level determination module is used to extract risk-related features based on classified intelligence data, use a dynamic risk determination model, and combine it with the OWASP risk matrix to assess the threat likelihood and impact, thereby achieving automated risk level determination. The dynamic risk determination model uses a time decay factor to dynamically adjust the risk weight.

10. The data security intelligence automatic classification and risk level automatic determination system based on AI algorithm according to claim 9 is characterized in that: The system also includes: The continuous learning and model optimization module is used to dynamically adjust the weights of the classification and risk assessment models through real-time feedback mechanisms and historical data, using the XGBoost+Transformer hybrid architecture.