Intelligent identification and automatic disposal method for intranet malicious traffic

Through the intelligent identification and automatic disposal of malicious traffic in the intranet, data cleaning, vectorization processing and large-scale model judgment are used to solve the problems of inefficiency and misjudgment in traditional methods, and the rapid identification and automatic disposal of malicious traffic in the intranet are realized, and security protection efficiency is improved.

CN120455105APending Publication Date: 2025-08-08ZUOYEBANG EDUCATION TECH (BEIJING) CO LTD

Patent Information

Application Number
CN202510654758.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Traditional intranet malicious traffic detection and processing methods rely on manual judgment, which is inefficient and can easily lead to misjudgment or misjudgment. The whitelist rules are costly to maintain, making it difficult to effectively identify and deal with malicious traffic.

Method used

Monitor intranet traffic through security sensors and file threat identifiers, perform data cleaning and formatting processing, extract key features and vectorized processing, calculate similarity using vector database, build prompt templates and input large models for judgment, and automatically execute disposal strategies.

Benefits of technology

It realizes rapid and intelligent identification and automatic disposal of malicious traffic in the intranet, reduces the cost of manual intervention and whitelist maintenance, improves the level of security protection, and avoids misjudgment and misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455105A_ABST
    Figure CN120455105A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent identification and automatic disposal method for intranet malicious traffic, which belongs to the technical field of network security and comprises the following steps: acquiring security alarm information by monitoring intranet traffic data; performing data cleaning on the security alarm information, and performing formatting processing according to a predefined format to obtain preprocessed alarm information; key features are extracted according to the preprocessed alarm information, vectorization processing is carried out, and vector alarm feature information is obtained; performing similarity calculation according to the vector alarm feature information and vector data information in a vector database to obtain similar vector data information; and constructing a prompt template according to the preprocessed alarm information, the similar vector data information and a constraint rule, inputting the prompt template into the large model for processing, obtaining a judgment result, executing a corresponding disposal strategy, and feeding back the judgment result and a disposal condition to safety management personnel. According to the method, intelligent identification of the safety alarm information is realized, and corresponding treatment is automatically executed according to the judgment result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to a method for intelligently identifying and automatically handling malicious traffic in an intranet. Background Art

[0002] As enterprises continue to increase their informatization, network security issues are becoming increasingly prominent. Malicious intranet traffic transmits malicious data within enterprise networks, stealing data, damaging systems, and spreading malware. Detecting and addressing malicious intranet traffic has become a critical component in ensuring enterprise network security. Traditional security systems rely primarily on manual judgment to process traffic logs and alarm logs submitted by sensors, as well as alerts from file threat identifiers. This approach presents numerous challenges.

[0003] Enterprises generate a massive number of alarms daily, making manual handling difficult and inefficient. Furthermore, due to factors like human subjectivity and fatigue, valid alarms can easily be overlooked or misjudged. Secondly, traditional methods use whitelisting rules to filter alarms, but this requires manual whitelisting. Due to the diverse nature of alarms, frequent whitelisting increases workload and can miss important security threats. Furthermore, whitelisting rules are expensive to update and maintain, requiring constant adjustments based on new threats. Furthermore, traditional methods are prone to misjudgments or omissions when processing large numbers of alarms, making it impossible to accurately identify valid malicious traffic.

[0004] Therefore, a method for intelligent identification and automatic disposal of malicious traffic in the intranet is proposed. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a method for intelligent identification and automatic disposal of malicious traffic in the intranet, which is used to solve the problems in traditional technologies of low efficiency in handling security alarm information and difficulty in effectively identifying malicious traffic.

[0006] An embodiment of the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, including:

[0007] Monitor intranet traffic data through security sensors, alarm logs, and file threat identifiers to obtain security alarm information;

[0008] Performing data cleaning on the security alarm information and formatting it according to a predefined format to obtain pre-processed alarm information;

[0009] Extract key features based on the pre-processed alarm information, and perform vector processing to obtain vector alarm feature information;

[0010] Calculate similarity between the vector alarm feature information and the vector data information in the vector database, and obtain similar vector data information in the vector database whose similarity to the vector alarm feature information is higher than a preset threshold;

[0011] Constructing a prompt template based on the pre-processed alarm information, the similar vector data information and the constraint rules, and inputting the prompt template into the large model for processing to obtain a determination result;

[0012] According to the judgment result output by the large model, the corresponding disposal strategy is executed, and the judgment result and disposal situation are fed back to the security management personnel.

[0013] Preferably, the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, the steps of: performing data cleaning on the security alarm information and formatting it according to a predefined format to obtain pre-processed alarm information; including:

[0014] Obtaining useless symbols in the security alarm information and filtering them out, replacing special symbols with a preset reference library, and obtaining alarm text information;

[0015] Detect duplicate values, missing values, and abnormal values in the alarm text information. When duplicate values are detected in the alarm text information, delete the values. When missing values are detected in the alarm text information, obtain relevant data of the missing values for prediction and recovery. When abnormal values are detected in the alarm text information, truncate and replace the abnormal values according to preset peak data to obtain pre-processing and cleaning alarm information.

[0016] The cleaning alarm information is standardized according to a predefined format, and data binning is performed based on time information to obtain pre-processed alarm information.

[0017] Preferably, the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, the steps of extracting key features based on the pre-processed alarm information and performing vector processing to obtain vector alarm feature information; including:

[0018] Extracting key feature information from the pre-processed alarm information according to preset feature fields;

[0019] Obtaining a keyword group in the key feature information, calculating the frequency of occurrence of the keyword group in the key feature information, and performing logarithmic processing to obtain logarithmic frequency data corresponding to the keyword group;

[0020] Calculating the inverse document frequency of the keyword phrase in the feature database, calculating the feature weight corresponding to the keyword phrase, and generating vector feature sub-information;

[0021]

[0022] Among them, w i is the feature weight corresponding to the i-th keyword group, f i is the frequency of occurrence of the i-th keyword group in the key feature information, N is the number of all data groups in the feature database, n i is the number of data groups containing the keyword group in the feature database, and ε is a very small constant;

[0023] The vector alarm feature information is obtained by calculating the vector feature sub-information corresponding to all the keyword groups in the key feature information and performing position coding.

[0024] Preferably, the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the steps of: performing similarity calculation based on the vector alarm feature information and the vector data information in the vector database, and obtaining similar vector data information in the vector database whose similarity to the vector alarm feature information is higher than a preset threshold; comprising:

[0025] According to the vector data information in the vector database, a number of vector centroid points are randomly selected; optionally, the method includes:

[0026] C i =C lb +(C ub -C lb )·R i

[0027] Among them, C i is the randomly selected i-th vector centroid point, C lb is the lower bound of the search space constructed based on the vector database, C ub is the upper bound of the search space constructed based on the vector database, R i is the i-th random number;

[0028]

[0029] Among them, mod is the modulo operation, R i-1 is the i-1th random number, a is the first control parameter, and b is the second control parameter;

[0030] By calculating the Euclidean distance between the vector centroid and the remaining vector data information in the vector database, when the Euclidean distance is lower than the Euclidean distance threshold requirement, constructing the vector centroid and the vector data information into a vector set, and calculating the mean vector of the vector set;

[0031] traversing and calculating the Euclidean distance between the vector set and the vector data information in the vector database that is not included in the vector set until all the vector data information in the vector database is divided into the corresponding vector set, and updating the mean vector of the vector set in real time;

[0032] Acquire the vector alarm feature information, calculate the vector set distance between the vector alarm feature information and the mean vector corresponding to each vector set, and select the vector set with the smallest vector set distance as the pre-search vector space;

[0033] Setting a number of search points to select vector data information in the pre-retrieval vector space, calculating the Euclidean distance with the vector alarm feature information, when the Euclidean distance between the vector data information selected by the search point and the vector alarm feature information is less than the Euclidean distance between the corresponding vector set in the pre-retrieval vector space and the vector alarm feature information, the search point randomly moves its position in the neighborhood until a maximum number of iterations is reached, obtaining the search point corresponding to the vector data information with the closest Euclidean distance to the vector alarm feature information as the optimal search point, and obtaining similar vector data information; when the Euclidean distance between the vector data information selected by the search point and the vector alarm feature information is greater than the Euclidean distance between the corresponding vector set in the pre-retrieval vector space and the vector alarm feature information, the search point quickly leaves the current neighborhood and updates the position of the search point;

[0034]

[0035] in, The position of the ith search point after t+1 iterations, is the position of the ith search point after t iterations, τ is the adjustment parameter, σ is a uniform random number, t max is the maximum number of iterations, p1, p2, p3 are random numbers in the interval [0,1], L t Search point location The Euclidean distance between the corresponding vector data information and the vector alarm feature information, L ave is the Euclidean distance between the corresponding vector set in the pre-retrieval vector space and the vector alarm feature information, The degrees of freedom are T distribution, m is the scale factor, is the iteration critical value, and k is a random integer.

[0036] Preferably, the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the steps of: constructing a prompt template based on the pre-processed alarm information, the similarity vector data information, and the constraint rules, and inputting the prompt template into a large model for processing to obtain a judgment result; comprising:

[0037] Obtaining a preset prompt word template can be implemented as "based on [] conditions, according to [] and [], give []";

[0038] The pre-processed alarm information, the similar vector data information and the constraint rules are input into the preset prompt word template to generate a prompt template; which can be implemented as "based on the [constraint rule] conditions, according to [pre-processed alarm information] and [similar vector data information], give [judgment result]".

[0039] Preferably, the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the steps of: constructing a prompt template based on the pre-processed alarm information, the similarity vector data information, and the constraint rules, and inputting the prompt template into a large model for processing to obtain a judgment result; comprising:

[0040] Obtain sample malicious traffic data and sample normal traffic data;

[0041] Based on the forward diffusion sub-model, noise data is added to the sample malicious traffic data and the sample normal traffic data according to a preset step size to obtain malicious traffic noise data and normal traffic noise data; based on the reverse diffusion sub-model, the malicious traffic noise data and the normal traffic noise data are restored according to the noise data to obtain restored sample malicious data and restored sample normal data;

[0042] Optimizing the reverse diffusion sub-model by analyzing the deviation between the recovered sample malicious data and the sample malicious traffic data, and the deviation between the recovered sample normal data and the sample normal traffic data;

[0043] The malicious traffic data of the samples are fused with the malicious traffic noise data to construct a malicious sample set; the normal traffic data of the samples are fused with the normal traffic noise data to construct a normal sample set;

[0044] Build a pre-trained large model based on the encoder, decoder and classifier;

[0045] The encoder extracts and analyzes sample features of the malicious sample set and the normal sample set respectively, and outputs sample classification information through the decoder. The classifier makes classification decisions based on the sample classification information to obtain predicted classification results. The pre-trained large model is trained based on the predicted classification results and the sample classification results to obtain a large model.

[0046] Preferably, the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the steps of: constructing a prompt template based on the pre-processed alarm information, the similarity vector data information, and the constraint rules, and inputting the prompt template into a large model for processing to obtain a judgment result; comprising:

[0047] The pre-processed alarm information is processed by the reverse diffusion sub-model to obtain restored alarm information, the alarm fusion information is subjected to random masking and position encoding by the encoder in the large model, key features are extracted based on the attention mechanism, and model feature information is obtained;

[0048] The encoder performs feature extraction on the similar vector data information to obtain similar feature information; the model feature information and the similar feature information are fused according to preset weights to obtain fused feature information; the decoder outputs feature classification information based on constraint rules, and the classifier makes classification decisions to obtain judgment results.

[0049] Preferably, the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the steps of: constructing a prompt template based on the pre-processed alarm information, the similarity vector data information, and the constraint rules, and inputting the prompt template into a large model for processing to obtain a judgment result; and further comprising:

[0050] Obtain several large models, process them respectively according to the prompt template, and obtain corresponding model judgment results;

[0051] According to the sample traffic data, the model judgment performance of the large model is evaluated based on the model evaluation index, and the corresponding model evaluation weight is obtained. According to the model judgment result and the model evaluation weight, the judgment result is obtained.

[0052] Preferably, the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, the steps of: executing a corresponding handling strategy based on the determination result output by the large model, and feeding back the determination result and handling status to security management personnel; including:

[0053] When the determination result is "ignore" or "false alarm", the security alarm information is marked as an invalid alarm;

[0054] When the judgment result is "antivirus recommended", the intranet traffic data is judged to be malicious traffic, and antivirus operations are performed on the intranet traffic data and related device terminals by calling an artificial intelligence program, and the operation log of the intranet traffic data and the called artificial intelligence program are recorded.

[0055] Compared with conventional technologies, the present invention has the following beneficial effects: a method for intelligently identifying and automatically handling malicious traffic in an intranet, which obtains security alarm information, performs data cleaning, formatting, and vectorization processing, obtains vector alarm feature information, and obtains similar vector data information from a vector database. Based on the pre-processed alarm information, similar vector data information, and constraint rules, a prompt template is constructed and input into a large model for processing to generate a more accurate and more reliable judgment result, thereby achieving effective identification of security alarm information and further achieving effective monitoring of malicious traffic data in the intranet, and automatically executing corresponding handling strategies based on the judgment results. The above method uses the large model to judge the processed security alarm information, achieving effective monitoring of security alarm information, and solving the problems of difficulty and low efficiency in manual alarm processing in conventional technologies. It can achieve rapid and intelligent identification of a large number of security alarm messages, while avoiding the problem of ignoring or misjudging valid alarms due to human factors. There is no need for manual whitelist addition, which reduces the cost required for maintaining and updating the whitelist, thereby achieving effective intelligent identification of malicious traffic in the intranet and automatically executing corresponding handling strategies based on the judgment results, achieving automatic handling of malicious traffic in the intranet without human intervention, and effectively improving the security protection level of intranet operation.

[0056] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.

[0057] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 The present invention provides a flow chart of a method for intelligently identifying and automatically handling malicious traffic in an intranet. DETAILED DESCRIPTION

[0059] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0060] Example 1:

[0061] The embodiment of the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet. Figure 1 ,include:

[0062] Monitor intranet traffic data through security sensors, alarm logs, and file threat identifiers to obtain security alarm information;

[0063] Clean the security alarm information and format it according to the predefined format to obtain pre-processed alarm information;

[0064] Extract key features based on pre-processed alarm information and perform vector processing to obtain vector alarm feature information;

[0065] Calculate similarity between the vector alarm feature information and the vector data information in the vector database, and obtain similar vector data information in the vector database whose similarity to the vector alarm feature information is higher than a preset threshold;

[0066] Based on the pre-processed alarm information, similar vector data information and constraint rules, a prompt template is constructed, and the prompt template is input into the large model for processing to obtain the judgment result;

[0067] Based on the judgment results output by the large model, the corresponding disposal strategy is executed, and the judgment results and disposal status are fed back to the security management personnel.

[0068] In the above embodiments, intranet traffic data is monitored through security sensors, alarm logs, and file threat identifiers to obtain security alarm information; the security alarm information is cleaned and formatted according to a predefined format to obtain preprocessed alarm information; key features are extracted from the preprocessed alarm information and vectorized to obtain vector alarm feature information; similarity is calculated between the vector alarm feature information and vector data information in a vector database to obtain similar vector data information in the vector database whose similarity to the vector alarm feature information is higher than a preset threshold;

[0069] Based on the pre-processed alarm information, similar vector data information and constraint rules, a prompt template is constructed, and the prompt template is input into the large model for processing to obtain the judgment result. Based on the judgment result, the corresponding disposal strategy is executed, and the judgment result and disposal situation are fed back to the safety management personnel.

[0070] In one embodiment, the security alarm information includes victim IP address information, attacking IP address information, alarm type, threat level, threat intelligence or rules, threat level, number of repeated attacks, victim IP asset ownership, host domain name, source address of the attacking IP, source IP address information of the request, destination IP address information, source port, destination port, protocol, API, packet content (after base64), request body (after base64), request header (after base64), response content (after base64), and response header (after base64) details.

[0071] The beneficial effects of the above technology are: by obtaining security alarm information, performing data cleaning, formatting and vectorization processing, obtaining vector alarm feature information, and obtaining similar vector data information in the vector database, constructing a prompt template based on the pre-processed alarm information, similar vector data information and constraint rules, and inputting it into the large model for processing to generate a more accurate and more reliable judgment result, thereby achieving effective identification of security alarm information, and thus achieving effective monitoring of malicious traffic data in the intranet, and automatically executing corresponding disposal strategies based on the judgment results; compared with traditional technologies, a method for intelligent identification and automatic disposal of malicious traffic in the intranet uses a large model to judge the processed security alarm information, achieves effective monitoring of security alarm information, solves the problem of difficulty and low efficiency in manual processing of alarm information in traditional technologies, can achieve rapid and intelligent identification of a large number of security alarm information, and avoids the problem of ignoring or misjudging valid alarms due to human factors, does not require manual addition of whitelists, and reduces the cost required for maintaining and updating whitelists, thereby achieving effective intelligent identification of malicious traffic in the intranet, and automatically executing corresponding disposal strategies based on the judgment results, achieving automatic disposal of malicious traffic in the intranet without human intervention, and effectively improving the security protection level of intranet operation.

[0072] Example 2:

[0073] The embodiment of the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the steps of: performing data cleaning on security alarm information and formatting it according to a predefined format to obtain pre-processed alarm information; including:

[0074] Obtain useless symbols from security alarm information and filter them out, replace special symbols with a preset reference library, and obtain alarm text information;

[0075] Detect duplicate values, missing values, and abnormal values in alarm text information. When duplicate values are detected in the alarm text information, delete them. When missing values are detected in the alarm text information, obtain relevant data of the missing values for prediction and recovery. When abnormal values are detected in the alarm text information, truncate and replace the abnormal values according to the preset peak data to obtain pre-processed cleaning alarm information.

[0076] The cleaning alarm information is standardized according to a predefined format, and data is binned based on time information to obtain pre-processed alarm information.

[0077] In the above embodiments, useless symbols in the security alarm information are screened out, special symbols are replaced through a preset reference library, the alarm text information is obtained, the repeated values in the alarm text information are deleted, the missing values are predicted and restored, the abnormal values are truncated and replaced, the pre-processed and cleaned alarm information is obtained, and standardized processing is performed to obtain the pre-processed alarm information.

[0078] The beneficial effects of the above technology are: by cleaning the security alarm information, removing invalid or redundant data in the security alarm information, ensuring the accuracy and completeness of the data, and performing standardized processing to obtain pre-processed alarm information, it is convenient for subsequent steps to process and analyze the pre-processed alarm information.

[0079] Example 3:

[0080] The embodiment of the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the steps of extracting key features based on pre-processed alarm information and performing vector processing to obtain vector alarm feature information; including:

[0081] Extract key feature information from pre-processed alarm information based on preset feature fields;

[0082] Obtaining a keyword group in the key feature information, calculating the frequency of occurrence of the keyword group in the key feature information, and performing logarithmic processing to obtain logarithmic frequency data corresponding to the keyword group;

[0083] Calculate the inverse document frequency of the keyword phrase in the feature database, calculate the feature weight corresponding to the keyword phrase, and generate vector feature sub-information;

[0084]

[0085] Among them, w i is the feature weight corresponding to the i-th keyword group, f i is the frequency of occurrence of the i-th keyword group in the key feature information, N is the number of all data groups in the feature database, n i is the number of data sets containing keyword phrases in the feature database, and ε is a minimum constant;

[0086] The vector alarm feature information is obtained by calculating the vector feature sub-information corresponding to all the key word groups in the key feature information and performing position encoding.

[0087] In the above embodiments, according to the preset feature fields, the key feature information in the pre-processed alarm information is extracted, the logarithmic frequency and inverse document frequency of the keyword groups in the key feature information are calculated, the corresponding feature weights are obtained, the vector feature sub-information is generated, the vector feature sub-information corresponding to all keyword groups in the key feature information is calculated, and position encoding is performed to obtain the vector alarm feature information.

[0088] The beneficial effects of the above technology are: feature encoding is performed based on the key features in the pre-processed alarm information, thereby realizing the acquisition of vector feature sub-information, and according to the position of the keyword group in the key feature information, realizing the acquisition of vector alarm feature information, thereby realizing vectorized processing of the pre-processed alarm information.

[0089] Example 4:

[0090] An embodiment of the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the steps of: calculating similarity between vector alarm feature information and vector data information in a vector database, and obtaining similar vector data information in the vector database whose similarity to the vector alarm feature information is higher than a preset threshold; and comprising:

[0091] According to the vector data information in the vector database, several vector centroid points are randomly selected; optional ones include:

[0092] C i =C lb +(C ub -C lb )·R i

[0093] Among them, C i is the randomly selected i-th vector centroid point, C lb is the lower bound of the search space constructed based on the vector database, C ub is the upper bound of the search space constructed based on the vector database, R i is the i-th random number;

[0094]

[0095] Among them, mod is the modulo operation, R i-1 is the i-1th random number, a is the first control parameter, and b is the second control parameter;

[0096] By calculating the Euclidean distance between the vector centroid and the rest of the vector data information in the vector database, when the Euclidean distance is lower than the Euclidean distance threshold requirement, the vector centroid and the vector data information are constructed into a vector set, and the mean vector of the vector set is calculated;

[0097] Traversing and calculating the Euclidean distance between the vector set and the vector data information in the vector database that is not included in the vector set, until all the vector data information in the vector database is divided into the corresponding vector set, and updating the mean vector of the vector set in real time;

[0098] Obtaining vector alarm feature information, calculating the vector set distance between the vector alarm feature information and the mean vector corresponding to each vector set, and selecting the vector set with the smallest vector set distance as the pre-search vector space;

[0099] Set several search points to select vector data information in the pre-retrieval vector space, calculate the Euclidean distance with the vector alarm feature information, when the Euclidean distance between the vector data information selected at the search point and the vector alarm feature information is less than the Euclidean distance between the corresponding vector set in the pre-retrieval vector space and the vector alarm feature information, the search point randomly moves its position in the neighborhood until the maximum number of iterations is reached, obtain the search point corresponding to the vector data information with the closest Euclidean distance to the vector alarm feature information as the optimal search point, and obtain similar vector data information; when the Euclidean distance between the vector data information selected at the search point and the vector alarm feature information is greater than the Euclidean distance between the corresponding vector set in the pre-retrieval vector space and the vector alarm feature information, the search point quickly leaves the current neighborhood and updates the position of the search point;

[0100]

[0101] in, The position of the ith search point after t+1 iterations, is the position of the ith search point after t iterations, τ is the adjustment parameter, σ is a uniform random number, t max is the maximum number of iterations, p1, p2, p3 are random numbers in the interval [0,1], L t Search point location The Euclidean distance between the corresponding vector data information and the vector alarm feature information, L ave is the Euclidean distance between the corresponding vector set in the pre-retrieval vector space and the vector alarm feature information, The degrees of freedom are T distribution, m is the scale factor, is the iteration critical value, and k is a random integer.

[0102] In the above embodiment, several vector sets are divided according to the vector data information in the vector database; specifically, a vector centroid point is randomly selected in the vector database based on a chaotic map, and the Euclidean distance between the vector centroid point and the remaining vector data information in the vector database is calculated. When the Euclidean distance is lower than the Euclidean distance threshold requirement, the vector centroid point and the vector data information are constructed into a vector set, and the Euclidean distance between the vector set and the vector data information in the vector database that has not entered the vector set is traversed and calculated until all vector data information in the vector database is divided into corresponding vector sets, and the mean vector of the vector set is updated.

[0103] In the above embodiment, based on the vector set distance between the vector alarm feature information and the mean vector corresponding to each vector set, the vector set with the smallest vector set distance is selected as the pre-retrieval vector space, and similar vector data information is searched in the pre-retrieval vector space.

[0104] In the above embodiment, several search points are set in the pre-retrieval vector space to select vector data information, and the Euclidean distance with the vector alarm feature information is calculated. When the Euclidean distance between the vector data information selected at the search point and the vector alarm feature information is less than the Euclidean distance between the corresponding vector set in the pre-retrieval vector space and the vector alarm feature information, the search point randomly moves its position in the neighborhood to search for similar vector data information; when the Euclidean distance between the vector data information selected at the search point and the vector alarm feature information is greater than the Euclidean distance between the corresponding vector set in the pre-retrieval vector space and the vector alarm feature information, the search point quickly leaves the current neighborhood based on the normal distribution to search for similar vector data information.

[0105] In one embodiment, similar vector data information is acquired using a vector retrieval-augmented generation (RAG) technique.

[0106] The beneficial effects of the above technology are: based on the similarity of vector data information in the vector database, the construction of the vector set is realized, and when searching for similar vector data information based on the vector alarm feature information, the vector set with the minimum Euclidean distance is selected as the pre-retrieval vector space, and the search point is set based on the algorithm to quickly search for the optimal position, which greatly improves the efficiency of searching for similar vector data information in the vector database based on the vector alarm feature information.

[0107] Example 5:

[0108] An embodiment of the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the following steps: constructing a prompt template based on pre-processed alarm information, similar vector data information, and constraint rules, and inputting the prompt template into a large model for processing to obtain a judgment result; including:

[0109] Obtaining a preset prompt word template can be implemented as "based on [] conditions, according to [] and [], give []";

[0110] The pre-processed alarm information, similar vector data information and constraint rules are input into the preset prompt word template to generate a prompt template; it can be implemented as "based on the [constraint rule] conditions, according to [pre-processed alarm information] and [similar vector data information], give [judgment result]".

[0111] In the above embodiments, a prompt word template is preset, and the pre-processed alarm information, similar vector data information and constraint rules are filled in to construct a prompt template.

[0112] In one embodiment, the prompt template may be implemented as "based on [constraint rule] conditions, according to [pre-processed alarm information] and [similar vector data information], giving [determination result]".

[0113] In one embodiment, the prompt template may be implemented as a prompt template.

[0114] In one embodiment, the constraint rules include enterprise security policies and industry standards.

[0115] The beneficial effects of the above technology are: based on the pre-processed alarm information, similar vector data information and constraint rules, a prompt template is constructed, and the large model is processed according to the prompt template, which effectively improves the efficiency and accuracy of the large model processing and reduces the uncertainty of the output judgment results.

[0116] Example 6:

[0117] An embodiment of the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the following steps: constructing a prompt template based on pre-processed alarm information, similar vector data information, and constraint rules, and inputting the prompt template into a large model for processing to obtain a judgment result; including:

[0118] Obtain sample malicious traffic data and sample normal traffic data;

[0119] Based on the forward diffusion sub-model, noise data is added to the sample malicious traffic data and the sample normal traffic data according to a preset step size to obtain malicious traffic noise data and normal traffic noise data; based on the reverse diffusion sub-model, the malicious traffic noise data and the normal traffic noise data are restored according to the noise data to obtain restored sample malicious data and restored sample normal data;

[0120] The reverse diffusion sub-model is optimized by analyzing the deviation between the recovered sample malicious data and the sample malicious traffic data, and the deviation between the recovered sample normal data and the sample normal traffic data;

[0121] The malicious traffic data of the samples are fused with the malicious traffic noise data to construct a malicious sample set; the normal traffic data of the samples are fused with the normal traffic noise data to construct a normal sample set;

[0122] Build a pre-trained large model based on the encoder, decoder and classifier;

[0123] The encoder extracts and analyzes sample features of the malicious sample set and the normal sample set respectively, and outputs sample classification information through the decoder. The classifier makes classification decisions based on the sample classification information to obtain predicted classification results; the pre-trained large model is trained based on the predicted classification results and the sample classification results to obtain a large model.

[0124] In the above embodiments, a forward diffusion sub-model and a reverse diffusion sub-model are constructed. The forward diffusion sub-model adds noise data to the sample malicious traffic data and the sample normal traffic data according to a preset step size to obtain malicious traffic noise data and normal traffic noise data. The reverse diffusion sub-model restores the malicious traffic noise data and the normal traffic noise data according to the noise data to obtain restored sample malicious data and restored sample normal data. The reverse diffusion sub-model is optimized by analyzing the deviation between the restored sample malicious data and the sample malicious traffic data, and the deviation between the restored sample normal data and the sample normal traffic data.

[0125] In the above embodiments, the malicious sample data and the malicious traffic noise data are fused to construct a malicious sample set; and the normal sample data and the normal traffic noise data are fused to construct a normal sample set.

[0126] In the above embodiments, the pre-trained large model is trained using a malicious sample set and a normal sample set. The encoder extracts and analyzes sample features of the malicious sample set and the normal sample set respectively, and outputs sample classification information through the decoder. The classifier makes classification decisions based on the sample classification information to obtain predicted classification results. The pre-trained large model is trained based on the predicted classification results and the sample classification results to obtain a large model.

[0127] The beneficial effects of the above technology are: through the forward diffusion sub-model, the sample data of sample malicious traffic data and sample normal traffic data are expanded, and malicious sample sets and normal sample sets are obtained for training the pre-trained large model, which solves the problem that the pre-trained model cannot be fully trained when the amount of sample malicious traffic data and sample normal traffic data is small, and the feature enhancement of sample malicious traffic data and sample normal traffic data based on the forward diffusion sub-model is achieved to improve the training effect of the pre-trained large model; and by constructing a reverse diffusion sub-model, training is performed based on analyzing the deviation between the recovered sample malicious data and the sample malicious traffic data, and the deviation between the recovered sample normal data and the sample normal traffic data to optimize the denoising performance of the reverse diffusion sub-model, which is used for noise elimination of pre-processed alarm information in subsequent steps.

[0128] Example 7:

[0129] An embodiment of the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the following steps: constructing a prompt template based on pre-processed alarm information, similar vector data information, and constraint rules, and inputting the prompt template into a large model for processing to obtain a judgment result; including:

[0130] The pre-processed alarm information is processed through the reverse diffusion sub-model to obtain the restored alarm information. The alarm fusion information is randomly masked and position-encoded through the encoder in the large model. Key features are extracted based on the attention mechanism to obtain model feature information.

[0131] The encoder extracts features from similar vector data information to obtain similar feature information; the model feature information and similar feature information are fused according to preset weights to obtain fused feature information; the decoder outputs feature classification information based on constraint rules, and the classifier makes classification decisions to obtain judgment results.

[0132] In the above embodiments, the pre-processed alarm information is processed by the reverse diffusion sub-model to obtain the restored alarm information, the alarm fusion information is randomly masked and position-encoded by the encoder in the large model, key features are extracted based on the attention mechanism, and model feature information is obtained; the encoder performs feature extraction on the similar vector data information to obtain similar feature information; the model feature information and the similar feature information are fused according to the preset weights to obtain fused feature information, the decoder outputs feature classification information based on the constraint rules, and the judgment result is obtained through the classifier classification decision.

[0133] The beneficial effects of the above technology are: denoising of the pre-processed alarm is achieved through the inverse diffusion sub-model, and model feature information is obtained through processing by the large model; similar feature information corresponding to the similar vector data information is fused with the model feature information to obtain fused feature information, which is processed by the decoder and then the classifier outputs the corresponding judgment result. The above technical solution enables the large model to perform feature extraction and analysis on the pre-processed alarm information and similar vector data information based on constraint rules, and obtain the corresponding judgment results, which facilitates the execution of corresponding disposal strategies in subsequent steps.

[0134] Example 8:

[0135] An embodiment of the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the steps of: constructing a prompt template based on pre-processed alarm information, similar vector data information, and constraint rules, and inputting the prompt template into a large model for processing to obtain a judgment result; and further comprising:

[0136] Obtain several large models, process them according to the prompt templates, and obtain the corresponding model judgment results;

[0137] According to the sample traffic data, the model judgment performance of the large model is evaluated based on the model evaluation index, the corresponding model evaluation weight is obtained, and the judgment result is obtained according to the model judgment result and the model evaluation weight.

[0138] In the above embodiment, several large models are called through a network interface to perform processing according to the prompt template to obtain the corresponding model judgment result, and the judgment result is obtained according to the model evaluation weight and model judgment result corresponding to the large model.

[0139] In the above embodiment, the large model processes the sample flow data to obtain the sample model determination result, and evaluates the model determination performance of the large model based on the model evaluation index according to the sample standard determination result corresponding to the sample flow data;

[0140]

[0141] Among them, F is the model evaluation index parameter, θ is the balance parameter, Pre is the parameter for calculating the model processing accuracy based on the sample model judgment results and the sample standard judgment results, and Re is the ratio parameter of the correct judgment of the large model on the sample malicious traffic data in the sample traffic data to all the sample malicious traffic data.

[0142] In one embodiment, the macro model is constructed based on a thought chain model.

[0143] The beneficial effect of the above technology is that by obtaining multiple large models, the analysis and processing of pre-processed alarm information is realized, and the corresponding model judgment results are obtained. Based on the model judgment performance corresponding to each major model, the judgment results are obtained. In the above technical solution, multiple large models are used for analysis and processing to obtain judgment results, which effectively improves the accuracy of the judgment of pre-processed alarm information.

[0144] Example 9:

[0145] The embodiment of the present invention provides a method for intelligently identifying and automatically handling malicious traffic in an intranet, comprising the following steps: executing a corresponding handling strategy based on the determination results output by a large model, and feeding back the determination results and handling status to security management personnel; including:

[0146] When the judgment result is "ignore" or "false alarm", the security alarm information is marked as an invalid alarm;

[0147] When the judgment result is "antivirus recommended", the intranet traffic data is judged to be malicious traffic, and antivirus operations are performed on the intranet traffic data and related device terminals by calling the artificial intelligence program, and the operation log of the intranet traffic data and the called artificial intelligence program are recorded.

[0148] In the above embodiment, when the determination result is "ignore" or "false alarm", the security alarm information is marked as an invalid alarm, and the intranet traffic data is not processed.

[0149] In the above embodiment, when the judgment result is "antivirus recommended", the intranet traffic data is judged to be malicious traffic, and the intranet traffic data and related device terminals are disinfected by calling the artificial intelligence program, and the operation log of the intranet traffic data and the called artificial intelligence program are recorded.

[0150] In the above embodiments, by recording the operation logs of the intranet traffic data and the called artificial intelligence program for subsequent query and analysis, it serves as the basis for learning and optimization of intranet malicious traffic monitoring, and continuously improves the monitoring performance and accuracy.

[0151] The beneficial effects of the above technology are: by obtaining the judgment results and executing the corresponding disposal strategy, the automatic disposal of intranet traffic data is realized. When the judgment result is "antivirus recommended", the artificial intelligence program is automatically called to issue antivirus instructions, effectively improving the timeliness and effectiveness of security protection.

[0152] The present invention proposes a method for intelligently identifying and automatically handling malicious traffic in an intranet. The method adopts RAG and large model technology, vector database technology, and security alarm determination and automated response technology, and combines technical means from multiple fields such as natural language processing, machine learning, and data retrieval to realize an automated process from alarm determination to security response, which is used to improve the efficiency and accuracy of intranet security threat detection and response, and also improves the intelligence level of the method.

[0153] The present invention proposes a method for intelligently identifying and automatically handling malicious intranet traffic, which reduces the reliance on manual addition of whitelists. Through automated similarity comparison and large-scale model judgment, it can more comprehensively and accurately identify effective alarms and avoid security risks caused by the limitations of whitelist rules.

[0154] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A method for intelligent identification and automatic handling of malicious intranet traffic, characterized in that: include: Monitor intranet traffic data through security sensors, alarm logs, and file threat identifiers to obtain security alarm information; Performing data cleaning on the security alarm information and formatting it according to a predefined format to obtain pre-processed alarm information; Extract key features based on the pre-processed alarm information, and perform vector processing to obtain vector alarm feature information; Calculate similarity between the vector alarm feature information and the vector data information in the vector database, and obtain similar vector data information in the vector database whose similarity to the vector alarm feature information is higher than a preset threshold; Constructing a prompt template based on the pre-processed alarm information, the similar vector data information and the constraint rules, and inputting the prompt template into the large model for processing to obtain a determination result; According to the judgment result output by the large model, the corresponding disposal strategy is executed, and the judgment result and disposal situation are fed back to the security management personnel.

2. The method for intelligently identifying and automatically handling malicious traffic in an intranet according to claim 1, characterized in that: The step of performing data cleaning on the security alarm information and formatting it according to a predefined format to obtain pre-processed alarm information includes: Obtaining useless symbols in the security alarm information and filtering them out, replacing special symbols with a preset reference library, and obtaining alarm text information; Detect duplicate values, missing values, and abnormal values in the alarm text information. When duplicate values are detected in the alarm text information, delete the values. When missing values are detected in the alarm text information, obtain relevant data of the missing values for prediction and recovery. When abnormal values are detected in the alarm text information, truncate and replace the abnormal values according to preset peak data to obtain pre-processing and cleaning alarm information. The cleaning alarm information is standardized according to a predefined format, and data binning is performed based on time information to obtain pre-processed alarm information.

3. The method for intelligently identifying and automatically handling malicious traffic in an intranet according to claim 1, characterized in that: The step of extracting key features based on the pre-processed alarm information and performing vector processing to obtain vector alarm feature information includes: Extracting key feature information from the pre-processed alarm information according to preset feature fields; Obtaining a keyword group in the key feature information, calculating the frequency of occurrence of the keyword group in the key feature information, and performing logarithmic processing to obtain logarithmic frequency data corresponding to the keyword group; Calculating the inverse document frequency of the keyword phrase in the feature database, calculating the feature weight corresponding to the keyword phrase, and generating vector feature sub-information; Among them, w i is the feature weight corresponding to the i-th keyword group, f i is the frequency of occurrence of the i-th keyword group in the key feature information, N is the number of all data groups in the feature database, n i is the number of data groups containing the keyword group in the feature database, and ε is a very small constant; The vector alarm feature information is obtained by calculating the vector feature sub-information corresponding to all the keyword groups in the key feature information and performing position coding.

4. The method for intelligently identifying and automatically handling malicious traffic in an intranet according to claim 1, characterized in that: The step of calculating similarity between the vector alarm feature information and the vector data information in the vector database, and obtaining similar vector data information in the vector database whose similarity to the vector alarm feature information is higher than a preset threshold, includes: According to the vector data information in the vector database, a number of vector centroid points are randomly selected; optionally, the method includes: C i =C lb +(C ub -C lb )·R i Among them, C i is the randomly selected centroid point of the i-th vector, C lb is the lower bound of the search space constructed based on the vector database, C ub is the upper bound of the search space constructed based on the vector database, R i is the i-th random number; Among them, mod is the modulo operation, R i-1 is the i-1th random number, a is the first control parameter, and b is the second control parameter; By calculating the Euclidean distance between the vector centroid and the remaining vector data information in the vector database, when the Euclidean distance is lower than the Euclidean distance threshold requirement, constructing the vector centroid and the vector data information into a vector set, and calculating the mean vector of the vector set; traversing and calculating the Euclidean distance between the vector set and the vector data information in the vector database that is not included in the vector set, until all the vector data information in the vector database is divided into the corresponding vector set, and updating the mean vector of the vector set in real time; Acquire the vector alarm feature information, calculate the vector set distance between the vector alarm feature information and the mean vector corresponding to each vector set, and select the vector set with the smallest vector set distance as the pre-search vector space; Setting a number of search points to select vector data information in the pre-retrieval vector space, calculating the Euclidean distance with the vector alarm feature information, when the Euclidean distance between the vector data information selected by the search point and the vector alarm feature information is less than the Euclidean distance between the corresponding vector set in the pre-retrieval vector space and the vector alarm feature information, the search point randomly moves its position in the neighborhood until a maximum number of iterations is reached, obtaining the search point corresponding to the vector data information with the closest Euclidean distance to the vector alarm feature information as the optimal search point, and obtaining similar vector data information; when the Euclidean distance between the vector data information selected by the search point and the vector alarm feature information is greater than the Euclidean distance between the corresponding vector set in the pre-retrieval vector space and the vector alarm feature information, the search point quickly leaves the current neighborhood and updates the position of the search point; in, The position of the ith search point after t+1 iterations, is the position of the ith search point after t iterations, τ is the adjustment parameter, σ is a uniform random number, t max is the maximum number of iterations, p1, p2, p3 are random numbers in the interval [0,1], L t Search point location The Euclidean distance between the corresponding vector data information and the vector alarm feature information, L ave is the Euclidean distance between the corresponding vector set in the pre-retrieval vector space and the vector alarm feature information, The degrees of freedom are T distribution, m is the scale factor, is the iteration critical value, and k is a random integer.

5. The method for intelligently identifying and automatically handling malicious traffic in an intranet according to claim 1, characterized in that: The step of constructing a prompt template based on the pre-processed alarm information, the similar vector data information and the constraint rules, and inputting the prompt template into the large model for processing to obtain a judgment result includes: Obtaining a preset prompt word template can be implemented as "based on [] conditions, according to [] and [], give []"; The pre-processed alarm information, the similar vector data information and the constraint rules are input into the preset prompt word template to generate a prompt template; which can be implemented as "based on the [constraint rule] condition, according to [pre-processed alarm information] and [similar vector data information], give [judgment result]".

6. The method for intelligently identifying and automatically handling malicious traffic in an intranet according to claim 1, characterized in that: The step of constructing a prompt template based on the pre-processed alarm information, the similar vector data information and the constraint rules, and inputting the prompt template into the large model for processing to obtain a judgment result includes: Obtain sample malicious traffic data and sample normal traffic data; Based on the forward diffusion sub-model, noise data is added to the sample malicious traffic data and the sample normal traffic data according to a preset step size to obtain malicious traffic noise data and normal traffic noise data; based on the reverse diffusion sub-model, the malicious traffic noise data and the normal traffic noise data are restored according to the noise data to obtain restored sample malicious data and restored sample normal data; Optimizing the reverse diffusion sub-model by analyzing the deviation between the recovered sample malicious data and the sample malicious traffic data, and the deviation between the recovered sample normal data and the sample normal traffic data; The malicious traffic data of the samples are fused with the malicious traffic noise data to construct a malicious sample set; the normal traffic data of the samples are fused with the normal traffic noise data to construct a normal sample set; Build a pre-trained large model based on the encoder, decoder and classifier; The encoder extracts and analyzes sample features of the malicious sample set and the normal sample set respectively, and outputs sample classification information through the decoder. The classifier makes classification decisions based on the sample classification information to obtain predicted classification results. The pre-trained large model is trained based on the predicted classification results and the sample classification results to obtain a large model.

7. The method for intelligently identifying and automatically handling malicious traffic in an intranet according to claim 6, characterized in that: The step of constructing a prompt template based on the pre-processed alarm information, the similar vector data information and the constraint rules, and inputting the prompt template into the large model for processing to obtain a judgment result includes: The pre-processed alarm information is processed by the reverse diffusion sub-model to obtain restored alarm information, the alarm fusion information is subjected to random masking and position encoding by the encoder in the large model, key features are extracted based on the attention mechanism, and model feature information is obtained; The encoder performs feature extraction on the similar vector data information to obtain similar feature information; the model feature information and the similar feature information are fused according to preset weights to obtain fused feature information; the decoder outputs feature classification information based on constraint rules, and the classifier makes classification decisions to obtain judgment results.

8. The method for intelligently identifying and automatically handling malicious traffic in an intranet according to claim 1, characterized in that: The steps include: constructing a prompt template based on the pre-processed alarm information, the similar vector data information and the constraint rules, and inputting the prompt template into the large model for processing to obtain a judgment result; Also includes: Obtain several large models, process them respectively according to the prompt template, and obtain corresponding model judgment results; According to the sample traffic data, the model judgment performance of the large model is evaluated based on the model evaluation index, and the corresponding model evaluation weight is obtained. According to the model judgment result and the model evaluation weight, the judgment result is obtained.

9. The method for intelligently identifying and automatically handling malicious traffic in an intranet according to claim 1, characterized in that: The step of executing a corresponding disposal strategy based on the determination result output by the large model and feeding back the determination result and disposal situation to the security management personnel includes: When the determination result is "ignore" or "false alarm", the security alarm information is marked as an invalid alarm; When the judgment result is "antivirus recommended", the intranet traffic data is judged to be malicious traffic, and antivirus operations are performed on the intranet traffic data and related device terminals by calling an artificial intelligence program, and the operation log of the intranet traffic data and the called artificial intelligence program are recorded.

Citation Information

Patent Citations

  • Malicious traffic detection method, system and device and storage medium

    CN116599683A

  • Universal method and system for identifying attack success through large model

    CN119728193A

  • Network risk assessment method and system based on multi-modal data pre-training model

    CN119814354A

  • Keyword generation method and apparatus, and electronic device and computer storage medium

    WO2022134759A1

Cited By

  • Intelligent processing method and system for intranet malicious traffic, storage medium and electronic equipment

    CN121727775A