Detection model training and challenge black hole attack detection method and device

By generating target training samples in segments and training detection models, the problem of challenging black hole attack detection in the existing technology is solved, and more efficient attack recognition and defense is achieved.

CN120151041APending Publication Date: 2025-06-13BEIJING BAIDU NETCOM SCI & TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510323311.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has poor accuracy when detecting challenging black hole attacks, making it difficult to effectively defend against such attacks.

Method used

By obtaining traffic data that challenges black hole attacks, generating target training samples in segments, and using deep learning technology to train detection models, and then detecting traffic data.

Benefits of technology

It improves the accuracy of challenging black hole attack detection, can more effectively identify and defend against such attacks, and improves network security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151041A_ABST
    Figure CN120151041A_ABST
Patent Text Reader

Abstract

The invention provides a detection model training and challenge black hole attack detection method and device, and relates to the field of artificial intelligence such as network security, deep learning and large models. The detection model training method comprises the following steps: acquiring first traffic data generated by a first port in a first time period, wherein challenge black hole attacks occur in part of time in the first time period; the first time period is averagely divided into M parts, M second time periods are obtained, M is a positive integer larger than 1, a target training sample is generated according to second traffic data corresponding to the second time periods, and the second traffic data are data fragments in the first traffic data; and training the detection model according to the target training sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to fields such as network security, deep learning, and large models, and particularly to a method and device for detecting model training and challenge collapsar attack detection. Background Art

[0002] Challenge Collapsar (CC) attack is a distributed denial of service (DDoS) attack mainly targeting web servers. By exploiting a large number of forged Hypertext Transfer Protocol (HTTP) requests, it consumes the computing resources of the server, making it unable to process normal requests. Summary of the Invention

[0003] The present disclosure provides a method and device for detecting model training and challenge collapsar attack detection.

[0004] A method for detecting model training includes:

[0005] Obtaining first traffic data generated by a first port within a first time period, where a challenge collapsar attack occurs during part of the first time period;

[0006] Dividing the first time period evenly into M parts to obtain M second time periods, where M is a positive integer greater than 1, and generating target training samples according to the second traffic data corresponding to each second time period, where the second traffic data is a data segment in the first traffic data;

[0007] Training the detection model according to the target training samples.

[0008] A method for challenge collapsar attack detection includes:

[0009] Obtaining third traffic data to be detected, where the third traffic data is traffic data generated by a second port within a third time period;

[0010] Detecting the third traffic data by using a detection model to obtain a detection result on whether the third traffic data has a challenge collapsar attack, where the detection model is a model trained according to the above method.

[0011] A device for detecting model training includes: a first acquisition module, a sample generation module, and a model training module;

[0012] The first acquisition module is configured to obtain first traffic data generated by a first port within a first time period, where a challenge collapsar attack occurs during part of the first time period;

[0013] The sample generation module is configured to divide the first time period into M equal parts to obtain M second time periods, where M is a positive integer greater than 1, and generate target training samples according to the second traffic data corresponding to each second time period, and the second traffic data is a data segment in the first traffic data;

[0014] The model training module is configured to train the detection model according to the target training samples.

[0015] A challenge black hole attack detection device includes: a second acquisition module and an attack detection module;

[0016] The second acquisition module is configured to acquire third traffic data to be detected, and the third traffic data is traffic data generated by a second port within a third time period;

[0017] The attack detection module is configured to use the detection model to detect the third traffic data to obtain a detection result on whether a challenge black hole attack occurs in the third traffic data, and the detection model is a model trained according to the above method.

[0018] An electronic device includes:

[0019] At least one processor; and

[0020] A memory communicatively connected to the at least one processor; wherein,

[0021] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0022] A non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method as described above.

[0023] A computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the method as described above is implemented.

[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0026] Figure 1 is a flowchart of the first embodiment of the detection model training method described in the present disclosure;

[0027] Figure 2 Flow chart of the second embodiment of the detection model training method described in the present disclosure;

[0028] Figure 3 Flow chart of the embodiment of the CC attack detection method described in the present disclosure;

[0029] Figure 4 Schematic diagram of the overall implementation process of the CC attack detection method described in the present disclosure;

[0030] Figure 5 Schematic diagram of the composition structure of the detection model training apparatus 500 according to the embodiment of the present disclosure;

[0031] Figure 6 Schematic diagram of the composition structure of the CC attack detection apparatus 600 according to the embodiment of the present disclosure;

[0032] Figure 7 Schematic block diagram of an electronic device 700 that can be used to implement the embodiments of the present disclosure. Detailed implementation manners

[0033] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0034] In addition, it should be understood that the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0035] Figure 1 Flow chart of the first embodiment of the detection model training method described in the present disclosure. As Figure 1 shown, it includes the following specific implementation manners.

[0036] In step 101, obtain first traffic data generated by a first port within a first time period, and a CC attack occurs during part of the first time period.

[0037] In step 102, the first time period is evenly divided into M parts to obtain M second time periods, where M is a positive integer greater than 1. A target training sample is generated according to the second traffic data corresponding to each second time period, and the second traffic data is a data segment in the first traffic data.

[0038] In step 103, the detection model is trained according to the target training sample.

[0039] CC attacks usually have the following characteristics: simulating real user behaviors, it is difficult to defend against by simply blocking Internet Protocol (IP) addresses; targeting the application layer, unlike other DDoS attacks that rely on large traffic, but attacking the computing resources of the server; the attack traffic is relatively small, but it can cause serious impacts; it is easy to hide the attack source, and attackers can use proxy servers, etc. to launch attacks, with IP addresses widely distributed and difficult to trace.

[0040] Currently, traditional DDoS attack detection methods are usually used for CC attack detection. However, due to the above characteristics of CC attacks, the accuracy of the detection results is relatively poor.

[0041] By adopting the solution described in the above method embodiment, the target training data can be constructed according to the obtained first traffic data, and then the detection model can be trained according to the target training data. In this way, the detection model can be used to detect the traffic data subsequently to determine whether a CC attack has occurred. Correspondingly, with the powerful reasoning ability of the detection model, the accuracy of the detection results can be improved, and corresponding defense measures can be taken to improve the security of the traffic data, etc.

[0042] For example, if it is determined in a certain way that a CC attack has occurred on the first port, assuming that the CC attack lasts for 20 minutes, and assuming that no CC attack has occurred in the 30 minutes before and 30 minutes after these 20 minutes, then this 80 (30 + 20 + 30) - minute period can be determined as the first time period, and the traffic data generated by the first port within the first time period can be obtained as the first traffic data.

[0043] After that, the first time period can be evenly divided into M parts, thus obtaining M second time periods, where M is a positive integer greater than 1, and the specific value can be determined according to actual needs. For example, the first time period can be evenly divided into multiple second time periods in such a way that the duration of each second time period is 5 seconds.

[0044] Furthermore, the target training sample of the detection model can be generated according to the second traffic data corresponding to each second time period. Each second time period respectively corresponds to a data segment in the first traffic data, and the data segment can be referred to as the second traffic data.

[0045] In some embodiments of the present disclosure, the manner of generating target training samples based on the second traffic data corresponding to each second time period may include: respectively obtaining the connection features of the second traffic data corresponding to each second time period, generating initial training samples corresponding to each second time period according to the connection features, and generating target training samples according to the initial training samples.

[0046] The connection features refer to the relevant information for establishing a connection with the client, and usually include multiple specific features, such as: the number of newly established Synchronize Sequence Numbers (SYN), the number of newly established Acknowledge Characters (ACK), the number of half-open connections, the number of concurrent connections, the number of active connections, the inbound packet flag bits, the outbound packet flag bits, the number of closed connections, and the number of failed connections, etc. Among them, the closed connections may include client close, service close, client reset, service reset, service flash break, and timeout close, etc., and the failed connections may include client ignore, client reject, and service reject, etc.

[0047] There is no limitation on how to obtain the connection features of the second traffic data. For example, it can be obtained by analyzing and statistically processing the second traffic data.

[0048] Although the changes brought by the CC attack at the traffic level are not significant enough, there will still be changes at the connection feature level. Therefore, in the solution described in the present disclosure, training samples are constructed based on the connection features so that the detection model can learn the changes in the connection features when a CC attack occurs, and thus accurately detect the CC attack.

[0049] In some embodiments of the present disclosure, the initial training samples may include: initial positive samples and initial negative samples. The manner of generating initial training samples corresponding to each second time period according to the connection features may include: for any second time period, in response to determining that no CC attack occurs within this second time period, generating an initial positive sample corresponding to this second time period, where the initial positive sample includes: the connection features corresponding to this second time period and a first label; in response to determining that a CC attack occurs within this second time period, generating an initial negative sample corresponding to this second time period, where the initial negative sample includes: the connection features corresponding to this second time period and a second label.

[0050] For example, for any second time period, if it is determined that this second time period is within the 20 minutes during which the aforementioned CC attack persists, an initial negative sample corresponding to this second time period can be generated. The initial negative sample may include the connection features corresponding to this second time period and a second label. On the contrary, an initial positive sample corresponding to this second time period can be generated, and the initial positive sample may include the connection features corresponding to this second time period and a first label.

[0051] Among them, the first tag is used to indicate that no CC attack has occurred, and its value can be 0. The second tag is used to indicate that a CC attack has occurred, and its value can be 1.

[0052] In the above manner, for each second time period, corresponding initial training samples can be generated respectively. Moreover, the initial training samples can be generated only based on the connection features and corresponding tags corresponding to each second time period, which is very simple and convenient, thereby improving the generation efficiency of the initial training samples and laying a good foundation for the construction of subsequent target training samples.

[0053] In some embodiments of the present disclosure, the target training samples may include: target positive samples and target negative samples. When generating the target training samples from the initial training samples, the initial training samples can be sorted first in the order of corresponding time from the earliest to the latest. Then, based on the sorting result and the first reference point selected from each initial training sample, feature optimization processing can be performed on each initial training sample other than the first reference point. Furthermore, the initial positive samples after feature optimization processing can be determined as target positive samples, and the initial negative samples after feature optimization processing can be determined as target negative samples.

[0054] Specifically, in some embodiments of the present disclosure, the initial training sample ranked first after sorting can be determined as the first reference point, and the first reference point is an initial positive sample. Then, each initial training sample after the first reference point can be traversed in sequence, and for each currently traversed sample, the following processing can be performed respectively: in response to determining that it does not meet the first update condition, obtain the change ratio of the connection feature in the current sample compared to the connection feature in the latest obtained first reference point, and use the change ratio to replace the connection feature in the current sample, and then continue to traverse the next initial training sample; in response to determining that it meets the first update condition, update the first reference point to the current sample, and then continue to traverse the next initial training sample.

[0055] For example, assume that there are a total of 1000 (the number is only for illustrative purposes) initial training samples, including initial positive samples and initial negative samples. For the sake of convenience in description, the 1000 sorted initial training samples are respectively called initial training sample 1 to initial training sample 1000. Initial training sample 1 is usually an initial positive sample. Then, initial training sample 1 can be determined as the first reference point. After that, for initial training sample 2 (the current sample), it can be determined whether it meets the first update condition. If it does not meet, the change ratio of the connection feature in initial training sample 2 compared to the connection feature in initial training sample 1 (the first reference point) can be obtained, and the connection feature in initial training sample 2 can be replaced with the change ratio to obtain the initial training sample 2 after feature optimization processing. If it meets, the first reference point can be updated to initial training sample 2. Similarly, for each initial training sample after initial training sample 2, the processing method of initial training sample 2 can be respectively adopted for processing.

[0056] The connection features in the initial training samples are usually in the form of specific values. If the training of the detection model is directly based on the values, it may affect the accuracy of the training result because there are differences in the connection features of different ports. For example, for port A, an increase in the number of half connections from 30 to 100 indicates a CC attack, while for port B, the normal number of half connections may be close to or greater than 100. Therefore, in the solution proposed in the present disclosure, a dynamic feature processing method can be adopted, that is, the change ratio of the connection feature with the first reference point is used to replace the connection features in each initial training sample other than the first reference point, so as to normalize the connection features in each initial training sample other than the first reference point, overcome the differences between the connection features of different ports as much as possible, and then improve the accuracy of the subsequent model training result, etc.

[0057] The number of connection features included in each initial training sample and the specific feature content are the same. For example, each includes 5 connection features, which are respectively referred to as connection feature 1, connection feature 2, connection feature 3, connection feature 4, and connection feature 5 for the sake of convenience of description. Taking the above initial training sample 2 as an example, change ratios 1, 2, 3, 4, and 5 can be respectively obtained. Among them, change ratio 1 is the change ratio of connection feature 1 in initial training sample 2 compared to connection feature 1 in initial training sample 1, change ratio 2 is the change ratio of connection feature 2 in initial training sample 2 compared to connection feature 2 in initial training sample 1, change ratio 3 is the change ratio of connection feature 3 in initial training sample 2 compared to connection feature 3 in initial training sample 1, change ratio 4 is the change ratio of connection feature 4 in initial training sample 2 compared to connection feature 4 in initial training sample 1, and change ratio 5 is the change ratio of connection feature 5 in initial training sample 2 compared to connection feature 5 in initial training sample 1. Correspondingly, change ratio 1 can be used to replace connection feature 1 in initial training sample 2, change ratio 2 can be used to replace connection feature 2 in initial training sample 2, change ratio 3 can be used to replace connection feature 3 in initial training sample 2, change ratio 4 can be used to replace connection feature 4 in initial training sample 2, and change ratio 5 can be used to replace connection feature 5 in initial training sample 2.

[0058] In addition, the change ratio can refer to the ratio of two corresponding connection features. Taking change ratio 1 as an example, change ratio 1 can refer to the ratio of the value of connection feature 1 in initial training sample 2 to the value of connection feature 1 in initial training sample 1. For example, if the value of connection feature 1 in initial training sample 2 is 40 and the value of connection feature 1 in initial training sample 1 is 50, then the ratio is 0.8.

[0059] In some embodiments of the present disclosure, determining that the first update condition is met may include: in response to determining that the number of initial training samples traversed since the last update of the first reference point is greater than the first threshold and determining that the current sample is an initial positive sample, determining that the first update condition is met.

[0060] The specific value of the first threshold can be determined according to actual needs, such as 99. Still taking the above initial training samples 1 to 1000 as an example, at the beginning, the initial training sample 1 is determined as the first reference point, and then each initial training sample after the initial training sample 1 is traversed in turn. Assuming that the initial training samples 2 to 100 do not meet the first update condition, then the change ratios of the connection features in the initial training samples 2 to 100 compared to the connection features in the initial training sample 1 can be obtained respectively, and the obtained change ratios can be used to replace the connection features in the initial training samples 2 to 100 respectively. After that, when traversing to the initial training sample 101, assuming that it meets the first update condition, that is, the number of initial training samples (initial training samples 2 to 101) traversed since the most recent update of the first reference point (setting the initial training sample 1 as the first reference point) is greater than 99, and the initial training sample 101 is an initial positive sample, then the first reference point can be updated from the initial training sample 1 to the initial training sample 101. After that, the subsequent initial training samples after the initial training sample 101 can be continued to be traversed. Assuming that when traversing to the initial training sample 201, at this time, the number of initial training samples (initial training samples 102 to 201) traversed since the most recent update of the first reference point (updating the first reference point to the initial training sample 101) is greater than 99, but the initial training sample 201 is an initial negative sample, then it can be determined that it does not meet the first update condition. Assuming that the initial training samples 201 to 250 are all initial negative samples, then when traversing to the initial training sample 251, it will be determined to meet the first update condition. Correspondingly, the first reference point can be updated from the initial training sample 101 to the initial training sample 251. After that, the subsequent initial training samples after the initial training sample 251 can be continued to be traversed.

[0061] The connection features at different times are constantly changing. If the change ratio is always calculated based on a first reference point, it may lead to the calculated change ratio not accurately reflecting the actual change situation. Therefore, the first reference point can be updated each time it is determined to meet the first update condition, so that the calculated change ratio can more accurately reflect the actual change situation.

[0062] In addition, in the solution described in the present disclosure, the initial negative sample will not be determined as the first reference point. This is to more accurately judge the change of the connection features during a CC attack compared to the connection features in the normal situation. Because if the initial negative sample is determined as the first reference point, then when calculating the change ratio of the subsequent initial negative samples compared to the connection features of the first reference point, there may be no change, that is, the change ratio is 1, thus losing the judgment value.

[0063] After performing feature optimization processing on each initial training sample other than the first reference point in the above manner, the initial positive sample after feature optimization processing can be determined as the target positive sample, and the initial negative sample after feature optimization processing can be determined as the target negative sample, thereby obtaining the target training samples.

[0064] After obtaining a sufficient number of target training samples, the detection model can be trained based on the target training samples. The detection model can be a neural network model adopting the Extreme Gradient Boosting (XGBoost) architecture.

[0065] Combined with the above introduction, Figure 2 is a flowchart of the second embodiment of the detection model training method described in this disclosure. As Figure 2 shown, it includes the following specific implementation manners.

[0066] In step 201, obtain the first traffic data generated by the first port within the first time period, and a CC attack occurs during part of the first time period.

[0067] In step 202, evenly divide the first time period into M parts to obtain M second time periods, where M is a positive integer greater than 1.

[0068] In step 203, respectively obtain the connection features of the second traffic data corresponding to each second time period.

[0069] In step 204, for each second time period, if it is determined that no CC attack occurs within the second time period, generate the initial positive sample corresponding to the second time period, and the initial positive sample includes: the connection feature corresponding to the second time period and the first label; if it is determined that a CC attack occurs within the second time period, generate the initial negative sample corresponding to the second time period, and the initial negative sample includes: the connection feature corresponding to the second time period and the second label; use the initial positive sample and the initial negative sample to form the initial training samples.

[0070] In step 205, sort the initial training samples in the order of corresponding time from the earliest to the latest, and perform feature optimization processing on each initial training sample other than the first reference point according to the sorting result and the first reference point selected from each initial training sample.

[0071] For example, the initial training sample ranked first after sorting can be determined as the first reference point, and the first reference point is the initial positive sample. Then, the initial training samples after the first reference point can be traversed in sequence, and for each currently traversed sample, the following processing can be performed respectively: in response to determining that the first update condition is not met, obtain the change ratio of the connection features in the current sample compared to the connection features in the latest obtained first reference point, and use the change ratio to replace the connection features in the current sample, and then continue to traverse the next initial training sample; in response to determining that the first update condition is met, update the first reference point to the current sample, and then continue to traverse the next initial training sample.

[0072] Among them, determining that the first update condition is met may include: in response to determining that the number of initial training samples traversed since the last update of the first reference point is greater than the first threshold, and determining that the current sample is an initial positive sample, it is determined that the first update condition is met.

[0073] In step 206, the processed initial positive sample is determined as the target positive sample, the processed initial negative sample is determined as the target negative sample, and the target positive sample and the target negative sample are used to form the target training sample.

[0074] In step 207, the detection model is trained according to the target training sample.

[0075] The detection model can be a large model. A large model is an ultra-large-scale language model built based on deep learning technology, which can generate natural language text or understand the meaning of natural language text, etc. After the training of the detection model is completed, it can be applied to the actual CC attack detection.

[0076] Correspondingly, Figure 3 This is the flowchart of the embodiment of the CC attack detection method described in this disclosure. As Figure 3 shown, it includes the following specific implementation manners.

[0077] In step 301, the third traffic data to be detected is obtained, and the third traffic data is the traffic data generated by the second port within the third time period.

[0078] In step 302, the detection model is used to detect the third traffic data to obtain the detection result of whether the third traffic data has a CC attack.

[0079] Among them, the detection model is the model trained according to the method shown in Figure 1 or Figure 2 shown.

[0080] By implementing the solution described in the above method embodiment, a pre-trained detection model can be used to detect CC attacks. Correspondingly, with the powerful inference ability of the detection model, the accuracy of the detection result can be improved, and corresponding defense measures can be taken to improve the security of traffic data, etc.

[0081] In some embodiments of the present disclosure, the connection features of the third traffic data can be obtained, and then, based on the connection features, a required detection result can be generated using the detection model.

[0082] Specifically, in some embodiments of the present disclosure, the change ratio of the connection features of the third traffic data compared to the connection features in the second reference point can be obtained. The second reference point includes: the connection features of the fourth traffic data generated by the second port within the fourth time period. The durations of the third time period and the fourth time period are the same, and the fourth time period is earlier than the third time period, and no CC attack occurs within the fourth time period. Then, the change ratio can be input into the detection model to obtain the output detection result.

[0083] The second port and the aforementioned first port can be the same port or different ports. Additionally, the durations of the third time period and the fourth time period can be the same as the duration of the second time period.

[0084] In practical applications, the third traffic data can be obtained in real time. For example, the traffic data generated within every 5 seconds can be used as the third traffic data respectively, and for each obtained third traffic data, it can be processed respectively according to Figure 3 the manner shown.

[0085] Suppose the latest obtained third traffic data is the traffic data generated within the time period from 10 to 20 seconds at 8:30 am on March 12, 2025 (the third time period). Then, the connection features of the third traffic data can be obtained. Additionally, the fourth time period corresponding to the second reference point can be a time period before the third time period when no CC attack occurs, such as from 0 to 10 seconds at 7:30 am on March 12, 2025. The fourth traffic data generated during this time period can be obtained, and the connection features of the fourth traffic data can be obtained as the second reference point. Correspondingly, the change ratio of the connection features of the third traffic data compared to the connection features in the second reference point can be obtained. The manner of obtaining the change ratio can be the same as the manner of obtaining the change ratio during the training of the detection model. Then, the obtained change ratio can be input into the detection model to obtain the output detection result.

[0086] That is, in the same processing manner as during the training of the detection model, the change ratio of the connection features of the traffic data compared to the connection features in the reference point can be obtained. Correspondingly, the change ratio can be input into the detection model, thereby improving the accuracy of the obtained detection result and further improving the security of the traffic data, etc.

[0087] In practical applications, the output detection result may refer to outputting a first value or a second value. For example, the first value can be 1, which is used to indicate that a CC attack has occurred in the third traffic data, and the second value can be 0, which is used to indicate that no CC attack has occurred in the detected third traffic data.

[0088] In some embodiments of the present disclosure, the second reference point can also be updated. For example, in response to determining that the second update condition is met, the connection characteristics of the fifth traffic data generated by the second port within the fifth time period can be obtained, and the second reference point can be updated to the connection characteristics of the fifth traffic data. The durations of the fourth time period and the fifth time period are the same, and the fifth time period is later than the fourth time period, and no CC attack has occurred within the fifth time period.

[0089] Meeting the second update condition can refer to passing a predetermined duration, etc. For example, every 2 hours, the second reference point can be updated once. Generally speaking, the time period corresponding to the second reference point and the time period corresponding to the third traffic data should not be too far apart.

[0090] By updating the second reference point, the accuracy of the obtained change ratio can be improved, thereby further improving the accuracy of the obtained detection result, and further improving the security of the traffic data, etc.

[0091] Combined with the above introduction, Figure 4 This is a schematic diagram of the overall implementation process of the CC attack detection method described in the present disclosure. As Figure 4 shown, the third traffic data to be detected can be obtained in real time, and the connection characteristics of the third traffic data can be obtained. Then, the change ratio of the connection characteristics of the third traffic data compared with the connection characteristics in the second reference point can be obtained. Furthermore, the change ratio can be input into a pre-trained detection model to obtain the output detection result.

[0092] Among them, when the first value is output, that is, when it is detected that a CC attack has occurred in the third traffic data, an alarm can also be made in a predetermined manner. There is no limitation on how to make the alarm. For example, a predetermined alarm sound can be emitted, or the alarm information can be notified to relevant personnel by email or text message, etc. When the first value is output, that is, when it is detected that no CC attack has occurred in the third traffic data, no processing is required. In addition, the second reference point can also be updated when the second update condition is met.

[0093] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited by the described action sequence, because according to the present disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present disclosure. In addition, for the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions in other embodiments.

[0094] The above is the introduction to the method embodiments. The following further illustrates the solution of the present disclosure through device embodiments.

[0095] Figure 5 It is a schematic structural diagram of the composition of the detection model training device embodiment 500 of the present disclosure. As Figure 5 shown, it includes: a first acquisition module 501, a sample generation module 502, and a model training module 503.

[0096] The first acquisition module 501 is configured to acquire first traffic data generated by a first port within a first time period, and a CC attack occurs during a part of the first time period.

[0097] The sample generation module 502 is configured to divide the first time period into M equal parts to obtain M second time periods, where M is a positive integer greater than 1, and generate target training samples according to the second traffic data corresponding to each second time period, and the second traffic data is a data segment in the first traffic data.

[0098] The model training module 503 is configured to train the detection model according to the target training samples.

[0099] In some embodiments of the present disclosure, the manner in which the sample generation module 502 generates target training samples according to the second traffic data corresponding to each second time period may include: respectively acquiring the connection features of the second traffic data corresponding to each second time period, respectively generating initial training samples corresponding to each second time period according to the connection features, and generating target training samples according to the initial training samples.

[0100] In some embodiments of the present disclosure, the initial training samples may include: initial positive samples and initial negative samples. The manner in which the sample generation module 502 generates the initial training samples corresponding to each second time period according to the connection features may include: for any second time period, in response to determining that no CC attack has occurred during this second time period, generating an initial positive sample corresponding to this second time period, where the initial positive sample includes: the connection features corresponding to this second time period and a first label; in response to determining that a CC attack has occurred during this second time period, generating an initial negative sample corresponding to this second time period, where the initial negative sample includes: the connection features corresponding to this second time period and a second label.

[0101] In some embodiments of the present disclosure, the target training samples may include: target positive samples and target negative samples. When the sample generation module 502 generates target training samples based on the initial training samples, it may first sort the initial training samples in the order of corresponding time from earliest to latest, and then, according to the sorting result and a first reference point selected from each of the initial training samples, perform feature optimization processing on each of the initial training samples other than the first reference point. Furthermore, the initial positive samples after the feature optimization processing may be determined as target positive samples, and the initial negative samples after the feature optimization processing may be determined as target negative samples.

[0102] Specifically, in some embodiments of the present disclosure, the sample generation module 502 may determine the initial training sample ranked first after sorting as the first reference point, and the first reference point is an initial positive sample. Then, it may sequentially traverse each of the initial training samples after the first reference point, and for each currently traversed sample, perform the following processing respectively: in response to determining that it does not meet the first update condition, obtain the change ratio of the connection features in the current sample compared to the connection features in the latest obtained first reference point, and use the change ratio to replace the connection features in the current sample, and then continue to traverse the next initial training sample; in response to determining that it meets the first update condition, update the first reference point to the current sample, and then continue to traverse the next initial training sample.

[0103] In some embodiments of the present disclosure, the sample generation module 502 may determine that it meets the first update condition in response to determining that the number of initial training samples traversed since the last update of the first reference point is greater than the first threshold and determining that the current sample is an initial positive sample.

[0104] Figure 6 It is a schematic structural diagram of the composition of the CC attack detection device embodiment 600 described in the present disclosure. As Figure 6 shown, it includes: a second acquisition module 601 and an attack detection module 602.

[0105] The second acquisition module 601 is configured to acquire third traffic data to be detected, where the third traffic data is traffic data generated by the second port within a third time period.

[0106] The attack detection module 602 is configured to detect the third traffic data by using a detection model to obtain a detection result on whether a CC attack has occurred in the third traffic data.

[0107] Wherein, the detection model is a model trained according to the method shown in Figure 1 or Figure 2 the method shown.

[0108] In some embodiments of the present disclosure, the attack detection module 602 may obtain connection features of the third traffic data, and then, according to the connection features, use the detection model to generate a required detection result.

[0109] Specifically, in some embodiments of the present disclosure, the attack detection module 602 may obtain a change ratio of the connection features of the third traffic data compared to the connection features in a second reference point. The second reference point includes: connection features of fourth traffic data generated by the second port within a fourth time period. The third time period and the fourth time period have the same duration, and the fourth time period is earlier than the third time period, and no CC attack has occurred within the fourth time period. Then, the change ratio may be input into the detection model to obtain an output detection result.

[0110] In addition, in some embodiments of the present disclosure, the attack detection module 602 may also update the second reference point. For example, in response to determining that the second update condition is met, it may obtain connection features of fifth traffic data generated by the second port within a fifth time period, and may update the second reference point to the connection features of the fifth traffic data. The fourth time period and the fifth time period have the same duration, and the fifth time period is later than the fourth time period, and no CC attack has occurred within the fifth time period.

[0111] Figure 5 and Figure 6 The specific working process of the device embodiment shown may refer to the relevant description in the foregoing method embodiment.

[0112] In summary, by adopting the solution of the present disclosure, the connection features of traffic data can be used as a driving force, and the detection model can be used to timely and accurately detect CC attacks, so as to quickly take corresponding defense measures subsequently, thereby improving the security of traffic data. Moreover, it is applicable to different service scenarios and has wide applicability.

[0113] The solutions described in this disclosure can be applied to the field of artificial intelligence, especially in the fields of network security, deep learning, and large models. Artificial intelligence is a discipline that studies how to make computers simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0114] In addition, the traffic data and the like in the embodiments described in this disclosure are not targeted at a specific user and do not reflect the personal information of a specific user. In the technical solutions of this disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0115] According to an embodiment of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0116] Figure 7 FIG. shows a schematic block diagram of an electronic device 700 that can be used to implement the embodiments of this disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of this disclosure described and / or claimed herein.

[0117] As Figure 7 shown, the electronic device 700 includes a computing unit 701, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM, Read-Only Memory) 702 or the computer program loaded from the storage unit 708 into the random access memory (RAM, Random Access Memory) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O, Input / Output) interface 705 is also connected to the bus 704.

[0118] Multiple components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0119] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphic processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the methods described in the present disclosure can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the methods described in the present disclosure by any other suitable means (e.g., by means of firmware).

[0120] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or a general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0121] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0122] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0124] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0125] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0126] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0127] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A detection model training method, comprising: Acquire first traffic data generated by a first port in a first time period, wherein a challenge black hole attack occurs during part of the first time period; The first time period is evenly divided into M parts to obtain M second time periods, where M is a positive integer greater than 1, and a target training sample is generated according to second flow data corresponding to each second time period, where the second flow data is a data segment in the first flow data; The detection model is trained according to the target training samples.

2. The method according to claim 1, wherein: Generating target training samples according to the second flow data corresponding to each second time period includes: respectively obtaining connection features of the second flow data corresponding to each second time period, and respectively generating initial training samples corresponding to each second time period according to the connection features; The target training sample is generated according to the initial training sample.

3. The method according to claim 2, wherein: The initial training samples include: initial positive samples and initial negative samples; Generating the initial training samples corresponding to each second time period according to the connection features comprises: For any second time period, in response to determining that the challenge black hole attack does not occur in the second time period, generating the initial positive sample corresponding to the second time period, the initial positive sample including: the connection feature corresponding to the second time period and the first label; In response to determining that the challenge black hole attack occurs within the second time period, the initial negative sample corresponding to the second time period is generated, and the initial negative sample includes: a connection feature corresponding to the second time period and a second label.

4. The method according to claim 3, wherein: The target training samples include: target positive samples and target negative samples; Generating the target training sample according to the initial training sample comprises: Sort the initial training samples in order of their corresponding time from first to last; According to the sorting result and the first reference point selected from each initial training sample, performing feature optimization processing on each initial training sample other than the first reference point; The initial positive sample after feature optimization processing is determined as the target positive sample, and the initial negative sample after feature optimization processing is determined as the target negative sample.

5. The method according to claim 4, wherein: The step of performing feature optimization processing on each initial training sample other than the first reference point according to the sorting result and the first reference point selected from each initial training sample includes: Determine the initial training sample that is at the first position after sorting as the first reference point, and the first reference point is the initial positive sample; The initial training samples after the first reference point are traversed in sequence, and the following processing is performed for each current sample traversed: in response to determining that the first update condition is not met, the change ratio of the connection features in the current sample compared to the connection features in the first reference point obtained most recently is obtained, and the connection features in the current sample are replaced with the change ratio, and then the next initial training sample is traversed continuously; in response to determining that the first update condition is met, the first reference point is updated to the current sample, and then the next initial training sample is traversed continuously.

6. The method according to claim 5, wherein: The determining that the first update condition is met includes: In response to determining that the number of initial training samples traversed since the last update of the first reference point is greater than a first threshold, and determining that the current sample is the initial positive sample, it is determined that the first update condition is met.

7. A challenge black hole attack detection method, comprising: Acquire third flow data to be detected, where the third flow data is flow data generated by the second port in a third time period; The third traffic data is detected by using a detection model to obtain a detection result of whether the third traffic data has undergone a black hole attack challenge, wherein the detection model is a model trained according to the method described in any one of claims 1 to 6.

8. The method according to claim 7, wherein: The detecting of the third traffic data by using the detection model to obtain a detection result of whether the third traffic data has a challenge black hole attack includes: Acquiring a connection characteristic of the third flow data; The detection result is generated using the detection model according to the connection feature.

9. The method according to claim 8, wherein: Generating the detection result by using the detection model according to the connection feature includes: Obtaining a change ratio of a connection feature of the third flow data compared to a connection feature in a second reference point, wherein the second reference point includes: a connection feature of fourth flow data generated by the second port in a fourth time period, the third time period and the fourth time period have the same duration, the fourth time period is earlier than the third time period, and the challenge black hole attack does not occur in the fourth time period; The change ratio is input into the detection model to obtain the output detection result.

10. The method according to claim 9, further comprising: In response to determining that the second update condition is met, the connection characteristics of the fifth traffic data generated by the second port in the fifth time period are obtained, and the second reference point is updated to the connection characteristics of the fifth traffic data, the fourth time period and the fifth time period are the same in length, and the fifth time period is later than the fourth time period, and the challenge black hole attack does not occur in the fifth time period.

11. A detection model training device, comprising: A first acquisition module, a sample generation module, and a model training module; The first acquisition module is used to acquire first traffic data generated by a first port in a first time period, and a challenge black hole attack occurs during part of the first time period; The sample generation module is used to divide the first time period into M parts on average to obtain M second time periods, where M is a positive integer greater than 1, and to generate a target training sample according to second traffic data corresponding to each second time period, where the second traffic data is a data segment in the first traffic data; The model training module is used to train the detection model according to the target training sample.

12. A challenge black hole attack detection device, comprising: A second acquisition module and an attack detection module; The second acquisition module is used to acquire third flow data to be detected, where the third flow data is flow data generated by the second port in a third time period; The attack detection module is used to detect the third traffic data using a detection model to obtain a detection result of whether the third traffic data has undergone a black hole attack. The detection model is a model trained according to the method described in any one of claims 1 to 6.

13. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1 to 10.

15. A computer program product, comprising a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 10 is implemented.