Data processing method and device, storage medium and program product

CN119968810APending Publication Date: 2025-05-09YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070132.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The construction of existing automobile intrusion detection samples is not comprehensive enough, resulting in insufficient accuracy of network security assessment and inability to effectively cover attacks in unknown intrusion scenarios.

Method used

By attacking the device to be detected, the first type of attack samples are obtained, and the second type of attack samples are generated by applying noise, and the two are combined to build an intrusion detection sample set to cover known and unknown intrusion scenarios.

Benefits of technology

It improves the comprehensiveness and accuracy of the intrusion detection sample set, enhances the network security assessment capability of the equipment to be detected, and can more accurately define the anti-attack effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119968810A_ABST
    Figure CN119968810A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, a storage medium and a program product, is suitable for the technical field of network security, and is used for constructing a more comprehensive intrusion detection sample set. The method comprises the steps of obtaining a first type of attack samples by attacking to-be-detected equipment, obtaining a second type of attack samples by applying noise to the first type of attack samples, and constructing an intrusion detection sample set according to the first type of attack samples and the second type of attack samples. Through the method, the intrusion detection sample set not only can cover the attack samples in the known intrusion scene, but also can cover the attack samples in the unknown intrusion scene, the intrusion detection samples in the intrusion detection sample set are more comprehensive, and then the to-be-detected equipment is evaluated by using the more comprehensive intrusion detection sample set, so that the detection accuracy of the to-be-detected equipment is improved. And the accuracy of carrying out safety evaluation on the to-be-detected equipment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method, device, storage medium and program product Technical Field

[0001] The present application relates to the field of network security technology and provides a data processing method, device, storage medium and program product. Background Art

[0002] In recent years, automotive cybersecurity attacks have become increasingly frequent. According to reports, in 2021 alone, there were 256 publicly reported automotive cybersecurity attacks worldwide, a 225% increase compared to 2018. Given the rapid development of smart cars, automotive cybersecurity issues should also be taken seriously.

[0003] To improve automotive cybersecurity, the industry has developed intrusion detection samples based on typical intrusion scenarios. These samples are then used to conduct attack tests on vehicles before they leave the factory to assess their attack resistance. While this approach ensures that only vehicles with high attack resistance are shipped, since the intrusion detection samples are only developed for certain typical intrusion scenarios, they are not comprehensive and therefore hinder the accuracy of vehicle cybersecurity assessments.

[0004] Therefore, further research is needed on the construction of automobile intrusion detection samples.

[0005] Summary of the Invention

[0006] The present application provides a data processing method, apparatus, storage medium, and program product for constructing a more comprehensive intrusion detection sample set to improve the accuracy of network security assessment of devices to be detected (such as vehicles).

[0007] In a first aspect, the present application provides a data processing method applicable to a data processing device, which can be any device with processing capabilities, such as a server or a server cluster consisting of multiple servers. The method includes: the data processing device attacks a device to be detected to obtain a first type of attack sample; applies noise to the first type of attack sample to obtain a second type of attack sample; and then constructs an intrusion detection sample set based on the first and second type of attack samples.

[0008] In the above design, the first type of attack samples are obtained by attacking the device to be detected, and belong to attack samples in known intrusion scenarios, while the second type of attack samples are obtained by adding noise to the first type of attack samples, and can be considered as attack samples in unknown intrusion scenarios obtained by performing some deformation on the attack samples in known intrusion scenarios. In this way, an intrusion detection sample set is constructed by combining the first type of attack samples and the second type of attack samples, so that the intrusion detection sample set can cover both attack samples in known intrusion scenarios and attack samples in unknown intrusion scenarios, so that the intrusion detection samples in the intrusion detection sample set are more comprehensive. Furthermore, using a more comprehensive intrusion detection sample set to evaluate the device to be detected can also improve the accuracy of network security assessment of the device to be detected.

[0009] In a possible design, the first type of attack samples may include real attack samples and simulated attack samples, wherein the real attack samples are obtained by manually attacking the device to be detected, and the simulated attack samples are obtained by attacking the device to be detected by an attack tool.

[0010] In the above design, the human-generated attack method can be used to construct highly concealed attack behaviors or those that rely on business logic, thereby obtaining attack samples that are easily labeled in a real attack environment. Meanwhile, the attack tool-generated attack method can simulate the generation of attack traffic to obtain attack samples that are difficult to label in a real attack environment. Thus, by combining the human-generated and tool-generated attack methods to construct the first type of attack samples, the first type of attack samples can fully cover the various types of attacks that may exist in a real attack environment, thereby improving the comprehensiveness of the first type of attack samples.

[0011] In further design, the real attack samples can correspond to one or more of the following attack types: identity document (ID) non-existence attack, replay attack, tampering attack, data length error attack, signal out of defined range attack, context error attack, ID source non-specified electronic control unit (ECU) attack, identical ID attack, controller area network (CAN) scanning attack, unified diagnostic services (UDS) sensitive operation attack, message authentication error attack, ECU identity spoofing attack, man-in-the-middle attack, ECU authentication error attack, brute force attack, application layer protocol error attack, unknown stack connection attack, unknown stack connection attack.

[0012] In the above design, by providing multiple attack types that are easy to mark in real attack environments, users can choose one or more attack types according to actual needs to construct real attack samples, so as to be suitable for different real attack scenarios and improve the flexibility and versatility of constructing real attack samples.

[0013] In further design, the simulated attack samples can correspond to one or more of the following attack types: ID fuzzy attack, data fuzzy attack, CAN denial-of-service (Dos) attack, Ethernet (ethnic, ETH) Dos attack, malformed packet injection attack, and port scanning attack.

[0014] In the above design, by providing multiple attack types that are difficult to mark in a real attack environment, users can easily select one or more attack types according to actual needs to construct simulated attack samples, so as to be suitable for different simulated attack scenarios and improve the flexibility and versatility of constructing simulated attack samples.

[0015] In one possible design, a data processing device obtains a first type of attack sample by attacking a device to be detected, including: the data processing device traverses each of a plurality of preset attack types, and when traversing each attack type: executes an attack behavior corresponding to the attack type on the device to be detected, and obtains traffic data generated by the device to be detected for the attack behavior; and then, when it is determined that the traffic data is attack traffic, marks the traffic data as a first type of attack sample.

[0016] In the above design, the preset multiple attack types can be, for example, the multiple attack types that are easy to mark in a real attack environment and the multiple attack types that are not easy to mark in a real attack environment given in the aforementioned design. In this way, by collecting and marking traffic for each known attack type, the first type of attack samples can fully cover various known attack types, thereby improving the richness and comprehensiveness of the first type of attack samples.

[0017] In a further design, after the data processing device obtains the traffic data generated by the device to be detected for any attack behavior, if it is determined that the traffic data is not attack traffic but normal traffic, the traffic data can be marked as a non-attack sample. Then, after traversing all attack types, an intrusion detection sample set is constructed based on the first type attack sample, the second type attack sample and the non-attack sample.

[0018] In the above design, the intrusion detection sample set includes both attack samples (including the aforementioned first type attack samples and second type attack samples) and non-attack samples. This not only makes the sample information in the intrusion detection sample set more complete, but also can more accurately define the anti-attack effect of the device to be detected when using the intrusion detection sample set to perform an attack test on the device to be detected, based on whether the device to be detected can intercept attack samples and whether it can not intercept non-attack samples.

[0019] In a further design, after the data processing device obtains the traffic data generated by the device to be tested for any attack behavior, if it is determined that the traffic data neither meets the characteristics of attack traffic nor the characteristics of normal traffic, the traffic data can be determined as an unlabeled sample. Among them, unlabeled samples are unknown samples, that is, samples that cannot be identified as attack samples or non-attack samples under current technical means. In some scenarios, unlabeled samples are abnormal samples. For example, when the attack test of the data processing device causes the software and hardware system of the device to be tested to malfunction, the device to be tested itself may generate some abnormal data. These abnormal data neither meet the characteristics of normal traffic nor the characteristics of attack traffic, but will still be collected by the data processing device. Furthermore, for unlabeled samples, the data processing device can directly discard them to save the amount of data in the intrusion detection sample set. It can also construct an intrusion detection sample set based on attack samples, unlabeled samples and non-attack samples after traversing all attack types, so as to add all samples that actually exist when attacking the device to be detected to the intrusion detection sample set, thereby improving the sample richness in the intrusion detection sample set, and facilitating the subsequent labeling of unlabeled samples through other analyses or performing other operations.

[0020] In a further design, attack samples and non-attack samples in the intrusion detection sample set occupy the same proportion, for example, attack samples and non-attack samples each account for 50% of the total samples. This way, by evenly dividing attack samples and non-attack samples, the attack and non-attack samples in the intrusion detection sample set can be balanced, making it easier to extract the same proportion of data for attack testing of the device to be tested, thereby improving the credibility of the attack test results.

[0021] In a further design, after constructing an intrusion detection sample set, if the data processing device determines that the proportion of attack samples in the intrusion detection sample set is different from the proportion of non-attack samples in the intrusion detection sample set, the data processing device can trim the samples with a higher proportion so that the proportion of attack samples and non-attack samples after trimming is the same. The samples with a higher proportion are generally non-attack samples, also known as context data. In this way, by trimming the non-attack samples, the balance of attack samples and non-attack samples in the intrusion detection sample set can be maintained.

[0022] In one possible design, a data processing device obtains a second type of attack sample by applying noise to a first type of attack sample, including: the data processing device first applies noise to the first type of attack sample to obtain a perturbation sample, then inputs the perturbation sample into an attack recognition model, and obtains a recognition result output by the attack recognition model; when the recognition result indicates that it is impossible to determine whether the perturbation sample is an attack sample, the perturbation sample is determined to be a second type of attack sample; otherwise, the perturbation sample is adjusted according to the recognition result, and the adjusted perturbation sample is re-input into the attack recognition model, and the above process is repeated until the recognition result corresponding to the adjusted perturbation sample indicates that it is impossible to determine whether the adjusted perturbation sample is an attack sample, and the adjusted perturbation sample is determined to be a second type of attack sample.

[0023] In the above design, the second type of attack sample is an unrecognizable sample obtained by adding noise on the basis of the first type of attack sample of the known attack type. It can be considered as an attack sample whose attack type cannot be determined under current technical means, that is, an attack sample of an unknown attack type. In this way, by adding the attack sample to the intrusion detection sample set, the phenomenon of misreporting these attack samples of unknown attack types as non-attack samples can be avoided in real attack test scenarios.

[0024] In one possible design, after constructing an intrusion detection sample set based on the first and second type attack samples, the data processing device may further extract features from the intrusion detection sample set to obtain an offline detection sample set. This offline detection sample set is then used to input into an intrusion detection model to evaluate the detection performance of the intrusion detection model based on the detection results output by the intrusion detection model. For example, if the intrusion detection model detects more attack samples in the offline detection sample set as attack samples, while the intrusion detection model detects more non-attack samples as non-attack samples, the intrusion detection model has a good detection performance. Conversely, if the intrusion detection model detects more attack samples in the offline detection sample set as non-attack samples, while the intrusion detection model detects more non-attack samples as attack samples, the intrusion detection model has a poor detection performance. Furthermore, if the detection performance is poor, the data processing device may adjust the parameters of the intrusion detection model and use the adjusted intrusion detection model to retest the offline detection sample set, repeating the above process until an intrusion detection model with good detection performance is obtained.

[0025] In the above design, the intrusion detection sample set can support the offline evaluation of the detection effect of the intrusion detection model (referred to as offline evaluation), so as to continuously optimize the intrusion detection model according to the detection effect, obtain an intrusion detection model with better detection effect, and improve the implementation effect of the intrusion detection model on the device to be detected.

[0026] In a further design, the data processing device performs feature extraction on the intrusion detection sample set to obtain an offline detection sample set, including: the data processing device first determines the message type of each intrusion detection sample in the intrusion detection sample set, and then, for each intrusion detection sample whose message type is a transmission control protocol (TCP) message, the data processing device obtains an offline detection sample by performing feature extraction on all intrusion detection samples belonging to the same TCP connection; and for each intrusion detection sample whose message type is a CAN (CAN with flexible data rate, CAN (FD)) message or a UDP message with a variable data rate, the data processing device obtains an offline detection sample by performing feature extraction on the intrusion detection sample of each CAN (FD) message or UDP message.

[0027] In the above design, by performing joint feature extraction on each intrusion detection sample belonging to the same TCP connection, all intrusion detection samples collected on a TCP connection can be aggregated into a single offline detection sample, effectively reducing the number of offline detection samples while making the features contained in each offline detection sample more relevant. In addition, by performing separate feature analysis on intrusion detection samples that do not belong to a TCP connection (such as intrusion detection samples belonging to CAN (FD) messages and UDP messages), each unrelated intrusion detection sample is not missed, ensuring that the offline detection sample set covers all intrusion detection samples.

[0028] In further designs, the features extracted by offline evaluation can include one or more of the following features: timestamp, frequency characteristics, protocol type, content characteristics, packet loss rate, number of error packets, connection duration, connection initiator, and connection receiver. For example, for intrusion detection samples belonging to TCP messages, the extracted features can include all the features shown here, while for intrusion detection samples belonging to CAN (FD) messages or UDP messages, the extracted features can include timestamp, frequency characteristics, protocol type, content characteristics, packet loss rate, and number of error packets.

[0029] In the above design, by providing a variety of extractable features corresponding to offline evaluation, users can easily select one or more features according to the application layer requirements in the actual offline scenario to convert the intrusion detection sample set into the offline detection sample set, so as to be suitable for different offline evaluation scenarios and improve the versatility of the intrusion detection sample set in the offline evaluation field.

[0030] In a further design, in order to maintain the format consistency of each offline detection sample, the data processing device can extract all the aforementioned features for each intrusion detection sample. Then, when a certain feature of a certain intrusion detection sample does not exist, the feature of the intrusion detection sample set is configured as a preset character. The preset character can be, for example, a number, a letter, a symbol, or a combination of one or more of them.

[0031] In one possible design, after the data processing device constructs an intrusion detection sample set based on the first type of attack samples and the second type of attack samples, it can also convert the format of the intrusion detection sample set to obtain an online detection sample set that matches the format of the test tool, and then input the online detection sample set into the device to be tested through the test tool. The online detection sample set is used to evaluate the detection performance of the device to be tested that is deployed with an intrusion detection model.

[0032] In the above design, the intrusion detection sample set can support the online evaluation of the detection effect of the device to be tested that is deployed with the intrusion detection model (referred to as online evaluation), so as to determine the anti-attack performance of the device to be tested based on the detection effect, and ensure that only the device to be tested with good anti-attack effect is shipped out of the factory.

[0033] In further design, the test tool can be CANoe, PCAN, Technica or other tools that can realize online evaluation, and the intrusion detection sample set after format conversion can be .PCAP, .ASC, .BLF or other formats corresponding to online test tools.

[0034] In the above design, by enabling intrusion detection samples to be converted into formats corresponding to multiple testing tools, users can easily select the corresponding format for conversion based on the testing tools in the actual online scenario, so as to adapt to different online evaluation scenarios and improve the versatility of the intrusion detection sample set in the online evaluation field.

[0035] In one possible design, after the data processing device constructs an intrusion detection sample set based on the first type of attack samples and the second type of attack samples, it can also determine the evaluation value corresponding to the intrusion detection sample set based on the values ​​of the intrusion detection sample set under various preset indicators. When the evaluation value is lower than the preset threshold, the intrusion detection sample set is adjusted.

[0036] In the above design, the preset indicators can be set according to the characteristics of the system architecture to which the device to be detected belongs. By evaluating the constructed intrusion detection sample set with reference to the preset indicators, the intrusion detection sample set can be effectively optimized according to the evaluation results, so that the optimized intrusion detection sample set is more suitable for the system architecture to which the device to be detected belongs.

[0037] In further design, the preset indicators may include one or more of the following indicators: data redundancy indicator, attack coverage indicator, protocol coverage indicator, service coverage indicator, data labeling indicator, balance indicator, feature independence indicator, and ease of use indicator. Among them, data redundancy indicator, attack coverage indicator, protocol coverage indicator, service coverage indicator, data labeling indicator, and balance indicator are quantitative indicators, while feature independence indicator and ease of use indicator are qualitative indicators. By combining quantitative and qualitative indicators to comprehensively evaluate the construction quality of the intrusion detection sample set, the evaluation results can be made more comprehensive and more convincing.

[0038] In further design, in order to make the measurements of various preset indicators comparable, the value ranges of various preset indicators can be configured to be the same range, such as [0,1].

[0039] In further design, the evaluation value corresponding to the intrusion detection sample set can be the weighted average of the values ​​of the intrusion detection sample set under each preset indicator, wherein the weights corresponding to each preset indicator can be the same or different. For example, in a specific example, each quantitative preset indicator can be configured to correspond to a first weight, and each qualitative preset indicator can correspond to a second weight, and the first weight is greater than the second weight.

[0040] In the above design, by reducing the weight of qualitative preset indicators, the impact of preset indicators that require human experience judgment on the evaluation value can be reduced, so that the evaluation process pays more attention to rational evaluation criteria, while not giving up the evaluation criteria that require human experience.

[0041] In one possible design, the first type of attack sample can be obtained by attacking any of the following areas: the entire device to be tested; one or more physical areas of the device to be tested; or one or more functional areas of the device to be tested. For example, when the device to be tested is a vehicle, the one or more physical areas may include one or more of the vehicle trunk area, left front body area, right front body area, left rear body area, or right rear body area. The one or more functional areas may include one or more of the vehicle central control gateway area, body control area, cockpit control area, power control area, chassis control area, or infotainment area.

[0042] In the above design, the data processing device can construct an intrusion detection sample set for the entire device to be detected, or construct an intrusion detection sample set for one or more physical areas in the device to be detected, or construct an intrusion detection sample set for one or more functional areas in the device to be detected, or construct an intrusion detection sample set for a combination of one or more physical areas and one or more functional areas. It can be seen that the method for constructing the intrusion detection sample set can be applied to various different construction scenarios, which helps to improve the flexibility, versatility and ease of use of the intrusion detection sample set.

[0043] In a second aspect, the present application provides a data processing device, which can be any device with processing capabilities, such as a server or a server cluster composed of servers. The data processing device includes: an attack unit for obtaining a first type of attack sample by attacking a device to be detected; a perturbation unit for obtaining a second type of attack sample by applying noise to the first type of attack sample; and a construction unit for constructing an intrusion detection sample set based on the first type of attack sample and the second type of attack sample.

[0044] In a possible design, the first type of attack samples may include real attack samples and simulated attack samples. The real attack samples are obtained by manually attacking the device to be detected, and the simulated attack samples are obtained by attacking the device to be detected with an attack tool.

[0045] In a possible design, the real attack samples may correspond to one or more of the following attack types: ID non-existence attack, replay attack, tampering attack, data length error attack, signal out-of-defined-range attack, context error attack, ID source non-specified ECU attack, identical ID attack, CAN scan attack, UDS execution sensitive operation attack, message authentication error attack, ECU identity spoofing attack, man-in-the-middle attack, ECU authentication error attack, brute force attack, application layer protocol error attack, unknown stack connection attack, and unknown stack connection attack.

[0046] In a possible design, the simulated attack samples may correspond to one or more of the following attack types: ID Fuzz attack, data Fuzz attack, CAN DoS attack, ETH DoS attack, malformed packet injection attack, and port scanning attack.

[0047] In one possible design, the attack unit is specifically used to: traverse each attack type among a plurality of preset attack types, and when traversing each attack type: execute the attack behavior corresponding to the attack type on the device to be detected, and obtain the traffic data generated by the device to be detected for the attack behavior; if the traffic data is attack traffic, mark the traffic data as a first type attack sample.

[0048] In one possible design, after obtaining the traffic data generated by the device to be detected in response to the attack behavior, the attack unit is also used to: if the traffic data is normal traffic, mark the traffic data as a non-attack sample; correspondingly, the construction unit is specifically used to: construct an intrusion detection sample set based on the first type of attack samples, the second type of attack samples and the non-attack samples.

[0049] In one possible design, the perturbation unit is specifically configured to: apply noise to a first-type attack sample to obtain a perturbed sample; input the perturbed sample into an attack recognition model; obtain a recognition result output by the attack recognition model; and adjust the perturbed sample based on the recognition result until the recognition result corresponding to the adjusted perturbation sample indicates that it is impossible to determine whether the perturbed sample is an attack sample; and then determine the adjusted perturbation sample as a second-type attack sample. The recognition result output by the attack recognition model indicates whether the perturbed sample is an attack sample.

[0050] In one possible design, the data processing device may further include a feature extraction unit, which is used to: extract features from the intrusion detection sample set to obtain an offline detection sample set, and the offline detection sample set is used to evaluate the detection effect of the intrusion detection model.

[0051] In one possible design, the feature extraction unit is specifically used to: determine the message type of each intrusion detection sample in the intrusion detection sample set; for each intrusion detection sample whose message type is a TCP message, an offline detection sample is obtained by performing feature extraction on all intrusion detection samples belonging to the same TCP connection; and for each intrusion detection sample whose message type is a CAN (FD) message or a UDP message, an offline detection sample is obtained by performing feature extraction on the intrusion detection sample of each CAN (FD) message or UDP message.

[0052] In one possible design, the features extracted from the aforementioned features include one or more of the following features: timestamp, frequency feature, protocol type, content feature, packet loss rate, number of error packets, connection duration, connection initiator, and connection receiver.

[0053] In one possible design, the data processing device may also include a format conversion unit, which is used to: convert the format of the intrusion detection sample set to obtain an online detection sample set that matches the format of the test tool, and input the online detection sample set into the device to be detected through the test tool, wherein the online detection sample set is used to evaluate the detection performance of the device to be detected on which the intrusion detection model is deployed.

[0054] In one possible design, the data processing device may further include an adjustment unit, which is used to: determine an evaluation value corresponding to the intrusion detection sample set based on the value of the intrusion detection sample set under each preset indicator, and adjust the intrusion detection sample set when the evaluation value is lower than a preset threshold.

[0055] In one possible design, the preset indicators include one or more of the following indicators: data redundancy indicator, attack coverage indicator, protocol coverage indicator, business coverage indicator, data labeling indicator, balance indicator, feature independence indicator, and ease of use indicator.

[0056] In one possible design, the first type of attack sample can be obtained by attacking any of the following areas: the entire device to be detected; one or more physical areas of the device to be detected; or one or more functional areas of the device to be detected.

[0057] In a third aspect, the present application provides a data processing device, comprising a processor, the processor being connected to a memory, the memory being used to store computer programs, and the processor being used to execute the computer programs stored in the memory, so that the data processing device executes the method described in any one of the designs of the first aspect above.

[0058] In a fourth aspect, the present application provides a data processing device, comprising a processor and a memory, wherein the memory stores computer program instructions, and the processor executes the computer program instructions to implement the method as described in any one of the designs of the first aspect above.

[0059] In a fifth aspect, the present application provides a data processing device comprising a processor, a memory and a transceiver, wherein the memory stores computer program instructions, and the processor runs the computer program instructions to call the transceiver to implement the method described in any one of the designs in the first aspect above.

[0060] In a sixth aspect, the present application provides a chip, which may include a processor and an interface, wherein the processor is used to read instructions through the interface to execute the method described in any one of the designs in the first aspect above.

[0061] In a seventh aspect, the present application provides a data processing system, which may include a device to be detected and a data processing device, and the data processing device is used to execute the method described in any one of the designs in the first aspect above.

[0062] In an eighth aspect, the present application provides a computer-readable storage medium storing a computer program. When the computer program is executed, the method described in any one of the above-mentioned first aspects is implemented.

[0063] In a ninth aspect, the present application provides a computer program product, which, when executed on a processor, implements the method as described in any one of the above-mentioned first aspects.

[0064] For the beneficial effects of the second to ninth aspects mentioned above, please refer to the technical effects that can be achieved by the corresponding designs in the first aspect mentioned above, and no further details will be given here. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] FIG1 exemplarily shows a possible system architecture diagram provided by an embodiment of the present application;

[0066] FIG2 exemplarily shows a schematic diagram of a vehicle partition architecture provided by an embodiment of the present application;

[0067] FIG3 exemplarily shows a flow chart of a data processing method provided in an embodiment of the present application;

[0068] FIG4 exemplarily shows a schematic diagram of a process for obtaining a first type of attack sample provided in an embodiment of the present application;

[0069] FIG5 exemplarily shows a schematic diagram of an application scenario of an intrusion detection sample set provided by an embodiment of the present application;

[0070] FIG6 exemplarily shows a flow chart of evaluating an intrusion detection sample set provided by an embodiment of the present application;

[0071] FIG7 exemplarily illustrates a design architecture diagram of a development data processing solution provided in an embodiment of the present application;

[0072] FIG8 exemplarily shows a structural diagram of a data processing device provided in an embodiment of the present application;

[0073] FIG9 exemplarily shows a structural diagram of another data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0074] The data processing method disclosed in this application can be used to construct an intrusion detection sample set, which can be used to train an intrusion detection model, or to detect the detection effect of an intrusion detection model or a device to be detected that is deployed with an intrusion detection model. Among them, the device to be detected can be any terminal device with communication capabilities, and in particular, it can be a terminal device that has certain requirements for network security. In some examples, the terminal device can include but is not limited to: intelligent transportation equipment, such as cars, ships, drones, trains, vans, trucks, flying cars, etc.; smart home devices, such as TVs, sweeping robots, smart desk lamps, audio systems, smart lighting systems, electrical control systems, home background music, home theater systems, intercom systems, video surveillance, etc.; intelligent manufacturing equipment, such as robots, industrial equipment, industrial computers, intelligent logistics, smart factories, etc. Alternatively, the terminal device can also be a computer device, such as a desktop computer, a personal computer, a server, etc. It should also be understood that the terminal device can also be a portable electronic device, such as a mobile phone, a tablet computer, a PDA, headphones, speakers, wearable devices (such as smart watches), vehicle-mounted devices, virtual reality devices, augmented reality devices, etc. Examples of portable electronic devices include but are not limited to those equipped with Or a portable electronic device with other operating systems. The portable electronic device may also be a laptop computer (Laptop) with a touch-sensitive surface (eg, a touch panel).

[0075] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.

[0076] It should be noted that the terms "system" and "network" in the embodiments of the present application can be used interchangeably. "At least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0077] Furthermore, unless otherwise specified, ordinal numbers such as "first" and "second" in the embodiments of this application are used to distinguish between multiple objects and are not used to define the priority or importance of multiple objects. For example, the terms "first type attack sample" and "second type attack sample" are used only to distinguish between different types of attack samples and do not indicate a difference in priority or importance between the two attack samples.

[0078] In addition, in the embodiments of the present application, "connection" can be understood as electrical connection, and the connection between two electrical components can be a direct or indirect connection between the two electrical components. For example, the connection between A and B can be either a direct connection between A and B, or an indirect connection between A and B through one or more other electrical components, such as A and B being connected, or A and C being directly connected, C being directly connected to B, and A and B being connected through C. In some scenarios, "connection" can also be understood as coupling, such as electromagnetic coupling between two inductors. In short, the connection between A and B enables electrical energy to be transmitted between A and B.

[0079] FIG1 exemplarily illustrates a possible system architecture diagram provided by an embodiment of the present application. As shown in FIG1 , the system architecture includes a device to be detected 100 and a data processing device 200. Among them, the device to be detected 100 can be a device with certain requirements for network security, such as a vehicle. The data processing device 200 can be any device with data processing capabilities, such as a server or a server cluster composed of multiple servers, or it can also be a chip or circuit, such as a chip or circuit arranged in a server or a server cluster. In a specific example, the data processing device 200 can be a cloud server, and the cloud server can be connected to the device to be detected 100 by wireless. It should be understood that the embodiment of the present application does not limit the number of devices to be detected 100 and the number of data processing devices 200 in the system architecture. For example, a data processing device 200 can be connected to only one device to be detected 100 as shown in FIG1 , or it can be connected to multiple devices to be detected 100 at the same time. In addition, the data processing device 200 in the embodiment of the present application can integrate all functions on an independent physical device, or it can distribute the functions on multiple independent physical devices, and this embodiment of the present application is not specifically limited.

[0080] As a further example, and with continued reference to FIG1 , the system architecture may further include one or more of a database 300, a model training device 400, and a test tool 500, or may further include other devices, such as routing devices, wireless relay devices, wireless backhaul devices, and operation, management, and maintenance equipment. The database 300 may be used to store the intrusion detection sample set constructed by the data processing device 200. The database 300 may be a storage unit independent of the data processing device 200, such as a database server, or may be an internal storage unit of the data processing device 200, such as a cache memory, random access memory, register, main memory, or read-only memory. The model training device 400 may be any device with algorithm development and verification capabilities, such as a model training server. The model training device 400 is deployed with an intrusion detection system (IDS), which is a test system released by the automotive open system architecture (AUTOSAR) for automotive network security. It is one of the most mainstream vehicle-mounted test systems at present. It can train a high-accuracy, high-timeliness and high-robustness intrusion detection model based on limited vehicle-mounted hardware resources and vehicle-mounted storage resources in a combination of rules and machine learning. The test tool 500 refers to a tool that can inject test traffic into the device to be tested 100, such as a CAN test tool, a TCP test tool or a UDP test tool. In addition, the test tool 500 can present a user interface to the outside world. By clicking on the corresponding test command (such as traffic injection type, traffic injection quantity and traffic injection frequency, etc.) on the user interface, the user can drive the test tool 500 to automatically inject test traffic into the device to be tested 100 according to the corresponding test command.

[0081] Further illustratively, when the system includes a device to be detected 100, a data processing device 200, a database 300, a model training device 400, and a testing tool 500, the data processing device 200 can be connected to the device to be detected 100, the database 300, the model training device 400, and the testing tool 500, respectively, and the device to be detected 100 can also be connected to the model training device 400 and the testing tool 500. In implementation, the data processing device 200 can construct an intrusion detection sample set for the device to be detected 100 and store the intrusion detection sample set in the database 300. Thereafter, when in use, the data processing device 200 can convert the intrusion detection sample set in the database 300 into an offline detection sample set, and can select a portion of the offline detection samples and send them to the model training device 400. After the model training device 400 uses the portion of offline detection samples to train an intrusion detection model, the intrusion detection model is deployed in the device to be detected 100, so that the device to be detected 100 can use the intrusion detection model to identify attack traffic. Furthermore, when offline evaluation is required, the data processing device 200 can also send another part of the converted offline detection samples to the model training device 400, and the model training device 400 uses the previously trained intrusion detection model to detect this part of the offline detection samples to obtain offline evaluation information. The data processing device 200 evaluates the detection effect of the intrusion detection model trained by the model training device 400 based on the offline evaluation information. Correspondingly, when online evaluation is required, the data processing device 200 can convert the intrusion detection sample set in the database 300 into an online detection sample set, and then input part or all of the online detection samples into the device to be detected 100 through an online tool, and obtain the online evaluation information generated by the device to be detected 100 for the online detection samples, and evaluate the detection effect of the device to be detected 100 deployed with the intrusion detection model based on the online evaluation information.

[0082] Based on the above, we can see that the intrusion detection sample set is not only used to train the intrusion detection model, but also serves as a test sample to test the quality of the intrusion detection model and the quality of the equipment to be detected deployed with the intrusion detection model. When the sample types in the intrusion detection sample set are sufficient, the detection effect of the intrusion detection model trained based on sufficient intrusion detection samples will be better, and the anti-attack performance of the equipment to be detected deployed with the intrusion detection model will also be better. Correspondingly, the reliability of the evaluation results obtained by evaluating the intrusion detection model or the equipment to be detected based on sufficient intrusion detection samples will also be better. Therefore, how to construct an intrusion detection sample set with sufficient sample types is crucial to improving the detection effect of the intrusion detection model, improving the anti-attack effect of the equipment to be detected, improving the reliability of the intrusion detection model evaluation, and improving the reliability of the equipment to be detected.

[0083] However, when constructing intrusion detection sample sets, the industry only uses attack tools to simulate some typical intrusion scenarios to attack the devices to be detected, resulting in the intrusion detection sample set only containing intrusion detection samples corresponding to typical intrusion scenarios. The sample types in the intrusion detection sample set are very limited, which is not conducive to improving the detection effect of the intrusion detection model, improving the anti-attack effect of the devices to be detected, improving the reliability of the intrusion detection model, and improving the reliability of the devices to be detected, which is also not conducive to the implementation of intrusion detection technology on the devices to be detected.

[0084] In view of this, an embodiment of the present application provides a data processing method for constructing an intrusion detection sample set with richer sample types, so that it can cover both intrusion detection samples in known intrusion scenarios and intrusion detection samples in unknown intrusion scenarios, so as to improve the detection effect of the intrusion detection model trained using the intrusion detection sample set, improve the anti-attack effect of the device to be detected deployed with the intrusion detection model, improve the reliability of evaluating the intrusion detection model using the intrusion detection sample set, and improve the reliability of evaluating the device to be detected using the intrusion detection sample set.

[0085] It should be noted that the intrusion detection sample set in the embodiments of the present application can be constructed for the entire device to be detected, or for one or more areas in the device to be detected. For example, when the device to be detected is a vehicle, Figure 2 shows a schematic diagram of a vehicle partition architecture provided by the embodiments of the present application, wherein:

[0086] Figure 2 (A) shows the physical partition architecture of the vehicle. This architecture divides the entire vehicle into the vehicle trunk area, the right front body area, the left front body area, the right rear body area, and the left rear body area according to different physical areas. The right front body area, the left front body area, the right rear body area, and the right rear body area exist as branch areas of the vehicle trunk area. Among them, the vehicle control unit (VCU) is deployed in the vehicle trunk area, and each branch area is deployed with its own control node (i.e., Z1, Z2, Z3, Z4) and several ECUs connected to the control node. All ECUs in each branch area are connected to the control node via CAN (FD), and the control node in each branch area is connected to the VCU in the vehicle trunk area via ETH.

[0087] Figure 2 (B) shows the functional partition architecture of the vehicle. This architecture divides the entire vehicle into the vehicle-mounted central control gateway functional area and K other functional areas according to the different functions to be implemented. The K other functional areas exist as branch areas of the vehicle-mounted central control gateway functional area, each of which implements different functions. For example, in some scenarios, the K other functional areas may include one or more of the body control domain, cockpit control domain, power control domain, chassis control domain or infotainment domain, etc. K is a positive integer. Among them, the GateWay gateway is deployed in the vehicle-mounted central control gateway functional area, and each other functional area is deployed with its own domain controller (i.e., D1, D2, ..., D 4K ) and several ECUs connected to the domain controller, all ECUs in each other functional area are connected to the domain controller through CAN (FD), and the domain controller in each other functional area is connected to the GateWay in the vehicle central control gateway functional area through ETH.

[0088] Based on the vehicle partition architecture shown in FIG2 (A) or FIG2 (B), in the embodiment of the present application, an intrusion detection sample set can be constructed for the entire vehicle, an intrusion detection sample set can be constructed for one or more physical areas in the vehicle, an intrusion detection sample set can be constructed for one or more functional areas in the vehicle, or an intrusion detection sample set can be constructed for a combination of one or more physical areas and one or more functional areas, etc. Among them, the physical area can be the vehicle trunk area, the right front body area, the left front body area, the right rear body area or the left rear body area, and the functional area can be the vehicle central control gateway functional area or other functional areas, without specific limitation.

[0089] It should be noted that the specific areas for constructing intrusion detection sample sets can be determined based on the processing capabilities of the data processing device and the actual needs of the user. For example, when the data processing device has strong processing capabilities, an intrusion detection sample set can be constructed for the entire vehicle, leveraging efficient processing capabilities to construct a relatively complete global intrusion detection sample set. This global intrusion detection sample set can be applied to both global and local attack scenarios, providing good versatility. Conversely, when the data processing device has weak processing capabilities, local intrusion detection sample sets can be constructed by specifically selecting areas involved in actual business operations. This local intrusion detection sample set is a subset of the global intrusion detection sample set. Although it is only applicable to local attack scenarios, the area it needs to process is smaller, thus significantly reducing the difficulty and complexity of constructing the intrusion detection sample set.

[0090] The following uses the construction of an intrusion detection model corresponding to the entire device to be detected as an example to describe how the above technical problems are solved through specific embodiments. It should be noted that in the following description, the data processing device can be the data processing device 200 shown in Figure 1, or it can be other communication nodes, communication devices, or communication systems that can support the data processing device to achieve the required functions, such as chips, chip systems, circuits, or circuit systems, without specific limitation.

[0091] Based on the system architecture shown in FIG1 , FIG3 exemplarily illustrates a flow chart of a data processing method provided in an embodiment of the present application. The data processing method can be executed by a data processing device, such as the data processing device 200 shown in FIG1 . As shown in FIG3 , the method includes:

[0092] In step 301, the data processing device obtains a first type of attack sample by attacking a device to be detected.

[0093] When a global intrusion detection sample set is constructed for the entire device to be detected, the data processing device may initiate an attack against the entire device to be detected to obtain a first-type attack sample corresponding to the entire device to be detected. Conversely, when a local intrusion detection sample set is constructed for one or more regions, the data processing device may initiate an attack only against the one or more regions to obtain a first-type attack sample corresponding to the one or more regions.

[0094] In one example, the device to be detected is a vehicle, which can be a vehicle with a distributed electrical and electronic architecture, a domain-centralized electrical and electronic architecture, a vehicle-centralized electrical and electronic architecture, or any other vehicle architecture that may emerge in the future. Thus, by attacking vehicles belonging to each vehicle architecture, the data processing device can construct an intrusion detection sample set corresponding to each vehicle architecture. This allows users to select one or more intrusion detection sample sets corresponding to each vehicle architecture based on their actual needs, thereby improving their user experience.

[0095] In a further example, in order to improve the accuracy of the intrusion detection sample set corresponding to any vehicle-mounted architecture, the data processing device can also attack multiple vehicles belonging to the vehicle-mounted architecture together, so that the first type of attack sample can cover multiple vehicles under the vehicle-mounted architecture, avoiding the problem of inaccurate sample collection due to problems with the vehicle when only one vehicle is attacked.

[0096] In a further example, when attacking multiple vehicles belonging to a single vehicle-mounted architecture, after acquiring a large number of first-type attack samples corresponding to the multiple vehicles, the data processing device can further filter these first-type attack samples through clustering or model recognition, for example, eliminating first-type attack samples that differ significantly from other first-type attack samples and retaining only relatively similar first-type attack samples. This preemptive removal of clearly problematic first-type attack samples can avoid subsequent pointless processing of these problematic first-type attack samples, effectively conserving computing resources for the data processing device.

[0097] In an embodiment of the present application, the first type of attack samples may include real attack samples and simulated attack samples. Real attack samples are obtained by manually attacking the device to be tested, for example, they may be collected during a penetration test. A penetration test involves an attacker actually attacking a vehicle from a hacker's perspective with the goal of breaking into the device to be tested. This type of attack can construct highly concealed attack behaviors and attack behaviors that rely on business logic, and can obtain attack samples that are easily labeled in a real attack environment. Correspondingly, simulated attack samples are obtained by attacking the device to be tested using an attack tool, for example, they may be collected from the device to be tested after simulating the generation of attack traffic using an attack tool and automatically injecting the attack traffic into the device to be tested. This type of attack can obtain attack samples that are not easily labeled in a real attack environment. Thus, by combining human attack methods and attack tool attack methods to comprehensively construct the first type of attack samples, the first type of attack samples can fully cover attack samples of various attack types that may exist in a real attack environment, thereby improving the comprehensiveness of the first type of attack samples.

[0098] In an optional embodiment, the first type of attack samples can be obtained by attacking the device to be detected according to various known attack types in currently known intrusion scenarios. The various known attack types include attack types that are easy to collect and annotate in a real attack environment (referred to as the first attack type) and attack types that are difficult to collect and annotate in a real attack environment (referred to as the second attack type). Therefore, by manually attacking the device to be detected according to the first attack type, the above-mentioned real attack samples can be collected. By using attack tools to simulate attacks on the device to be detected according to the second attack type, the above-mentioned simulated attack samples can be collected in a centralized manner.

[0099] For example, assuming that the device to be detected is a vehicle, and the various ECUs in the vehicle communicate through the CAN (FD) bus and the ETH bus, please refer to Table 1, which shows the corresponding relationship table of possible attack types in the device to be detected and the attack type corresponding to each attack sample.

[0100] Table 1

[0101] As shown in Table 1, the first attack type may include one or more of the following attack types:

[0102] ID non-existent attack refers to attacking the vehicle by changing the ID in the CAN message to a non-existent ID;

[0103] A replay attack is a method of deceiving a vehicle by sending CAN messages that it has previously received.

[0104] Tampering attack refers to attacking the vehicle by tampering with the data carried in the CAN message;

[0105] Data length error attack refers to attacking the vehicle by modifying the data length of the CAN message;

[0106] Signal out-of-range attack refers to attacking the vehicle by modifying the signal value in the CAN message to a value greater than the maximum specified value or less than the minimum specified value;

[0107] Context error attack refers to attacking the vehicle by publishing specific messages or signals on the CAN network that are not suitable for a certain state, such as sending an acceleration signal when the vehicle is braking;

[0108] An attack with a non-specified ECU as the source of the ID is an attack on a vehicle by using a non-specified ECU to publish CAN messages that should have been published by the designated ECU.

[0109] The same ID attack occurs when two ECUs send the same ID in CAN messages to attack the vehicle.

[0110] CAN scanning attack refers to invading the vehicle by sending CAN scanning messages;

[0111] UDS performs sensitive operation attacks, which means attacking vehicles by carrying sensitive information in UDS messages;

[0112] Message authentication error attack refers to modifying the message authentication process of the CAN bus so that the vehicle cannot successfully authenticate the message;

[0113] ECU identity spoofing attack refers to deceiving the vehicle by changing the ECU identity in the ETH message;

[0114] A man-in-the-middle attack is a method of attacking a vehicle by virtually placing a device controlled by an intruder between two ECUs connected via ETH.

[0115] ECU authentication error attack refers to modifying the ETH ECU authentication process to prevent the vehicle from successfully ECU authentication;

[0116] Brute force attack refers to attacking the vehicle by deciphering the ETH key in an exhaustive manner;

[0117] Application layer protocol error attacks attack vehicles by modifying the application layer protocol so that the vehicle cannot obtain the correct application layer protocol for message exchange.

[0118] Unknown outbound connection attack refers to attacking the vehicle by giving a fake outbound connection;

[0119] Unknown push connection attack refers to attacking a vehicle by giving a false push connection.

[0120] Among the first type of attacks mentioned above, ID non-existence attack, replay attack, tampering attack, data length error attack, signal out-of-defined range attack, context error attack, ID source non-specified ECU attack, identical ID attack, CAN scan attack, UDS execution sensitive operation attack and message authentication error attack belong to the attack types that exist under the CAN (FD) communication mode, while ECU identity spoofing attack, man-in-the-middle attack, ECU authentication error attack, brute force attack, application layer protocol error attack, unknown stack connection attack and unknown stack connection attack belong to the attack types that exist under the ETH communication mode.

[0121] Continuing with Table 1, the second attack type may include one or more of the following attack types:

[0122] ID fuzzing attack refers to attacking vehicles by obfuscating the ID in CAN messages;

[0123] Data fuzz attack refers to attacking vehicles by obfuscating the data in CAN messages;

[0124] A CAN DoS attack is an attack on a vehicle by stopping the sending and receiving services of a certain part of the CAN bus.

[0125] The DoS attack on ETH refers to attacking the vehicle by stopping the sending and receiving services of a certain part of the ETH bus;

[0126] Malformed packet injection attacks attack vehicles by injecting malformed packets;

[0127] A port scan attack is an attempt to intrude into a vehicle by sending port scan messages.

[0128] Among the second attack types mentioned above, ID fuzzing (Fuzz) attacks, data fuzzing attacks, and CAN DoS attacks belong to the attack types that exist in the CAN (FD) communication mode, while ETH DoS attacks, malformed packet injection attacks, and port scanning attacks belong to the attack types that exist in the ETH communication mode.

[0129] In the above embodiment, by providing various attack types that may exist in known intrusion scenarios, it is convenient for users to select one or more attack types according to actual needs to construct first-type attack samples, and real attacks or simulated attacks can also be used to obtain the required real attack samples or simulated attack samples to adapt to different attack scenarios, thereby improving the flexibility and versatility of constructing first-type attack samples.

[0130] In a further example, the data processing device may attack the device to be detected in a variety of ways to obtain the first type of attack sample. For example, referring to FIG4 , which shows a schematic diagram of a specific process for obtaining the first type of attack sample provided in an embodiment of the present application, the process includes:

[0131] In step 401, the data processing device obtains multiple preset attack types.

[0132] The preset multiple attack types may illustratively include all attack types shown in Table 1 above, so that the first type attack samples can fully cover various known attack types, thereby improving the richness and comprehensiveness of the first type attack samples.

[0133] In step 402 , the data processing device determines whether there is an attack type that has not been traversed among the preset multiple attack types. If so, step 403 is executed; if not, step 409 is executed.

[0134] In step 403 , the data processing device executes an attack behavior corresponding to an untraversed attack type on the device to be detected, and obtains traffic data generated by the device to be detected in response to the attack behavior.

[0135] For example, for the first attack type, the required attack code can be manually pre-written and mapped to the corresponding first attack type and stored in the data processing device. In this way, after the data processing device obtains an attack type that has not been traversed, if it determines that the attack type belongs to the first attack type, it can directly obtain the attack code corresponding to the first attack type from the local computer, automatically generate the corresponding attack behavior according to the attack code, and then attack the device to be detected. Conversely, if it determines that the attack type belongs to the second attack type, the data processing device can invoke an attack tool, use the attack tool to generate the attack behavior corresponding to the second attack type, and then attack the device to be detected. In this way, the entire attack process can be automatically implemented by the data processing device, which helps improve the unified management of the entire attack process and eliminates the need for on-site manual programming, helping to reduce sample collection delays.

[0136] In a further example, when the device to be detected is a vehicle, the data processing device can access the on-board diagnostics (OBD) interface of the device to be detected through a data cable. The OBD interface can be, for example, an interface of a Type II on-board diagnostic system, i.e., an OBD-II interface. Initially, since all attack types have not been traversed, the data processing device can select an attack type from all attack types in a random manner, a sequential manner, or other manner, and then attack the device to be detected according to the attack type, and obtain the traffic data generated by the device to be detected for the attack type through the OBD interface. Furthermore, after analyzing the traffic data, the data processing device can select an attack type from the attack types that have not been traversed to continue attacking the device to be detected, and repeat the above process until all attack types have been traversed.

[0137] As a further example, the traffic collection operation for the entire attack process described above can be automatically performed by a data processing device. Specifically, a collection duration can be configured in the data processing device. After the attack begins, the data processing device can start a timer. During the timer, the data processing device continuously collects traffic data output from the OBD interface of the device under test until the configured collection duration is reached, at which point the collection ends. The collection duration can range from several hours, days, weeks, or even months, and can be configured by those skilled in the art based on actual needs. For example, when the number of pre-set attack types is large, the time required to attack the device under test is longer, and the collection duration can be configured to be longer, such as several weeks. Conversely, when the number of pre-set attack types is small, the time required to attack the device under test is shorter, and the collection duration can be configured to be shorter, such as several days.

[0138] In step 404 , the data processing device determines whether the traffic data is attack traffic. If so, step 405 is executed; if not, step 406 is executed.

[0139] In step 405 , the data processing device marks the traffic data as a first type attack sample, and then executes step 402 .

[0140] Exemplarily, after determining that certain traffic data meets the characteristics of attack traffic, if the traffic data is collected under the first attack type, the data processing device can mark the attack traffic as a real attack sample; conversely, if the traffic data is collected under the second attack type, the data processing device can mark the attack traffic as a simulated attack sample.

[0141] In step 406 , the data processing device determines whether the traffic data is normal traffic. If so, step 407 is executed; if not, step 408 is executed.

[0142] In step 407 , the data processing device marks the traffic data as a non-attack sample, and then executes step 402 .

[0143] In some scenarios, non-attack samples are also called context data or context samples.

[0144] In step 408 , the data processing device determines that the traffic data is an unlabeled sample, and then executes step 402 .

[0145] For example, when traffic data meets neither the characteristics of attack traffic nor the characteristics of normal traffic, it means that the traffic data cannot be identified as a sample type using current technical means. In this case, the data processing device may not mark the traffic data, but instead treat it as an unmarked sample. In some scenarios, this unmarked sample is an abnormal sample. For example, when the attack test of the data processing device causes the hardware and software system of the device under test to malfunction, the device under test may generate some abnormal data. This abnormal data does not meet the characteristics of normal traffic nor attack traffic, but it will still be collected by the data processing device.

[0146] Step 409: The data processing device ends the attack process.

[0147] In the above example, the data processing device can obtain attack samples, such as real attack samples and simulated attack samples, as well as non-attack samples and even unlabeled samples by attacking the device to be detected. This acquisition method can obtain multiple types of samples, which is convenient for improving the sample richness of the subsequent construction of the intrusion detection sample set.

[0148] It should be noted that FIG4 is merely an illustrative example of one possible method for obtaining a first-type attack sample, and the embodiments of the present application are not limited to using this method to obtain a first-type attack sample. For example, in another possible acquisition method, the data processing device may combine at least two of the preset multiple attack types, execute the attack behaviors corresponding to the at least two attack types on the device to be detected at one time, and then, after obtaining the traffic data generated by the device to be detected, separate the traffic data corresponding to the at least two attack types from the traffic data, and obtain the first-type attack sample based on the at least two traffic data. It should be understood that there are many possible acquisition methods, which will not be listed one by one here.

[0149] In step 302 , the data processing device applies noise to the first type of attack samples to obtain the second type of attack samples.

[0150] Exemplarily, after obtaining the first type of attack sample, the data processing device may convert the format of the first type of attack sample to obtain the first type of attack sample in text format or binary format, and store the first type of attack sample in text format or binary format in the original database.

[0151] Further illustratively, the data processing device may traverse all first-type attack samples in the original database. When traversing each first-type attack sample: noise is applied to the first-type attack sample to obtain a perturbation sample, the perturbation sample is then input into the attack recognition model, and a recognition result output by the attack recognition model is obtained; when the recognition result indicates that it is impossible to determine whether the perturbation sample is an attack sample, the perturbation sample is determined to be a second-type attack sample; otherwise, the perturbation sample is adjusted according to the recognition result, and the adjusted perturbation sample is input into the attack recognition model again. The above process is repeated until the recognition result corresponding to the adjusted perturbation sample indicates that it is impossible to determine whether the adjusted perturbation sample is an attack sample, and the adjusted perturbation sample is determined to be a second-type attack sample.

[0152] The attack recognition model can be any model with recognition capabilities. Specifically, it can be a neural network model trained using an artificial intelligence (AI) algorithm, such as a generative adversarial network (GAN) algorithm. The GAN algorithm can learn the features of known attack and non-attack samples and construct an attack recognition model based on the learned features, enabling the attack recognition model to determine the probability that an input sample is an attack sample. Specifically, the attack recognition model can include a generator and a discriminator. Any first-type attack sample input to the attack recognition model is first received by the generator. When the first-type attack sample has an N-dimensional feature space (N is a positive integer), the generator can generate a perturbed sample by applying an N-dimensional noise vector corresponding to the N-dimensional feature space to the first-type attack sample, and then input the perturbed sample into the discriminator. After the discriminator determines the probability that the perturbed sample is an attack sample, if the probability is greater than 50%, it notifies the generator to adjust the N-dimensional noise vector in a first direction; if the probability is less than 50%, it notifies the generator to adjust the N-dimensional noise vector in a second direction. Furthermore, the generator uses the adjusted N-dimensional noise vector to re-scramble the first type of attack sample to generate a new perturbation sample, and sends it to the discriminator. The discriminator re-identifies the probability that the new perturbation sample belongs to the attack sample. If the probability is 50%, the new perturbation sample can be used as a second type of attack sample. Otherwise, the above process is repeated until a perturbation sample with a probability of 50% is obtained.

[0153] In this way, through the game between the generator and the discriminator, the attack identification model can generate samples whose sample types cannot be identified. However, since this sample is obtained by scrambling the attack samples obtained from the real attack on the device to be detected (i.e., the real attack sample and the simulated attack sample), it is likely to be an attack sample. Therefore, this sample can be considered as an attack sample whose attack type cannot be determined under current technical means. This attack sample is prone to false positives during the real identification process. Based on this, the data processing device can treat this sample as a second-type attack sample and subsequently store it in the intrusion detection sample set, so as to avoid the phenomenon of these attack samples of unknown attack types being falsely reported as non-attack samples in real attack test scenarios.

[0154] In addition, since the second type of attack sample is generated by the adversarial between the generator and the discriminator in the AI ​​algorithm, the second type of sample can also be called an AI adversarial sample, or can have other names, which is not specifically limited in the embodiments of the present application.

[0155] Step 303: The data processing device constructs an intrusion detection sample set based on the first type attack sample and the second type attack sample.

[0156] For example, after obtaining the first type of attack samples (including real attack samples and simulated attack samples) and non-attack samples according to step 301, and obtaining the second type of attack samples according to step 302, the data processing device can save these samples in a unified .pcap format, and then construct an intrusion detection sample set based on all the samples in the .pcap format. The .pcap format is a format that interfaces with existing IDS systems. If it interfaces with other systems, the data processing device can also save these samples in the format that interfaces with other systems. This embodiment of the present application does not specifically limit this.

[0157] Further exemplarily, for the unlabeled samples obtained in the above step 301, the data processing device can directly discard them to save the data volume of the intrusion detection sample set, or after traversing all attack types, construct an intrusion detection sample set based on attack samples, unlabeled samples and non-attack samples, so as to add all samples that actually exist when attacking the device to be detected to the intrusion detection sample set, thereby improving the sample richness in the intrusion detection sample set, and facilitating the subsequent labeling of unlabeled samples through other analyses or performing other operations, or in some special cases, they can also be marked as abnormal samples and added to the intrusion detection sample set, without specific limitation.

[0158] For further example, after constructing an intrusion detection sample set, if the data processing device determines that the proportion of attack samples (including first-type attack samples and second-type attack samples) in the intrusion detection sample set is different from the proportion of non-attack samples, the data processing device can trim the samples with a higher proportion so that the proportion of attack samples and non-attack samples after trimming is the same, such as attack samples and non-attack samples each accounting for 50% of all samples. The samples with a higher proportion are usually non-attack samples, but may also be attack samples in certain special cases. By trimming the samples with a higher proportion, the balance between attack samples and non-attack samples in the intrusion detection sample set can be maintained, making it easier to subsequently extract data with the same proportion for attack testing, thereby improving the credibility of the attack test results.

[0159] In the above construction method, the intrusion detection sample set can cover the first type of attack samples in known intrusion scenarios, the second type of attack samples in unknown intrusion scenarios, and non-attack samples. In this way, not only can the sample information in the intrusion detection sample set be more complete, but also when the intrusion detection sample set is used to perform an attack test on the device to be detected, the anti-attack effect of the device to be detected can be more accurately defined based on whether the device to be detected can intercept the attack samples and whether it can not intercept the non-attack samples.

[0160] The above content introduces the specific construction process of the intrusion detection sample set. The following is a detailed introduction to the application of the constructed intrusion detection sample set.

[0161] FIG5 exemplarily illustrates an application scenario diagram of an intrusion detection sample set provided by an embodiment of the present application. As shown in FIG5 , the intrusion detection sample set can be applied to one or more of the following scenarios: a model training scenario, an offline evaluation scenario, or an online evaluation scenario. The solid line in the figure illustrates the application process for the model training scenario, the dotted line in the figure illustrates the application process for the offline evaluation scenario, and the double-node line in the figure illustrates the application process for the online evaluation scenario. The following detailed description of these three application scenarios is provided with reference to FIG5 .

[0162] Model training scenario

[0163] Referring to the solid line portion illustrated in Figure 5 , model training refers to the use of an intrusion detection sample set to train an intrusion detection model during early algorithm development and verification. Because the training samples required for the intrusion detection model have their own unique feature format, and the intrusion detection samples in the intrusion detection sample set may not necessarily match this feature format, the data processing device may also perform feature extraction on the intrusion detection sample set before training the intrusion detection model, and construct an offline detection sample set based on the extracted features. In this way, when it is necessary to train the intrusion detection model, the data processing device can directly select a portion of the offline detection samples from the offline detection sample set as training set data, input them into the model training device, and the model training device uses the training set data to train the intrusion detection model. The training set data may include both attack-type offline detection samples and normal-type offline detection samples. The attack-type offline detection samples are obtained by performing feature extraction on first-type attack samples and / or second-type attack samples in the intrusion detection sample set, while the normal-type offline detection samples are obtained by performing feature extraction on non-attack samples in the intrusion detection sample set. The number of attack-type offline detection samples and normal-type offline detection samples can also be consistent, so that a more effective intrusion detection model can be trained based on balanced data.

[0164] It should be noted that each offline detection sample in the offline detection sample set can be saved in a text format, such as .csv. This .csv format is compatible with the model training device in existing IDS systems. If other types of model training devices are connected, the data can be saved in a format compatible with the other system, without limitation.

[0165] In the embodiment of the present application, when performing feature extraction, the data processing device may analyze each intrusion detection sample in isolation, that is, extract features from each intrusion detection sample to obtain an offline detection sample, or may combine multiple intrusion detection samples for centralized analysis, such as extracting features from multiple intrusion detection samples with a correlation to obtain an offline detection sample. Multiple intrusion detection samples with a correlation may, for example, be multiple intrusion detection samples belonging to the same connection, multiple intrusion detection samples whose traffic data originates from the same ECU, or multiple intrusion detection samples whose traffic data is sent to the same ECU.

[0166] For further example, taking the case where multiple intrusion detection samples with an associated relationship are multiple intrusion detection samples belonging to the same connection, when the device to be detected is a vehicle, the message types in the vehicle may generally include TCP messages, CAN (FD) messages, and UDP messages. TCP messages are messages transmitted on a connection after a connection is established between at least two ECUs, while CAN (FD) messages and UDP messages are messages sent on a corresponding bus via broadcast and then retrieved from the bus by the desired ECU node. The content of these three messages includes both a source Internet Protocol (IP) address and a destination IP address. It can be seen that TCP messages have the concept of a connection, while CAN (FD) messages and UDP messages do not. Therefore, after constructing an intrusion detection sample set, the data processing device can first determine the message type of each intrusion detection sample in the intrusion detection sample set. Then, for each intrusion detection sample whose message type is a TCP message, all intrusion detection samples belonging to the same TCP connection are obtained based on the source IP address and destination IP address contained in the message content. Feature extraction is performed on these intrusion detection samples to obtain an offline detection sample. Conversely, for each intrusion detection sample whose message type is CAN(FD) message or UDP message, feature extraction is performed on each intrusion detection sample separately to obtain an offline detection sample. In this way, by combining all intrusion detection samples on a TCP connection line to obtain an offline detection sample, the number of samples in the offline detection sample set can be effectively reduced while making the features contained in each offline detection sample more relevant. Moreover, by performing separate feature analysis on each intrusion detection sample that does not have a connection concept, any unrelated intrusion detection samples are not missed, ensuring that the offline detection sample set covers all intrusion detection samples.

[0167] Further illustratively, the features extracted by the above feature extraction may include one or more of the following features: timestamp, frequency feature, protocol type, content feature, packet loss rate, number of error packets, connection duration, connection initiator, and connection receiver. Among them, the timestamp refers to the time when the message is collected, including the date and time. The frequency feature refers to the communication frequency of sending and receiving messages between the source IP and the destination IP. The protocol type refers to the protocol to which the message is adapted, such as any of the protocol types in Table 3 above. The content feature refers to the actual content carried in the message, such as data or instructions. The packet loss rate refers to the proportion of messages lost when sending and receiving messages between the source IP and the destination IP. The number of error packets refers to the number of error packets that occur when sending and receiving messages between the source IP and the destination IP. The connection duration refers to the time interval from the start of a connection to its disconnection. The connection initiator refers to the ECU that requests to establish a connection. The connection receiver refers to the ECU that receives the request sent by the connection initiator.

[0168] Further exemplarily, in order to maintain the format consistency of each offline detection sample, the data processing device can also extract all the above-mentioned features for each type of intrusion detection sample. Specifically, for intrusion detection samples belonging to TCP messages, since there is a concept of connection, all the above-mentioned features can be extracted. However, for intrusion detection samples belonging to CAN (FND) messages or UDP messages, since there is no concept of connection, only the timestamp, frequency feature, protocol type, content feature, packet loss rate and number of error packets among the above-mentioned features can be extracted, but the connection duration, connection initiator and connection receiver cannot be extracted. In this case, in order to maintain the format consistency of offline detection samples, the data processing device can also configure the features that cannot be extracted as preset characters, which can be, for example, numbers, letters, symbols or a combination of one or more of them.

[0169] It should be noted that the above content only provides an example of feature extraction. As for which features need to be extracted in actual applications, those skilled in the art can set them based on their experience, and the embodiments of this application do not limit this.

[0170] In the model training scenario, by providing a variety of extractable features, users can easily select one or more features according to the application layer requirements in the actual model training scenario to convert the intrusion detection sample set into the offline detection sample set, so as to be suitable for different model training scenarios and improve the versatility of the intrusion detection sample set in the model training field.

[0171] Offline evaluation scenario

[0172] Referring to the dotted portion shown in FIG5 , offline evaluation refers to the use of an intrusion detection sample set to test the detection effectiveness of an intrusion detection model trained in a model training scenario during early algorithm development and verification. Since an offline detection sample set has already been extracted from the model training scenario, after training an intrusion detection model using a portion of the offline detection samples in the offline detection sample set, the data processing device can also select another portion of offline detection samples from the offline detection sample set as test set data and input them into the model training device. The model training device then uses the test set data to test the intrusion detection model to obtain offline evaluation information, which is then sent to the data processing device. The data processing device then evaluates the detection effectiveness of the intrusion detection model based on the offline evaluation information.

[0173] The test set data may also include offline detection samples of attack types and offline detection samples of normal types, and the number of offline detection samples of attack types and normal types may also be consistent. Thus, after the model training device obtains the detection results of the intrusion detection model for each offline detection sample, it combines the number of offline detection samples of attack types identified as attack samples and the number of offline detection samples of normal types identified as non-attack samples to obtain the total number of offline detection samples correctly identified, and combines the number of offline detection samples of attack types identified as non-attack samples and the number of offline detection samples of normal types identified as attack samples to obtain the total number of offline detection samples incorrectly identified. The total number of offline detection samples correctly identified and the total number of offline detection samples incorrectly identified are then carried in offline evaluation information and sent to the data processing device. Furthermore, if the total number of offline detection samples correctly identified is greater and the total number of offline detection samples incorrectly identified is smaller, the data processing device can determine that the detection effect of the intrusion detection model is better; otherwise, the detection effect is worse.

[0174] It should be noted that the specific implementation process of the above-mentioned offline evaluation is only an example. There may be other implementation methods in actual operation. For example, the model training device can also directly send the detection results of each offline detection sample as offline evaluation information to the data processing device. The data processing device will automatically count the total number of offline detection samples that are correctly identified and the total number of offline detection samples that are incorrectly identified to complete the offline evaluation. There are many possible implementation methods, which will not be listed one by one here.

[0175] In the offline evaluation scenario, the intrusion detection sample set can support the offline evaluation of the detection effect of the intrusion detection model. In this way, the model training device can continuously optimize the intrusion detection model according to the detection effect, obtain an intrusion detection model with better detection effect, and provide a basis for the implementation of the intrusion detection model on the equipment to be detected.

[0176] Online evaluation scenario

[0177] Referring to the double-node line shown in Figure 5, online evaluation refers to the use of an intrusion detection sample set to test the detection effectiveness of the device to be detected, where the intrusion detection model is deployed, during the later stages of algorithm deployment. After the aforementioned model training and offline evaluation, the intrusion detection model can achieve good detection results. However, this detection effect is only measured without the device to be detected and cannot represent the actual effect after actual application on the device to be detected. Therefore, it is necessary to deploy the intrusion detection model on the device to be detected, conduct a real attack on the device to be detected, and then determine the actual detection effect of the intrusion detection model applied to the device to be detected based on the device's response.

[0178] In a specific implementation, the data processing device can perform format conversion on the intrusion detection samples in the intrusion detection sample set to obtain online detection samples that match the format of the test tool, and then input the online detection samples into the device to be detected through the test tool, and obtain the online evaluation information generated by the intrusion detection model deployed in the device to be detected for the online detection samples, and evaluate the detection performance of the device to be detected deployed with the intrusion detection model based on the online evaluation information. Among them, the format conversion can also be performed on a connection basis. For example, for each intrusion detection sample belonging to the same TCP connection, these intrusion detection samples are first aggregated to obtain a preliminary online detection sample, and then the format of the preliminary online detection sample is converted to a format adapted by the test tool in the current scenario to obtain an online detection sample. For each intrusion detection sample belonging to a CAN (FD) message or a UDP message, the format of each intrusion detection sample can be directly converted to a format adapted by the test tool in the current scenario to obtain a corresponding online detection sample.

[0179] For example, when the device to be tested is a vehicle, the test tool can inject different online detection samples into the vehicle in real time through the vehicle's OBD interface. For each online detection sample, if the vehicle identifies it as an attack sample, it can alarm the security operations center (SOC) in the cloud through the VCU. Then, after all the online detection samples are injected, the SOC combines the number of all online detection samples and the alarm record information of the vehicle during this period to determine the total number of online detection samples that are correctly identified and the total number of online detection samples that are incorrectly identified. Then, based on these two total numbers, online evaluation information is generated and sent to the data processing device. Then, if the total number of online detection samples that are correctly identified is greater and the total number of online detection samples that are incorrectly identified is smaller, the data processing device can determine that the detection effect of the device to be tested with the intrusion detection model deployed is better, and vice versa.

[0180] Furthermore, considering that the test injection of CAN (FD) message type requires test tools such as CANoe or PCAN, and the test injection of a data set requires test tools such as CANoe or Technica, the test tool in the embodiment of the present application can be specifically CANoe, PCAN, Technica or other tools that can realize online testing. Correspondingly, the online detection sample after format conversion can be in the format corresponding to .PCAP, .ASC, .BLF or other tools that can realize online testing. After the data processing device converts the intrusion detection sample in .pcap format into an online detection sample in the format corresponding to the test tool, it injects the online detection sample into the device to be tested through the test tool. In this way, by enabling the intrusion detection sample set to support conversion into formats corresponding to multiple test tools, it is convenient for users to select the corresponding format for conversion according to the test tool in the actual online scenario, so as to be suitable for different online evaluation scenarios and improve the versatility of the intrusion detection sample set in the field of online evaluation.

[0181] In the online evaluation scenario, the intrusion detection sample set can support the online evaluation of the detection effect of the device to be tested that is deployed with the intrusion detection model. This allows users to determine the anti-attack performance of the device to be tested based on the detection effect, ensuring that only devices to be tested with good anti-attack performance are shipped.

[0182] The above content details how the data processing device constructs and applies the intrusion detection sample set. In addition, in some scenarios, the data processing device can also evaluate the quality of the constructed intrusion detection sample set. The following exemplifies a specific evaluation process.

[0183] FIG6 is a flow chart of evaluating an intrusion detection sample set according to an embodiment of the present application. The flow includes:

[0184] Step 601: The data processing device obtains an intrusion detection sample set.

[0185] In step 602, the data processing device calculates the value of the intrusion detection sample set under each preset indicator, and calculates the evaluation value corresponding to the intrusion detection sample set according to the value of the intrusion detection sample set under each preset indicator.

[0186] Among them, each preset indicator can be set according to the characteristics of the system architecture to which the device to be tested belongs, and can exemplarily include quantitative indicators and qualitative indicators. Quantitative indicators refer to evaluation indicators that can be defined by accurate quantities, and qualitative indicators refer to evaluation indicators that cannot be directly quantified and need to be quantified through other means.

[0187] For example, Table 2 shows a schematic table of possible preset indicators provided in an embodiment of the present application:

[0188] Table 2

[0189] As shown in Table 2, in this example, each preset indicator can include one or more of a data redundancy indicator, an attack coverage indicator, a protocol coverage indicator, a service coverage indicator, a balance indicator, a feature independence indicator, and an ease of use indicator. The data redundancy indicator, attack coverage indicator, protocol coverage indicator, service coverage indicator, and balance indicator are quantitative indicators, while the feature independence indicator and ease of use indicator are qualitative indicators. Furthermore, to maintain consistency in the measurement of each preset indicator, the value range of each preset indicator can also be consistent, for example, all set to [0, 1]. That is, the value of each preset indicator can be any real number between 0 and 1, inclusive.

[0190] The following is a detailed introduction to the various preset indicators listed in Table 2.

[0191] Data redundancy index

[0192] The data redundancy index is used to indicate the degree of non-redundancy in an intrusion detection sample set. For example, it can be expressed as the ratio of the number of non-redundant intrusion detection samples in the intrusion detection sample set to the total number of intrusion detection samples. For example, if there are 100 intrusion detection samples in the intrusion detection sample set, and 5 of them correspond to the same ID, and the remaining 95 correspond to different IDs, then the data redundancy index value for the intrusion detection sample set is 95 / 100.

[0193] It should be noted that the data redundancy index can also be expressed in other forms, as long as it can be guaranteed to be negatively correlated with the ratio of the number of redundant intrusion detection samples to the total number of intrusion detection samples. Thus, when all intrusion detection samples are different, the value of the data redundancy index is the largest, the number of valid samples in the intrusion detection sample set is the largest, and the sample adequacy is the best. As the number of identical intrusion detection samples increases, the value of the data redundancy index gradually decreases, the number of valid samples in the intrusion detection sample set gradually decreases, and the sample adequacy gradually deteriorates. Until all intrusion detection samples are identical, the value of the data redundancy index is the smallest, the number of valid samples in the intrusion detection sample set is the smallest, and the sample adequacy is the worst.

[0194] Attack coverage indicators

[0195] The attack coverage metric indicates the degree of coverage of the attack types included in the intrusion detection sample set. It can be expressed, for example, as the ratio of the number of attack types covered by the intrusion detection sample set to the total number of attack types that may exist on the device to be detected. For example, assuming that the total number of attack types that may exist on the device to be detected is the 25 attack types shown in Table 1, and the first type of attack sample in the intrusion detection sample set is obtained by attacking the device to be detected using 10 of these attack types, then the attack coverage metric for the intrusion detection sample set is 10 / 25.

[0196] It should be noted that the attack coverage index can also be expressed in other forms, as long as it can ensure a positive correlation with the ratio of the number of covered attack types to the number of all possible attack types. In this way, when the intrusion detection sample set covers all possible attack types, the attack coverage index value is the largest, the intrusion detection sample set contains the most sample types, and the sample diversity is the best. As the number of covered attack types decreases, the attack coverage index value gradually decreases, the sample types in the intrusion detection sample set gradually decreases, and the sample diversity gradually deteriorates. Until no attack type is covered, the attack coverage index value is the smallest, the intrusion detection sample set contains the fewest sample types, and the sample diversity is the worst.

[0197] Protocol coverage indicators

[0198] The protocol coverage index is used to indicate the coverage level of the communication protocols included in the intrusion detection sample set. For example, it can be expressed as the ratio of the number of communication protocols covered by the intrusion detection sample set to the total number of communication protocols that may exist in the device to be detected. For example, please refer to Table 3, which is a schematic table of communication protocols that may exist in the field of Internet of Vehicles provided in an embodiment of the present application. It should be understood that with the development of Internet of Vehicles technology, new communication protocols may appear in the future. Therefore, the communication protocols in Table 3 may also be updated accordingly, and this embodiment of the present application does not specifically limit this.

[0199] Table 3

[0200] Assuming that the device to be detected is a vehicle, all possible communication protocols in the vehicle are the eight protocols shown in Table 3, and the communication protocols covered by the intrusion detection sample set include CAN (FD), DoCAN, DDS, MQTT, and HTTP (S). The value of the intrusion detection sample set under the protocol coverage index is 5 / 8.

[0201] It should be noted that the protocol coverage index can also be expressed in other forms, as long as it can be guaranteed to be positively correlated with the ratio of the number of covered communication protocols to the number of all possible communication protocols. In this way, when the intrusion detection sample set covers all possible communication protocols, the value of the protocol coverage index is the largest. The samples in the intrusion detection sample set are obtained by attacking all communication protocol messages. The intrusion detection sample set has the most sample sources and the richness of the samples is the best. As the number of covered protocols decreases, the value of the protocol coverage index gradually decreases, the source of samples in the intrusion detection sample set gradually decreases, and the richness of the samples gradually deteriorates. Until no protocol is covered, the value of the protocol coverage index is the smallest, the source of samples in the intrusion detection sample set is the least, and the richness of the samples is the worst.

[0202] Business coverage indicators

[0203] The service coverage index is used to indicate the coverage of the services included in the intrusion detection sample set. It can be expressed as the ratio of the number of services included in the intrusion detection sample set to the total number of services that may exist in the device to be detected. For example, assuming that all possible services in the device to be detected include remote control, log transmission, over-the-air (OTA) updates, diagnostic services, video transmission, and network management, and the services included in the intrusion detection sample set are remote control, OTA updates, and diagnostic services, then the value of the service coverage index for the intrusion detection sample set is 3 / 6.

[0204] It should be noted that the service coverage indicator can also be expressed in other forms, as long as it maintains a positive correlation with the ratio of the number of covered services to the number of all possible services. Thus, when the intrusion detection sample set covers all services in the device to be detected, the service coverage indicator is at its maximum value, and the sample has the best service applicability. As the number of covered services decreases, the service coverage indicator value gradually decreases, and the service applicability of the intrusion detection sample set gradually deteriorates, until no service is covered, the protocol coverage indicator value is at its minimum, and the service applicability of the intrusion detection sample set is at its worst.

[0205] Data labeling metrics

[0206] The data labeling index is used to indicate the degree of labeling of intrusion detection samples in the intrusion detection sample set, and can be expressed, for example, as the ratio of the number of labeled intrusion detection samples to the total number of intrusion detection samples in the intrusion detection sample set. Labeled intrusion detection samples may include the real attack samples, simulated attack samples, and non-attack samples obtained in step 301, as well as the second-type attack samples obtained in step 302. Unlabeled intrusion detection samples include the unlabeled samples obtained in step 301. For example, when there are 100 intrusion detection samples in the intrusion detection sample set, if 20 of them are first-type attack samples, 20 are second-type attack samples, 55 are non-attack samples, and 5 are unlabeled samples, then the number of labeled intrusion detection samples in the intrusion detection sample set is 95. Therefore, the value of the intrusion detection sample set under the data labeling index is 95 / 100.

[0207] It should be noted that the data labeling indicator can also be expressed in other forms, as long as it can ensure a positive correlation with the ratio of the number of labeled intrusion detection samples to the total number of all intrusion detection samples. In this way, when all intrusion detection samples in the intrusion detection sample set are labeled, it means that all intrusion detection samples have been clearly divided into non-attack samples and attack samples, and do not contain uncertain samples. The sample clarity in the intrusion detection sample set is the highest. As the number of labeled intrusion detection samples decreases, the number of uncertain samples contained in the intrusion detection sample set gradually increases, and the sample clarity gradually deteriorates, until all intrusion detection samples are unlabeled, the intrusion detection sample set contains the most uncertain samples, and the sample clarity is the worst.

[0208] Balance Index

[0209] The balance index is used to indicate the degree of balance between attack samples and non-attack samples in an intrusion detection sample set. For example, it can be expressed as the ratio of the difference between the total number of all intrusion detection samples in the intrusion detection sample set and the difference between the number of attack samples and non-attack samples to the total number of all intrusion detection samples. For example, when there are 100 intrusion detection samples in the intrusion detection sample set, if 45 of them are attack samples and 55 are non-attack samples, then the difference between the number of attack samples and non-attack samples in the intrusion detection sample set is 10. Therefore, the value of the balance index for the intrusion detection sample set can be (100-10) / 100.

[0210] It should be noted that the balance index can also be expressed in other forms, as long as it can ensure a negative correlation with the difference in the number of attack samples and non-attack samples. In this way, when the attack samples and non-attack samples in the intrusion detection sample set each account for 50%, the value of the balance index is the largest, the sample balance in the intrusion detection sample set is the best, and it is easier to subsequently extract an equal number of attack samples and non-attack samples from the intrusion detection sample set for testing. As the difference in the number of attack samples and non-attack samples increases, the sample balance in the intrusion detection sample set gradually deteriorates, and it becomes more difficult to extract an equal number of attack samples and non-attack samples from the intrusion detection sample set for testing. Until the intrusion detection sample set is entirely attack samples or entirely non-attack samples, the sample balance in the intrusion detection sample set is the worst, and it is impossible to extract an equal number of attack samples and non-attack samples from the intrusion detection sample set for testing.

[0211] Feature independence index

[0212] The feature independence index is used to indicate the independence of the features extracted when using the intrusion detection sample set for offline evaluation. The more independent features there are, the larger the value of the feature independence index is, and the fewer independent features there are, the smaller the value of the feature independence index is.

[0213] It should be noted that whether a feature is independent can be judged by technical personnel in this field based on experience, for example, it can be judged by at least two of engineers, experts or third-party organizations, so as to obtain a more accurate evaluation result by combining the experience of all parties.

[0214] Usability indicators

[0215] The usability index is used to indicate the versatility of application scenarios that the intrusion detection sample set is compatible with. For example, the application scenario can be the format of the test tools that can be supported when the intrusion detection sample set is used for online evaluation. The more formats of test tools that can be supported, the larger the value of the usability index. The fewer formats of test tools that can be supported, the smaller the value of the usability index.

[0216] It should be noted that the application scenarios that the intrusion detection sample set is compatible with can also be judged by those skilled in the art based on their experience, for example, by at least two of engineers, experts or third-party organizations, so as to obtain a more accurate evaluation result by combining the experience of all parties.

[0217] Furthermore, according to the various preset indicators shown in Table 2 above, after calculating the value of the intrusion detection sample set under each preset indicator, the data processing device can also perform weighted averaging on the values ​​of the intrusion detection sample set under each preset indicator according to the weight corresponding to each preset indicator, and use the calculated weighted average as the evaluation value of the intrusion detection sample set. The weights corresponding to the various preset indicators can be the same or different. For example, in one example, each quantitative preset indicator can be configured to correspond to a first weight, and each qualitative preset indicator can be configured to correspond to a second weight, and the first weight is greater than the second weight. In this way, by reducing the weight of the qualitative preset indicator, the influence of the preset indicator that needs to be judged by human experience on the evaluation value corresponding to the intrusion detection sample can be reduced, so that the evaluation process pays more attention to rational judgment criteria.

[0218] In step 603 , the data processing device determines whether the evaluation value corresponding to the intrusion detection sample set is lower than a preset threshold value. If so, step 604 is executed; if not, step 605 is executed.

[0219] In step 604 , the data processing device adjusts the intrusion detection sample set, and then executes step 602 .

[0220] For example, when the evaluation value corresponding to the intrusion detection sample set is lower than a preset threshold, it means that the quality of the intrusion detection sample set cannot meet the requirements. In this case, the data processing device can adjust the intrusion detection sample set according to the values ​​of the intrusion detection sample set under various preset indicators, so that the values ​​of the adjusted intrusion detection sample set under one or more preset indicators become larger, thereby increasing the evaluation value corresponding to the intrusion detection sample set. For example, when the value of the intrusion detection sample set under the balance indicator is low, the data processing device can trim a large number of samples so that the attack samples and non-attack samples are close to each other, thereby improving the value of the intrusion detection sample set under the balance indicator. Alternatively, when the value of the intrusion detection sample set under the data redundancy indicator is low, the data processing device can delete redundant intrusion detection samples to reduce the ratio of the number of redundant intrusion detection samples to the total number of intrusion detection samples, thereby improving the value of the intrusion detection sample set under the data redundancy indicator.

[0221] It should be noted that the above content is only an example of a possible adjustment method. In actual operation, those skilled in the art may also choose other methods to adjust the intrusion detection sample set based on experience, and the embodiments of the present application do not specifically limit this.

[0222] Step 605: The data processing device stores the intrusion detection sample set.

[0223] For example, when the evaluation value corresponding to the intrusion detection sample set is not lower than a preset threshold, it means that the quality of the intrusion detection sample set meets the requirements. In this case, the data processing device can store the intrusion detection sample set in a database in a text format. The text format can be, for example, .cpap or other formats supported by the IDS.

[0224] In the above implementation, by evaluating the constructed intrusion detection sample set against preset indicators, the intrusion detection sample set can be continuously optimized based on the evaluation results, making the optimized intrusion detection sample set more suitable for the system architecture of the device to be detected. Furthermore, by comprehensively evaluating the construction quality of the intrusion detection sample set using a combination of quantitative and qualitative indicators, the evaluation results can be made more comprehensive and persuasive.

[0225] It should be understood that evaluating the quality of the intrusion detection sample set through preset indicators is only an optional evaluation method. In actual operation, there may be other evaluation methods, such as direct evaluation through human experience, indirect evaluation through a third-party organization, or evaluation by comparing historical intrusion detection sample sets, etc. The embodiments of the present application do not make specific limitations on this.

[0226] Based on the above, FIG7 illustrates a design architecture diagram for developing the aforementioned data processing solution provided by an embodiment of the present application. This design architecture diagram can be, for example, an interface presented to development and testing personnel or operations and maintenance personnel, who then write corresponding program code based on the various functions that need to be implemented in this interface. Specifically, as shown in FIG7 , this design architecture includes a hardware tool layer, a data generation layer, a call interface layer, and an application scenario layer. The contents of each layer are described in detail below.

[0227] The hardware tool layer is responsible for providing hardware interface tools for accessing the device to be tested and supporting the collection of traffic data from the device to be tested. For example, if the device to be tested is a vehicle, the hardware tool layer may primarily include a signal-oriented CAN (FD) tool and a service-oriented ETH tool. These two tools can be connected to the vehicle's OBD port to support the data collection module in collecting traffic data generated by the vehicle in response to attack behaviors.

[0228] The data generation layer is responsible for constructing and reviewing intrusion detection sample sets. It primarily includes a data acquisition module, an AI sample generation module, a feature extraction module, a format conversion module, a sample set evaluation module, and a data review module. It may also include a database. During implementation, the data acquisition module can collect traffic data generated by the device under test in response to attack behaviors through the hardware tool layer. By analyzing this traffic data, it labels the data to identify first-type attack samples and non-attack samples, and stores these first-type attack and non-attack samples in the database. The AI ​​sample generation module can add noise to the first-type attack samples labeled by the data acquisition module, process them using an AI adversarial algorithm to obtain second-type attack samples prone to false positives, and then store these second-type attack samples in the database. In this way, the first-type attack samples, second-type attack samples, and non-attack samples in the database constitute the intrusion detection sample set. Furthermore, the feature extraction module can extract features from the intrusion detection sample set in the database to form an offline detection sample set, and the format conversion module can convert the format of the intrusion detection sample set in the database to form an online detection sample set. The sample set evaluation module calculates the evaluation value of the intrusion detection sample set under various preset indicators. If the evaluation value falls below the preset threshold, it adjusts the intrusion detection sample set until the corresponding evaluation value of the intrusion detection sample set is no less than the preset threshold. The data viewing module can display some or all intrusion detection samples to developers, testers, or operators based on their commands.

[0229] The call interface layer is responsible for providing an application programming interface (API). Specifically, it can call corresponding data from the data generation layer through the API and provide it to the upper-layer application according to the usage requirements of the upper-layer application. For example, the call interface layer can provide some offline detection samples from the offline detection sample set in the data generation layer to the upper-layer application through the API to train the intrusion detection model. It can also provide another part of the offline detection samples from the offline detection sample set in the data generation layer to the upper-layer application through the API to implement offline evaluation of the intrusion detection model. It can also provide online detection samples from the data generation layer to the upper-layer application through the API to implement online evaluation of the device to be detected that is deployed with the intrusion detection model.

[0230] The application scenario layer is responsible for interacting with external devices to apply intrusion detection sample sets to various possible application scenarios. In some typical application scenarios, the application scenario layer can use the offline detection samples provided by the call interface layer to train intrusion detection models during IDS development. It can also use the online detection samples provided by the call interface layer to evaluate the detection effectiveness of devices deployed with intrusion detection models during vehicle-cloud operations and maintenance. It can also use the offline detection samples provided by the call interface layer to evaluate the detection effectiveness of intrusion detection models during IDS testing. It can also provide the intrusion detection sample sets provided by the call interface layer to third-party testing and certification organizations to ensure that the devices under test are certified by third-party certification organizations before they can be shipped smoothly.

[0231] In the above-mentioned embodiments of the present application, by combining the first type of attack samples of known attack types, the second type of attack samples of unknown attack types and non-attack samples to construct an intrusion detection sample set, the sample types in the intrusion detection sample set can be made richer and more comprehensive. In this way, when applied in the field of Internet of Vehicles, even if there are significant differences in the network topology and traffic characteristics of different vehicle models, this construction method can also construct adaptive and rich intrusion detection samples for each type of vehicle model, effectively improving the accuracy of using intrusion detection samples to evaluate vehicle network security.

[0232] It should be understood that the data processing method provided in this application can also be extended to any information system that has a demand for network security. For example, it can also be applied in the field of smart homes. By attacking smart home products and scrambling them to obtain a rich set of attack samples, the attacker profile in the smart home scenario can be studied. Alternatively, it can also be applied in the field of industrial control. By introducing attack samples into the digital twin model to study intrusion defense, the robustness of the industrial control system can be enhanced. Specifically, a vehicle or a component to be tested on a vehicle can be virtually constructed through a cloud server, and then the above-mentioned data processing operation can be performed on this virtual vehicle or vehicle component to obtain an intrusion detection sample set. In this way, the anti-attack capability of the vehicle can be known in advance before the actual assembly of the vehicle, so that the vehicle can be actually built only when it is determined that the vehicle can better defend against attacks, effectively saving manpower and material costs.

[0233] In addition, with the evolution of system architecture and the emergence of new scenarios, the data processing method provided in this application is also applicable to similar technical problems, and this application does not make specific limitations on this.

[0234] It should be noted that the names of the above-mentioned information are merely examples. With the evolution of communication technology, the names of any of the above-mentioned information may change. However, no matter how the names change, as long as their meanings are the same as those of the above-mentioned information in this application, they fall within the scope of protection of this application.

[0235] The above mainly introduces the solution provided by the present application from the perspective of the interaction between various network elements. It can be understood that in order to realize the above functions, the above-mentioned network elements include hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software-driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0236] According to the aforementioned method, FIG8 is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. The data processing device can be any device with processing capabilities, such as a server, or can also be a chip or circuit, such as a chip or circuit that can be set in a server, or can also be a server cluster composed of multiple servers. As shown in FIG8, the data processing device 800 can include a processor 801, a memory 802, and a transceiver 803, and can further include a bus system, and the processor 801, the memory 802, and the transceiver 803 can be connected via the bus system.

[0237] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 801 or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor 801. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 802, and the processor 801 reads the information in the memory 802 and completes the steps of the above method in conjunction with its hardware.

[0238] It should be understood that the processor 801 may be a chip. For example, the processor 801 may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD), or other integrated chips.

[0239] It is understood that the memory 802 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0240] In the embodiment of the present application, the memory 802 is used to store instructions, and the processor 801 is used to execute the instructions stored in the memory 802 to implement the method corresponding to the data processing device in any one or more of Figures 3, 4, or 6. Specifically, the processor 801 attacks the device to be detected by calling the transceiver 803 to obtain a first type of attack sample, obtains a second type of attack sample by applying noise to the first type of attack sample, and then constructs an intrusion detection sample set based on the first type of attack sample and the second type of attack sample.

[0241] In an optional implementation, the first type of attack samples may include real attack samples and simulated attack samples. The real attack samples are obtained by manually attacking the device to be detected, and the simulated attack samples are obtained by attacking the device to be detected with an attack tool.

[0242] In an optional embodiment, the real attack sample may correspond to one or more of the following attack types: identity identification number ID non-existence attack, replay attack, tampering attack, data length error attack, signal out of defined range attack, context error attack, ID source non-specified electronic control unit ECU attack, identical ID attack, controller area network CAN scan attack, unified diagnostic service UDS execution sensitive operation attack, message authentication error attack, ECU identity spoofing attack, man-in-the-middle attack, ECU authentication error attack, brute force attack, application layer protocol error attack, unknown stack connection attack, unknown stack connection attack.

[0243] In an optional implementation, the simulated attack sample may correspond to one or more of the following attack types: ID Fuzz attack, data Fuzz attack, DoS attack of CAN, DoS attack of ETH, malformed packet injection attack, and port scanning attack.

[0244] In an optional embodiment, the processor 801 is specifically used to: traverse each attack type of a preset plurality of attack types, and when traversing each attack type: call the transceiver 803 to execute the attack behavior corresponding to the attack type on the device to be detected, and obtain the traffic data generated by the device to be detected for the attack behavior; if the traffic data is attack traffic, mark the traffic data as a first type attack sample.

[0245] In an optional implementation, after obtaining the traffic data generated by the device to be detected in response to the attack behavior, if the processor 801 determines that the traffic data is normal traffic, it marks the traffic data as a non-attack sample, and then constructs an intrusion detection sample set based on the first type of attack sample, the second type of attack sample and the non-attack sample.

[0246] In an optional embodiment, the processor 801 is specifically configured to: apply noise to the first type of attack sample to obtain a perturbed sample, input the perturbed sample into an attack recognition model, obtain a recognition result output by the attack recognition model, and then adjust the perturbed sample based on the recognition result until the recognition result corresponding to the adjusted perturbed sample indicates that it is impossible to determine whether the adjusted perturbed sample is an attack sample, and then determine the adjusted perturbed sample as a second type of attack sample. The recognition result is used to indicate whether the perturbed sample is an attack sample.

[0247] In an optional implementation, after constructing an intrusion detection sample set based on the first type of attack samples and the second type of attack samples, the processor 801 may also perform feature extraction on the intrusion detection sample set to obtain an offline detection sample set, which is used to evaluate the detection performance of the intrusion detection model.

[0248] In an optional implementation, the processor 801 is specifically used to: determine the message type of each intrusion detection sample in the intrusion detection sample set; for each intrusion detection sample whose message type is a TCP message, obtain an offline detection sample by performing feature extraction on all intrusion detection samples belonging to the same TCP connection; for each intrusion detection sample whose message type is a CAN (FD) message or a UDP message, obtain an offline detection sample by performing feature extraction on the intrusion detection sample of each CAN (FD) message or UDP message.

[0249] In an optional embodiment, the extracted features include one or more of the following features: timestamp, frequency feature, protocol type, content feature, packet loss rate, number of error packets, connection duration, connection initiator, and connection receiver.

[0250] In an optional implementation, after the processor 801 constructs an intrusion detection sample set based on the first type of attack samples and the second type of attack samples, it can also convert the format of the intrusion detection sample set to obtain an online detection sample set that matches the format of the test tool, and then input the online detection sample set into the device to be tested through the test tool. The online detection sample is used to evaluate the detection performance of the device to be tested that is deployed with an intrusion detection model.

[0251] In an optional implementation, after constructing the intrusion detection sample set based on the first type of attack samples and the second type of attack samples, the processor 801 can also determine the evaluation value corresponding to the intrusion detection sample set based on the values ​​of the intrusion detection sample set under various preset indicators, and adjust the intrusion detection sample set when the evaluation value is lower than the preset threshold.

[0252] In an optional implementation, the preset indicators may include one or more of the following indicators: data redundancy indicator, attack coverage indicator, protocol coverage indicator, service coverage indicator, data labeling indicator, balance indicator, feature independence indicator, and ease of use indicator.

[0253] In an optional implementation, the first type of attack sample may be obtained by attacking any of the following areas: the entire device to be detected; one or more physical areas of the device to be detected; or one or more functional areas of the device to be detected.

[0254] For the concepts, explanations, detailed descriptions and other steps involved in the data processing device 800 and related to the technical solutions of the data processing device provided in the embodiments of the present application, please refer to the descriptions of these contents in the aforementioned methods or other embodiments, which are not repeated here.

[0255] Based on the above embodiments and the same concept, FIG9 is a schematic diagram of another data processing device provided in an embodiment of the present application. The data processing device 900 can be, for example, the data processing device described in any of the above embodiments, or a chip or circuit, such as a chip or circuit that can be provided in a data processing device. The data processing device 900 can implement the steps performed by the data processing device in any one or more of the corresponding methods shown in FIG3, FIG4, or FIG6.

[0256] As shown in Figure 9, the data processing device 900 may include an attack unit 901, a perturbation unit 902, and a construction unit 903. It may also illustratively include one or more of a feature extraction unit 904, a format conversion unit 905, and an adjustment unit 906. The attack unit 901 is configured to obtain a first type of attack sample by attacking the device under detection; the perturbation unit 902 is configured to obtain a second type of attack sample by applying noise to the first type of attack sample; and the construction unit 903 is configured to construct an intrusion detection sample set based on the first and second type of attack samples. In some scenarios, the feature extraction unit 904 is configured to perform feature extraction on the intrusion detection sample set to obtain an offline detection sample set, which is used to evaluate the detection performance of the intrusion detection model. In other scenarios, the format conversion unit 905 is configured to convert the format of the intrusion detection sample set to obtain an online detection sample set that matches the format of the test tool. The online detection sample set is then input into the device under detection through the test tool. The online detection sample set is used to evaluate the detection performance of the device under detection deployed with the intrusion detection model. In other scenarios, the adjustment unit 906 is configured to determine an evaluation value corresponding to the intrusion detection sample set according to the values ​​of the intrusion detection sample set under various preset indicators, and adjust the intrusion detection sample set when the evaluation value is lower than a preset threshold.

[0257] During implementation, the attack unit 901 attacks the device to be detected by injecting traffic into the device to be detected. The attack unit can be a sending unit, a transmitter, an output interface, a pin, or a circuit when it comes to traffic. When the data processing device 900 includes a storage unit, the storage unit is used to store computer instructions. The attack unit 901, the perturbation unit 902, the construction unit 903, the feature extraction unit 904, the format conversion unit 905, and the adjustment unit 906 are respectively connected to the storage unit for communication and respectively execute the computer instructions stored in the storage unit, so that the data processing device 900 can be used to execute the method executed by the data processing device in any of the above embodiments. Among them, the attack unit 901, the perturbation unit 902, the construction unit 903, the feature extraction unit 904, the format conversion unit 905, and the adjustment unit 906 can be a general-purpose central processing unit (CPU), a microprocessor, or an application-specific integrated circuit (ASIC). The storage unit is a storage unit within the chip, such as a register, cache, etc. The storage unit can also be a storage unit within the data processing device 900 located outside the chip, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0258] For the concepts, explanations, detailed descriptions and other steps involved in the data processing device 900 and related to the technical solutions provided in the embodiments of the present application, please refer to the descriptions of these contents in the aforementioned methods or other embodiments, which will not be repeated here.

[0259] It should be understood that the above division of the units of data processing device 900 is merely a division of logical functions. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. In the embodiment of the present application, the attack unit 901, the perturbation unit 902, the construction unit 903, the feature extraction unit 904, the format conversion unit 905, and the adjustment unit 906 can be implemented by the processor 801 of Figure 8 above.

[0260] According to the data processing method provided in the embodiments of the present application, the present application also provides a data processing device, which includes a processor, the processor is connected to a memory, the memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory, so that the data processing device implements the method described in any one of the embodiments in Figures 3, 4 or 6.

[0261] According to the data processing method provided in the embodiments of the present application, the present application also provides a data processing device, which includes a processor and a memory, the memory is used to store computer program instructions, and the processor is used to run the computer program instructions to implement the method described in any one of the embodiments in Figures 3, 4 or 6.

[0262] According to the data processing method provided in the embodiments of the present application, the present application also provides a chip, which may include a processor and an interface, and the processor is used to read instructions through the interface to execute the method described in any embodiment of Figure 3, Figure 4 or Figure 6.

[0263] According to the data processing method provided in the embodiment of the present application, the present application also provides a data processing system, which may include the aforementioned device to be detected and a data processing device.

[0264] According to the data processing method provided in the embodiments of the present application, the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed, the method described in any one of the embodiments in Figures 3, 4 or 6 is implemented.

[0265] According to the data processing method provided in the embodiments of the present application, the present application also provides a computer program product, which, when running on a processor, implements the method described in any one of the embodiments in FIG3 , FIG4 or FIG6 .

[0266] As used in this specification, the terms "component," "unit," "system," and the like are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and a computing device can be a component. One or more components can reside in a process and / or an execution thread, and a component can be located on one computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).

[0267] Those skilled in the art will appreciate that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0268] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0269] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0270] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0271] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0272] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0273] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A data processing method, characterized in that: include: By attacking the device to be detected, a first type of attack sample is obtained; Obtaining a second type of attack sample by applying noise to the first type of attack sample; An intrusion detection sample set is constructed based on the first type attack samples and the second type attack samples.

2. The method according to claim 1, characterized in that The first type of attack samples includes real attack samples and simulated attack samples. The real attack samples are obtained by artificially attacking the device to be detected, and the simulated attack samples are obtained by attacking the device to be detected with an attack tool.

3. The method according to claim 2, characterized in that The real attack samples correspond to one or more of the following attack types: Identity identification number ID does not exist attack, replay attack, tampering attack, data length error attack, signal out of defined range attack, context error attack, ID source non-specified electronic control unit ECU attack, identical ID attack, controller area network CAN scanning attack, unified diagnostic service UDS execution sensitive operation attack, message authentication error attack, ECU identity spoofing attack, man-in-the-middle attack, ECU authentication error attack, brute force attack, application layer protocol error attack, unknown stack connection attack, unknown stack connection attack.

4. The method according to claim 2 or 3, characterized in that The simulated attack sample corresponds to one or more of the following attack types: ID fuzzy fuzz attack, data fuzz attack, CAN denial of service DoS attack, Ethernet DoS attack, malformed packet injection attack, port scanning attack.

5. The method according to any one of claims 1 to 4, characterized in that The step of obtaining a first type of attack sample by attacking the device to be detected includes: Each of the preset multiple attack types is traversed, and when traversing each of the attack types: Executing an attack behavior corresponding to the attack type on the device to be detected; Acquire the traffic data generated by the device to be detected in response to the attack behavior; If the traffic data is attack traffic, the traffic data is marked as a first-type attack sample.

6. The method according to claim 5, characterized in that After acquiring the flow data generated by the device to be detected for the attack behavior, the method further includes: If the traffic data is normal traffic, marking the traffic data as a non-attack sample; The step of constructing an intrusion detection sample set according to the first type of attack samples and the second type of attack samples includes: The intrusion detection sample set is constructed according to the first type of attack samples, the second type of attack samples and the non-attack samples.

7. The method according to any one of claims 1 to 6, characterized in that The step of applying noise to the first type attack sample to obtain the second type attack sample comprises: Applying noise to the first type of attack sample to obtain a disturbance sample; Inputting the disturbance sample into an attack recognition model to obtain a recognition result output by the attack recognition model, wherein the recognition result is used to indicate whether the disturbance sample is an attack sample; The disturbance sample is adjusted according to the recognition result until the recognition result corresponding to the adjusted disturbance sample indicates that it is impossible to determine whether the adjusted disturbance sample is the attack sample, and the adjusted disturbance sample is determined to be This is the second type of attack sample.

8. The method according to any one of claims 1 to 7, characterized in that After constructing the intrusion detection sample set according to the first type attack sample and the second type attack sample, the method further includes: Extracting features from the intrusion detection sample set to obtain an offline detection sample set; The offline detection sample set is used to evaluate the detection performance of the intrusion detection model.

9. The method according to claim 8, characterized in that The step of extracting features from the intrusion detection sample set to obtain an offline detection sample set includes: Determining the message type of each intrusion detection sample in the intrusion detection sample set; For each intrusion detection sample whose message type is a transmission control protocol TCP message, an offline detection sample is obtained by extracting features from all intrusion detection samples belonging to the same TCP connection; For each intrusion detection sample whose message type is a CAN (FD) message or a UDP message, one offline detection sample is obtained by performing feature extraction on the intrusion detection sample of each CAN (FD) message or UDP message.

10. The method according to claim 8 or 9, characterized in that The features extracted from the feature extraction include one or more of the following features: Timestamp, frequency characteristics, protocol type, content characteristics, packet loss rate, number of error packets, connection duration, connection initiator, and connection receiver.

11. The method according to any one of claims 1 to 10, characterized in that After constructing the intrusion detection sample set according to the first type attack sample and the second type attack sample, the method further includes: Converting the format of the intrusion detection sample set to obtain an online detection sample set that matches the format of the test tool; Inputting the online detection sample set into the device to be detected through the testing tool; The online detection sample set is used to evaluate the detection performance of the device to be detected on which the intrusion detection model is deployed.

12. The method according to any one of claims 1 to 11, characterized in that After constructing the intrusion detection sample set according to the first type attack sample and the second type attack sample, the method further includes: According to the values ​​of the intrusion detection sample set under various preset indicators, an evaluation value corresponding to the intrusion detection sample set is determined, and when the evaluation value is lower than a preset threshold, the intrusion detection sample set is adjusted.

13. The method according to claim 12, characterized in that The preset indicators include one or more of the following indicators: Data redundancy index, attack coverage index, protocol coverage index, business coverage index, data labeling index, balance index, feature independence index, and ease of use index.

14. The method according to any one of claims 1 to 13, characterized in that The first type of attack sample is obtained by attacking any of the following areas: The entire device to be tested; One or more physical areas of the device to be detected; or, One or more functional areas of the device to be detected.

15. A data processing device, characterized in that: include: An attack unit, used to obtain a first type of attack sample by attacking a device to be detected; a perturbation unit, configured to obtain a second type of attack sample by applying noise to the first type of attack sample; A construction unit is used to construct an intrusion Detection sample set.

16. The device according to claim 15, characterized in that The first type of attack samples includes real attack samples and simulated attack samples. The real attack samples are obtained by artificially attacking the device to be detected, and the simulated attack samples are obtained by attacking the device to be detected with an attack tool.

17. The device according to claim 16, characterized in that The real attack samples correspond to one or more of the following attack types: Identity ID does not exist attack, replay attack, tampering attack, data length error attack, signal out of defined range attack, context error attack, ID source is not specified electronic control unit ECU attack, same ID attack, controller area network CAN scanning attack, unified diagnostic service UDS execution sensitive operation attack, message authentication error attack, ECU identity spoofing attack, man-in-the-middle attack, ECU authentication error attack, brute force attack, application layer protocol error attack, unknown stack connection attack, unknown stack connection attack.

18. The device according to claim 16 or 17, characterized in that The simulated attack sample corresponds to one or more of the following attack types: ID fuzzy fuzz attack, data fuzz attack, CAN denial of service DoS attack, Ethernet DoS attack, malformed packet injection attack, port scanning attack.

19. The device according to any one of claims 15 to 18, characterized in that The attack unit is specifically used for: Each of the preset multiple attack types is traversed, and when traversing each of the attack types: Executing an attack behavior corresponding to the attack type on the device to be detected; Acquire the traffic data generated by the device to be detected in response to the attack behavior; If the traffic data is attack traffic, the traffic data is marked as a first-type attack sample.

20. The device according to claim 19, characterized in that After acquiring the flow data generated by the device to be detected for the attack behavior, the attack unit is further configured to: if the flow data is normal flow, mark the flow data as a non-attack sample; The construction unit is specifically used to construct the intrusion detection sample set according to the first type of attack samples, the second type of attack samples and the non-attack samples.

21. The device according to any one of claims 15 to 20, characterized in that The disturbance unit is specifically used for: Applying noise to the first type of attack sample to obtain a disturbance sample; Inputting the disturbance sample into an attack recognition model to obtain a recognition result output by the attack recognition model, wherein the recognition result is used to indicate whether the disturbance sample is an attack sample; The disturbance sample is adjusted according to the recognition result until the recognition result corresponding to the adjusted disturbance sample indicates that it is impossible to determine whether the disturbance sample is the attack sample, and the adjusted disturbance sample is determined as the second type attack sample.

22. The device according to any one of claims 15 to 21, characterized in that Also includes a feature extraction unit; The feature extraction unit is used for: Extracting features from the intrusion detection sample set to obtain an offline detection sample set; The offline detection sample set is used to evaluate the detection effect of the intrusion detection model.

23. The device according to claim 22, characterized in that The feature extraction unit is specifically used for: Determining the message type of each intrusion detection sample in the intrusion detection sample set; For each intrusion detection sample whose message type is a transmission control protocol TCP message, an offline detection sample is obtained by extracting features from all intrusion detection samples belonging to the same TCP connection; For each intrusion detection sample whose message type is a CAN (FD) message or a UDP message, one offline detection sample is obtained by performing feature extraction on the intrusion detection sample of each CAN (FD) message or UDP message.

24. The device according to claim 22 or 23, characterized in that The features extracted from the feature extraction include one or more of the following features: Timestamp, frequency characteristics, protocol type, content characteristics, packet loss rate, number of error packets, connection duration, connection initiator, and connection receiver.

25. The device according to any one of claims 15 to 24, characterized in that Also includes a format conversion unit; The format conversion unit is used for: Converting the format of the intrusion detection sample set to obtain an online detection sample set that matches the format of the test tool; Inputting the online detection sample set into the device to be detected through the testing tool; The online detection sample set is used to evaluate the detection performance of the device to be detected on which the intrusion detection model is deployed.

26. The device according to any one of claims 15 to 25, characterized in that Also includes a regulating unit; The regulating unit is used for: According to the values ​​of the intrusion detection sample set under various preset indicators, an evaluation value corresponding to the intrusion detection sample set is determined, and when the evaluation value is lower than a preset threshold, the intrusion detection sample set is adjusted.

27. The device according to claim 26, characterized in that The preset indicators include one or more of the following indicators: Data redundancy index, attack coverage index, protocol coverage index, business coverage index, data labeling index, balance index, feature independence index, and ease of use index.

28. The device according to any one of claims 15 to 27, characterized in that The first type of attack sample is obtained by attacking any of the following areas: The entire device to be tested; One or more physical areas of the device to be detected; or, One or more functional areas of the device to be detected.

29. A data processing device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer program instructions, and the processor executes the computer program instructions to implement the method according to any one of claims 1 to 14.

30. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 14 is implemented.

31. A computer program product, characterized in that When the computer program product is run on a processor, the method according to any one of claims 1 to 14 is implemented.