Data processing method and device and computer equipment

Through the SBC algorithm and SMOTE technology of multi-layer binary classifier, the accuracy and efficiency problems of existing IDS when identifying packet categories are solved, and fast and accurate packet type recognition is achieved, reducing the false positive rate.

CN120579015APending Publication Date: 2025-09-02CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510620621.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

When identifying data packet categories, existing intrusion detection systems find it difficult to quickly and accurately determine unseen signatures or abnormal patterns, resulting in high false alarm rates and low recognition efficiency.

Method used

The SBC algorithm of multi-layer binary classifier is used to identify the five-tuple of data packets layer by layer, use the trained N-layer target classifier to identify the data packet category, and use the synthetic minority class oversampling technology (SMOTE) to deal with the data set imbalance problem to ensure the accuracy and efficiency of the classifier.

Benefits of technology

Improves the accuracy and efficiency of packet type identification, can quickly identify unseen signatures or exception patterns, and reduces the false positive rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579015A_ABST
    Figure CN120579015A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device and computer equipment, and the method comprises the steps: obtaining a target data packet, determining a quintuple of the target data packet, and recognizing the type of the target data packet through employing a target SBC algorithm according to the quintuple of the target data packet, and the target SBC algorithm comprises N layers of target classifiers, one output of the ith layer of target classifier in the N layers of target classifiers is the ith category, the other output of the ith layer of target classifier is the input of the (i + 1) th layer of target classifier, i = 1,..., N-1, and N is an integer greater than 1. By adopting the method, the category of the data packet can be quickly and accurately identified by using the SBC algorithm with the multi-layer binary classifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to a data processing method, apparatus, and computer equipment. Background Art

[0002] With the widespread adoption of fifth-generation mobile communication technology (5G), networks are becoming increasingly vulnerable to new vulnerabilities, making intrusion detection systems (IDS) increasingly important. IDS monitors transmitted data packets in real time and can identify potential attacks based on packet classification. Therefore, quickly and accurately determining the packet classification is crucial. Summary of the Invention

[0003] Based on this, it is necessary to provide a data processing method, device and computer equipment to address the above technical problems, which can use the SBC algorithm with a multi-layer binary classifier to quickly and accurately identify the category of the data packet.

[0004] In a first aspect, the present application provides a data processing method, comprising:

[0005] Get the target data packet;

[0006] Determining a quintuple of the target data packet;

[0007] According to the quintuple, a target binary classification SBC algorithm is used to identify the category of the target data packet, wherein the target SBC algorithm includes N layers of target classifiers, wherein an output of the i-th layer target classifier in the N layers of target classifiers is the i-th category, and another output of the i-th layer target classifier is the input of the i+1-th layer target classifier, where i=1,…,N-1, and N is an integer greater than 1.

[0008] In one embodiment, the method further comprises:

[0009] Acquire training data comprising N data sets, where each of the N data sets comprises data packets of one category;

[0010] Arrange the N data sets in descending order according to the number of data packets included to obtain a data set list;

[0011] Mark the kth data set in the data set list as the kth category, k=1,…,N;

[0012] According to the data set list and the corresponding categories, the N layers of initial classifiers in the initial SBC algorithm are trained in sequence to obtain the target SBC algorithm.

[0013] In one embodiment, the training of N layers of initial classifiers in the initial SBC algorithm in sequence according to the dataset list and the corresponding categories to obtain the target SBC algorithm includes:

[0014] According to the N data sets and corresponding categories in the data set list, the first layer initial classifier in the initial SBC algorithm is trained to obtain a first SBC algorithm;

[0015] According to the jth to Nth datasets in the dataset list and the corresponding categories, the jth layer initial classifier in the j-1th SBC algorithm is trained to obtain the jth SBC algorithm, where j=2, ..., N-1;

[0016] According to the Nth data set and the target data set in the data set list and the corresponding categories, the Nth layer initial classifier in the N-1th SBC algorithm is trained to obtain the target SBC algorithm, and the target data set is one or more data sets in the N data sets except the Nth data set.

[0017] In one embodiment, the training of the first layer initial classifier in the initial SBC algorithm according to the N data sets and corresponding categories in the data set list to obtain the first SBC algorithm includes:

[0018] The first data set in the data set list is used as a positive sample, and the second to Nth data sets in the data set list are used as negative samples, and the first layer initial classifier in the initial SBC algorithm is trained to obtain a first SBC algorithm.

[0019] In one embodiment, the training of the j-th layer initial classifier in the j-1 SBC algorithm based on the j-th to N-th datasets and the corresponding categories in the dataset list to obtain the j-th SBC algorithm includes:

[0020] Performing data expansion on the jth data set in the data set list, wherein the absolute value of the difference between the number of data packets included in the jth data set after expansion and the first data set in the data set list is less than or equal to a threshold;

[0021] According to the expanded jth data set and the corresponding category, as well as the j+1th to Nth data sets in the data set list and the corresponding categories, the jth layer initial classifier in the j-1th SBC algorithm is trained to obtain the jth SBC algorithm.

[0022] In one embodiment, the data expansion of the jth data set in the data set list includes:

[0023] A synthetic minority oversampling technique (SMOTE) algorithm may be used to perform data expansion on the j-th data set in the data set list.

[0024] In a second aspect, the present application also provides a data processing method, comprising:

[0025] Acquire training data comprising N data sets, each of the N data sets comprising data packets of one category, where N is an integer greater than 1;

[0026] Arrange the N data sets in descending order according to the number of data packets included to obtain a data set list;

[0027] Mark the kth data set in the data set list as the kth category, k=1,…,N;

[0028] According to the data set list and the corresponding categories, the N layers of initial classifiers in the initial SBC algorithm are trained in sequence to obtain a target SBC algorithm including N layers of target classifiers, where an output of the i-th layer classifier in the N-layer classifiers is the i-th category, and another output of the i-th layer classifier is the input of the i+1-th layer classifier, where i=1,…,N-1.

[0029] In a third aspect, the present application further provides a data processing device, comprising:

[0030] An acquisition unit, configured to acquire a target data packet;

[0031] a determining unit, configured to determine a quintuple of the target data packet;

[0032] An identification unit is used to identify the category of the target data packet using a target binary classification (SBC) algorithm based on the quintuple, wherein the target SBC algorithm includes N layers of target classifiers, wherein an output of an i-th layer target classifier in the N layers of target classifiers is the i-th category, and another output of the i-th layer target classifier is an input of an i+1-th layer target classifier, where i=1,…,N-1, and N is an integer greater than 1.

[0033] In a fourth aspect, the present application further provides a data processing device, comprising:

[0034] an acquisition unit, configured to acquire training data comprising N data sets, each of the N data sets comprising a data packet of a category, where N is an integer greater than 1;

[0035] an arranging unit, configured to arrange the N data sets in descending order according to the number of data packets included therein, to obtain a data set list;

[0036] a marking unit, configured to mark the kth data set in the data set list as the kth category, where k=1, ..., N;

[0037] A training unit is used to train N layers of initial classifiers in the initial SBC algorithm in sequence according to the data set list and the corresponding categories to obtain a target SBC algorithm including N layers of target classifiers, where an output of the i-th layer classifier in the N layers of classifiers is the i-th category, and another output of the i-th layer classifier is the input of the i+1-th layer classifier, where i=1,…,N-1.

[0038] In a fifth aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above methods when executing the computer program.

[0039] In a sixth aspect, the present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above methods are implemented.

[0040] In a seventh aspect, the present application also provides a computer program product, comprising a computer program, which implements the steps of the above methods when executed by a processor.

[0041] In this application, a target data packet is obtained, a quintuple of the target data packet is determined, and based on the quintuple of the target data packet, a target SBC algorithm is used to identify the category of the target data packet. The target SBC algorithm includes N layers of target classifiers, wherein one output of the target classifier of the i-th layer in the N layers is the i-th category, and another output of the target classifier of the i-th layer is the input of the target classifier of the i+1-th layer, where i=1,…,N-1, and N is an integer greater than 1. It can be seen that by using an SBC algorithm with multiple layers of binary classifiers to identify the category of a data packet, multiple binary classifiers can be used to accurately identify the category of the data packet. If the binary classifier of the previous layer fails to identify the category of the data packet, it can be automatically or directly passed to the binary classifier of the next layer for identification, thereby improving the efficiency of data packet type identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 This is a schematic diagram of a network architecture disclosed in an embodiment of the present application;

[0044] Figure 2 This is a flow chart of a data processing method provided in an embodiment of the present application;

[0045] Figure 3 This is a schematic diagram of the structure of a target SBC algorithm provided in an embodiment of the present application;

[0046] Figure 4 This is a flow chart of another data processing method provided in an embodiment of the present application;

[0047] Figure 5 This is a flow chart of another data processing method provided in an embodiment of the present application;

[0048] Figure 6 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;

[0049] Figure 7 A schematic diagram of the structure of another data processing device provided in an embodiment of the present application;

[0050] Figure 8 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0052] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0053] The present application provides a data processing method, apparatus, and computer equipment, which can use an SBC algorithm with a multi-layer binary classifier to quickly and accurately identify the category of a data packet.

[0054] In order to better understand the embodiments of the present application, the network architecture of the present application is first described below.

[0055] Figure 1 This is a network architecture diagram disclosed in the embodiment of this application. Figure 1As shown, the network architecture may include passive Internet of Things (AIoT) devices, User Equipment (UE) readers, Radio Access Network (RAN) devices, User Plane Function (UPF) network elements, Access and Mobility Management Function (AMF) network elements, Charging Function (CHF) network elements, Application Function (AF) network elements, Policy Control Function (PCF) network elements, Session Management Function (SMF) network elements, Unified Data Management (UDM) network elements, Network Data Analytics Function (NWDAF) network elements, Network Repository Function (NRF) network elements, Network Exposure Function (NEF) network elements, and Authentication Server Function (AUSF) network elements.

[0056] The UE Reader can communicate directly with the RAN device. The UE Reader can communicate with the AMF network element through the N1 interface. The RAN device can communicate with the AMF network element through the N2 interface. The RAN device can communicate with the UPF network element through the N3 interface. The UPF network element can communicate with the SMF network element through the N4 interface. The AMF network element can provide the service interface Namf, the CHF network element can provide the service interface Nchf, the AF network element can provide the service interface Naf, the PCF network element can provide the service interface Npcf, the SMF network element can provide the service interface Nsmf, the UDM network element can provide the service interface Nudm, the NWDAF network element can provide the service interface Nnwdaf, the NRF network element can provide the service interface Nnrf, the NEF network element can provide the service interface Nnef, and the AUSF network element can provide the service interface Nausf. AMF network elements, CHF network elements, AF network elements, PCF network elements, SMF network elements, UDM network elements, NWDAF network elements, NRF network elements, NEF network elements and AUSF network elements can communicate through service-oriented interfaces.

[0057] The UE Reader is a UE with reader functionality. It is responsible for communication between AIoT devices and AF network elements and can report UE capability information to the UDM network element. A UE, also known as a terminal device, mobile station (MS), or mobile terminal (MT), is a device that provides voice and / or data connectivity to users. The terminal device may be a handheld terminal, a laptop computer, a subscriber unit, a cellular phone, a smart phone, a wireless data card, a personal digital assistant (PDA) computer, a tablet computer, a wireless modem, a handheld device, a laptop computer, a cordless phone or a wireless local loop (WLL) station, a machine type communication (MTC) terminal, a wearable device (such as a smart watch, a smart bracelet, a pedometer, etc.), an in-vehicle device (such as a car, a bicycle, an electric car, an airplane, a ship, a train, a high-speed train, etc.), a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a smart home device (such as a refrigerator, a television, an air conditioner, an electric meter, etc.), an intelligent robot, a workshop device, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a smart grid, etc. Wireless terminals in the Internet of Things (IoT) grid, transportation safety, smart cities, smart homes, flying devices (such as smart robots, hot air balloons, drones, airplanes, etc.), or other devices that can access the Internet. Figure 1 The terminal device is shown as UE Reader, which is only an example and does not limit the terminal device.

[0058] RAN equipment provides wireless access for UE readers and is primarily responsible for air interface functions such as radio resource management, Quality of Service (QoS) flow management, data compression, and encryption. RAN equipment can include various base stations, such as macro base stations, micro base stations (also known as small cells), relay stations, and access points. RAN equipment can also include Wireless Fidelity (WiFi) access points (APs) and Worldwide Interoperability for Microwave Access (WiMax) base stations (BSs).

[0059] The UPF network element is mainly responsible for user plane related content, such as data packet routing and transmission, mobility anchor point, uplink classifier to support routing service flows to the data network, branch point to support multi-homing Protocol Data Unit (PDU) sessions, packet inspection, service usage reporting, QoS processing, lawful interception, downlink packet storage, etc.

[0060] AMF network elements can be responsible for processing AIoT business logic, tag registration, connection management, registration process, mobility management, access authentication and authorization management, reachability management, security context management, SMF network element selection and other access and mobility related functions.

[0061] The CHF network element may be responsible for billing-related content.

[0062] AF network elements primarily support interaction with the 3rd Generation Partnership Project (3GPP) core network to provide services or services, influence traffic routing, access network capability exposure, policy control, etc., and can interact with NEF network elements. AF network elements can be responsible for sending service requirements to AIoT devices.

[0063] The PCF network element is responsible for unified policy formulation, policy control, and other policy-related functions such as obtaining contract information related to policy decisions from the Unified Data Repository (UDR) network element. It also receives data flow policies sent by the intelligent module. Policy control can include service data flow and application detection, gating, QoS, and flow-based charging control.

[0064] The SMF network element is primarily responsible for session management in mobile networks, selection and control of UPF network elements, service and session continuity (SSC) mode selection, roaming, and other session-related functions. Session management can include session establishment, modification, release, and update. Session management can also include tunnel maintenance between UPF network elements and access network (AN) equipment.

[0065] The UDM network element is responsible for tag contract data management and security management, UE Reader authorization information, generating UE Reader capability lists, and storing UE Reader capability lists.

[0066] The NWDAF network element provides network data collection and analysis capabilities based on technologies such as big data and artificial intelligence. It can provide network analysis services based on network service request data and provide data flow strategies to the PCF network element.

[0067] The NRF network element is mainly responsible for service discovery and maintaining the NF context of available Network Function (NF) instances and the services they support.

[0068] The NEF network element, located between the core network and external third-party application functions (and possibly some internal AF network elements), is primarily responsible for securely exposing the services and capabilities provided by 3GPP network functions, either internally or to third parties. The NEF network element provides security guarantees to ensure the security of external applications accessing the 3GPP network, including the exposure of external application QoS customization capabilities, mobility status event subscription, and AF request distribution.

[0069] The AUSF network element can be responsible for authenticating and authorizing the access of the UE Reader, such as generating intermediate keys.

[0070] The above network elements in the core network can also be called functional entities, which can be network elements implemented on dedicated hardware, software instances running on dedicated hardware, or instances of virtualized functions on an appropriate platform. For example, the above virtualization platform can be a cloud platform.

[0071] It should be noted that Figure 1 The network architecture shown is not limited to including only the network elements and devices shown in the figure, but may also include other network elements or devices not shown in the figure, which will not be listed one by one in the present invention.

[0072] It should be noted that the embodiment of the present invention does not limit the distribution form of each network element in the core network. Figure 1The distribution form shown is only exemplary and is not limiting to the present invention.

[0073] It should be understood that the names of all network elements in this disclosure are merely examples. In future communications, such as the 6th Generation Mobile Networks (6G), they may be referred to by other names. Alternatively, in future communications, such as 6G, the network elements described in this disclosure may be replaced by other entities or devices with the same functions. This disclosure is not intended to limit these terms. This is a general description and will not be further elaborated upon.

[0074] It should be noted that Figure 1 The fifth generation mobile communication technology (5G) network architecture shown does not limit the 5G network. Optionally, the method of the embodiment of the present invention is also applicable to various future communication systems, such as 6G or other communication networks.

[0075] In order to better understand the embodiments of the present application, the relevant technologies are described below.

[0076] The proliferation of core network devices has significantly increased network vulnerability, creating an urgent need for effective IDS. Signature-Based IDS (SIDS) can accurately identify attack patterns based on a database of previously identified intrusion "signatures," but they struggle to generalize to unseen signatures. Anomaly-Based IDS (AIDS) can identify possible intrusions based on deviations from an established baseline of normal network activity, but they suffer from high false positive rates. Machine learning (ML)-based IDS are generally more robust and reliable, but IDS datasets are often unbalanced, potentially leading to overfitting or data that does not represent the original distribution, reducing the accuracy of packet data processing.

[0077] Based on the above network architecture, Figure 2 1 is a flow chart of a data processing method provided by an embodiment of the present application. The data processing method can be applied to the above-mentioned NWDAF network element, and can also be applied to other network elements or devices. Figure 2 As shown, the data processing method may include the following steps.

[0078] 201. Obtain target data packet.

[0079] The target data packet is any data packet transmitted or any data packet received.

[0080] When there is data packet transmission, the target data packet can be obtained.

[0081] 202. Determine the quintuple of the target data packet.

[0082] The five-tuple of the target data packet can be determined. The five-tuple of the target data packet may include the source Internet Protocol (IP) address, destination IP address, source port number, destination port number, and transport layer protocol of the target data packet.

[0083] The source IP address of the target data packet is the IP address of the device that sends the target data packet. The destination IP address of the target data packet is the IP address of the device that ultimately receives the target data packet. The source port number of the target data packet is the port number of the device that sends the target data packet. The destination port number of the target data packet is the port number of the device that ultimately receives the target data packet. The transport layer protocol of the target data packet is the protocol used by the target data packet. The protocol can be Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Internet Control Message Protocol (ICMP), etc.

[0084] 203. According to the quintuple of the target data packet, a target binary classification (SBC) algorithm is used to identify the category of the target data packet.

[0085] Figure 3 This is a schematic diagram of the structure of a target SBC algorithm provided in an embodiment of the present application. Figure 3 As shown, the target SBC algorithm includes N layers of target classifiers, one output of the i-th target classifier in the N layers is the i-th category, and another output of the i-th target classifier in the N layers is the input of the i+1-th target classifier, i=1,…,N-1, and N is an integer greater than 1.

[0086] The target SBC algorithm is a pre-trained SBC algorithm, that is, the N-layer target classifiers included in the target SBC algorithm are pre-trained classifiers, and the N-layer target classifiers are all binary classifiers.

[0087] Because the quintuples of the data packets are different, the categories of the data packets are different. Therefore, the target SBC algorithm can be used to identify the category of the target data packet based on the quintuple of the target data packet. The difference in the quintuple of the data packet can be a difference in one or more tuples of the quintuple of the data packet, that is, a difference in one or more of the source IP address, destination IP address, source port number, destination port number, and transport layer protocol of the data packet.

[0088] The target data packet and the five-tuple of the target data packet can be input into the target SBC algorithm, that is, the first-layer target classifier. When the first-layer target classifier determines that the target data packet is of the first category according to the five-tuple of the target data packet, the first-layer target classifier outputs the target data packet as the first category. When the first-layer target classifier determines that the target data packet is not of the first category according to the five-tuple of the target data packet, the first-layer target classifier can input the target data packet and the five-tuple of the target data packet into the second-layer target classifier. When the second-layer target classifier determines that the target data packet is of the second category according to the five-tuple of the target data packet, the second-layer target classifier outputs the target data packet as the first category. The classifier outputs the target data packet as the second category. When the second-layer target classifier determines that the target data packet is not the second category based on the quintuple of the target data packet, the second-layer target classifier can input the target data packet and the quintuple of the target data packet into the third-layer target classifier. And so on. When the N-layer target classifier determines that the target data packet is the Nth category based on the quintuple of the target data packet, the N-layer target classifier outputs the target data packet as the Nth category. When the N-layer target classifier determines that the target data packet is not the Nth category based on the quintuple of the target data packet, the N-layer target classifier outputs the target data packet as other categories.

[0089] exist Figure 2 In the data processing method shown, an SBC algorithm with multiple layers of binary classifiers is used to identify the category of a data packet. Multiple binary classifiers can be used to accurately identify the category of a data packet. If the binary classifier in the previous layer fails to identify the category of the data packet, it can be automatically or directly passed to the binary classifier in the next layer for identification, which can improve the efficiency of data packet type identification.

[0090] Based on the above network architecture, Figure 4 1 is a flow chart of another data processing method provided by an embodiment of the present application. The data processing method can be applied to the above-mentioned NWDAF network element, and can also be applied to other network elements or devices. Figure 3 As shown, the data processing method may include the following steps.

[0091] 401. Obtain training data including N data sets.

[0092] When training the initial SBC algorithm is required, training data can be obtained first. This training data can be obtained locally or from other network elements, devices, or servers. Training data is the data used to train the initial SBC algorithm.

[0093] The training data may include N data sets. Each of the N data sets may include data packets of one category. Different data sets may include data packets of different categories. The N data sets may also include quintuples of data packets.

[0094] It can be seen that the number of data sets included in the training data is the same as the number of layers or the number of initial classifiers included in the initial SBC algorithm to be trained, which can ensure that each initial classifier has a data set of the corresponding category, the validity of the training data, and the accuracy of the trained classifier.

[0095] 402. Arrange the N data sets in descending order according to the number of data packets included to obtain a data set list.

[0096] After the training data is obtained, the N data sets may be arranged in descending order according to the number of data packets included to obtain a data set list, ie, a descending order data set list.

[0097] 403. Mark the kth dataset in the dataset list as the kth category.

[0098] k=1,…,N.

[0099] The k-th dataset in the dataset list can be marked as the k-th category, that is, each dataset in the dataset list is labeled according to the order of arrangement, and the label is the category of the data packets included in each dataset.

[0100] It can be seen that the categories are arranged in descending order according to the number of corresponding data packets, that is, in descending order according to the frequency of occurrence.

[0101] 404. According to the data set list and the corresponding categories, the N layers of initial classifiers in the initial SBC algorithm are trained in sequence to obtain a target SBC algorithm.

[0102] The initial SBC algorithm is an untrained SBC algorithm, and the target SBC algorithm is an SBC algorithm obtained by training the initial SBC algorithm.

[0103] The initial SBC algorithm includes N layers of initial classifiers, where one output of the i-th layer of initial classifiers is the i-th category, and another output of the i-th layer of initial classifiers is the input of the i+1-th layer of initial classifiers, where i=1,…,N-1, and N is an integer greater than 1. The structure of the initial SBC algorithm is similar to Figure 3 The structure of the target classifier shown is similar, except that the classifier in the target SBC algorithm is the target classifier, which is the classifier after the initial classifier is trained.

[0104] According to the data set list and the corresponding categories, the N layers of initial classifiers in the initial SBC algorithm can be trained in sequence to obtain the target SBC algorithm.

[0105] It can be seen that when training the N layers of initial classifiers in the initial SBC algorithm, the initial classifiers are trained layer by layer, that is, the first layer of initial classifiers can be trained to obtain the first layer of target classifiers, and then the second layer of initial classifiers can be trained to obtain the second layer of target classifiers, and so on, until the N layer of initial classifiers are trained to obtain the N layer of target classifiers.

[0106] It should be understood that the classifier involved in this application is a binary classifier.

[0107] In some embodiments, step 404 may include:

[0108] 4041. According to the N data sets in the data set list and the corresponding categories, the first layer initial classifier in the initial SBC algorithm is trained to obtain a first SBC algorithm.

[0109] The first layer initial classifier in the initial SBC algorithm can be trained according to the N data sets in the data set list and the corresponding (or labeled) categories to obtain the first SBC algorithm.

[0110] The dataset whose marked category is category 1 among the N datasets can be determined as a positive sample, and the dataset whose marked category is not category 1 among the N datasets can be determined as a negative sample. Then, the first layer initial classifier in the initial SBC algorithm can be trained based on the positive samples and the negative samples to obtain the first SBC algorithm.

[0111] The first dataset in the dataset list can be used as a positive sample, and the second to Nth datasets in the dataset list can be used as negative samples to train the first-layer initial classifier in the initial SBC algorithm to obtain a first SBC algorithm. The first-layer classifier in the N-layer classifier included in the first SBC algorithm is the target classifier, and the classifiers in the other layers are initial classifiers.

[0112] The data packets included in the N data sets can be input into the first-layer initial classifier respectively, and the loss value of the loss function can be determined according to the classification result of the first-layer initial classifier and the category of the data packet label, and then the parameters of the first-layer initial classifier can be optimized according to the loss value.

[0113] The loss function can be a cross entropy loss function, a hinge loss function, a focal loss function, or any other loss function that can be used for a binary classifier.

[0114] The parameters of the initial classifier in the first layer can be optimized through algorithms such as gradient descent based on the loss value.

[0115] 4042. According to the j-th dataset to the N-th dataset in the dataset list and the corresponding categories, the j-th layer initial classifier in the j-1-th SBC algorithm is trained to obtain the j-th SBC algorithm, where j=2, ..., N-1.

[0116] After the first-layer initial classifier is trained as the first-layer target classifier, the second-layer initial classifier, ..., the Nth-layer initial classifier can be trained in sequence.

[0117] The j-th layer initial classifier in the j-1 SBC algorithm can be trained according to the j-th to N-th datasets in the dataset list and the corresponding categories to obtain the j-th SBC algorithm.

[0118] For example, assuming that N is 4, the second-layer initial classifier in the first SBC algorithm can be trained according to the second to fourth data sets in the data set list and the corresponding categories to obtain the second SBC algorithm. The third-layer initial classifier in the second SBC algorithm can be trained according to the third to fourth data sets in the data set list and the corresponding categories to obtain the third SBC algorithm.

[0119] The jth data set in the data set list can be used as a positive sample, and the j+1th to Nth data sets in the data set list can be used as negative samples to train the jth layer initial classifier in the j-1th SBC algorithm to obtain the jth SBC algorithm.

[0120] Because different datasets in the dataset list contain different numbers of data packets, that is, the dataset sizes are unbalanced, this may cause the trained classifier to overfit. To solve this problem, before training the j-th initial classifier, the j-th dataset in the dataset list can be expanded. This can avoid the problem of unbalanced training data and, in turn, the problem of overfitting the trained classifier.

[0121] Data expansion can be performed on the jth dataset in the dataset list. Then, the jth layer initial classifier in the j-1th SBC algorithm can be trained based on the expanded jth dataset and its corresponding category, as well as the j+1th to Nth datasets in the dataset list and their corresponding categories, to obtain the jth SBC algorithm. The absolute value of the difference between the number of data packets included in the expanded jth dataset and the first dataset in the dataset list is less than or equal to a threshold.

[0122] A Synthetic Minority Oversampling Technique (SMOTE) algorithm can be used to perform data augmentation on the j-th dataset in the dataset list. The SMOTE algorithm can be a Border-line SMOTE algorithm, a Support Vector Machine (SVM) SMOTE algorithm, a Geometric SMOTE (G-SMOTE) algorithm, a SMOTE-Edited Nearest Neighbor (ENN) algorithm, or another SMOTE algorithm.

[0123] For example, two closest data packets can be randomly selected from the jth data set in the data set list, and then the midpoint of the Euclidean distance between the two data packets can be calculated to obtain a new data packet. The above process is repeated until the absolute value of the difference between the number of data packets included in the expanded jth data set and the first data set in the data set list is less than or equal to the threshold.

[0124] Since the absolute value of the difference between the number of data packets included in the expanded j-th dataset and the first dataset in the dataset list is less than or equal to the threshold, the balance of the dataset can be guaranteed, thereby avoiding overfitting of the trained classifier.

[0125] Since only the jth data set in the data set list is expanded when training the jth layer initial classifier, and the j+1th to Nth data sets in the data set list are not expanded, the training data can be reduced while ensuring data balance, thereby improving training efficiency.

[0126] 4043. According to the Nth dataset and the target dataset in the dataset list and the corresponding categories, the Nth layer initial classifier in the N-1th SBC algorithm is trained to obtain a target SBC algorithm, where the target dataset is one or more datasets in the N datasets except the Nth dataset.

[0127] If the Nth layer initial classifier is trained in the manner of step 4024, the Nth data set in the data set list can be determined as a positive sample, but there is no negative sample, so that the trained Nth layer classifier cannot classify the Nth category and other categories.

[0128] To address the above issue, the target dataset can be determined as a negative sample. The target dataset is one or more datasets in the N datasets, excluding the Nth dataset. Exemplarily, the target dataset is the first dataset in the dataset list. Exemplarily, the target dataset is the first dataset to the N-1th dataset in the dataset list.

[0129] Data augmentation may be performed on the Nth dataset in the dataset list. An Nth layer initial classifier in the N-1th SBC algorithm may be trained based on the augmented Nth dataset and its corresponding category, as well as the target dataset and its corresponding category, to obtain a target SBC algorithm. The absolute value of the difference between the number of data packets included in the augmented Nth dataset and the number of data packets included in the first dataset in the dataset list is less than or equal to a threshold.

[0130] The Nth layer initial classifier in the N-1th SBC algorithm can be trained with the expanded Nth data set as a positive sample and the target data set as a negative sample to obtain the target SBC algorithm.

[0131] As can be seen, the SBC algorithm uses a divide-and-conquer algorithm, which helps improve the classification of packets corresponding to less frequent categories. By transforming a multi-class problem into a series of progressively smaller binary classification problems, the SBC algorithm can efficiently handle large-scale training data and multi-class situations, thereby improving training efficiency. As the classifier hierarchy deepens, the dataset size continues to shrink, reducing the training data and model size of subsequent classifiers, thereby improving training efficiency.

[0132] 405. Obtain the target data packet.

[0133] The detailed description of step 405 can refer to the description of step 201 and will not be repeated here.

[0134] 406. Determine the quintuple of the target data packet.

[0135] The detailed description of step 406 can refer to the description of step 202 and will not be repeated here.

[0136] 407. Identify the category of the target data packet using the target SBC algorithm according to the quintuple of the target data packet.

[0137] The detailed description of step 407 can refer to the description of step 203 and will not be repeated here.

[0138] It can be seen that the target SBC algorithm follows the order of the hierarchy to identify the packet category. When a packet enters the target SBC algorithm, the packet will be classified by each layer of the target classifier in turn until it is identified as a specific category.

[0139] After the category of the target data packet is identified, it can be determined whether the target data packet is attack data or abnormal data based on the category of the target data packet.

[0140] In some embodiments, when multiple data packets belong to the same category, sequence numbers of the multiple data packets in the data stream may be determined according to the sending time of the multiple data packets, and the multiple data packets include the target data packet.

[0141] Since the categories of data packets in the same data stream are definitely the same, the data packets of the same data stream can be determined by category, and then multiple data packets of the same data stream can be sorted according to the sending time to ensure the correct order of the data packets in the data stream.

[0142] exist Figure 4 In the data processing method shown, the SBC algorithm can be trained first, and then the trained SBC algorithm with multiple layers of binary classifiers can be used to identify the category of the data packet. Multiple binary classifiers can be used to accurately identify the category of the data packet. If the binary classifier in the previous layer fails to identify the category of the data packet, it can be automatically or directly passed to the binary classifier in the next layer for identification, which can improve the efficiency of data packet type identification.

[0143] Based on the above network architecture, Figure 5 1 is a flow chart of a data processing method provided by an embodiment of the present application. The data processing method can be applied to the above-mentioned NWDAF network element, and can also be applied to other network elements or devices. Figure 5 As shown, the data processing method may include the following steps.

[0144] 501. Obtain training data including N data sets.

[0145] The detailed description of step 501 can refer to the description of step 401 and will not be repeated here.

[0146] 502. Arrange the N data sets in descending order according to the number of data packets included, to obtain a data set list.

[0147] The detailed description of step 502 can refer to the description of step 402 and will not be repeated here.

[0148] 503. Mark the kth dataset in the dataset list as the kth category.

[0149] The detailed description of step 503 can refer to the description of step 403 and will not be repeated here.

[0150] 504. According to the data set list and the corresponding categories, the N layers of initial classifiers in the initial SBC algorithm are trained in sequence to obtain a target SBC algorithm including N layers of target classifiers.

[0151] In an N-layer classifier, one output of the i-th layer classifier is the i-th category, and another output of the i-th layer classifier is the input of the i+1-th layer classifier, i=1,…,N-1, and N is an integer greater than 1.

[0152] The initial SBC algorithm includes N layers of initial classifiers, where one output of the i-th layer initial classifier in the N layers is the i-th category, and another output of the i-th layer initial classifier in the N layers is the input of the i+1-th layer initial classifier, i=1,…,N-1.

[0153] The target SBC algorithm includes N layers of target classifiers, where one output of the i-th layer target classifier is the i-th category, and another output of the i-th layer target classifier is the input of the i+1-th layer target classifier, i=1,…,N-1.

[0154] The detailed description of step 504 may refer to the description of step 404 and will not be repeated here.

[0155] exist Figure 5 In the data processing method shown, the N layers of initial classifiers in the initial SBC algorithm can be trained in sequence according to a data set list including N data sets arranged in descending order and the corresponding labeled categories to obtain a target SBC algorithm including N layers of target classifiers. Since the recognition problem of multiple categories is converted into the training of an SBC algorithm with multiple layers of binary classifiers, the training process can be simplified, thereby improving the training efficiency.

[0156] It should be understood that the same or corresponding information in different embodiments may be referenced to each other, and the contents and / or steps in different embodiments may be combined with each other.

[0157] It should be understood that in the above embodiments, some steps may be omitted and some steps may be combined.

[0158] Based on the same inventive concept, embodiments of the present application also provide a data processing device for implementing the aforementioned data processing method. The implementation solution provided by the data processing device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations in one or more data processing device embodiments provided below can be found in the above-mentioned limitations on the data processing method and will not be repeated here.

[0159] Based on the above network structure, Figure 6This is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. The data processing device can be applied to the above-mentioned NWDAF network element, or to other network elements or devices. The data processing device may include:

[0160] An acquisition unit 601 is configured to acquire a target data packet;

[0161] A determining unit 602 is configured to determine a quintuple of a target data packet;

[0162] The identification unit 603 is used to identify the category of the target data packet using the target SBC algorithm based on the quintuple of the target data packet. The target SBC algorithm includes N layers of target classifiers, where one output of the i-th layer target classifier in the N layers is the i-th category, and another output of the i-th layer target classifier is the input of the i+1-th layer target classifier, where i=1,…,N-1, and N is an integer greater than 1.

[0163] In some embodiments, the acquisition unit 601 is further configured to acquire training data comprising N data sets, each of the N data sets comprising a category of data packets;

[0164] The data processing device may further include:

[0165] an arranging unit, configured to arrange the N data sets in descending order according to the number of data packets included therein, to obtain a data set list;

[0166] A labeling unit is used to label the kth dataset in the dataset list as the kth category, k = 1, ..., N;

[0167] The training unit is used to train the N layers of initial classifiers in the initial SBC algorithm in sequence according to the data set list and the corresponding categories to obtain the target SBC algorithm.

[0168] In some embodiments, the training unit is specifically configured to:

[0169] According to the N data sets and corresponding categories in the data set list, the first layer initial classifier in the initial SBC algorithm is trained to obtain the first SBC algorithm;

[0170] According to the jth to Nth datasets in the dataset list and their corresponding categories, the jth layer initial classifier in the j-1th SBC algorithm is trained to obtain the jth SBC algorithm, j=2,…,N-1;

[0171] According to the Nth dataset and the target dataset in the dataset list and the corresponding categories, the Nth layer initial classifier in the N-1th SBC algorithm is trained to obtain the target SBC algorithm. The target dataset is one or more datasets in the N datasets except the Nth dataset.

[0172] In some embodiments, the training unit trains the first-layer initial classifier in the initial SBC algorithm according to the N data sets and corresponding categories in the data set list to obtain a first SBC algorithm, including:

[0173] The first data set in the data set list is used as a positive sample, and the second to Nth data sets in the data set list are used as negative samples. The first layer initial classifier in the initial SBC algorithm is trained to obtain the first SBC algorithm.

[0174] In some embodiments, the training unit trains the j-th layer initial classifier in the j-1 SBC algorithm according to the j-th to N-th datasets in the dataset list and the corresponding categories to obtain the j-th SBC algorithm, including:

[0175] Perform data expansion on the jth data set in the data set list, and the absolute value of the difference between the number of data packets included in the jth data set after expansion and the first data set in the data set list is less than or equal to the threshold;

[0176] According to the expanded j-th data set and the corresponding category, as well as the j+1-th data set to the N-th data set and the corresponding category in the data set list, the j-th layer initial classifier in the j-1-th SBC algorithm is trained to obtain the j-th SBC algorithm.

[0177] In some embodiments, the training unit performs data expansion on the j-th data set in the data set list, including:

[0178] The SMOTE algorithm can be used to perform data expansion on the j-th dataset in the dataset list.

[0179] Based on the above network structure, Figure 7 This is a schematic diagram of the structure of another data processing device provided in an embodiment of the present application. The data processing device can be applied to the above-mentioned NWDAF network element, or to other network elements or devices. The data processing device may include:

[0180] An acquisition unit 701 is configured to acquire training data comprising N data sets, each of the N data sets comprising a data packet of a certain category, where N is an integer greater than 1.

[0181] an arranging unit 702, configured to arrange the N data sets in descending order according to the number of data packets included therein, to obtain a data set list;

[0182] A marking unit 703 is used to mark the k-th data set in the data set list as the k-th category, where k=1, ..., N;

[0183] The training unit 704 is used to train the N layers of initial classifiers in the initial SBC algorithm in sequence according to the data set list and the corresponding categories to obtain a target SBC algorithm including N layers of target classifiers, where one output of the i-th layer classifier in the N-layer classifier is the i-th category, and another output of the i-th layer classifier is the input of the i+1-th layer classifier, where i=1, ..., N-1.

[0184] In some embodiments, the training unit 704 is specifically configured to:

[0185] According to the N data sets and corresponding categories in the data set list, the first layer initial classifier in the initial SBC algorithm is trained to obtain the first SBC algorithm;

[0186] According to the jth to Nth datasets in the dataset list and their corresponding categories, the jth layer initial classifier in the j-1th SBC algorithm is trained to obtain the jth SBC algorithm, j=2,…,N-1;

[0187] According to the Nth dataset and the target dataset in the dataset list and the corresponding categories, the Nth layer initial classifier in the N-1th SBC algorithm is trained to obtain the target SBC algorithm. The target dataset is one or more datasets in the N datasets except the Nth dataset.

[0188] In some embodiments, the training unit 704 trains the first-layer initial classifier in the initial SBC algorithm according to the N data sets and corresponding categories in the data set list to obtain a first SBC algorithm, including:

[0189] The first data set in the data set list is used as a positive sample, and the second to Nth data sets in the data set list are used as negative samples. The first layer initial classifier in the initial SBC algorithm is trained to obtain the first SBC algorithm.

[0190] In some embodiments, the training unit 704 trains the j-th layer initial classifier in the j-1 SBC algorithm according to the j-th to N-th datasets and the corresponding categories in the dataset list to obtain the j-th SBC algorithm, including:

[0191] Perform data expansion on the jth data set in the data set list, and the absolute value of the difference between the number of data packets included in the jth data set after expansion and the first data set in the data set list is less than or equal to the threshold;

[0192] According to the expanded j-th data set and the corresponding category, as well as the j+1-th data set to the N-th data set and the corresponding category in the data set list, the j-th layer initial classifier in the j-1-th SBC algorithm is trained to obtain the j-th SBC algorithm.

[0193] In some embodiments, the training unit 704 performs data expansion on the j-th data set in the data set list, including:

[0194] The SMOTE algorithm can be used to perform data expansion on the j-th dataset in the dataset list.

[0195] Each module in the above-mentioned data processing device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of the processor in the data processing device in the form of hardware, or may be stored in the memory of the data processing device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0196] In an exemplary embodiment, Figure 8 A structural diagram of a computer device provided in an embodiment of the present application. The computer device can be a NWDAF network element, and can also be applied to other network elements or devices. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface is used to exchange information between the processor and an external device. The communication interface is used to communicate with an external device via a network connection. When the computer program is executed by the processor, the above method is implemented.

[0197] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0198] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0199] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0200] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0201] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0202] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0203] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A data processing method, characterized in that: include: Get the target data packet; Determining a quintuple of the target data packet; According to the quintuple, a target binary classification SBC algorithm is used to identify the category of the target data packet, wherein the target SBC algorithm includes N layers of target classifiers, wherein an output of the i-th layer target classifier in the N layers of target classifiers is the i-th category, and another output of the i-th layer target classifier is the input of the i+1-th layer target classifier, where i=1,…,N-1, and N is an integer greater than 1.

2. The method according to claim 1, characterized in that The method further comprises: Acquire training data comprising N data sets, where each of the N data sets comprises data packets of one category; Arrange the N data sets in descending order according to the number of data packets included to obtain a data set list; Mark the kth data set in the data set list as the kth category, k=1,…,N; According to the data set list and the corresponding categories, the N layers of initial classifiers in the initial SBC algorithm are trained in sequence to obtain the target SBC algorithm.

3. The method according to claim 2, characterized in that The method of training the N layers of initial classifiers in the initial SBC algorithm in sequence according to the data set list and the corresponding categories to obtain the target SBC algorithm includes: According to the N data sets and corresponding categories in the data set list, the first layer initial classifier in the initial SBC algorithm is trained to obtain a first SBC algorithm; According to the jth to Nth datasets in the dataset list and the corresponding categories, the jth layer initial classifier in the j-1th SBC algorithm is trained to obtain the jth SBC algorithm, where j=2, ..., N-1; According to the Nth data set and the target data set in the data set list and the corresponding categories, the Nth layer initial classifier in the N-1th SBC algorithm is trained to obtain the target SBC algorithm, and the target data set is one or more data sets in the N data sets except the Nth data set.

4. The method according to claim 3, characterized in that The step of training the first layer initial classifier in the initial SBC algorithm according to the N data sets and the corresponding categories in the data set list to obtain a first SBC algorithm includes: The first data set in the data set list is used as a positive sample, and the second to Nth data sets in the data set list are used as negative samples, and the first layer initial classifier in the initial SBC algorithm is trained to obtain a first SBC algorithm.

5. The method according to claim 3, characterized in that The method of training the j-th layer initial classifier in the j-1 SBC algorithm according to the j-th to N-th datasets and the corresponding categories in the dataset list to obtain the j-th SBC algorithm includes: Performing data expansion on the jth data set in the data set list, wherein the absolute value of the difference between the number of data packets included in the jth data set after expansion and the first data set in the data set list is less than or equal to a threshold; According to the expanded jth data set and the corresponding category, as well as the j+1th to Nth data sets in the data set list and the corresponding categories, the jth layer initial classifier in the j-1th SBC algorithm is trained to obtain the jth SBC algorithm.

6. The method according to claim 5, characterized in that The data expansion of the j-th data set in the data set list includes: A synthetic minority oversampling technique (SMOTE) algorithm may be used to perform data expansion on the j-th data set in the data set list.

7. A data processing method, characterized in that: The method further comprises: Acquire training data comprising N data sets, each of the N data sets comprising data packets of one category, where N is an integer greater than 1; Arrange the N data sets in descending order according to the number of data packets included to obtain a data set list; Mark the kth data set in the data set list as the kth category, k=1,…,N; According to the data set list and the corresponding categories, the N layers of initial classifiers in the initial SBC algorithm are trained in sequence to obtain a target SBC algorithm including N layers of target classifiers, where an output of the i-th layer classifier in the N-layer classifiers is the i-th category, and another output of the i-th layer classifier is the input of the i+1-th layer classifier, where i=1,…,N-1.

8. A data processing device, characterized in that: include: An acquisition unit, configured to acquire a target data packet; a determining unit, configured to determine a quintuple of the target data packet; An identification unit is used to identify the category of the target data packet using a target binary classification (SBC) algorithm based on the quintuple, wherein the target SBC algorithm includes N layers of target classifiers, wherein an output of an i-th layer target classifier in the N layers of target classifiers is the i-th category, and another output of the i-th layer target classifier is an input of an i+1-th layer target classifier, where i=1,…,N-1, and N is an integer greater than 1.

9. A data processing device, characterized in that: include: an acquisition unit, configured to acquire training data comprising N data sets, each of the N data sets comprising a data packet of a category, where N is an integer greater than 1; an arranging unit, configured to arrange the N data sets in descending order according to the number of data packets included therein, to obtain a data set list; a marking unit, configured to mark the kth data set in the data set list as the kth category, where k=1, ..., N; A training unit is used to train N layers of initial classifiers in the initial SBC algorithm in sequence according to the data set list and the corresponding categories to obtain a target SBC algorithm including N layers of target classifiers, where an output of the i-th layer classifier in the N layers of classifiers is the i-th category, and another output of the i-th layer classifier is the input of the i+1-th layer classifier, where i=1,…,N-1.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.