A High-Dimensional Anomaly Data Detection Method

The parameters in the isolated forest algorithm are optimized through the ant colony optimization algorithm, which solves the problem of increasing difficulty in detecting abnormal data in high-dimensional data, and achieves efficient and stable abnormal detection effects.

CN119848744BActive Publication Date: 2025-06-27BEIJING BIG DATA ADVANCED TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510322371.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-27
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

In high-dimensional data, the difficulty of detecting abnormal data increases, and the prior art is difficult to effectively improve the effect of abnormal detection of high-dimensional data.

Method used

Ant colony optimization algorithm is used to optimize the isolated tree parameters in the isolated forest algorithm through ants' position coding, including feature subset selection, exception threshold and weight optimization.

Benefits of technology

It effectively improves the accuracy and stability of abnormal detection of high-dimensional data, and overcomes the problem of insufficient detection capabilities of traditional isolated forests on high-dimensional data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848744B_ABST
    Figure CN119848744B_ABST
Patent Text Reader

Abstract

The present disclosure provides a high-dimensional anomaly data detection method, aiming to improve the effect of high-dimensional data anomaly detection. The method includes: constructing M isolation forests according to the position encodings of M first ants in the current iteration; based on high-dimensional anomaly data detection, respectively evaluating the effectiveness of the M isolation forests to determine the fitness values corresponding to the M first ants in the current iteration; determining the elite ants in the current iteration according to the fitness values corresponding to the M first ants in the current iteration and the fitness values corresponding to the M second ants in the previous iteration; performing multiple iterations according to the above steps, and determining the final target position encoding according to the position encoding corresponding to the elite ants determined in the last iteration; constructing an anomaly detection model according to the final target position encoding, and performing segmentation on the high-dimensional data set based on the data segmentation method corresponding to the anomaly detection model to obtain the anomaly detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of abnormal data detection, and particularly to a high-dimensional abnormal data detection method. Background Art

[0002] With the rapid development of big data and artificial intelligence technologies and their wide applications in various industries, the amount of data has increased rapidly. A large amount of high-dimensional data usually contains redundant features or irrelevant features, which greatly increases the difficulty of applying data processing and data mining algorithms to high-dimensional data. On the other hand, abnormal (anomaly) data, or outlier data, is a relatively common phenomenon in different industrial fields. Due to the existence of these data, the overall statistical distribution characteristics of the data are affected and the data quality is reduced, resulting in a decline or even unusability of the performance of the model obtained by training the application learning algorithm on the data, and thus may lead to incorrect conclusions.

[0003] Therefore, improving data quality through data cleaning has become an important step before data analysis applications, and how to detect abnormal data is the key to accurately and efficiently realizing data cleaning. Among them, in the case of an increasing amount of high-dimensional data, there is a more extensive need for effective anomaly detection of high-dimensional data. However, the complex feature space significantly increases the difficulty of anomaly detection, and how to effectively improve the effect of high-dimensional data anomaly detection remains a major technical difficulty. Summary of the Invention

[0004] To overcome the problems existing in the related art, the present disclosure provides a high-dimensional abnormal data detection method. The technical solution of the present disclosure is as follows:

[0005] According to the first aspect of the embodiments of the present disclosure, a high-dimensional abnormal data detection method is provided, including:

[0006] Determine the position encodings of M first ants in this round of iteration. The position encoding of each ant includes: a discrete encoding group for determining the isolated tree feature subset, and a continuous encoding group for determining the isolated tree anomaly threshold and the isolated tree weight. Each feature in the isolated tree feature subset represents a way of splitting the high-dimensional data set;

[0007] Construct M isolated forests according to the position encodings of the M first ants in this round of iteration;

[0008] Based on the high-dimensional abnormal data detection, evaluate the effectiveness of each of the M isolated forests respectively, and determine the fitness value corresponding to each of the M first ants in this round of iteration. The fitness value represents the detection effect on the abnormal data in the high-dimensional data set;

[0009] Determine the elite ants of the current iteration according to the fitness values corresponding to the M first ants of the current iteration and the fitness values corresponding to the M second ants of the previous iteration. The elite ants are used to determine the position encoding of the M first ants of the next iteration;

[0010] Perform multiple iterations according to the above steps. According to the position encoding corresponding to the elite ants determined in the last iteration, determine the final target position encoding, which includes: the optimal feature subset selection, the isolation tree anomaly threshold, and the isolation tree weight;

[0011] Construct an anomaly detection model according to the final target position encoding, and based on the data splitting method corresponding to the anomaly detection model, split the high-dimensional data set to obtain the anomaly detection result.

[0012] Optionally, determining the elite ants of the current iteration according to the fitness values corresponding to the M first ants of the current iteration and the fitness values corresponding to the M second ants of the previous iteration includes:

[0013] Sort the M first ants of the current iteration and the M second ants of the previous iteration according to the fitness values corresponding to the M first ants of the current iteration and the fitness values corresponding to the M second ants of the previous iteration to obtain the first sorting result;

[0014] Determine the first M ants in the first sorting result as the M second ants of the current iteration;

[0015] Determine the ant with the highest fitness value among the M second ants of the current iteration as the elite ant of the current iteration.

[0016] Optionally, it further includes:

[0017] Sort the M first ants of the current iteration according to the fitness values corresponding to the M first ants of the current iteration to obtain the second sorting result;

[0018] Take the half of the ants with higher fitness values in the second sorting result corresponding to the M first ants of the current iteration as the better ant pool, and take the half of the ants with lower fitness values in the second sorting result as the worse ant pool;

[0019] Select ants from the better ant pool and the worse ant pool respectively according to the preset matching quantity to obtain matching quantity of matching ant groups; each matching ant group includes one ant from the better ant pool and one ant from the worse ant pool;

[0020] Exchange the consecutive encodings of the two ants included in each matching ant group to obtain new ants;

[0021] Sort the M first ants in the current iteration and the M second ants in the previous iteration according to the fitness values corresponding to each of the M first ants in the current iteration and the fitness values corresponding to each of the M second ants in the previous iteration, to obtain a first sorting result, including:

[0022] Sort the new ant, the M first ants in the current iteration, and the M second ants in the previous iteration according to the fitness values corresponding to each of the new ant, the M first ants in the current iteration, and the M second ants in the previous iteration, to obtain a first sorting result.

[0023] Optionally, construct M isolation forests according to the position encodings of the M first ants in the current iteration, including: constructing an isolation forest according to the position encoding of each first ant;

[0024] Each isolation forest is constructed in the following manner:

[0025] Decode each of the sorted discrete encodings in the discrete encoding group to obtain each of the sorted features;

[0026] Determine the feature subsets of each of the isolation trees in the isolation forest corresponding to the first ant according to the order of the features;

[0027] Decode the first continuous encoding in the continuous encoding group to obtain the isolation tree anomaly threshold of the first ant;

[0028] Decode the remaining sorted continuous encodings in the continuous encoding group to obtain each of the sorted weight values;

[0029] Determine the weights of each of the isolation trees in the isolation forest corresponding to the first ant according to the order of the weight values;

[0030] Construct the isolation forest corresponding to the first ant according to the feature subsets corresponding to each of the isolation trees, the weights corresponding to each of the isolation trees, and the isolation tree anomaly threshold.

[0031] Optionally, determine the position encodings of the M first ants in the current iteration, including:

[0032] According to the elite ants obtained in the previous iteration, the target ant selects the feature subsets respectively corresponding to each of the isolation trees in the isolation forest from the feature selection space, where the feature subsets include multiple features; wherein, the feature subsets are sorted in order, and each of the features in the feature subsets is sorted in order, and the target ant is any one of the M second ants in the previous iteration;

[0033] Encode the features in the feature subset according to the first encoding format to obtain the discrete encoding of the features in the discrete encoding group;

[0034] The target ant updates the position vector of the target ant based on the elite ant obtained in the previous iteration according to the wandering range and wandering step size of the current iteration;

[0035] Encode according to the position vector of the target ant according to the second encoding format to obtain the continuous encoding group corresponding to the target ant;

[0036] Determine the position encoding of any one of the M first ants in the current iteration according to the discrete encoding group composed of the discrete encodings and the continuous encoding group.

[0037] Optionally, according to the elite ant obtained in the previous iteration, the target ant selects the feature subset corresponding to each isolated tree in the isolation forest from the feature selection space, including:

[0038] Determine the pheromone concentration of each feature in the feature selection space in the current iteration according to the discrete encoding group corresponding to the elite ant obtained in the previous iteration;

[0039] Determine the selection steps of the target ant, and the selection steps are used to determine the number of features included in the feature subset;

[0040] When the target ant selects a feature from the optional features in the feature search space at each step, determine the selection probability of each optional feature according to the pheromone concentration and heuristic information of each optional feature; the optional feature represents an unselected feature;

[0041] Based on the selection probability of the optional feature, select an optional feature as the feature corresponding to this step according to the roulette strategy;

[0042] Determine the feature subset of the isolated tree according to the feature corresponding to each step;

[0043] After determining the feature subset of an isolated tree, clear the selected state of the feature in the feature selection space.

[0044] Optionally, when the target ant selects a feature from the optional features in the feature search space at each step, determine the selection probability of each optional feature according to the pheromone value and heuristic information of each optional feature, including:

[0045] Determine the multiple selection coefficient of the optional feature, and the multiple selection coefficient is used to control the diversity in the feature selection process; the value corresponding to the multiple selection coefficient of the feature is determined according to the number of times the feature is selected in the current iteration;

[0046] Determine the selection probability of each optional feature according to the multiple selection coefficients, pheromone values, and heuristic information of the optional features.

[0047] Optionally, determine the pheromone concentration of each feature in the feature selection space in this iteration according to the discrete coding group corresponding to the elite ant obtained in the previous iteration, including:

[0048] Determine the feature subset of the elite ant according to the discrete coding group corresponding to the elite ant obtained in the previous iteration;

[0049] Perform pheromone enhancement processing on the features of the feature subset corresponding to the elite ant, and perform pheromone evaporation processing on the features in the feature subsets corresponding to the remaining ants obtained in the previous iteration;

[0050] Determine the pheromone concentration of each feature in the feature selection space in this iteration through the pheromone enhancement processing and the pheromone evaporation processing.

[0051] Optionally, the target ant updates the position vector of the target ant based on the elite ant obtained in the previous iteration according to the wandering range and wandering step length in this iteration, including:

[0052] Select a roulette ant from the M second ants in the previous iteration in the way of roulette;

[0053] According to the wandering range and wandering step length in this iteration, the target ant determines the elite position vector and the roulette position vector respectively around the elite ant and the roulette ant obtained in the previous iteration;

[0054] Update the position vector of the target ant according to the elite position vector and the roulette position vector.

[0055] Optionally, the wandering step length and the wandering range are determined through the following steps:

[0056] Determine the upper and lower bounds of the dimension of the preset target ant;

[0057] Determine the wandering range of the target ant according to the upper and lower bounds of the dimension of the target ant, the current iteration number corresponding to this iteration, and the maximum iteration number;

[0058] Determine the wandering step length of the target ant in this iteration according to the current iteration number corresponding to this iteration.

[0059] According to the second aspect of the embodiments of the present disclosure, there is provided a high-dimensional abnormal data detection device, including:

[0060] A determination module, configured to determine the position encodings of M first ants in the current iteration. The position encoding of each ant includes: a discrete encoding group for determining an isolated tree feature subset, and a continuous encoding group for determining an isolated tree anomaly threshold and an isolated tree weight. Each feature in the isolated tree feature subset characterizes a way of partitioning a high-dimensional data set;

[0061] A construction module, configured to construct M isolated forests according to the position encodings of the M first ants in the current iteration;

[0062] An evaluation module, configured to evaluate the effectiveness of the M isolated forests respectively based on high-dimensional anomaly data detection, and determine the fitness values corresponding to the M first ants in the current iteration. The fitness value characterizes the detection effect on the anomaly data in the high-dimensional data set;

[0063] An update module, configured to determine the elite ants in the current iteration according to the fitness values corresponding to the M first ants in the current iteration and the fitness values corresponding to the M second ants in the previous iteration. The elite ants are used to determine the position encodings of the M first ants in the next iteration;

[0064] An iteration module, configured to perform multiple iterations according to the above steps, and determine the final target position encoding according to the position encoding corresponding to the elite ants determined in the last iteration. The target position encoding includes: the optimal feature subset selection, the isolated tree anomaly threshold, and the isolated tree weight;

[0065] An end module, configured to construct an anomaly detection model according to the final target position encoding, and perform partitioning on the high-dimensional data set based on the data partitioning method corresponding to the anomaly detection model to obtain an anomaly detection result.

[0066] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the high-dimensional anomaly data detection method described in the first aspect are implemented.

[0067] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the high-dimensional anomaly data detection method described in the first aspect are implemented.

[0068] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the high-dimensional anomaly data detection method described in the first aspect are implemented.

[0069] The present disclosure optimizes the parameters of the isolation trees in the isolation forest algorithm through the position encoding of ants, which can effectively improve the anomaly detection effect of high-dimensional data using the isolation forest and overcome the problem of insufficient anomaly detection ability of traditional isolation forests for high-dimensional anomaly data. The position encoding of each ant includes two parts: discrete encoding and continuous encoding. This encoding method enables ants to search and optimize in two different spaces (discrete space and continuous space). Through this dual encoding method, the present disclosure can perform feature selection and model parameter optimization simultaneously, thereby improving the anomaly detection effect. The anomaly data detection effects of the M isolation forests determined in each round of iteration for the high-dimensional data set are evaluated respectively to obtain evaluation results. According to the evaluation results, the ant with the highest fitness value among the M first ants in this round of iteration and the M second ants in the previous round of iteration is selected as the elite ant, which can retain the information of the M second ants in the previous round of iteration, ensure that the elite ant determined in this round of iteration is the optimal ant, and guiding the next round of iteration according to the elite ant can help the ant colony approach a better solution space. Through multiple iterations, the ant colony continuously adjusts the position encoding under the guidance of the updated elite ant in the previous round, gradually optimizing the feature subset, threshold, and weight, enabling the model to continuously approach the global optimal solution (i.e., continuously determining the optimal segmentation method for the high-dimensional data set), obtaining an anomaly detection model. The anomaly detection model contains the optimal segmentation method for the high-dimensional data set. Therefore, by detecting the anomaly data in the high-dimensional data set using the anomaly detection model, the anomaly data in the high-dimensional data set can be efficiently identified, thereby improving the accuracy and stability of anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for the description of the embodiments of the present disclosure will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0071] Figure 1 is a schematic diagram of the steps of a high-dimensional anomaly data detection method shown in an embodiment of the present disclosure;

[0072] Figure 2 is a schematic diagram of a position encoding shown in an embodiment of the present disclosure;

[0073] Figure 3 is a schematic diagram of a feature selection process shown in an embodiment of the present disclosure;

[0074] Figure 4 is a schematic diagram of the steps of a high-dimensional anomaly data detection algorithm shown in an embodiment of the present disclosure;

[0075] Figure 5 It is a schematic diagram of a high-dimensional abnormal data detection device shown in an embodiment of the present disclosure.

[0076] Figure 6 It is a schematic diagram of an electronic device shown in an embodiment of the present disclosure. Detailed implementation manners

[0077] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0078] Terms such as "first" and "second" in the specification and claims of the present disclosure are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0079] First, the professional terms related to the present disclosure are explained:

[0080] Isolation forest: The isolation forest does not need to rely on metrics such as distance and density to measure the differences between samples. Instead, it constructs isolation trees to perform binary tree splitting on the dataset samples to characterize the degree of isolation of abnormal samples from the main trunk of the tree structure. It has the characteristics of being easy to understand, simple and efficient. The final decision result of the isolation forest is obtained by integrating the decision results of several isolation trees. The threshold of the isolation tree, that is, the proportion or quantity of abnormal data in the training data, and the integration weight parameter of the isolation tree are important parameters for training and integrating the isolation trees, which have a key impact on the performance of the final isolation forest model.

[0081] Ant colony optimization (ACO): Ant colony optimization is an evolutionary algorithm inspired by the behavior of ant colonies searching for food in nature. It simulates the natural behavior of ants releasing pheromones during the process of searching for food. In this way, indirect communication is achieved among ants, enabling them to cooperate to find the shortest path from the ant nest to the food source.

[0082] Ant Lion Optimizer (ALO): The ant lion optimizer simulates the behavior of ant lions preying on ants in nature. The ant lion digs a conical pit in the sand and hides at the bottom of the cone waiting for prey. When an ant enters the trap, the ant lion throws sand towards the edge, causing the ant to slide down the cone towards the bottom of the pit. When the ant slides to the bottom of the pit, the ant lion preys on the ant. Subsequently, the trap is repaired to prepare for the next hunt.

[0083] In the wave of the rise of big data, the data scale has increased significantly. The amount of data generated daily has rapidly jumped from the PB (Petabyte) and EB (Exabyte) levels to the ZB (Zettabyte) level, and even the YB (Yottabyte) level. High dimensionality is an important feature of a large amount of this data. This kind of high-dimensional data usually contains redundant or irrelevant features, increasing the complexity of data processing and data analysis.

[0084] Isolation forest can be used to detect abnormal data. However, the classic isolation forest method uses all features to train isolation trees, and its applicability to high-dimensional data is relatively low. To solve this problem, the present disclosure provides a method for detecting high-dimensional abnormal data, which can reduce the feature space complexity of isolation trees and enhance the diversity of isolation trees.

[0085] Figure 1 It is a schematic diagram of the steps of a method for detecting high-dimensional abnormal data shown in an embodiment of the present disclosure. According to Figure 1 as shown, the method may specifically include the following steps S11 - step S16.

[0086] Step S11: Determine the position encodings of M first ants in this round of iteration. The position encoding of each ant includes: a discrete encoding group for determining the feature subset of the isolation tree, and a continuous encoding group for determining the abnormal threshold and weight of the isolation tree. Each feature in the feature subset of the isolation tree represents a way of splitting the high-dimensional data set.

[0087] The input for each round of iteration is M second ants from the previous round of iteration. The M second ants from the previous round of iteration are the M ants with the best detection effect for high-dimensional abnormal data obtained in the previous round. The M second ants in this round of iteration can each update their corresponding position encodings. The M second ants from the previous round of iteration that update the position encodings are the M first ants in this round of iteration. Different position encodings corresponding to ants mean different ants.

[0088] The position encoding of an ant consists of two parts, including a discrete encoding group and a continuous encoding group. M is an integer greater than 0.

[0089] The discrete coding group is used to determine the feature subsets corresponding to each isolated tree in the isolation forest. Each isolated tree corresponds to a feature subset. In the case of having I isolated trees, the discrete coding group includes I feature subsets. When the isolation forest detects abnormal data in a high-dimensional dataset, each isolated tree selects a feature from its corresponding feature subset as the splitting method to split the data in the high-dimensional dataset, so as to obtain the abnormal data determined by the isolated tree. Each feature included in the feature subset corresponding to an isolated tree represents a splitting method for a dimension. For example, in the case where the feature subset corresponding to an isolated tree includes feature A, feature B, and feature C, feature A, feature B, and feature C each represent a splitting method. When the isolated tree selects feature A, the high-dimensional dataset is split based on the splitting method corresponding to feature A.

[0090] By establishing feature subsets for isolated trees, the feature space complexity of isolated trees can be effectively reduced, avoiding isolated trees from selecting features as splitting methods from all the features corresponding to the high-dimensional dataset, but rather selecting features as splitting methods from a finite feature subset. In addition, constructing different feature subsets for different isolated trees can also effectively improve the diversity of isolated trees.

[0091] The discrete coding represents which features are selected from all features for an arbitrary isolated tree to construct the feature subset of the isolated tree. The continuous coding group is used to determine the isolated tree anomaly threshold and the weight of the isolated tree. The anomaly threshold and the weight of the isolated tree can be used to distinguish abnormal data from normal data. Among them, the isolated tree anomaly threshold can be used to determine the proportion or quantity of abnormal data in the dataset, and the weight of the isolated tree is used to obtain the final decision result of the isolation forest for abnormal data.

[0092] Step S12: Construct M isolation forests according to the position encodings of the M first ants in this round of iteration.

[0093] In each round of iteration, M isolation forests are constructed according to the position encodings of the M first ants obtained in this round of iteration. Each ant can generate an isolation forest that matches its encoding, that is, each ant represents a candidate anomaly detection model. The feature subset of the isolated tree in the isolation forest is determined by the discrete coding part, and the anomaly threshold and weight are determined by the continuous coding part.

[0094] Step S13: Based on high-dimensional abnormal data detection, evaluate the effectiveness of the M isolation forests respectively, and determine the fitness values corresponding to the M first ants in this round of iteration. The fitness value characterizes the detection effect of abnormal data in the high-dimensional dataset.

[0095] Based on high-dimensional abnormal data, evaluate the performance of each isolation forest, specifically, evaluate the effect of the isolation forest in detecting abnormal data on the high-dimensional data set. Use the isolation forest to detect the abnormal data in the high-dimensional data set, determine the effect of the isolation forest in detecting abnormal data, so as to evaluate the effectiveness of M isolation forests. An isolation forest with high effectiveness indicates that it can effectively identify the abnormal data in the high-dimensional data set.

[0096] By evaluating the effectiveness of the isolation forest corresponding to each first ant, a fitness value can be assigned to each first ant. The fitness value represents the quality of the model represented by the ant in the current high-dimensional abnormal data detection. The fitness value reflects the performance of the corresponding isolation forest. The higher the fitness value of the ant, the better the effect of the isolation forest corresponding to the position encoding of the ant in detecting high-dimensional abnormal data.

[0097] Step S14: Determine the elite ants in this round of iteration according to the fitness values corresponding to the M first ants in this round of iteration and the fitness values corresponding to the M second ants in the previous round of iteration. The elite ants are used to determine the position encodings of the M first ants in the next round of iteration.

[0098] The M second ants in the previous round of iteration are the input of this round of iteration. The M second ants were selected based on the fitness corresponding to each ant in the previous round of iteration. Therefore, when the M second ants in the previous round of iteration are input into this round of iteration, the M second ants carry their corresponding fitness values and do not need to be recalculated in this round of iteration.

[0099] In each round of iteration, after determining the fitness values corresponding to the M first ants in this round of iteration, the elite ants will be updated according to the fitness values corresponding to the M first ants and the M second ants in the previous round of iteration. The elite ants represent the optimal model in the current iteration, that is, the best feature selection and parameter configuration. The parameter configuration includes the isolation tree anomaly threshold and the isolation tree weight. The position encodings of the M first ants represent the solution space explored in the current iteration. By introducing the M second ants in the previous round of iteration, the fitness values of the ants in this round and the previous round can be compared, and the individuals with better fitness can be selected as elite ants, ensuring that even if the solutions of individual ants in a certain round of iteration perform poorly, useful ant information can be supplemented by selecting excellent solutions from the previous round of iteration, thus avoiding the loss of useful ant information.

[0100] The performance of the elite ants provides the best guidance for the next iteration. In each iteration, the position encodings of the M first ants in this iteration are determined by using the elite ants updated in the previous iteration. Among them, for the first iteration, the position encodings of M ants are initialized, and the fitness values of all initialized ants are calculated in advance. The ant with the optimal fitness value is used as the elite ant for guiding the determination of the position encoding in the first iteration.

[0101] Specifically, the position encoding of the elite ants can be used as a reference to guide the determination of the position encodings of the M first ants in the next iteration. In each iteration, based on the elite ants re-determined in the previous iteration, the position encodings of the M ants input in this iteration (i.e., the M second ants in the previous iteration) are updated, which can enable the ant colony to explore a better solution space along the direction guided by the elite ants.

[0102] By updating the elite ants, it can be ensured that the ant colony approaches the optimal solution. The elite ants guide the search direction of the ant colony and effectively avoid falling into the local optimum.

[0103] Step S15: Perform multiple iterations according to the above steps. According to the position encoding corresponding to the elite ants determined in the last iteration, determine the final target position encoding, which includes: the optimal feature subset selection, the isolation tree anomaly threshold, and the isolation tree weight.

[0104] The above process is repeated for multiple iterations. In each iteration, the position encodings of the M ants input in this iteration (i.e., the M second ants in the previous iteration before this iteration) are continuously updated and evaluated.

[0105] The position encodings of the M ants are continuously updated in multiple iterations, and new M ants with different position encodings are continuously generated. The elite ants determined in each iteration will continuously provide guidance to the M ants input in the next iteration, enabling the M ants updated in each iteration to explore the better solution space in the next iteration and obtain the M ants with updated position encodings.

[0106] As the iteration progresses, the position encodings of the M first ants and the M second ants will be highly similar, and the algorithm corresponding to the method converges. Based on the position encodings of the M first ants determined in the last round of iteration and the position encodings of the M second ants input in the last round of iteration (i.e., the M second ants in the previous round of iteration before the last round of iteration), the final target position encoding is obtained. Specifically, the position encoding of the elite ant updated in the last round of iteration can be determined as the target position encoding. The elite ant refers to the ant with the best effect in detecting abnormal data in the high-dimensional dataset among the M first ants in the last round of iteration and the M second ants input in the last round of iteration. The position encoding corresponding to the elite ant includes the optimal feature subset selection, the isolation tree anomaly threshold, and the isolation tree weight. Among them, the optimal feature subset selection, the isolation tree anomaly threshold, and the isolation tree weight refer to that the detection effect of abnormal data in the high-dimensional dataset meets the preset requirements.

[0107] Step S16: According to the final target position encoding, construct an anomaly detection model, and based on the data splitting method corresponding to the anomaly detection model, split the high-dimensional dataset to obtain the anomaly detection result.

[0108] Based on the target position encoding, construct an anomaly detection model, and the anomaly detection model detects abnormal data in the high-dimensional dataset based on the isolation forest corresponding to the target position encoding.

[0109] In the last round of iteration, based on the position encodings of the M first ants in the last round of iteration and the M second ants input in the last round of iteration, determine the final target position encoding. The target position encoding represents the optimal feature subset selection, anomaly threshold, and isolation tree ensemble weight configuration. Thus, it can be realized that each isolation tree in the isolation forest can more effectively detect abnormal data in the high-dimensional dataset based on the feature subset, and the isolation forest integrates the anomaly data detection results of each isolation tree based on the optimal isolation tree anomaly threshold and isolation tree weight to obtain the final anomaly data detection result.

[0110] Based on this optimal configuration, construct the final anomaly detection model. Use the finally constructed anomaly detection model to detect anomalies in the dataset.

[0111] Embodiments of the present disclosure implement parallel processing of feature selection and isolation forest parameter optimization by combining dual coding of discrete coding groups and continuous coding groups, and can select the correct feature subset and optimize the anomaly threshold and weight. Each ant can obtain an independent isolation forest, and each isolation forest can be constructed based on different feature subsets and parameter configurations, which enables each isolation forest to have a different perspective to identify abnormal data. The combination of multiple isolation forests obtained in one round of iteration enhances the robustness and detection ability of the model, enabling the present disclosure to perform anomaly detection on the data set from multiple perspectives. After evaluating the effectiveness of each isolation forest, the elite ants are updated. The feature subset selection and isolation forest parameter configuration of the elite ants provide good guidance for the next round of iteration. Guided by the elite ants, the present disclosure can focus the search process on the area most likely to find the global optimal solution, avoiding the problem of being likely to fall into the local optimal solution; on the other hand, guided by the elite ants, the M ants will adjust their position coding based on the features and parameters of the elite ants in the next round of iteration. On the basis that the ant colony gradually approaches the optimal solution, the efficiency of the search is ensured.

[0112] Among them, in an optional embodiment, determining the elite ants of the current round of iteration according to the fitness values corresponding to the M first ants in the current round of iteration and the fitness values corresponding to the M second ants in the previous round of iteration includes: sorting the M first ants in the current round of iteration and the M second ants in the previous round of iteration according to the fitness values corresponding to the M first ants in the current round of iteration and the fitness values corresponding to the M second ants in the previous round of iteration, to obtain a first sorting result; determining the first M ants in the first sorting result as the M second ants in the current round of iteration; and determining the ant with the highest fitness among the M second ants in the current round of iteration as the elite ant of the current round of iteration.

[0113] The fitness value represents the quality of the ant's solution in the current task. Specifically, the fitness value reflects the effect of the isolated tree feature subset, threshold, and weight selected by the ant in the anomaly detection of the high-dimensional data set. By comparing the fitness of the ants in the current round and the previous round, it can be ensured that the current and previous best solutions are effectively fused, providing more comprehensive information for the subsequent selection of elite ants and avoiding the limitation of making decisions based only on the performance of the current round.

[0114] Combine the fitness values of the M first ants in this round with the fitness values of the M second ants in the previous round, and sort them in descending order according to their respective fitness values to obtain the first sorting result. Select the top M ants from the first sorting result and use them as the M second ants in this round of iteration. Use them as the input for the next round of iteration. In the next round of iteration, update the position encoding of the M second ants in the input to obtain the M first ants in the next round of iteration. This screening method ensures that the optimal solution (i.e., the ant with a higher fitness value) in each round of iteration can be retained and passed to the next round. By selecting ants with higher fitness values, it ensures the effective exploration of the search space and reduces possible "information loss".

[0115] The elite ant represents the optimal solution in the current round of iteration. The isolated tree feature subset, anomaly threshold, and weight corresponding to its solution can optimize the anomaly data detection effect to the greatest extent. By selecting the ant with the best fitness as the elite, it ensures that the model converges to the optimal solution after each round of iteration. Therefore, determine the ant with the highest fitness value among the M second ants in this round of iteration as the elite ant in this round of iteration.

[0116] Adopt the embodiments of the present disclosure. By combining the M first ants in this round of iteration and the M second ants in the previous round of iteration, sorting and selecting the M second ants with the optimal fitness in this round of iteration, it ensures the effective retention and optimization of the optimal solution in each round of iteration. Furthermore, it enables the ant algorithm to continuously converge to the global optimal solution, not only improving the search efficiency but also avoiding information loss, and ensuring the accuracy and stability of high-dimensional anomaly data detection. Update the ant population by eliminating the strategy of low fitness values. This dynamic update mechanism ensures that the ant population can continuously evolve and approach a better solution. Select the ant with the highest fitness value from the updated ant population as the elite ant, which helps to avoid the algorithm falling into a local optimal solution and ensures that the algorithm can continuously approach the global optimal solution.

[0117] Among them, in an optional embodiment, it further includes: sorting the M first ants in the current iteration according to the fitness values respectively corresponding to the M first ants in the current iteration to obtain a second sorting result; taking the half of the ants with higher fitness values in the second sorting result corresponding to the M first ants in the current iteration as the better ant pool, and taking the half of the ants with lower fitness values in the second sorting result as the worse ant pool; selecting ants from the better ant pool and the worse ant pool respectively according to a preset matching number to obtain a matching number of matching ant groups; each matching ant group includes one ant from the better ant pool and one ant from the worse ant pool; swapping the consecutive codes of the two ants included in each matching ant group to obtain new ants; sorting the M first ants in the current iteration and the M second ants in the previous iteration according to the fitness values respectively corresponding to the M first ants in the current iteration and the fitness values respectively corresponding to the M second ants in the previous iteration to obtain a first sorting result, including: sorting the new ants, the M first ants in the current iteration, and the M second ants in the previous iteration according to the fitness values respectively corresponding to the new ants, the M first ants in the current iteration, and the M second ants in the previous iteration to obtain a first sorting result.

[0118] In the current iteration, after obtaining the fitness values respectively corresponding to the M first ants in the current iteration, the M first ants in the current iteration can be sorted in descending order of fitness value to obtain a second sorting result, and the half of the ants with higher fitness values in the second sorting result are classified into the better ant pool; the half of the ants with lower fitness values are classified into the worse ant pool.

[0119] Ants are selected from the better ant pool and the worse ant pool respectively according to a preset matching number to obtain a corresponding number of matching ant groups. For example, when the matching number is 2, 2 groups of matching ant groups are required. For each group of matching ant groups, one ant needs to be selected from the better ant pool and one ant from the worse ant pool for matching to obtain the matching ant group. By pairing the better and worse ants, it is ensured that the worse ants can learn from the advantages of the better ants, thereby realizing the optimization of the worse ants.

[0120] Perform consecutive coding exchanges on two ants in each matching ant group, that is, exchange the parts used to determine the isolated tree anomaly threshold and the isolated tree weight in the two ants. After the exchange, new ants are obtained. The number of new ants obtained from each group of matching ant groups is the same as the number of ants included in the matching ant group, that is, two new ants. For example, a certain matching ant group includes a worse ant A and a better ant B. The position coding of the worse ant A includes discrete coding A and consecutive coding A, and the position coding of the better ant B includes discrete coding B and consecutive coding B. After exchanging the consecutive coding parts of the worse ant A and the better ant B, new ants C and new ants D are obtained. The position coding of new ant C includes discrete coding A and consecutive coding B, and the position coding of new ant D includes discrete coding B and consecutive coding A. By allowing the better and worse ants to exchange the consecutive coding parts, the excellent characteristics of the better ants can be transmitted to the worse ants, thereby improving the performance of the worse ants.

[0121] After the new ants are generated, based on the fitness values of the new ants, the M first ants in this round of iteration, and the M second ants in the previous round of iteration, sort the new ants, the M first ants in this round of iteration, and the M second ants in the previous round of iteration in descending order of fitness value to obtain the first sorting result.

[0122] Adopting the embodiments of the present disclosure to sort the M first ants in this round of iteration according to the fitness value can clearly divide the ants into two groups with better and worse performance. The better ants and the worse ants exchange information (that is, exchange the consecutive coding parts), which helps the worse ants learn the excellent characteristics of the better ants, thereby improving the quality of the overall ant population and accelerating the convergence process of the algorithm. In each round of iteration, determine the first sorting result according to the fitness values of the newly generated ants, the M first ants in this round of iteration, and the M second ants in the previous round of iteration.

[0123] Among them, in an optional embodiment, according to the position encodings of the M first ants in this round of iteration, M isolation forests are constructed, including: according to the position encoding of each first ant, one isolation forest is constructed; each isolation forest is constructed in the following manner: decoding each discrete encoding sorted in order in the discrete encoding group to obtain each feature sorted in order; determining the feature subsets of each isolation tree in the isolation forest corresponding to the first ant according to the order of each feature; decoding the first continuous encoding in the continuous encoding group to obtain the anomaly threshold of the isolation tree of the first ant; decoding the remaining continuous encodings sorted in order in the continuous encoding group to obtain each weight value sorted in order; determining the weights of each isolation tree in the isolation forest corresponding to the first ant according to the order of each weight value; constructing the isolation forest corresponding to the first ant according to the feature subsets corresponding to each isolation tree, the weights corresponding to each isolation tree, and the isolation tree anomaly threshold.

[0124] The structure of the position encoding can be determined according to the anomaly detection model to be constructed. Specifically, when the isolation forest corresponding to the anomaly detection model needs to include I isolation trees and the feature subset corresponding to each isolation tree needs to include q features, the position encoding at this time needs to include discrete encodings corresponding to I*q features, continuous encodings corresponding to I isolation tree weights, and a continuous encoding corresponding to an isolation tree anomaly threshold.

[0125] Figure 2 is a schematic diagram of a position encoding shown in an embodiment of the present disclosure. According to Figure 2 shown, each discrete encoding represents a feature selected to enter the feature subset. When the isolation forest includes I isolation trees and the feature subset corresponding to each isolation tree needs to include q features, the discrete encodings 1 - discrete encoding q in the discrete encoding group represent the q features included in the feature subset corresponding to one of the isolation trees in the isolation forest. The continuous encoding group includes two parts, including an isolation tree threshold encoding representing the isolation tree anomaly threshold and an isolation tree weight encoding representing the isolation tree weight. In the continuous encoding part, the isolation tree anomaly threshold is represented by one continuous encoding, and each isolation tree can distinguish normal data and abnormal data through this continuous encoding. The number of isolation tree weight encodings in the continuous encoding part is the same as the number of isolation trees in the isolation forest. When there are I isolation trees in the isolation forest, I isolation tree weight encodings are required in the discrete encoding group.

[0126] The position encoding of each ant can be represented in the form of Figure 2 , and each ant can construct a corresponding isolation forest. For the position encodings of the M first ants in this round of iteration, M corresponding isolation forests can be constructed.

[0127] For each discrete encoding group of the first ants, first decode it in the order of the encoding to obtain a series of features. Each discrete encoding represents a specific feature. By decoding the discrete encodings in the discrete encoding group, the features used by each isolated tree can be determined. The number of features obtained by decoding the discrete encoding group needs to conform to the number of features required by the isolation forest. For example, when the isolation forest includes I isolated trees and each isolated tree's corresponding feature subset needs to include q features, the number of features obtained by decoding the discrete encoding group needs to satisfy I * q.

[0128] Decoding each continuous encoding group of the first ants can obtain an anomaly threshold and several weight values. Specifically, decoding the first continuous encoding in the continuous encoding group can obtain the anomaly threshold of the isolated tree corresponding to the isolation forest. Continuing to decode the remaining continuous encodings, a series of weight values are obtained.

[0129] For a series of features and a series of weight values obtained from the position encoding, determine the feature subset and weight used by each isolated tree. The feature order determines which features will be applied to the isolated tree, and the weight value determines the influence and importance of each isolated tree in the isolation forest.

[0130] Specifically, the features corresponding to the same isolated tree are connected end to end among the features sorted in order obtained by decoding. The series of features sorted in order obtained by decoding are sequentially assigned to different isolated trees from the first feature to the last feature. Among them, after the feature subset of an isolated tree is determined, the remaining features are sequentially assigned to the next isolated tree to obtain the feature subset of the next isolated tree.

[0131] The series of weight values obtained by decoding are sequentially assigned to different isolated trees.

[0132] After decoding a position encoding and assigning the decoded features, isolated tree anomaly threshold, and isolated tree weights to each isolated tree in the isolation forest, an isolation forest corresponding to the position encoding can be constructed based on the assignment results.

[0133] Using the embodiments of the present disclosure, the isolation forests constructed by each ant will have different feature selections, weights, and anomaly thresholds, so that the models of different ants have differences in the detection ability of abnormal data, ensuring that the present disclosure can be optimized simultaneously in multiple dimensions, enhancing the exploration ability of the swarm algorithm, and enabling different ants to find the optimal solution in different local solution spaces. The decoding process of the discrete coding group determines the features used by the isolation trees in the isolation forest, enabling each isolation tree to customize a feature subset for a specific data set and problem, avoiding the interference of redundant features, and improving the efficiency of the model. By setting different anomaly thresholds and weights for each isolation tree, the contribution of each isolation tree can be adjusted, so that the model can better adapt to the changes of high-dimensional complex data and improve the detection effect.

[0134] Among them, in an optional embodiment, determining the position encodings of the M first ants in this round of iteration includes: according to the elite ants obtained in the previous round of iteration, the target ant selects the feature subsets corresponding to the respective isolation trees in the isolation forest from the feature selection space, where the feature subsets include multiple features; wherein, the feature subsets are sorted in order, and the features in the feature subsets are sorted in order, and the target ant is any one of the M second ants in the previous round of iteration; encoding the features in the feature subset according to the first encoding format to obtain the discrete encodings of the features in the discrete coding group; the target ant updates the position vector of the target ant based on the elite ants obtained in the previous round of iteration according to the roaming range and roaming step length of this round of iteration; encoding according to the position vector of the target ant according to the second encoding format to obtain the continuous coding group corresponding to the target ant; determining the position encoding of any one of the M first ants in this round of iteration according to the discrete coding group composed of the discrete encodings and the continuous coding group.

[0135] The elite ants obtained in the previous round of iteration will affect the pheromones of each feature in the feature selection space. Specifically, in the case where a feature is selected into the discrete coding group of the elite ants, this feature will carry stronger pheromones in the feature selection space. When the target ant selects features from the feature selection space, the selection is based on the pheromones carried by the features. The stronger the pheromones carried by a feature, the greater the probability that the feature will be selected.

[0136] The elite ants determined in the previous iteration can affect the feature selection of M ants input in this iteration (i.e., the M second ants in the previous iteration). The features selected based on the elite ants determined in the previous iteration can approximate the optimal solution more closely. The target ant is any one of the M second ants in the previous iteration. Taking the target ant as an example, this embodiment illustrates how the M second ants in the previous iteration update their position encodings in this iteration to obtain the M first ants in this iteration.

[0137] When the target ant determines the feature subset for each isolated tree in its corresponding isolation forest, the determination of the feature subset is carried out in the order of the isolated trees. After the feature subset of one isolated tree is determined, the feature subset of the next isolated tree is determined. The features selected by each ant are also sorted in the order of being selected. Among them, the feature subset corresponding to each isolated tree is determined from the complete feature selection space.

[0138] Encode the selected features according to the first encoding format to obtain Figure 2 the discrete encoding shown in, and each selected feature corresponds to a discrete encoding. Combine all the discrete encodings in the order of the corresponding features being selected to update the discrete encoding group of the target ant.

[0139] In the part of determining the isolation tree anomaly threshold and the isolation tree weight, the present disclosure adopts the search method of ALO.

[0140] The position of the target ant in this iteration will be updated based on the performance of the elite ants obtained in the previous iteration. The target ant has a position vector representing its position in the continuous search space. The target ant will update the position vector corresponding to its current position based on the elite ant according to its wandering range and step size. Specifically, it is manifested that the target ant wanders around the elite ant within the wandering range according to the wandering step size, thereby updating its position. The target ant updates the continuous encoding group of the target ant according to the updated position vector.

[0141] Among them, as the iteration progresses, the wandering range and step size corresponding to each iteration will change. The wandering range and step size determine the distance and direction of the ant's movement in the search space. Specifically, as the iteration progresses, the wandering range of the ant will shrink, and the ant will recalculate a wandering step size in each iteration.

[0142] By updating the continuous encoding group and the discrete encoding group of the target ant, the position encoding of the target ant is updated. The target ant after updating the position encoding will be used as one of the M first ants in this iteration.

[0143] Among them, in an optional embodiment, the wandering step length and the wandering range are determined through the following steps: determining the upper and lower bounds of the dimensions of the preset target ant; determining the wandering range of the target ant according to the upper and lower bounds of the dimensions of the target ant, the current iteration number corresponding to the current iteration, and the maximum iteration number; determining the wandering step length of the target ant in the current iteration according to the current iteration number corresponding to the current iteration.

[0144] As the iteration progresses, the wandering range of the ant will gradually shrink, thereby prompting the algorithm to gradually focus on the refined search of the solution. The maximum iteration number determines the overall running time of the algorithm and the depth of the search. The wandering range should be adjusted according to different stages of the maximum iteration number so that the search behavior of the ant adapts to the iteration progress.

[0145] The wandering range of the target ant in each iteration can be determined by the following formula:

[0146]

[0147] c and d represent the upper and lower bounds of the values of each dimension of the preset target ant, and and respectively represent the upper and lower bounds of the search range of the values of each dimension of the ant in the t-th iteration, and the value of R is determined according to the current iteration number.

[0148] R can be determined by the following formula:

[0149]

[0150] The value of w depends on the current iteration number t, and T represents the maximum iteration number. When, w = 0; When, w = 2; Then w = 3; Then w = 4; when When, w = 5; when When, w = 6, so that shows a piecewise exponential increasing trend.

[0151] The wandering step length determines the moving amplitude of the ant when updating its position each time. As each iteration progresses, the wandering step length of the ant changes with the change of the random number. The step length r is calculated according to the generated random number. In each iteration, the step length r affects the current position X(t). Specifically, the current position X(t) is the sum of the step lengths generated in all iterations, representing the overall change from the initial position to the current iteration position, that is, the current position of the ant is obtained by adding the random wandering step lengths corresponding to each iteration.

[0152] The current position of the target ant can be determined by the following formula:

[0153]

[0154] Among them, t is the current iteration number, T is the maximum iteration number, X(t) represents the random walk position, cumsum represents the cumulative sum of the random walk step lengths in the previous t - 1 rounds of iteration, r is the generation function of the random walk step length, which is used to calculate the walk step length corresponding to each round of iteration, and its calculation is , where rand represents a random number between (0, 1).

[0155] By adopting the embodiments of the present disclosure, the range of the walk is reduced as the number of iterations increases, such that the range of the walk is larger in the initial stage, which is beneficial for extensive exploration. As the number of iterations increases, the range of the walk gradually shrinks, helping the ant to perform local refinement optimization after finding a better solution.

[0156] In this round of iteration, the new position vector of the target ant will be encoded according to the second encoding format to obtain a continuous encoding group of the target ant. The continuous encoding group usually contains the updated parameters of the target ant, such as thresholds and weights, etc.

[0157] After determining the discrete encoding group of the target ant through ACO and the continuous encoding group of the target ant through ALO respectively, the position encoding of the target ant is obtained according to the discrete encoding group and the continuous encoding group.

[0158] By adopting the embodiments of the present disclosure, based on the feature selection method guided by elite ants, the efficiency of the search process can be improved. The target ant can inherit its excellent feature selection strategy by imitating the selection of elite ants, thereby avoiding blindly searching the feature space from scratch. By adjusting the range of the walk and the step length, the target ant can flexibly adjust the search direction and pace within the feature space, effectively avoiding falling into local optimal solutions or premature convergence. The guiding role of the elite ants makes the search of the target ant more directional and enables it to maintain the ability of global optimization in a large search space, thereby enhancing the global search ability and optimization efficiency of the algorithm. The combination of the position encoding provides a comprehensive and accurate representation method, and the isolation forest corresponding to the target ant is fully described. This encoding form enables each ant to flexibly search and adjust in the feature space, improving the efficiency and search quality of the algorithm.

[0159] Among them, in an optional embodiment, according to the elite ants obtained in the previous iteration, the target ant selects the feature subsets corresponding to the individual isolation trees in the isolation forest from the feature selection space, including: determining the pheromone concentration of each feature in the feature selection space in the current iteration according to the discrete coding group corresponding to the elite ants obtained in the previous iteration; determining the selection steps of the target ant, where the selection steps are used to determine the number of features included in the feature subset; in the case where the target ant selects a feature from the optional features in the feature search space at each step, determining the selection probability of each optional feature according to the pheromone concentration and heuristic information of each optional feature; the optional features represent the unselected features; based on the selection probability of the optional features, selecting an optional feature as the feature corresponding to this step according to the roulette strategy; determining the feature subset of the isolation tree according to the feature corresponding to each step; after determining the feature subset of an isolation tree, clearing the selected status of the features in the feature selection space.

[0160] In the ACO algorithm, ants communicate experience with each other through pheromones to influence the decisions of other ants. The elite ants are the ants with the highest fitness determined in the previous iteration. Therefore, the pheromone concentration of each feature in the feature selection space can be updated according to the discrete coding group corresponding to the elite ants, so as to enable the elite ants determined in the previous iteration to guide the M ants in the current iteration to determine their own corresponding position codes.

[0161] Among them, in an optional embodiment, determining the pheromone concentration of each feature in the feature selection space in the current iteration according to the discrete coding group corresponding to the elite ants obtained in the previous iteration includes: determining the feature subset of the elite ants according to the discrete coding group corresponding to the elite ants obtained in the previous iteration; performing pheromone enhancement processing on the features in the feature subset corresponding to the elite ants, and performing pheromone evaporation processing on the features in the feature subsets corresponding to the remaining ants obtained in the previous iteration; determining the pheromone concentration of each feature in the feature selection space in the current iteration through the pheromone enhancement processing and the pheromone evaporation processing.

[0162] For the feature subset corresponding to the elite ants, it is necessary to enhance the pheromone of these features, that is, increase the weight of the features selected by the elite ants, so that in the next iteration, these features are more likely to be selected. By enhancing the pheromone, more search resources can be concentrated on those features with good performance, thus accelerating the process of finding the optimal feature subset.

[0163] Opposite to pheromone enhancement is pheromone evaporation, which means reducing the weight of the features selected by those ants with poor performance. Pheromone evaporation makes those features that fail to perform well gradually lose their attractiveness, thus encouraging ants to select new feature subsets and maintaining the diversity of the search space.

[0164] The pheromone concentration of each feature in the feature selection space is obtained by processing the selection results of the elite ants and other ants obtained in the previous iteration. Through the pheromone enhancement process of the elite ant feature subset and the pheromone evaporation process of other ant feature subsets, the updated pheromone concentration of all features in this iteration is finally determined. These pheromone concentration values will affect the decisions of M ants in the next round of feature selection.

[0165] The pheromone of each feature in the feature selection space can be updated by the following formula:

[0166]

[0167] Where: represents the evaporation coefficient of pheromone; represents the path solution (feature subset) to be updated; represents the evaluation value of this path solution on the objective function; Q is a constant, which is a scaling factor of the objective function evaluation value. The value of this factor is jointly determined by the initial pheromone value, evaporation coefficient and objective function evaluation value. Among them, the objective function is preset and can be used to measure the performance of elite ants in a certain task.

[0168] Adopting the embodiments of the present disclosure, the selection results of elite ants are usually the best solutions in the current search space. By strengthening the pheromone concentration of these features, it is possible to accelerate the convergence towards the optimal feature subset, and it is possible to find a relatively excellent feature subset in fewer iterations, avoiding the situation of a large number of ineffective searches. This balance can enable the algorithm to ensure local search optimization while not losing the global search ability, thereby avoiding falling into local optimal solutions and being able to discover global optimal solutions.

[0169] The selection step of the target ant determines how many features the target ant needs to select when performing feature selection for a certain isolated tree. The selection step is usually set according to the requirements of the problem or the maximum number of features, ensuring that the scale of the feature subset is appropriate and avoiding excessive or insufficient feature selection.

[0170] The optional features represent the features not selected in the feature selection space, that is, each feature included in the feature subset is non-repetitive. And the selection of features in each feature subset is made from the complete feature selection space. For example, for the feature subset of a certain isolated tree, when selecting the first feature included in the feature subset from the feature selection space, none of the features in the feature selection space have been selected. After the feature subset of an isolated tree is determined, the selected status of each feature in the feature selection space can be cleared. When selecting features from the feature selection space for the feature subset of the next isolated tree, feature selection is performed again from the complete feature selection space.

[0171] Figure 3 is a schematic diagram of a feature selection process shown in an embodiment of the present disclosure. According to Figure 3 shown, f represents a feature, the feature selection space includes n features, s represents the number of steps selected by the ant; q represents the maximum number of steps the ant can select. Initially, the selected features in the feature selection space are empty. The ant starts selecting features from the first step and stops at the qth step. Each time, an unvisited feature is selected from the optional features according to the roulette wheel strategy.

[0172] Each time the target ant selects a feature from the optional features, the selection probability is calculated by comprehensively considering the pheromone concentration and heuristic information. The selection probability of the jth feature in the feature selection space can be determined by the following formula:

[0173]

[0174] where represents the pheromone value of the jth feature; represents the heuristic information of the jth feature; represents the set of features that ant a has visited; and represent the influence factors of pheromone and heuristic information.

[0175] Among them, the heuristic information uses the information gain of the feature, which measures the reduction in uncertainty before and after dividing the dataset by the feature. The greater the information gain, the stronger the correlation of the feature. Uncertainty can be represented by entropy. Information gain is an index to measure the importance of a feature in data division. The greater the information gain of a feature, the more effectively the feature can reduce the uncertainty of the data, and thus the stronger the correlation.

[0176] The target ant selects a feature using the roulette wheel strategy according to the selection probability of the optional features. In this process, the pheromone concentration and heuristic information work together to ensure that the selected feature is the most important for constructing the isolation tree. Based on the features selected in each step, the target ant constructs a feature subset of a certain isolation tree. When determining the feature subset corresponding to the next isolation tree, features are reselected from the feature selection space according to the number of selection steps.

[0177] Adopting the embodiments of the present disclosure, according to the distribution of the pheromone concentration, the target ant can be more inclined to select those features that performed well in the previous iteration when selecting features, which can effectively guide the feature selection to efficient and important features, avoid the selection of irrelevant features, thereby accelerating the search process and improving the accuracy of the selection. The roulette wheel strategy maintains the diversity of the search by introducing randomness and avoids falling into local optimal solutions. By setting the number of selection steps of the target ant, the number of features used for each isolation tree can be precisely controlled, thereby ensuring the quality of each isolation tree in the isolation forest, contributing to improving the effect of the isolation forest model, and avoiding overfitting or underfitting. By ensuring that the features in each feature subset are selected from the feature selection space and are non-repetitive, the addition of redundant features can be avoided, the independence and diversity of each feature selection can be guaranteed, and the feature subset of each isolation tree is not affected by the selection results of other isolation trees, avoiding the situation of feature overlap between different trees or selection bias towards specific features.

[0178] Among them, in an optional embodiment, when the target ant selects a feature from the optional features in the feature search space in each step, the selection probability of each optional feature is determined according to the pheromone value and heuristic information of each optional feature, including: determining the multiple selection coefficient of the optional feature, where the multiple selection coefficient is used to control the diversity in the feature selection process; the value corresponding to the multiple selection coefficient of the feature is determined according to the number of times the feature is selected in the current iteration; and the selection probability of each optional feature is determined according to the multiple selection coefficient, pheromone value and heuristic information of the optional feature.

[0179] To increase the diversity of the isolation trees, a multiple selection coefficient can be introduced. Diversity refers to the difference in the feature selection by ants during the search process. Increasing the diversity of the isolation trees can avoid always selecting similar feature subsets for the isolation trees. The multiple selection coefficient is used to control the possibility of each feature being selected multiple times in the feature selection process. Specifically, if a certain feature is selected multiple times, its multiple selection coefficient will increase, which means that the possibility of this feature being selected in subsequent selections will decrease.

[0180] The multiple selection coefficient of a specific feature is dynamically calculated and is related to the number of times the specific feature is selected in the current iteration. Specifically, in each round of iteration, the number of times each feature is selected is recorded, and the multiple selection coefficient of each feature is calculated based on the number of times.

[0181] After introducing the multiple selection coefficient, the selection probability of the optional features in the feature selection space can be determined by the following formula:

[0182]

[0183] where, represents the pheromone value of the j-th feature; represents the heuristic information of the j-th feature; represents the set of features that ant a has visited; and represent the influence factors of pheromone and heuristic information, is the multiple selection coefficient of the c-th isolated tree, represents the number of times the first c - 1 isolated trees have selected the j-th feature. For example, represents the number of times the first 10 isolated trees have selected the 8th feature.

[0184] By adopting the embodiment of the present disclosure, through the introduction of the multiple selection coefficient, it is possible to effectively avoid over-relying on certain features during the search process while ignoring other potentially important features, ensuring that the ants have sufficient exploration ability in feature selection and avoiding the dilemma of local optimal solutions. The multiple selection coefficient helps the algorithm maintain sufficient diversity during the search process, and at the same time strengthens the selection of high-quality features through pheromone and heuristic information, so that the exploitation and exploration are balanced, so that the algorithm will neither converge prematurely to the local optimal solution during the search process nor be able to continuously adjust the search strategy to explore new feature combinations.

[0185] Among them, in an optional embodiment, the target ant updates the position vector of the target ant based on the elite ant obtained in the previous iteration according to the wandering range and wandering step length of the current iteration, including: selecting a roulette ant from the M second ants in the previous iteration in the way of roulette; according to the wandering range and wandering step length of the current iteration, the target ant respectively determines an elite position vector and a roulette position vector around the elite ant and the roulette ant obtained in the previous iteration; and updates the position vector of the target ant according to the elite position vector and the roulette position vector.

[0186] The roulette ant is selected from the M ants input in this round (i.e., the M second ants in the previous iteration) based on the roulette method. The roulette ant represents that it performed better in the previous iteration, and the roulette ant plays a guiding role in updating the target ant.

[0187] When updating the position vector of the target ant in this iteration around the elite ant, it is also necessary to synchronously determine a position vector around the roulette ant. The position vector obtained around the elite ant is determined as the elite position vector, and the position vector determined around the roulette ant is determined as the roulette position vector.

[0188] When wandering around the elite ant and the roulette ant respectively, the wandering step size and wandering range are the same. The same wandering step size and wandering range are used to determine the position vectors around the elite ant and the roulette ant respectively.

[0189] After obtaining the elite position vector and the roulette position vector, combining the elite position vector and the roulette position vector to update the position vector of the target ant in this iteration, and this position vector is used to obtain the continuous coding part in the position coding of the target ant in this iteration.

[0190] Specifically, the following method can be used to determine the position vector of the target ant in this iteration:

[0191]

[0192] Where At is the position vector of the ant in the t-th iteration, e and r represent the elite ant and the roulette ant respectively, is the elite position vector obtained by the target ant in the t-th iteration through random wandering around e, is obtained by the target ant in the t-th iteration through random wandering around r; the roulette vector.

[0193] By adopting the embodiments of the present disclosure, by simultaneously considering the influences of the elite ant and the roulette ant in position update, the target ant can explore in multiple directions, not only limited to the area of the elite ant. The guidance of the roulette ant enables the ant to maintain appropriate diversity around the current solution, avoiding the search process from being too concentrated and reducing the risk of falling into local optima. Simultaneously considering the influences of the elite ant and the roulette ant enables the ant to approach the optimal solution while not ignoring other potential high-quality solutions, enhancing the global optimization ability of the algorithm and being able to better discover the global optimal solution in a large-scale or high-dimensional search space.

[0194] To implement the high-dimensional abnormal data detection method provided by the present disclosure, the present disclosure proposes a high-dimensional abnormal data detection algorithm, Figure 4It is a schematic diagram of the steps of a high-dimensional anomaly data detection algorithm shown in an embodiment of the present disclosure.

[0195] As shown Figure 4 The input of the algorithm is the maximum number of iterations T, the objective function F for updating pheromone, and the number of ants M. After T rounds of iteration, an anomaly detection model for high-dimensional data detection is output. M ants mean that in each round of iteration, M ants will update their position encodings respectively.

[0196] Step S1: Initialize the position encodings of M ants.

[0197] The position encodings of the ants are updated in each round of iteration. Based on the position encodings of the ants, a corresponding isolation forest can be obtained. The M ants after initializing the position encodings are used as the M ants input in the first round of iteration, and their initialized position encodings represent the position encodings before the first round of iteration update.

[0198] Step S2: Calculate the fitness values of all ants, and select the individual with the best fitness value from the population as the elite ant.

[0199] Calculate the fitness values of the initialized ants, and this fitness value is used to judge the effect of the isolation forest corresponding to the ants when performing anomaly data detection. The elite ant selected before entering the iteration process is the elite ant that provides guidance in the first round of iteration.

[0200] Start the iteration process from step S3.

[0201] Step S3: Judge whether the current iteration number t is less than the maximum iteration number T.

[0202] If the judgment is yes, enter step S4; if the judgment is no, enter step S10.

[0203] Step S4: Judge whether the current m-th ant is less than the maximum number of ants M.

[0204] If the judgment is yes, enter step S5; if the judgment is no, enter step S8.

[0205] Step S5: Judge whether the current i-th isolation tree is less than the maximum number of isolation trees I.

[0206] If the judgment is yes, enter step S6; if the judgment is no, enter step S7.

[0207] Step S6: The m-th ant calculates the path probability value that meets the conditions for a certain step of the i-th isolation tree, and selects a certain path using the roulette wheel strategy, and finally generates a path solution for this isolation tree. The path solution represents the feature subset corresponding to the current isolation tree i.

[0208] The m-th ant determines the feature subset corresponding to the i-th isolated tree. The feature subset includes q features. The m ants make q selections. In any selection step, the path probability value that meets the conditions is calculated, that is, the selection probability of the optional features in the feature selection space is calculated. Subsequently, the roulette wheel strategy is used to select a certain path, that is, a feature, and finally a path solution is generated for the i-th isolated tree. The path solution represents the feature subset of the isolated tree.

[0209] The m-th ant continues to determine the feature subset corresponding to the (i + 1)-th isolated tree, returns to step S5, and continues to judge whether the i = i + 1-th isolated tree is less than the maximum number of isolated trees I.

[0210] After I rounds of step S6, the m-th ant completes the determination of the feature subsets of I isolated trees, encodes the features in the I feature subsets, and obtains the discrete coding group of the m-th ant.

[0211] Step S7: Select a roulette wheel ant through roulette wheel selection, and the ant updates the continuous coding group around the elite ant and the roulette wheel ant.

[0212] Select a roulette wheel ant from M ants through the roulette wheel strategy. The ant m moves around the roulette wheel ant and the elite ant respectively, so as to update the continuous coding group in the position coding of the m-th ant.

[0213] After steps S5 - S7, the position coding of the m-th ant is obtained, the position coding of the m-th ant is updated, and the m-th first ant in this round of iteration is obtained. Continue to determine the position coding of the (m + 1)-th ant, return to step S4, and continue to judge whether the m = m + 1-th ant is less than the maximum number of ants M.

[0214] Step S8: Decode the position codings of the M ants obtained in the t-th round of iteration respectively, obtain the corresponding isolated forests and evaluate them to obtain the fitness values of the M ants. Apply the matching and exchange strategy to generate new ants and evaluate them to obtain the fitness values of the new ants.

[0215] The position codings of the M ants obtained in the t-th round of iteration are the updated position codings. The M ants after the t-th round of position coding update represent the M first ants in the t-th round of iteration.

[0216] After decoding the position encodings of the M first ants in the t-th iteration respectively, M corresponding isolation forests are obtained, and based on the anomaly data detection effect of the isolation forests, the fitness values of the M first ants in the t-th iteration are obtained. Based on the fitness values of the M first ants in the t-th iteration, the M first ants in the t-th iteration are divided into a better ant pool and a worse ant pool. The matching exchange strategy is applied to exchange the consecutive encodings of the better ants and the worse ants in the matching ant group to obtain new ants. The fitness values of the new ants are also determined.

[0217] Step S9: Update the elite ants based on the fitness values of the ants, and update the pheromone using the elite ants.

[0218] Based on the fitness values of the new ants, the fitness values of the M first ants in the t-th iteration, and the fitness values of the M ants (the M second ants in the previous iteration) before the update of the t-th position encoding, the M ants input to the next round are updated to obtain the updated M ants (i.e., the M second ants in this iteration). Among them, the position encodings of the M second ants in this iteration are updated in the (t + 1)-th iteration.

[0219] Based on the fitness values of the updated M ants, the elite ants are updated, and the pheromone is updated using the elite ants. The updated pheromone and the elite ants are used to determine the position encodings of the M ants in the (t + 1)-th iteration.

[0220] Enter the (t + 1)-th iteration, return to step S3, and continue to determine whether t = t + 1 is less than the maximum number of iterations T.

[0221] Step S10: Generate an anomaly detection model according to the elite ants determined in the last iteration.

[0222] Decode the position encoding of the elite ant obtained in the last iteration to obtain the corresponding isolation forest, and generate an anomaly detection model based on this isolation forest.

[0223] Based on the same technical concept, the present disclosure provides a high-dimensional anomaly data detection device. Figure 5 It is a schematic diagram of a high-dimensional anomaly data detection device shown in an embodiment of the present disclosure. As Figure 5 shown, the device includes:

[0224] A determination module 510, configured to determine the position encodings of the M first ants in this iteration. The position encoding of each ant includes: a discrete encoding group for determining an isolation tree feature subset, and a continuous encoding group for determining an isolation tree anomaly threshold and an isolation tree weight. Each feature in the isolation tree feature subset represents a way of splitting a high-dimensional data set;

[0225] A construction module 520, configured to construct M isolation forests according to the position encodings of the M first ants in the current iteration.

[0226] An evaluation module 530, configured to evaluate the effectiveness of the M isolation forests respectively based on high-dimensional anomaly data detection, and determine the fitness values corresponding to the M first ants in the current iteration, where the fitness values represent the detection effects on the anomaly data in the high-dimensional dataset.

[0227] An update module 540, configured to determine the elite ants in the current iteration according to the fitness values corresponding to the M first ants in the current iteration and the fitness values corresponding to the M second ants in the previous iteration, where the elite ants are used to determine the position encodings of the M first ants in the next iteration.

[0228] An iteration module 550, configured to perform multiple iterations according to the above steps, and determine the final target position encoding according to the position encoding corresponding to the elite ants determined in the last iteration. The target position encoding includes: optimal feature subset selection, isolation tree anomaly threshold, and isolation tree weight.

[0229] An end module 560, configured to construct an anomaly detection model according to the final target position encoding, and perform segmentation on the high-dimensional dataset based on the data segmentation method corresponding to the anomaly detection model, so as to obtain an anomaly detection result.

[0230] The embodiment of the present disclosure further provides an electronic device. Refer to Figure 6 , Figure 6 which is a schematic diagram of an electronic device shown in the embodiment of the present disclosure. As Figure 6 shown, the electronic device 600 includes: a memory 610 and a processor 620. The memory 610 is communicatively connected to the processor 620 through a bus. A computer program is stored in the memory 610, and the computer program can run on the processor 620, so as to implement the steps in the high-dimensional anomaly data detection method disclosed in the embodiment of the present disclosure.

[0231] The embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the high-dimensional anomaly data detection method disclosed in the embodiment of the present disclosure are implemented.

[0232] The embodiment of the present disclosure further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps in the high-dimensional anomaly data detection method disclosed in the embodiment of the present disclosure are implemented.

[0233] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.

[0234] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, apparatus, or computer program product. Therefore, the embodiments of the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0235] The embodiments of the present disclosure are described with reference to the flowcharts and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0236] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0237] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal devices, such that a series of operation steps are executed on the computer or other programmable terminal devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable terminal devices provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0238] Although some embodiments of the present disclosure have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the embodiments of the present disclosure.

[0239] The above provides a detailed introduction to a high-dimensional anomaly data detection method provided by the present disclosure. Specific examples are used in this article to elaborate on the principle and implementation manner of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure; at the same time, for those of ordinary skill in the art, according to the idea of the present disclosure, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present disclosure.

Claims

1. A high-dimensional abnormal data detection method, characterized in that: include: Determine the position codes of the M first ants in this round of iteration, where the position code of each ant includes: a discrete code group for determining an isolated tree feature subset, and a continuous code group for determining an isolated tree abnormality threshold and an isolated tree weight, where each feature in the isolated tree feature subset represents a segmentation method for a high-dimensional data set; Constructing M isolation forests according to the position codes of the M first ants in the current iteration; Based on high-dimensional abnormal data detection, the effectiveness of the M isolation forests is evaluated respectively, and the fitness value corresponding to each of the M first ants in the current iteration is determined, and the fitness value represents the detection effect of abnormal data in the high-dimensional data set; Determine the elite ant of this iteration according to the fitness values ​​corresponding to the M first ants of this iteration and the fitness values ​​corresponding to the M second ants of the previous iteration, wherein the elite ant is used to determine the position codes of the M first ants of the next iteration; Perform multiple iterations according to the above steps, and determine the final target position code based on the position code corresponding to the elite ant determined in the last round of iteration. The target position code includes: optimal feature subset selection, isolated tree anomaly threshold and isolated tree weight; According to the final target position code, an anomaly detection model is constructed, and based on the data segmentation method corresponding to the anomaly detection model, the high-dimensional data set is segmented to obtain an anomaly detection result.

2. The method according to claim 1, characterized in that Determining the elite ant of the current iteration according to the fitness values ​​corresponding to the M first ants of the current iteration and the fitness values ​​corresponding to the M second ants of the previous iteration comprises: Sorting the M first ants in the current iteration and the M second ants in the previous iteration according to the fitness values ​​corresponding to the M first ants in the current iteration and the fitness values ​​corresponding to the M second ants in the previous iteration to obtain a first sorting result; Determine the first M ants in the first sorting result as the M second ants in this round of iteration; The ant with the highest fitness among the M second ants in the current iteration is determined as the elite ant in the current iteration.

3. The method according to claim 2, characterized in that Also includes: Sorting the M first ants in the current iteration according to the fitness values ​​corresponding to the M first ants in the current iteration to obtain a second sorting result; Taking half of the ants with high fitness values ​​in the second sorting result corresponding to the M first ants in the current iteration as a better ant pool, and taking half of the ants with low fitness values ​​in the second sorting result as a worse ant pool; According to a preset matching number, ants are selected from the better ant pool and the worse ant pool respectively to obtain a matching number of matching ant groups; each matching ant group includes an ant from the better ant pool and an ant from the worse ant pool; For each matching ant group, the two ants included in the group exchange their continuous codes to obtain new ants; According to the fitness values ​​corresponding to the M first ants of the current iteration and the fitness values ​​corresponding to the M second ants of the previous iteration, the M first ants of the current iteration and the M second ants of the previous iteration are sorted to obtain a first sorting result, including: The new ants, the M first ants in the current iteration, and the M second ants in the previous iteration are sorted according to their respective corresponding fitness values ​​to obtain a first sorting result.

4. The method according to claim 1, characterized in that: According to the position codes of the M first ants in the current iteration, M isolation forests are constructed, including: according to the position code of each first ant, an isolation forest is constructed; Each isolation forest is constructed as follows: Decoding each discrete code in the discrete code group that is sorted in sequence to obtain each feature that is sorted in sequence; According to the order of each feature, determine the feature subset of each isolated tree in the isolated forest corresponding to the first ant; Decoding the first continuous code in the continuous code group to obtain the isolated tree abnormality threshold of the first ant; Decoding the remaining sequentially ordered continuous codes in the sequentially ordered group to obtain respective weight values ​​that are sequentially ordered; Determine the weight of each isolated tree in the isolated forest corresponding to the first ant according to the order of each weight value; An isolated forest corresponding to the first ant is constructed according to the feature subsets corresponding to the isolated trees, the weights corresponding to the isolated trees, and the isolated tree abnormality threshold.

5. The method according to claim 1, characterized in that: Determine the position codes of the M first ants in this round of iteration, including: According to the elite ants obtained in the previous round of iteration, the target ant selects feature subsets corresponding to each isolated tree in the isolated forest from the feature selection space, wherein the feature subsets include multiple features; wherein the feature subsets are sorted in order, and each feature in the feature subsets is sorted in order, and the target ant is any one of the M second ants in the previous round of iteration; Encoding the features in the feature subset according to a first encoding format to obtain discrete codes of the features in the discrete code group; The target ant updates the position vector of the target ant according to the wandering range and wandering step length of this iteration and based on the elite ant obtained in the previous iteration; According to the position vector of the target ant, encoding is performed according to a second encoding format to obtain a continuous encoding group corresponding to the target ant; The position code of any one of the M first ants in the current iteration is determined according to the discrete code group composed of the discrete codes and the continuous code group.

6. The method according to claim 5, characterized in that According to the elite ants obtained in the previous iteration, the target ants select feature subsets corresponding to each isolated tree in the isolation forest from the feature selection space, including: According to the discrete coding group corresponding to the elite ants obtained in the previous iteration, the pheromone concentration of each feature in the feature selection space in this iteration is determined; Determine the number of selection steps of the target ant, where the number of selection steps is used to determine the number of features included in the feature subset; In the case where the target ant selects a feature from the optional features in the feature search space at each step, determining the selection probability of each optional feature according to the pheromone concentration and heuristic information of each optional feature; the optional feature represents the feature that is not selected; Based on the selection probability of the optional feature, select an optional feature as the feature corresponding to the step according to the roulette strategy; According to the features corresponding to each step, determine the feature subset of the isolated tree; After determining a feature subset of an isolated tree, the selected state of the features in the feature selection space is cleared.

7. The method according to claim 6, characterized in that In the case where the target ant selects a feature from the optional features in the feature search space at each step, determining the selection probability of each optional feature according to the pheromone value and heuristic information of each optional feature, including: Determine a multiple selection coefficient of the optional feature, wherein the multiple selection coefficient is used to control the diversity in the feature selection process; a value corresponding to the multiple selection coefficient of the feature is determined according to the number of times the feature is selected in this round of iteration; The selection probability of each optional feature is determined according to the multiple selection coefficients, pheromone values ​​and heuristic information of the optional features.

8. The method according to claim 6, characterized in that According to the discrete coding group corresponding to the elite ants obtained in the previous iteration, the pheromone concentration of each feature in the feature selection space in this iteration is determined, including: Determine a feature subset of the elite ant according to the discrete coding group corresponding to the elite ant obtained in the previous iteration; Performing pheromone enhancement processing on the features of the feature subset corresponding to the elite ants, and performing pheromone volatilization processing on the features of the feature subsets corresponding to the remaining ants obtained in the previous round of iterations; The pheromone concentration of each feature in the feature selection space of this round of iteration is determined through the pheromone enhancement process and the pheromone volatilization process.

9. The method according to claim 5, characterized in that The target ant updates the position vector of the target ant according to the wandering range and wandering step length of this iteration based on the elite ant obtained in the previous iteration, including: Selecting a roulette ant from the M second ants of the previous iteration in a roulette manner; According to the wandering range and wandering step length of this round of iteration, the target ant respectively surrounds the elite ant and the roulette ant obtained in the previous round of iteration to determine the elite position vector and the roulette position vector; The position vector of the target ant is updated according to the elite position vector and the roulette position vector.

10. The method according to claim 5 or 9, characterized in that: The walking step length and the walking range are determined by the following steps: Determine preset upper and lower bounds of the dimension of the target ant; Determine the wandering range of the target ant according to the upper and lower bounds of the dimension of the target ant, the current number of iterations corresponding to this round of iterations, and the maximum number of iterations; According to the current iteration number corresponding to the current iteration, the walking step length of the target ant in the current iteration is determined.

Citation Information

Patent Citations

  • Group intelligence optimization rolling force prediction method for strip steel cold continuous rolling

    CN117609722A

  • Intelligent data driving optimization method for electric reactor vibration analysis

    CN117909721A