Anomaly detection apparatus and method using data augmentation

The anomaly detection device and method utilize data augmentation to enhance accuracy in identifying gray behavior by generating mathematical combinations of behavioral features, addressing inefficiencies in existing security systems and improving detection in zero-trust environments.

JP2026002763AActive Publication Date: 2026-01-08RUIKONG NETWORK SECURITY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025068927
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-21
Filing Date
2025-04-18
Publication Date
2026-01-08
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Existing information security systems struggle to accurately identify gray behavior or grayware, leading to false positives or false negatives, and are inefficient in zero-trust environments due to the complexity of verifying access to software resources and assets.

Method used

An anomaly detection device and method using data augmentation that collects data records from supervised training environments, performs activation state determination, and generates mathematical combinations of behavioral features to enhance anomaly detection accuracy.

Benefits of technology

Significantly reduces human resource requirements for data collection and improves the accuracy of detecting gray behavior anomalies, enhancing security in zero-trust environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026002763000001_ABST
    Figure 2026002763000001_ABST
Patent Text Reader

Abstract

To provide an abnormality detection device and an abnormality detection method using data extension.SOLUTION: The abnormality detection device 100 using data augmentation includes a continuous data collection module and a processor. The processor performs, on the collected data records, an activation state determination on the plurality of behavior information using a plurality of activation functions associated with a plurality of behavior types; Generating a set of behavior types specific to each subject, performing data augmentation by enumerating, for each subject, a plurality of mathematical combinations of the behavior types, training a machine learning model of a plurality of baseline behavior features with an output of the data augmentation as an input of a learning function, and performing anomaly detection with a new output of the data augmentation performed based on all new data records captured from the test subject by the continuous data collection module within a predefined time range as an input of a prediction function.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of information security technology, and more particularly to an anomaly detection device and method using data augmentation. [Background technology]

[0002] In the field of traditional information security, in addition to computer virus attacks and Trojan horse program attacks, various types of grey behavior and greyware exist in various network environments, network programs, and processing systems. Grey behavior and greyware refer to any behavior or software that exhibits suspicious behavior that is not a computer virus or Trojan horse program but that adversely affects the effectiveness of various network environments, network programs, and processing systems and may cause damage to information security. Summary of the Invention [Problem to be solved by the invention]

[0003] However, in the past, the information security field has been unable to effectively identify whether a subject or data having gray behavior or grayware is benign or abnormal.

[0004] Furthermore, in a zero trust environment, each access to a system's individual software resources (e.g., software programs, firmware programs, etc.), the systems used by users (e.g., network systems, wireless communication systems, etc.), and the company's assets themselves (e.g., the company itself, factories, aircraft, etc.) is considered as the subject, and each access to the subject is verified. However, the ownership and access location of the individual software resources, the systems used by users, and the company's assets themselves are not simply verified. In other words, in a zero trust environment, each user must verify each of these individual access actions before accessing each company's resources. However, in the past, the information security field was unable to effectively identify or verify whether gray activities or grayware in a zero trust environment passed verification.

[0005] The present invention has been made in view of the above circumstances, and aims to solve the above problems. That is, an object of the present invention is to provide an anomaly detection device and method that utilizes data augmentation, which significantly reduces the human resources required for collecting data and significantly improves the accuracy of detecting anomalies in gray behavior. [Means for solving the problem]

[0006] In order to achieve the above object, an anomaly detection device using data augmentation according to one aspect of the present invention comprises: a continuous data collection module configured to capture a plurality of data records of a plurality of subjects in a training environment or a normal work environment under supervision to perform a normal task, each data record including a time, a subject, a plurality of behavioral types, and a plurality of behavioral information corresponding to each of the behavioral types; configured to store machine learning models of the plurality of instructions and the plurality of baseline behavioral features; a) a learning function step that takes as input a set of said behavior types of a predefined length N and adds said set to a plurality of said baseline behavior features; b) a prediction function step that takes as input the new set of behavior types of the defined length N and outputs a match state; a memory connected to the continuous data collection module and the memory, for each data record: performing activation state determination for the plurality of behavioral information using a plurality of activation functions associated with the plurality of behavioral types, and combining all of the behavioral types in activation states for each of the subjects to generate the set of behavioral types specific to each of the subjects, each of the sets representing a plurality of behavioral features of the corresponding subject; performing data augmentation by enumerating, for each subject, all mathematical combinations of the predefined lengths N in the set of behavior types based on the predefined lengths N used in the machine learning model of the plurality of baseline behavioral features, each mathematical combination being a subset of the predefined lengths N of the plurality of behavioral features of the subject; In the training environment, after data collection training is completed, training the machine learning model of the plurality of baseline behavioral features during a model training period using the output of the data augmentation as an input to the learning function; and a processor configured to execute a plurality of said instructions to perform the steps of: in said normal working environment, taking as input to said prediction function a new output of data augmentation based on all new data records captured by said continuous data collection module from a test subject within a defined time range; and indicating as an anomalous event if the return of said prediction function indicates an abnormal said match state.

[0007] In order to achieve the above object, another aspect of the present invention is a method for anomaly detection using data augmentation, comprising: capturing, by a continuous data collection module, a plurality of data records of a plurality of subjects in a training environment where the subject is supervised to perform a normal task or in a normal work environment, each of the data records including a time, a subject, a plurality of behavioral types, and a plurality of behavioral information corresponding to each of the plurality of behavioral types; and performing, by a processor, for each of the data records, a determination of activation state for the plurality of the behavioral information using a plurality of activation functions associated with a plurality of the behavioral types, and combining all the behavioral types in activation state for each of the subjects to generate a set of the behavioral types specific to each of the subjects, each of the sets representing a plurality of behavioral characteristics of the corresponding subject; performing data augmentation by the processor for each data record by enumerating, for each subject, all mathematical combinations of the predefined lengths N in the set of behavior types based on predefined lengths N used in a machine learning model of a plurality of baseline behavioral features, each mathematical combination being a subset of the predefined lengths N of the plurality of behavioral features of the subject; In the training environment, after completing data collection training, the processor trains the machine learning model on the plurality of baseline behavioral features during a model training period using the output of the data augmentation as input to a learning function, wherein the learning function takes the set of behavior types of the defined length N as input and adds the set to the plurality of baseline behavioral features; In the normal working environment, the processor uses a new output of the data augmentation performed by the continuous data collection module based on all new data records captured from the test subject within a defined time range as input to a prediction function, and if the return of the prediction function indicates an abnormal match state, indicate it as an anomalous event, and the prediction function takes a new set of the behavior types of the defined length N as input and outputs the match state. [Effects of the Invention]

[0008] The present invention is configured as described above and therefore provides the following effects. Compared with related technologies, the technical effect achieved by the present disclosure is to directly use combination processing to perform data augmentation, thereby significantly saving the human resources required for data collection and significantly improving the accuracy of gray behavior anomaly detection.

[0009] At least the following points will become clear from the description and drawings to be described later. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram illustrating an anomaly detection apparatus using data augmentation in accordance with some embodiments of the present disclosure. [Figure 2] FIG. 1 is a schematic diagram illustrating a baseline set in some embodiments of the present disclosure. [Figure 3] FIG. 1 is a schematic diagram illustrating a validation set in some embodiments of the present disclosure. [Figure 4] 1 is a flowchart illustrating a method for anomaly detection using data augmentation in accordance with some embodiments of the present disclosure. [Figure 5] FIG. 1 is a schematic diagram illustrating determining an activation state from multiple behavioral data in some embodiments of the present disclosure. [Figure 6] FIG. 1 is a schematic diagram illustrating the generation of a set of mathematical combinations in some embodiments of the present disclosure. [Figure 7] 10 is a flowchart illustrating detailed steps of further steps included in an anomaly detection method using data augmentation according to some embodiments of the present disclosure. [Figure 8] FIG. 10 is a block diagram illustrating an anomaly detection device using data augmentation in some other embodiments of the present disclosure. [Figure 9] 10 is a flowchart illustrating detailed steps of further steps included in an anomaly detection method using data augmentation according to some other embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] The following describes in detail the embodiments of the present invention, but the present invention is not limited to these, and various modifications are possible within the scope of the description. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

[0012] In traditional information security, gray behavior is extremely difficult to identify. Furthermore, many benign subjects and data often exhibit gray behavior, leading to false positives. Because it is difficult to determine whether a subject's gray behavior is benign or malicious, information security systems often adopt overly lenient or overly strict policies to identify gray behavior. However, when information security systems adopt overly lenient policies (e.g., whitelists), benign behaviors in the whitelist can be exploited as loopholes or transformed into malicious software and used as a tool for network attacks. When information security systems adopt overly strict policies, they often generate a large number of false alarms, causing system administrators to waste time checking for false alarms.

[0013] In addition, in traditional zero-trust architectures, information security systems reduce information security risks by verifying the subjects that generate gray behavior. However, the items that information security systems must verify for gray behavior are often very complex. In addition, if an information security system adopts verification settings that are either too lenient or too strict, this can become a weakness in zero-trust architectures.

[0014] To solve the above-mentioned problems, the anomaly detection device and method using data augmentation according to the present disclosure uses the concept of mathematical combination to perform data augmentation. Specifically, compared to focusing on individual gray behaviors, the present disclosure collects all data records of benign subjects at a specific time, performs activation state determination on all behavioral information in these data records, and performs data augmentation by using a mathematical combination method to generate all mathematical combinations based on all behavior types in the activation state. In this way, the anomaly detection device and method utilize a very large number of mathematical combinations to detect anomalies in a zero-trust environment, thereby achieving the effect of accurately detecting whether gray behaviors are anomalous. Furthermore, the anomaly detection device and method further utilize the technical feature of detecting verification values ​​to perform anomaly detection on test subjects, thereby achieving the effect of accurately detecting whether the test subject is anomalous.

[0015] [Example of an anomaly detection device using data augmentation] FIG. 1 is a block diagram illustrating an anomaly detection device 100 using data augmentation according to some embodiments of the present disclosure. The anomaly detection device 100 may be implemented as any electronic device or server (e.g., an individual user's or company's processing device, a cloud device, a server, a cloud server, etc.), and the anomaly detection device 100 may be connected to a network in a zero trust environment. The anomaly detection device 100 includes a continuous data collection module 110, a processor 120, and a memory 130 (see FIG. 1). The processor 120 is connected to the continuous data collection module 110 and the memory 130.

[0016] In this embodiment, the continuous data collection module 110 is used to capture multiple data records d1-dM of multiple subjects in a supervised training environment or a regular operation environment. Each data record d1-dM includes a time, a subject, multiple behavioral types, and multiple behavioral information corresponding to each of the behavioral types, with each behavioral type simultaneously including multiple behavioral information. The supervised training environment is used to collect multiple data records d1-dM that are recognized as benign subjects, while the regular operation environment is used to detect anomalies in test subjects whose benign or malignant status is unknown. Note that M may be any positive integer and is not particularly limited.

[0017] Specifically, the continuous data collection module 110 is used only to capture data records d1-dM in a training environment or a normal work environment where the user is supervised to perform normal tasks, and does not have the ability to distinguish whether these data records d1-dM are benign or malicious. The present disclosure enables the continuous data collection module 110 to operate normally, be controllable, and continuously capture data records d1-dM in an environment where only benign subjects exist (i.e., the training environment where the user is supervised to perform normal tasks). In this way, the data records d1-dM captured by the continuous data collection module 110 all belong to data related to gray behavior that has been identified as benign subjects. In some embodiments, the benign subjects are individual software resources of a system (e.g., software programs, firmware programs, etc.) that are free of anomalies (i.e., free of attack behavior), systems used by users (e.g., network systems, wireless communication systems, etc.), or corporate assets themselves (e.g., the company itself, a factory, an aircraft, etc.).

[0018] In some embodiments, the continuous data collection module 110 captures data records d1-dM from the sample database SD. In some embodiments, the time of each data record d1-dM is the time when the behavioral information was captured. In some embodiments, the subject of each data record d1-dM is the subject of the behavioral information generated. In some embodiments, each behavioral information of each data record d1-dM is displayed as numerical data. In some embodiments, each behavior type of each data record d1-dM is displayed as a gray behavior policy (e.g., a policy shared across the entire network), and the gray behavior policy was used to generate the behavioral information. In some embodiments, each behavior type of each data record d1-dM has a corresponding activation state after processing the activation function described below, and the activation state of each behavior type is displayed as a Boolean value.

[0019] In some embodiments, the sample database SD may be any database storing data records d1-dM of multiple benign subjects (e.g., an open-source whitelist database storing behavioral characteristics of application programs or corporate assets). In some embodiments, the continuous data collection module 110 is installed in an environment where only benign subjects exist, and pre-collects data records d1-dM from multiple benign subjects and stores them in the sample database SD.

[0020] In some embodiments, the gray behavior policy may be a policy shared across all networks to which the anomaly detection device 100 is connected (e.g., the policy is stored on all routers in the network), and the gray behavior policy may be a behavioral rule, such as the number of times a user touches something within an hour, the number of times a user opens a web page within a day, or the number of times a host is continuously connected to the network. Multiple gray behavior policies may be used to identify multiple pieces of behavioral information, each corresponding to a time, a subject, multiple behavior types, and multiple behavioral types, for all data records from multiple benign subject operation behaviors (e.g., touching the interface of a specific software program five times within an hour at 10:00 on March 1, 2020). The behavioral information for each data record d1 through dM can be considered a feature vector.

[0021] For example, the plurality of pieces of behavioral information may include numerical data such as a user touching 10 times within an hour, a user opening a web page 20 times in a day, and a host continuously connecting to a network 80 times, and these numerical data may be regarded as a feature vector. The time is the time when these numerical data were generated (e.g., 10:00 on March 1, 2020). The subject is the subject generated by these numerical data (e.g., a specific software program). The plurality of behavioral types may represent gray behavior policies such as the number of times a user touched within an hour, the number of times a user opened a web page in a day, and the number of times a host was continuously connected to a network.

[0022] In some embodiments, each piece of behavioral information corresponds to one gray behavioral policy, and one gray behavioral policy is used to generate the corresponding piece of behavioral information.

[0023] For example, the plurality of gray behavior policies may include a first gray behavior policy for the number of times a user touches within an hour, a second gray behavior policy for the number of times a user opens a web page within a day, and a third gray behavior policy for the number of times a host is continuously connected to a network. In the first embodiment, for example, the two pieces of behavioral information for a benign subject detected include that the user touches 30 times within an hour and that the user opens a web page 20 times within a day. That is, the behavior of this benign subject is associated with the first gray behavior policy and the second gray behavior policy.

[0024] In the second embodiment, for example, the other benign subject is detected as two other behavioral information items, such as a user opening a web page 50 times in one day and a host continuously connecting to a network 80 times, and the behavior of this benign subject is associated with the second gray behavior policy and the third gray behavior policy.

[0025] In this example, memory 130 stores a machine learning model BM of multiple baseline behavioral features and multiple instructions, and processor 120 executes the detailed steps described below based on the instructions. In some embodiments, the machine learning model BM is trained to identify whether a gray behavior is anomalous based on the multiple baseline behavioral features, and therefore can be considered a machine learning model BM having multiple baseline behavioral features. The machine learning model BM of multiple baseline behavioral features and the instructions may be corresponding software or firmware instruction programs. In some embodiments, the instructions may be software or firmware instruction programs of gray behavior resident programs, which may be implemented as operating system hooks. In some embodiments, the machine learning model BM of multiple baseline behavioral features may be any machine learning model (e.g., an artificial neural network model, a convolutional neural network model, etc.).

[0026] In some embodiments, memory 130 includes a validation database AD, which stores multiple baseline behavioral features (not shown). The multiple baseline behavioral features include multiple sets (hereinafter referred to as baseline sets) each containing multiple behavioral types detected from multiple other benign subjects (i.e., other subjects that have already been determined not to be abnormal). Each baseline set detects a set of behavioral types generated by one of the other benign subjects based on multiple predetermined gray behavior policies. Specifically, continuous data collection module 110 pre-captures multiple other data records of the multiple other benign subjects. The other data records also include time, subject, multiple behavioral types, and multiple behavioral information corresponding to each of the multiple behavioral types. Processor 120 then compiles the multiple behavioral types of each of the other data records into a baseline set. In some embodiments, the other benign subjects may also be individual software resources of a system that are not abnormal, a system used by a user, or an enterprise asset itself. In some embodiments, the other benign subjects may be the same as or different from the benign subjects described above.

[0027] First Example of a Baseline Set of Baseline Behavioral Features The baseline set will be described below with reference to an actual example. Figure 2 is a schematic diagram showing a baseline set rc in some embodiments of the present disclosure. As shown in Figure 2, the baseline set rc includes multiple behavior types rf1 to rf20. The behavior types rf1 to rf20 represent multiple corresponding gray behavior policies.

[0028] In some embodiments, each of the multiple baseline sets includes a validation set. In some embodiments, each validation set includes multiple validation values. In some embodiments, the multiple validation values ​​for each validation set may be determined by detecting other benign subjects corresponding to each baseline set based on multiple validation policies. For example, a management server in a network or a distributed management server may detect other benign subjects based on multiple validation policies to generate a single validation set, generate detection results, and generate the multiple validation values ​​for the single validation set based on the detection results.

[0029] In some embodiments, a validation policy may be a policy for managing and validating a subject, such as whether the subject has a trusted digital signature, whether the subject was built at a particular time, or whether an administrator is logged into the subject's management console. Thus, a validation value may be a Boolean value that is logically true or false (a validation value is a logical "true" or "false").

[0030] Second Example of a Baseline Set of Baseline Behavioral Features The following describes the validation set with a practical example. Figure 3 is a schematic diagram showing a validation set VC in some embodiments of the present disclosure. As shown in Figure 3, the baseline set rc further includes a validation set VC, and multiple reference validation values ​​of the validation set VC are generated by detecting other benign subjects corresponding to the baseline set based on multiple validation policies.

[0031] Returning to FIG. 1 , in this embodiment, the memory 130 executes the learning function and the prediction function using the machine learning model BM. The learning function trains the machine learning model BM by taking a set of behavior types of a predefined length N for a known subject and adding the set to multiple baseline behavioral features. The prediction function inputs a new set of behavior types of a predefined length N for an unknown test subject and outputs a match status to determine whether the test subject's behavior is abnormal. In some embodiments, N may be equal to or less than the number of behavior types in the data record, and the present disclosure performs data augmentation by enumerating the subject's behavior types using the predefined length N. The learning function and the prediction function will be further described in subsequent paragraphs, so their description will be omitted here.

[0032] In some embodiments, the continuous data collection module 110 may be any data capture software, firmware, hardware, or combination thereof. For example, the continuous data collection module 110 may be one or a combination of a network interface that continuously monitors a network, a network card driver program, a network application program, a transmission circuit, an A / D converter, a D / A converter, a low-noise amplifier, a mixer, a filter, an impedance matcher, a transmission cord, a power amplifier, one or more antenna circuits, and a local storage media element, but the present invention is not limited thereto. In some embodiments, the continuous data collection module 110 may continuously observe and record the gray behavior of any subject during an activity period.

[0033] In some embodiments, memory 130 may be implemented as, but is not limited to, a memory unit, flash memory, read-only memory, a hard disk, or a storage kit having the same. In some embodiments, processor 120 may be implemented as, but is not limited to, a central processing unit (CPU), a microcontrol unit (MCU), a programmable logic controller (PLC), a system on chip (SoC), or a field programmable gate array (FPGA).

[0034] [Example of anomaly detection method using data augmentation] 4 is a flowchart illustrating an anomaly detection method using data augmentation according to some embodiments of the present disclosure. This anomaly detection method is applied to the anomaly detection device 100 shown in FIG.

[0035] As shown in FIG. 4, the anomaly detection method includes steps S410 to S450. First, in step S410, the continuous data collection module 110 captures data records d1 to dM of multiple subjects in a training environment or a normal work environment where subjects are supervised to perform tasks normally. In this embodiment, as described above, the data records d1 to dM each include time, a subject, multiple behavioral types, and multiple behavioral information corresponding to the multiple behavioral types. In the training environment where subjects are supervised to perform tasks normally, the continuous data collection module 110 captures data records d1 to dm of the sample database SD or data records d1 to dM of benign subjects to train the machine learning model BM. In the normal work environment, the continuous data collection module 110 captures data records d1 to dM of unknown subjects to detect whether the behavior of the unknown subjects is abnormal.

[0036] In step S420, for each captured data record, processor 120 performs activation determinations on the multiple behavioral information using multiple activation functions associated with multiple behavioral types (i.e., one behavioral type associated with one activation function), and combines all active behavioral types for each subject to generate a subject-specific set of behavioral types. In this example, the sets represent multiple behavioral characteristics of the corresponding subject. In other words, all active behavioral types for each data record can be considered as a subject-specific set of behavioral types corresponding to the data record. In some embodiments, each activation function is used to perform activation determinations (i.e., one-to-one activation determinations) on the behavioral information corresponding to the associated behavioral type.

[0037] In some embodiments, the activation function may be any type of activation function (e.g., a unit step function, a rectified linear unit (ReLU), a Softmax function, or a behavioral feature comparison function (i.e., matching a predetermined behavioral feature means matching the activation state), etc.). In other words, the activation function may be a function that compares thresholds, a function that compares whether a predetermined behavioral feature (e.g., 10 activations) is matched, etc. It should be noted that the activation function acts to convert the numerical data of observed gray behaviors (different gray behaviors have heterogeneous numerical distributions) into a data that can represent the homogeneity of a set of behavior types.

[0038] For example, if the behavioral information indicates that a user touches a web page 10 times within an hour and the activation function is a unit step function that sets a value of 1 for values ​​greater than 5 and a value of 0 for values ​​less than 5, the behavioral information is processed by the unit step function and converted to 1. Thus, the behavior type corresponding to the behavioral information (i.e., the number of times the user touches a web page within an hour) is considered to be in an activated state (i.e., the activation state of the behavior type is represented by a Boolean value with a logical value of 1) and is included in the set. In another example, if the behavioral information indicates that a user opens a web page 2 times within a day and the activation function is a unit step function that sets a value of 1 for values ​​greater than 3 and a value of 0 for values ​​less than 3, the behavioral information is processed by the unit step function and converted to 0. Thus, the behavior type corresponding to the behavioral information (i.e., the number of times the user opens a web page within a day) is considered to be in an inactivated state (i.e., the activation state of the behavior type is represented by a Boolean value with a logical value of 0) and is not included in the set. In another example, if the activation function is a behavioral feature comparison function, the behavioral feature comparison function compares predefined behavioral features with the behavioral information. If the predefined behavioral features match the behavioral information, the behavior type corresponding to the behavioral information is considered to be activated and included in the set. Conversely, if the predefined behavioral features do not match the behavioral information, the behavior type corresponding to the behavioral information is considered to be in an unactivated state and is not included in the set.

[0039] The following provides an example to explain the relationship between the behavior information and a set of behavior types specific to a subject. Figure 5 is a schematic diagram illustrating determining the activation state from a plurality of pieces of behavior information bf1 to bf20 in some embodiments of the present disclosure. As shown in Figure 5, when determining the activation state using a plurality of activation functions, the processor 120 obtains the behavior types tf2, tf8, tf10, and tf17 to tf18 in the activation state from the plurality of pieces of behavior information bf1 to bf20, and the behavior types tf2, tf8, tf10, and tf17 to tf18 correspond to the behavior information bf2, bf8, bf10, and bf17 to bf18, respectively. In other words, the 15 behavior types corresponding to the behavior information bf1, bf3-bf7, bf9, bf11-bf16, and bf19-bf20, respectively, are removed by the corresponding activation functions, and the behavior types tf2, tf8, tf10, and tf17-tf18 are taken as the set of behavior types specific to the subject. That is, the set includes the five behavior types tf2, tf8, tf10, tf17, and tf18 that are in the activated state.

[0040] Returning to FIG. 4 , in step S430, for each data record, processor 120 performs data augmentation by enumerating all mathematical combinations of the predefined length N in the set of behavior types for each subject based on the predefined length N used in the machine learning model BM of the multiple baseline behavioral features. In this example, each mathematical combination is a subset of the predefined length N of the multiple behavioral features of the subject. In some embodiments, N is any positive integer greater than or equal to 1. For example, if the set of behavior types includes five behavior types and the predefined length N is 3, processor 120 selects all mathematical combinations of any three behavior types from the five active behavior types (i.e., the quantity of all mathematical combinations is 5C3).

[0041] In some embodiments, all mathematical combinations are considered as one hybrid combination, i.e., gray behaviors can be directly predicted to occur in any order in the future, and fragments of gray behaviors can be observed. In some embodiments, processor 120 updates the plurality of baseline behavioral features by adding all of the mathematical combinations. In some embodiments, processor 120 updates the plurality of baseline behavioral features by adding mathematical combinations that are different from the plurality of baseline behavioral features.

[0042] It should be noted that the larger the predefined length N, the higher the similarity between all behavior types in the new data record and data records d1 to dM must be in the subsequent stage before it can be determined that there is no abnormality (if the predefined length N is appropriate, the detection rate will improve and the false alarm rate will decrease in the subsequent abnormality determination).The smaller the predefined length N, the lower the similarity between all behavior types in the new data record and data records d1 to dM, and it may still be determined that there is no abnormality (i.e., sensitivity).

[0043] In some embodiments, for each data record, processor 120 performs data augmentation by enumerating each mathematical combination of the plurality of defined lengths N of the set of behavior types for each subject based on the plurality of defined lengths N used in the machine learning model BM of the plurality of baseline behavioral features. In some embodiments, the plurality of defined lengths N may be all positive integers in any one numerical interval (e.g., the plurality of defined lengths N may be all positive integers less than 10, all positive integers greater than 2 and less than the quantity of all active behavior types, or all positive integers with the selected quantity less than 10 and greater than 1).

[0044] It should be noted that the higher the upper limit of the numerical range, the higher the similarity between all the behavioral types of the new data record and the data records d1 to dM must be in the subsequent stage in order to determine that there is no abnormality (if the upper limit is appropriate in the subsequent abnormality determination, the detection rate will improve and the false alarm rate will decrease). The lower the lower limit of the numerical range, the lower the similarity between all the behavioral types of the new data record and the data records d1 to dM, and there is still a possibility that there is no abnormality (i.e., sensitivity). The process of enumerating the above-mentioned combinations is a conventional technique in the field of mathematics, so its explanation will be omitted here.

[0045] For example, if the numerical range is all positive integers less than 4 and the number of all active behavior types is 5, processor 120 may enumerate all mathematical combinations of three active behavior types among the five active behavior types as a first set (i.e., the number of mathematical combinations is 5C3). Next, processor 120 may enumerate all mathematical combinations of two active behavior types among the five active behavior types as a second set (i.e., the number of mathematical combinations is 5C2). Then, processor 120 may enumerate all mathematical combinations of one active behavior type among the five active behavior types as a third set (i.e., the number of mathematical combinations is 5C1). In this case, the total number of all mathematical combinations (i.e., the plurality of sets) is 25 (i.e., 5C3 + 5C2 + 5C1). In subsequent anomaly detection, this selection method has the best detection rate and the lowest false alarm rate.

[0046] For another example, if the numerical range is all positive integers less than 2 and the number of active behavior types is 5, the processor 120 will list all mathematical combinations of one active behavior type from the five active behavior types as one set (i.e., the number of mathematical combinations is 5C1). In this case, the total number of all mathematical combinations of the active behavior types is 5 (i.e., 5C1). This selection method is similar to the effect of a conventional whitelist.

[0047] [Example of all mathematical combinations of defined length N] Below, we will give practical examples to illustrate all mathematical combinations of the defined length N. 6 is a schematic diagram illustrating the generation of a set of mathematical combinations Cs in some embodiments of the present disclosure. As shown in FIG. 6, when the example in FIG. 5 is extended and the defined length N is 3, the processor 120 enumerates mathematical combinations bc1 to bc10 of 10 active behavior types from the behavior types tf2, tf8, tf10, and tf17 to tf18 (i.e., the active behavior types) of the set of mathematical combinations Cs by performing a 3 / 5 mathematical combination process (i.e., the following quantity is 5C3).

[0048] The mathematical combination bc1 of the action types in the activated state includes action types tf2, tf8, and tf10, the mathematical combination bc2 of the action types in the activated state includes action types tf2, tf8, and tf17, the mathematical combination bc3 of the action types in the activated state includes action types tf2, tf10, and tf17, the mathematical combination bc4 of the action types in the activated state includes action types tf8, tf10, and tf17, the mathematical combination bc5 of the action types in the activated state includes action types tf2, and tf17 to tf18, The mathematical combination bc6 of behavior types in the activated state includes behavior types tf8 and tf17 to tf18, the mathematical combination bc7 of behavior types in the activated state includes behavior types tf10 and tf17 to tf18, the mathematical combination bc8 of behavior types in the activated state includes behavior types tf2, tf10, and tf18, the mathematical combination bc9 of behavior types in the activated state includes behavior types tf8, tf10, and tf18, and the mathematical combination bc10 of behavior types in the activated state includes behavior types tf2, tf8, and tf18.

[0049] In some embodiments, each of these mathematical combinations also has a validation set, and the validation set of the mathematical combination includes multiple validation values, which are also generated by performing detection on the subjects corresponding to all of the mathematical combinations (i.e., the corresponding benign subjects) based on the validation policy. It should be noted that the generation of these validation values ​​for the mathematical combinations is similar to the generation of these validation values ​​for the baseline set, and therefore will not be described here.

[0050] [Example of updating mathematical combinations in a validation database] FIG. 7 illustrates a flowchart for executing steps S710 to S720 after step S430 in some embodiments of the present disclosure. As shown in FIG. 7, in step S710, the processor 120 adds a mathematical combination that does not match the multiple baseline sets to the multiple baseline sets. Specifically, if the processor 120 cannot find a mathematical combination that is the same as one of the baseline sets in the verification database AD, the processor 120 sets the mathematical combination as a new baseline set. In this manner, the processor 120 stores the new baseline set and the multiple verification values ​​of the new baseline set in the verification database AD. By doing so, the processor 120 adds the new baseline set and its verification values ​​to the verification database AD. Conversely, if the processor 120 determines that the mathematical combination matches one of the stored baseline sets, the processor 120 updates the multiple verification values ​​of the matching baseline set in a subsequent stage.

[0051] In step S720, the processor 120 performs a logical operation (e.g., an "AND" operation) on the multiple verification values ​​of the mathematical combination that matches a baseline set and the multiple verification values ​​of the matching baseline set to generate a result of the operation and set the result as the multiple verification values ​​of the matching baseline set. Specifically, each time the processor 120 finds a baseline set that matches a mathematical combination in the verification database AD, the processor 120 uses the multiple verification values ​​of the baseline set to perform a logical operation on the multiple verification values ​​of the mathematical combination to generate a result and set the result as the multiple verification values ​​of the baseline set. That is, based on inductive generalization, the processor 120 uses the result of the logical operation to verify the "minimum and necessary" verification item list of these subsets of gray behaviors (i.e., subsets of the subject's multiple behavioral features of a predefined length N) using the result.

[0052] Through the above step, the anomaly detection device 100 further updates the multiple verification values ​​of the baseline set by comparing all mathematical combinations with the baseline set. In this way, the multiple baseline sets of multiple baseline behavioral features in the verification database AD perform further verification on the subject of the network in the zero trust environment (i.e., access verification is completed).

[0053] [Example of training a machine learning model] Returning to FIG. 4 , in step S440, after data collection training is completed in a supervised training environment, the processor 120 trains a machine learning model BM of multiple baseline behavioral features during a model training period, using the output of data augmentation as input to a learning function. In other words, the entire training period includes data collection training and model training. In some embodiments, after training (i.e., the data collection and data augmentation) is completed, the processor 120 trains a machine learning model BM of multiple baseline behavioral features during a model training period, using a learning function to generate all mathematical combinations generated by data augmentation as multiple training samples. The processor 120 then uses the trained machine learning model BM of multiple baseline behavioral features to identify whether captured data (i.e., the new data record in the subsequent stage) is abnormal. In other words, the trained machine learning model BM of multiple baseline behavioral features is used to identify whether undetermined gray behavior is gray behavior belonging to a benign subject. In some embodiments, the model training period is a phase in which collection of data records d1 to dM is stopped and training of machine learning model BM is performed.

[0054] [Example of using machine learning models] In step S450, in a normal working environment, the processor 120 uses the new output of the data augmentation performed by the continuous data collection module 110 based on all new data records captured from the test subject within a predefined time range (e.g., preset by a user) as input to a prediction function, and if the return based on the new output of the prediction function indicates an abnormal match, it indicates an anomalous event (i.e., one new data record is data belonging to abnormal gray behavior or there is an abnormality in the test subject). In some embodiments, the processor 120 utilizes the machine learning model BM of multiple baseline behavioral features to perform the test subject detection and provide an abnormality alert (e.g., an audio or informational alert).

[0055] In some embodiments, the anomaly detection apparatus 100 using data augmentation further comprises a user interface 140. The user interface 140 is used to configure a gray behavior policy and select a learning mode, a detection mode, or an offline mode as the operating mode of the processor 120 (i.e., to adjust the operation of the anomaly detection apparatus 100).

[0056] In some embodiments, the user interface 140 selects the offline mode as the operating mode and the processor 120 stops operating. In other words, once the operating mode is switched to the offline mode, the anomaly detection device 100 stops operating. In some embodiments, the user interface 140 selects the learning mode or the detection mode as the operating mode and the processor 120 activates the continuous data collection module 110.

[0057] In some embodiments, the operation of selecting the learning mode as the operation mode through the user interface 140 indicates that the processor 120 operates in a training environment where the processor 120 is supervised to perform a task normally and has begun training a machine learning model BM of a plurality of baseline behavioral features. In other words, when the learning mode is selected as the operation mode, the processor 120 operates in a training environment where the processor 120 is supervised to perform a task normally (i.e., the user operates the anomaly detection device 100 in a training environment where the processor 120 is supervised to perform a task normally) and has begun training the machine learning model BM.

[0058] In some embodiments, changing the operation mode from the learning mode to the detection mode or the offline mode through the user interface 140 indicates that one data collection training session has been completed, and the processor 120 performs data augmentation and training of the machine learning model BM using the data records collected during the training period. In other words, when the operation mode is switched from the learning mode to the detection mode or the offline mode, the processor 120 completes the data collection training session (i.e., the processor 120 collects data records through the continuous data collection module 110) and trains the machine learning model BM during the model training period (i.e., starts training the machine learning model BM using the data from the data augmentation).

[0059] In some embodiments, selecting the detection mode as the operating mode through the user interface 140 indicates that the processor 120 will operate in a normal working environment, and the processor 120 will use the trained machine learning model BM to detect anomalies. In other words, when the detection mode is selected as the operating mode, the processor 120 begins operating in a normal working environment (i.e., the user operates the anomaly detection device 100 in a normal working environment), and the trained machine learning model BM is used to detect anomalies for new data records in the normal working environment. In some embodiments, when the learning mode is selected again through the user interface 140 and the processor 120 performs new training, the processor 120 adds the new baseline behavioral feature to the multiple baseline behavioral features (i.e., accumulated in the machine learning model BM) in the new training.

[0060] The detection of an abnormal event will be described below with reference to an example. 8 is a block diagram showing an anomaly detection device 100 using data augmentation according to some other embodiments of the present disclosure. The anomaly detection device 100 in FIG. 8 is the same as the anomaly detection device 100 in FIG. 1, and therefore a description thereof will be omitted here.

[0061] As shown in FIG. 8, the continuous data collection module 110 can be connected to multiple test subjects TS1-TSm. In some embodiments, m can be any positive integer and is not particularly limited. In this example, in a normal working environment, the continuous data collection module 110 captures a new data record td from any one of the test subjects. In some embodiments, each test subject TS1-TSm can be an individual software resource of a system with or without an abnormality, a system used by a user, or an enterprise asset itself. It should be noted that the content base of the new data record td is similar to the content of the data records d1-dM, and its description will be omitted here.

[0062] The embodiment of Figure 8 differs from the embodiment of Figure 1 in that Figure 8 is applied to an unspecified environment (e.g., the normal working environment) in which benign and malicious subjects may exist. In other words, the test subjects TS1 to TSm are subjects whose anomalies are uncertain, that is, the test subjects TS1 to TSm may or may not have aggressive behavior. Therefore, it is necessary to identify whether the behavior of these test subjects TS1 to TSm is anomalous by capturing a new data record td.

[0063] In some embodiments, the processor 120 utilizes a prediction function to input the new data record td into a machine learning model BM of multiple baseline behavioral features to identify whether the new data record td is abnormal, and generates an abnormality alert if the new data record td is abnormal. In other words, once the match status returned by the prediction function is abnormal, the processor 120 knows that the new data record td belongs to abnormal gray behavior, and issues an abnormality alert to notify the user.

[0064] [Example of performing detection on a test subject] FIG. 9 shows a flowchart of steps S910 to S920, which are executed after step S450, according to some embodiments of the present disclosure. As shown in FIG. 9, in step S910, processor 120 selects a baseline set that matches new data record td. In some embodiments, processor 120 compares whether multiple behavioral types of the new data record are the same as multiple behavioral types of the baseline set. If multiple behavioral types of the new data record are the same as multiple behavioral types of the baseline set, processor 120 determines that the baseline set matches the new data record. If multiple behavioral types of the new data record are different from multiple behavioral types of the baseline set, processor 120 determines that the test subject that generated the new data record is a malignant subject.

[0065] In step S920, processor 120 verifies whether the test subject is anomalous using multiple validation values ​​from the baseline set that match the new data record. In some embodiments, processor 120 validates the test subject using validation policies that correspond to validation values ​​from the matching baseline set that are true (i.e., the validation values ​​have a logical value of "true"). If at least one validation value generated from the test subject is logically false, processor 120 determines that the test subject is anomalous. Conversely, if all validation values ​​generated by the test subject are logically true, processor 120 determines that the test subject is not anomalous.

[0066] For example, if a first logically true verification value and a second logically true verification value exist, the first logically true verification value and the second logically true verification value correspond to whether the subject has a trusted digital signature (i.e., a first verification policy) and whether the subject was constructed at a particular time (e.g., 10:00 AM) (i.e., a second verification policy), respectively. For example, the first logically true verification value indicates that the subject has a trusted digital signature, and the second logically true verification value indicates that the subject was constructed at a particular time.

[0067] Therefore, processor 120 determines whether the test subject generates two logically true verification values ​​based on the first verification policy and the second verification policy. In other words, processor 120 generates a first verification value by verifying whether the test subject has a trusted digital signature, and then generates a second verification value by verifying whether the test subject was constructed at a specific time. Next, if both the first verification value and the second verification value are logically true, processor 120 determines that the test subject conforms to the first logically true verification value and the first logically true verification value, and determines that the test subject is not abnormal. Conversely, if either the first verification value or the second verification value is not logically true, processor 120 determines that the test subject does not conform to the first logically true verification value and the first logically true verification value, and determines that the test subject is abnormal.

[0068] Through this step, the anomaly detection device 100 further verifies whether the test subject is an anomaly using the multiple verification values ​​of the baseline set (i.e., determines whether the test subject is a benign subject with no offensive behavior, or a malignant subject with the possibility of offensive behavior). Therefore, the multiple verification values ​​of the baseline set in the verification database AD are used to perform further verification on the test subject in the network of the zero trust environment. This significantly improves the accuracy of anomaly detection for subjects with gray behavior.

[0069] In summary, the present disclosure uses a combinatorial processing method to generate multiple mathematical combinations based on one or more active behavior types, thereby generating a very large number of mathematical combinations from a very small set of behavior types, thereby achieving the purpose of data augmentation. In this way, the present disclosure uses a very large number of mathematical combinations to detect anomalies in a zero trust environment. Furthermore, in the present disclosure, a machine learning model of multiple baseline behavioral features is used to perform further validation on data in a network in a zero trust environment, and the validation values ​​of the multiple baseline behavioral features are used to perform further validation on subjects in the network in a zero trust environment. As described above, the anomaly detection device and method using data augmentation according to the present disclosure significantly reduces human resources required for data collection and significantly improves the accuracy of detecting anomalies in gray behavior.

[0070] The above description is for the purpose of explaining the present invention, and should not be construed as limiting the invention described in the claims or narrowing its scope. Furthermore, the configuration of each part of the present invention is not limited to the above-mentioned embodiment, and various modifications are possible within the technical scope described in the claims. [Explanation of symbols]

[0071] 100 Anomaly detection device 110 Continuous Data Collection Module 120 processors 130 memory 140 User Interface SD sample database BM Machine learning model of multiple baseline behavioral features AD Verification Database d1~dM data recording rc baseline set rf1~rf20 Action type tf1~tf20 Action type VC Validation Set bf1~bf20 behavior information bc1~bc10 Mathematical Combinations cs set TS1~TSm Test Subjects td New data record S410~S450 Step S710~S720 Step S910~S920 Step

Claims

1. a continuous data collection module configured to capture a plurality of data records of a plurality of subjects in a training environment or a normal work environment under supervision to perform a normal task, each data record including a time, a subject, a plurality of behavioral types, and a plurality of behavioral information corresponding to each of the behavioral types; configured to store machine learning models of the plurality of instructions and the plurality of baseline behavioral features; a) a learning function step that takes as input a set of said behavior types of a predefined length N and adds said set to a plurality of said baseline behavior features; b) a prediction function step that takes as input the new set of behavior types of the defined length N and outputs a match state; a memory connected to the continuous data collection module and the memory, for each data record: performing activation state determination for the plurality of behavioral information using a plurality of activation functions associated with the plurality of behavioral types, and combining all of the behavioral types in activation states for each of the subjects to generate the set of behavioral types specific to each of the subjects, each of the sets representing a plurality of behavioral features of the corresponding subject; performing data augmentation by enumerating, for each subject, all mathematical combinations of the predefined lengths N in the set of behavior types based on the predefined lengths N used in the machine learning model of the plurality of baseline behavioral features, each mathematical combination being a subset of the predefined lengths N of the plurality of behavioral features of the subject; In the training environment, after data collection training is completed, training the machine learning model of the plurality of baseline behavioral features during a model training period using the output of the data augmentation as an input to the learning function; and a processor configured to execute a plurality of the instructions to perform the steps of: in the normal working environment, using as input to the prediction function a new output of data augmentation performed based on all new data records captured from a test subject by the continuous data collection module within a defined time range; and indicating as an anomalous event if a return of the prediction function indicates an abnormal match state.

2. 2. The anomaly detection device using data augmentation of claim 1, wherein the continuous data collection module is configured to capture a plurality of the data records of a plurality of benign subjects in the training environment where the subjects are supervised to perform their work normally, and the plurality of benign subjects are individual software resources of a system that is not abnormal, a system used by a user, or an asset of an enterprise.

3. 2. The anomaly detection device using data augmentation described in claim 1, characterized in that each of the behavioral information is displayed as numerical data, each of the behavior types is displayed as a gray behavior policy, the activation state of each of the behavior types is displayed as a Boolean value, and the gray behavior policy is used to generate multiple pieces of behavioral information.

4. a user interface configured to set a grey behavior policy and to select a learning mode, a detection mode, or an offline mode as an operating mode of the processor, wherein the processor stops operating when the offline mode is selected as the operating mode via the user interface; selecting the learning mode or the detection mode as the operating mode through the user interface causes the processor to activate the continuous data collection module; selecting the learning mode as the operating mode in the user interface indicates that the processor is operating in the supervised training environment to successfully perform a task and has begun training the machine learning model of the plurality of baseline behavioral features; changing the operation mode from the learning mode to the detection mode or the offline mode through the user interface indicates that one session of the data collection training has been completed, and the processor performs data augmentation and training of the machine learning model using the data records collected during the data collection training session; an operation of selecting the detection mode as the operation mode through the user interface indicates that the processor is operating in the normal working environment, and the processor detects anomalies using the trained machine learning model; 2. The anomaly detection device using data augmentation according to claim 1, wherein when the user interface reselects the learning mode as the operating mode and the processor performs new training, the processor adds new baseline behavioral features to the plurality of baseline behavioral features during the new training.

5. the new output of the data augmentation performed based on the new data record of the test subject as the input to the prediction function, and indicating as the anomalous event if the return of the prediction function indicates an abnormal match condition, the processor:

2. The apparatus for anomaly detection using data augmentation of claim 1 , configured to perform the steps of utilizing the machine learning model of a plurality of the baseline behavioral features to detect the test subjects and provide an anomaly alert.

6. The plurality of baseline behavioral features comprises a plurality of baseline sets, and the processor: adding the mathematical combinations that do not fit into a plurality of said baseline sets to a plurality of said baseline sets; 2. The anomaly detection apparatus using data augmentation of claim 1, further configured to perform a logical operation on the plurality of verification values ​​of the mathematical combination that matches one of the baseline sets and the plurality of verification values ​​of the matching baseline sets to generate an operation result, and setting the operation result as the plurality of verification values ​​of the matching baseline set.

7. The plurality of baseline behavioral features comprises a plurality of baseline sets, and the processor: selecting one of the baseline sets that matches the new data record; 2. The apparatus for anomaly detection using data augmentation of claim 1, further configured to perform the step of verifying whether the test subject is anomalous using a plurality of verification values ​​of the baseline set that match the new data record.

8. capturing, by a continuous data collection module, a plurality of data records of a plurality of subjects in a training environment where the subject is supervised to perform a normal task or in a normal work environment, each of the data records including a time, a subject, a plurality of behavioral types, and a plurality of behavioral information corresponding to each of the plurality of behavioral types; and performing, by a processor, for each of the data records, a determination of activation state for the plurality of the behavioral information using a plurality of activation functions associated with a plurality of the behavioral types, and combining all the behavioral types in activation state for each of the subjects to generate a set of the behavioral types specific to each of the subjects, each of the sets representing a plurality of behavioral characteristics of the corresponding subject; performing data augmentation by the processor for each data record by enumerating, for each subject, all mathematical combinations of predefined lengths N in the set of behavior types based on predefined lengths N used in a machine learning model of a plurality of baseline behavioral features, each mathematical combination being a subset of the predefined lengths N of a plurality of behavioral features of the subject; In the training environment, after completing data collection training, the processor trains the machine learning model on the plurality of baseline behavioral features during a model training period using the output of the data augmentation as an input to a learning function, wherein the learning function takes the set of behavior types of the defined length N as an input and adds the set to the plurality of baseline behavioral features; and in the normal working environment, the processor uses a new output of the data augmentation performed by the continuous data collection module based on all new data records captured from the test subject within a defined time range as input to a prediction function, and if the return of the prediction function indicates an abnormal match state, indicate an anomalous event, wherein the prediction function takes as input a new set of the behavior types of the defined length N and outputs the match state.

9. 9. The method of claim 8, wherein the continuous data collection module is configured to capture a plurality of the data records of a plurality of benign subjects in the training environment under supervision to perform normal work, the plurality of benign subjects being individual software resources of a system that is not abnormal, a system used by a user, or an asset of an enterprise.

10. 9. The anomaly detection method using data augmentation described in claim 8, wherein each of the behavioral information is displayed as numerical data, each of the behavioral types is displayed as a gray behavior policy, and the activation state of each of the behavior types is displayed as a Boolean value, and the gray behavior policy is used to generate multiple of the behavioral information.

11. setting a gray behavioral policy through a user interface and selecting a learning mode, a detection mode, or an offline mode as an operation mode of the processor for the plurality of baseline behavioral features; When the offline mode is selected as the operation mode through the user interface, the processor stops operation; selecting the learning mode or the detection mode as the operating mode via the user interface, the processor activating the continuous data collection module; selecting, through the user interface, the learning mode as the operating mode, indicating that the processor is operating in a supervised training environment to successfully perform a task and has begun training the machine learning model of the plurality of baseline behavioral features; changing the operational mode from the learning mode to the detection mode or the offline mode through the user interface indicates that one session of the data collection training has been completed, and the processor performs data augmentation and training of the machine learning model using the data records collected during the data collection training session; selecting the detection mode as the operation mode through the user interface indicates that the processor is operating in the normal working environment, and the processor performs anomaly detection using the trained machine learning model; 10. The method of claim 8, further comprising: when the user interface reselects the learning mode as the operating mode and the processor performs new training, the processor adds new baseline behavioral features to the plurality of baseline behavioral features during the new training.

12. the step of using the new output of the data augmentation performed based on the new data record of the test subject as the input to the prediction function, and indicating as the anomalous event if the return of the prediction function indicates an abnormal match condition, comprising:

10. The method of claim 8, further comprising utilizing the machine learning model of the plurality of baseline behavioral features to detect the test subjects and provide an anomaly alert.

13. The plurality of baseline behavioral features comprises a plurality of baseline sets, and the anomaly detection method comprises: the processor adding the mathematical combination that does not match the plurality of baseline sets to the plurality of baseline sets; 9. The method for anomaly detection using data augmentation of claim 8, further comprising: the processor performing a logical operation on the plurality of verification values ​​of the mathematical combination that matches one of the baseline sets and the plurality of verification values ​​of the matching baseline sets to generate an operation result, and setting the operation result as the plurality of verification values ​​of the matching baseline sets.

14. The plurality of baseline behavioral features comprises a plurality of baseline sets, and the anomaly detection method comprises: the processor selecting one of the baseline sets that matches the new data record; 10. The method for anomaly detection using data augmentation of claim 8, further comprising the step of: the processor verifying whether the test subject is anomalous using a plurality of validation values ​​of the baseline set that match the new data record.

Citation Information

Patent Citations

  • Model training method, and method and device for optimizing training data set

    CN113204614A

  • Classification device

    JP2019057016A