Model training method, apparatus, device, and medium

By acquiring the identifiers and features of the target service, constructing a historical dataset, training samples, and labels, the problem of low efficiency and strong subjectivity in the identification of abnormal network services in existing technologies is solved, and efficient and accurate anomaly identification is achieved.

CN117194654BActive Publication Date: 2025-11-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210584158.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-26
Publication Date
2025-11-25
Estimated Expiration
2042-05-26

AI Technical Summary

Technical Problem

In existing technologies, relying on risk control experts to identify abnormal network services is inefficient and yields subjective results, and there is a lack of effective machine learning models for identification.

Method used

By acquiring the identifiers and characteristics of the target business, a historical dataset is constructed, training samples and labels are determined, an anomaly recognition model is trained, and a machine learning model is used to identify abnormal business.

Benefits of technology

It improves the efficiency of anomaly identification, reduces subjectivity, and provides an objective and accurate solution for identifying abnormal network services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117194654B_ABST
    Figure CN117194654B_ABST
Patent Text Reader

Abstract

The application discloses a model training method and device, equipment and medium, relates to the technical field of big data, and in particular to the field of network risk control. The method comprises the following steps: obtaining the identification of a target service, determining the characteristics matched with the target service according to the identification of the target service, and determining the training characteristics of an abnormality identification model (for identifying whether the target service is abnormal according to the traffic data of the target service) according to the characteristics matched with the target service; obtaining a historical data set of the target service, determining the training sample of the abnormality identification model according to at least two historical traffic data in the historical data set, determining the label of the training sample according to the evaluation result of the at least two historical traffic data, and training the abnormality identification model based on the training sample, the label of the training sample and the training characteristics. The abnormality identification model can be trained, the problem existing in model training of the identification model of abnormal network service is overcome, and support is provided for the application of the machine learning model in the identification of abnormal network service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of big data technology, specifically to the field of network risk control technology, and in particular to a model training method, apparatus, device, and medium. Background Technology

[0002] With the rapid development of internet big data technology, various internet-based services have experienced explosive growth. However, some abnormal network services have emerged, employing illegal or irregular means, posing a significant threat to network security.

[0003] Currently, the identification of abnormal network transactions mainly relies on risk control experts for manual verification. Experts identify key information from network traffic data, such as account information and device information. They can also check the credit scores of accounts and devices, and finally use manual judgment rules to determine whether the network transaction is abnormal. However, relying on risk control experts for manual identification of abnormal network transactions has the problems of low efficiency and subjective results.

[0004] To address these issues, researchers hope to leverage machine learning models to identify abnormal network traffic, thereby improving identification efficiency and ensuring the objectivity and impartiality of the results. However, currently, there are no machine learning models suitable for identifying abnormal network traffic, and training such models presents significant challenges. Summary of the Invention

[0005] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide a model training method, apparatus, device and medium that can train a model for identifying abnormal network services, thereby providing support for the application of machine learning models in the identification of abnormal network services.

[0006] Firstly, a model training method is provided, including:

[0007] Obtain the identifier of the target service, determine the features that match the target service based on the identifier, and determine the training features of the anomaly detection model based on the features that match the target service. The features are used to identify whether the target service is an abnormal service, and the anomaly detection model is used to identify whether the target service is an abnormal service based on the traffic data of the target service.

[0008] The historical dataset of the target business is obtained according to the data selection strategy. The historical dataset includes at least two historical traffic data of the target business and the evaluation result corresponding to each historical traffic data. The evaluation result is used to characterize whether the target business is an abnormal business.

[0009] The training sample of the anomaly recognition model is determined according to at least two historical traffic data, and the label of the training sample is determined according to the evaluation result of the at least two historical traffic data;

[0010] The anomaly recognition model is trained based on the training sample, the label of the training sample, and the training feature.

[0011] In the present application, the historical data set of the target service is obtained according to the data selection strategy, and the historical data set includes at least two historical traffic data with evaluation results (used to represent whether the target service is an abnormal service). The training sample of the anomaly recognition model can also be determined based on the at least two historical traffic data. Since the historical traffic data has a clear evaluation result, the label of the training sample can be determined according to the evaluation result of the historical traffic data, so that the training sample with the label can be constructed. The problem of blindly obtaining training data without obtaining positive feedback of the data and being unable to create training samples with labels can be avoided. In addition, the feature used to identify whether the target service is an abnormal service (i.e., the feature matching the target service) is determined according to the identification of the target service, and the training feature of the model is determined based on the above-mentioned feature, so that the model can extract some features for identifying whether the training sample is abnormal. The model can also judge whether the training sample is abnormal based on the extracted features, so as to gradually learn the ability to identify whether the target service is an abnormal service based on the traffic data of the target service. Based on the above-mentioned training feature, training sample, and label of the training sample, the anomaly recognition model of the target service can be trained, which can identify whether the target service is an abnormal service according to the traffic data of the target service. The problem of the anomaly network service recognition model in model training is overcome, and support is provided for the application of machine learning model in anomaly network service recognition.

[0012] In a second aspect, an anomaly recognition method is provided, comprising:

[0013] Traffic data of a target service is obtained, the traffic data is input into an anomaly recognition model of the target service, and whether the target service is an abnormal service is identified according to the output of the anomaly recognition model;

[0014] The training feature of the anomaly recognition model is determined according to the feature matching the target service, and the feature is used to identify whether the target service is an abnormal service. The training sample of the anomaly recognition model is determined according to at least two historical traffic data of the target service, and the label of the training sample is determined according to the evaluation result of the at least two historical traffic data. The evaluation result is used to represent whether the target service is an abnormal service.

[0015] In the present application, the traffic data can be identified by means of the anomaly identification model. Compared with the prior art which relies on risk control experts to manually identify abnormal network services, the efficiency of anomaly identification is greatly improved, and the subjectivity of anomaly identification can be avoided, thereby providing an objective and accurate anomaly network service identification scheme.

[0016] In a third aspect, a model training apparatus is provided, comprising:

[0017] The feature determination unit is configured to obtain an identifier of the target service, determine a feature matched with the target service according to the identifier of the target service, and determine a training feature of the anomaly identification model according to the feature matched with the target service; the feature is used to identify whether the target service is an abnormal service, and the anomaly identification model is used to identify whether the target service is an abnormal service according to traffic data of the target service.

[0018] The obtaining unit is configured to obtain a historical data set of the target service according to a data selection strategy; the historical data set comprises at least two historical traffic data of the target service and an evaluation result corresponding to each historical traffic data, and the evaluation result is used to represent whether the target service is an abnormal service.

[0019] The sample determination unit is configured to determine a training sample of the anomaly identification model according to the at least two historical traffic data, and determine a label of the training sample according to the evaluation result of the at least two historical traffic data.

[0020] The training unit is configured to train the anomaly identification model based on the training sample, the label of the training sample, and the training feature.

[0021] In a fourth aspect, an anomaly identification apparatus is provided, comprising:

[0022] The obtaining unit is configured to obtain traffic data of a target service.

[0023] The identification unit is configured to input the traffic data into an anomaly identification model of the target service, and identify whether the target service is an abnormal service according to an output of the anomaly identification model.

[0024] The training feature of the anomaly identification model is determined according to a feature matched with the target service, and the feature is used to identify whether the target service is an abnormal service; the training sample of the anomaly identification model is determined according to at least two historical traffic data of the target service, and the label of the training sample is determined according to an evaluation result of the at least two historical traffic data, and the evaluation result is used to represent whether the target service is an abnormal service.

[0025] In a fifth aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the program, the method of the first aspect or the second aspect is implemented.

[0026] In a sixth aspect, a computer readable storage medium is provided, having stored thereon a computer program, wherein the program, when executed by a processor, implements the method of the first aspect or the second aspect.

[0027] In a seventh aspect, a computer program product is provided, comprising instructions which, when executed by a processor, implement the method of the first aspect or the second aspect.

[0028] Additional aspects and advantages of the application will be made apparent by the following description. BRIEF DESCRIPTION OF DRAWINGS

[0029] Other features, objects, and advantages of the application will become apparent from the following detailed description of non-limiting embodiments:

[0030] Figure 1 A model training schematic diagram provided for the embodiments of the application;

[0031] Figure 2 An implementation environment schematic diagram provided for the embodiments of the application;

[0032] Figure 3 A flowchart schematic diagram of a model training method provided for the embodiments of the application;

[0033] Figure 4a A structure schematic diagram of an initial network model provided for the embodiments of the application;

[0034] Figure 4b Another structure schematic diagram of an initial network model provided for the embodiments of the application;

[0035] Figure 5 A training feature configuration interface provided for the embodiments of the application;

[0036] Figure 6 A data source configuration interface provided for the embodiments of the application;

[0037] Figure 7 A model update board provided for the embodiments of the application;

[0038] Figure 8 An artificial judgment rule configuration interface provided for the embodiments of the application;

[0039] Figure 9 A rule microservice calling schematic diagram provided for the embodiments of the application;

[0040] Figure 10 A model evaluation schematic diagram provided for the embodiments of the application;

[0041] Figure 11 A negative sample acquisition schematic diagram provided for an embodiment of the present application is provided.

[0042] Figure 12 A schematic diagram of an anomaly recognition system provided for an embodiment of the present application is provided.

[0043] Figure 13 A schematic diagram of a model evaluation system provided for an embodiment of the present application is provided.

[0044] Figure 14 A flowchart of an anomaly recognition method provided for an embodiment of the present application is provided.

[0045] Figure 15 A structural schematic diagram of a model training device provided for an embodiment of the present application is provided.

[0046] Figure 16 A structural schematic diagram of an anomaly recognition device provided for an embodiment of the present application is provided.

[0047] Figure 17 A structural schematic diagram of a computer device provided for an embodiment of the present application is provided. DETAILED DESCRIPTION

[0048] The present application will be further described below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related application, and not to limit the application. In addition, it should be noted that only the parts related to the application are shown in the drawings for ease of description.

[0049] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and embodiments.

[0050] First, the terms related to the present application are explained.

[0051] (1) Model training: input a large amount of data with known classification results into an initial model, let the algorithm inside the model learn the classification rules of these data, so that the trained model can classify unknown data.

[0052] The data of a large number of known classification results input into the model can be referred to as a training sample of the model. Taking supervised machine learning as an example, the true classification result of the training sample can be used as a label of the training sample. In the model training process, feature extraction is performed on the training sample, the extracted features are input into the classification function of the model, and the output result of the model can be obtained. Further, the loss between the output result of the model and the label of the training sample can be determined according to the loss function, and the model is iteratively trained according to the loss until the output result of the model approaches the label of the training sample. That is, the model has the ability to accurately classify data. In addition, the features obtained by the model through feature extraction on the model input can be referred to as training features of the model, which are usually key features for identifying and classifying the model input. For example, when identifying a face in a picture, the key features to be referred to can be "face contour", "eyes", "mouth", etc. The model can extract face contour features, eye features, mouth features, etc. from the input picture, and the classification function of the model can output the identification result of the model based on these features.

[0053] Figure 1 Taking the training of a picture anomaly recognition model as an example, the process of model training is introduced. Assuming that the training sample is a picture containing the "32E8" character, the true label of the training sample is "32E8". The picture is input into the neural network model, and the model can output the identification result of the picture. For example, the identification result output by the model is "32E0". Further, the loss value between "32E0" and the label "32E8" of the training sample can be determined according to the loss function, and the parameters of the neural network are adjusted according to the loss value to iteratively train the model.

[0054] (2) Internet abnormal business: It can be a network business that uses the Internet as a medium and adopts abnormal technical means. For example, malicious registration, pornography, malicious single brushing, fraud, gambling, sheep shearing, password stealing, external hanging and other network businesses.

[0055] (3) Flow data: It can be data generated by a certain network business. For example, the flow data of the user registration business can include the information of the terminal initiating the registration request, the user name, the user password, the verification code and other data related to the user registration business.

[0056] Currently, in the field of Internet risk control, a risk control expert mainly refers to an artificial judgment rule to identify the traffic data of network services to determine whether the network service is an abnormal service. On the one hand, the amount of network traffic data is huge, and relying on a risk control expert to identify Internet abnormal services leads to low efficiency of abnormal identification. In the face of rapidly growing traffic data, the efficiency of manual identification has very limited room for improvement. On the other hand, the artificial judgment rule referred to by the expert for abnormal identification is often determined according to historical experience, which has a certain subjectivity, and also leads to the result of abnormal identification being greatly affected by subjective factors, and the accuracy is also limited.

[0057] The above two aspects become problems that need to be solved, and the idea of identifying abnormal services with the help of a machine learning model emerges as the times require. However, due to the particularity of abnormal services in the risk control scene, the traffic data of many network services has no positive feedback, that is, it cannot be determined whether the traffic data is generated by an abnormal service. That is, it is difficult to construct training samples with labels, which brings great difficulty to the training of the abnormal identification model of abnormal network services.

[0058] Based on this, the present application proposes a model training method, device and storage medium, which can determine the training samples and training features suitable for this type of network service according to the characteristics of the network service, provide guidance for the model training of the abnormal identification model (i.e. the model for identifying whether the network service is an abnormal service), greatly reduce the training difficulty of the abnormal identification model, and provide support for the application of machine learning models in abnormal network service identification.

[0059] Figure 2 An implementation environment for the embodiments of the present application is shown in the schematic diagram. Reference is made to Figure 2 In the field of Internet risk control, the risk control platform 10 can provide risk control services to the risk control demander 20. The risk control demander 20 can support the implementation of multiple network services, such as "user registration", "add friends", "create group chat", etc. The risk control demander 20 can obtain the traffic data of the network service and send the traffic data to the risk control platform 10. The risk control platform 10 can perform risk control based on the traffic data to identify whether the network service is abnormal. Further, the risk control platform 10 can also send the identification result to the risk control demander 20, such as "malicious registration", "external plug-in", etc.

[0060] The risk control platform 10 includes a background server 101 and a configuration interface 102. The configuration interface 102 can be an application programming interface (API) or a graphical user interface (GUI). The management personnel of the risk control platform can specifically configure the risk control business of the risk control platform 10 through the configuration interface 102. The background server 101 can support the background implementation of the risk control business of the risk control platform 10.

[0061] The risk control demander 20 includes a background server 201 and a client 202. The client 202 can use various network services provided by the risk control demander 20, and the background server 201 supports the background implementation of various network services.

[0062] The background servers 101 and 201 can be independent physical servers, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0063] The client 202 can be a terminal or an application installed on a terminal. The terminal can be a device including but not limited to a personal computer, a platform computer, a smart phone, a vehicle-mounted terminal, and the like, and the embodiments of the present application do not limit the terminal.

[0064] The embodiments of the present application provide a model training method. The execution subject of the method can be the background server 101 described above. The method can provide guidance for model training of an anomaly recognition model, greatly reducing the training difficulty of the anomaly recognition model. Referring to the method, the method includes the following steps: Figure 3

[0065] 301, obtaining an identifier of a target business, and determining a feature matched with the target business according to the identifier of the target business, the feature being used to identify whether the target business is an abnormal business.

[0066] As known from the foregoing, one element of model training is the training feature of the model. For example, the embodiments of the present application can determine a plurality of key features for identifying abnormal network services, so as to determine the training feature of the model based on the key features, so that the model can learn the ability to identify abnormal target businesses.

[0067] ​In a possible implementation, the target service can be any network service, for example, any network service provided by the risk control demander 20. The feature matched with the target service can be a key feature (or key indicator) for identifying whether the target service is an abnormal service. For example, in the network service "user registration", the key feature for identifying an abnormal service can be "number of login failures", "number of verification code errors", "number of password errors", and the like, and whether the "user registration" service is abnormal can be identified according to the "number of login failures", "number of verification code errors", and "number of password errors". For example, when the "number of login failures", "number of verification code errors", and "number of password errors" exceed a preset threshold, the "user registration" service can be identified as abnormal.

[0068] In a possible implementation, the network service and the key feature for identifying an abnormal service in the service scenario have a corresponding relationship. For example, the service identifier of the network service and the key feature of the network service have a corresponding relationship, and the risk control platform can store the service identifier and the key feature correspondingly. When the key feature of the target service is determined, the identifier of the target service can be acquired first, so that the key feature (i.e., the feature matched with the target service) of the target service can be determined according to the identifier of the target service and the corresponding relationship. Different features are matched for different services, different abnormality identification models can be trained for different network services, and this is also conducive to realizing fine prediction of abnormal services.

[0069] For example, the feature library can store the identifier and the key feature of the network service correspondingly. In step 301, the feature matched with the target service can be determined according to the identifier of the target service and the feature library.

[0070] Table 1 below is an implementation of the feature library:

[0071] Table 1

[0072] Business identification Key features 11 "Username" "Number of failed verifications" "Device identification" 12 "Transaction amount" "Parameter group chat number" "Discount amount" … …

[0073] 302、According to the feature matched with the target service, the training feature of the abnormality identification model is determined.

[0074] The abnormality identification model is a machine learning model capable of identifying whether the target service is an abnormal service. Specifically, the input of the abnormality identification model is the traffic data of the target service, and the output of the abnormality identification model can indicate whether the traffic data is abnormal, thereby identifying whether the target service is an abnormal service. It can be understood that, when identifying the abnormality of the target service, the traffic data of the target service can be obtained, the traffic data is input into the abnormality identification model, and the abnormality identification model can identify whether the target service is an abnormal service according to the input traffic data.

[0075] In a possible implementation, in order to enable the model to learn the ability of identifying the abnormality according to the traffic data of the target service, the model can learn some key features for identifying whether the target service is an abnormal service, and therefore the training features of the abnormality identification model can be determined according to the key features of the target service. For example, the features matched with the target service are taken as the training features of the abnormality identification model. Alternatively, part of the features matched with the target service can be selected as the training features of the abnormality identification model. In a specific implementation, the training features can be selected from the features matched with the target service in an artificial manner, or the training features of the abnormality identification model can be selected from the features matched with the target service according to the characteristics of the target service by the background server of the risk control platform, but it is necessary to ensure that the selected features can achieve the abnormality identification of the target service, and the necessary features for the abnormality identification of the target service cannot be missing.

[0076] 303、According to the data selection strategy, the historical data set of the target service is obtained; the historical data set includes at least two historical traffic data of the target service and an evaluation result corresponding to each historical traffic data, and the evaluation result is used to represent whether the target service is an abnormal service.

[0077] As can be known from the foregoing introduction to the model training process, in addition to the training features, another element of the model training is the training sample. In the embodiment of the present application, in order to overcome the problem that the training sample has no explicit label, the historical data set related to the target service can be obtained. The historical traffic data in the historical data set has an explicit evaluation result, that is, the historical traffic data is data with positive feedback, and can be used to construct a labeled training sample.

[0078] It should be noted that the evaluation result can explicitly represent that the historical traffic data is abnormal data, that is, the target service is an abnormal service; or the evaluation result can explicitly represent that the historical traffic data is normal data, that is, the target service is a normal service.

[0079] For example, the evaluation result can be the result of manually identifying the historical traffic data based on an existing artificial judgment rule. For example, the result of the evaluation of the historical traffic data by an expert of the risk control platform.

[0080] In a possible implementation, the historical data set of the target service can be obtained according to a data selection strategy. For example, the data selection strategy can indicate a plurality of data sources from which historical traffic data of the target service is obtained, and a proportion of data volume of each data source. The historical traffic data of the proportion can be obtained from the data sources indicated by the data selection strategy. For example, the data sources can be a virtual abnormal object, an online detection data stream, and the like.

[0081] It should be noted that the data selection strategy can be generated in advance by a background server of the risk control platform according to characteristics of the target service. Alternatively, one or more data sources can be manually set for the target service, and the data selection strategy can be generated according to the one or more manually set data sources.

[0082] 304. Determine training samples of the abnormality recognition model according to at least two historical traffic data of the target service, and determine labels of the training samples according to evaluation results of the at least two historical traffic data.

[0083] It should be noted that, in order to overcome the problem that the training data has no positive feedback and cannot construct labeled training samples, training samples can be constructed according to historical traffic data in the historical data set, and evaluation results of the historical traffic data are determined as labels of the training samples.

[0084] Specifically, for each historical traffic data in the historical data set, the historical traffic data is determined as a training sample, the evaluation result of the historical traffic data is determined as a label of the training sample, and a training sample set is generated according to all training samples and labels of all training samples.

[0085] For example, assuming that the target service is “user registration”, the obtained historical traffic data of the user registration service includes: data 1 “username 123, number of incorrect verification codes x1, number of registration failures y1”, data 2 “username 456, number of incorrect verification codes x2, number of registration failures y2”, and data 3 “username 789, number of incorrect verification codes x3, number of registration failures y3”. The evaluation result of data 1 is “abnormal”, the evaluation result of data 2 is “normal”, and the evaluation result of data 3 is “abnormal”. The constructed training sample set includes data 1, data 2, and data 3, wherein the label of data 1 is “abnormal”, the label of data 2 is “normal”, and the label of data 3 is “abnormal”.

[0086] 305. Train the abnormality recognition model of the target service based on the training samples, the labels of the training samples, and the training features.

[0087] In a specific implementation, the initial network model can be determined according to the training features. Then, the training sample set (including the training samples and the labels of the training samples) can be input into the initial network model for feature extraction to obtain the training features of each training sample in the training sample set, and the prediction result of the initial network model for the training sample can be obtained based on the training features corresponding to the training sample. Further, the initial network model can be trained to obtain the abnormality recognition model according to the loss between the prediction result corresponding to each training sample and the label of the training sample.

[0088] In a possible implementation, the initial network model can be a basic model for training a classification model, such as a binary classification model or a multi-classification model. For example, the initial network model is a model for binary classification, that is, the model can predict two types, and the output result of the model is one of two categories. For example, the output of the model is “abnormal service” or “normal service”. Alternatively, the initial network model can also be a multi-classification model, that is, the model can predict multiple types, and the output result of the model is one or more of multiple categories, and the one or more types output by the model can be an abnormal service type predicted by the model, such as “sheep shearing”, “malicious registration”, “malicious promotion”, and the like.

[0089] In a possible implementation, the network structure of the initial network model is related to the training features, and the initial network model can be determined according to the training features. The network structure inside the initial network model has the ability to extract features of the training samples to obtain the training features of the training samples. For example, the initial network model can include a feature extraction network and a classification network. Figure 4a is a structural diagram of the initial network model, Figure 4a The initial network model can include a feature extraction network and a classification network. After the training sample is input into the initial network model, the feature extraction network first extracts features of the training sample to obtain the training features of the training sample. The training features can also be input into the classification network, and the classification function of the classification network is used to operate the training features to finally obtain the prediction result output by the model.

[0090] It should be noted that the algorithm (i.e., network structure) of the feature extraction network can be determined according to the training features described above, so that the feature extraction network has the ability to extract the training features from the training samples (i.e., traffic data of the business). In an embodiment of the present application, the feature extraction network can be at least one of a convolutional network, an embedding layer, and a text recognition network. Among them, when the training sample contains an image, the convolutional network can perform convolution processing on the training sample input to the model to extract the image features of the training sample; the embedding layer can perform vectorization processing on the input of the model, for example, converting words in the traffic data into vectors; the text recognition network can perform semantic recognition on the text input to the model, for example, performing word segmentation processing on the text, and recognizing the word segmentation result to extract the features of the input text.

[0091] Suppose the target business is a user registration business, and the key features for identifying abnormal network businesses in the user registration scenario can be "username", "registration failure times", and "verification code error times". The model training features can also be "username", "registration failure times", and "verification code error times". After the historical traffic data as training samples are input to the initial network model, the feature extraction network can extract the features "username w", "registration failure times t", and "verification code error times k" of the training samples, and then input "username w", "registration failure times t", and "verification code error times k" to the classification network. The classification network outputs a prediction result according to "username w", "registration failure times t", and "verification code error times k", indicating whether the training sample is traffic data of an abnormal business.

[0092] For example, the following Figure 4b is a possible structure of the initial network model. Referring to Figure 4b , the initial network model includes a text recognition network, an image recognition network, and a classification network. That is, the feature extraction network of the initial network model includes the text recognition network and the image recognition network. When the training sample is input to the initial network model, the data in the training sample can be first shunted, and the text therein is input to the text recognition network. The text recognition network can perform semantic recognition on the text, and can also perform word segmentation processing on the semantic recognition result. Further, the obtained word segmentation is recognized to extract candidate word segmentation matching the semantic of the training features.

[0093] In addition, the image in the training sample can be input into the image recognition network through the above-mentioned shunt processing. The image recognition network can use a convolution network to extract features of the input image and obtain a feature map. Specifically, the convolution network can be determined according to the training features, and the convolution network has the ability to extract the above-mentioned training features from the image. Different convolution networks correspond to different convolution channels, for example, a channel for extracting eye features, a channel for extracting mouth features, and the like. In the embodiments of the present application, in order to improve the efficient extraction of local detail features by the anomaly recognition model, different weight coefficients can be assigned to each channel according to the different attention levels. In addition, for different regions in the same feature map, different weight coefficients can also be set for each region according to the different attention levels, to achieve more detailed feature extraction. That is, when using the convolution network to extract the feature map, the feature map can be adjusted according to the weight coefficient of the corresponding convolution channel, and the feature map can be further updated according to the weight coefficient of each region in the feature map. Finally, the feature map output by each convolution channel is taken as the candidate feature map extracted by the image recognition network. Wherein, updating the feature map according to the weight coefficient can be multiplying the weight coefficient and the value of each pixel in the feature map, and taking the result as the value of the corresponding pixel position.

[0094] Finally, the above-mentioned candidate word segmentation and candidate feature map can be input into a classification network. The classification network can respectively perform vectorization processing on the candidate word segmentation and the candidate feature map, and then fuse the vectors obtained by the vectorization processing to obtain a fusion vector. Then, the fusion vector is input into a classification function of the classification network to perform operation, and the prediction result of the initial network model for the training sample is obtained. Wherein, the classification function can be a sigmoid function for realizing binary classification prediction, or can be a softmax function for realizing multi-classification.

[0095] It should be noted that if the training sample contains an image and does not contain text, the candidate feature map can be obtained by using the above-mentioned image recognition network, the vectorization processing is performed on the candidate feature map, the obtained vector is input into the classification function to perform operation, and the prediction result is obtained. Or, the training sample contains text and does not contain image, then the candidate word segmentation can be obtained by using the above-mentioned text recognition network, the vectorization processing is performed on the candidate word segmentation, the obtained vector is input into the classification function to perform operation, and the prediction result is obtained.

[0096] The following describes the prediction result of the initial network model by taking a binary classification model as an example. Specifically, if the initial network model is a binary classification model, and the types that can be predicted by the initial network model are represented by 0 and 1, where type “1” represents “abnormal service” and type “0” represents “normal service”. The output of the initial network model can be a 2*1 vector (i.e., the output of the classification function), where each element of the vector corresponds to a type. For example, in the order of row arrangement, the elements correspond to type 1 and type 0 in turn. The value (score) of an element represents the probability that the input sample is of the corresponding type, and the sum of the scores is equal to 1. The higher the score, the higher the probability that the input sample hits the type corresponding to the score, so the type corresponding to the highest score can be selected as the type predicted by the model. Assuming that the output of the initial network model is [0.33, 0.67], it means that the probability of the input sample being of type 1 is 0.33 and the probability of the input sample being of type 0 is 0.67. At this time, the type with the highest score, i.e., type 0, can be taken as the prediction result of the initial network model, i.e., the model identifies the input sample as “normal service”.

[0097] It should be noted that there is usually a certain difference between the prediction of the initial network model on the training sample and the true label of the training sample. Therefore, it is necessary to determine the loss between the prediction result of the model and the true label by means of a loss function, so as to adjust the parameters of the model according to the loss between the prediction result and the true label. The training sample can also be input into the adjusted model, and the loss between the prediction result of the model and the true label is calculated and the model is adjusted according to the loss. Such iteration is performed until the loss between the prediction result of the model and the true label reaches a minimum value, the prediction result of the model approaches the true label of the training sample, and the model learns the ability to accurately classify the training sample.

[0098] Referring to the foregoing, assume that the true label of the training sample is “abnormal service”, and the output obtained by inputting the training sample into the initial network model is a 2*1 vector [x, y]. If x is greater than y, it indicates that the model predicts the training sample as “type 0”, i.e., the prediction result of the model on the training sample is “normal service”; if x is less than y, it indicates that the model predicts the training sample as “type 1”, i.e., abnormal service. When x in the output vector of the model is greater than y, it indicates that the model prediction is wrong, and there is a loss between the output of the model and the true label of the training sample, which needs to be iteratively trained according to the loss between the two until the loss between the output of the model and the true label reaches a minimum value.

[0099] In the embodiments of the present application, the historical data set of the target service is obtained according to the data selection strategy, and the historical data set includes at least two historical traffic data with evaluation results (used to represent whether the target service is an abnormal service). The training sample of the abnormality recognition model can also be determined based on the at least two historical traffic data. Since the historical traffic data has a clear evaluation result, the label of the training sample can be determined according to the evaluation result of the historical traffic data, so that the training sample with the label can be constructed. The problem that the positive feedback of the data cannot be obtained by blindly obtaining the training data, and the training sample with the label cannot be created, can be avoided. In addition, the feature used to identify whether the target service is an abnormal service (i.e., the feature matched with the target service) is determined, the training feature of the model is determined based on the above-mentioned feature, so that the model can extract some features for identifying whether the training sample is abnormal, and the model can also judge whether the training sample is abnormal based on the extracted features, thereby gradually learning the ability to identify whether the target service is an abnormal service based on the traffic data of the target service. The abnormality recognition model of the target service can be trained based on the training feature, the training sample, and the label of the training sample. The model can identify whether the target service is an abnormal service according to the traffic data of the target service, which overcomes the difficulty in model training of the abnormal network service identification model, and provides support for the application of the machine learning model in the abnormal network service identification.

[0100] In another embodiment of the present application, training instruction information can also be generated based on the training feature and the training sample set. The training instruction information can be used to guide the training of the abnormality recognition model. The risk control platform can send the training instruction information to the third-party model training platform, and the third-party model training platform can train the abnormality recognition model of the target service based on the above-mentioned training instruction information. Specifically, the training instruction information can indicate the training sample set of the abnormality recognition model and the training feature of the abnormality recognition model. After the training sample is input into the model, the training sample can be feature-extracted according to the training feature indicated by the training instruction information, so as to obtain the output of the model according to the extracted feature. The model is iteratively trained according to the loss between the output and the label of the training sample, so that the model gradually has the ability to identify whether the target service is an abnormal service based on the traffic data of the target service.

[0101] In a possible implementation manner, the training instruction information can include the training feature and the training sample set, and the training sample set includes the training sample and the label of the training sample. Alternatively, in order to save the data amount overhead, the training instruction information can include the training feature and the indication information of the training sample set. For example, the indication information can indicate the resource address of the training sample set. For example, the indication information can be a uniform resource locator (URL), and the training sample set of the abnormality recognition model can be obtained by accessing the URL address.

[0102] Of course, the training features can also be indicated by the indication information, which can include indication information of the training sample set (hereinafter referred to as first indication information) and indication information of the training features (hereinafter referred to as second indication information). For example, the indication information of the training features can be an index number. For example, N features matched with the target business are numbered as 1, 2, 3, …, N in turn, and it is assumed that the determined training features are features numbered as i, j, k, o (all are integers less than N), and the second indication information can be "i, j, k, o".

[0103] In the embodiment of the present application, the training process of the abnormality recognition model and the process of determining the model training indication information are decoupled, the model training object is flexibly configured, and the model training platform or other model training platform can be executed. In the specific training process of the abnormality recognition model, the model can be trained according to the training sample, the label of the training sample, and the training feature. The training feature is a key feature for identifying whether the target business is an abnormal business. The training sample and the label are input into the initial network model, so that the model can extract the key feature, and obtain the model prediction result based on the key feature. In addition, the initial network model is iteratively trained according to the loss between the prediction result and the real label, so that the model gradually has the ability to identify abnormal businesses based on business traffic data. The problem of difficult training of the abnormality recognition model is fundamentally solved, and the application of machine learning for abnormality recognition is realized.

[0104] In another embodiment of the present application, the features matched with the target business can be displayed through the GUI mode, so as to determine the training features of the abnormality recognition model in the manner of artificial configuration. For example, the specific implementation of determining the training features according to the features matched with the target business in the foregoing includes: displaying the first configuration item corresponding to the above-mentioned features. The first operation instruction for the first configuration item can also be received, and it is determined whether the feature corresponding to the first configuration item is the training feature according to the first operation instruction. For example, the first operation instruction for the first configuration item is a selected instruction, and it is determined that the feature corresponding to the first configuration item is the training feature.

[0105] It should be noted that the first configuration item corresponding to the above-mentioned features can be displayed through the risk control platform, so that the risk control platform can determine the multiple features configured by the artificial configuration according to the operation instruction of the first configuration item, thereby determining the training features of the abnormality recognition model. In a possible implementation manner, if the configuration interface of the background server of the risk control platform 10 includes a GUI, the configuration item of the above-mentioned features can be displayed through the GUI. For example, the risk control platform 10 provides a training feature configuration interface, and the training feature configuration interface can be logged in to perform the artificial configuration of the training features.

[0106] Figure 5 is a possible training feature configuration interface. Referring toFigure 5 The training feature configuration interface includes a plurality of candidate features (which can be features matched with the target service) and a first configuration item corresponding to each candidate feature. For example, the training feature configuration interface includes a candidate feature "username", a configuration item t1 corresponding to the candidate feature "username", a candidate feature "number of registration failures", a configuration item t2 corresponding to the candidate feature "number of registration failures", a candidate feature "number of verification code errors", and a configuration item t3 corresponding to the candidate feature "number of verification code errors". The configuration item corresponding to the candidate feature is used to select the candidate feature. For example, if the configuration item t1 and the configuration item t3 are selected, the training features configured by the human are "username" and "number of verification code errors".

[0107] It should be noted that, Figure 5 In the foregoing embodiment, the first configuration item is a check item, but the specific form of the first configuration item is not limited in the embodiment of the present application, and can be any configuration item for receiving the human configured feature. In addition, the first operation instruction is not limited, and the specific form of the first operation instruction is related to the first configuration item, and the selection of the configuration item corresponding feature can be realized. For example, the first configuration item is a check item, and the first operation instruction can be a click instruction.

[0108] In the embodiment of the present application, the training features of the anomaly recognition model can also be flexibly configured by the human, and the training features configured by the human can flexibly match the changes of the business scenario, so that the training features contain the key features of the abnormal service recognition while covering various changes of the business scenario, thereby ensuring the correct guidance of the training features to the anomaly recognition model training process.

[0109] In another embodiment of the present application, the data selection strategy mentioned in the foregoing can be configured by the human. For example, a plurality of candidate data sources are displayed through the GUI, and the data source configured by the human is obtained through the GUI, and then the data selection strategy can be determined according to the data source configured by the human, so as to select the historical traffic data with the positive feedback according to the data selection strategy, thereby providing a data basis for determining the training sample.

[0110] For example, before the historical data set of the target service is obtained according to the data selection strategy, a second configuration item corresponding to each candidate data source can also be displayed; a second operation instruction for the second configuration item is received, the data source for determining the historical data set from the plurality of candidate data sources is determined according to the second operation instruction, and the data selection strategy is generated according to the data source of the historical data set. The data selection strategy can indicate a plurality of data sources and the data amount proportion of each data source.

[0111] It should be noted that the second configuration item corresponding to the plurality of candidate data sources can be displayed by the risk control platform, so that the risk control platform can determine the plurality of data sources configured manually according to the operation instruction of the second configuration item, and generate the data selection strategy according to the data sources configured manually. In a possible implementation, if the configuration interface of the background server of the risk control platform 10 includes a GUI, the configuration item of the candidate data source can be displayed through the GUI. For example, the risk control platform 10 provides a data source configuration interface, and logging into the data source configuration interface can perform manual configuration of the candidate data source.

[0112] Figure 6 is a possible data source configuration interface. Referring to Figure 6 , the data source configuration interface includes a plurality of candidate data sources, and a second configuration item corresponding to each candidate data source. For example, the candidate data sources include: random risk control request samples, blue army verification samples, and expert review samples. The “random risk control request samples” refer to traffic data obtained from risk control requests initiated by a risk control demand party, and the risk control platform has a clear evaluation result for the traffic data. The “blue army verification samples” refer to traffic data generated by virtual abnormal objects (for example, virtual abnormal devices, virtual abnormal accounts), and the evaluation result of the traffic data is “abnormal”. The “expert review samples” refer to traffic data obtained from historical evaluation data of risk control experts, and the traffic data has a clear evaluation result. Referring to Figure 6 , the configuration items corresponding to the candidate data sources “blue army verification samples”, “random risk control request samples”, and “expert review samples” are p1, p2, and p3, respectively. For example, if the configuration item p2 and the configuration item p3 are selected, the data sources configured manually are “blue army verification samples” and “expert review samples”.

[0113] In a possible implementation, each candidate data source is further provided with a proportion setting item for inputting the proportion of the data of the data source, which can be the proportion of the data of the data source in the training sample set. For example, after the “blue army verification samples” and the “expert review samples” are selected, the proportion corresponding to the “blue army verification samples” is set to “30%”, and the proportion corresponding to the “expert review samples” is set to “70%”, that is, the obtained historical traffic data includes 30% of “blue army verification samples” and 70% of “expert review samples”.

[0114] It should be noted that Figure 6 the second configuration item is a check item, but the specific form of the second configuration item is not limited in the embodiments of the present application, and can be any configuration item receiving the data source configured manually. In addition, the second operation instruction is not limited, and the second operation instruction is related to the specific form of the second configuration item, and can realize the selection of the configuration item corresponding to the data source. For example, the second configuration item is a check item, and the second operation instruction can be a click instruction.

[0115] In the embodiments of the present application, the data source of the training sample set can also be flexibly configured in an artificial manner. The artificially configured data source can flexibly match changes in the business scenario, so that the training sample set can cover various changes in the business scenario, thereby improving the model performance.

[0116] In another embodiment of the present application, the trained anomaly recognition model can also be updated. Specifically, when the model update condition is met, the anomaly recognition model is updated. The model update condition includes at least one of the following:

[0117] (1) there is a difference between the output result of the anomaly recognition model and the artificial judgment result;

[0118] The output result of the anomaly recognition model is a prediction result obtained by inputting the traffic data of the target business into the anomaly recognition model. The artificial judgment result is a result obtained by manually identifying the same traffic data based on an artificial judgment rule. The artificial judgment rule can be a public anomaly recognition rule.

[0119] It should be noted that for the same prediction object, when there is a difference between the prediction result of the anomaly recognition model and the artificial judgment result, i.e., the prediction result of the model is different from the artificial judgment result. For example, for the same traffic data of the prediction object, the anomaly recognition model outputs a result indicating that the traffic data is normal, and the artificial judgment result indicates that the traffic data is abnormal, i.e., there is a difference between the prediction result of the anomaly recognition model and the artificial judgment result. The same prediction object can be the same traffic data, or data generated by the same object. For example, the same object can be the same account, the same device, etc.

[0120] Alternatively, the identification results of the anomaly recognition model and the artificial judgment on quantitative data can be collected regularly. For example, 1000 pieces of traffic data are collected every month, and the identification results of the anomaly recognition model and the artificial judgment on the 1000 pieces of traffic data are obtained. The number X of traffic data with identification differences is determined, and if the proportion of the number X is greater than a preset threshold (for example, X / 1000 is greater than 20%), it is determined that there is a difference between the prediction result of the anomaly recognition model and the artificial judgment result.

[0121] When the prediction result of the model and the artificial judgment result exist difference, it is possible that the model performance has declined and cannot correctly identify the traffic data of the target business, and the model can be updated.

[0122] (2) the anomaly recognition rule of the target business is changed;

[0123] It should be noted that the abnormality identification rule of the target service can include a judgment condition for identifying a key feature of an abnormal service, that is, the abnormality identification rule can represent "when the key feature meets which condition, the target service is determined to be an abnormal service". The abnormality identification rule is applicable to manual identification of abnormal services and an abnormality identification model.

[0124] The current abnormality identification model learns the ability to identify abnormalities based on the original identification rule. When the abnormality identification rule changes, the current abnormality identification model does not have the ability to identify abnormalities based on the new identification rule, the accuracy of model prediction will decrease, and the model cannot be compatible with the scene of abnormality identification rule change. The abnormality identification model for the target service needs to be updated according to the changed abnormality identification rule.

[0125] (3) The online duration of the abnormality identification model exceeds a preset duration;

[0126] It should be noted that after the initial network model is trained according to the training instruction information of the abnormality identification model to obtain the abnormality identification model of the target service, the abnormality identification model can be published online so that the abnormality identification model can be widely applied to the prediction of various traffic data of the target service. After the abnormality identification model is published online, the model can be updated and trained at regular intervals. Through regular updating, the performance of the model can be gradually improved. That is, when the online duration of the abnormality identification model exceeds the preset duration, the abnormality identification model can be updated. The preset duration can be the periodic length of model regular updating, for example, it can be one month, 15 days, etc.

[0127] In a possible implementation, the training instruction information of the abnormality identification model can also include the model update condition described above, so that the model training party can update the model when the model update condition is met.

[0128] In a possible implementation, updating the abnormality identification model specifically includes: determining a model update rule, the model update rule including updating a training sample, or updating a training feature, or updating both the training sample and the training feature.

[0129] Further, the abnormality recognition model is updated based on a model updating rule. Specifically, the updated training sample is input into the abnormality recognition model, the feature extracted by the convolution network of the abnormality recognition model is the original training feature, and the abnormality recognition model is iteratively trained based on the loss between the prediction result of the new training sample and the real label. Alternatively, the original training sample is input into the abnormality recognition model, the feature extracted by the convolution network of the abnormality recognition model is the changed training feature, and the abnormality recognition model is iteratively trained based on the loss between the prediction result of the training sample and the real label. Alternatively, the updated training sample is input into the abnormality recognition model, the feature extracted by the convolution network of the abnormality recognition model is the changed training feature, and the abnormality recognition model is iteratively trained based on the loss between the prediction result of the new training sample and the real label.

[0130] It should be noted that when there is a difference between the output result of the abnormality recognition model and the artificial judgment result, the training sample set (including the training sample and the label) and the training feature can be updated, or only the training sample set or the training feature can be updated. When the abnormality recognition rule of the target service changes, the training feature can be updated. When the online time of the abnormality recognition model exceeds the preset time, the training sample set and the training feature can be updated, or only the training sample set or the training feature can be updated.

[0131] The embodiments of the present application provide specific implementation methods for updating the abnormality recognition model, including the model updating time and the specific updating scheme, which can update the model in time according to various changes, and ensure the stability of the model performance and the accuracy of the model prediction.

[0132] In another embodiment of the present application, after updating the abnormality recognition model, the model training object (for example, a risk control platform or other model training platform) can also present the iterative updating situation of the abnormality recognition model, so as to facilitate the management personnel (for example, a model designer) of the risk control platform to check the iterative updating situation of the abnormality recognition model. For example, model updating information can be generated based on the abnormality recognition model and the updated abnormality recognition model, and the model updating information can be used to indicate the iterative relationship between the abnormality recognition model and the updated abnormality recognition model. For example, the model updating information can indicate the current version number and the historical version number of the abnormality recognition model. A browsing entrance of the model updating information can also be created, so that the management personnel of the risk control platform can check the model updating information through the browsing entrance and understand the updating and iterative situation of the model.

[0133] For example, the browsing entrance of the model updating information can be a function board as shown in the drawing. Figure 7 Figure 7 ​The function board includes version numbers v1.0, v1.1, and v2.0 of the anomaly identification model, where v2.0 is the currently online version, and v1.0 and v1.1 are historical versions of the anomaly identification model. The function board can also show the iteration order of the model versions using arrows, for example, v1.0, v1.1, and v2.0 in sequence.

[0134] In a possible implementation, the function board can also display other information of the anomaly identification model, for example, improvement points (for example, 20% improvement in accuracy) of the new version of the anomaly identification model, applicable scenarios, and the like.

[0135] In another embodiment of the present application, an artificial judgment rule can also be generated according to the features matched with the target business. For example, the features matched with the target business and the corresponding configuration items can be displayed through the risk control platform, and the artificial configuration of the artificial judgment rule can be performed through the configuration items. For example, the method described above further includes: displaying a third configuration item corresponding to the features matched with the target business; receiving a third operation instruction for the third configuration item, and determining a target feature from the features matched with the target business according to the third operation instruction; where the target feature can be a feature selected through the operation instruction of the third configuration item. Further, an artificial judgment rule can be generated according to the target feature and the execution logic of the target feature.

[0136] In a possible implementation, if the configuration interface of the background server of the risk control platform 10 includes a GUI, the artificial judgment rule configuration interface can be displayed through the GUI, and the features participating in the artificial judgment rule can be selected by logging into the artificial judgment rule configuration interface, so as to generate the artificial judgment rule.

[0137] Figure 8 is a possible artificial judgment rule configuration interface. Referring to Figure 8 The artificial judgment rule configuration interface includes a plurality of candidate features (which can be features matched with the target business) and configuration items (i.e., the third configuration items described above) corresponding to each candidate feature. For example, the candidate features “username”, “username” correspond to the configuration items R1, the candidate features “registration failure times”, “registration failure times” correspond to the configuration items R2, the candidate features “verification code error times”, and “verification code error times” correspond to the configuration items R3. The configuration items corresponding to the candidate features are used to select the candidate features. For example, if the configuration items R1 and R3 are selected, the target features used to generate the artificial judgment are “username” and “verification code error times”.

[0138] The artificial judgment rule configuration interface can further include a parameter configuration item of each feature. The parameter configuration item is used to obtain a judgment condition of the feature participating in the artificial judgment, but a specific attribute or parameter of the feature. For example, a specific value x of "verification code error times" and a specific character "xabijpo" of "username".

[0139] The artificial judgment rule configuration interface can further include a "confirm" button used to trigger generation of the artificial judgment rule. When the "confirm" button is triggered, the artificial judgment rule can be generated according to the "username", "verification code error times", execution logic of the two, and the specific attribute parameters of the two. For example, the artificial judgment rule is: if the "verification code error times" of the traffic data exceeds x, and the "username" contains xabijpo, the traffic data is determined as abnormal traffic.

[0140] It should be noted that, Figure 8 The third configuration item is a check item in this embodiment of the present application, but the present application does not limit the specific form of the third configuration item. It can be any configuration item receiving artificial configuration of the feature. In addition, the third operation instruction is not limited. The third operation instruction is related to the specific form of the third configuration item, and can realize selection of the configuration item corresponding to the feature. For example, the third configuration item is a check item, and the third operation instruction can be a click instruction.

[0141] In a possible implementation manner, after the artificial judgment rule is generated, the artificial judgment rule can be presented. The artificial judgment rule can be used for artificial identification of the traffic data, that is, the traffic data is artificially identified by referring to the artificial judgment rule. For example, a browsing entrance of the artificial judgment rule can be created. The browsing entrance is used for a user of the artificial judgment rule to obtain the artificial judgment rule. The user of the artificial judgment rule can be an audit personnel (for example, a risk assessment expert) of a network risk control service. The audit personnel can artificially identify the traffic data by referring to the artificial judgment rule.

[0142] For example, the browsing entrance can be a microservice. Referring to Figure 9 After the terminal device of the audit personnel calls the microservice, the terminal device can output the artificial judgment rule. For example, the display interface of the terminal device can display the artificial judgment rule: if the "verification code error times" of the traffic data exceeds x, and the "username" contains xabijpo, the traffic data is determined as abnormal traffic.

[0143] In this embodiment of the present application, reference features are provided for the artificial judgment rule. Compared with completely relying on artificial experience to generate the artificial judgment rule, this embodiment of the present application can provide more reference features, can cover more business scenarios, and can improve the applicability of the artificial judgment rule.

[0144] In another embodiment of the present application, after the anomaly identification model of the target service is trained, the performance of the model can also be evaluated. After the performance evaluation of the anomaly identification model, the model is released online. For example, an evaluation sample set of the anomaly identification model is obtained; the evaluation sample set includes a positive sample set and a negative sample set. The evaluation sample set is input into the anomaly identification model, and the performance of the anomaly identification model is evaluated according to the output of the anomaly identification model.

[0145] It can be understood that when the performance of the model is evaluated, the positive sample can be a sample belonging to a specified type, and the negative sample is a sample not belonging to the specified type. The specified type can be a type that the model can identify. For example, when the anomaly identification model is evaluated, the positive sample set includes a plurality of positive samples, and the positive sample is a sample with a label of "normal service". The negative sample set includes a plurality of negative samples, and the negative sample is a sample with a label of "abnormal service".

[0146] It should be noted that the evaluation samples can be classified according to the true label of the evaluation samples and the prediction result of the anomaly identification model on the evaluation samples. For specific classification, refer to Table 2:

[0147] Table 2

[0148] Positive Negative TRUE True Positive (TP) True Negative (TN) FALSE False Positive (FP) False Negative (FN)

[0149] Wherein, Positive, Negative represent the prediction type obtained by the anomaly identification model for predicting the evaluation sample, Positive is the model prediction of "normal service", and Negative is the model prediction of "abnormal service". TRUE, FALSE represent the true type (true label) of the evaluation sample, TRUE is the evaluation sample with a label of "normal service", and FALSE is the evaluation sample with a label of "abnormal service".

[0150] TP represents that the label of the evaluation sample is "normal service", and the anomaly identification model also predicts the evaluation sample as "normal service"; TN represents that the label of the evaluation sample is "abnormal service", and the anomaly identification model also predicts the evaluation sample as "abnormal service"; FP represents that the label of the evaluation sample is "abnormal service", but the anomaly identification model predicts the evaluation sample as "normal service"; FN represents that the label of the evaluation sample is "normal service", but the anomaly identification model predicts the evaluation sample as "abnormal service".

[0151] For example, the following indicators of the anomaly identification model can be determined according to the prediction result output by the model, and the performance of the anomaly identification model is evaluated based on these indicators:

[0152] (1) Accuracy: the proportion of samples correctly classified by the anomaly recognition model among all evaluation samples. The higher the accuracy, the better the performance of the anomaly recognition model.

[0153] For example,

[0154] (2) Error rate: contrary to accuracy, it represents the proportion of samples incorrectly classified by the anomaly recognition model among all evaluation samples. The lower the error rate, the better the performance of the anomaly recognition model.

[0155] Error rate = (FP + FN) / (TP + TN + FP + FN) = 1 - Accuracy.

[0156] (3) Sensitivity: represents the proportion of samples correctly classified among all positive samples, measuring the ability of the anomaly recognition model to identify positive examples. For example, sensitivity = TP / P.

[0157] (4) Specificity: represents the proportion of samples correctly classified among all negative samples, measuring the ability of the anomaly recognition model to identify negative examples. Specificity = TN / N.

[0158] (5) Precision: represents the proportion of samples with actual label "normal business" among samples classified as positive examples "normal business" by the anomaly recognition model. Precision = TP / (TP + FP).

[0159] (6) Recall: used to represent how many positive samples are classified as positive examples, recall = TP / (TP + FN) = TP / P = sensitivity.

[0160] Take accuracy as an example to introduce the performance evaluation process of the anomaly recognition model. Referring to Figure 10 , input positive and negative samples into the anomaly recognition model, and evaluate the performance of the anomaly recognition model according to the output results of the anomaly recognition model. If the prediction result of the anomaly recognition model for the evaluation sample is the same as the true label of the evaluation sample, it indicates that the sample is correctly classified by the anomaly recognition model. Assuming that the total number of samples in the evaluation sample set is Y, and the number of samples correctly classified by the model is X, the accuracy of the anomaly recognition model is X / Y. The larger the value of X / Y, the better the performance of the model, otherwise, the smaller the value of X / Y, the worse the performance of the model.

[0161] The method provided in the embodiments of the present application can further perform performance evaluation on the model after the abnormality identification model of the target service is trained, and publish the model online in the case that the model passes the performance evaluation. The performance advantage of the online model is guaranteed, and an abnormality identification model with better performance is provided.

[0162] It can be understood that, due to the particularity of abnormal services in the risk control scene, the historical data often has no positive feedback, which not only causes certain difficulties in constructing training samples, but also causes certain difficulties in model performance evaluation. In another embodiment of the present application, accurate negative samples can be acquired to construct an evaluation sample set, and the evaluation sample set is constructed in combination with positive samples, which not only guarantees that the qualified recall rate (i.e., accuracy, recall rate) of the evaluation sample is qualified, but also guarantees the balance of the number distribution of the positive samples and the negative samples. The accurate negative sample can be a sample that is rechecked from the initially acquired negative sample and the result is still a negative sample.

[0163] For example, referring to Figure 11 The negative sample set referred to in the foregoing can be acquired through the following process: first, an initial negative sample set is determined according to a historical data set, wherein the historical traffic data in the initial negative sample set is different from the historical traffic data in the training sample. That is, the negative example data with an evaluation result of “abnormal” can also be acquired from the historical traffic data of the target service as the initial negative sample set. In addition, in order to accurately evaluate the model performance, the data used for model training is different from the data used for model evaluation, that is, the data used as negative samples in the historical traffic data of the target service is different from the negative example data participating in model training.

[0164] That is, according to the data selection strategy when step 402 is performed, a large amount of historical traffic data with positive feedback can be acquired. The historical traffic data can be used as training samples to train the model, or can be used as evaluation samples to evaluate the performance of the trained abnormality identification model. It can be understood that, in order to improve the accuracy of performance evaluation, the data used for model training and model evaluation is often different. For example, 70% of the historical traffic data can be used as training samples, and the remaining 30% can be used as evaluation samples. Alternatively, the initial negative sample set is determined according to the negative samples in the remaining 30% samples, and the samples that are still negative examples after rechecking are used as negative samples for model evaluation.

[0165] It can be understood that, in order to improve the accuracy of negative samples, the initial negative samples in the historical data set can be rechecked to determine the samples with a rechecking result of negative samples, that is, the accurate negative samples described in the foregoing. Specifically, the initial negative samples can be divided into two parts, one part is rechecked in an artificial manner, and the other part is rechecked by using a known model.

[0166] For example, the manual review samples are determined from the initial negative sample set, the manual review samples are reviewed by a human (e.g., a risk control expert), and if the manual review result of the sample is a negative sample, i.e., the human identifies the sample as "abnormal business", the sample is determined as an accurate negative sample. The review personnel can also upload the accurate negative samples obtained by manual review to the risk control platform, so that the accurate negative samples in the manual review samples can be obtained.

[0167] In addition, the remaining samples (i.e., the remaining negative samples in the initial negative sample set except for the manual review samples) in the initial negative sample set can also be verified based on the early warning model. Specifically, the remaining samples can be input into the early warning model, and if the output result of the early warning model represents that the remaining samples are risk samples, the remaining samples are determined as accurate negative samples. Thus, the accurate negative samples in the remaining samples can be obtained through repeated verification of the early warning model. In one possible implementation, the early warning model can be a machine learning model for risk early warning in the network risk control field, and the early warning model is used to identify network risk business according to network traffic data. For example, the input of the early warning model is network traffic data, and the output is used to represent whether the input network traffic data has network risk. For example, the early warning model can be a model for identifying illegal network transactions.

[0168] Finally, the negative sample set in the evaluation sample set can be determined according to the accurate negative samples in the manual review samples and the accurate negative samples in the remaining samples. For example, the negative sample set is constructed according to the accurate negative samples in the manual review samples and the accurate negative samples in the remaining samples.

[0169] It should be noted that the traffic data can be obtained from the risk control request as positive samples. The risk control request can be initiated by the risk control demand party to the risk control platform, carrying the traffic data of the business, and used to request the risk control platform to perform abnormal identification based on the traffic data in the risk control request.

[0170] In one possible implementation, the negative sample set can also include other negative example data. For example, the negative sample set further includes: historical traffic data generated by a virtual abnormal object, historical traffic data identified by a human (e.g., a risk control expert) as abnormal business, and data missed by a human judgment rule.

[0171] It should be noted that the virtual abnormal object refers to a virtual abnormal account or a virtual abnormal device, etc. The historical traffic data generated by the virtual abnormal object is often explicit negative example data, that is, the historical traffic data generated by the virtual abnormal object is definitely abnormal data, which can be used as negative samples for evaluating the abnormal identification model. The data missed by the human judgment rule can be historical data whose human identification result is normal, but whose review result after online complaint by the risk control demand party is abnormal.

[0172] In a possible implementation, the proportion of each type of negative sample in the negative sample set can be adjusted according to an actual business scenario. For example, the proportion of accurate negative samples in the artificial review samples in the negative sample set is 30%, the proportion of accurate negative samples in the remaining samples of the initial sample set is 30%, the proportion of the number of historical traffic data generated by the virtual abnormal object is 20%, the proportion of the historical traffic data with the artificial identification result as abnormal business is 10%, and the proportion of the data missed by the artificial judgment rule is 10%.

[0173] The embodiments of the present application also provide an abnormality identification system, and the background server 101 of the risk control platform 10 supports the operation of the system. The system can include training of the abnormality identification model, publishing of the artificial judgment rule, iterative updating of the abnormality identification model, and the like.

[0174] Reference Figure 12 First, a plurality of features matched with the target business can be determined from the feature library according to the business identifier of the target business. In addition, the online of the artificial judgment rule (referred to as rule development) and the training of the abnormality identification model (referred to as model development) can be performed based on the plurality of features.

[0175] Specifically, the rule canvas is performed based on the matched features to obtain the artificial judgment rule. The artificial judgment rule can also be published in the form of microservices, and the microservices can serve as a browsing entry of the artificial judgment rule. The rule canvas includes feature selection and feature arrangement, that is, selecting features participating in artificial judgment from the matched features, and arranging the selected features according to the execution logic of each feature, thereby generating the artificial judgment rule.

[0176] In the training process of the abnormality identification model, the risk control platform 10 can determine the training features of the model according to the matched features, and can also set the data selection strategy of the training sample and the model update condition. Then, the data selection strategy of the training sample can obtain the training sample set from the evaluated historical traffic data, and the model training instruction information can be generated according to the training features of the model, the training sample set, and the model update condition, and the model training instruction information is sent to the model training platform. Of course, the model training can also be performed by the risk control platform 10 according to the training features of the model and the training sample set to train the abnormality identification model of the target business.

[0177] The model training platform can perform model training according to the training sample set and the training features. After the model training, the performance of the abnormality identification model is evaluated by the evaluation system. For example, referring to Figure 12 After the model training, the model microservice can be published, and the pre-published model can be obtained by calling the microservice. The evaluation system can evaluate the pre-published model, and after the evaluation, the model is formally published online.

[0178] The evaluation system can construct negative samples based on accurate negative samples of manual review, historical data generated by virtual abnormal objects, accurate negative samples verified by early warning models, and missed data of manual judgment rules. The evaluation system can also construct an evaluation sample set according to the negative samples and the positive samples, input the evaluation sample set into the abnormal identification model, and evaluate the performance of the model according to the output of the model.

[0179] If the abnormal identification model passes the performance evaluation of the evaluation system, the abnormal identification model is released online. The released abnormal identification model can identify whether the target business is an abnormal business according to the traffic data of the target business. The traffic data of the target business is input into the abnormal identification model, and the output of the model can be used to determine whether the current traffic data is abnormal. If the abnormal identification model does not pass the performance evaluation of the evaluation system, the model training instruction information can be updated so that the model training platform updates the abnormal identification model according to the updated training instruction information. For example, the updated training instruction information can include updated training features and updated training sample sets (including training samples and labels of the training samples).

[0180] In addition, the system can also analyze the difference between the judgment results of the manual and the model, that is, determine the difference between the identification results of the manual judgment rules and the prediction results of the model. When the difference between the two reaches a certain degree, the abnormal identification model can be updated.

[0181] It should be noted that the update conditions of the abnormal identification model can include that the difference between the identification results of the manual judgment rules and the prediction results of the model is greater than a threshold, that the update is performed at a fixed time, and that the abnormal identification rules of the target business change.

[0182] The system can also provide Figure 5 The training feature configuration interface shown in FIG. 10 supports the implementation of the system to "select training features". Specifically, receiving operation instructions of configuration options of the training feature configuration interface can determine the training features of the abnormal identification model.

[0183] The system can also provide Figure 6 The data source configuration interface shown in FIG. 11 supports the implementation of the system to "select training samples". Specifically, receiving operation instructions of configuration options of the data source configuration interface can determine the data selection strategy of the training sample set.

[0184] In a possible implementation, the system can also provide a browsing entry of model update information. For example, the system can provide the function board shown in FIG. 12 to display the version update of the abnormal identification model. Figure 7

[0185] In another possible implementation, the system can provide​Figure 8 The artificial judgment rule configuration interface shown supports the implementation of the rule development of the system. Specifically, an operation instruction of a configuration option of the artificial judgment rule configuration interface can determine the features participating in the artificial judgment, so as to generate the artificial judgment rule based on the features.

[0186] The embodiment of the application further provides a specific implementation manner of the evaluation system for identifying the performance of the anomaly identification model. Referring to Figure 13 , the flow data of the risk control request initiated by the risk control platform in the risk control demand direction may have no positive feedback, and the evaluation sample set used by the evaluation system is no longer taken from random risk control requests, but an initial negative sample set is obtained from the historical flow data hit by the artificial judgment rule. For example, the initial negative sample set is constructed by selecting negative samples from the negative samples hit by the artificial judgment rule in a certain proportion.

[0187] Further, the initial negative sample set can be divided into two parts, one part is manually reviewed, and the other part is cross-validated by the existing early warning model, so as to obtain accurate negative samples. Thus, samples whose artificial review results are still negative samples (referred to as artificial review negative samples) and samples whose cross-validation results are still negative samples (referred to as model validation negative samples) can be obtained.

[0188] Specifically, the proportion of each type of negative sample is first determined, and then each type of negative sample is determined according to the respective proportion. For example, the artificial review negative sample is 30%, the model validation negative sample is 40%, and the blue army negative sample is 30%. Among them, the blue army negative sample is the abnormal data in the blue army data, and the blue army can be a virtual abnormal account or an abnormal device.

[0189] According to the artificial review negative sample, the model validation negative sample, and the blue army negative sample, an accurate negative sample set can be constructed, and a positive sample set can also be obtained. Further, based on the accurate negative sample set and the positive sample set, an evaluation sample set is constructed, and the evaluation sample set is input into the anomaly identification model to evaluate the performance of the model.

[0190] The anomaly identification system provided by the embodiment of the application can identify anomalies by means of a machine learning model, improve the objectivity and accuracy of anomaly identification, and greatly improve the efficiency of anomaly identification in the field of risk control. In addition, the key features of the business are matched by the business identification to identify anomalies, so that the artificial judgment rule and the anomaly identification model can cover as many key features related to the business as possible. Further, the artificial configuration function of the features is provided, which is conducive to extracting key header features related to the target business anomaly identification by means of priori, and improving the accuracy of the artificial judgment rule and the anomaly identification model.

[0191] It should be noted that the effect observation of the online abnormality identification model can also be combined with specific business scenarios. The effect observation is different from the performance evaluation described above, and there is no fixed quantitative index for measuring the effect. Instead, the model performance is observed according to the application of the model in a specific scenario. For example, the abnormality identification model is applied to the "add friend" business for abnormality identification. The performance of the model can be evaluated according to the changes in specific business indicators such as the number of failed friend adding and the number of account nurturing after the model is online for a period of time. For example, if the values of the business indicators such as the number of failed friend adding and the number of account nurturing decrease, it indicates that the abnormality identification model can identify the abnormal business in the "add friend" business and has a certain effect on business risk control.

[0192] In addition, the specific business indicators for model effect observation are not fixed, and the change factors are complex, which can include business promotion activities, business processing measures that cannot be monitored and fed back, and the effects of multiple strategy rules.

[0193] The embodiment of the present application also provides an abnormality identification method, which refers to Figure 14 The method comprises the following steps:

[0194] 1401, obtaining traffic data of a target business;

[0195] In the implementation, in the field of risk control, when a risk control demander needs to perform abnormality identification on the traffic data of a target business, the risk control demander can send a risk control request to a risk control platform. The risk control request can include the traffic data of the target business, and the risk control platform can obtain the traffic data of the target business by receiving the risk control request. For example, for an "add friend" business, the traffic data can include "requester username", "requester device information", and "number of failed friend adding".

[0196] 1402, inputting the traffic data of the target business into an abnormality identification model of the target business, and identifying whether the target business is an abnormal business according to the output of the abnormality identification model.

[0197] It should be noted that the abnormality identification model can be a classification model, for example, a binary classification model or a multi-classification model. In one possible implementation, the abnormality identification model is a binary classification model, which can predict two types, and the output result of the model is one of the two categories. For example, the output of the model is "abnormal business" or "normal business".

[0198] The above abnormality identification model can include a convolution network and a classification network. The processing flow of the traffic data to be identified after inputting into the abnormality identification model can include: first, the convolution network extracts features of the traffic data. The extracted features can also be input into the classification network, and the classification function of the classification network performs operations on the features to finally obtain the prediction result of the model output.

[0199] For example, assuming that the target service is a user registration service, after the traffic data is input into the anomaly identification model, the convolutional network can extract the features of the traffic data, i.e., the username w, the number of registration failures t, and the number of verification code errors k. Then, the username w, the number of registration failures t, and the number of verification code errors k are input into the classification network, which predicts whether the traffic data is abnormal based on the username w, the number of registration failures t, and the number of verification code errors k.

[0200] The following also provides another processing flow of the anomaly identification model for the input traffic data. Specifically, when the traffic data is input into the anomaly identification model, the traffic data can first be processed by shunting, and the text therein is input into the text recognition network. The text recognition network can perform semantic recognition on the text, and can also perform word segmentation on the semantic recognition result. The obtained word segmentation is further recognized, and candidate word segmentation matching the training feature semantics is extracted.

[0201] In addition, through the above shunting processing, the image in the traffic data can be input into the image recognition network. The image recognition network can use a convolutional network to extract features from the image and obtain a feature map. Specifically, the convolutional network can be determined according to the training features, and the convolutional network has the ability to extract the above training features from the image. Different convolutional networks correspond to different convolutional channels, such as a channel for extracting eye features, a channel for extracting mouth features, and the like. In the embodiments of the present application, in order to improve the efficient extraction of local detailed features by the anomaly identification model, different weight coefficients can be assigned to each channel according to the degree of attention. In addition, for different regions within the same feature map, different weight coefficients can also be set for each region according to the degree of attention, to achieve more detailed feature extraction. That is, when using the convolutional network to extract the feature map, the feature map can be adjusted according to the weight coefficient of the corresponding convolutional channel, and the feature map can be further updated according to the weight coefficient of each region in the feature map. Finally, the image recognition network outputs multiple candidate feature maps extracted therefrom.

[0202] Finally, the above candidate word segmentation and candidate feature map can be input into the classification network. The classification network can respectively perform vectorization processing on the candidate word segmentation and the candidate feature map, and then fuse the vectors obtained by the vectorization processing to obtain a fusion vector. Then, the fusion vector is input into the classification function of the classification network for operation, and the prediction result of the anomaly identification model for the traffic data is obtained.

[0203] Exemplarily, the prediction result of the anomaly identification model is introduced taking binary classification as an example. The types that can be predicted by the anomaly identification model are represented by 0 and 1, where "1" represents "abnormal service" and "0" represents "normal service". The output of the anomaly identification model can be a 2*1 vector, each element in the vector corresponds to a type, for example, in the order of row arrangement, it corresponds to type 1 and type 0 in turn. The value (score value) of the element represents the probability that the input data is of the corresponding type, and the sum of the score values is equal to 1. The higher the score value, the higher the probability that the traffic data hits the type corresponding to the score value, so the type corresponding to the highest score value can be selected as the type predicted by the model. Assuming that the output of the anomaly identification model is [0.33, 0.67], it means that the probability that the input traffic data is of type 1 is 0.33 and the probability that the traffic data is of type 0 is 0.67. At this time, "type 0" with the highest score value can be taken as the prediction result of the anomaly identification model, that is, the model identifies the traffic data as "normal service" data.

[0204] In a possible implementation, the training sample of the anomaly identification model is determined according to at least two historical traffic data of the target service. Specifically, the historical traffic data is a training sample in the training sample set, and the evaluation result of the historical traffic data is the label of the training sample. The evaluation result of the historical traffic data can represent whether the target service is an abnormal service. In addition, the training feature of the anomaly identification model is determined according to the feature matched with the target service, and the feature is a key indicator for identifying whether the target service is an abnormal service. It should be noted that the training process of the anomaly identification model is described in the foregoing method, which will not be repeated here. Figure 3

[0205] In the embodiments of the present application, the traffic data can be identified by means of the anomaly identification model. Compared with the prior art that relies on risk control experts to manually identify abnormal network services, the efficiency of anomaly identification is greatly improved, and the subjectivity of anomaly identification can be avoided, thereby providing an objective and accurate anomaly identification scheme.

[0206] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. The program is executed by a processor to implement the steps of the method described in the embodiments of the present application. Figure 3 or Figure 14

[0207] The embodiments of the present application provide a computer program product, which contains instructions. The instructions are executed by a processor to implement the steps of the method described in the foregoing Figure 3 or Figure 14

[0208] ​​​It is to be noted that, although the operations of the method of the present application are described in a specific order in the accompanying drawings, this does not require or imply that the operations must be performed in that specific order, or that all of the illustrated operations must be performed to achieve the desired result.

[0209] Figure 15 A block schematic diagram of a model training apparatus according to an embodiment of the present application. Referring to Figure 15 The apparatus comprises a feature determining unit 1501, an obtaining unit 1502, a sample determining unit 1503, and a training unit 1504.

[0210] In a specific implementation, the feature determining unit 1501 is configured to obtain an identifier of a target service, determine a feature matched with the target service according to the identifier of the target service, and determine a training feature of an anomaly recognition model according to the feature matched with the target service; the feature is used to identify whether the target service is an abnormal service, and the anomaly recognition model is used to identify whether the target service is an abnormal service according to traffic data of the target service.

[0211] The obtaining unit 1502 is configured to obtain a historical data set of the target service according to a data selection strategy; the historical data set comprises at least two historical traffic data of the target service and an evaluation result corresponding to each historical traffic data, and the evaluation result is used to represent whether the target service is an abnormal service.

[0212] The sample determining unit 1503 is configured to determine a training sample of the anomaly recognition model according to the at least two historical traffic data, and determine a label of the training sample according to the evaluation result of the at least two historical traffic data.

[0213] The training unit 1504 is configured to train the anomaly recognition model based on the training sample, the label of the training sample, and the training feature.

[0214] In one embodiment, the training unit 1504 is specifically configured to determine an initial network model according to the training feature.

[0215] The training unit 1504 is configured to train the anomaly recognition model based on the training sample, the label of the training sample, and the training feature.

[0216] The training unit 1504 is configured to train the anomaly recognition model based on the training sample, the label of the training sample, and the training feature.

[0217] In one embodiment, the feature determining unit 1501 is specifically configured to display a first configuration item corresponding to the feature.

[0218] The feature determining unit 1501 is configured to receive a first operation instruction for the first configuration item, and determine, according to the first operation instruction, that the feature corresponding to the first configuration item is the training feature.

[0219] In an embodiment, the obtaining unit 1502 is specifically configured to display the second configuration items corresponding to the plurality of candidate data sources.

[0220] The second operation instruction for the second configuration item is received, the data source of the historical data set is determined from the plurality of candidate data sources according to the second operation instruction, and the data selection strategy is generated according to the data source of the historical data.

[0221] In an embodiment, the training unit 1504 is further configured to update the anomaly identification model if a model update condition is met.

[0222] The model update condition includes at least one of the following: a difference between the output result of the anomaly identification model and the artificial judgment result, a change in the anomaly identification rule of the target business, and an online duration of the anomaly identification model exceeding a preset duration; and the artificial judgment result is a result obtained by artificially identifying the traffic data of the target business based on an artificial judgment rule.

[0223] In an embodiment, the training unit 1504 is specifically configured to determine a model update rule, and the model update rule includes at least one of the following: updating a training sample set and updating a training feature.

[0224] The anomaly identification model is updated based on the model update rule.

[0225] In an embodiment, the training unit 1504 is further configured to generate model update information based on the anomaly identification model and the updated anomaly identification model; and the model update information is used to indicate an iteration relationship between the anomaly identification model and the updated anomaly identification model.

[0226] A browsing entry of the model update information is created.

[0227] In an embodiment, the feature determination unit 1501 is further configured to display a third configuration item corresponding to the feature.

[0228] The third operation instruction for the third configuration item is received, and the target feature is determined from the features matched with the target business according to the third operation instruction.

[0229] The artificial judgment rule is generated according to the target feature and the execution logic of the target feature; and the artificial judgment rule is used to artificially identify whether the traffic data is abnormal.

[0230] In an embodiment, the feature determination unit 1501 is further configured to create a browsing entry of the artificial judgment rule; and the browsing entry is used for a use object of the artificial judgment rule to obtain the artificial judgment rule.

[0231] In an embodiment, the training unit 1501 is further configured to obtain an evaluation sample set of the anomaly identification model, wherein the evaluation sample set comprises a positive sample set and a negative sample set.

[0232] The evaluation sample set is input into the anomaly identification model, and performance of the anomaly identification model is evaluated according to an output of the anomaly identification model.

[0233] In an embodiment, the training unit 1501 is specifically configured to determine an initial negative sample set according to the historical data set, wherein historical traffic data in the initial negative sample set is different from historical traffic data in the training sample.

[0234] An artificial review sample is determined from the initial negative sample set, and an accurate negative sample in which an artificial review result is a negative sample is obtained from the artificial review sample.

[0235] The remaining samples in the initial negative sample set are verified based on a pre-warning model to obtain accurate negative samples in the remaining samples, wherein the pre-warning model is used to identify network risk services according to network traffic data.

[0236] The negative sample set is determined according to the accurate negative samples in the artificial review sample and the accurate negative samples in the remaining samples.

[0237] In an embodiment, the training unit 1501 is specifically configured to input the remaining samples into the pre-warning model, and if an output result of the pre-warning model indicates that the remaining samples are risk samples, the remaining samples are determined as accurate negative samples.

[0238] It should be understood that the units described in the model training method correspond to the steps in the method described above. Figure 3 The operations and features described above for the method also apply to the model training apparatus and the units included therein, and will not be described here. The model training apparatus can be pre- implemented in a browser or other security application of a computer device, or can be loaded into the browser or security application thereof of the computer device through downloading or the like. The corresponding units in the model training apparatus can cooperate with the units in the computer device to realize the schemes of the embodiments of the present application.

[0239] The model training apparatus provided in the embodiments of the present application, in the present application, obtains a historical data set of a target service according to a data selection strategy, and the historical data set includes at least two historical traffic data with evaluation results (used to represent whether the target service is an abnormal service). The training sample of the abnormality identification model can also be determined based on the at least two historical traffic data. Since the historical traffic data has a clear evaluation result, the label of the training sample can be determined according to the evaluation result of the historical traffic data, so that the training sample with the label can be constructed. The problem that the positive feedback of the training data cannot be obtained and the training sample with the label cannot be created can be avoided. In addition, the feature used to identify whether the target service is an abnormal service (i.e., the feature matched with the target service) is determined, the training feature of the model is determined based on the feature, so that the model can extract some features for identifying whether the training sample is abnormal. The model can also judge whether the training sample is abnormal based on the extracted features, so as to gradually learn the ability to identify whether the target service is an abnormal service based on the traffic data of the target service. The abnormality identification model of the target service can be trained based on the training feature, the training sample and the label of the training sample. The model can identify whether the target service is an abnormal service according to the traffic data of the target service, which overcomes the problem of the model training of the abnormal network service identification model, and provides support for the application of the machine learning model in the abnormal network service identification.

[0240] The embodiments of the present application also provide an abnormality identification apparatus, referring to Figure 16 The apparatus includes an obtaining unit 1601 and an identification unit 1602.

[0241] The obtaining unit 1601 is configured to obtain traffic data of a target service.

[0242] The identification unit 1602 is configured to input the traffic data into an abnormality identification model of the target service, and identify whether the target service is an abnormal service according to the output of the abnormality identification model.

[0243] The training feature of the abnormality identification model is determined according to the feature matched with the target service, and the feature is used to identify whether the target service is an abnormal service. The training sample of the abnormality identification model is determined according to at least two historical traffic data of the target service, and the label of the training sample is determined according to the evaluation result of the at least two historical traffic data. The evaluation result is used to represent whether the target service is an abnormal service.

[0244] The abnormality identification apparatus provided in the embodiments of the present application can identify the traffic data by means of the abnormality identification model. Compared with the prior art which relies on the manual identification of abnormal network services by risk control experts, the efficiency of abnormality identification is greatly improved, and the subjectivity of abnormality identification can be avoided. An objective and accurate abnormality identification scheme is provided.

[0245] It should be understood that the units described in the anomaly identification apparatus correspond to the units described in the method Figure 14 The individual steps in the described method correspond. Thus, the operations and features described above in relation to the method also apply to the anomaly identification apparatus and the units comprised therein, which will not be described again. The anomaly identification apparatus can be pre-implemented in a browser or other security application of a computer device, or can be loaded into the browser or security application thereof of a computer device by downloading or the like. The respective units in the anomaly identification apparatus can cooperate with the units in the computer device to realize the solutions of the embodiments of the present application.

[0246] In the foregoing detailed description, several modules or units are mentioned. The division into such modules or units is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into a plurality of modules or units.

[0247] It should be noted that the details of the anomaly identification apparatus and the model training apparatus of the embodiments of the present application that are not disclosed are described with reference to the details disclosed in the above embodiments of the present application, which will not be described again.

[0248] The following refers to Figure 17 , Figure 17 A structural schematic diagram of a computer device suitable for implementing the embodiments of the present application is shown. As shown in Figure 17 The computer system 1700 includes a central processing unit (CPU) 1701, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1702 or programs loaded from a storage portion 1708 into a random access memory (RAM) 1703. Various programs and data required for operation instructions of the system are also stored in the RAM 1703. The CPU 1701, the ROM 1702, and the RAM 1703 are connected to each other through a bus 1704. An input / output (I / O) interface 1705 is also connected to the bus 1704.

[0249] The following components are connected to the I / O interface 1705; an input part 1706 including a keyboard, a mouse, etc.; an output part 1707 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 1708 including a hard disk, etc.; and a communication part 1709 including a network interface card such as a LAN card, a modem, etc. The communication part 1709 performs communication processing via a network such as the Internet. A drive 1710 is also connected to the I / O interface 1705 as necessary. A removable media 1711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 1710 as necessary, so that a computer program read therefrom is installed in the storage part 1708 as necessary.

[0250] In particular, the processes described above with reference to flowchart Figure 3 , FIG. 4 or Figure 15 may be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such an embodiment, the computer program contains program code for executing the methods illustrated by the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network by the communication part 1709, and / or installed from the removable media 1711. When the computer program is executed by the central processing unit (CPU) 1701, the above-described functions defined in the system of the present application are executed.

[0251] It should be noted that the computer-readable medium shown in the application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component. In this application, the computer-readable signal medium can include a data signal carried in a baseband or as a carrier wave part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or component. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0252] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions and operation instructions of the systems, methods and computer program products according to various embodiments of the application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can also occur in different order from that noted in the drawings. For example, two connected blocks can actually be executed substantially in parallel, and sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operation instructions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0253] The units or modules described in the embodiments of the present application can be implemented in the form of software, or can be implemented in the form of hardware. The described units or modules can also be arranged in a processor, for example, can be described as: a processor includes a first collection module, a second collection module and a sending module. Among them, the name of these units or modules does not constitute a limitation of the unit or module itself in some cases.

[0254] As another aspect, the present application also provides a computer readable storage medium, which can be contained in the electronic device described in the above embodiments, or can exist independently without being assembled into the electronic device. The above computer readable storage medium stores one or more programs, when the above programs are used by one or more processors to execute the model training method and the anomaly identification method described in the present application.

[0255] The above description is merely preferred embodiments of the present application and a description of the principles of the technology used. Those skilled in the art should understand that the disclosed range in the present application is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features disclosed in the present application (but not limited to) having similar functions to form technical solutions.

Claims

1. A model training method, characterized in that, include: Obtain the identifier of the target service, determine the features that match the target service based on the identifier of the target service, and determine the training features of the anomaly recognition model based on the features that match the target service; The feature is used to identify whether the target service is an abnormal service, and the abnormality identification model is used to identify whether the target service is an abnormal service based on the traffic data of the target service; The historical dataset of the target service is obtained according to the data selection strategy; the historical dataset includes at least two historical traffic data of the target service and an evaluation result corresponding to each historical traffic data, and the evaluation result is used to characterize whether the target service is an abnormal service; The training samples for the anomaly identification model are determined based on the at least two historical traffic data, and the labels of the training samples are determined based on the evaluation results of the at least two historical traffic data. The anomaly detection model is trained based on the training samples, the labels of the training samples, and the training features; An initial negative sample set is determined based on the historical dataset; the historical traffic data in the initial negative sample set is different from the historical traffic data in the training samples. From the initial negative sample set, manually reviewed samples are determined, and precise negative samples from the manually reviewed samples whose manually reviewed results are negative are obtained; The remaining samples in the initial negative sample set are verified based on the early warning model to obtain the precise negative samples in the remaining samples; the early warning model is used to identify network risk services based on network traffic data; The negative sample set is determined based on the exact negative samples in the manually reviewed samples and the exact negative samples in the remaining samples; The evaluation sample set is input into the anomaly detection model, and the performance of the anomaly detection model is evaluated based on its output.

2. The method according to claim 1, characterized in that, Training the anomaly detection model based on the training samples, the labels of the training samples, and the training features includes: The initial network model is determined based on the training features; The training samples are input into the initial network model for feature extraction to obtain the training features of the training samples. Based on the training features of the training samples, the prediction results of the initial network model for the training samples are obtained. The anomaly detection model is obtained by training the initial network model based on the loss between the prediction result corresponding to each training sample and the label of the training sample.

3. The method according to claim 1 or 2, characterized in that, The step of determining the training features of the anomaly detection model based on the features matching the target service includes: Display the first configuration item corresponding to the feature; Receive a first operation instruction for the first configuration item, and determine the feature as the training feature based on the first operation instruction.

4. The method according to claim 1 or 2, characterized in that, The method further includes: Displays the second configuration item corresponding to multiple candidate data sources; Receive a second operation instruction for the second configuration item, determine the data source of the historical dataset from the plurality of candidate data sources according to the second operation instruction, and generate the data selection strategy according to the data source of the historical data.

5. The method according to claim 1 or 2, characterized in that, The method further includes: If the model update conditions are met, the anomaly identification model is updated. The model update conditions include at least one of the following: the output of the anomaly identification model differs from the manual judgment result, the anomaly identification rules of the target service are changed, and the online duration of the anomaly identification model exceeds a preset duration; the manual judgment result is the result obtained by manually identifying the traffic data of the target service based on the manual judgment rules.

6. The method according to claim 5, characterized in that, The update of the anomaly detection model includes: Determine the model update rule, which includes at least one of the following: updating training samples and updating training features; The anomaly detection model is updated based on the model update rules.

7. The method according to claim 5, characterized in that, The method further includes: Model update information is generated based on the anomaly detection model and the updated anomaly detection model; the model update information is used to indicate the iterative relationship between the anomaly detection model and the updated anomaly detection model; Create a browsing portal for the model update information.

8. The method according to claim 1 or 2, characterized in that, The method further includes: Display the third configuration item corresponding to the feature; Receive a third operation instruction for the third configuration item, and determine the target feature from the features that match the target service according to the third operation instruction; Manual judgment rules are generated based on the target features and the execution logic of the target features; the manual judgment rules are used to manually identify whether traffic data is abnormal; Create a browsing entry point for the manual judgment rules; the browsing entry point is used by the users of the manual judgment rules to obtain the manual judgment rules.

9. The method according to claim 1 or 2, characterized in that, The step of validating the remaining samples in the initial negative sample set based on the early warning model to obtain the precise negative samples in the remaining samples includes: The remaining samples are input into the early warning model. If the output of the early warning model indicates that the remaining samples are risky samples, then the remaining samples are determined to be the exact negative samples.

10. An anomaly identification method, characterized in that, include: Obtain traffic data of the target service, input the traffic data into the anomaly detection model of the target service, and identify whether the target service is an abnormal service based on the output of the anomaly detection model; The anomaly recognition model is trained using the model training method described in any one of claims 1 to 9.

11. A model training device, characterized in that, include: The feature determination unit is used to acquire the identifier of the target service, determine the features that match the target service based on the identifier of the target service, and determine the training features of the anomaly recognition model based on the features that match the target service. The feature is used to identify whether the target service is an abnormal service, and the abnormality identification model is used to identify whether the target service is an abnormal service based on the traffic data of the target service; The acquisition unit is used to acquire the historical dataset of the target service according to the data selection strategy; the historical dataset includes at least two historical traffic data of the target service and an evaluation result corresponding to each historical traffic data, the evaluation result being used to characterize whether the target service is an abnormal service; A sample determination unit is used to determine training samples for the anomaly identification model based on the at least two historical traffic data, and to determine the labels of the training samples based on the evaluation results of the at least two historical traffic data. The training unit is used to train the anomaly recognition model based on the training samples, the labels of the training samples, and the training features, and to determine an initial negative sample set based on the historical dataset; the historical traffic data in the initial negative sample set is different from the historical traffic data in the training samples. From the initial negative sample set, manually reviewed samples are determined, and precise negative samples with negative manually reviewed results are obtained from the manually reviewed samples. The remaining samples in the initial negative sample set are verified based on the early warning model to obtain precise negative samples from the remaining samples. The early warning model is used to identify network risk services based on network traffic data. The negative sample set is determined based on the precise negative samples in the manually reviewed samples and the precise negative samples in the remaining samples. The evaluation sample set is input into the anomaly identification model, and the performance of the anomaly identification model is evaluated based on the output of the anomaly identification model.

12. The apparatus according to claim 11, characterized in that, The feature determination unit is specifically used to display the first configuration item corresponding to the feature; Receive a first operation instruction for a first configuration item, and determine the feature corresponding to the first configuration item as a training feature based on the first operation instruction.

13. The apparatus according to claim 11 or 12, characterized in that, The acquisition unit is specifically used to display the second configuration items corresponding to multiple candidate data sources; Receive a second operation instruction for the second configuration item, determine the data source of the historical dataset from multiple candidate data sources according to the second operation instruction, and generate a data selection strategy based on the data source of the historical data.

14. The apparatus according to claim 11 or 12, characterized in that, The training unit is also used to update the anomaly recognition model if the model update conditions are met. The model update conditions include at least one of the following: the output of the anomaly identification model differs from the manual judgment result, the anomaly identification rules of the target business are changed, and the online duration of the anomaly identification model exceeds the preset duration; the manual judgment result is the result obtained by manually identifying the traffic data of the target business based on the manual judgment rules.

15. The apparatus according to claim 14, characterized in that, The training unit is specifically used to determine the model update rules, which include at least one of the following: updating the training sample set and updating the training features; The anomaly detection model is updated based on model update rules.

16. The apparatus according to claim 15, characterized in that, The training unit is specifically used to generate model update information based on the anomaly detection model and the updated anomaly detection model; the model update information is used to indicate the iterative relationship between the anomaly detection model and the updated anomaly detection model. Create an entry point for browsing model update information.

17. The apparatus according to claim 11 or 12, characterized in that, The feature determination unit is also used to display the third configuration item corresponding to the feature; Receive a third operation instruction for the third configuration item, and determine the target feature from the features that match the target service according to the third operation instruction; Manual judgment rules are generated based on the target characteristics and the execution logic of the target characteristics; these rules are used to manually identify whether traffic data is abnormal.

18. The apparatus according to claim 11 or 12, characterized in that, The feature determination unit is also used to create a browsing entry for manual judgment rules; the browsing entry is used by the users of the manual judgment rules to obtain the manual judgment rules.

19. The apparatus according to claim 11 or 12, characterized in that, The training unit is specifically used to input the remaining samples into the early warning model. If the output of the early warning model indicates that the remaining samples are risky samples, then the remaining samples are determined to be accurate negative samples.

20. An anomaly detection device, characterized in that, include: The acquisition unit is used to acquire traffic data for the target business. The identification unit is used to input the traffic data into the anomaly identification model of the target service, and identify whether the target service is an abnormal service based on the output of the anomaly identification model; The anomaly recognition model is trained using the model training method described in any one of claims 1 to 9.

21. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1-10.

22. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-10.

23. A computer program product, the computer program product comprising instructions, characterized in that, The instructions are executed by the processor to implement the method as described in any one of claims 1-10.

Citation Information

Patent Citations

  • Traffic segmentation recognition method and system, and electronic device and storage medium

    WO2022078042A1