A risk identification method and related device

CN122548730APending Publication Date: 2026-08-11GUANGZHOU TENCENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]在对当前技术的研究和实践过程中,本申请的发明人发现实时风险识别策略在风险识别的过程中引入了人工专家经验对具有共性的样本数据进行分析和总结,从而得到用于风险识别的风险识别策略,随着对抗的深入,安全的画像建设可能得到成百上千维,人工在这样量级的特征下进行有效性分析和组合是非常困难的,因此,导致风险识别的效率和准确性较低

Benefits of technology

[0038]此外,本申请实施例还提供一种计算机程序产品,包括计算机程序或指令,该计算机程序或指令被处理器执行时实现本申请实施例提供的风险识别方法中的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548730A_ABST
    Figure CN122548730A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a risk identification method and related equipment, which can include a risk identification device and an electronic device. After obtaining object behavior data of a sample object and constructing at least one behavior sequence, the embodiments of the present application collect feature data of the sample object to obtain a training data set, train a preset tree model using the training data set, traverse nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence, filter a candidate risk identification strategy from the initial risk identification strategy, test the candidate risk identification strategy at least once according to historical traffic data and current to-be-identified data, determine a target risk identification strategy from the candidate risk identification strategy based on a test result, and perform risk identification using the target risk identification strategy. The present scheme can improve the efficiency and accuracy of risk identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of risk identification, and more specifically to a risk identification method and related equipment, which may include risk identification devices and electronic equipment. Background Technology

[0002] In recent years, with the rapid development of internet technology, cybersecurity has become increasingly important. In some object-oriented interaction platforms, there are often risky entities engaging in risky behaviors. To improve cybersecurity, it is necessary to identify these risky entities. Current risk identification methods often rely on real-time risk identification strategies.

[0003] In the process of researching and practicing current technologies, the inventors of this application discovered that the real-time risk identification strategy introduces human expert experience to analyze and summarize sample data with commonalities in the risk identification process, thereby obtaining a risk identification strategy for risk identification. As the confrontation deepens, the security profile may become hundreds or thousands of dimensions. It is very difficult for humans to perform effective analysis and combination under such a scale of features, thus resulting in low efficiency and accuracy of risk identification. Summary of the Invention

[0004] This application provides a risk identification method and related equipment, which may include a risk identification device and electronic equipment, and can improve the efficiency and accuracy of risk identification.

[0005] A risk identification method, comprising:

[0006] Obtain object behavior data of a sample object, and construct at least one behavior sequence based on the object behavior data, the behavior sequence indicating the sample object having at least one object behavior;

[0007] Collect at least one feature data of the sample object under the object behavior to obtain the training dataset corresponding to the behavior sequence;

[0008] The training dataset is used to train at least one preset tree model, and the nodes in the trained tree model are traversed to obtain at least one initial risk identification strategy corresponding to the behavior sequence.

[0009] Based on the training dataset, at least one candidate risk identification strategy is selected from the initial risk identification strategy.

[0010] The candidate risk identification strategy is tested at least once based on historical traffic data and current data to be identified;

[0011] Based on the test results, the target risk identification strategy corresponding to the behavior sequence is determined from the candidate risk identification strategies, and the target risk identification strategy is used for risk identification.

[0012] Accordingly, embodiments of this application provide a risk identification device, including:

[0013] An acquisition unit is configured to acquire object behavior data of a sample object and, based on the object behavior data, construct at least one behavior sequence, wherein the behavior sequence indicates the sample object having at least one object behavior;

[0014] The acquisition unit is used to acquire at least one feature data of the sample object under the object behavior to obtain the training dataset corresponding to the behavior sequence;

[0015] The training unit is used to train at least one preset tree model using the training dataset and to traverse the nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence.

[0016] A filtering unit is used to filter out at least one candidate risk identification strategy from the initial risk identification strategy based on the training dataset.

[0017] The testing unit is used to test the candidate risk identification strategy at least once based on historical traffic data and current data to be identified;

[0018] The identification unit is used to determine the target risk identification strategy corresponding to the behavior sequence from the candidate risk identification strategies based on the test results, and to use the target risk identification strategy to perform risk identification.

[0019] In some embodiments, the acquisition unit may be specifically used to identify the object behavior of the sample object within a preset historical time range in the object behavior data; construct at least one candidate behavior sequence corresponding to the sample object based on the object behavior, the candidate behavior sequence including at least one of the object behaviors; based on the object behaviors in the candidate behavior sequence, count the behavior frequency of the candidate behavior sequence, and filter at least one behavior sequence from the candidate behavior sequence based on the behavior frequency.

[0020] In some embodiments, the acquisition unit may be specifically configured to: construct an interaction behavior sequence corresponding to the interaction behavior when the object behavior includes at least one of the interaction behaviors, and use the interaction behavior sequence as the candidate behavior sequence; construct a single-point behavior sequence corresponding to the single-point behavior when the object behavior includes at least one of the single-point behaviors, and use the single-point behavior sequence as the candidate behavior sequence; and construct an interaction behavior sequence corresponding to the interaction behavior and a single-point behavior sequence corresponding to the single-point behavior when the object behavior includes at least one of the interaction behaviors and at least one single-point behavior, and add the single-point behavior sequence to the interaction behavior sequence respectively to obtain at least one candidate behavior sequence.

[0021] In some embodiments, the testing unit may be specifically used to evaluate the candidate risk identification strategies using historical traffic data, so as to select at least one risk identification strategy to be tested from the candidate risk identification strategies; to obtain the current data to be identified in the current time range, and to identify the current object data of at least one object to be identified in the current data to be identified; and to use the risk identification strategy to be tested to identify the object to be identified, so as to obtain the test result of the risk identification strategy to be tested.

[0022] In some embodiments, the testing unit may be specifically used to perform risk identification on objects in the historical traffic data using the candidate risk identification strategy to obtain a first identification result; obtain risk object information in the historical traffic data, and sort the candidate risk identification strategies based on the risk object information and the first identification result; and identify at least one risk identification strategy to be tested from the candidate risk identification strategies according to the sorting result.

[0023] In some embodiments, the testing unit may be specifically used to perform risk identification on the object to be identified based on the current object data and the risk identification strategy to be tested, to obtain a second identification result; send the second identification result and the current object data to a verification server, so that the verification server verifies the second identification result based on the current object data; receive the verification result of the second identification result returned by the verification server, and use the verification result as the test result of the risk identification strategy to be tested.

[0024] In some embodiments, the testing unit may be specifically used to: count the number of correctly identified second identification results in the second identification results based on the test results to obtain the target identification number; determine the identification accuracy of the risk identification strategy to be tested within the current time range based on the target identification number and the number of objects to be identified; and select at least one risk identification strategy to be verified from the risk identification strategies to be verified to obtain the target risk identification strategy corresponding to the behavior sequence.

[0025] In some embodiments, the identification unit may be specifically used to detect the target risk identification strategy when the time for risk identification using the target risk identification strategy exceeds a preset time threshold; based on the detection result, determine the current identification accuracy of the target risk identification strategy, and calculate the difference between the current identification accuracy and the historical identification accuracy of the target risk identification strategy; when the difference is greater than or equal to a preset difference threshold, stop using the target risk identification strategy for risk identification.

[0026] In some embodiments, the acquisition unit may be specifically used to: bucket the sample objects according to the object behaviors in the behavior sequence to obtain at least one target sample object corresponding to the behavior sequence; collect at least one feature data of the target sample object when performing the object behavior to obtain the initial feature data of the target sample object; and filter out at least one target feature data from the initial feature data to obtain the training dataset corresponding to the behavior sequence.

[0027] In some embodiments, the acquisition unit may be specifically used to acquire at least one object data of the target sample object when it performs the object behavior, to obtain the original object data of the target sample object; to sample original object data less than or equal to a preset data quantity threshold from the original object data, to obtain the target object data of the target sample object; and to perform feature parsing on the target object data to obtain the initial feature data of the target sample object.

[0028] In some embodiments, the acquisition unit can be specifically used to: when the first feature type does not include a preset first feature type, generate current feature data corresponding to the preset first feature type based on the object attribute information of the target sample object, and add the current feature data to the target object data to obtain the initial feature data of the target sample object; when the first feature type includes a feature type to be converted, convert the object data of the feature type to be converted into feature data of a preset second feature type, and update the target object data based on the converted feature data to obtain the initial feature data of the target sample object; when the first feature type includes a feature type to be encoded, perform feature encoding on the object data corresponding to the feature type to be encoded, and update the target object data based on the encoded feature data to obtain the initial feature data of the target sample object.

[0029] In some embodiments, the acquisition unit may be specifically used to filter out feature data of at least one continuous feature from the initial feature data to obtain continuous feature data, and to bin the continuous feature data to obtain target continuous feature data; to update the initial feature data based on the target continuous feature data to obtain updated feature data, wherein the updated feature data includes feature data of at least one second feature type, the second feature type including the continuous feature; and to filter out target feature data of at least one target feature type from the updated feature data to obtain the training dataset corresponding to the behavior sequence.

[0030] In some embodiments, the acquisition unit may be specifically used to discretize the continuous feature data based on at least one preset data distribution, and evaluate the discrete confidence level corresponding to each preset data distribution; according to the discrete confidence level, select a target data distribution from the preset data distribution, and bin the discrete feature data corresponding to the target data distribution to obtain at least one first feature data bin corresponding to each continuous feature; based on the feature data in the first feature data bin, determine a first information value parameter of the continuous feature, the first information value parameter being used to measure the contribution of the feature data of the continuous feature to risk identification; and according to the first information value parameter, select target continuous feature data corresponding to the target continuous feature from the continuous feature data.

[0031] In some embodiments, the acquisition unit may be specifically used to calculate the object ratio of sample objects corresponding to feature data in the first feature data box, the object ratio including the risk object ratio and the normal object ratio; calculate the ratio between the risk object ratio and the normal object ratio to obtain the good-to-bad ratio of the first feature data box; calculate the ratio difference between the risk object ratio and the normal object ratio, and determine the initial information value parameter of the first feature data box based on the ratio difference and the good-to-bad ratio; and fuse the initial information value parameters of the first feature data box corresponding to the continuous feature to obtain the first information value parameter of the continuous feature.

[0032] In some embodiments, the acquisition unit may be specifically used to bin the feature data corresponding to each second feature type in the updated feature data to obtain at least one second feature data bin corresponding to the second feature type; based on the feature data in the second feature data bin, determine a second information value parameter for the second feature type, the second information value parameter being used to measure the contribution of the feature data of the second feature type to risk identification; according to the second information value parameter, select target feature data of at least one target feature type from the updated feature data, and filter the target feature data to obtain the training dataset corresponding to the behavior sequence.

[0033] In some embodiments, the training unit may be specifically used to identify at least one node in the tree model, the node including a root node and leaf nodes, the node indicating the feature type corresponding to the feature data in the training dataset; obtain the execution path of the tree model, and according to the execution path, traverse from the root node to the leaf node to obtain at least one node path; based on the node path, determine an initial risk identification strategy corresponding to the behavior sequence, the initial risk identification strategy representing a combination of feature data of at least one feature type used for risk identification.

[0034] In some embodiments, the filtering unit may be specifically used to perform risk identification on the sample objects based on the training dataset using the initial risk identification strategy, to obtain a third identification result, wherein the third identification result includes at least one candidate risk sample object indicated by the leaf node corresponding to the initial risk identification strategy; obtain the labeled risk category of the candidate risk sample object, and calculate at least one identification performance parameter of the initial risk identification strategy based on the labeled risk category; and filter out at least one candidate risk identification strategy from the initial risk identification strategy whose identification performance parameter is greater than a preset parameter threshold.

[0035] In some embodiments, the acquisition unit may be specifically used to acquire negative feedback information for at least one object in the object interaction platform, and based on the negative feedback information, to filter out at least one risk object from the objects; to collect behavioral data of the risk object within a preset historical time range to obtain risk behavior data; to filter out at least one normal object from the normal object set of the object interaction platform, and to collect behavioral data of the normal object within the preset historical time range to obtain normal behavior data; and to use the risk object and the normal object as sample objects, and to use the risk behavior data and the normal behavior data as object behavior data of the sample objects.

[0036] Furthermore, this application also provides an electronic device, including a processor and a memory, wherein the memory stores an application program, and the processor is used to run the application program in the memory to execute the risk identification method provided in this application.

[0037] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the risk identification methods provided in embodiments of this application.

[0038] Furthermore, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in the risk identification method provided in embodiments of this application.

[0039] In this embodiment, after acquiring object behavior data of a sample object and constructing at least one behavior sequence based on the object behavior data, at least one feature data of the sample object under the object behavior in the behavior sequence is collected to obtain a training dataset corresponding to the behavior sequence. Then, at least one preset tree model is trained using the training dataset, and the nodes in the trained tree model are traversed to obtain at least one initial risk identification strategy corresponding to the behavior sequence. Then, based on the training dataset, at least one candidate risk identification strategy is selected from the initial risk identification strategies. The candidate risk identification strategies are tested at least once based on historical traffic data and current data to be identified. Then, based on the test results, the target risk identification strategy corresponding to the behavior sequence is determined from the candidate risk identification strategies, and then... The target risk identification strategy is used for risk identification. This solution trains the data model using training data of specific behavioral sequences, automatically generating an initial risk identification strategy for each sequence. Then, through the training dataset, historical traffic data, and current data to be identified, the initial risk identification strategy undergoes multiple stages of quality evaluation. This effectively improves the efficiency and accuracy of risk identification strategy discovery from a large number of features. Furthermore, the generated target risk identification strategy for each behavioral sequence is primarily used to identify risks in objects with specific object behaviors (object behaviors within the behavioral sequence), without needing to identify risks in all objects, thus reducing the total number of objects to be identified and improving the efficiency of risk identification. Therefore, it can improve the efficiency and accuracy of risk identification. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a schematic diagram of a scenario for the risk identification method provided in the embodiments of this application;

[0042] Figure 2 This is a flowchart illustrating the risk identification method provided in an embodiment of this application;

[0043] Figure 3 This is a schematic diagram of the overall framework of the automatic generation target risk identification strategy provided in the embodiments of this application;

[0044] Figure 4 This is another schematic diagram of the risk identification method provided in the embodiments of this application;

[0045] Figure 5This is a schematic diagram of the risk identification device provided in the embodiments of this application;

[0046] Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0047] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] This application provides a risk identification method and related equipment, which may include a risk identification device and an electronic device. The risk identification device may be integrated into the electronic device, which may be a server or a terminal, etc.

[0049] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0050] For example, see Figure 1Taking the integration of a risk identification device into an electronic device as an example, after acquiring object behavior data of a sample object and constructing at least one behavior sequence based on the object behavior data, the electronic device collects at least one feature data of the sample object under the object behavior in the behavior sequence to obtain a training dataset corresponding to the behavior sequence. Then, the training dataset is used to train at least one preset tree model, and the nodes in the trained tree model are traversed to obtain at least one initial risk identification strategy corresponding to the behavior sequence. Then, based on the training dataset, at least one candidate risk identification strategy is selected from the initial risk identification strategies. Based on historical traffic data and current data to be identified, the candidate risk identification strategies are tested at least once. Then, based on the test results, the target risk identification strategy corresponding to the behavior sequence is determined from the candidate risk identification strategies, and the target risk identification strategy is used for risk identification, thereby improving the efficiency and accuracy of risk identification.

[0051] It is understood that, in the specific implementation of this application, the object behavior data, object data, historical traffic data, or current data to be identified of the sample object / object to be identified are involved. When the following embodiments of this application are applied to specific products or technologies, permission or consent is required, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0052] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0053] This embodiment will be described from the perspective of a risk identification device, which can be integrated into an electronic device, such as a server or a terminal. The terminal can include tablet computers, laptops, personal computers (PCs), wearable devices, virtual reality devices, or other smart devices capable of risk identification.

[0054] A risk identification method, comprising:

[0055] Obtain object behavior data of sample objects, and construct at least one behavior sequence based on the object behavior data. The behavior sequence indicates a sample object with at least one object behavior. Collect at least one feature data of the sample object under the object behavior to obtain a training dataset corresponding to the behavior sequence. Use the training dataset to train at least one preset tree model, and traverse the nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence. Based on the training dataset, select at least one candidate risk identification strategy from the initial risk identification strategies. Test the candidate risk identification strategies at least once based on historical traffic data and current data to be identified. Based on the test results, determine the target risk identification strategy corresponding to the behavior sequence from the candidate risk identification strategies, and use the target risk identification strategy to perform risk identification.

[0056] like Figure 2 As shown, the specific process of this risk identification method is as follows:

[0057] 101. Obtain object behavior data of the sample object, and construct at least one behavior sequence based on the object behavior data.

[0058] In this context, object behavior data can be understood as the data on the object behaviors generated by sample objects within the object interaction platform. A sample object can be understood as the object that generates the risk identification strategy. Sample objects can include risky objects and normal objects. A risky object can be understood as an object on the object interaction platform that poses a risk, while a normal object can be understood as an object on the object interaction platform that does not pose a risk. Object behavior can include interactive behavior and single-point behavior. Interactive behavior can be understood as the behavior that occurs between multiple objects, such as adding friends, deleting, blocking, commenting, liking, sending messages, or other behaviors. For example, object A sends a document to object B, or object A likes, comments, or performs other interactive behaviors on content published by object B on the object interaction platform, or object A... Single-point interaction can be the behavior generated by an object itself within the object interaction platform, such as registering, logging in, posting on social media, etc. The object interaction platform can be understood as a platform used for object interaction. Object interaction platforms can be of various types, such as instant messaging platforms, content publishing platforms, content sharing platforms, content playback platforms, or other platforms where object interaction is possible, etc.

[0059] Here, a behavior sequence indicates a sample object that has at least one object behavior. A behavior sequence can be understood as a sequence of one or more object behaviors. A behavior sequence can include one object behavior or multiple object behaviors. For example, a behavior sequence can include the behavior of object A registering an account, and it can also include the behavior of object A requesting to add object B as a friend.

[0060] There are several ways to obtain object behavior data of sample objects, including the following:

[0061] For example, negative feedback information for at least one object can be obtained from the object interaction platform. Based on the negative feedback information, at least one risk object can be selected from the objects. Behavioral data of the risk object within a preset historical time range can be collected to obtain risk behavior data. At least one normal object can be selected from the normal object set of the object interaction platform. Behavioral data of the normal object within a preset historical time range can be collected to obtain normal behavior data. The risk object and the normal object are used as sample objects, and the risk behavior data and the normal behavior data are used as the object behavior data of the sample objects.

[0062] Negative feedback information can be understood as information generated by the object interaction platform when it performs negative feedback actions on an object. There are various types of negative feedback actions, such as reporting, deleting, blocking, or other feedback actions indicating that the object poses a risk. Based on negative feedback information, there are several ways to filter out at least one risky object from the pool of objects. For example, negative feedback information can be sent to an audit server so that the audit server can audit the object based on the negative feedback information, receive the audit results returned by the audit server, and filter out at least one risky object based on the audit results.

[0063] After identifying at least one risky object from the pool of objects, behavioral data of that risky object within a preset historical timeframe can be collected to obtain risk behavior data. There are several ways to collect this data; for example, one can obtain a set of historical behavioral data within the preset timeframe from the object interaction platform, and then filter out the historical behavioral data corresponding to the risky object from this set to obtain the risk behavior data.

[0064] There are several ways to select at least one normal object from the object collection of the object interaction platform. For example, one can obtain the normal object collection of the object interaction platform and randomly select at least one normal object from the normal object collection.

[0065] After selecting at least one normal object from the set of normal objects, historical behavioral data of the normal object within a preset time range can be collected to obtain normal behavioral data. The method for collecting historical behavioral data of normal objects within a preset time range is similar to the method for collecting behavioral data of risky objects within a preset time range, as detailed above, and will not be repeated here.

[0066] After identifying risky and normal objects and collecting risky and normal behavior data, the risky and normal objects can be used as sample objects, and the risky and normal behavior data can be used as the object behavior data of the sample objects.

[0067] After obtaining the object behavior data of the sample object, at least one behavior sequence can be constructed based on the object behavior data. There are several ways to construct the behavior sequence. For example, the object behavior of the sample object within a preset historical time range can be identified in the object behavior data. Based on the object behavior, at least one candidate behavior sequence corresponding to the sample object can be constructed. The candidate behavior sequence includes at least one object behavior. Based on the object behavior in the candidate behavior sequence, the behavior frequency of the candidate behavior sequence is counted, and at least one behavior sequence is selected from the candidate behavior sequences based on the behavior frequency.

[0068] Object behaviors include interactive behaviors and single-point behaviors. Single-point behaviors represent behaviors generated by the sample object itself. There are various types of single-point behaviors, such as registration, login, content posting, or other self-generated behaviors. Based on object behaviors, there are several ways to construct at least one candidate behavior sequence corresponding to the sample object. For example, when the object behaviors include at least one interactive behavior, construct the interactive behavior sequence corresponding to the interactive behavior and use it as a candidate behavior sequence; when the object behaviors include at least one single-point behavior, construct the single-point behavior sequence corresponding to the single-point behavior and use it as a candidate behavior sequence; when the object behaviors include at least one interactive behavior and at least one single-point behavior, construct the interactive behavior sequence corresponding to the interactive behavior and the single-point behavior sequence corresponding to the single-point behavior sequence, and add the single-point behavior sequence to the interactive behavior sequence to obtain at least one candidate behavior sequence.

[0069] Taking object A as an example, if object A interacts with three objects (x (add friend), y (delete), and z (block)) within a preset historical time range, then three interaction behavior sequences (Ax (add friend), Ay (delete), and Az (block)) can be constructed, and the interaction behavior sequences are strongly correlated with the interacting objects. If object A does not interact with other objects within the preset historical time range, but only logs into the object interaction platform, then object A's behavior can be considered a single-point interaction behavior (login), and the single-point behavior sequence (a (login)) can be constructed. If object A first logs in within the preset historical time range, and then interacts with objects x, y, and z respectively, then the constructed single-point behavior sequences can be added to the interaction behavior sequences and sorted according to time order, becoming (a (login) + Ax (add friend)), (a (login) + Ay (add / delete)), and (a (login) + Az (block)), and these behavior sequences are used as candidate behavior sequences.

[0070] After constructing at least one candidate behavior sequence corresponding to the sample object, the behavior frequency of the candidate behavior sequence can be counted based on the object behavior in the candidate behavior sequence. Behavior frequency can be understood as the number of times the candidate behavior sequence appears within a preset historical time range.

[0071] After calculating the frequency of behaviors in the candidate behavior sequences, at least one behavior sequence can be selected from the candidate behavior sequences based on the behavior frequency. There are several ways to select at least one behavior sequence from the candidate behavior sequences based on the behavior frequency. For example, the candidate behavior sequences can be sorted according to the behavior frequency, and the top N candidate behavior sequences can be selected based on the sorting results, thus obtaining at least one behavior sequence.

[0072] In this context, a behavior sequence can be understood as a high-frequency behavior sequence selected from candidate behavior sequences. It should also be noted that a behavior sequence can include multiple sub-behavior sequences. For example, if the high-frequency behavior sequence is (login + file transfer), then any candidate behavior sequence containing (login + file transfer) can be considered as this behavior sequence.

[0073] Among them, normal objects and risk objects in the sample objects are actually very different in type, but groups with the same behavioral sequence (behavioral pattern) are more likely to be a group, and it is also possible to better extract the actual common strategies, thereby improving the accuracy of risk identification.

[0074] 102. Collect at least one feature data of the sample object under the object behavior to obtain the training dataset corresponding to the behavior sequence.

[0075] Feature data can be understood as data in at least one feature dimension under the object's behavior. There are several ways to collect at least one feature data of a sample object under its behavior, including the following:

[0076] For example, sample objects can be bucketed according to the object behavior in the behavior sequence to obtain at least one target sample object corresponding to the behavior sequence. At least one feature data of the target sample object under the execution of the object behavior can be collected to obtain the initial feature data of the target sample object. At least one target feature data can be selected from the initial feature data to obtain the training dataset corresponding to the behavior sequence.

[0077] There are several ways to bucket sample objects based on the object behavior in the behavior sequence. For example, you can identify the object identifier of at least one sample object with the object behavior in the behavior sequence, and use the sample object corresponding to the object identifier as a bucket to obtain at least one target sample object corresponding to the behavior sequence.

[0078] It should be noted that, since a sample object can have at least one object behavior, a sample object can correspond to multiple behavior sequences or a single behavior sequence.

[0079] After binning the sample objects according to the object behaviors in the behavior sequence, at least one feature data of the target sample object under the execution of the object behavior can be collected, thereby obtaining the initial feature data of the target sample object. There are several ways to collect at least one feature data of the target sample object under the execution of the object behavior. For example, the object data of the target sample object under the execution of the object behavior can be collected to obtain the original object data of the target sample object. From the original object data, a portion of the original object data less than or equal to a preset data quantity threshold can be sampled to obtain the target object data of the target sample object. Feature parsing can then be performed on the target object data to obtain the initial feature data of the target sample object.

[0080] There are several ways to collect at least one object data of the target sample object under the execution of object behavior. For example, at least one data snapshot of the target sample object under the execution of object behavior can be collected and used as object data. For instance, taking login as an example, at least one dimension of data snapshots under the login behavior can be collected. These data snapshots may include login time, login device, or login method, etc.; taking sending an image as an example, at least one dimension of data snapshots under sending an image can be collected. These data snapshots may include sent image, sending time, sending object, receiving object, sending device, image attribute information, or other dimension of data snapshots, etc.

[0081] After collecting at least one object data point of the target sample object under the execution of object behavior, the target object data of the target sample object can be obtained by sampling original object data that is less than or equal to a preset data quantity threshold. There are several ways to sample original object data that is less than or equal to the preset data quantity threshold from the original object data. For example, original object data that is less than or equal to the preset data quantity threshold can be randomly sampled from the original object data to obtain the target object data of the target sample object.

[0082] It should be noted that the purpose of downsampling the original object data is to prevent some target sample objects from having a large amount of homogeneous original object data, which would pollute the credibility of the training dataset. Therefore, during the data downsampling process, a hyperparameter n (a preset data quantity threshold) can be limited for each target sample object to ensure that each target sample object can retain at most n original object data, and these n original object data are randomly collected.

[0083] After sampling a subset of original object data less than or equal to a preset data quantity threshold from the original object data, feature parsing can be performed on the sampled target object data to obtain the initial feature data of the target sample object. The target object data includes object data of at least one first feature type. There are several ways to perform feature parsing on the target object data. For example, when the first feature type does not include a preset first feature type, current feature data corresponding to the preset first feature type is generated based on the object attribute information of the target sample object, and the current feature data is added to the target object data to obtain the initial feature data of the target sample object. When the first feature type includes a feature type to be converted, the object data of the feature type to be converted is converted into feature data of a preset second feature type, and the target object data is updated based on the converted feature data to obtain the initial feature data of the target sample object. When the first feature type includes a feature type to be encoded, the object data corresponding to the feature type to be encoded is feature-encoded, and the target object data is updated based on the encoded feature data to obtain the initial feature data of the target sample object.

[0084] When the first feature type does not include a preset first feature type, it indicates that data for certain feature types of the target sample object has not been collected. In this case, current feature data corresponding to the preset first feature type can be generated based on the object attribute information of the target sample object. There are several ways to generate the current feature data corresponding to the first feature type. For example, one can obtain the object attribute information of the target sample object when performing object behavior, obtain the feature data mapping table corresponding to the preset first feature type, and map the object data information according to the feature data mapping table to obtain the current feature data corresponding to the preset first feature type. After generating the current feature data corresponding to the preset first feature type, the current feature data can be added to the target object data to obtain the initial feature data of the target sample object.

[0085] When the first feature type contains the feature type to be converted, it indicates that feature conversion is needed for the unique data corresponding to the feature type to be converted. The feature type to be converted can be various, such as floating-point, string, or other feature types that need to be converted. After feature conversion of the object data corresponding to the feature type to be converted, the target object data can be updated based on the converted feature data to obtain the initial feature data of the target sample object. There are several ways to update the target object data; for example, the object data corresponding to the feature type to be converted can be replaced with the converted feature data to obtain the initial feature data of the target sample object.

[0086] When the first feature type includes the feature type to be encoded, feature encoding can be performed on the object data corresponding to the feature to be encoded. There are several ways to perform feature encoding on the object data corresponding to the feature to be encoded. For example, the encoding method for the object data corresponding to the feature type to be encoded can be determined based on the data type. Based on the encoding method, feature encoding is performed on the object data corresponding to the feature type to be encoded, thus obtaining the encoded feature data. For instance, taking a data type of country or province as an example, the encoding method in this case could be one-hot encoding (a data encoding method). It should also be noted that different data types can correspond to different encoding methods.

[0087] After feature encoding is performed on the object data corresponding to the feature type to be encoded, the target object data can be updated based on the encoded feature data to obtain the initial feature data of the target sample object. The method of updating the target object data based on the encoded feature data is similar to the method of updating the target object data based on the transformed feature data, as described above, and will not be repeated here.

[0088] After performing feature parsing on the target object data, at least one target feature data can be selected from the parsed initial feature data to obtain the training dataset corresponding to the behavior sequence. There are several ways to select at least one target feature data from the initial feature data. For example, at least one continuous feature data can be selected from the initial feature data to obtain continuous feature data. This continuous feature data can then be binned to obtain target continuous feature data. The initial feature data can then be updated based on the target continuous feature data to obtain updated feature data. This updated feature data includes at least one feature data of a second feature type, which includes continuous features. Finally, at least one target feature data of the target feature type can be selected from the updated feature data to obtain the training dataset corresponding to the behavior sequence.

[0089] Continuous features can be understood as feature data that is continuous. For example, this could include age; 17 and 18 are two relatively close ages, but for the model input, they are two completely different inputs. There are various ways to bin continuous feature data. For instance, based on at least one preset data distribution, the continuous feature data can be discretized, and the discrete confidence level corresponding to each preset data distribution can be evaluated. Based on the discrete confidence level, a target data distribution can be selected from the preset data distributions, and the discrete feature data corresponding to the target data distribution can be binned to obtain at least one first feature data bin for each continuous feature. Based on the feature data in the first feature data bin, a first information value parameter for the continuous feature can be determined. Based on the first information value parameter, target continuous feature data corresponding to the target continuous feature can be selected from the continuous feature data.

[0090] Here, the preset data distribution can be understood as at least one pre-defined data distribution. This data distribution can be of various types, such as chi-square distributions or other data distributions, etc.

[0091] The first information value parameter can be understood as a measure of the contribution of continuous feature data to risk identification, and can be regarded as an IV (Information Value). There are several ways to determine the first information value parameter of continuous features based on the feature data in the first feature data bin. For example, the proportion of sample objects corresponding to the feature data in the first feature data bin can be calculated. This proportion can include the proportion of risky objects and the proportion of normal objects. The ratio between the proportion of risky objects and the proportion of normal objects is calculated to obtain the good-to-bad ratio of the first feature data bin. The difference between the proportion of risky objects and the proportion of normal objects is calculated, and based on the difference and the good-to-bad ratio, the initial information value parameter of the first feature data bin is determined. The initial information value parameters of the first feature data bins corresponding to continuous features are then fused to obtain the first information value parameter of the continuous features.

[0092] There are several ways to calculate the proportion of sample objects corresponding to the feature data in the first feature data bin. For example, the number of sample objects corresponding to the feature data can be counted in the first feature data bin, the number of risky objects corresponding to the feature data can be counted in the first feature data bin, the number of normal objects corresponding to the feature data can be counted in the first feature data bin, the ratio of the number of risky objects to the total number of objects can be calculated to obtain the risky object proportion (Bad Rate), and the ratio of the number of normal objects to the total number of objects can be calculated to obtain the normal object proportion (Good Rate).

[0093] There are several ways to determine the initial information value parameters of the first feature data box based on the ratio difference and the good-to-bad ratio. For example, based on the good-to-bad ratio, the good-to-bad ratio encoding parameter corresponding to the first feature data box can be determined, and the good-to-bad ratio encoding parameter can be fused with the ratio difference to obtain the initial information value parameters of the first feature data box.

[0094] The good-to-bad ratio coding parameter can be understood as a coding form for the good-to-bad ratio (the original independent variable), which measures the influence of different groups on the target variable. Based on the good-to-bad ratio, there are several ways to determine the good-to-bad ratio coding parameter corresponding to the first feature data bin. For example, the good-to-bad ratio can be calculated by comparing the values ​​of the good-to-bad ratios to obtain the good-to-bad ratio coding parameter, as shown in formula (1). Specifically, it can be as follows:

[0095]

[0096] Wherein, WOE is the good-to-bad ratio encoding parameter, Good Rate is the proportion of normal objects, and Bad Rate is the proportion of risky objects.

[0097] After determining the good-to-bad ratio encoding parameters for the first feature data bin based on the good-to-bad ratio, the good-to-bad ratio encoding parameters can be fused with the ratio difference to obtain the initial information value parameters of the first feature data bin. There are several ways to fuse the good-to-bad ratio encoding parameters with the ratio difference; for example, the good-to-bad ratio encoding parameters can be multiplied by the ratio difference to obtain the initial information value parameters of the first feature data bin.

[0098] After determining the initial information value parameters of the first feature data box based on the ratio difference and the good-to-bad ratio, the initial information value parameters of the first feature data boxes corresponding to continuous features can be fused to obtain the first information value parameters of the continuous features. There are several ways to fuse the initial information value parameters of the first feature data boxes corresponding to continuous features. For example, the initial value information parameters of the first feature data boxes belonging to the same continuous feature can be accumulated to obtain the first information value parameters of the continuous feature, as shown in formula (2). Specifically, it can be as follows:

[0099] IV=∑(Good Rate-Bad Rate)×WOE (2)

[0100] Wherein, IV is the first information value parameter, Good Rate is the proportion of normal objects, Bad Rate is the proportion of risky objects, and WOE is the good-bad ratio coding parameter.

[0101] After determining the first information value parameter of the continuous feature based on the feature data in the first feature data bin, the target continuous feature data corresponding to the target continuous feature can be filtered out from the continuous feature data according to the first information value parameter. There are several ways to filter out the target continuous feature data corresponding to the target continuous feature from the continuous feature data according to the first information value parameter. For example, at least one continuous feature whose first information value parameter exceeds a preset parameter threshold can be filtered out from the continuous features to obtain the target continuous feature. Then, at least one continuous feature data corresponding to the target continuous feature can be filtered out from the continuous feature data to obtain the target continuous feature data.

[0102] After binning the continuous feature data to obtain the target continuous feature data, the initial feature data can be updated based on the target continuous feature data to obtain the updated feature data. There are several ways to update the initial feature data based on the target continuous feature data. For example, the continuous feature data in the initial feature data can be replaced with the target continuous feature data to obtain the updated feature data.

[0103] After updating the initial feature data based on the target continuous feature data, at least one target feature type can be selected from the updated feature data to obtain the training dataset corresponding to the behavior sequence. The updated feature data includes at least one second feature type, which includes continuous features. The first feature type can be the same as or different from the second feature type. There are several ways to select at least one target feature type from the updated feature data. For example, the feature data corresponding to each second feature type in the updated feature data can be binned to obtain at least one second feature data bin for each second feature type. Based on the feature data in the second feature data bins, a second information value parameter for the second feature type is determined. According to the second information value parameter, at least one target feature type can be selected from the updated feature data, and the target feature data can be filtered to obtain the training data corresponding to the behavior sequence.

[0104] There are several ways to bin the feature data corresponding to each second feature type in the updated feature data. For example, the updated feature data can be adjusted based on the second feature type to obtain adjusted feature data, and the adjusted feature data can be binned to obtain at least one second feature data bin corresponding to the second feature type.

[0105] There are several ways to adjust the updated feature data based on the second feature type. For example, the feature data corresponding to the continuous variables in the second feature type can be divided into several intervals, and the feature data corresponding to the categorical variables in the second feature type can be merged to obtain the adjusted feature data.

[0106] After adjusting the updated feature data, it can be binned to obtain at least one second feature data bin corresponding to the second feature type. There are several ways to bin the adjusted feature data. For example, a wideband binning method can be used; an equal-frequency binning method can also be used; or, based on business knowledge, the adjusted feature data can be binned to obtain at least one second feature data bin corresponding to the second feature type, and so on.

[0107] After binning the feature data corresponding to each second feature type in the updated feature data, the second information value parameter of the second feature type can be determined based on the feature data in the binned second feature data bins. The second information value parameter is used to measure the contribution of the feature data of the second feature type to risk identification, that is, the predictive ability of the feature. The higher the second information value parameter, the stronger the predictive ability of the feature and the higher the information contribution. The method for determining the first information value parameter of the second feature type can be similar to the method for determining the first information value parameter of continuous features, as detailed above, and will not be repeated here.

[0108] After determining the second information value parameter of the second feature type, target feature data of at least one target feature type can be selected from the updated feature data based on the second information value parameter. There are several ways to select target feature data of at least one target feature type from the updated candidate feature data based on the second information value parameter. For example, at least one target feature type whose second information value parameter exceeds a preset threshold can be selected from the second feature types, and the feature data corresponding to the target feature type can be selected from the updated feature data to obtain the target feature data.

[0109] After selecting target feature data of at least one target feature type from the updated feature data, the target feature data can be filtered to obtain the training dataset corresponding to the behavior sequence. There are several ways to filter the target feature data. For example, one can obtain the generality parameter corresponding to each target feature type, and filter out feature data with a generality parameter less than a preset threshold from the target feature data to obtain the training dataset corresponding to the behavior sequence. Alternatively, based on expert experience information, one can select at least one preset feature type with weak generality from the target feature types, and filter out the feature data corresponding to the preset feature type from the target feature data to obtain the training dataset corresponding to the behavior sequence, and so on.

[0110] 103. Train at least one preset tree model using a training dataset, and traverse the nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence.

[0111] In this context, a pre-defined tree model can be understood as a pre-defined tree model or tree-like model. There are various types of pre-defined tree models, such as decision tree models, random forest models, XGBoost (a type of tree model), and so on.

[0112] There are several ways to train at least one pre-defined tree model using a training dataset, as follows:

[0113] For example, the training dataset can be fitted and trained by adjusting multiple parameters of the preset tree model, such as tree depth, number of trees, and number of randomly sampled features, until the preset tree model converges, thus obtaining the trained tree model.

[0114] After training at least one pre-defined tree model using a training dataset, the nodes in the trained tree model can be traversed to obtain at least one initial risk identification strategy corresponding to the behavior sequence. The nodes in the trained tree model can include a root node and at least one leaf node. There are several ways to traverse the nodes in the trained tree model. For example, at least one node can be identified in the tree model, which may include a root node and a leaf node. This node indicates the feature type corresponding to the feature data in the training dataset. The execution path of the tree model can be obtained, and based on the execution path, traversal can be performed from the root node to the leaf node to obtain at least one node path. Based on the node path, the initial risk identification strategy corresponding to the behavior sequence can be determined. This initial risk identification strategy represents a combination of feature data of at least one feature type used for risk identification.

[0115] The trained tree model can contain at least one node. Each node indicates the feature type corresponding to the feature data in the training dataset. Nodes can include root nodes and leaf nodes.

[0116] The execution path can be understood as the data path of feature data from input to output in the trained tree model. The node path can be understood as the path from the root node to the leaf node. After traversing the node paths, the initial risk identification strategy corresponding to the behavior sequence can be determined based on the node paths. There are several ways to determine the initial risk identification strategy corresponding to the behavior sequence based on the node paths. For example, at least one candidate feature data corresponding to the node path can be identified in the training dataset, and these candidate feature data can be combined to obtain the initial risk identification strategy corresponding to the behavior sequence. The initial risk identification strategy is essentially a combination of multiple features used for risk identification. During the risk identification process, feature data of the object to be identified is collected, and the collected feature data of the object to be identified is matched with the combination of these features. If the match is successful, the object to be identified can be determined as a risk object; if the match fails, the object to be identified can be determined as a normal object.

[0117] 104. Based on the training dataset, select at least one candidate risk identification strategy from the initial risk identification strategy.

[0118] For example, based on the training dataset, an initial risk identification strategy can be used to identify the risk of sample objects and obtain a third identification result. The third identification result includes at least one candidate risk sample object indicated by the leaf node corresponding to the initial risk identification strategy. The labeled risk category of the candidate risk sample object is obtained, and based on the labeled risk category, at least one identification performance parameter of the initial risk identification strategy is calculated. At least one candidate risk identification strategy with an identification performance parameter greater than a preset parameter threshold is selected from the initial risk identification strategy.

[0119] There are several ways to identify risk in sample objects based on the training dataset and the initial risk identification strategy. For example, the feature data in the training dataset can be matched with the feature data in the initial risk identification strategy, and at least one sample object that successfully matches the feature data corresponding to the leaf node (i.e. the sample object hit by the leaf node) can be used as a candidate risk sample object.

[0120] In this context, the labeled risk category can be understood as the true risk category of the candidate risk sample objects, which can include either risky objects or normal objects. The identification performance parameter can be understood as a parameter indicating the identification performance of the initial risk identification strategy. There are various types of identification performance parameters, such as precision, recall, or other parameters that reflect identification performance. Based on the labeled risk category, there are several ways to calculate at least one identification performance parameter of the initial risk identification strategy. For example, based on the labeled risk category, the number of successfully identified sample objects among the candidate risk sample objects can be determined to obtain the current identification count. Based on the current identification count, the precision and recall of the initial risk identification strategy can be determined, and these precision and recall can be used as identification performance parameters.

[0121] After calculating at least one identification performance parameter of the initial risk identification strategy based on the labeled risk category, at least one candidate risk identification strategy with an identification performance parameter greater than a preset parameter threshold can be selected from the initial risk identification strategy.

[0122] 105. Based on historical traffic data and current data to be identified, test the candidate risk identification strategy at least once.

[0123] Historical traffic data can be understood as the traffic data of the object interaction platform up to the current moment.

[0124] The data to be identified can be understood as the data of the objects to be identified that need to be risk-identified within the current time frame.

[0125] There are several ways to test candidate risk identification strategies at least once based on historical traffic data and current data to be identified, as follows:

[0126] For example, historical traffic data can be used to evaluate candidate risk identification strategies, so as to select at least one risk identification strategy to be tested from the candidate risk identification strategies, obtain the current data to be identified in the current time range, identify the current object data of at least one object to be identified in the data to be identified, and use the risk identification strategy to be tested to identify the object to be identified, so as to obtain the test results of the risk identification strategy to be tested.

[0127] There are several ways to evaluate candidate risk identification strategies using historical traffic data. For example, candidate risk identification strategies can be used to identify objects in historical traffic data to obtain a first identification result, obtain risk object information in historical traffic data, and rank candidate risk identification strategies based on risk object information and the first identification result. Based on the ranking result, at least one risk identification strategy to be tested can be identified among the candidate risk identification strategies.

[0128] The method of using a candidate risk identification strategy to identify risks in historical traffic data is similar to the method of using an initial risk identification strategy to identify risks in sample objects, as described above, and will not be repeated here.

[0129] In this context, risk object information can be understood as information about risk objects in historical traffic data, indicating at least one risk object in the historical traffic data. After using candidate risk identification strategies to identify objects in historical traffic data and obtain risk object information, the candidate risk identification strategies can be ranked based on the risk object information and the first identification result. There are several ways to rank candidate risk identification strategies based on risk object information and the first identification result. For example, based on the first identification result, the number of first risk objects identified by each candidate risk identification strategy in historical traffic data, the number of second risk objects identified in the risk object information that have been reported at least once, and the number of third risk objects that have at least one negative feedback are calculated. Based on the number of first, second, and third risk objects, the identification accuracy rate in at least one dimension is calculated, and the candidate risk identification strategies are ranked by quality based on the identification accuracy rate.

[0130] After ranking the candidate risk identification strategies based on the risk object information and the initial identification result, at least one risk identification strategy to be tested can be identified from the candidate risk identification strategies according to the ranking result. There are several ways to identify at least one risk identification strategy to be tested from the candidate risk identification strategies according to the ranking result. For example, the top N candidate risk identification strategies can be identified from the candidate risk identification strategies according to the ranking result, thereby obtaining at least one risk strategy to be tested.

[0131] Taking the sample object as the object account as an example, evaluating the candidate risk identification strategy through historical traffic data can be regarded as backtracking the account data of the object interaction platform within the scope of historical events. After comprehensively calculating multiple indicators such as the number of backtracked accounts, the number of reported accounts, and the number of negative feedback (deleted, blacklisted, etc.), the candidate risk identification strategies are ranked, and the high-priority candidate risk identification strategies are used as the risk identification strategies to be tested.

[0132] There are several ways to obtain the current data to be identified within the current time range. For example, you can obtain the current traffic data of the object interaction platform, filter out the object data within the current time range from the current traffic data, and then filter out the target object data that has at least one object behavior in the behavior sequence corresponding to the candidate risk identification strategy from the object data, thereby obtaining the current data to be identified.

[0133] After acquiring the current data to be identified within the current time range, the current object data of at least one object to be identified can be identified from the current data. There are several ways to identify the current object data of at least one object to be identified from the current data. For example, the object identifier of at least one object to be identified can be identified from the current data, and the current object data of each object to be identified can be filtered out from the current data based on the object identifier.

[0134] After identifying at least one object's current data from the current data to be identified, the risk identification strategy to be tested can be used to identify the risk of the object, thus obtaining the test result of the risk identification strategy. There are several ways to use the risk identification strategy to identify the object. For example, based on the current object data, the risk identification strategy to be tested can be used to identify the risk of the object, obtaining a second identification result. The second identification result and the current object data are then sent to a verification server, whereby the verification server verifies the second identification result based on the current object data. The verification result returned by the verification server is then received and used as the test result for the risk identification measurement.

[0135] There are several ways to identify risks in the target object based on the current object data and using the risk identification strategy to be tested. For example, the strategy engine can read the normalized strategy from the script in the risk identification strategy to be tested, and deploy the normalized strategy corresponding to the extracted risk identification strategy to be tested. The deployed initial risk identification strategy can then be used to identify risks in the sample object to obtain the second identification result.

[0136] The method of using the initial risk identification strategy after deployment to identify risks in sample objects is similar to that of using the initial risk identification strategy, as detailed above, and will not be repeated here.

[0137] After identifying the risk of the target object using the risk identification strategy to be tested, the second identification result and the current object data can be sent to the verification server so that the verification server can verify the second identification result based on the current object data. The verification server can also manually verify the second identification result. The verification result is understood as information indicating whether the second identification result is correct.

[0138] After sending the second identification result and the current object data to the verification server so that the verification server can verify the second identification result based on the current object data, the verification result of the second identification result returned by the verification server can be received and used as the test result of the risk identification strategy to be tested.

[0139] Taking the object account in the object interaction platform of the sample object as an example, for the risk identification strategy to be tested, the standardized test strategy can be read directly from the script through the strategy engine. After automatic deployment, it is sent to the account review platform for manual sampling and verification, so as to further determine the accuracy of the risk identification strategy to be tested in the current time period.

[0140] 106. Based on the test results, determine the target risk identification strategy corresponding to the behavior sequence from the candidate risk identification strategies, and use the target risk identification strategy to identify risks.

[0141] There are several ways to determine the target risk identification strategy corresponding to the behavior sequence from the candidate risk identification strategies based on the test results, as follows:

[0142] For example, based on the test results, the number of correctly identified second identification results can be counted to obtain the target identification number. Based on the target identification number and the number of objects to be identified, the identification accuracy of the risk identification strategy to be tested within the current time range can be determined. Based on the identification accuracy, at least one risk identification strategy to be verified can be selected from the risk identification strategies to be verified to obtain the target risk identification strategy corresponding to the behavior sequence.

[0143] There are several ways to determine the accuracy of the risk identification strategy under test within the current time range based on the number of target identifications and the number of objects to be identified. For example, the number of objects to be identified that need to be identified can be counted in the current data to be identified, and the ratio between the number of target identifications and the number of objects can be calculated to obtain the accuracy of the risk identification strategy under test within the current time range.

[0144] After determining the recognition accuracy of the risk identification strategy to be tested within the current time range based on the number of targets to be identified and the number of objects to be identified, at least one risk identification strategy to be verified can be selected from the risk identification strategies to be verified based on the recognition accuracy, thus obtaining the target risk identification strategy corresponding to the behavior sequence. There are several ways to select at least one risk identification strategy to be verified based on the recognition accuracy. For example, the risk identification strategies to be verified can be ranked based on the recognition accuracy, and the top N risk identification strategies to be verified can be selected based on the ranking results to obtain the target risk identification strategy corresponding to the behavior sequence. Alternatively, at least one risk identification strategy to be verified with a recognition accuracy exceeding a preset accuracy threshold can be selected from the risk identification strategies to obtain the target risk identification strategy corresponding to the behavior sequence, and so on.

[0145] It should be noted that the target risk identification strategy corresponding to the behavioral sequence may include one risk identification strategy to be verified, or multiple risk identification strategies to be verified, and so on.

[0146] After determining the target risk identification strategy corresponding to the behavioral sequence from the candidate risk identification strategies based on the test results, the target risk identification strategy can be used for risk identification. There are several ways to use the target risk identification strategy for risk identification. For example, current traffic data can be obtained, and target object data of target objects with at least one object behavior in the behavioral sequence corresponding to the target risk identification strategy can be filtered out from the current traffic data. At least one feature data of the object behavior can be identified in the target object data to obtain target feature data. The target feature data can be matched with at least one feature data corresponding to the target risk identification strategy. When the match is successful, the target object can be determined to be a risk object. When the match fails, the target object can be determined to be a normal object, and so on.

[0147] Optionally, in some embodiments, after risk identification using a target risk identification strategy, the target risk identification strategy can also be tested. There are various ways to test the target risk identification strategy. For example, when the time spent using the target risk identification strategy for risk identification exceeds a preset time threshold, the target risk identification strategy is tested. Based on the test results, the current identification accuracy of the target risk identification strategy is determined, and the difference between the current identification accuracy and the accuracy of the target risk identification strategy during deployment is calculated. When the difference is greater than a preset difference threshold, the use of the target risk identification strategy for risk identification is stopped.

[0148] When the time for risk identification using the target risk identification strategy exceeds a preset time threshold, there are multiple ways to detect the target risk identification strategy. For example, risk feature data of at least one feature type of at least one current risk object identified by the target risk identification strategy can be obtained and sent to the verification server so that the verification server can review the current risk object based on the risk feature data. The verification server can then receive the verification information returned by the verification server regarding whether the current risk object is a risk object and use this verification information as the detection result of the target risk identification strategy.

[0149] After testing the target risk identification strategy, the current identification accuracy of the strategy can be determined based on the test results. There are several ways to determine the current identification accuracy based on the test results. For example, the number of real risk objects can be counted from the test results to obtain the current number of risks. The ratio between the current number of risks and the current number of risk objects can then be calculated to obtain the current identification accuracy of the target risk identification strategy.

[0150] After determining the current identification accuracy of the target risk identification strategy, the difference between the current accuracy and the historical accuracy of the target risk identification strategy can be calculated. There are several ways to calculate this difference. For example, if the target risk identification strategy has not been tested since deployment, the identification accuracy at the time of deployment can be used as the historical accuracy, and the difference between the current and historical accuracy can be calculated. If the target risk identification strategy has been tested at least once since deployment, the identification accuracy at the time of the last test can be obtained to obtain the historical accuracy, and the difference between the current and historical accuracy can be calculated, and so on.

[0151] After calculating the difference between the current identification accuracy and the historical identification accuracy of the target risk, the target risk identification strategy can be stopped when the difference is greater than or equal to a preset difference threshold.

[0152] Optionally, in some embodiments, when the difference is less than a preset difference threshold, the target risk identification strategy can continue to be used for risk identification. When it is detected that the time between the target risk identification strategy and the last detection exceeds a preset time distance, the step of continuing to detect the target risk identification strategy can be returned until the difference is greater than or equal to the preset difference threshold, thereby stopping the target risk identification strategy from performing risk identification.

[0153] In this solution, a three-stage quality assessment is conducted to identify the target risk identification strategy for each behavioral sequence. This strategy is then deployed and converted to an audit mode, eliminating the need for manual review. However, over time, these target risk identification strategies will still undergo manual sampling review / verification via a verification server. Strategies showing a significant decrease in accuracy will be identified as low-quality risk identification strategies and automatically removed from the system (i.e., their use for risk identification will cease).

[0154] In this example, taking the sample object as the object account in the object interaction platform, the risk object as the black sample, and the normal object as the white sample, the overall framework for automatically generating the target risk identification strategy for behavioral sequences in this solution can be as follows: Figure 3 As shown, the main steps include constructing a black and white sample set, constructing at least one behavioral sequence (pattern), feature snapshot acquisition, data downsampling, feature parsing, feature processing, multi-model training, candidate risk identification strategy generation, traffic replay to evaluate the quality of candidate risk identification strategies, priority strategy review and verification, and automatic deployment of high-priority strategies. Specifically, it can be described as follows:

[0155] (1) Construct a set of black and white samples: Based on the negative feedback experience, at least one black sample is generated. At the same time, these black samples can be classified by type. Some normal samples are randomly sampled from the full disk of the object interaction platform to enter the subsequent training. This ensures that they are mutually exclusive with the black samples and that there will be no duplicate accounts.

[0156] (2) Construct at least one behavioral sequence (pattern): The types of the constructed black and white sample sets are actually very different, but groups with the same behavioral sequence are more likely to be a team, and it is also easier to extract common risk identification strategies. Therefore, for this type of group, a behavioral sequence or pattern corresponding to this type of group can be constructed. The process of constructing a behavioral sequence or pattern may include automatically reporting and storing account behaviors (such as registration, login, adding friends, blocking, deleting, reporting, or other behaviors), and then backtracking and constructing their respective behavioral sequences for black and white samples. Here, the backtracking of the sequence is based on account pairs. In addition, the constructed behaviors can mainly include interactive behavior sequences corresponding to interactive behaviors and single-point behavior sequences corresponding to single-point behaviors. Interactive behavior sequences are strongly related to the interactive objects and are complementary and identical. Single-point behavior sequences are shared and can be added to the interactive behavior sequences. At the same time, these behavior sequences are sorted in chronological order.

[0157] (3) Feature Snapshot Collection: Feature collection is performed on a sample set of samples with the same behavior sequence. This feature collection mainly involves collecting feature snapshots of black and white samples in the sample set when they perform the behavior in the behavior sequence. Different behavior types will have feature snapshots (e.g., logging in, adding friends, joining or leaving groups, etc.). For black and white samples in the sample set involving the above behaviors, feature snapshots will be collected within a specified time window to obtain feature data.

[0158] (4) Data downsampling: Before feature processing and model training, the data needs to be sampled to prevent some accounts from having a large number of snapshots that are homogeneous, thus polluting the credibility of the dataset. A hyperparameter n is limited to ensure that each user can retain at most n snapshots, and these n snapshots are randomly collected.

[0159] (5) Feature parsing: Each snapshot data involves thousands of features, which need to be parsed. For example, some features were not collected, so default values ​​need to be added to the features according to the situation when the features were generated, or the features need to be converted according to their type, such as floating point or string. Each feature has a clear type when it is generated, so different processing methods are needed for different features. For example, one-hot encoding is needed for features. For character data such as country and province, it is not possible to directly input them into the model. At the same time, it is also unreasonable to directly encode them as 1, 2, 3 (there is no size relationship between countries and provinces). Therefore, one-hot encoding is needed for processing, and so on.

[0160] (6) Feature Processing: For many continuous feature data, direct input may not yield good distinctions. For example, ages 17 and 18 are quite similar, but for the model, they are two different inputs. Therefore, we designed an automatic binning strategy for continuous data. We automatically discretize continuous features based on the chi-square distribution, evaluate the confidence level, and select the optimal distribution for final binning, thus obtaining the target continuous feature data. The initial feature data is then updated based on the target continuous feature data to obtain the updated feature data. The updated feature data is then filtered using the Information Value (IV) to obtain the target feature data. IV binning is an effective feature selection and processing method that can understand the relationship between variables and target variables and provide valuable input for subsequent modeling. Through reasonable binning and IV value calculation, the model's performance and interpretability can be improved. Finally, combined with expert experience, some feature data with weak generality are filtered out, thus obtaining the training dataset corresponding to each row sequence.

[0161] (7) Multi-model training: After processing the feature data, the data is fitted and trained using models such as decision trees and random forests. Multiple parameters such as tree depth, number of trees, and number of randomly sampled features are adjusted until the model converges, thus obtaining at least one trained tree model.

[0162] (8) Generation of candidate strategies (risk identification strategies): For the tree model that has been trained, each path from the root node to the leaf node can be traversed according to the execution path of the tree model to obtain at least one initial risk identification strategy. Then, based on the accounts hit by the leaf nodes, the precision and recall of each initial risk identification strategy are calculated. A specific threshold is set to filter out low-value initial risk identification strategies, and the remaining ones are the actual candidate risk identification strategies.

[0163] (9) Traffic replay evaluation of candidate strategies (candidate risk identification strategies): For candidate risk identification strategies, it is actually a combination of multiple features. The market data of the past x days is used to backtrack the accounts. After comprehensive calculation of multiple indicators such as the number of backtracked accounts, the number of reported accounts, and the number of negative feedback (deleted, blacklisted, etc.), the candidate risk identification strategies are ranked to obtain the high-priority risk identification strategies for testing.

[0164] (10) Priority Strategy Review and Verification: For the high-priority risk identification strategies extracted for testing, the strategy engine can directly read the standardized strategy from the script, automatically deploy it online, and send it to the account review platform (verification server) for manual spot checks and verification. Based on the verification results, the accuracy of the risk identification strategies to be tested in the current time period can be further confirmed.

[0165] (11) Automatic deployment of high-quality strategies: For the target risk identification strategies that have been reviewed in the three stages, the best performing strategy is adjusted to obtain the target risk identification strategy corresponding to each behavior sequence. The target risk identification strategy is automatically converted into audit mode, and no manual review is required. Moreover, as time goes by, each target risk identification strategy will still be sampled and reviewed. Target risk identification strategies with significantly reduced accuracy will be judged as low-quality strategies and automatically taken offline.

[0166] This solution utilizes an automated risk identification strategy generation framework, moving away from reliance on expert experience for strategy extraction. Instead, it improves efficiency by generating multiple candidate risk identification strategies through standardized sample collection, standardized data feature processing, efficient model training, and model parsing. Furthermore, a three-stage strategy quality assessment is designed for the model-generated strategy process: a quality assessment and screening of strategies for the current dataset, a secondary screening based on traffic replay over a past period, and a simple manual sampling screening of strategies for the current time period after a no-load run. This ensures reliable, effective, and high-quality target risk identification strategies for automatic deployment. Finally, a data script-driven approach achieves greater automation, eliminating the need for manual strategy analysis, design, and deployment, transforming these processes into readily accessible and self-deploying tools.

[0167] As described above, in this embodiment, after acquiring object behavior data of a sample object and constructing at least one behavior sequence based on the object behavior data, at least one feature data of the sample object under the object behavior in the behavior sequence is collected to obtain a training dataset corresponding to the behavior sequence. Then, at least one preset tree model is trained using the training dataset, and the nodes in the trained tree model are traversed to obtain at least one initial risk identification strategy corresponding to the behavior sequence. Then, based on the training dataset, at least one candidate risk identification strategy is selected from the initial risk identification strategies. The candidate risk identification strategies are tested at least once based on historical traffic data and current data to be identified. Finally, based on the test results, the target risk identification strategy corresponding to the behavior sequence is determined from the candidate risk identification strategies. This approach employs target risk identification for risk assessment. Since the data model can be trained using training data from specific behavioral sequences, an initial risk identification strategy for each behavioral sequence can be automatically generated. Then, through the training dataset, historical traffic data, and current data to be identified, the initial risk identification strategy undergoes multiple stages of quality evaluation. This effectively improves the efficiency and accuracy of risk identification strategy discovery from a large number of features. Furthermore, the generated target risk identification strategy for each behavioral sequence is primarily used to identify risks in objects with specific object behaviors (object behaviors within the behavioral sequence), without requiring risk identification for all objects, thus reducing the total number of objects to be identified and improving the efficiency of risk identification. Therefore, it can enhance the efficiency and accuracy of risk identification.

[0168] Based on the method described in the above embodiments, the following examples will provide further detailed explanations.

[0169] In this embodiment, the risk identification device will be specifically integrated into an electronic device, which will be a server, as an example for explanation.

[0170] like Figure 4 As shown, a risk identification method has the following specific process:

[0171] 201. The server obtains the object behavior data of the sample object.

[0172] For example, the server can obtain negative feedback information for at least one object in the object interaction platform, send the negative feedback information to the review server, so that the review server can review the object based on the negative feedback information, receive the review results returned by the review server, and filter out at least one risky object from the objects based on the review results.

[0173] The server can obtain a set of historical behavior data within a preset historical time range from the object interaction platform, and filter out the historical behavior data corresponding to the risk object from the set of historical behavior data to obtain the risk behavior data.

[0174] The server obtains a set of normal objects from the object interaction platform and randomly selects at least one normal object from this set. It then collects historical behavior data of these normal objects within a preset time range to obtain normal behavior data.

[0175] After the server filters out risky objects and normal objects, and collects risky behavior data and normal behavior data, it can use the risky objects and normal objects as sample objects, and use the risky behavior data and normal behavior data as the object behavior data of the sample objects.

[0176] 202. The server constructs at least one sequence of behaviors based on object behavior data.

[0177] For example, the server can identify the object behavior of a sample object within a preset historical time range from object behavior data. When the object behavior includes at least one interactive behavior, an interactive behavior sequence corresponding to the interactive behavior is constructed, and the interactive behavior sequence is used as a candidate behavior sequence. When the object behavior includes at least one single-point behavior, a single-point behavior sequence corresponding to the single-point behavior is constructed, and the single-point behavior sequence is used as a candidate behavior sequence. When the object behavior includes at least one interactive behavior and at least one single-point behavior, an interactive behavior sequence corresponding to the interactive behavior and a single-point behavior sequence corresponding to the single-point behavior sequence are constructed, and the single-point behavior sequence is added to the interactive behavior sequence to obtain at least one candidate behavior sequence.

[0178] The server can count the frequency of behaviors in the candidate behavior sequences based on the object behaviors in the candidate behavior sequences. Based on the behavior frequency, the candidate behavior sequences are sorted, and the top N candidate behavior sequences are selected from the sorting results, thus obtaining at least one behavior sequence.

[0179] 203. The server collects at least one feature data of the sample object under the object behavior to obtain the training dataset corresponding to the behavior sequence.

[0180] For example, the server can identify the object identifier of at least one sample object with the object behavior in the behavior sequence, and use the sample object corresponding to the object identifier as a bucket to obtain at least one target sample object corresponding to the behavior sequence.

[0181] The server can collect at least one data snapshot of the target sample object when it performs object behavior, and use the data snapshot as object data to obtain the original object data of the target sample object.

[0182] The server can randomly sample a portion of the original object data that is less than or equal to a preset data quantity threshold, thereby obtaining the target object data of the target sample object.

[0183] When the first feature type does not include a preset first feature type, the server obtains the object attribute information of the target sample object when performing object behavior, obtains the feature data mapping table corresponding to the preset first feature type, maps the object data information according to the feature data mapping table, thereby obtaining the current feature data corresponding to the preset first feature type, and adds the current feature data to the target object data to obtain the initial feature data of the target sample object. When the first feature type includes a feature type to be converted, the object data of the feature type to be converted is converted into feature data of a preset second feature type, and the object data corresponding to the feature type to be converted is replaced with the converted feature data in the target object data to obtain the initial feature data of the target sample object. When the first feature type includes a feature type to be encoded, the server determines the encoding method of the object data corresponding to the feature type to be encoded according to the data type of the object data, performs feature encoding on the object data corresponding to the feature type to be encoded based on the encoding method, thereby obtaining the encoded feature data, and updates the target object data based on the encoded feature data to obtain the initial feature data of the target sample object.

[0184] The server can filter out feature data of at least one continuous feature from the initial feature data to obtain continuous feature data. Based on at least one preset data distribution, the continuous feature data is discretized, and the discrete confidence level corresponding to each preset data distribution is evaluated. Based on the discrete confidence level, a target data distribution is selected from the preset data distribution, and the discrete feature data corresponding to the target data distribution is binned to obtain at least one first feature data bin corresponding to each continuous feature.

[0185] The server can count the number of sample objects corresponding to the feature data in the first feature data bin, count the number of risk objects corresponding to the feature data in the first feature data bin, obtain the number of risk objects, count the number of normal objects corresponding to the feature data in the first feature data bin, obtain the number of normal objects, calculate the ratio of the number of risk objects to the number of objects, obtain the risk object ratio (Bad Rate), calculate the ratio of the number of normal objects to the number of objects, obtain the normal object ratio (Good Rate). Calculate the ratio between the risk object ratio and the normal object ratio to obtain the good-bad ratio of the first feature data bin, and calculate the ratio difference between the risk object ratio and the normal object ratio. Calculate the comparison value of the good-bad ratio to obtain the good-bad ratio encoding parameters, as shown in formula (1).

[0186] The server multiplies the good-to-bad ratio encoding parameter by the ratio difference to obtain the initial information value parameter of the first feature data box. The initial value information parameters of the first feature data box belonging to the unified continuous feature are accumulated to obtain the first information value parameter of the continuous feature, as shown in formula (2). At least one continuous feature whose first information value parameter exceeds a preset parameter threshold is selected from the continuous features to obtain the target continuous feature. At least one continuous feature data corresponding to the target continuous feature is selected from the continuous feature data to obtain the target continuous feature data.

[0187] The server can replace continuous feature data in the initial feature data with target continuous feature data to obtain updated feature data. The updated feature data includes feature data of at least one second feature type, which includes continuous features. The first feature type may be the same as or different from the second feature type.

[0188] The server can divide the feature data corresponding to the continuous variables in the second feature type into several intervals, and merge the feature data corresponding to the categorical variables in the second feature type to obtain the adjusted feature data.

[0189] The server can use a wideband binning method to bin the adjusted feature data, thereby obtaining at least one second feature data bin corresponding to the second feature type. Alternatively, it can use an equal-frequency binning method to bin the adjusted feature data, thereby obtaining at least one second feature data bin corresponding to the second feature type. Or, it can bin the adjusted feature data based on business knowledge, thereby obtaining at least one second feature data bin corresponding to the second feature type, and so on.

[0190] The server can determine the second information value parameter of the second feature type based on the feature data in the second feature data bin, which is split into two bins. It then filters out at least one target feature type whose second information value parameter exceeds a preset threshold from the second feature type, and finally filters out the feature data corresponding to the target feature type from the updated feature data, thereby obtaining the target feature data.

[0191] The server can obtain the generality parameter corresponding to each target feature type, filter out the feature data with the generality parameter less than the preset generality parameter threshold from the target feature data, and thus obtain the training dataset corresponding to the behavior sequence. Alternatively, based on expert experience information, it can select at least one preset feature type with weak generality from the target feature types, filter out the feature data corresponding to the preset feature type from the target feature data, and thus obtain the training dataset corresponding to the behavior sequence, and so on.

[0192] 204. The server uses the training dataset to train at least one preset tree model.

[0193] For example, the server can perform fitting training on the training dataset, adjust multiple parameters of the preset tree model such as tree depth, number of trees, and number of randomly sampled features, and finally make the preset tree model converge, thus obtaining the trained tree model.

[0194] 205. The server traverses the nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence.

[0195] For example, the server can identify at least one node in the tree model, which may include a root node and leaf nodes. The node indicates the feature type corresponding to the feature data in the training dataset. The server can obtain the execution path of the tree model and traverse from the root node to the leaf node according to the execution path to obtain at least one node path. The server can determine at least one candidate feature data corresponding to the node path in the training dataset and combine these candidate feature data to obtain the initial risk identification strategy corresponding to the behavior sequence. The initial risk identification strategy represents the combination of feature data of at least one feature type used for risk identification.

[0196] 206. Based on the training dataset, the server selects at least one candidate risk identification strategy from the initial risk identification strategy.

[0197] For example, the server can match the feature data in the training dataset with the feature data in the initial risk identification strategy, and select at least one sample object that successfully matches the feature data corresponding to the leaf node (i.e., the sample object hit by the leaf node) as a candidate risk sample object.

[0198] The server obtains the labeled risk categories of candidate risk sample objects. Based on the labeled risk categories, it determines the number of successfully identified sample objects among the candidate risk sample objects, obtaining the current identification count. Based on the current identification count, it determines the precision and recall of the initial risk identification strategy, and uses precision and recall as identification performance parameters. At least one candidate risk identification strategy with identification performance parameters greater than a preset threshold is selected from the initial risk identification strategy.

[0199] 207. The server tests the candidate risk identification strategies at least once based on historical traffic data and current data to be identified.

[0200] For example, the server can employ candidate risk identification strategies to identify objects in historical traffic data, obtaining a first identification result and acquiring risk object information from the historical traffic data. Based on the first identification result, the server counts the number of first-risk objects identified by each candidate risk identification strategy in the historical traffic data, the number of second-risk objects identified as having at least one reported risk object, and the number of third-risk objects identified as having at least one negative feedback risk object. Based on the number of first-risk objects, second-risk objects, and third-risk objects, the server calculates the identification accuracy in at least one dimension. Based on the identification accuracy, the candidate risk identification strategies are ranked by quality. According to the ranking results, the top N candidate risk identification strategies are identified, thus obtaining at least one risk strategy to be tested.

[0201] The server can obtain the current traffic data of the object interaction platform, filter object data within the current time range from the current traffic data, and then filter target object data that exhibits at least one object behavior in the behavioral sequence corresponding to the candidate risk identification strategy, thereby obtaining the current data to be identified. From the current data to be identified, at least one object identifier for the object to be identified is determined. Based on the object identifier, the current object data for each object to be identified is then filtered out.

[0202] The server can use a policy engine to read the normalized policy from the script in the risk identification policy to be tested, and deploy the normalized policy corresponding to the extracted risk identification policy. The deployed initial risk identification policy is then used to identify risks in sample objects, thus obtaining a second identification result. This second identification result and the current object data are sent to a verification server so that the verification server can verify the second identification result based on the current object data. The server receives the verification result of the second identification result returned by the verification server and uses this verification result as the test result for the risk identification measurement.

[0203] 208. Based on the test results, the server determines the target risk identification strategy corresponding to the behavior sequence from the candidate risk identification strategies.

[0204] For example, the server can count the number of correctly identified second recognition results based on the test results to obtain the target recognition count.

[0205] The server can count the number of objects that need to be identified for risk identification in the current data to be identified, calculate the ratio between the target number of objects to be identified, and thus obtain the identification accuracy of the risk identification strategy to be tested within the current time range.

[0206] The server can sort the risk identification strategies to be verified based on the recognition accuracy. Based on the sorting results, it can select the top N risk identification strategies to be verified to obtain the target risk identification strategy corresponding to the behavior sequence. Alternatively, it can select at least one risk identification strategy to be verified with an accuracy exceeding a preset accuracy threshold to obtain the target risk identification strategy corresponding to the behavior sequence, and so on.

[0207] 209. The server uses a target risk identification strategy to identify risks.

[0208] For example, the server can obtain current traffic data, filter out target object data of target objects that have at least one object behavior in the behavior sequence corresponding to the target risk identification strategy, identify at least one feature data of the object that performs the object behavior in the target object data, obtain target feature data, match the target feature data with at least one feature data corresponding to the target risk identification strategy, when the match is successful, the target object can be determined to be a risk object, when the match fails, the target object can be determined to be a normal object, and so on.

[0209] Optionally, in some embodiments, when the server detects that the time for risk identification using the target risk identification strategy exceeds a preset time threshold, it acquires risk feature data of at least one feature type of at least one current risk object identified by the target risk identification strategy, and sends the risk feature data to the verification server so that the verification server can review the current risk object based on the risk feature data, receive verification information returned by the verification server regarding whether the current risk object is a risk object, and use the verification information as the detection result of the target risk identification strategy.

[0210] When the target risk identification strategy has not been tested after deployment, the server can use the identification accuracy rate of the target risk identification strategy at the time of deployment as the historical identification accuracy rate, and calculate the difference between the current identification accuracy rate and the historical identification accuracy rate. When the target risk identification strategy has been tested at least once after deployment, the server can obtain the identification accuracy rate of the target risk identification strategy at the time of the last test, obtain the historical identification accuracy rate, and calculate the difference between the current identification accuracy rate and the historical identification accuracy rate, and so on.

[0211] When the difference is greater than or equal to a preset difference threshold, the server stops using the target risk identification strategy for risk identification.

[0212] Optionally, in some embodiments, the server can continue to use the target risk identification strategy for risk identification when the difference is less than a preset difference threshold. When it is detected that the time between the target risk identification strategy and the last detection exceeds a preset time distance, the server can return to the step of continuing to detect the target risk identification strategy until the difference is greater than or equal to the preset difference threshold, thereby stopping the target risk identification strategy from performing risk identification.

[0213] As described above, in this embodiment, after the server acquires object behavior data of a sample object and constructs at least one behavior sequence based on the object behavior data, it collects at least one feature data of the sample object under the object behavior in the behavior sequence to obtain a training dataset corresponding to the behavior sequence. Then, it uses the training dataset to train at least one preset tree model and traverses the nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence. Then, based on the training dataset, it selects at least one candidate risk identification strategy from the initial risk identification strategies. Based on historical traffic data and current data to be identified, it tests the candidate risk identification strategies at least once. Finally, based on the test results, it determines the target risk identification corresponding to the behavior sequence from the candidate risk identification strategies. The strategy employs target risk identification for risk identification. Since this solution can train the data model using training data from specific behavioral sequences, it automatically generates an initial risk identification strategy for each behavioral sequence. Then, through the training dataset, historical traffic data, and current data to be identified, the initial risk identification strategy undergoes multiple stages of quality evaluation. This effectively improves the efficiency and accuracy of risk identification strategy discovery from a large number of features. Furthermore, the generated target risk identification strategy for each behavioral sequence is primarily used to identify risks in objects with specific object behaviors (object behaviors within the behavioral sequence), without needing to identify risks in all objects, thus reducing the total number of objects to be identified and improving the efficiency of risk identification. Therefore, it can improve the efficiency and accuracy of risk identification.

[0214] To better implement the above methods, this application also provides a risk identification device, which can be integrated into an electronic device, such as a server or terminal. The terminal may include a tablet computer, a laptop computer, and / or a personal computer.

[0215] For example, such as Figure 5 As shown, the risk identification device may include an acquisition unit 301, a collection unit 302, a training unit 303, a screening unit 304, a testing unit 305, and an identification unit 306, as follows:

[0216] (1) Obtain unit 301;

[0217] The acquisition unit 301 is used to acquire object behavior data of the sample object and construct at least one behavior sequence based on the object behavior data, the behavior sequence indicating the sample object having at least one object behavior.

[0218] For example, the acquisition unit 301 can be specifically used to acquire negative feedback information for at least one object in the object interaction platform, and based on the negative feedback information, to select at least one risk object from the objects, collect the behavior data of the risk object within a preset historical time range to obtain risk behavior data, select at least one normal object from the normal object set of the object interaction platform, collect the behavior data of the normal object within a preset historical time range to obtain normal behavior data, use the risk object and the normal object as sample objects, and use the risk behavior data and the normal behavior data as the object behavior data of the sample objects, identify the object behavior of the sample objects within the preset historical time range in the object behavior data, construct at least one candidate behavior sequence corresponding to the sample object based on the object behavior, the candidate behavior sequence includes at least one object behavior, count the behavior frequency of the candidate behavior sequence based on the object behavior in the candidate behavior sequence, and select at least one behavior sequence from the candidate behavior sequence based on the behavior frequency.

[0219] (2) Acquisition unit 302;

[0220] The acquisition unit 302 is used to acquire at least one feature data of the sample object under the object behavior to obtain a training dataset corresponding to the behavior sequence.

[0221] For example, the acquisition unit 302 can be specifically used to: bucket sample objects according to object behaviors in the behavior sequence to obtain at least one target sample object corresponding to the behavior sequence; collect at least one feature data of the target sample object under the execution of object behaviors to obtain initial feature data of the target sample object; filter at least one continuous feature data from the initial feature data to obtain continuous feature data; bin the continuous feature data to obtain target continuous feature data; update the initial feature data based on the target continuous feature data to obtain updated feature data, which includes at least one feature data of a second feature type, which includes continuous features; and filter at least one target feature data of a target feature type from the updated feature data to obtain the training dataset corresponding to the behavior sequence.

[0222] (3) Training Unit 303;

[0223] Training unit 303 is used to train at least one preset tree model using a training dataset and traverse the nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence.

[0224] For example, training unit 303 can be used to train at least one preset tree model using a training dataset, identify at least one node in the tree model, the node may include a root node and a leaf node, the node indicates the feature type corresponding to the feature data in the training dataset, obtain the execution path of the tree model, and traverse from the root node to the leaf node according to the execution path to obtain at least one node path, and determine the initial risk identification strategy corresponding to the behavior sequence based on the node path, the initial risk identification strategy characterizing the combination of feature data of at least one feature type used for risk identification.

[0225] (4) Filtering unit 304;

[0226] The screening unit 304 is used to select at least one candidate risk identification strategy from the initial risk identification strategy based on the training dataset.

[0227] For example, the screening unit 304 can be used to identify the risk of sample objects based on the training dataset using an initial risk identification strategy, and obtain a third identification result. The third identification result includes at least one candidate risk sample object indicated by the leaf node corresponding to the initial risk identification strategy, obtain the labeled risk category of the candidate risk sample object, and calculate at least one identification performance parameter of the initial risk identification strategy based on the labeled risk category, and screen at least one candidate risk identification strategy whose identification performance parameter is greater than a preset parameter threshold from the initial risk identification strategy.

[0228] (5) Test unit 305;

[0229] Test unit 305 is used to test the candidate risk identification strategy at least once based on historical traffic data and current data to be identified.

[0230] For example, test unit 305 can be used to identify objects in historical traffic data using candidate risk identification strategies, obtain a first identification result, acquire risk object information in historical traffic data, and sort candidate risk identification strategies based on risk object information and the first identification result. According to the sorting result, at least one risk identification strategy to be tested is identified among the candidate risk identification strategies. The current data to be identified within the current time range is acquired, and the current object data of at least one object to be identified is identified in the data to be identified. Based on the current object data, the risk identification strategy to be tested is used to identify the object to be identified, and a second identification result is obtained. The second identification result and the current object data are sent to the verification server so that the verification server can verify the second identification result based on the current object data. The verification result of the second identification result returned by the verification server is received, and the verification result is used as the test result of the risk identification measurement to be tested.

[0231] (6) Identification unit 306;

[0232] The identification unit 306 is used to determine the target risk identification strategy corresponding to the behavior sequence from the candidate risk identification strategies based on the test results, and to use the target risk identification strategy to identify risks.

[0233] For example, the identification unit 306 can be used to count the number of correctly identified second identification results based on the test results, obtain the target identification number, determine the identification accuracy of the risk identification strategy to be tested within the current time range based on the target identification number and the number of objects to be identified, select at least one risk identification strategy to be verified from the risk identification strategies to be verified based on the identification accuracy, obtain the target risk identification strategy corresponding to the behavior sequence, and use the target risk identification strategy to perform risk identification.

[0234] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0235] As can be seen from the above, in this embodiment of the application, after the acquisition unit 301 acquires the object behavior data of the sample object and constructs at least one behavior sequence based on the object behavior data, the collection unit 302 collects at least one feature data of the sample object under the object behavior in the behavior sequence to obtain the training dataset corresponding to the behavior sequence. Then, the training unit 303 uses the training dataset to train at least one preset tree model and traverses the nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence. Then, the filtering unit 304 filters at least one candidate risk identification strategy from the initial risk identification strategies based on the training dataset. The testing unit 305 tests the candidate risk identification strategies at least once based on historical traffic data and current data to be identified. Finally, the identification unit 306 identifies the candidate risk identification strategy based on the test results. The strategy identifies the target risk identification strategy corresponding to the behavior sequence and uses target risk identification for risk identification. Since this solution can automatically generate an initial risk identification strategy for each behavior sequence by training the data model with training data of specific behavior sequences, and then conducts multi-stage quality evaluation of the initial risk identification strategy using the training dataset, historical traffic data, and current data to be identified, it can effectively improve the efficiency and accuracy of risk identification strategy discovery from a large number of features. Furthermore, the generated target risk identification strategy for each behavior sequence is mainly used to identify risks for objects with specific object behaviors (object behaviors in the behavior sequence), without needing to identify risks for all objects, thus reducing the total number of objects to be identified and improving the efficiency of risk identification. Therefore, it can improve the efficiency and accuracy of risk identification.

[0236] This application also provides an electronic device, such as... Figure 6 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically:

[0237] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0238] The processor 401 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes software programs and / or modules stored in the memory 402, and calls data stored in the memory 402, to perform various functions and process data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.

[0239] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0240] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0241] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0242] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:

[0243] Obtain object behavior data of sample objects, and construct at least one behavior sequence based on the object behavior data. The behavior sequence indicates a sample object with at least one object behavior. Collect at least one feature data of the sample object under the object behavior to obtain a training dataset corresponding to the behavior sequence. Use the training dataset to train at least one preset tree model, and traverse the nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence. Based on the training dataset, select at least one candidate risk identification strategy from the initial risk identification strategies. Test the candidate risk identification strategies at least once based on historical traffic data and current data to be identified. Based on the test results, determine the target risk identification strategy corresponding to the behavior sequence from the candidate risk identification strategies, and use the target risk identification strategy to perform risk identification.

[0244] For example, an electronic device can acquire negative feedback information for at least one object in an object interaction platform, and based on the negative feedback information, filter out at least one risk object from the objects, collect behavioral data of the risk object within a preset historical time range to obtain risk behavior data, filter out at least one normal object from the normal object set of the object interaction platform, collect behavioral data of the normal object within a preset historical time range to obtain normal behavior data, use the risk object and normal object as sample objects, and use the risk behavior data and normal behavior data as object behavior data of the sample objects, identify the object behavior of the sample objects within the preset historical time range in the object behavior data, construct at least one candidate behavior sequence corresponding to the sample object based on the object behavior, the candidate behavior sequence includes at least one object behavior, count the behavior frequency of the candidate behavior sequence based on the object behavior in the candidate behavior sequence, and filter out at least one behavior sequence from the candidate behavior sequence based on the behavior frequency. Based on the object behaviors in the behavior sequence, the sample objects are bucketed to obtain at least one target sample object corresponding to the behavior sequence. At least one feature data of the target sample object under the execution of the object behavior is collected to obtain the initial feature data of the target sample object. At least one continuous feature data is selected from the initial feature data to obtain continuous feature data. The continuous feature data is binned to obtain target continuous feature data. The initial feature data is updated based on the target continuous feature data to obtain updated feature data. The updated feature data includes at least one feature data of a second feature type, which includes continuous features. At least one target feature data of a target feature type is selected from the updated feature data to obtain the training dataset corresponding to the behavior sequence. The training dataset is used to train at least one preset tree model. At least one node is identified in the tree model. The node may include a root node and a leaf node. The node indicates the feature type corresponding to the feature data in the training dataset. The execution path of the tree model is obtained. According to the execution path, traversal is performed from the root node to the leaf node to obtain at least one node path. Based on the node path, an initial risk identification strategy corresponding to the behavior sequence is determined. The initial risk identification strategy represents the combination of feature data of at least one feature type used for risk identification. Based on the training dataset, an initial risk identification strategy is used to identify the risk of sample objects, resulting in a third identification result. This third identification result includes at least one candidate risk sample object indicated by the leaf node corresponding to the initial risk identification strategy. The labeled risk category of the candidate risk sample object is obtained, and based on the labeled risk category, at least one identification performance parameter of the initial risk identification strategy is calculated. At least one candidate risk identification strategy with an identification performance parameter greater than a preset parameter threshold is selected from the initial risk identification strategy.A candidate risk identification strategy is used to identify objects in historical traffic data to obtain a first identification result. Risk object information in the historical traffic data is obtained, and the candidate risk identification strategies are ranked based on the risk object information and the first identification result. According to the ranking result, at least one risk identification strategy to be tested is identified from the candidate risk identification strategies. The current data to be identified within the current time range is obtained, and the current object data of at least one object to be identified is identified from the data to be identified. Based on the current object data, the risk identification strategy to be tested is used to identify the object to be identified to obtain a second identification result. The second identification result and the current object data are sent to the verification server so that the verification server can verify the second identification result based on the current object data. The verification result of the second identification result returned by the verification server is received, and the verification result is used as the test result of the risk identification measurement to be tested. Based on the test results, the number of correctly identified second identification results is counted in the second identification results to obtain the target identification number. Based on the target identification number and the number of objects to be identified, the identification accuracy of the risk identification strategy to be tested in the current time range is determined. Based on the identification accuracy, at least one risk identification strategy to be verified is selected from the risk identification strategies to be verified to obtain the target risk identification strategy corresponding to the behavior sequence. The target risk identification strategy is then used for risk identification.

[0245] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0246] As described above, in this embodiment, after acquiring object behavior data of a sample object and constructing at least one behavior sequence based on the object behavior data, at least one feature data of the sample object under the object behavior in the behavior sequence is collected to obtain a training dataset corresponding to the behavior sequence. Then, at least one preset tree model is trained using the training dataset, and the nodes in the trained tree model are traversed to obtain at least one initial risk identification strategy corresponding to the behavior sequence. Then, based on the training dataset, at least one candidate risk identification strategy is selected from the initial risk identification strategies. The candidate risk identification strategies are tested at least once based on historical traffic data and current data to be identified. Finally, based on the test results, the target risk identification strategy corresponding to the behavior sequence is determined from the candidate risk identification strategies. This approach employs target risk identification for risk assessment. Since the data model can be trained using training data from specific behavioral sequences, an initial risk identification strategy for each behavioral sequence can be automatically generated. Then, through the training dataset, historical traffic data, and current data to be identified, the initial risk identification strategy undergoes multiple stages of quality evaluation. This effectively improves the efficiency and accuracy of risk identification strategy discovery from a large number of features. Furthermore, the generated target risk identification strategy for each behavioral sequence is primarily used to identify risks in objects with specific object behaviors (object behaviors within the behavioral sequence), without requiring risk identification for all objects, thus reducing the total number of objects to be identified and improving the efficiency of risk identification. Therefore, it can enhance the efficiency and accuracy of risk identification.

[0247] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0248] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the risk identification methods provided in embodiments of this application. For example, the instructions can execute the following steps:

[0249] Obtain object behavior data of sample objects, and construct at least one behavior sequence based on the object behavior data. The behavior sequence indicates a sample object with at least one object behavior. Collect at least one feature data of the sample object under the object behavior to obtain a training dataset corresponding to the behavior sequence. Use the training dataset to train at least one preset tree model, and traverse the nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence. Based on the training dataset, select at least one candidate risk identification strategy from the initial risk identification strategies. Test the candidate risk identification strategies at least once based on historical traffic data and current data to be identified. Based on the test results, determine the target risk identification strategy corresponding to the behavior sequence from the candidate risk identification strategies, and use the target risk identification strategy to perform risk identification.

[0250] For example, acquire negative feedback information for at least one object in the object interaction platform, and based on the negative feedback information, filter out at least one risk object from the objects, collect the behavioral data of the risk object within a preset historical time range to obtain risk behavior data, filter out at least one normal object from the normal object set of the object interaction platform, collect the behavioral data of the normal object within a preset historical time range to obtain normal behavior data, use the risk object and normal object as sample objects, and use the risk behavior data and normal behavior data as the object behavior data of the sample objects, identify the object behavior of the sample objects within the preset historical time range in the object behavior data, construct at least one candidate behavior sequence corresponding to the sample object based on the object behavior, the candidate behavior sequence includes at least one object behavior, count the behavior frequency of the candidate behavior sequence based on the object behavior in the candidate behavior sequence, and filter out at least one behavior sequence from the candidate behavior sequence based on the behavior frequency. Based on the object behaviors in the behavior sequence, the sample objects are bucketed to obtain at least one target sample object corresponding to the behavior sequence. At least one feature data of the target sample object under the execution of the object behavior is collected to obtain the initial feature data of the target sample object. At least one continuous feature data is selected from the initial feature data to obtain continuous feature data. The continuous feature data is binned to obtain target continuous feature data. The initial feature data is updated based on the target continuous feature data to obtain updated feature data. The updated feature data includes at least one feature data of a second feature type, which includes continuous features. At least one target feature data of a target feature type is selected from the updated feature data to obtain the training dataset corresponding to the behavior sequence. The training dataset is used to train at least one preset tree model. At least one node is identified in the tree model. The node may include a root node and a leaf node. The node indicates the feature type corresponding to the feature data in the training dataset. The execution path of the tree model is obtained. According to the execution path, traversal is performed from the root node to the leaf node to obtain at least one node path. Based on the node path, an initial risk identification strategy corresponding to the behavior sequence is determined. The initial risk identification strategy represents the combination of feature data of at least one feature type used for risk identification. Based on the training dataset, an initial risk identification strategy is used to identify the risk of sample objects, resulting in a third identification result. This third identification result includes at least one candidate risk sample object indicated by the leaf node corresponding to the initial risk identification strategy. The labeled risk category of the candidate risk sample object is obtained, and based on the labeled risk category, at least one identification performance parameter of the initial risk identification strategy is calculated. At least one candidate risk identification strategy with an identification performance parameter greater than a preset parameter threshold is selected from the initial risk identification strategy.A candidate risk identification strategy is used to identify objects in historical traffic data to obtain a first identification result. Risk object information in the historical traffic data is obtained, and the candidate risk identification strategies are ranked based on the risk object information and the first identification result. According to the ranking result, at least one risk identification strategy to be tested is identified from the candidate risk identification strategies. The current data to be identified within the current time range is obtained, and the current object data of at least one object to be identified is identified from the data to be identified. Based on the current object data, the risk identification strategy to be tested is used to identify the object to be identified to obtain a second identification result. The second identification result and the current object data are sent to the verification server so that the verification server can verify the second identification result based on the current object data. The verification result of the second identification result returned by the verification server is received, and the verification result is used as the test result of the risk identification measurement to be tested. Based on the test results, the number of correctly identified second identification results is counted in the second identification results to obtain the target identification number. Based on the target identification number and the number of objects to be identified, the identification accuracy of the risk identification strategy to be tested in the current time range is determined. Based on the identification accuracy, at least one risk identification strategy to be verified is selected from the risk identification strategies to be verified to obtain the target risk identification strategy corresponding to the behavior sequence. The target risk identification strategy is then used for risk identification.

[0251] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0252] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0253] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the risk identification methods provided in the embodiments of this application, the beneficial effects that any of the risk identification methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0254] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in the various optional implementations of the risk identification aspect or the automatic generation of risk identification strategies described above.

[0255] The above provides a detailed description of a risk identification method and related equipment provided in the embodiments of this application. The related equipment may include risk identification devices and electronic devices. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A risk identification method, characterized by, include: Obtain object behavior data of a sample object, and construct at least one behavior sequence based on the object behavior data, the behavior sequence indicating the sample object having at least one object behavior; Collect at least one feature data of the sample object under the object behavior to obtain the training dataset corresponding to the behavior sequence; The training dataset is used to train at least one preset tree model, and the nodes in the trained tree model are traversed to obtain at least one initial risk identification strategy corresponding to the behavior sequence. Based on the training dataset, at least one candidate risk identification strategy is selected from the initial risk identification strategy. The candidate risk identification strategy is tested at least once based on historical traffic data and current data to be identified; Based on the test results, the target risk identification strategy corresponding to the behavior sequence is determined from the candidate risk identification strategies, and the target risk identification strategy is used for risk identification.

2. The risk identification method of claim 1, wherein, The construction of at least one behavior sequence based on the object behavior data includes: Identify the object behavior of the sample object within a preset historical time range from the object behavior data; Based on the object behavior, at least one candidate behavior sequence corresponding to the sample object is constructed, and the candidate behavior sequence includes at least one of the object behaviors; Based on the object behaviors in the candidate behavior sequences, the frequency of the behaviors in the candidate behavior sequences is counted, and at least one behavior sequence is selected from the candidate behavior sequences according to the behavior frequency.

3. The risk identification method of claim 2, wherein, The object behavior includes interactive behavior or single-point behavior, whereby the single-point behavior characterizes the behavior generated by the sample object itself. Constructing at least one candidate behavior sequence corresponding to the sample object based on the object behavior includes: When the object behavior includes at least one of the interaction behaviors, an interaction behavior sequence corresponding to the interaction behavior is constructed, and the interaction behavior sequence is used as the candidate behavior sequence. When the object behavior includes at least one of the single-point behaviors, a single-point behavior sequence corresponding to the single-point behavior is constructed, and the single-point behavior sequence is used as the candidate behavior sequence. When the object behavior includes at least one interactive behavior and at least one single-point behavior, an interactive behavior sequence corresponding to the interactive behavior and a single-point behavior sequence corresponding to the single-point behavior are constructed, and the single-point behavior sequence is added to the interactive behavior sequence respectively to obtain at least one candidate behavior sequence.

4. The risk identification method of claim 1, wherein, The step of testing the candidate risk identification strategy at least once based on historical traffic data and current data to be identified includes: Historical traffic data is used to evaluate the candidate risk identification strategies in order to select at least one risk identification strategy to be tested from the candidate risk identification strategies. Obtain the current data to be identified within the current time range, and identify the current object data of at least one object to be identified from the current data to be identified; The risk identification strategy to be tested is used to identify the risk of the object to be identified, so as to obtain the test results of the risk identification strategy to be tested.

5. The risk identification method of claim 4, wherein, The step of evaluating the candidate risk identification strategies using historical traffic data to select at least one risk identification strategy to be tested includes: The candidate risk identification strategy is used to identify the objects in the historical traffic data to obtain a first identification result; Obtain risk object information from the historical traffic data, and rank the candidate risk identification strategies based on the risk object information and the first identification result; Based on the ranking results, at least one risk identification strategy to be tested is identified from the candidate risk identification strategies.

6. The risk identification method of claim 4, wherein, The step of using the risk identification strategy to identify the target object to obtain the test results of the risk identification strategy includes: Based on the current object data, the risk identification strategy to be tested is used to identify the object to be identified, and a second identification result is obtained. The second identification result and the current object data are sent to the verification server so that the verification server can verify the second identification result based on the current object data; The verification result of the second identification result returned by the verification server is received, and the verification result is used as the test result of the risk identification strategy to be tested.

7. The risk identification method according to claim 6, characterized in that, The step of determining the target risk identification strategy corresponding to the behavior sequence from the candidate risk identification strategies based on the test results includes: Based on the test results, the number of correctly identified second recognition results is counted in the second recognition results to obtain the target recognition count; Based on the number of targets identified and the number of objects to be identified, the identification accuracy of the risk identification strategy to be tested within the current time range is determined. Based on the recognition accuracy, at least one risk recognition strategy to be verified is selected from the risk recognition strategies to be verified, and the target risk recognition strategy corresponding to the behavior sequence is obtained.

8. The risk identification method of claim 1, wherein, After employing the target risk identification strategy for risk identification, the method further includes: When the time for risk identification using the target risk identification strategy exceeds a preset time threshold, the target risk identification strategy is detected. Based on the detection results, the current identification accuracy of the target risk identification strategy is determined, and the difference between the current identification accuracy and the historical identification accuracy of the target risk identification strategy is calculated. When the difference is greater than or equal to a preset difference threshold, the target risk identification strategy is stopped from being used for risk identification.

9. The risk identification method of claim 1, wherein, The step of collecting at least one feature data of the sample object under the object's behavior to obtain a training dataset corresponding to the behavior sequence includes: Based on the object behaviors in the behavior sequence, the sample objects are bucketed to obtain at least one target sample object corresponding to the behavior sequence; Collect at least one feature data of the target sample object when it performs the object behavior to obtain the initial feature data of the target sample object; At least one target feature data is selected from the initial feature data to obtain the training dataset corresponding to the behavior sequence.

10. The risk identification method of claim 9, wherein, The step of collecting at least one feature data of the target sample object when performing the object behavior to obtain the initial feature data of the target sample object includes: Collect at least one object data of the target sample object when the object behavior is executed, and obtain the original object data of the target sample object; Sample raw object data less than or equal to a preset data quantity threshold from the raw object data to obtain the target object data of the target sample object; The target object data is analyzed to obtain the initial feature data of the target sample object.

11. The risk identification method of claim 10, wherein, The target object data includes object data of at least one first feature type. The step of performing feature parsing on the target object data to obtain the initial feature data of the target sample object includes: When the first feature type does not include the preset first feature type, the current feature data corresponding to the preset first feature type is generated based on the object attribute information of the target sample object, and the current feature data is added to the target object data to obtain the initial feature data of the target sample object; When the first feature type includes the feature type to be converted, the object data of the feature type to be converted is converted into feature data of a preset second feature type, and the target object data is updated based on the converted feature data to obtain the initial feature data of the target sample object; When the first feature type includes a feature type to be encoded, feature encoding is performed on the object data corresponding to the feature type to be encoded, and the target object data is updated based on the encoded feature data to obtain the initial feature data of the target sample object.

12. The risk identification method of claim 9, wherein, The step of selecting at least one target feature data from the initial feature data to obtain the training dataset corresponding to the behavior sequence includes: Feature data with at least one continuous feature is selected from the initial feature data to obtain continuous feature data, and the continuous feature data is binned to obtain target continuous feature data. The initial feature data is updated based on the target continuous feature data to obtain updated feature data, wherein the updated feature data includes at least one feature data of a second feature type, and the second feature type includes the continuous feature. Target feature data of at least one target feature type is selected from the updated feature data to obtain the training dataset corresponding to the behavior sequence.

13. The risk identification method of claim 12, wherein, The binning of the continuous feature data to obtain the target continuous feature data includes: Based on at least one preset data distribution, the continuous feature data is discretized, and the discrete confidence level corresponding to each preset data distribution is evaluated. Based on the discrete confidence level, a target data distribution is selected from the preset data distribution, and the discrete feature data corresponding to the target data distribution is binned to obtain at least one first feature data bin corresponding to each continuous feature. Based on the feature data in the first feature data box, a first information value parameter of the continuous feature is determined. The first information value parameter is used to measure the contribution of the feature data of the continuous feature to risk identification. Based on the first information value parameter, target continuous feature data corresponding to the target continuous feature is selected from the continuous feature data.

14. The risk identification method of claim 13, wherein, The step of determining the first information value parameter of the continuous feature based on the feature data in the first feature data bin includes: Calculate the proportion of sample objects corresponding to the feature data in the first feature data box, where the proportion of objects includes the proportion of risky objects and the proportion of normal objects; Calculate the ratio between the proportion of risky objects and the proportion of normal objects to obtain the good-to-bad ratio of the first feature data box; Calculate the ratio difference between the proportion of risky objects and the proportion of normal objects, and determine the initial information value parameters of the first feature data box based on the ratio difference and the good-to-bad ratio; The initial information value parameters of the first feature data box corresponding to the continuous feature are fused to obtain the first information value parameter of the continuous feature.

15. The risk identification method of claim 12, wherein, The step of selecting target feature data of at least one target feature type from the updated data features to obtain the training dataset corresponding to the behavior sequence includes: The feature data corresponding to each second feature type in the updated feature data is binned to obtain at least one second feature data bin corresponding to the second feature type. Based on the feature data in the second feature data box, a second information value parameter for the second feature type is determined. The second information value parameter is used to measure the contribution of the feature data of the second feature type to risk identification. Based on the second information value parameter, at least one type of target feature data is selected from the updated feature data, and the target feature data is filtered to obtain the training dataset corresponding to the behavior sequence.

16. The risk identification method of claim 1, wherein, The step of traversing the nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence includes: At least one node is identified in the tree model, the node including a root node and leaf nodes, and the node indicates the feature type corresponding to the feature data in the training dataset; Obtain the execution path of the tree model, and traverse from the root node to the leaf node according to the execution path to obtain at least one node path; Based on the node path, an initial risk identification strategy corresponding to the behavior sequence is determined, wherein the initial risk identification strategy characterizes a combination of feature data of at least one feature type used for risk identification.

17. The risk identification method of claim 16, wherein, The step of selecting at least one candidate risk identification strategy from the initial risk identification strategy based on the training dataset includes: Based on the training dataset, the initial risk identification strategy is used to identify the risk of the sample objects to obtain a third identification result, which includes at least one candidate risk sample object indicated by the leaf node corresponding to the initial risk identification strategy. Obtain the labeled risk category of the candidate risk sample object, and calculate at least one identification performance parameter of the initial risk identification strategy based on the labeled risk category; In the initial risk identification strategy, at least one candidate risk identification strategy whose identification performance parameter is greater than a preset parameter threshold is selected.

18. The risk identification method of claim 1, wherein, The acquisition of object behavior data of the sample object includes: Obtain negative feedback information for at least one object in the object interaction platform, and based on the negative feedback information, filter out at least one risk object from the objects; Collect behavioral data of the risk object within a preset historical time range to obtain risk behavior data; At least one normal object is selected from the normal object set of the object interaction platform, and the behavior data of the normal object is collected within the preset historical time range to obtain normal behavior data; The risky object and the normal object are used as sample objects, and the risky behavior data and the normal behavior data are used as the object behavior data of the sample objects.

19. A risk identification device, characterized in that, include: An acquisition unit is configured to acquire object behavior data of a sample object and, based on the object behavior data, construct at least one behavior sequence, wherein the behavior sequence indicates the sample object having at least one object behavior; The acquisition unit is used to acquire at least one feature data of the sample object under the object behavior to obtain the training dataset corresponding to the behavior sequence; The training unit is used to train at least one preset tree model using the training dataset and to traverse the nodes in the trained tree model to obtain at least one initial risk identification strategy corresponding to the behavior sequence. A filtering unit is used to filter out at least one candidate risk identification strategy from the initial risk identification strategy based on the training dataset. The testing unit is used to test the candidate risk identification strategy at least once based on historical traffic data and current data to be identified; The identification unit is used to determine the target risk identification strategy corresponding to the behavior sequence from the candidate risk identification strategies based on the test results, and to use the target risk identification strategy to perform risk identification.

20. An electronic device, comprising: It includes a processor and a memory, the memory storing an application program, and the processor running the application program within the memory to perform the steps of the risk identification method according to any one of claims 1 to 18.