Abnormal object set determination method and apparatus, and electronic device

By extracting the attribute information of candidate exception objects from multi-source data and performing feature vector clustering, combining the behavior sequence of exception accounts, a collection of associated exception objects is generated, and a problem of low efficiency in determining exception object sets in the prior art is solved, and efficient and adaptive abnormal object set determination is achieved.

CN119961746APending Publication Date: 2025-05-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311505268.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, the efficiency of determining the set of abnormal objects is low, the single-point traffic data analysis method lacks overall environmental perception, while the modeling and analysis method relies on manual processing and cannot adapt to traffic changes.

Method used

By determining candidate exception objects from multi-source data, obtaining their attribute information and converting them into feature vectors for clustering, combining the behavior sequence and similarity of the exception account, a collection of associated exception objects is generated to reduce manual intervention.

Benefits of technology

Improves the determination efficiency of abnormal object collections, reduces manual participation, and adapts to changes in traffic or business concerns without remodeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961746A_ABST
    Figure CN119961746A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for determining an abnormal object set and electronic equipment, and the method comprises the steps: determining candidate abnormal objects which are possibly abnormal and are determined from data sets of different sources; the attribute information corresponding to the mined candidate abnormal object and the attribute information of the object associated with the candidate abnormal object are converted to obtain a feature vector of each candidate abnormal object, and clustering of similar objects is carried out based on the feature vectors; the method comprises the steps of clustering an object set, analyzing an account set corresponding to the clustered object set according to a preset condition to obtain an abnormal account set of which the similarity is greater than a threshold value and the behavior sequence is abnormal, and obtaining an associated abnormal object set by combining an object associated with each account in the abnormal account set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device and electronic device for determining a set of abnormal objects. Background Art

[0002] In related technologies, in order to maintain the healthy development of the network ecology, it is necessary to dig out the abnormal objects (for example, accounts, IP (Internet Protocol), devices, etc.) corresponding to these abnormal behaviors and the abnormal object collection associated with the abnormal objects for abnormal behaviors such as taking advantage of platform / business loopholes to make money, increase traffic, and create public opinion.

[0003] However, the commonly used methods for mining abnormal objects are mainly single-point traffic data analysis and mining methods or modeling and analysis traffic mining methods. Among them, the single-point traffic data analysis and mining methods mainly analyze traffic at single point data, statistically change, and formulate expert rule strategies. They lack the perception of the overall environmental traffic, and the traffic analysis results are not accurate enough, and they cannot efficiently mine the information of the abnormal object set associated with the abnormal object. The modeling and analysis traffic mining method mainly collects traffic data, relies on feature modeling analysis and review after manual processing. Due to the high degree of manual participation, the data analysis consumes a long time. In addition, the analysis models in the prior art are mostly fixed input models. When the traffic changes or the business focus changes, it is necessary to re-model in order to analyze the changed data, but re-modeling also takes a certain amount of time, which will affect the mining and analysis efficiency of the associated abnormal object set.

[0004] It can be seen that the method of determining the abnormal object set in the related art has a technical problem of low efficiency in determining the abnormal object set.

[0005] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0006] The embodiments of the present application provide a method, device and electronic device for determining an abnormal object set, so as to at least solve the technical problem that the method of determining an abnormal object set in the related art has low efficiency in determining the abnormal object set.

[0007] According to one aspect of an embodiment of the present application, a method for determining an abnormal object set is provided, comprising: determining N candidate abnormal objects from data sets from different sources, wherein N is a positive integer greater than or equal to 1; obtaining attribute information of each candidate abnormal object in the N candidate abnormal objects to obtain a first attribute information set, and obtaining attribute information of an object associated with each candidate abnormal object in the N candidate abnormal objects to obtain a second attribute information set; converting the first attribute information set and the second attribute information set into feature vectors to obtain N feature vectors corresponding to the N candidate abnormal objects, and using the N feature vectors to cluster the N candidate abnormal objects to obtain M candidate abnormal objects. The method comprises the following steps: a first step of searching an abnormal account set that satisfies a preset condition and a second step of searching an abnormal account set that satisfies a preset condition; a first step of searching an abnormal account set that satisfies a preset condition and a second step of searching an abnormal account set that satisfies a preset condition; a second step of searching an abnormal account set that satisfies a preset condition and ...;

[0008] According to another aspect of an embodiment of the present application, a device for determining a set of abnormal objects is further provided, comprising: a first determining unit, configured to determine N candidate abnormal objects from data sets from different sources, wherein N is a positive integer greater than or equal to 1; an acquiring unit, configured to acquire attribute information of each of the N candidate abnormal objects to obtain a first attribute information set, and acquire attribute information of an object associated with each of the N candidate abnormal objects to obtain a second attribute information set; a first executing unit, configured to convert the first attribute information set and the second attribute information set into feature vectors to obtain N feature vectors corresponding to the N candidate abnormal objects, and cluster the N candidate abnormal objects using the N feature vectors to obtain to M candidate abnormal object sets, wherein the distance between the feature vectors corresponding to each candidate abnormal object in the j-th candidate abnormal object set in the M candidate abnormal object sets is less than a preset distance threshold, and j is a positive integer greater than or equal to 1 and less than or equal to M; a second execution unit, used to use the M candidate abnormal object sets to find an abnormal account set that meets a preset condition, and use each abnormal account in the abnormal account set and the object associated with each abnormal account to generate an associated abnormal object set, wherein the preset condition includes: the behavior sequence of each abnormal account in the abnormal account set is abnormal, and the similarity between each abnormal account and at least one abnormal account in the abnormal account set is greater than or equal to a preset first similarity threshold.

[0009] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned method for determining a set of abnormal objects when running.

[0010] According to another aspect of the embodiments of the present application, a computer program product is provided, including a computer program / instruction, which implements the steps of the above method when executed by a processor.

[0011] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the method for determining the set of abnormal objects through the computer program.

[0012] Through the above-mentioned embodiments provided by the present application, for objects that may be abnormal and are perceived from multi-source data, similar objects are clustered by combining the attribute information associated with the object and the relevant attribute information of other objects associated with the object, and then the behavior sequence of the account corresponding to each abnormal object is analyzed to obtain a group of similar abnormal accounts, and the objects associated with each account in the abnormal account set are combined to obtain a set of associated abnormal objects. Since the change in traffic or business focus affects the category of abnormal objects, the impact on the information category contained in the attribute information of the abnormal objects is relatively small, and the attribute information of each object associated is used in the analysis and judgment of the abnormal objects. No matter how the traffic or business focus changes, there is no need to re-establish the model. At the same time, since the analysis and judgment of abnormal objects mainly rely on the associated attribute information, there is no need to manually search for a large amount of data, which can reduce the degree of manual participation, thereby achieving the technical effect of improving the efficiency of determining the abnormal object set. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0014] Figure 1 It is a schematic diagram of an application scenario of an optional method for determining a set of abnormal objects according to an embodiment of the present application;

[0015] Figure 2 is a flowchart of an optional method for determining a set of abnormal objects according to an embodiment of the present application;

[0016] Figure 3 is a schematic diagram of an optional method for determining attribute information of an abnormal object according to an embodiment of the present application;

[0017] Figure 4 is a schematic diagram of an optional method for determining a set of abnormal objects according to an embodiment of the present application;

[0018] Figure 5 is a schematic diagram of another optional method for determining a set of abnormal objects according to an embodiment of the present application;

[0019] Figure 6 is a flowchart of another optional method for determining a set of abnormal objects according to an embodiment of the present application;

[0020] Figure 7 is a flowchart of another optional method for determining a set of abnormal objects according to an embodiment of the present application;

[0021] Figure 8is a schematic diagram of an optional method for determining a set of abnormal accounts according to an embodiment of the present application;

[0022] Fig. 9 is a flowchart of another optional method for determining a set of abnormal objects according to an embodiment of the present application;

[0023] Fig.10 is a schematic diagram of an optional method for dividing an abnormal object graph according to an embodiment of the present application;

[0024] Fig.11 is a schematic diagram of another optional method for determining an abnormal object association set according to an embodiment of the present application;

[0025] Fig.12 is a schematic structural diagram of an optional device for determining a set of abnormal objects according to an embodiment of the present application;

[0026] Fig.13 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] The technical solutions in the embodiments of the present application will comply with legal provisions during implementation. When performing operations according to the technical solutions in the embodiments, the data used will not involve user privacy. While ensuring that the operation process is compliant and legal, the security of the data is guaranteed.

[0030] According to one aspect of the embodiments of the present application, a method for determining an abnormal object set is provided. As an optional implementation, the above-mentioned method for determining an abnormal object set can be applied to, but is not limited to, Figure 1 The application scenario shown in Figure 1 In the application scenario shown, the terminal device 102 may, but is not limited to, communicate with the server 106 via the network 104, and the server 106 may, but is not limited to, perform operations on the database 108, such as write data operations or read data operations. The terminal device 102 may, but is not limited to, include a display, a processor, and a memory. The display may, but is not limited to, be used to display relevant information of the abnormal object collection information on the terminal device 102. The processor may, but is not limited to, be used to send object attribute information to the server 106. The memory is used to store relevant processing data.

[0031] As an optional method, the following steps in the method for determining a set of abnormal objects can be executed on the server 106: step S102, determining candidate abnormal objects based on the received data set; step S104, obtaining attribute information corresponding to the selected abnormal objects; step S106, analyzing and processing the attribute information to obtain an abnormal account set and an abnormal object set.

[0032] By adopting the above method, for objects that may be abnormal and are perceived from multi-source data, similar objects are clustered by combining the attribute information associated with the object and the relevant attribute information of other objects associated with the object, and then the behavior sequence of the account corresponding to each abnormal object is analyzed to obtain a group of similar abnormal accounts, and the objects associated with each account in the abnormal account set are combined to obtain a set of associated abnormal objects. Since the changes in traffic or business focus affect the category of abnormal objects, the impact on the information category contained in the attribute information of abnormal objects is relatively small, and the attribute information of each associated object is used in the analysis and judgment of abnormal objects. No matter how the traffic or business focus changes, there is no need to re-establish the model. At the same time, since the analysis and judgment of abnormal objects mainly rely on the associated attribute information, there is no need to manually search for a large amount of data, which can reduce the degree of manual participation, thereby achieving the technical effect of improving the efficiency of determining the abnormal object set.

[0033] In order to solve the problem of low efficiency in determining the above abnormal object set, a method for determining the abnormal object set is proposed in an embodiment of the present application. Figure 2 1 is a flow chart of an optional method for determining a set of abnormal objects according to an embodiment of the present application, the flow chart comprising the following steps:

[0034] Step S202, determining N candidate abnormal objects from data sets from different sources, where N is a positive integer greater than or equal to 1;

[0035] Step S204, acquiring attribute information of each candidate abnormal object among the N candidate abnormal objects to obtain a first attribute information set, and acquiring attribute information of an object associated with each candidate abnormal object among the N candidate abnormal objects to obtain a second attribute information set;

[0036] Step S206, converting the first attribute information set and the second attribute information set into feature vectors to obtain N feature vectors corresponding to the N candidate abnormal objects, and clustering the N candidate abnormal objects using the N feature vectors to obtain M candidate abnormal object sets, wherein the distance between the feature vectors corresponding to each candidate abnormal object in the j-th candidate abnormal object set in the M candidate abnormal object sets is less than a preset distance threshold, and j is a positive integer greater than or equal to 1 and less than or equal to M;

[0037] Step S208, using M candidate abnormal object sets to search for an abnormal account set that meets preset conditions, and using each abnormal account in the abnormal account set and the objects associated with each abnormal account to generate an associated abnormal object set, wherein the preset conditions include: the behavior sequence of each abnormal account in the abnormal account set is abnormal, and the similarity between each abnormal account and at least one abnormal account in the abnormal account set is greater than or equal to a preset first similarity threshold.

[0038] The method for determining the abnormal object set in this embodiment can be applied to the scenario of analyzing and mining fraudulent traffic on the Internet platform. Here, fraudulent traffic can refer to the traffic generated by the black and gray industries taking advantage of platform / business loopholes to commit fraud, such as stealing wool, brushing volume, creating public opinion, etc. The traffic can be Internet data related to the platform / business, such as account, IP (Internet Protocol), device, etc.

[0039] In related technologies, the analysis and mining of fraudulent traffic mainly rely on single-point traffic data analysis and mining methods and modeling analysis traffic mining methods. Among them, the single-point traffic data analysis and mining method mainly analyzes traffic of a single dimension (such as IP), statistically changes, and formulates expert rule strategies. The modeling analysis traffic mining method mainly collects traffic data, processes it manually, and relies on feature modeling analysis and review.

[0040] However, the single-point traffic data analysis and mining process lacks the perception of the overall environmental traffic, and in the early stage of data mining, manual data search is required, and modeling analysis is performed after manual data processing. The analysis results need to be manually reviewed, and the entire mining process takes a long time. In addition, the models relied on by the analysis methods in the above methods are mostly fixed input models, and the models cannot be adaptively adjusted to cope with changes caused by traffic changes or business focus switching, and often require re-modeling.

[0041] In order to at least partially solve the above problems, in this embodiment, for objects that may be perceived to be abnormal, clustering can be performed in combination with the attribute information associated with the object and the relevant attribute information of other objects associated with the object to determine candidate abnormal objects with relatively similar attribute information and the account corresponding to each candidate abnormal object. At the same time, by analyzing the behavior sequence of each account and combining the similarity calculation between the accounts, a similar abnormal account set can be further screened out, and then combined with the objects associated with the abnormal account set, an abnormal object set can be mined. The abnormal object set here can be a set of objects with a large correlation. In other words, the abnormal object set can be abnormal objects belonging to an organization, and in this set, the abnormal objects can have the same or similar abnormal behaviors or behavioral purposes.

[0042] The above-mentioned similarity calculation between accounts can effectively solve the many complex factors existing in actual business (such as different black market organizations using the same or similar black market tools, account buying and selling transactions reuse, and the use of tools such as virtual devices, etc.), which cause objects in different organizations to have relatively similar attribute information characteristics, and thus lead to the problem that multiple accounts in the determined abnormal account set do not actually belong to the same organization.

[0043] It should be noted that if Figure 3 As shown, taking the candidate abnormal object as an abnormal individual as an example, the data sets from different sources mentioned above may include data of the environmental dimension corresponding to each object, such as login IP, login time, login device, login account, etc. Through the acquired data sets from different sources, individuals that may be abnormal (i.e., candidate abnormal objects) can be sensed, and then based on the sensed abnormal individuals, relevant information of the attribute dimensions corresponding to the abnormal individuals can be obtained, such as account information, registration information, login information, activity information, transaction information, relationship information, etc.

[0044] It should be noted that the abnormal object in this embodiment can be a fraudulent traffic individual, that is, all single points related to fraudulent traffic in the entire business traffic. Whatever single point is needed for business analysis and detection, the fraudulent traffic individual is what it is. Traffic data can refer to data that is analyzed and integrated to generate various data forms with business value, such as indicators, content, statistics, tags, reports, etc.

[0045] The above-mentioned data sets from different sources may include, but are not limited to, multi-source data obtained from open source data on the Internet, business data (i.e., the platform's own data) and third-party data (i.e., external cooperation data).

[0046] The candidate abnormal objects mentioned above can be objects marked as abnormal in network data. They can be objects with abnormal behaviors on the Internet platform or data generated on the Internet. Objects, i.e. individuals, can be abnormal accounts, abnormal devices, abnormal IP addresses, abnormal mobile phone numbers, etc.

[0047] For example, if an account browses pages 24 hours a day, then this account may be a candidate abnormal object. Or if a device logs in to multiple accounts at the same time, then this device may be a candidate abnormal object. In addition, if there is a large amount of request traffic under an IP, then this IP may also be considered a candidate abnormal object.

[0048] In this embodiment, the N candidate abnormal objects may be objects marked as abnormal in the network data. Here, marking abnormalities may be based on the objects that need to be marked determined based on information obtained from the outside (such as feedback from relevant regulatory units, information reported by users, etc.), or based on information analyzed from within the platform (analysis of the business indicator dashboard within the relevant platform, as well as abnormal detection risk control models plus information obtained through manual review, etc.).

[0049] The above-mentioned attribute information can be information of dimensions associated with the object that can be queried on the Internet, and can include the identity information of the object and various activity information performed by the object on the platform, such as account attributes, IP registration and filing information, device login information, business activity information, business payment transaction information, etc. In addition, the attribute information of each candidate abnormal object can also include relationship information. The business activity information here can be the activities performed by the object on the Internet platform. Taking the object as an account as an example, the actions of adding shopping carts, browsing products, collecting, liking, etc. performed by the account on the website can be considered as the activities of the account on the platform. Relationship information can refer to the information of other objects associated with each candidate abnormal object. It should be noted that the objects associated with each candidate abnormal object can be determined based on relationship information.

[0050] When the first attribute information set and the second attribute information set are determined, the first attribute information and the second attribute information corresponding to each candidate abnormal object can be converted into a feature vector, and similarity clustering is performed through the feature vector to obtain M candidate abnormal object sets. It can be considered that the correlation between the candidate abnormal objects in each candidate abnormal object set is greater than a certain threshold.

[0051] The clustering using the feature vectors can be clustering based on the distance between the feature vectors. The clustering method can be a K-means (a partitioning clustering method), a DBSCAN (a density-based clustering method), or a combination of K-means and DBSCAN, that is, the same clusters in the K-means clustering result and the DBSCAN clustering result are determined as the candidate abnormal object set.

[0052] For the determined M candidate abnormal object sets, the account corresponding to each set can be determined based on the objects in each set. When the objects in each candidate abnormal object set are objects other than account numbers, the account set corresponding to the candidate abnormal object set can be determined based on the account number associated with each abnormal object in the candidate abnormal object set. When the objects in the candidate abnormal object set are account numbers, the candidate abnormal object set is the account number set.

[0053] Considering that there may be normal accounts and abnormal accounts in the M account sets corresponding to the M candidate abnormal object sets, and there may be accounts with less actual correlation in each set, in this embodiment, the accounts in the M account sets can be screened, and the similarity of the screened accounts can be calculated again. The screening of normal accounts and abnormal accounts in the account set can be based on the behavior sequence corresponding to the account. It should be noted that the behavior sequence of each account can be displayed in the form of a set of strings, and different behaviors correspond to different identification strings.

[0054] For example, the behaviors of an account on a video platform may include logging in, viewing videos, liking, commenting, swiping away, etc. Different behaviors correspond to different identification strings. The behavior sequence of an account can be a combination of multiple strings in the order in which the behaviors occur.

[0055] After the abnormal accounts are identified by screening the abnormal accounts and normal accounts in the account set, the similarities between the abnormal accounts in each set are determined respectively to find abnormal accounts whose similarity is greater than or equal to the first preset threshold, and finally obtain the abnormal account set.

[0056] Taking the IP address of an abnormal individual as an example, through the behavior sequence analysis of the account, a set of suspicious accounts is mined. Combined with the information corresponding to the account, the output data corresponding to the abnormal individual can be sorted and generated. The style format of the output data can be shown in Table 1.

[0057] Table 1

[0058]

[0059] Based on the determined set of abnormal accounts, abnormal objects belonging to an organization can be mined according to the abnormal accounts and other objects associated with each abnormal account. Figure 4 As shown, taking the perceived candidate abnormal object as a mobile phone number and the abnormal behavior as a malicious ordering behavior as an example, through the aforementioned method of determining the abnormal account set and the associated abnormal object set, after the abnormal mobile phone number is perceived, the abnormal account set corresponding to the malicious ordering behavior can be further determined, and then other objects associated with the abnormal behavior (such as equipment, platform business accounts, and IP, etc.) can be mined, and the connection relationship of the abnormal object set can be determined through information such as the IP corresponding to the abnormal event and the order address of the account, and the address corresponding to the abnormal event and the mobile phone number information of the abnormal event can be further mined.

[0060] Taking into account that some abnormal individuals are abnormal events carried out in the form of organizations, based on determining the associated abnormal object set in the manner in the above embodiment, as shown in FIG. Figure 5 As shown, the organizer account corresponding to the associated abnormal object set and the account of the related industrial chain of the abnormal event on the platform can be additionally associated. Based on the analysis of the relevant information of the account, the industrial chain resources can be further determined, such as account merchants specializing in account transactions, accounts that provide technical support for renting and selling servers, payment channel accounts involved in abnormal transactions, and accounts responsible for providing drainage services for abnormal businesses.

[0061] Based on the determined set of abnormal objects, an abnormal analysis result can be generated, such as Figure 6 As shown in the figure, the steps of generating abnormal analysis results can be divided into four modules: data acquisition, data processing, data analysis, and result output. Data acquisition refers to multi-source data obtained from Internet open source data, business data, and third-party data. By integrating data from different sources, multi-source heterogeneous data fusion governance can be achieved. Based on the data sets obtained from different sources, objects that may have abnormalities can be determined, that is, the N candidate abnormal objects mentioned above.

[0062] Data processing mainly refers to processing the data corresponding to the determined candidate abnormal objects. The core of data processing is to select high-value data from all the data corresponding to the candidate abnormal objects for analysis. It can include two submodules, namely, a module for automatically analyzing and processing the collected raw data through automatic sensing technology, and a module for processing key business hotspot data. By integrating the data output by the two modules, high-value data required for abnormal object analysis can be mined. The data processing process can correspond to the processing processes of obtaining the attribute information of each candidate abnormal object and the attribute information of the object associated with each candidate abnormal object, as well as determining the feature vector corresponding to the attribute information in the aforementioned embodiment.

[0063] Data analysis mainly analyzes and mines candidate abnormal objects based on attribute information, that is, determines the correlation between data, combines the data features corresponding to the attribute information, analyzes and calculates through the corresponding strategy model, classifies the data, and obtains data related to abnormal individuals and organizations (including the abnormal account set associated with the abnormal individual, and the related data of the abnormal object set). In this process, the data can also be further screened in combination with the expert manual analysis experience (that is, the screening conditions formulated in combination with manual experience).

[0064] Result output, that is, output the analyzed result. Here, the result can be Figure 4 The output in the format shown can also be used to output the final analysis results and other related information in the form of data tables, analysis reports, etc. The data obtained from the aforementioned analysis process can also be the result of processing the data generated by the analysis according to business needs, such as outputting professional report effect evaluation, multiple abnormal indicators, abnormal situation awareness, and recording the abnormal data into the database, and summarizing it as existing data in the abnormal traffic database.

[0065] Through the above steps S202 to S208, N candidate abnormal objects are determined from data sets from different sources, where N is a positive integer greater than or equal to 1; attribute information of each candidate abnormal object in the N candidate abnormal objects is obtained to obtain a first attribute information set, and attribute information of objects associated with each candidate abnormal object in the N candidate abnormal objects is obtained to obtain a second attribute information set; the first attribute information set and the second attribute information set are converted into feature vectors to obtain N feature vectors corresponding to the N candidate abnormal objects, and the N candidate abnormal objects are clustered using the N feature vectors to obtain M candidate abnormal object sets, where each candidate in the j-th candidate abnormal object set in the M candidate abnormal object set has a first attribute information set. The distance between the feature vectors corresponding to the selected abnormal objects is less than a preset distance threshold, j is a positive integer greater than or equal to 1 and less than or equal to M; the M candidate abnormal object sets are used to find an abnormal account set that meets the preset conditions, and each abnormal account in the abnormal account set and the object associated with each abnormal account are used to generate an associated abnormal object set, wherein the preset conditions include: the behavior sequence of each abnormal account in the abnormal account set is abnormal, and the similarity between each abnormal account and at least one abnormal account in the abnormal account set is greater than or equal to a preset first similarity threshold, which solves the technical problem that the determination method of the abnormal object set in the related art has a low efficiency in determining the abnormal object set, and improves the efficiency in determining the abnormal object set.

[0066] As an optional example, obtaining attribute information of each candidate abnormal object among N candidate abnormal objects to obtain a first attribute information set, and obtaining attribute information of an object associated with each candidate abnormal object among the N candidate abnormal objects to obtain a second attribute information set, including:

[0067] S11, acquiring attribute information of one or more dimensions of each of the N candidate abnormal objects, and obtaining N groups of self attribute information, wherein the first attribute information set includes the N groups of self attribute information, the i-th group of self attribute information in the N groups of self attribute information includes the attribute information of one or more dimensions of the i-th candidate abnormal object in the N candidate abnormal objects, and i is a positive integer greater than or equal to 1 and less than or equal to N;

[0068] S12, obtaining attribute information of one or more dimensions of one or more objects associated with each of the N candidate abnormal objects, and obtaining N groups of associated attribute information, wherein the second attribute information set includes N groups of associated attribute information, and the i-th group of associated attribute information in the N groups of associated attribute information includes attribute information of one or more dimensions of one or more objects associated with the i-th candidate abnormal object in the N candidate abnormal objects.

[0069] In this embodiment, the acquired attribute information of each candidate abnormal object may be attribute information of one or more dimensions, and the attribute information of the object associated therewith may also be attribute information of one or more dimensions.

[0070] For N candidate abnormal objects, N groups of attribute information can be obtained, each group of attribute information corresponds to a candidate abnormal object, and each group of attribute information can be attribute information of one dimension or multi-dimensional attribute information. Considering that each candidate abnormal object can be associated with one or more objects (e.g., the candidate abnormal object is an IP, which can be associated with one or more accounts, devices, etc.), the second attribute information of the N candidate abnormal objects can be N groups of associated attribute information, and each group of associated attribute information can be dimensional information of one or more objects associated with each candidate abnormal object.

[0071] Through this embodiment, objects that may have abnormalities are associated with their attribute information and the attribute information of their associated objects, and analyzed from multi-dimensional factors, which can improve the comprehensiveness and accuracy of abnormal object mining.

[0072] As an optional example, the first attribute information set and the second attribute information set are converted into feature vectors to obtain N feature vectors corresponding to the N candidate abnormal objects, including:

[0073] S21, obtaining content information included in each group of self-attribute information in N groups of self-attribute information, and obtaining N groups of self-content information, wherein the i-th group of self-attribute information in the N groups of self-attribute information includes attribute information of one or more dimensions of the i-th candidate abnormal object in the N candidate abnormal objects, i is a positive integer greater than or equal to 1 and less than or equal to N, and the i-th group of self-content information of the i-th candidate abnormal object is used to describe the i-th candidate abnormal object;

[0074] S22, when the second attribute information set includes N groups of associated attribute information, obtain content information included in each group of associated attribute information in the N groups of associated attribute information to obtain N groups of associated content information, wherein the i-th group of associated attribute information in the N groups of associated attribute information includes attribute information of one or more dimensions of one or more objects associated with the i-th candidate abnormal object among the N candidate abnormal objects, and the i-th group of associated content information of the one or more objects associated with the i-th candidate abnormal object is used to describe the one or more objects associated with the i-th candidate abnormal object;

[0075] S23, perform vector concatenation on the feature vectors corresponding to the N groups of own content information and the feature vectors corresponding to the N groups of associated content information to obtain N feature vectors corresponding to the N candidate abnormal objects, wherein the i-th feature vector among the N feature vectors is used to represent the i-th candidate abnormal object among the N candidate abnormal objects, and the i-th feature vector among the N feature vectors is a feature vector determined based on the i-th group of own content information among the N groups of own content information and the i-th group of associated content information among the N groups of associated content information.

[0076] In this embodiment, content information can be extracted from the acquired attribute information to perform clustering operations based on the features corresponding to the content information. The content information here can be information directly obtained from the information associated with the object, and can include information used to describe the identity of the object. The content information can be divided into text information and image information. Taking the object as an account as an example, the content information can include the account's nickname, avatar, comments sent by the account, dissemination links, etc.

[0077] When the first attribute information set includes N sets of own attribute information and the second attribute information set includes N sets of associated attribute information, N sets of own content information and N sets of associated content information can be extracted. By converting the N sets of own content information into feature vectors and converting the N sets of associated content information into feature vectors, the feature vectors corresponding to the own content information and the feature vectors corresponding to the associated content information of each candidate abnormal object are concatenated respectively, and N feature vectors corresponding to the N candidate abnormal objects can be obtained, which are then used for feature clustering.

[0078] Optionally, the feature vectors corresponding to the N groups of own content information and the feature vectors corresponding to the N groups of associated content information are concatenated to obtain N feature vectors corresponding to the N candidate abnormal objects, which can be implemented in the following manner:

[0079] The i-th feature vector among N feature vectors is determined by the following steps: each content information in the i-th group of own content information is converted into a corresponding feature vector to obtain a first group of feature vectors; each content information in the i-th group of associated content information is converted into a corresponding feature vector to obtain a second group of feature vectors; the first group of feature vectors and the second group of feature vectors are concatenated to obtain the i-th feature vector.

[0080] like Figure 7 As shown, each content information can be divided into account information, behavior information and relationship information. Each type of information corresponds to a feature vector. The feature information corresponding to different information can be spliced. The content information of each candidate abnormal object itself and the content information of its associated objects can also be spliced.

[0081] Optionally, when the content information is divided into image information and text information, the corresponding feature vector can be obtained by using an embedding technique, such as using a Resnet (i.e., Residual Network, a deep convolutional neural network architecture) model to obtain image features, and using a pre-trained BERT (Bidirectional Encoder Representations from Transformers) model to obtain text features, and then concatenating the obtained image feature vector and text feature vector. In addition, a multi-input network can also be used to input image information and text information at the same time to obtain image feature vectors and text feature vectors, and then concatenate the feature vectors.

[0082] For the concatenated feature vector, such as Figure 7 As shown, it can be input into a DNN model (Deep Neural Network). In the neural network, the feature vector is processed and converted by multiple layers to obtain a feature vector containing different types of features. The clustering in this embodiment can be clustering of the feature vector corresponding to each candidate abnormal object output by the neural network.

[0083] Through this embodiment, by performing feature fusion on the feature vectors of the attribute information associated with each object, and then clustering the candidate abnormal objects according to the feature vectors containing different types of feature information, the accuracy of the clustering result can be improved.

[0084] As an optional example, using M candidate abnormal object sets to search for abnormal account sets that meet preset conditions includes:

[0085] S31, when a target abnormal object set in the M candidate abnormal object sets meets a preset screening condition, determining a target account set corresponding to the target abnormal object set;

[0086] S32, obtaining a behavior sequence of each account in the target account set, and determining an account in the target account set with an abnormal behavior sequence as an abnormal account;

[0087] S33, when a group of abnormal accounts is determined in the target account set, determine the abnormal account set in the group of abnormal accounts.

[0088] Considering that there may be normal objects in the candidate abnormal object set, for multiple candidate abnormal objects in the candidate abnormal object set, the candidate abnormal objects and normal objects can also be analyzed and judged in combination with preset screening conditions. Candidate abnormal objects that meet the preset screening conditions can be considered as abnormal objects.

[0089] The above-mentioned screening conditions may be one or more rules in the business rule set corresponding to the currently determined candidate abnormal object. Figure 7 As shown, attribute information can be divided into account attributes, behavior attributes, and relationship attributes. Each attribute can correspond to multiple policy rules. Different policy rules can be combined together to obtain multiple policy combinations. If a candidate abnormal object can hit at least one policy combination among multiple policy combinations, it can be considered that the candidate abnormal object is an abnormal object, that is, the abnormal object meets the preset screening conditions.

[0090] Optionally, a risk control model (such as a scoring card, a decision tree, etc.) may also be used to determine the screening conditions.

[0091] Optionally, after determining the clustering results, the objects contained in each cluster obtained by clustering can be analyzed and judged in turn in combination with preset filtering conditions. Alternatively, after determining the first attribute information set and the second attribute information set, clustering and filtering conditions can be judged based on the two attribute information sets at the same time. Finally, the clustering results and the filtering results can be combined to determine the target abnormal object set.

[0092] like Figure 7 As shown, the captured candidate abnormal individuals are associated with relevant attribute dimensions, and the attribute dimension information of other individuals related to these individuals is recorded. Through clustering and analysis and judgment of rule strategy sets, relatively similar abnormal individuals can be obtained, and then the account dimension of the platform / business can be obtained. Through the behavior sequence analysis of the account, an abnormal account set can be obtained, thereby outputting the abnormal individual data corresponding to the abnormal object. The information contained in the abnormal individual data can be as shown in the above Table 1, and this embodiment will not be repeated here.

[0093] Optionally, the above-mentioned target abnormal object set that meets the preset screening conditions may refer to a set in which the number of candidate abnormal objects that meet at least one of the preset screening conditions accounts for a proportion of the total number of objects in the set that is greater than a certain threshold.

[0094] It should be noted that after determining the target abnormal object set, if the objects in the target abnormal object set are objects other than account numbers, the target account set corresponding to the target abnormal object set can be determined based on the account number associated with each abnormal object in the target abnormal object set. If the objects in the target abnormal object set are account numbers, the target abnormal object set is the target account number set.

[0095] Through this embodiment, the clustered results are further screened using screening conditions, which can improve the accuracy of the determined candidate abnormal object set.

[0096] As an optional example, before determining the target account set corresponding to the target abnormal object set, the method further includes:

[0097] S41, in the case where the first attribute information set includes N groups of own attribute information and the second attribute information set includes N groups of associated attribute information, the N groups of own attribute information and the N groups of associated attribute information are combined into N attribute information sets, wherein the i-th group of own attribute information in the N groups of own attribute information includes attribute information of one or more dimensions of the i-th candidate abnormal object in the N candidate abnormal objects, the i-th group of associated attribute information in the N groups of associated attribute information includes attribute information of one or more dimensions of one or more objects associated with the i-th candidate abnormal object in the N candidate abnormal objects, i is a positive integer greater than or equal to 1 and less than or equal to N, and the i-th attribute information set in the N attribute information sets includes the i-th group of own attribute information and the i-th group of associated attribute information;

[0098] S42, performing statistics on information in each of the N attribute information sets to obtain N statistical information, wherein the i-th statistic in the N statistical information corresponds to the i-th candidate abnormal object in the N candidate abnormal objects;

[0099] S43, using N statistical information, determine a target abnormal object set from M candidate abnormal object sets, wherein the number of candidate abnormal objects whose statistical information meets the screening rule in the target abnormal object set is greater than or equal to a preset number threshold, and the screening condition includes the preset screening rule and the preset number threshold, or, the proportion of candidate abnormal objects whose statistical information meets the screening rule in the target abnormal object set is greater than or equal to a preset proportion threshold, and the screening condition includes the preset screening rule and the preset proportion threshold.

[0100] When judging whether an object meets the preset screening conditions, statistical information can be extracted from the object's attribute information and used for judgment. The statistical information here can be information indirectly obtained from the information associated with the object, that is, information obtained after a certain statistical analysis of the associated information, which can include the object's activity information and behavior data. Taking the object as an account as an example, the statistical information can include the number of devices that the account has logged in, the number of times the account has performed a certain behavior in the past 7 days, the number of times the account has been abnormally marked this month, the number of orders paid by the account, the number of coupons received by the account, etc.

[0101] For each determined candidate abnormal object, statistical information can be extracted from its attribute information and the attribute information of other related objects. The statistical information can be divided into account attributes, behavior attributes and relationship attributes, and each attribute can correspond to different filtering conditions. When different types of attribute information of an object meet different types of filtering conditions or at least meet one type of filtering condition, it can be determined that the object meets the preset filtering condition.

[0102] For the N attribute information of N candidate abnormal objects, each candidate abnormal object set in the M candidate abnormal object sets obtained after clustering can be directly selected and judged using the preset screening conditions respectively.

[0103] like Figure 7 As shown, taking the preset filtering conditions as rules as an example, account attributes can correspond to multiple policy rules, behavior attributes can correspond to multiple policy rules, and relationship attributes can also correspond to multiple policy rules. The statistical information of an object can hit at least one policy rule in multiple different attribute rules at the same time. Therefore, if the statistical information of an object can hit one or more of the policy combinations 1, ..., n, it can be considered that the object meets the preset filtering conditions and is an abnormal object. If the statistical information of an object meets one or more of the filtering policy rules 1, ..., n, it can be considered that the object can be directly filtered.

[0104] Optionally, using statistical information corresponding to each candidate abnormal object set in the M candidate abnormal object sets in the N statistical information to analyze whether each candidate abnormal object set meets a preset screening condition, including:

[0105] The following steps are used to determine whether the target abnormal object set in the M candidate abnormal object sets meets the preset screening conditions:

[0106] In the case where the target abnormal object set includes K candidate abnormal objects, determining K statistical information corresponding to the K candidate abnormal objects from the N statistical information, where K is a positive integer greater than or equal to 2;

[0107] Determine whether each of the K pieces of statistical information satisfies a screening rule included in the screening condition;

[0108] When the statistical information corresponding to P candidate abnormal objects among the K candidate abnormal objects meets the screening rules, and P is greater than or equal to the preset quantity threshold, or P / K is greater than or equal to the preset ratio threshold, it is determined that the target abnormal object set meets the preset screening conditions.

[0109] It should be noted that, among the M clustered candidate abnormal object sets, as long as the ratio of the number of candidate abnormal objects in the set that meet the preset screening conditions to the total number of objects in the set is greater than the preset ratio threshold, the set can be considered as the target abnormal object set.

[0110] Through this embodiment, the object's own attribute information and the statistical information in the attribute information of other objects associated with it are used to determine whether the object meets the screening condition to determine the abnormal object set, which can improve the accuracy of determining the abnormal object set.

[0111] As an optional example, determining a target account set corresponding to a target abnormal object set includes:

[0112] S51, when the target abnormal object set includes K candidate abnormal objects, the p-th group account in K groups of accounts corresponding to the p-th candidate abnormal object among the K candidate abnormal objects is determined by the following steps, wherein K is a positive integer greater than or equal to 2, p is a positive integer greater than or equal to 1 and less than or equal to K, and the target account set includes K groups of accounts:

[0113] In the case where the pth candidate abnormal object is an account, determining the pth group of accounts to include the pth candidate abnormal object;

[0114] In the case that the p-th candidate abnormal object is not an account, the account associated with the p-th candidate abnormal object is obtained, and the p-th group of accounts is determined to include the account associated with the p-th candidate abnormal object.

[0115] In this embodiment, after determining the target abnormal object set, if the objects in the target abnormal object set are accounts, the target account set is the target abnormal object set; if the objects in the target abnormal object set are not accounts, the target account set is the accounts associated with each object in the target abnormal object set.

[0116] It should be noted that the abnormal object set in this embodiment can be a set composed of objects of the same class, that is, the aforementioned clustering and analysis and judgment of the screening conditions for the candidate abnormal objects can be performed on the same class of candidate abnormal objects. In addition, the abnormal object set in this embodiment can also be a set composed of objects of different classes.

[0117] In the case where the objects in the target abnormal object set are not accounts, the accounts directly associated with each object in the target abnormal object set can be used as accounts in the target account set. Here, the accounts directly associated with the objects can mean that the objects and accounts do not need to be connected through a third object.

[0118] Optionally, when the p-th candidate abnormal object is not an account, obtaining an account associated with the p-th candidate abnormal object, and determining the p-th group of accounts to include the account associated with the p-th candidate abnormal object, comprises:

[0119] In the case where the pth candidate abnormal object is an IP address, obtaining an account using the IP address, and determining the pth group of accounts to include the account using the IP address;

[0120] When the pth candidate abnormal object is a device, the account of the device logged in is obtained, and the pth group of accounts is determined to include the account of the device logged in.

[0121] It should be noted that the above-mentioned candidate abnormal object can also be a communication identification number, such as a mobile phone number. In the case where the pth candidate abnormal object is a communication identification number, the account registered by the communication identification number is obtained, and the pth group of accounts is determined to include the account registered by the communication identification number.

[0122] Through this embodiment, the corresponding account set is determined based on the determined abnormal object set. Regardless of whether the first perceived possible abnormal object is an account, the set of abnormal accounts can be mined, which can improve the efficiency of mining abnormal accounts.

[0123] As an optional example, obtaining a behavior sequence of each account in the target account set, and determining an account with an abnormal behavior sequence in the target account set as an abnormal account, includes:

[0124] S61, when the target account set includes Q accounts, obtain a behavior sequence of each account in the Q accounts, wherein Q is a positive integer greater than or equal to 1, and the behavior sequence of each account in the Q accounts includes a time behavior sequence and / or a frequency behavior sequence, the time behavior sequence of the qth account in the Q accounts represents a group of operations performed by the qth account in chronological order, the frequency behavior sequence of the qth account in the Q accounts represents the number of times each operation in a group of operations performed by the qth account is performed, and q is a positive integer greater than or equal to 1 and less than or equal to Q;

[0125] S62, determining whether the behavior sequence of each of the Q accounts is abnormal;

[0126] S63, determining the accounts with abnormal behavior sequences among the Q accounts as abnormal accounts.

[0127] In this embodiment, the behavior sequence of each account may be the behavior performed by each account within a period of time, and may include a time behavior sequence, a frequency behavior sequence, or both. The time behavior sequence may be a behavior ID (Identity Document) corresponding to a set of operations (i.e., behaviors) of an account displayed in chronological order, and the frequency behavior sequence may be a spectrum behavior sequence (i.e., frequency behavior sequence) constructed by transforming the time behavior sequence into the time domain and the frequency domain.

[0128] like Figure 8 As shown, the horizontal axis of the time behavior sequence is time, and the vertical axis is the behavior ID. The horizontal axis of the frequency behavior sequence is the behavior ID, and the vertical axis is the behavior frequency. Pull out the time behavior sequence for the associated suspicious account. Each account corresponds to a time behavior sequence. Through sequence analysis, abnormal time behavior sequences can be found. Then convert them into frequency behavior sequences. Each account corresponds to a frequency behavior sequence. Through sequence analysis, abnormal frequency behavior sequences can be found.

[0129] It should be noted that when determining whether the behavior sequence of each of the Q accounts is abnormal, if the time behavior sequence and frequency behavior sequence of each account are determined, the time behavior sequence and the frequency behavior sequence can be analyzed for abnormal behavior separately, and the analysis can be performed separately by two models.

[0130] The two models for abnormal behavior analysis of time behavior series and frequency behavior series can be pre-trained models. When training the models, the data of normal behavior series in the business data are collected as positive samples, and the data of abnormal behavior series are collected as negative samples. The two models are trained respectively through the time behavior series and frequency behavior series corresponding to the positive and negative samples. The trained model can directly output relevant indication information of whether the input data is abnormal behavior.

[0131] Optionally, determining whether the behavior sequence of each of the Q accounts is abnormal includes:

[0132] Determine whether the time behavior sequence of the qth account among Q accounts is abnormal by following the steps below:

[0133] Obtaining a first similarity between the time behavior sequence of the qth account and a preset abnormal time behavior sequence;

[0134] When the first similarity is greater than or equal to a preset second similarity threshold, it is determined that the time behavior sequence of the qth account is abnormal.

[0135] It should be noted that the first similarity between the time behavior sequence of the qth account and the preset abnormal time behavior sequence may refer to the number of sequence segments in the time behavior sequence of the qth account that are identical to the preset abnormal behavior sequence, or may refer to the ratio of the number of sequence segments in the time behavior sequence of the qth account that are identical to the preset abnormal behavior sequence to the total number of segments in the time behavior sequence of the qth account.

[0136] The above sequence segment may be a segment composed of several adjacent behavior IDs. For example, the sequence segment corresponding to a normal account is: login - browse video 1 - swipe - browse video 2 - like / comment - swipe - browse video 3, while the sequence segment corresponding to an abnormal account is: login - browse video 1 - like - comment - browse video 2 - like - comment.

[0137] When the first similarity is greater than or equal to a preset second similarity threshold, it can be considered that the time behavior series is abnormal.

[0138] Optionally, determining whether the behavior sequence of each of the Q accounts is abnormal includes:

[0139] Determine whether the frequency behavior sequence of the qth account among Q accounts is abnormal by following the steps below:

[0140] Obtaining a second similarity between the frequency behavior sequence of the qth account and a preset abnormal frequency behavior sequence;

[0141] When the second similarity is greater than or equal to a preset third similarity threshold, it is determined that the frequency behavior sequence of the qth account is abnormal.

[0142] It should be noted that the second similarity between the frequency behavior sequence of the qth account and the preset abnormal frequency behavior sequence can be related to the frequency difference of the same behavior ID in the frequency behavior sequence of the qth account and the preset abnormal frequency behavior sequence, and can be the reciprocal of the sum of the frequency differences of the same behavior ID in the two sequences.

[0143] When the second similarity is greater than or equal to a preset third similarity threshold, the frequency behavior sequence may be considered abnormal. The third similarity threshold in this embodiment may be a value completely different from the second similarity threshold.

[0144] Optionally, determining an account with an abnormal behavior sequence among the Q accounts as an abnormal account includes:

[0145] When the behavior sequence of each account in the Q accounts includes a time behavior sequence, an account in the Q accounts whose time behavior sequence is abnormal is determined as an abnormal account;

[0146] When the behavior sequence of each account in the Q accounts includes a frequency behavior sequence, an account in the Q accounts having an abnormal frequency behavior sequence is determined as an abnormal account;

[0147] When the behavior sequence of each account among the Q accounts includes a time behavior sequence and a frequency behavior sequence, the account among the Q accounts having an abnormal time behavior sequence and an abnormal frequency behavior sequence is determined as an abnormal account.

[0148] It should be noted that, for the same account, there may be a situation where its time behavior sequence is abnormal but its frequency behavior sequence is normal, or vice versa. That is, when the behavior sequence of each account includes both the time behavior sequence and the frequency behavior sequence, four results may be obtained: both the time behavior sequence and the frequency behavior sequence are normal, the time behavior sequence is abnormal but the frequency behavior sequence is normal, the time behavior sequence is normal but the frequency behavior sequence is abnormal, and both the time behavior sequence and the frequency behavior sequence are abnormal.

[0149] Therefore, when the behavior sequence of each account includes both the time behavior sequence and the frequency behavior sequence, the account can be considered normal only when both the time behavior sequence and the frequency behavior sequence are normal.

[0150] Through this embodiment, by analyzing the time behavior sequence and / or frequency behavior sequence of each account, it is determined whether the account has abnormalities, which can improve the comprehensiveness of account abnormality analysis and improve the accuracy of account analysis.

[0151] As an optional example, determining an abnormal account set from a group of abnormal accounts includes:

[0152] S71, obtaining account description information of each abnormal account in a group of abnormal accounts to obtain a group of account description information;

[0153] S72, determining a feature vector of each abnormal account in a group of abnormal accounts according to a group of account description information, to obtain a group of feature vectors;

[0154] S73, obtaining the similarity between every two feature vectors in a set of feature vectors to obtain a set of similarities;

[0155] S74, determining an abnormal account set in a group of abnormal accounts based on a set of similarities, wherein the similarity between each abnormal account in the abnormal account set and at least one abnormal account in the abnormal account set is greater than or equal to a preset first similarity threshold means: the similarity between the feature vector of each abnormal account in the abnormal account set and at least one abnormal account in the abnormal account set is greater than or equal to a third similarity threshold.

[0156] In order to improve the accuracy of the judgment of the abnormal account set, after determining a group of abnormal accounts from the target account set, the similarity judgment can be performed on the group of abnormal accounts, such as Figure 8 As shown, according to the similarity determination result, the abnormal account set (i.e., Figure 8 Suspicious black market accounts).

[0157] In this embodiment, the similarity determination of abnormal accounts can be performed by calculating the similarity between the feature vectors corresponding to the abnormal accounts. A model can be used to determine the feature probability distribution of each abnormal account, where the feature probability distribution can be represented in the form of a normalized feature vector (e.g., a multidimensional vector).

[0158] The similarity between feature vectors of different accounts may be determined by a cosine similarity algorithm between vectors. The above model may be XGBoost (eXtreme Gradient Boosting, a machine learning model based on gradient boosting tree).

[0159] Optionally, a deep neural network model (e.g., Transformer (a neural network model based on a self-attention mechanism)) can be used to replace the XGBoost model for determining the similarity of abnormal accounts, and the deep neural network model can be an integrated feature processing model that can be used to analyze time behavior series and to determine the similarity of abnormal accounts, thereby avoiding building separate models for different parts.

[0160] Through this embodiment, by calculating the similarity between the feature vectors of abnormal accounts, accounts with less similarity in a group of accounts can be eliminated, thereby improving the accuracy of the determined abnormal account set.

[0161] As an optional example, each abnormal account in the abnormal account set and the object associated with each abnormal account are used to generate an associated abnormal object set, including:

[0162] S81, using each abnormal account in the abnormal account set and the object associated with each abnormal account, constructing an association graph, wherein the association graph includes a group of nodes, a group of nodes has a one-to-one correspondence with a group of objects, a group of objects includes each abnormal account and the object associated with each abnormal account, each node in the association graph has a connection relationship with at least one node, and two nodes with a connection relationship represent two objects in a group of objects that are associated with each other;

[0163] S82, clustering a group of nodes using the structural feature vector of each node in a group of nodes to obtain W node sets, wherein W is a positive integer greater than or equal to 1, the structural feature vector of each node is used to represent the position of each node in the association graph and the connection relationship between each node and other nodes, each node set in the W node sets includes multiple nodes in a group of nodes and the connection relationship between each node in the multiple nodes and at least one node, and the distance between the structural feature vectors of each node in each node set in the W node sets is less than or equal to a threshold;

[0164] S83, performing multiple rounds of iterative division of the association graph until the change in the modularity value of the node sets divided twice adjacently is less than a preset threshold, wherein the node set divided last time includes Z node sets, each node set in the Z node sets includes multiple nodes in a group of nodes and a connection relationship between each node in the multiple nodes and at least one node, Z is a positive integer greater than or equal to 1, and the modularity value is a value determined according to the number of nodes in all the divided node sets and the weight set on the edge connected to each node;

[0165] S84, determining the same node set in the W node sets and the Z node sets as the target node set, and determining the object corresponding to each node in the target node set to obtain an associated abnormal object set.

[0166] In this embodiment, after determining the abnormal account set, a correlation graph can be constructed based on the abnormal accounts and other objects associated with each abnormal account. The correlation graph can be a heterogeneous graph. In the graph, each object is a node, and nodes are connected based on actual object association relationships.

[0167] Take the above-mentioned association map as an abnormal map as an example, Fig. 9 As shown, the abnormal account set determined in the aforementioned embodiment can be directly combined with the associated objects to construct an abnormal graph. The constructed abnormal graph can contain multiple different types of objects, including but not limited to accounts, IPs, devices, mobile phone numbers, etc.

[0168] Considering that objects associated with different accounts may have the same objects, such as Figure 5As shown, account a is associated with IP1, and account b is associated with both IP1 and IP2. Therefore, IP1 exists in the objects associated with account a and account b. In this embodiment, after determining the objects associated with each abnormal account in the abnormal account set and obtaining the first group of objects, the first group of objects and the second group of objects corresponding to the abnormal account set can be deduplicated, that is, only one of the same multiple objects associated with different accounts is retained. The third group of objects obtained after deduplicating the first group of objects and the second group of objects is the above-mentioned group of objects, and each object in the third group of objects corresponds to a node.

[0169] For the constructed association graph (i.e., the above-mentioned heterogeneous graph), the structural feature vector of each node can be determined based on the connection relationship between each node and other nodes in the graph, and then a group of nodes in the association graph can be clustered based on the structural feature vectors of different nodes to obtain W node sets. Here, the multiple nodes in each node set can be considered as nodes with greater structural correlation, and the objects corresponding to the multiple nodes can be considered as objects with greater correlation in abnormal features and abnormal behaviors.

[0170] The HetGNN (Heterogeneous Graph Neural Network) algorithm can be used to determine the structural feature vector of each node. Each structural feature vector can represent the position of each node in the association graph and the connection relationship between each node and other nodes, and can also represent information such as the characteristics of other nodes around each node.

[0171] In this embodiment, the association graph can also be directly subjected to sub-graph segmentation using a sub-graph segmentation algorithm such as fast unfolding (Fast unfolding of communities in large networks, a fast unfolding algorithm for communities in large networks, i.e., a modularity optimization algorithm) to obtain Z node sets.

[0172] like Fig.10 As shown in the figure, through the graph construction method, a heterogeneous graph containing all objects is obtained. By dividing the sub-graph, multiple closely connected sub-heterogeneous graphs are obtained, and some nodes in the heterogeneous graph are removed ( Fig.10 Some nodes are not connected by connecting lines).

[0173] It should be noted that the above-mentioned partitioning operation on the association graph can be a multi-round iterative partitioning process. Taking the fastunfolding algorithm as an example, it can include the following process:

[0174] Step 1: Treat each node in the association graph as an independent sub-graph, that is, the number of sub-graphs initially divided is the same as the number of all nodes;

[0175] Step 2: For each node, try to divide the node into its adjacent sub-graph in turn, and calculate the change in the modularity value of the associated graph before and after the node is divided. If the modularity of a node belonging to sub-graph a after being divided into other sub-graphs does not increase compared to the value before the division, then the node is considered to belong to sub-graph a. If the modularity of a node after being divided into other sub-graphs (such as sub-graph b) increases by a certain value compared to the value before the division, then the node is considered to belong to sub-graph b and is divided into sub-graph b.

[0176] Step 3, repeat step 2, and continue to calculate the node transfer and modularity change between sub-graphs until all nodes have been operated in step 2. If the nodes in the divided sub-graphs do not change during the whole process, it can be considered that the optimal division of the sub-graphs under greedy is achieved, that is, the sub-graph division of the associated graph is completed. If the nodes in the divided sub-graphs have changed in the above steps, continue to step 4;

[0177] Step 4: Treat all the nodes that belong to the same sub-graph as a new node (i.e., package the nodes that currently belong to the same sub-graph as a single node), generate a new heterogeneous graph, and execute step 1 on the new heterogeneous graph;

[0178] Step 5, repeat steps 1-4 until the modularity no longer changes or reaches the preset number of iterations, and use the current partitioning result as the output result.

[0179] The above modularity value can be used to measure the ratio of the connection strength between the node in the sub-graph corresponding to each node set and other nodes in the sub-graph to the connection strength of the node in the associated graph. The calculation formula of modularity is shown in formula (1):

[0180]

[0181] Among them, c i is the subgraph label of node i, k i is the degree of point i, A ij For node n i and n j The edge weight of the network (connection graph) is δ(c i ,c j ) is 1 if the value is not set, otherwise it is 0.

[0182] It should be noted that the larger the modularity value is, the higher the connection density of the nodes in each node set currently divided is, that is, the better the division effect of the node set is.

[0183] For the W node sets and Z node sets determined by the above two methods, the same node sets in the two sets can be determined as target node sets, and the objects corresponding to the nodes in the target node sets can form an associated abnormal object set.

[0184] Through this embodiment, a graph is constructed by combining the abnormal account set and other objects associated with it to generate a set containing multiple types of objects, which can improve the richness of the determined abnormal object set.

[0185] The following is an explanation of the method for determining the abnormal object set in the embodiment of the present application in conjunction with an optional example. In this optional example, the abnormal object is an abnormal individual.

[0186] This optional example provides a data mining method for fraudulent traffic, such as Fig.11 As shown in the figure, the black and gray industry data of the Internet open source data, business data and third-party data are used as data sources. Data is mined from two dimensions: organization and single point. Through the perception and mining of abnormal individuals, the analysis of the behavior sequence of individual associated accounts, the analysis of organizational maps (as well as the mining of black industry organizations and the analysis of industrial chain resources), the abnormal individual data is comprehensively judged, and various types of abnormal results are output, such as threat event data, confrontation technology data, portrait mode data, abnormal data and black industry price data.

[0187] The process of the method for determining a set of abnormal objects in this optional example may include the following steps:

[0188] Step 1: From the data sets obtained from different sources, determine the candidate abnormal objects with abnormal data, such as abnormal accounts, abnormal devices, abnormal IPs, etc.

[0189] Step 2: Determine the attribute dimension information of the candidate abnormal object and the attribute dimension information of other objects associated with the candidate abnormal object.

[0190] Step 3: directly extract content information used to describe the identity of the object from the determined attribute dimension information, and perform statistics on the attribute information to obtain statistical information.

[0191] Step 4: Determine the feature vector corresponding to the content information of each candidate abnormal object, use the feature vector to perform clustering, and obtain a set of candidate abnormal objects; then use a decision tree or risk control model, combined with the statistical information of each candidate abnormal object in the set of candidate abnormal objects, to filter the set of candidate abnormal objects to obtain a set of target objects, and determine a set of accounts corresponding to the set of target objects.

[0192] Step 5: multiple accounts in the account set with abnormal behavior sequences and high similarity are determined as an abnormal account set.

[0193] Step 6: Take each account in the abnormal account set and the object associated with each account as a graph node, and combine the association relationships between different objects to generate an association graph containing multiple types of objects.

[0194] Step 7, divide the association graph into subgraphs to obtain multiple association subgraphs, and combine the results of clustering the structural feature vectors corresponding to each node in the association graph to obtain a set of abnormal objects.

[0195] Through this optional example, relying on this method, it is possible to achieve the ability to integrate and process and analyze massive amounts of Internet open source data, business data, and third-party data, ensuring data reliability and analysis accuracy. At the same time, multiple abnormal individuals can be combined with the overall traffic environment perception for data mining, which can improve the perception of abnormal individuals. In the process of mining fraud data, there is no need for manual search of large amounts of data. By combining its own business data with external data, the data is manually processed and then modeled and analyzed. Manual processing is only required for the data output form, which can reduce manual participation and improve mining efficiency. In addition, traffic data indicators, feature indicators, and modeling quantitative indicators can be adaptively analyzed, and can be adaptively adjusted due to changes in traffic or switching of business focus points, without the need to repeatedly re-model different abnormal individuals.

[0196] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0197] According to another aspect of the embodiment of the present application, there is also provided Fig.12 A device for determining a set of abnormal objects is shown, the device comprising:

[0198] A first determining unit 1202 is used to determine N candidate abnormal objects from data sets from different sources, where N is a positive integer greater than or equal to 1;

[0199] The acquiring unit 1204 is connected to the first determining unit 1202, and is used to acquire attribute information of each candidate abnormal object in the N candidate abnormal objects to obtain a first attribute information set, and acquire attribute information of an object associated with each candidate abnormal object in the N candidate abnormal objects to obtain a second attribute information set;

[0200] A first execution unit 1206 is connected to the acquisition unit 1204 and is used to convert the first attribute information set and the second attribute information set into a feature vector to obtain N feature vectors corresponding to the N candidate abnormal objects, and cluster the N candidate abnormal objects using the N feature vectors to obtain M candidate abnormal object sets, wherein a distance between feature vectors corresponding to each candidate abnormal object in a j-th candidate abnormal object set in the M candidate abnormal object sets is less than a preset distance threshold, and j is a positive integer greater than or equal to 1 and less than or equal to M;

[0201] The second execution unit 1208 is connected to the first execution unit 1206, and is used to use the M candidate abnormal object sets to find an abnormal account set that meets the preset conditions, and use each abnormal account in the abnormal account set and the object associated with each abnormal account to generate an associated abnormal object set, wherein the preset conditions include: the behavior sequence of each abnormal account in the abnormal account set is abnormal, and the similarity between each abnormal account and at least one abnormal account in the abnormal account set is greater than or equal to a preset first similarity threshold.

[0202] It should be noted that the first determination unit 1202 in this embodiment can be used to execute the above step S202, the acquisition unit 1204 in this embodiment can be used to execute the above step S204, the first execution unit 1206 in this embodiment can be used to execute the above step S206, and the second execution unit 1208 in this embodiment can be used to execute the above step S208.

[0203] Through the above modules, for objects that may be abnormal and are perceived from multi-source data, similar objects are clustered by combining the attribute information associated with the object and the relevant attribute information of other objects associated with the object, and then the behavior sequence of the account corresponding to each abnormal object is analyzed to obtain a group of similar abnormal accounts. By combining the objects associated with each account in the abnormal account set, a set of associated abnormal objects is obtained. Since changes in traffic or business focus affect the category of abnormal objects, the impact on the information category contained in the attribute information of abnormal objects is relatively small, and the attribute information of each associated object is used in the analysis and judgment of abnormal objects. No matter how the traffic or business focus changes, there is no need to re-establish the model. At the same time, since the analysis and judgment of abnormal objects mainly rely on the associated attribute information, there is no need to manually search for a large amount of data, which can reduce the degree of manual participation, thereby achieving the technical effect of improving the efficiency of determining the set of abnormal objects.

[0204] Optionally, the acquiring unit includes:

[0205] A first acquisition module is used to acquire attribute information of one or more dimensions of each candidate abnormal object in N candidate abnormal objects, and obtain N groups of self attribute information, wherein the first attribute information set includes the N groups of self attribute information, and the i-th group of self attribute information in the N groups of self attribute information includes the attribute information of one or more dimensions of the i-th candidate abnormal object in the N candidate abnormal objects, and i is a positive integer greater than or equal to 1 and less than or equal to N;

[0206] The second acquisition module is used to obtain attribute information of one or more dimensions of one or more objects associated with each of the N candidate abnormal objects, and obtain N groups of associated attribute information, wherein the second attribute information set includes N groups of associated attribute information, and the i-th group of associated attribute information in the N groups of associated attribute information includes attribute information of one or more dimensions of one or more objects associated with the i-th candidate abnormal object in the N candidate abnormal objects.

[0207] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0208] Optionally, the first execution unit includes:

[0209] A third acquisition module is used for acquiring content information included in each group of self-attribute information in the N groups of self-attribute information when the first attribute information set includes N groups of self-attribute information, so as to obtain N groups of self-content information, wherein the i-th group of self-attribute information in the N groups of self-attribute information includes attribute information of one or more dimensions of the i-th candidate abnormal object in the N candidate abnormal objects, i is a positive integer greater than or equal to 1 and less than or equal to N, and the i-th group of self-content information of the i-th candidate abnormal object is used to describe the i-th candidate abnormal object;

[0210] a fourth acquisition module, configured to, when the second attribute information set includes N groups of associated attribute information, acquire content information included in each group of associated attribute information in the N groups of associated attribute information, and obtain N groups of associated content information, wherein the i-th group of associated attribute information in the N groups of associated attribute information includes attribute information of one or more dimensions of one or more objects associated with the i-th candidate abnormal object among the N candidate abnormal objects, and the i-th group of associated content information of the one or more objects associated with the i-th candidate abnormal object is used to describe the one or more objects associated with the i-th candidate abnormal object;

[0211] The splicing module is used to perform vector splicing on feature vectors corresponding to N groups of own content information and feature vectors corresponding to N groups of associated content information to obtain N feature vectors corresponding to N candidate abnormal objects, wherein the i-th feature vector in the N feature vectors is used to represent the i-th candidate abnormal object in the N candidate abnormal objects, and the i-th feature vector in the N feature vectors is a feature vector determined based on the i-th group of own content information in the N groups of own content information and the i-th group of associated content information in the N groups of associated content information.

[0212] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0213] Optionally, the splicing module includes:

[0214] The execution submodule is used to determine the i-th feature vector among the N feature vectors by the following steps:

[0215] Convert each content information in the i-th group of self-content information into a corresponding feature vector to obtain a first group of feature vectors;

[0216] Convert each piece of content information in the i-th group of associated content information into a corresponding feature vector to obtain a second group of feature vectors;

[0217] The first set of eigenvectors and the second set of eigenvectors are concatenated to obtain the i-th eigenvector.

[0218] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0219] Optionally, the second execution unit includes:

[0220] A first determination module is used to determine a target account set corresponding to the target abnormal object set when the target abnormal object set in the M candidate abnormal object sets meets a preset screening condition;

[0221] A fifth acquisition module, configured to acquire a behavior sequence of each account in the target account set, and determine an account in the target account set with an abnormal behavior sequence as an abnormal account;

[0222] The second determining module is used to determine an abnormal account set in the group of abnormal accounts when a group of abnormal accounts is determined in the target account set.

[0223] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0224] Optionally, the above device further includes:

[0225] a combining unit, for combining the N groups of own attribute information and the N groups of associated attribute information into N attribute information sets before determining the target account set corresponding to the target abnormal object set, when the first attribute information set includes N groups of own attribute information and the second attribute information set includes N groups of associated attribute information, wherein the i-th group of own attribute information in the N groups of own attribute information includes attribute information of one or more dimensions of the i-th candidate abnormal object in the N candidate abnormal objects, the i-th group of associated attribute information in the N groups of associated attribute information includes attribute information of one or more dimensions of one or more objects associated with the i-th candidate abnormal object in the N candidate abnormal objects, i is a positive integer greater than or equal to 1 and less than or equal to N, and the i-th attribute information set in the N attribute information sets includes the i-th group of own attribute information and the i-th group of associated attribute information;

[0226] A statistical unit, used for performing statistics on information in each of the N attribute information sets to obtain N statistical information, wherein the i-th statistic in the N statistical information corresponds to the i-th candidate abnormal object in the N candidate abnormal objects;

[0227] The second determination unit is used to use N statistical information to determine a target abnormal object set from M candidate abnormal object sets, wherein the number of candidate abnormal objects whose statistical information meets the screening rule in the target abnormal object set is greater than or equal to a preset number threshold, and the screening condition includes the preset screening rule and the preset number threshold, or the proportion of candidate abnormal objects whose statistical information meets the screening rule in the target abnormal object set is greater than or equal to a preset proportion threshold, and the screening condition includes the preset screening rule and the preset proportion threshold.

[0228] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0229] Optionally, the second determining unit includes:

[0230] The third determination module is used to determine whether the target abnormal object set in the M candidate abnormal object sets meets the preset screening condition through the following steps:

[0231] In the case where the target abnormal object set includes K candidate abnormal objects, determining K statistical information corresponding to the K candidate abnormal objects from the N statistical information, where K is a positive integer greater than or equal to 2;

[0232] Determine whether each of the K pieces of statistical information satisfies a screening rule included in the screening condition;

[0233] When the statistical information corresponding to P candidate abnormal objects among the K candidate abnormal objects meets the screening rules, and P is greater than or equal to the preset quantity threshold, or P / K is greater than or equal to the preset ratio threshold, it is determined that the target abnormal object set meets the preset screening conditions.

[0234] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0235] Optionally, the first determining module includes:

[0236] The first determination submodule is used to determine the p-th group of accounts in K groups of accounts corresponding to the p-th candidate abnormal object in the K candidate abnormal objects by the following steps when the target abnormal object set includes K candidate abnormal objects, wherein K is a positive integer greater than or equal to 2, p is a positive integer greater than or equal to 1 and less than or equal to K, and the target account set includes K groups of accounts:

[0237] In the case where the pth candidate abnormal object is an account, determining the pth group of accounts to include the pth candidate abnormal object;

[0238] In the case that the p-th candidate abnormal object is not an account, the account associated with the p-th candidate abnormal object is obtained, and the p-th group of accounts is determined to include the account associated with the p-th candidate abnormal object.

[0239] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0240] Optionally, the first determining submodule includes:

[0241] A first acquisition subunit is used to acquire an account using the IP address when the p-th candidate abnormal object is an IP address, and determine the p-th group of accounts as including the account using the IP address;

[0242] The second acquisition subunit is used to obtain the account of the logged-in device when the p-th candidate abnormal object is a device, and determine the p-th group of accounts as including the account of the logged-in device.

[0243] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0244] Optionally, the fifth acquisition module includes:

[0245] A first acquisition submodule is used to acquire a behavior sequence of each of the Q accounts when the target account set includes Q accounts, wherein Q is a positive integer greater than or equal to 1, and the behavior sequence of each of the Q accounts includes a time behavior sequence and / or a frequency behavior sequence, and the time behavior sequence of the qth account among the Q accounts represents a group of operations performed by the qth account in chronological order, and the frequency behavior sequence of the qth account among the Q accounts represents the number of times each operation is performed in a group of operations performed by the qth account, and q is a positive integer greater than or equal to 1 and less than or equal to Q;

[0246] The second determination submodule is used to determine whether the behavior sequence of each account in the Q accounts is abnormal;

[0247] The third determination submodule is used to determine the accounts with abnormal behavior sequences among the Q accounts as abnormal accounts.

[0248] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0249] Optionally, the second determining submodule includes:

[0250] The second determination subunit is used to determine whether the time behavior sequence of the qth account among the Q accounts is abnormal by the following steps:

[0251] Obtaining a first similarity between the time behavior sequence of the qth account and a preset abnormal time behavior sequence;

[0252] When the first similarity is greater than or equal to a preset second similarity threshold, it is determined that the time behavior sequence of the qth account is abnormal.

[0253] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0254] Optionally, the second determining submodule includes:

[0255] The third determination subunit is used to determine whether the frequency behavior sequence of the qth account among the Q accounts is abnormal by the following steps:

[0256] Obtaining a second similarity between the frequency behavior sequence of the qth account and a preset abnormal frequency behavior sequence;

[0257] When the second similarity is greater than or equal to a preset third similarity threshold, it is determined that the frequency behavior sequence of the qth account is abnormal.

[0258] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0259] Optionally, the third determining submodule includes:

[0260] A fourth determining subunit is configured to, when the behavior sequence of each account in the Q accounts includes a time behavior sequence, determine an account in the Q accounts whose time behavior sequence is abnormal as an abnormal account;

[0261] A fifth determining subunit is configured to, when the behavior sequence of each account in the Q accounts includes a frequency behavior sequence, determine an account in the Q accounts whose frequency behavior sequence is abnormal as an abnormal account;

[0262] The sixth determination subunit is used to determine the account with abnormal time behavior sequence and abnormal frequency behavior sequence among the Q accounts as an abnormal account when the behavior sequence of each account among the Q accounts includes a time behavior sequence and a frequency behavior sequence.

[0263] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0264] Optionally, the second determining module includes:

[0265] The second acquisition submodule is used to acquire the account description information of each abnormal account in a group of abnormal accounts to obtain a group of account description information;

[0266] A fourth determination submodule is used to determine a feature vector of each abnormal account in a group of abnormal accounts according to a group of account description information to obtain a group of feature vectors;

[0267] A third acquisition submodule is used to obtain the similarity between every two feature vectors in a set of feature vectors to obtain a set of similarities;

[0268] The fifth determination submodule is used to determine a set of abnormal accounts in a group of abnormal accounts based on a set of similarities, wherein the similarity between each abnormal account in the abnormal account set and at least one abnormal account in the abnormal account set is greater than or equal to a preset first similarity threshold, which means that the similarity between the feature vector of each abnormal account in the abnormal account set and at least one abnormal account in the abnormal account set is greater than or equal to a third similarity threshold.

[0269] Optional examples of this embodiment can refer to the examples shown in the above-mentioned method for determining a set of abnormal objects, which will not be described in detail in this embodiment.

[0270] Optionally, the second execution unit includes:

[0271] A construction module, used to construct an association graph using each abnormal account in the abnormal account set and the object associated with each abnormal account, wherein the association graph includes a group of nodes, a group of nodes has a one-to-one correspondence with a group of objects, a group of objects includes each abnormal account and the object associated with each abnormal account, each node in the association graph has a connection relationship with at least one node, and two nodes with a connection relationship represent two mutually related objects in a group of objects;

[0272] A clustering module, used to cluster a group of nodes using a structural feature vector of each node in a group of nodes to obtain W node sets, wherein W is a positive integer greater than or equal to 1, the structural feature vector of each node is used to represent the position of each node in the association map and the connection relationship between each node and other nodes, each node set in the W node sets includes multiple nodes in a group of nodes and the connection relationship between each node in the multiple nodes and at least one node, and the distance between the structural feature vectors of each node in each node set in the W node sets is less than or equal to a threshold value;

[0273] A partitioning module is used to perform multiple rounds of iterative partitioning of the association graph until the change in the modularity value of the node sets divided from two adjacent partitions is less than a preset threshold, wherein the node set divided from the last partition includes Z node sets, each of the Z node sets includes multiple nodes in a group of nodes and a connection relationship between each node in the multiple nodes and at least one node, Z is a positive integer greater than or equal to 1, and the modularity value is a value determined based on the number of nodes in all the partitioned node sets and the weight set on the edge connected to each node;

[0274] The fourth determination module is used to determine the same node set in the W node sets and the Z node sets as the target node set, and determine the object corresponding to each node in the target node set to obtain the associated abnormal object set.

[0275] It should be noted that the embodiments of the apparatus for determining a set of abnormal objects herein may refer to the embodiments of the method for determining a set of abnormal objects described above, which will not be described in detail herein.

[0276] According to another aspect of the embodiment of the present application, an electronic device for implementing the above-mentioned method for determining a set of abnormal objects is also provided. The electronic device may be Fig.13 The terminal device shown in FIG. This embodiment is described by taking the electronic device as a background device as an example. Fig.13 As shown, the electronic device includes a memory 1302 and a processor 1304. The memory 1302 stores a computer program, and the processor 1304 is configured to execute the steps in any of the above method embodiments through the computer program.

[0277] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0278] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0279] S1, determining N candidate abnormal objects from data sets from different sources, where N is a positive integer greater than or equal to 1;

[0280] S2, obtaining attribute information of each candidate abnormal object among the N candidate abnormal objects to obtain a first attribute information set, and obtaining attribute information of an object associated with each candidate abnormal object among the N candidate abnormal objects to obtain a second attribute information set;

[0281] S3, converting the first attribute information set and the second attribute information set into feature vectors to obtain N feature vectors corresponding to the N candidate abnormal objects, and clustering the N candidate abnormal objects using the N feature vectors to obtain M candidate abnormal object sets, wherein the distance between the feature vectors corresponding to each candidate abnormal object in the j-th candidate abnormal object set in the M candidate abnormal object sets is less than a preset distance threshold, and j is a positive integer greater than or equal to 1 and less than or equal to M;

[0282] S4, using M candidate abnormal object sets to search for an abnormal account set that meets preset conditions, and using each abnormal account in the abnormal account set and the objects associated with each abnormal account to generate an associated abnormal object set, wherein the preset conditions include: the behavior sequence of each abnormal account in the abnormal account set is abnormal, and the similarity between each abnormal account and at least one abnormal account in the abnormal account set is greater than or equal to a preset first similarity threshold.

[0283] Alternatively, a person skilled in the art may understand that: Fig.13 The structure shown is for illustration only, and the electronic device may also be a target terminal such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, and a mobile Internet device (MID), a PAD, etc. Fig.13 The electronic device and the electronic equipment described above are not limited in structure. Fig.13 More or fewer components (such as network interfaces, etc.) as shown in, or with Fig.13 Different configurations shown.

[0284] Among them, the memory 1302 can be used to store software programs and modules, such as the program instructions / modules corresponding to the method and device for determining the abnormal object set in the embodiment of the present application. The processor 1304 executes various functional applications and data processing by running the software programs and modules stored in the memory 1302, that is, the above-mentioned method for determining the abnormal object set is realized. The memory 1302 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1302 may further include a memory remotely arranged relative to the processor 1304, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1302 can be specifically used, but is not limited to, for storing a splash screen page, a startup page of a first application, and description information of a free installation program. As an example, if Fig.13As shown, the memory 1302 may include, but is not limited to, the first determination unit 1202, the acquisition unit 1204, the first execution unit 1206, and the second execution unit 1208 in the device for determining the abnormal object set. In addition, other module units in the device for determining the abnormal object set may also be included but are not limited to, which will not be repeated in this example.

[0285] Optionally, the transmission device 1306 is used to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one example, the transmission device 1306 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 1306 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0286] In addition, the electronic device further includes: a display 1308 for displaying the direction prompt information of the target sound; and a connection bus 1310 for connecting various module components in the electronic device.

[0287] In other embodiments, the target terminal or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. Among them, a peer-to-peer network may be formed between the nodes, and any form of computing device, such as a server, terminal or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.

[0288] According to another aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for determining a set of abnormal objects provided in various optional implementations of the above-mentioned server verification processing and other aspects, wherein the computer program is configured to execute the steps of any of the above-mentioned method embodiments when running.

[0289] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:

[0290] S1, determining N candidate abnormal objects from data sets from different sources, where N is a positive integer greater than or equal to 1;

[0291] S2, obtaining attribute information of each candidate abnormal object among the N candidate abnormal objects to obtain a first attribute information set, and obtaining attribute information of an object associated with each candidate abnormal object among the N candidate abnormal objects to obtain a second attribute information set;

[0292] S3, converting the first attribute information set and the second attribute information set into feature vectors to obtain N feature vectors corresponding to the N candidate abnormal objects, and clustering the N candidate abnormal objects using the N feature vectors to obtain M candidate abnormal object sets, wherein the distance between the feature vectors corresponding to each candidate abnormal object in the j-th candidate abnormal object set in the M candidate abnormal object sets is less than a preset distance threshold, and j is a positive integer greater than or equal to 1 and less than or equal to M;

[0293] S4, using M candidate abnormal object sets to search for an abnormal account set that meets preset conditions, and using each abnormal account in the abnormal account set and the objects associated with each abnormal account to generate an associated abnormal object set, wherein the preset conditions include: the behavior sequence of each abnormal account in the abnormal account set is abnormal, and the similarity between each abnormal account and at least one abnormal account in the abnormal account set is greater than or equal to a preset first similarity threshold.

[0294] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the target terminal through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.

[0295] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0296] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to execute all or part of the steps of the methods of each embodiment of the present application.

[0297] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0298] In the several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0299] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0300] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0301] The above are only preferred implementations of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for determining a set of abnormal objects, characterized in that: include: Determine N candidate abnormal objects from data sets from different sources, where N is a positive integer greater than or equal to 1; Acquire attribute information of each candidate abnormal object among the N candidate abnormal objects to obtain a first attribute information set, and acquire attribute information of an object associated with each candidate abnormal object among the N candidate abnormal objects to obtain a second attribute information set; Converting the first attribute information set and the second attribute information set into feature vectors to obtain N feature vectors corresponding to the N candidate abnormal objects, and clustering the N candidate abnormal objects using the N feature vectors to obtain M candidate abnormal object sets, wherein the distance between the feature vectors corresponding to each candidate abnormal object in the j-th candidate abnormal object set in the M candidate abnormal object sets is less than a preset distance threshold, and j is a positive integer greater than or equal to 1 and less than or equal to M; The M candidate abnormal object sets are used to search for an abnormal account set that meets preset conditions, and each abnormal account in the abnormal account set and the objects associated with each abnormal account are used to generate an associated abnormal object set, wherein the preset conditions include: the behavior sequence of each abnormal account in the abnormal account set is abnormal, and the similarity between each abnormal account and at least one abnormal account in the abnormal account set is greater than or equal to a preset first similarity threshold.

2. The method according to claim 1, characterized in that The step of obtaining the attribute information of each candidate abnormal object in the N candidate abnormal objects to obtain a first attribute information set, and obtaining the attribute information of an object associated with each candidate abnormal object in the N candidate abnormal objects to obtain a second attribute information set includes: Acquire attribute information of one or more dimensions of each of the N candidate abnormal objects to obtain N groups of self-attribute information, wherein the first attribute information set includes the N groups of self-attribute information, and the i-th group of self-attribute information in the N groups of self-attribute information includes attribute information of one or more dimensions of the i-th candidate abnormal object in the N candidate abnormal objects, where i is a positive integer greater than or equal to 1 and less than or equal to N; Acquire attribute information of one or more dimensions of one or more objects associated with each of the N candidate abnormal objects to obtain N groups of associated attribute information, wherein the second attribute information set includes the N groups of associated attribute information, and the i-th group of associated attribute information in the N groups of associated attribute information includes attribute information of one or more dimensions of one or more objects associated with the i-th candidate abnormal object in the N candidate abnormal objects.

3. The method according to claim 1, characterized in that: The converting the first attribute information set and the second attribute information set into feature vectors to obtain N feature vectors corresponding to the N candidate abnormal objects includes: In the case where the first attribute information set includes N groups of own attribute information, content information included in each group of own attribute information in the N groups of own attribute information is obtained to obtain N groups of own content information, wherein the i-th group of own attribute information in the N groups of own attribute information includes attribute information of one or more dimensions of the i-th candidate abnormal object in the N candidate abnormal objects, i is a positive integer greater than or equal to 1 and less than or equal to N, and the i-th group of self content information of the i-th candidate abnormal object is used to describe the i-th candidate abnormal object; In the case where the second attribute information set includes N groups of associated attribute information, content information included in each group of associated attribute information in the N groups of associated attribute information is obtained to obtain N groups of associated content information, wherein the i-th group of associated attribute information in the N groups of associated attribute information includes attribute information of one or more dimensions of one or more objects associated with the i-th candidate abnormal object among the N candidate abnormal objects, and the i-th group of associated content information of the one or more objects associated with the i-th candidate abnormal object is used to describe the one or more objects associated with the i-th candidate abnormal object; Performing vector concatenation on the feature vectors corresponding to the N groups of own content information and the feature vectors corresponding to the N groups of associated content information to obtain N feature vectors corresponding to the N candidate abnormal objects, wherein the i-th feature vector among the N feature vectors is used to represent the i-th candidate abnormal object among the N candidate abnormal objects, and the i-th feature vector among the N feature vectors is a feature vector determined according to the i-th group of own content information among the N groups of own content information and the i-th group of associated content information among the N groups of associated content information.

4. The method according to claim 3, characterized in that The step of respectively concatenating the feature vectors corresponding to the N groups of self-content information and the feature vectors corresponding to the N groups of associated content information to obtain N feature vectors corresponding to the N candidate abnormal objects includes: The i-th eigenvector among the N eigenvectors is determined by the following steps: Convert each piece of content information in the i-th group of self-content information into a corresponding feature vector to obtain a first group of feature vectors; Convert each piece of content information in the i-th group of associated content information into a corresponding feature vector to obtain a second group of feature vectors; The first group of feature vectors and the second group of feature vectors are concatenated to obtain the i-th feature vector.

5. The method according to claim 1, characterized in that: The using the M candidate abnormal object sets to search for an abnormal account set that meets a preset condition includes: When a target abnormal object set in the M candidate abnormal object sets meets a preset screening condition, determining a target account set corresponding to the target abnormal object set; Acquire a behavior sequence of each account in the target account set, and determine an account in the target account set with an abnormal behavior sequence as an abnormal account; When a group of abnormal accounts is determined in the target account set, the abnormal account set is determined in the group of abnormal accounts.

6. The method according to claim 5, characterized in that Before determining the target account set corresponding to the target abnormal object set, the method further includes: In a case where the first attribute information set includes N groups of own attribute information, and the second attribute information set includes N groups of associated attribute information, the N groups of own attribute information and the N groups of associated attribute information are combined into N attribute information sets, wherein the i-th group of own attribute information in the N groups of own attribute information includes attribute information of one or more dimensions of the i-th candidate abnormal object in the N candidate abnormal objects, the i-th group of associated attribute information in the N groups of associated attribute information includes attribute information of one or more dimensions of one or more objects associated with the i-th candidate abnormal object in the N candidate abnormal objects, i is a positive integer greater than or equal to 1 and less than or equal to N, and the i-th attribute information set in the N attribute information sets includes the i-th group of own attribute information and the i-th group of associated attribute information; Performing statistics on information in each of the N attribute information sets to obtain N statistical information, wherein an i-th statistic in the N statistical information corresponds to an i-th candidate abnormal object in the N candidate abnormal objects; Using the N statistical information, the target abnormal object set is determined from the M candidate abnormal object sets, wherein the number of candidate abnormal objects whose statistical information satisfies the screening rule in the target abnormal object set is greater than or equal to a preset number threshold, and the screening condition includes the preset screening rule and the preset number threshold, or, the proportion of candidate abnormal objects whose statistical information satisfies the screening rule in the target abnormal object set is greater than or equal to a preset proportion threshold, and the screening condition includes the preset screening rule and the preset proportion threshold.

7. The method according to claim 6, characterized in that The using the N pieces of statistical information to determine the target abnormal object set from the M candidate abnormal object sets includes: Determine whether the target abnormal object set in the M candidate abnormal object sets meets the preset screening condition by the following steps: In a case where the target abnormal object set includes K candidate abnormal objects, determining K statistical information corresponding to the K candidate abnormal objects from the N statistical information, wherein K is a positive integer greater than or equal to 2; Determine whether each of the K pieces of statistical information satisfies the screening rule included in the screening condition; When the statistical information corresponding to P candidate abnormal objects among the K candidate abnormal objects meets the screening rule, and P is greater than or equal to the preset quantity threshold, or P / K is greater than or equal to the preset ratio threshold, it is determined that the target abnormal object set meets the preset screening condition.

8. The method according to claim 7, characterized in that The determining the target account set corresponding to the target abnormal object set includes: In the case where the target abnormal object set includes K candidate abnormal objects, the p-th group account in K groups of accounts corresponding to the p-th candidate abnormal object among the K candidate abnormal objects is determined by the following steps, wherein K is a positive integer greater than or equal to 2, p is a positive integer greater than or equal to 1 and less than or equal to K, and the target account set includes the K groups of accounts: In the case where the p-th candidate abnormal object is an account, determining the p-th group of accounts to include the p-th candidate abnormal object; In the case that the p-th candidate abnormal object is not an account, an account associated with the p-th candidate abnormal object is obtained, and the p-th group of accounts is determined to include the account associated with the p-th candidate abnormal object.

9. The method according to claim 8, characterized in that In the case where the p-th candidate abnormal object is not an account, obtaining an account associated with the p-th candidate abnormal object, and determining the p-th group of accounts to include the account associated with the p-th candidate abnormal object, comprises: In the case where the p-th candidate abnormal object is an IP address, obtaining an account using the IP address, and determining the p-th group of accounts to include the account using the IP address; In the case that the p-th candidate abnormal object is a device, an account for logging into the device is obtained, and the p-th group of accounts is determined to include the account for logging into the device.

10. The method according to claim 5, characterized in that The acquiring the behavior sequence of each account in the target account set, and determining the account with abnormal behavior sequence in the target account set as an abnormal account, includes: In the case where the target account set includes Q accounts, a behavior sequence of each account in the Q accounts is obtained, wherein Q is a positive integer greater than or equal to 1, and the behavior sequence of each account in the Q accounts includes a time behavior sequence and / or a frequency behavior sequence, the time behavior sequence of the qth account in the Q accounts represents a group of operations performed by the qth account in chronological order, the frequency behavior sequence of the qth account in the Q accounts represents the number of times each operation in a group of operations performed by the qth account is performed, and q is a positive integer greater than or equal to 1 and less than or equal to Q; Determining whether a behavior sequence of each of the Q accounts is abnormal; An account having an abnormal behavior sequence among the Q accounts is determined as the abnormal account.

11. The method according to claim 10, characterized in that The determining whether the behavior sequence of each of the Q accounts is abnormal includes: Determine whether the time behavior sequence of the qth account among the Q accounts is abnormal by the following steps: Obtaining a first similarity between the time behavior sequence of the qth account and a preset abnormal time behavior sequence; When the first similarity is greater than or equal to a preset second similarity threshold, it is determined that the time behavior sequence of the qth account is abnormal.

12. The method according to claim 10, characterized in that The determining whether the behavior sequence of each of the Q accounts is abnormal includes: Determine whether the frequency behavior sequence of the qth account among the Q accounts is abnormal by the following steps: Obtaining a second similarity between the frequency behavior sequence of the qth account and a preset abnormal frequency behavior sequence; When the second similarity is greater than or equal to a preset third similarity threshold, it is determined that the frequency behavior sequence of the qth account is abnormal.

13. The method according to claim 10, characterized in that The determining the account with abnormal behavior sequence among the Q accounts as the abnormal account includes: In the case where the behavior sequence of each account in the Q accounts includes the time behavior sequence, determining the account in the Q accounts whose time behavior sequence is abnormal as the abnormal account; In the case where the behavior sequence of each account in the Q accounts includes the frequency behavior sequence, determining the account in the Q accounts whose frequency behavior sequence is abnormal as the abnormal account; In the case where the behavior sequence of each account in the Q accounts includes the time behavior sequence and the frequency behavior sequence, the account in the Q accounts having abnormal time behavior sequence and abnormal frequency behavior sequence is determined as the abnormal account.

14. The method according to claim 5, characterized in that The determining the abnormal account set from the group of abnormal accounts includes: Obtaining account description information of each abnormal account in the group of abnormal accounts to obtain a group of account description information; Determine, according to the set of account description information, a feature vector of each abnormal account in the set of abnormal accounts to obtain a set of feature vectors; Obtaining the similarity between every two feature vectors in the set of feature vectors to obtain a set of similarities; Based on the set of similarities, the abnormal account set is determined in the set of abnormal accounts, wherein the similarity between each abnormal account in the abnormal account set and at least one abnormal account in the abnormal account set is greater than or equal to a preset first similarity threshold, which means that the similarity between the feature vector of each abnormal account in the abnormal account set and at least one abnormal account in the abnormal account set is greater than or equal to a third similarity threshold.

15. The method according to any one of claims 1 to 14, characterized in that The step of using each abnormal account in the abnormal account set and an object associated with each abnormal account to generate an associated abnormal object set includes: Using each abnormal account in the abnormal account set and the object associated with each abnormal account, construct an association graph, wherein the association graph includes a group of nodes, the group of nodes has a one-to-one correspondence with a group of objects, the group of objects includes each abnormal account and the object associated with each abnormal account, each node in the association graph has a connection relationship with at least one node, and two nodes with a connection relationship represent two mutually related objects in the group of objects; Using the structural feature vector of each node in the group of nodes, clustering the group of nodes to obtain W node sets, wherein W is a positive integer greater than or equal to 1, the structural feature vector of each node is used to represent the position of each node in the association graph and the connection relationship between each node and other nodes, each node set in the W node sets includes multiple nodes in the group of nodes and the connection relationship between each node in the multiple nodes and at least one node, and the distance between the structural feature vectors of each node in each node set in the W node sets is less than or equal to a threshold value; Perform multiple rounds of iterative division on the association graph until the change in the modularity value of the node sets divided from two adjacent divisions is less than a preset threshold, wherein the node set divided from the last division includes Z node sets, each of the Z node sets includes multiple nodes in a group of nodes and a connection relationship between each node in the multiple nodes and at least one node, Z is a positive integer greater than or equal to 1, and the modularity value is a value determined based on the number of nodes in all the divided node sets and the weight set on the edge connected to each node; The same node sets in the W node sets and the Z node sets are determined as target node sets, and the object corresponding to each node in the target node set is determined to obtain the associated abnormal object set.

16. A device for determining a set of abnormal objects, characterized in that: include: A first determining unit is used to determine N candidate abnormal objects from data sets from different sources, where N is a positive integer greater than or equal to 1; an acquiring unit, configured to acquire attribute information of each of the N candidate abnormal objects to obtain a first attribute information set, and acquire attribute information of an object associated with each of the N candidate abnormal objects to obtain a second attribute information set; A first execution unit is configured to convert the first attribute information set and the second attribute information set into feature vectors to obtain N feature vectors corresponding to the N candidate abnormal objects, and cluster the N candidate abnormal objects using the N feature vectors to obtain M candidate abnormal object sets, wherein a distance between feature vectors corresponding to each candidate abnormal object in a j-th candidate abnormal object set in the M candidate abnormal object sets is less than a preset distance threshold, and j is a positive integer greater than or equal to 1 and less than or equal to M; The second execution unit is used to use the M candidate abnormal object sets to search for an abnormal account set that meets a preset condition, and use each abnormal account in the abnormal account set and the object associated with each abnormal account to generate an associated abnormal object set, wherein the preset condition includes: the behavior sequence of each abnormal account in the abnormal account set is abnormal, and the similarity between each abnormal account and at least one abnormal account in the abnormal account set is greater than or equal to a preset first similarity threshold.

17. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program can be executed by a terminal device or a computer to execute the method described in any one of claims 1 to 15.

18. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method described in any one of claims 1 to 15 are implemented.

19. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 15 through the computer program.