Method, apparatus and terminal device for risk identification

By establishing edge relationships between media vectors and behavior vectors among users, combining community division and edge attribute information, the problem of low reliability of risk identification in the existing technology is solved, and more efficient and accurate risk identification is achieved.

CN114781517BActive Publication Date: 2025-07-18JINGDONG TECH HLDG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210431364.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-22
Publication Date
2025-07-18
Estimated Expiration
2042-04-22

AI Technical Summary

Technical Problem

In the prior art, risk identification methods rely on a large number of labeled training data, resulting in misjudgment or misjudgment of the model, and the recognition reliability is low.

Method used

By obtaining user service data within the preset time period, determining the media vector and behavior vectors, establishing edge relationships between users based on similarity, and dividing communities, and using edge attribute information to judge community risks.

Benefits of technology

It simplifies the complexity of risk identification and improves the accuracy of risk identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114781517B_ABST
    Figure CN114781517B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for risk identification. The method includes: obtaining business data sets corresponding to each user within a preset time period; preprocessing the business data corresponding to each user to determine a media vector and a behavior vector corresponding to each user; determining an edge relationship between users according to the similarity between each media vector and the similarity between each behavior vector; partitioning the relationship graph into communities according to the edge relationship between users to determine each community included in the relationship graph; and determining whether each community is a community with risks according to the attribute information of the edges included in each community. Thus, by establishing an edge relationship between users based on the media vector and the behavior vector, and then determining whether each community is a community with risks according to the attribute information of the edges included in each community, the complexity of risk identification is simplified and the accuracy of risk identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of artificial intelligence recognition and classification, and particularly to a method, device and terminal device for risk recognition. Background Art

[0002] With the rapid development of artificial intelligence technology, the demand for risk control is increasing.

[0003] In related technologies, a classification model is usually trained based on user time-series behavior events, and based on this classification model, it is determined whether there is a risk in the corresponding service. This method requires a large amount of labeled training data, but due to the difficulty of obtaining the labeled training data set, there are phenomena of misjudgment or missed judgment in the model. Therefore, how to provide a reliable method for risk recognition is an urgent problem to be solved at present. Summary of the Invention

[0004] The present disclosure provides a method, device and terminal device for risk recognition to at least solve the problem of low reliability of risk recognition in related technologies. The technical solutions of the present disclosure are as follows:

[0005] According to the first aspect of the embodiments of the present disclosure, the embodiments of the present disclosure provide a method for risk recognition, including:

[0006] Obtain the service data sets corresponding to each user within a preset time period, where each piece of the service data includes media data and behavior data;

[0007] Preprocess the service data corresponding to each user to determine the media vector and behavior vector corresponding to each user;

[0008] Determine the edge relationship between the users according to the similarity between the media vectors and the similarity between the behavior vectors;

[0009] According to the edge relationship between the users, perform community partitioning on the relationship graph to determine each community included in the relationship graph;

[0010] Determine whether each community is a community with risks according to the attribute information of the edges included in each community.

[0011] In the present disclosure, after the server obtains the business data sets corresponding to each user within a preset time period, it can preprocess the business data corresponding to each user to determine the media vector and behavior vector corresponding to each user. Then, according to the similarity between each media vector and the similarity between each behavior vector, the edge relationship between each user is determined, and according to the edge relationship between each user, the relationship graph is divided into communities to determine each community included in the relationship graph. Then, according to the attribute information of the edges included in each community, it is determined whether each community is a risky community. Thus, by establishing the edge relationship between each user based on the media vector and the behavior vector, and then determining whether each community is a risky community according to the attribute information of the edges included in each community, the complexity of risk identification is simplified and the accuracy of risk identification is improved.

[0012] In a possible implementation manner of the first aspect embodiment of the present disclosure, after determining the edge relationship between each user, it further includes:

[0013] Determine the attribute information of the operation object corresponding to each piece of the behavior data;

[0014] According to the attribute information of the operation object, determine the extended vector corresponding to each user;

[0015] Update the edge relationship between each user according to the similarity between each extended vector.

[0016] In a possible implementation manner of the first aspect embodiment of the present disclosure, the preprocessing of the business data corresponding to each user to determine the media vector and behavior vector corresponding to each user includes:

[0017] Perform vector mapping on the media data and behavior data distribution in each piece of business data corresponding to the user to determine each media vector and each behavior vector corresponding to the user.

[0018] In a possible implementation manner of the first aspect embodiment of the present disclosure, the determining the edge relationship between each user according to the similarity between each media vector and the similarity between each behavior vector includes:

[0019] When the similarity between any one media vector and another media vector is greater than a threshold, it is determined that there is a first edge between the first user corresponding to the any one media vector and the second user corresponding to the another media vector, where the attribute information of the first edge is the media data corresponding to the any one media vector.

[0020] In a possible implementation manner of the first aspect of the present disclosure, determining whether each of the communities is a community with risks according to the attribute information of the edges included in each of the communities includes:

[0021] Determining whether the community is a community with risks according to the matching degrees between the attribute information of each edge in each of the communities and preset reference information.

[0022] According to the second aspect of the embodiments of the present disclosure, an apparatus for risk identification is provided, including:

[0023] An acquisition module, configured to acquire a service data set corresponding to each user within a preset time period, where each piece of the service data includes media data and behavior data;

[0024] A determination module, configured to preprocess the service data corresponding to each user to determine a media vector and a behavior vector corresponding to each user;

[0025] An edge building module, further configured to determine the edge relationships between the users according to the similarities between the media vectors and the similarities between the behavior vectors;

[0026] A partitioning module, configured to perform community partitioning on the relationship graph according to the edge relationships between the users to determine each community included in the relationship graph;

[0027] The determination module is further configured to determine whether each of the communities is a community with risks according to the attribute information of the edges included in each of the communities.

[0028] In a possible implementation manner of the second aspect of the embodiments of the present disclosure, the determination module is further configured to:

[0029] Determine the attribute information of the operation object corresponding to each piece of the behavior data; and determine an extended vector corresponding to each user according to the attribute information of the operation object.

[0030] The apparatus further includes:

[0031] An update module, configured to update the edge relationships between the users according to the similarities between the extended vectors.

[0032] In a possible implementation manner of the second aspect of the embodiments of the present disclosure, the determination module is specifically configured to:

[0033] Perform vector mapping on the distributions of the media data and the behavior data in each piece of the service data corresponding to the user to determine each media vector and each behavior vector corresponding to the user.

[0034] In a possible implementation manner of the second aspect embodiment of the present disclosure, the edge building module is specifically configured to:

[0035] When the similarity between any media vector and another media vector is greater than a threshold, determine that there is a first edge between a first user corresponding to the any media vector and a second user corresponding to the another media vector, where the attribute information of the first edge is the media data corresponding to the any media vector.

[0036] In a possible implementation manner of the second aspect embodiment of the present disclosure, the determination module is specifically configured to:

[0037] Determine whether the community is a risky community according to the matching degree between the attribute information of each edge in each community and preset reference information.

[0038] According to a third aspect of the embodiments of the present disclosure, there is provided a terminal device, including:

[0039] A processor;

[0040] A memory for storing processor-executable instructions;

[0041] Wherein, the processor is configured to execute instructions to implement the risk identification method as described in the first aspect embodiment above.

[0042] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by the processor of the terminal device, enabling the terminal device to execute the risk identification method as described in the above-mentioned one aspect embodiment.

[0043] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, implementing the risk identification method as described in the above-mentioned one aspect embodiment.

[0044] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: In the present disclosure, after the server obtains the business data sets corresponding to each user within a preset time period, it can preprocess the business data corresponding to each user to determine the media vector and behavior vector corresponding to each user. Then, according to the similarity between each media vector and the similarity between each behavior vector, the edge relationship between each user is determined, and according to the edge relationship between each user, the relationship graph is divided into communities to determine each community included in the relationship graph. Then, according to the attribute information of the edges included in each community, it is determined whether each community is a risky community. Thus, by establishing the edge relationship between each user based on the media vector and the behavior vector, and then determining whether each community is a risky community according to the attribute information of the edges included in each community, the complexity of risk identification is simplified and the accuracy of risk identification is improved.

[0045] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.

[0047] Figure 1 It is a schematic flowchart of a method for risk identification provided by the first embodiment of the present disclosure;

[0048] Figure 2 It is a schematic flowchart of another method for risk identification provided by the second embodiment of the present disclosure;

[0049] Figure 3 It is a schematic flowchart of another method for risk identification provided by the third embodiment of the present disclosure

[0050] Figure 4 It is a schematic structural diagram of a risk identification processing device provided by the fourth embodiment of the present disclosure;

[0051] Figure 5 It is a block diagram of a terminal device for risk identification processing shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0053] It should be noted that the terms "first", "second", etc. in the description, claims and the above drawings of the present disclosure are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0054] In the present disclosure, mainly aiming at the problem in the related art that a large amount of labeled training data is required, but due to the difficulty in obtaining the labeled training data set, the model has problems of misjudgment or missed judgment, a risk identification method is proposed. In the method provided by the present disclosure, only based on the user service data within a period of time, the relationship graph of the user is determined, and then according to the attribute information of the edges included in each community, it can be determined whether each community is a community with risks, thereby simplifying the complexity of risk identification while improving the accuracy of risk identification.

[0055] Figure 1 It is a flowchart of a risk identification processing method provided by an embodiment of the present disclosure, including the following steps:

[0056] Step 101, obtain the service data sets corresponding to each user within a preset time period, where each piece of service data includes media data and behavior data.

[0057] Among them, the service corresponding to the user can be any service that the service provider can provide. For example, if the service provider is an e-commerce service provider, the services corresponding to the user can include services such as registration, login, transaction, refund, etc.

[0058] The media data can be the media information used by the user when requesting a service. For example, the media data can be the IP address of a computer device, the Media Access Control Address (MAC) of a mobile terminal device, the MAC address of the affiliated mobile hotspot (Wi-Fi), etc., and the present disclosure does not limit this.

[0059] The behavior data can be the operation data generated by the user when applying for a certain service. For example, the time of applying for a registered account, the transaction number, the transaction time, the transaction items, etc., and the present disclosure does not limit this.

[0060] In the present disclosure, to ensure the accuracy of risk identification, it is possible to determine whether a community is a risk community based on the business data sets corresponding to each user over a period of time. For example, every other day, week, or month, risk identification analysis is performed based on the business data sets within the last day, week, or month.

[0061] In the present disclosure, when any user applies for any business service, the server can store each piece of business data in the corresponding business table according to the business type, and store each business table in the data warehouse. Thus, when performing risk identification, the business data of each user within a preset time period can be extracted from each business table in the data warehouse for risk analysis.

[0062] Step 102: Preprocess the business data corresponding to each user to determine the media vector and behavior vector corresponding to each user.

[0063] In the present disclosure, the server can extract the media data of each user from the business data set, and integrate various media data corresponding to each user into a string as the media vector corresponding to the user. For example, when the media data of a user includes the MAC address of a mobile terminal device and the MAC address of the WIFI used, the MAC address of the mobile terminal device and the MAC address of the affiliated WIFI can be concatenated as the media vector corresponding to the user.

[0064] In the present disclosure, the server can extract the behavior data of each user from the business data set, and concatenate the business number and occurrence time in each behavior data into a string as the behavior vector corresponding to the user. It can be understood that since each behavior data has a corresponding occurrence time, the generated behavior vector can be a time series vector.

[0065] Step 103: Determine the edge relationship between users based on the similarity between each media vector and the similarity between each behavior vector.

[0066] In the present disclosure, the server can calculate the distance between each pair of media vectors corresponding to each user, and use the distance between the two vectors to represent the similarity. When the distance is larger, the corresponding similarity is smaller; when the distance is smaller, the corresponding similarity is larger. When the similarity is greater than the preset threshold, an edge corresponding to the media vector can be established between the two users. Similarly, the edges corresponding to the behavior vectors between each user can be established in the same way. Thus, through the similarity between each media vector and the similarity between each behavior vector, the edge relationship between each user is established, filtering out some invalid data for subsequent community division, which is conducive to improving the efficiency of risk identification.

[0067] Optionally, the server can also compare each media vector and each behavior vector corresponding to each user pairwise to determine the edge relationship between two users. When the media vectors corresponding to two users are the same, an edge corresponding to the media vector can be established between the two users. When the behavior vectors corresponding to two users are the same, an edge corresponding to the behavior vector can be established between the two users.

[0068] It can be understood that through the above edge establishment method, there may be 0 - 2 edges between two users, namely no associated edge, or there is an edge corresponding to the media vector, or there is an edge corresponding to the behavior vector, or there are both an edge corresponding to the media vector and an edge corresponding to the behavior vector.

[0069] Step 104: According to the edge relationship between each user, perform community partitioning on the relationship graph to determine each community included in the relationship graph.

[0070] Among them, the relationship graph can include multiple nodes, the connection edges between nodes, and the attribute information of each edge, etc. Among them, a node can represent a user, and the present disclosure does not limit this.

[0071] In the present disclosure, the server can input the relationship graph into community partitioning algorithms such as Infomap. The community partitioning algorithm can initialize multiple starting points and, according to the edge relationship between each user, divide the users into multiple communities by means of random walk.

[0072] Step 105: According to the attribute information of the edges included in each community, determine whether each community is a risky community.

[0073] Among them, the attribute information of the edge can include the type of the edge, the attribute information corresponding to the edge, and the corresponding similarity value, etc., and the present disclosure does not limit this. Among them, the type of the edge can be determined by the vector based on which the edge is established. For example, for an edge determined by calculating the similarity of the media vectors of two users, the type of the corresponding edge can be a media edge. For an edge determined by calculating the similarity of the behavior vectors of two users, the type of the corresponding edge can be a behavior edge. In addition, when the type of the edge is a media edge, the attribute information of the edge can also include the media data of any user connected by this edge. When the type of the edge is a behavior edge, the attribute information of the edge can also include the behavior data of any user connected by this edge.

[0074] In the present disclosure, according to the characteristics of each risk behavior, the community reference edge attributes corresponding to various risk behaviors can be determined, and then, according to the relationship between the attribute information of the edges in the actual community and the reference edge attributes, it can be determined whether the community is a risky community.

[0075] For example, in the scenario of malicious brush orders, the operation behaviors of users are similar. Then, in a certain community, if the attribute information of the behavior edges is the same, it can be considered that there may be malicious brush order behaviors in this community.

[0076] In the present disclosure, after the server obtains the business data sets corresponding to each user within a preset time period, it can preprocess the business data corresponding to each user to determine the media vector and behavior vector corresponding to each user. Then, according to the similarity between each media vector and the similarity between each behavior vector, the edge relationship between each user is determined, and according to the edge relationship between each user, the relationship graph is divided into communities to determine each community included in the relationship graph. Then, according to the attribute information of the edges included in each community, it is determined whether each community is a community with risks. Thus, by establishing the edge relationship between each user based on the media vector and the behavior vector, and then determining whether each community is a community with risks according to the attribute information of the edges included in each community, the complexity of risk identification is simplified and the accuracy of risk identification is improved.

[0077] Figure 2 It is a flowchart of a risk identification processing method provided by an embodiment of the present disclosure, including the following steps:

[0078] Step 201, obtain the business data sets corresponding to each user within a preset time period, where each piece of business data includes media data and behavior data.

[0079] Step 202, preprocess the business data corresponding to each user to determine the media vector and behavior vector corresponding to each user.

[0080] Step 203, determine the edge relationship between each user according to the similarity between each media vector and the similarity between each behavior vector.

[0081] In the present disclosure, for the specific implementation processes of steps 201-step 203, reference can be made to the detailed description of the above embodiments, which will not be elaborated here.

[0082] Step 204, determine the attribute information of the operation object corresponding to each piece of behavior data.

[0083] In the present disclosure, in addition to identifying risks based on behavior data related to time sequences such as business numbers and occurrence times, risks can also be identified based on the attribute information of the operation object corresponding to each piece of behavior data, further improving the accuracy of risk identification.

[0084] For example, in a transaction scenario, the operation object corresponding to each piece of behavior data can be any commodity, and the attribute information of the operation object can be commodity identification, store identification to which the commodity belongs, activity identification participated by the commodity, etc., and the present disclosure does not limit this.

[0085] Among them, the product identifier can be any information that can uniquely identify a product, such as a product number. The store identifier to which the product belongs can be any information that can uniquely identify a store, such as a store number. The activity identifier participated by the product can be any information that can uniquely identify an activity, such as an activity application number.

[0086] In the present disclosure, when any user applies for any business service, the server can correspond the attribute information of the operation object with the behavior data and store them in the corresponding business table. The server can obtain the attribute information of the operation object corresponding to each behavior data by querying the corresponding business table in the database.

[0087] Step 205: Determine the extended vector corresponding to each user according to the attribute information of the operation object.

[0088] In the present disclosure, after determining the attribute information of the operation object, the attribute information of the operation object can be converted into a string vector, and this string vector can be used as the extended vector corresponding to the user. For example, when the attribute information of the operation object includes a product identifier, a store identifier to which the product belongs, and an activity identifier participated by the product, the product identifier, the store identifier to which the product belongs, and the activity identifier participated by the product can be concatenated into a string vector as the extended vector corresponding to the user.

[0089] Optionally, it is also possible to determine the region to which the account ID corresponding to each behavior data belongs, and further expand the attribute information of the operation object according to the region to which the account ID belongs, so as to further improve the accuracy of risk identification.

[0090] In the present disclosure, the region to which the account ID belongs can be numbered, and the number of the region to which each account ID belongs can be concatenated with the corresponding attribute information of the operation object into a string vector as the extended vector corresponding to the user.

[0091] Step 206: Update the edge relationship between users according to the similarity between the extended vectors.

[0092] In the present disclosure, the server can calculate the distance between the extended vectors corresponding to each user in pairs, and use the distance between the two vectors to represent the similarity. When the distance is larger, the corresponding similarity is smaller; when the distance is smaller, the corresponding similarity is larger. When the similarity is greater than the preset threshold, an edge corresponding to the extended vector can be established between the two users.

[0093] It can be understood that after updating the edge relationships between users according to the similarity between each extended vector, there may be 0-3 edges between two users, namely, no associated edge, or an edge corresponding to the media vector, or an edge corresponding to the behavior vector, or an edge corresponding to the extended vector, or any two of the edges corresponding to the media vector, the behavior vector, and the extended vector, or three edges, namely, an edge corresponding to the media vector, an edge corresponding to the behavior vector, and an edge corresponding to the extended vector.

[0094] Step 207: Partition the relationship graph into communities according to the edge relationships between users to determine each community included in the relationship graph.

[0095] Step 208: Determine whether each community is a risky community according to the attribute information of the edges included in each community.

[0096] In the present disclosure, for the specific implementation processes of steps 207-step 208, reference may be made to the detailed description of the above embodiments, which will not be elaborated herein.

[0097] In the present disclosure, after the server determines the edge relationships between users according to the media vector and the behavior vector related to time series, it can also determine the extended vector corresponding to each user according to the attribute information of the operation object, and then update the edge relationships between users according to the similarity between each extended vector, and then partition the relationship graph into communities according to the edge relationships between users to determine each community included in the relationship graph, and determine whether each community is a risky community according to the attribute information of the edges included in each community. Thus, by means of the media vector, the behavior vector, and the extended vector, edge relationships are established between users, and then it is determined whether each community is a risky community according to the attribute information of the edges included in each community, further improving the accuracy of risk identification.

[0098] Figure 3 The flowchart of a risk identification processing method provided by an embodiment of the present disclosure includes the following steps:

[0099] Step 301: Obtain the business data sets corresponding to each user within a preset time period, where each piece of business data includes media data and behavior data.

[0100] Among them, for the specific implementation process of step 301, reference may be made to the detailed description of the above embodiments, which will not be elaborated herein.

[0101] Step 302: Perform vector mapping on the media data and behavior data distributions in each piece of business data corresponding to the user to determine each media vector and each behavior vector corresponding to the user.

[0102] In practical applications, within a certain period of time, a user may request services multiple times. Therefore, a user may generate multiple pieces of service data. In the present disclosure, according to each user identifier, the service data corresponding to each user within the set time period can be filtered out. Then, in the order of the occurrence time, the service numbers and occurrence times in each piece of service data of each user within the preset time period are concatenated into a string, so as to determine the behavior vector corresponding to the user within the preset time period. Among them, the user identifier can be any information such as a user number that can uniquely identify a user. Similarly, the media vectors corresponding to each user within the preset time period can be determined.

[0103] For example, the service data of a certain user within the preset time period includes: registration - 9; login - 10; transaction - 11. Then the behavior vector corresponding to the user within the preset time period can be "registration - 9 - login - 10 - transaction - 11". Among them, the numbers 9, 10, and 11 are the service occurrence times.

[0104] By concatenating the behavior data of each user within the preset time period in the order of the occurrence time, the behavior vector contains the timing information of the service occurrence. Subsequently, risk identification is performed based on this behavior vector, which can improve the accuracy of risk identification.

[0105] Step 303, when the similarity between any one media vector and another media vector is greater than the threshold, it is determined that there is a first edge between the first user corresponding to any one media vector and the second user corresponding to the another media vector. Among them, the attribute information of the first edge is the media data corresponding to any one media vector.

[0106] In the present disclosure, the server can take any one media vector corresponding to the first user and another media vector corresponding to the second user, calculate the distance, and use this distance to represent the similarity between the two vectors. When the distance is larger, the corresponding similarity is smaller; when the distance is smaller, the corresponding similarity is larger. When the similarity is greater than the preset threshold, an edge corresponding to this media vector can be established between the two users, and this edge is determined as the first edge. Similarly, in the same way, each media vector and each behavior vector between pairwise users can be compared and edges can be established to determine the corresponding edges between each user.

[0107] It can be understood that since both the first user and the second user may correspond to multiple media vectors, when there are multiple groups of media vectors with high similarity between the first user and the second user, there can be multiple media edges between the first user and the second user. Similarly, there may also be multiple edges corresponding to behavior vectors between the two users.

[0108] Step 304, according to the edge relationships between each user, the relationship graph is partitioned into communities to determine each community included in the relationship graph.

[0109] Among them, for the specific implementation process of step 304, reference can be made to the detailed description of the above embodiments, which will not be elaborated here.

[0110] Step 305: Determine whether a community is a risky community according to the matching degree between the attribute information of each edge in each community and the preset reference information.

[0111] In the present disclosure, the reference information can be set manually based on experience, or can be automatically generated by the system by statistically analyzing the characteristics of various risky behaviors. The present disclosure does not limit this.

[0112] In addition, the reference information can be the characteristics of any risky behavior. For example, the reference information corresponding to the brushing behavior can be: the attribute information of the behavior edges of each user is the same; the reference information corresponding to the scalper behavior can be: the attribute information of the media edges of each user is the same, etc. The present disclosure does not limit this.

[0113] In the present disclosure, the attribute information of each edge in each community can be statistically analyzed according to the preset reference information. When the statistical result matches the preset reference information, it can be determined that this community is a risky community.

[0114] For example, the preset reference information is: when the media edges of each user are the same. The server queries whether the attribute information of all media edges in a certain community is the same. If so, it can be determined that this community is a risky community.

[0115] In the present disclosure, after the server obtains the business data sets corresponding to each user within a preset time period, it can perform vector mapping on the media data and behavior data distributions in each piece of business data corresponding to the user to determine each media vector and each behavior vector corresponding to the user. Then, when the similarity between any media vector and another media vector is greater than the threshold, it is determined that there is a first edge between the first user corresponding to any media vector and the second user corresponding to the other media vector. Then, according to the edge relationship between each user, the relationship graph is divided into communities to determine each community included in the relationship graph, and according to the matching degree between the attribute information of each edge in each community and the preset reference information, it is determined whether the community is a risky community. Thus, by establishing the edge relationship between each user through the media vector and the behavior vector, and then determining whether each community is a risky community according to the attribute information of the edges included in each community, while simplifying the complexity of the algorithm, the accuracy of risk identification is improved.

[0116] Figure 4 It is a block diagram of a processing device for a service request shown according to an exemplary embodiment. Refer to Figure 4, the apparatus includes an acquisition module 410, a determination module 420, an edge building module 430, and a partitioning module 440.

[0117] The acquisition module 410 is configured to acquire a service data set corresponding to each user within a preset time period, where each piece of the service data includes media data and behavior data;

[0118] The determination module 420 is configured to preprocess the service data corresponding to each user to determine a media vector and a behavior vector corresponding to each user;

[0119] The edge building module 430 is further configured to determine an edge relationship between the users according to the similarity between each of the media vectors and the similarity between each of the behavior vectors;

[0120] The partitioning module 440 is configured to partition the relationship graph according to the edge relationship between the users to determine each community included in the relationship graph;

[0121] The determination module 420 is further configured to determine whether each community is a community with risks according to the attribute information of the edges included in each community.

[0122] In a possible implementation manner of the embodiment of the present disclosure, the above determination module 420 is further configured to:

[0123] Determine the attribute information of the operation object corresponding to each piece of the behavior data; determine an extended vector corresponding to each user according to the attribute information of the operation object;

[0124] The above apparatus further includes:

[0125] An update module, configured to update the edge relationship between the users according to the similarity between each of the extended vectors.

[0126] In a possible implementation manner of the embodiment of the present disclosure, the above determination module 420 is specifically configured to:

[0127] Perform vector mapping on the distribution of the media data and the behavior data in each piece of service data corresponding to the user to determine each media vector and each behavior vector corresponding to the user.

[0128] In a possible implementation manner of the embodiment of the present disclosure, the above edge building module 430 is specifically configured to:

[0129] When the similarity between any one media vector and another media vector is greater than a threshold, determine that there is a first edge between a first user corresponding to the any one media vector and a second user corresponding to the another media vector, where the attribute information of the first edge is the media data corresponding to the any one media vector.

[0130] In a possible implementation manner of the embodiments of the present disclosure, the determining module 420 is specifically configured to:

[0131] Determine whether the community is a risky community according to the matching degree between the attribute information of each edge in each community and the preset reference information.

[0132] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0133] In the present disclosure, after the server obtains the service data sets corresponding to each user within a preset time period, it can preprocess the service data corresponding to each user to determine the media vector and behavior vector corresponding to each user. Then, according to the similarity between each media vector and the similarity between each behavior vector, the edge relationship between each user is determined, and according to the edge relationship between each user, the relationship graph is divided into communities to determine each community included in the relationship graph. Then, according to the attribute information of the edges included in each community, it is determined whether each community is a risky community. Thus, by establishing the edge relationship between each user based on the media vector and the behavior vector, and then determining whether each community is a risky community according to the attribute information of the edges included in each community, the complexity of risk identification is simplified and the accuracy of risk identification is improved.

[0134] Figure 5 It is a block diagram of a terminal device for risk identification shown according to an exemplary embodiment.

[0135] As Figure 5 shown, the terminal device 500 includes:

[0136] A memory 510, a processor 520, and a bus 530 connecting different components (including the memory 510 and the processor 520). The memory 510 stores a computer program, and when the processor 520 executes the program, the processing method of the service request described in the embodiments of the present disclosure is implemented.

[0137] The bus 530 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0138] The terminal device 500 typically includes a variety of electronically readable media. These media can be any available media accessible by the terminal device 600, including volatile and non-volatile media, removable and non-removable media.

[0139] The memory 510 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 540 and / or cache memory 550. The terminal device 500 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 560 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 5 not shown, commonly referred to as a "hard disk drive"). Although Figure 5 not shown in the figure, a disk drive for reading and writing on a removable non-volatile disk (such as a "floppy disk"), and an optical disk drive for reading and writing on a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to the bus 530 through one or more data media interfaces. The memory 510 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present disclosure.

[0140] A program / utility 580 having a set (at least one) of program modules 570 can be stored, for example, in the memory 510. Such program modules 570 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, and the implementation of a network environment may be included in each or some combination of these examples. The program modules 570 generally perform the functions and / or methods in the embodiments described in the present disclosure.

[0141] The terminal device 500 can also communicate with one or more external devices 590 (such as a keyboard, a pointing device, a display 591, etc.), and can also communicate with one or more devices that enable a user to interact with the terminal device 500, and / or communicate with any device that enables the terminal device 500 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 592. Moreover, the terminal device 500 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 593. As shown in the figure, the network adapter 593 communicates with other modules of the terminal device 500 through the bus 530. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the terminal device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0142] The processor 520 executes various functional applications and data processing by running the programs stored in the memory 510.

[0143] It should be noted that for the implementation process and technical principle of the terminal device in this embodiment, refer to the foregoing explanation of the method for processing service requests in the embodiments of the present disclosure, which will not be elaborated here.

[0144] In the present disclosure, after the server obtains the service data sets corresponding to each user within a preset time period, it can preprocess the service data corresponding to each user to determine the media vector and behavior vector corresponding to each user. Then, according to the similarity between each media vector and the similarity between each behavior vector, the edge relationship between each user is determined, and according to the edge relationship between each user, the relationship graph is divided into communities to determine each community included in the relationship graph. Then, according to the attribute information of the edges included in each community, it is determined whether each community is a community with risks. Thus, by establishing the edge relationship between each user based on the media vector and the behavior vector, and then determining whether each community is a community with risks according to the attribute information of the edges included in each community, the complexity of risk identification is simplified and the accuracy of risk identification is improved.

[0145] In an exemplary embodiment, the present disclosure also provides a computer-readable storage medium including instructions, such as a memory including instructions, and the above instructions can be executed by the processor of the terminal device to complete the above method. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0146] To implement the above embodiments, the present disclosure also provides a computer program product. When the computer program is executed by a processor of a terminal device, the terminal device can execute the method for processing a service request as described above.

[0147] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0148] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A method for risk identification, characterized in that, Including: Obtain the business data sets corresponding to each user within a preset time period, where each piece of the business data includes media data and behavior data; Preprocess the business data corresponding to each user to determine the media vector and behavior vector corresponding to each user, where it includes: extracting the media data of each user, integrating various media data corresponding to each user into a string as the media vector corresponding to the user, extracting the behavior data of each user, and concatenating the business number and occurrence time in each behavior data into a string as the behavior vector corresponding to the user; Determine the edge relationships among the users according to the similarities among the respective media vectors and the similarities among the respective behavior vectors; Perform community division on the relationship graph according to the edge relationships among the users to determine each community included in the relationship graph; Determine whether each community is a risky community according to the attribute information of the edges included in each community; The determining the edge relationships among the users according to the similarities among the respective media vectors and the similarities among the respective behavior vectors includes: Compare each pair of the respective media vectors and the respective behavior vectors corresponding to the users to determine the edge relationships among the users, where the edge relationships include edges corresponding to media vectors, or edges corresponding to behavior vectors, or both edges corresponding to media vectors and edges corresponding to behavior vectors.

2. The method according to claim 1, wherein After the determining the edge relationships among the users, it further includes: Determine the attribute information of the operation object corresponding to each piece of the behavior data; Determine the extended vector corresponding to each user according to the attribute information of the operation object; Update the edge relationships among the users according to the similarities among the respective extended vectors.

3. The method according to claim 1, characterized in that, The preprocessing the business data corresponding to each user to determine the media vector and behavior vector corresponding to each user includes: Perform vector mapping on the media data and behavior data distributions in each piece of the business data corresponding to the user to determine each media vector and each behavior vector corresponding to the user.

4. The method according to any one of claims 1-3, characterized in that, The determining whether each community is a risky community according to the attribute information of the edges included in each community includes: Determine whether the community is a risky community according to the matching degrees between the attribute information of each edge in each community and preset reference information.

5. A device for risk identification, characterized in that, Including An obtaining module, configured to obtain the business data sets corresponding to each user within a preset time period, where each piece of the business data includes media data and behavior data; A determining module, configured to preprocess the business data corresponding to each user to determine the media vector and behavior vector corresponding to each user, where it includes: extracting the media data of each user, integrating various media data corresponding to each user into a string as the media vector corresponding to the user, extracting the behavior data of each user, and concatenating the business number and occurrence time in each behavior data into a string as the behavior vector corresponding to the user; The edge building module is further configured to determine the edge relationships among the users according to the similarities among the respective media vectors and the similarities among the respective behavior vectors; The partitioning module is configured to partition the relationship graph according to the edge relationships among the users to determine each community included in the relationship graph; The determining module is further configured to determine whether each community is a risky community according to the attribute information of the edges included in each community; The edge building module is specifically configured to: Compare each pair of the respective media vectors and the respective behavior vectors corresponding to the users to determine the edge relationships among the users, where the edge relationships include edges corresponding to the media vectors, or edges corresponding to the behavior vectors, or both edges corresponding to the media vectors and edges corresponding to the behavior vectors.

6. The device according to claim 5, wherein The determining module is further configured to: Determine the attribute information of the operation object corresponding to each piece of the behavior data; and determine the extended vector corresponding to each user according to the attribute information of the operation object; The apparatus further includes: The updating module is configured to update the edge relationships among the users according to the similarities among the respective extended vectors.

7. The device according to claim 5, characterized in that, The determining module is specifically configured to: Perform vector mapping on the media data and the behavior data distributions in each piece of service data corresponding to the user to determine each media vector and each behavior vector corresponding to the user.

8. The device according to any one of claims 5 to 7, characterized in that The determining module is specifically configured to: Determine whether the community is a risky community according to the matching degrees between the attribute information of each edge in each community and preset reference information, respectively.

9. A terminal device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the risk identification method according to any one of claims 1-4.

10. A computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of a terminal device, enabling the terminal device to execute the risk identification method according to any one of claims 1-4.

11. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the risk identification method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Community partitioning method and device based on characteristic matching network

    CN106709800A

  • Group identification method and device, electronic equipment and storage medium

    CN113487109A