Method for identifying groups of users of interest within a set of users of a transactional service, computer program product, storage medium and corresponding device
The method addresses inefficiencies in identifying user groups by pruning irrelevant users and interactions based on characteristic comparisons and deletion rules, enabling efficient and scalable detection of fraudsters and other groups of interest in transactional services.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2026-03-06
AI Technical Summary
Existing methods for identifying groups of users of interest in transactional services, such as fraudsters, are inefficient when dealing with small data samples, require extensive computational resources, and fail to consider the relational dimension of user interactions, especially in large networks.
A method involving pruning users and interactions that do not share characteristics with a reference group, using feature comparisons and deletion rules to identify groups of interest, which includes associating topological and semantic characteristics with users and interactions, and constructing a graph to remove irrelevant nodes and edges.
Enables effective identification of user groups without requiring large reference groups, scalable across various network sizes, and provides real-time detection and action on suspected fraudsters or other user groups of interest.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method for identifying groups of users of interest within a set of users of a transactional service, computer program and corresponding device. Technical field
[0001] The invention falls within the field of social networks of users sharing a transactional service, such as a telecommunications service or a payment service, for example, via their communication terminal. More particularly, the invention relates to a technique for identifying one or more groups of users of interest within a set of users of a transactional service.
[0002] The invention applies to all contexts of transactional service implementing telecommunications or financial transactions for which it is possible to establish a social relationship between the different users of this service. Technological background
[0003] In recent years, the analysis of socio-transactional networks has received increasing attention. It is of major interest in understanding the habits of service users through the exploitation of relational data. Such data is obtained, in particular, from interactions between users of a service via their communication devices. Whether they originate from social media, telecommunications services, or financial transactions, companies in these sectors have understood the importance of extracting and using information from this relational data in order to enrich their understanding of user behavior, improve service personalization, and / or enhance the security of the services offered, for example, by combating fraud and misuse.These interactions often reveal complex social structures and dynamics, such as user groups linked by common interests, which are worth analyzing in order to improve the functioning of the service in question.
[0004] In the context of implementing a financial service, such as, for example, a mobile banking service (“Mobile Banking” in English), the data created by the relationships that the users of this service maintain with each other can constitute relevant data for detecting the presence of banking fraud or misuse of this service.
[0005] There are in the state of the art various known methods for identifying users of interest, such as fraudsters for example, within a network of users of a service.
[0006] A first known method relies on the implementation of machine learning (or "Machine Learning") using a mathematical model and predefined rules to analyze relational data, interpret it, and identify groups of interest based on characteristics common to the service's users. However, to be effective, this method requires a large amount of labeled data. It is therefore not suitable for small samples. Furthermore, it is not very sensitive to the topological and temporal variability of the groups of interest sought, which limits its implementation to a relatively small number of application contexts.Indeed, in an operational context, it is important to be able to identify groups of users of interest whose topology may vary depending on time and users, as is the case, for example, in bank fraud risk management, where some individuals are particularly agile in their strategies for circumventing countermeasures implemented by banking services.
[0007] A second known method relies on user community detection. This method involves exploiting transactional data by searching for the highest density of interactions within a group of users relative to the entire user network, and grouping them into one or more user communities (family, friends, professional or business contacts, etc.). Variations allow for associating several user groups with the same entity; these are referred to as "overlapping communities." However, this method is complex to implement, particularly on large user networks. Indeed, to be effective, it requires processing a substantial volume of data (several million users and several hundred million transactions).Furthermore, relatively expensive post-processing is often necessary to enable accurate identification of groups of interest. Indeed, community detection algorithms are capable of identifying generic groups across the entire transactional network but are unable to classify these groups into precise categories without post-processing, which is not optimal.
[0008] Finally, a third known method relies on graph matching, more commonly called "Graph Matching". This approach consists of transforming the data into transactional graphical representations and analyzing, from these graphical representations, the topological structure of the groups of users of interest sought, the objective being to identify service users with similar topological characteristics. One of the limitations of this approach The problem lies in its failure to consider the environment of the target interest group and all its interconnections with all users of the service—in other words, the relational dimension of service interactions. Furthermore, this approach is also complex to implement on large user networks.
[0009] There is therefore a need to provide a technique for identifying groups of users of interest that is effective even when there are a small number of users and / or a small amount of data available. In particular, there is a need to provide a technique for identifying groups of users of interest that is scalable and as generic as possible. Description of the invention
[0010] The present technique addresses this need by proposing a method for identifying at least a subset of users, referred to as a group of interest, within a set of users of a transactional service via terminals, said method being implemented in an identification device and comprising the steps of: - receipt of a request to detect said at least one group of interest relative to at least one reference group of users previously identified among the users of said set, - removal of at least one user and / or interaction, within the framework of said transactional service, between two users of said set, delivering a pruned set of users, said removal taking into account at least one removal rule and the result of a comparison of at least one characteristic associated with said at least one user and / or said at least one interaction with at least one characteristic representative of said reference group, - identification of said at least one interest group among the users of said pruned set.
[0011] Thus, the present technique relies on removing, from a set of users, those users who do not share characteristics with a reference group, in order to highlight the remaining users as being, conversely, close to the reference group. This technique therefore differs from known techniques that rely on searching for common characteristics to identify, within a set of users, those users who are close to a reference group.
[0012] Thanks to this pruning principle based on the characteristics of all users in the set, the present technique does not require a large reference group to be effective, even in the presence of a very large number of users in the set research, unlike known techniques based on learning for example, which require a lot of reference data to be effective.
[0013] Furthermore, one or more deletion rules are implemented in addition to feature comparisons, in order to improve the accuracy of the present technique, these deletion rules being able, for example, to take into account the topology of the users intended to be deleted.
[0014] Once irrelevant users have been removed from the starting set, the remaining user groups correspond to the subgroups sought.
[0015] According to a particular aspect of the present technique, the method includes a preliminary step, implemented for at least one user of a plurality of users of said set and / or for at least one interaction between two users of said plurality of users, of associating at least one feature with said user and / or with said interaction, delivering a plurality of features associated with said at least one user and / or with said at least one interaction.
[0016] According to this approach, one or more characteristics specific to one or more users of the set, or a subset thereof, are extracted, including characteristics associated with the interactions between these users, so that these characteristics can then be compared with those representative of the reference group. Indeed, this phase of associating characteristics with at least one user and / or at least one interaction is also applied to the users of the reference group. In this way, it is possible to characterize one or more users of the set with both topological and semantic characteristics, unlike some known techniques that rely solely on the topology of a set to perform user selection.
[0017] According to a particular implementation, said associated characteristics are saved in at least one database. This has the effect of improving the performance of the pruning step.
[0018] According to a particular aspect of the present technique, the deletion comprises: • for at least one user and at least one interaction between two users of a plurality of users of said set: • comparison of at least one characteristic associated with at least one user or with at least one interaction with at least one characteristic representative of said reference group, providing a proximity distance, • when the proximity distance exceeds a predetermined threshold, selection of said at least one user or said at least one interaction and addition to a subgroup of no interest, • removal of said set, according to said at least one removal rule, of said at least one user or of said at least one interaction selected in the subgroup of no interest, delivering said pruned set.
[0019] Thus, according to this aspect, pruning within the user set is based on distance measures between the characteristics associated with a given user and / or interaction and the representative characteristics of the reference group. These measures are then compared to a predefined threshold to determine whether that given user and / or interaction should be removed. Subsequently, one or more removal rules are used to refine the decisions made based on the distance measures between characteristics, removing only users / interactions that are irrelevant to the reference group.
[0020] According to another aspect of the present technique, the method comprises assigning a label to at least one user of said identified interest group, forming a subgroup of interest, said label being representative of the implementation of an action belonging to the group comprising: - the sending of an alert or information message to the terminals of users of said set or subgroup of non-interest; - the exclusion of said transactional service from users of the subgroup of interest.
[0021] This technique is implemented in an application context aimed at optimizing the operation of the transactional service used by users of the set, for example, by detecting groups of fraudsters of a certain type of fraud among all users of the service, starting from an identified group of fraudsters constituting the reference group. In such circumstances, all users of the service, not identified as fraudsters, have a vested interest in being informed of the presence of fraudsters in the user set to which they belong. Once the groups of fraudsters are detected, this technique therefore also makes it possible to inform other users of the service. It is also possible to deny access to the service to groups of users identified as fraudsters through the implementation of this technique.
[0022] According to another aspect of the present technique, the method comprises the creation of a graph representing said set of users, said graph comprising a plurality of nodes corresponding to said users of said transactional service and a plurality of arcs corresponding to the interactions between said users, and where: - said comparison includes, for at least one node and at least one arc of said graph: • a measure of the distance between a feature vector associated with at least one node and / or at least one arc and a feature vector associated with that reference group, providing a proximity distance, • a comparison of said proximity distance delivered with a predetermined threshold, delivering a positive comparison result when said proximity distance is greater than said threshold, - said deletion includes a deletion according to said at least one deletion rule, of said graph, of said at least one node and / or edge when said comparison result is positive, delivering a pruned graph, - said identification includes the determination of at least one subgraph of interest in said pruned graph.
[0023] According to this aspect, the "graph" approach to representing a set of users of a service makes it possible to take into account the "social / relational" dimension of interactions within a transactional service. The characteristics associated with the nodes, representing users, and the edges, representing interactions between users, of the graph are, for example, extracted as characteristic vectors associated with the nodes / edges of the graph. Comparisons are then made in the form of distance measurements between characteristic vectors, and the removal of users / interactions consists of pruning the graph by deleting nodes and the edges between nodes. The remaining connected nodes in the pruned graph then correspond to the groups of interest being sought.
[0024] According to a particular implementation, the characteristics associated with the users and interactions of said set belong to the group comprising: - topological characteristics, - intrinsic characteristics of the service, - characteristics calculated by learning.
[0025] This list of features is not exhaustive.
[0026] According to a particular implementation, said at least one deletion rule belongs to the group comprising: - a so-called node rule, removing all users of said set present in the group of no interest, - a so-called arc rule, removing all interactions of said set present in the group of no interest, - a so-called simple rule, removing all users and interactions from said set present in the group of no interest, - a so-called arc majority rule, removing all users and interactions from said set present in the group of no interest except the nodes having a ratio of absent neighbors from the no-interest group greater than a predefined threshold.
[0027] This list of rules is not exhaustive.
[0028] In another embodiment of the invention, a computer program product is proposed which includes program code instructions for implementing the aforementioned method in any of its various implementation modes, when said program is executed on a computer.
[0029] In another embodiment of the present technique, a computer-readable and non-transient storage medium is proposed, storing a computer program comprising a set of instructions executable by a computer to implement the aforementioned process in any of its various implementation modes.
[0030] In another embodiment of the present technique, a device is proposed for identifying at least a subset of users, referred to as a group of interest, within a set of users of a transactional service via terminals, said identification device comprising at least one processor configured to: - receive a request to detect said at least one group of interest relative to at least one reference group of users previously identified among the users of said set, - to remove at least one user and / or interaction, within the framework of said transactional service, between two users of said set, delivering a pruned set of users, said removal taking into account at least one removal rule and the result of a comparison of at least one characteristic associated with said at least one user and / or said at least one interaction with at least one characteristic representative of said reference group, - identify said at least one interest group among the users of said pruned set.
[0031] Advantageously, the identification device includes means for implementing the steps it performs in the process as described above, in any of its different modes of implementation. List of figures
[0032] Other features and advantages of the invention will become apparent from the following description, given by way of illustrative and non-limiting example, and the accompanying drawings, in which:
[0033] [Fig.1] is a schematic view of an example of a communication system in which the process of the invention is implemented according to a particular embodiment;
[0034] [Fig.2] is a simplified example of graphical representation of the social network of users illustrated in [Fig.1];
[0035] [Fig.3] presents a flowchart of a particular embodiment of the process according to the invention;
[0036] [Fig.4] represents the simplified structure of a device implementing the process according to a particular embodiment of the invention. Detailed description of the invention
[0037] The general principle of the present technique is based on the use of a pruning mechanism for users who do not share common characteristics with a reference group of users identified within a user set (or user network) of a transactional service, in order to highlight the remaining users as being, conversely, close to the reference group. Thanks to this ingenious principle, the present technique does not require the creation of large reference groups to be effective, even with a large number of users in the search set, unlike approaches proposed in the prior art.
[0038] In the following description, an example of the implementation of this technique is considered in the context of a transactional service for managing anti-fraud financial transactions. This technique is, of course, not limited to this particular application context and can be applied to any type of transactional service subscribed to by a group of users where the identification of a subset of users within that group must be implemented via their communication terminals, for the purpose of optimizing the operation of the transactional service.
[0039] Figure 1 shows an example of a communication system SC in which the method of the invention is implemented. Such a communication system SC comprises a management server SG and a set of user terminals UT1-UTn capable of being connected within a communication network, for example, a network operating according to the IP protocol (Internet Protocol). The UT1-UTn terminals are standard communication terminals, such as mobile phones, tablets, or computers, each having subscribed to the same transactional service, for example, a payment or banking transaction service supported by the IP network. This transactional service, referred to as service "ST" in the following description, enables the connection of the users of UTl-UT-n terminals connected to the Internet for the purpose of establishing telecommunications and / or financial transactions between these users via their communication terminals: bank transfer, bank direct debit, credit, mandate, advice and assistance in financial management, electronic money management for example.
[0040] This ST transactional service may be an opportunity for fraudsters to attempt to divert some of these exchanges (telecommunications and / or transactions) to their advantage by contacting their targets in a fortuitous manner (for example: sending messages, widespread calls, etc.) or by pretending to be a trusted relationship (for example: targeted scam attempt, misappropriation or identity theft, etc.).
[0041] Based on past interactions between users, it is possible to construct membership groups (also called "social groups") forming a set of users (also called a "socio-transactional" network). These membership groups consist of individuals sharing relationships that lead them to communicate in various forms. These relationships between individuals (for example: relational, interest-based, or geographical proximity) are identifiable and quantifiable (for example: call frequency, call duration, transaction amount, transaction frequency, etc.). A user can therefore be assigned one or more membership groups based on the relationships they maintain with other users of the ST transactional service.
[0042] To construct this user network, the SG management server uses relational data from interactions between users of the ST service (for example: identity of sending / receiving users, amount of a financial transaction, duration of calls, frequency of calls, frequency of financial transactions, date, IB / AN "International Bank Account Number", etc.). This relational data is stored in a database, which records the history of exchanges since a given date (for example, since the user subscribed to the ST service). From this data, the SG management server establishes a graphical representation of the service's user base as a tree graph of nodes representing users and links (or arcs) representing interactions between users. An arc defines the relationship between two nodes of the graph. [Fig.[2] gives an example of a simplified graph representing the social network of users illustrated in [Fig. 1]. This socio-transactional graph is stored and administered by the SG management server in addition to the relational data saved in the database.
[0043] It is understood that the number of users represented here is intentionally limited, for purely educational purposes, so as not to overload the figures and the associated description. A larger number of terminals can of course be envisaged without departing from the scope of the invention.
[0044] The SG management server periodically, for example when a new user subscribes to the ST service or when a request to update the ST service user set network is received, constructs or updates the user set network and the associated socio-transactional graph.
[0045] The construction of socio-transactional graphs relies on a set of algorithms known from the state of the art and optimized to adapt to the structure, the volume of relational data and the context of application.
[0046] A flowchart of a particular embodiment of the identification method according to the invention is now shown in relation to [Fig. 3]. This flowchart illustrates the main steps in implementing the method in the specific context of a subscription to the ST transactional service discussed above. The purpose of this flowchart is to identify users suspected of fraud and to implement appropriate actions. These steps are carried out by an identification device, the principle of which is described later in relation to [Fig. 4].
[0047] The process is triggered upon receipt of a request to detect a group of interest relative to a reference group of users previously identified within the set of users of the ST service. This request is transmitted, for example, by the supervisor of the ST service after having identified the reference group via its human-machine interface.
[0048] For illustrative purposes, the group of interest sought within the aforementioned set of users of the ST service is considered to be a group of users classified as "fraudsters committing bank account fraud". The group of users previously identified by the supervisor constituting the reference group is hereafter referred to as the "fraudster group".
[0049] In a step E0 (noted "ASS_CA" in the figure), the identification device delivers a list of characteristics associated with each user from the set of users of the ST service and a list of characteristics associated with each interaction between two users of this set.
[0050] To this end, the socio-transactional graph representing all users of the ST service is retrieved beforehand by the identification device from the database. The nodes of the graph correspond to the users of the service and the arcs correspond to the interactions between users (transactions or communications). The identification device thus associates, for each node (i.e., for each user), one or more characteristics of said node, and for each arc (i.e., for each interaction between two users), one or more characteristics of said arc. in order to deliver the two lists of characteristics mentioned above, one associated with the nodes and the other associated with the arcs of the socio-transactional graph.
[0051] The types of characteristics associated with nodes and interactions, and of which the identification device can be aware, are as follows (without being exhaustive): - topological characteristics, such as the degree or coefficient of "local clustering" of the nodes for example; - intrinsic characteristics of the ST service, such as the average balance on an account, the number of bank transactions or the frequency of exchanges for example; - characteristics calculated by machine learning implementing a node2Vec, fastRP or GIN type algorithm (for "Graph Isomorphism Network" in English).
[0052] This step E0 thus makes it possible to extract one or more characteristics specific to each node of the graph (i.e., to each user in the set), including characteristics associated with the arcs (i.e., the interactions between these users), to then compare them with the representative characteristics of the reference group of fraudsters. Indeed, this phase of associating characteristics with each user and / or each interaction is applied a fortiori to the users of the reference group. In this way, it is possible to characterize each user of the ST user network with both topological and semantic characteristics, which makes it possible to take into account the "relational" dimension of the service's interactions.
[0053] Once the association phase is complete, the identification device stores the characteristics associated with the nodes and interactions in the database. Such storage improves the performance of the pruning step described below.
[0054] Once the detection request is received and processed, the identification device proceeds, in a step El (denoted "COMP_CA"), to compare the characteristics associated with users and / or interactions between users of the entire ST service with the representative characteristics of the reference group of fraudsters, based on an analysis of the topological and semantic characteristics of the graph obtained in the previous step. This step determines whether users (nodes) and / or interactions (arcs) between users (nodes) can be removed.
[0055] To this end, the identification device performs, for each node and each arc of the graph: - a measure of the distance between a feature vector associated with a node and / or an arc and a feature vector associated with the reference group, called the proximity distance, and - a comparison of the proximity distance with a predetermined threshold, delivering a positive or negative result.
[0056] The comparison result is positive when the proximity distance is greater than the predetermined threshold, while it is negative when the proximity distance is less than the predetermined threshold. When the comparison result is positive, the device selects the relevant node (i.e., user) and / or the relevant arc (i.e., interaction) and adds it to a subgroup of no interest, which is subsequently removed. This subgroup of no interest represents the users and / or user interactions intended to be excluded from the identification applied to all users of the ST service.
[0057] The proximity distance calculation and comparison phases rely on the use of at least one of the following methods: cosine similarity, Manhattan distance or any other statistical learning method.
[0058] This proximity distance represents, in a way, a level of similarity between the user characteristics / interactions of a given user of the service and the representative user characteristics / interactions of the reference group of fraudsters. This proximity distance evolves over time according to the interactions between these users.
[0059] Then, in a step E2 (denoted "SUP_U / I"), the identification device removes users from the entire ST service and / or user interactions, so as to deliver a pruned set of users based on a comparison of characteristics. This removal step is carried out taking into account, on the one hand, one or more predefined removal rules, and on the other hand, the results from the comparison step EL
[0060] More specifically, the identification device removes from the initial socio-transactional graph each node and / or each edge for which the result from the comparison step El is positive, according to the selected deletion rule(s). At the end of this step, the identification device delivers a pruned socio-transactional graph. This step amounts to removing from the set of users of the ST service, according to the selected deletion rule(s), the users and / or interactions selected in the subgroup of no interest obtained at the end of step El, to deliver a pruned set of users. In other words, the removal of users and / or interactions consists of pruning the graph by removing nodes and edges between nodes; the remaining connected nodes in the pruned graph then correspond to the groups of interest being sought.
[0061] The deletion rules that the identification device can take into account are as follows (without being exhaustive): - a rule, called a node rule, whose principle is to remove all users of said set present in the subgroup of no interest, - a rule, called the arc rule, whose principle is to remove all interactions from said set present in the subgroup of no interest, - a rule, called the simple rule, removing all users and interactions from said set present in the subgroup of no interest, - a rule, called the arc majority rule, removing all users and interactions of said set present in the subgroup of no interest, except for nodes with a ratio of neighbors absent from the subgroup of no interest greater than a predefined threshold.
[0062] These rules allow the identification system to refine its decisions based on distance measurements between feature vectors associated with the nodes and edges in the graph, and to remove only the nodes and edges that are irrelevant to the reference group of fraudsters. This amounts to pruning the entire user set by removing only the users and / or interactions included in the subgroup of no interest (users and / or interactions that are irrelevant to the reference group).
[0063] In the case of applying a simple rule, which consists for example of carrying out a "coarse" pruning of users and interactions present in the subgroup of no interest, it may be interesting, as a complement, for the identification device to carry out a new pruning phase by applying a more precise deletion rule.
[0064] According to an advantageous implementation, the pruning phase applied to the nodes and edges of the graph is carried out at least partially concurrently. This implementation is particularly well suited to large user networks.
[0065] Once the pruned graph has been established by the identification device, the latter determines, in a step E3 (labeled "IDE_NI" in the figure), a group of nodes of interest from among the nodes of the pruned graph. This group of nodes of interest corresponds to the users of the ST service identified as fraudsters (i.e., a subset of users from among the users of the pruned set of users constituting the group of interest being sought). This group of nodes can take the form of a subgraph of interest calculated from the pruned graph. This subgraph of interest is therefore representative of the group of users suspected of fraud.
[0066] Let us take, as an illustrative example, the aforementioned edge majority rule. In this example, the identification device is assumed to be configured to calculate the following "edgeMajority" indicator for each of the nodes v belonging to the subgroup of no interest:
[0067] edgeMajority^v) =
[0068] With:
[0069] E(v) , a vector representing the set of arcs connected to node v
[0070] E(P) \ Ebad, a representative ratio of all arcs connected to node v and which have not been removed.
[0071] If the result of the "edgeMajority" indicator is greater than a fixed threshold (0.5 for example), the node r in question is removed from the subgroup of no interest. Thus, after calculating the set of all nodes r in the subgroup of no interest, the identification device removes from the graph all the nodes that have not been removed from the subgroup of no interest, as well as all the arcs in the subgroup of no interest and the arcs connected (to / from) said nodes that have not been removed.
[0072] One of the advantages of the "graph" approach to representing users of the socio-transactional service discussed above lies in its ability to take into account the "social / relational" dimension of interactions within the service through the integration of topological and semantic characteristics of this service. The characteristics associated with the nodes (representing users) and the edges (representing interactions between users) of the graph are extracted as feature vectors associated with the nodes / edges of the graph. Comparisons are then made in the form of distance measurements between feature vectors, and user / interaction deletions consist of pruning the graph by removing nodes and edges.
[0073] Once the group of fraudsters is identified, the identification device provides the service supervisor, via its human-machine interface, with information relating to the identified group of fraudsters in the form of a list of the individuals concerned or a graphical representation of the subgraph of interest. This allows the supervisor to know which users of the ST service are suspected of bank account fraud and to take appropriate action (reporting information to a bank fraud expert, implementing informational and / or corrective actions, etc.).
[0074] The identification device then proceeds, in step E4 (labeled "M0_ACT" in the figure), to assign a label to one or more users of the identified group of fraudsters, for the purpose of implementing actions. This label assignment can be performed automatically based on at least one predefined assignment criterion in the algorithm or manually by the service supervisor via their human-machine interface. The users to whom a label has been assigned form a subgroup of interest concerned by the implementation of the following actions: - the sending of an alert or information message to the terminals of users in the subgroup of no interest (users not identified as fraudsters) or, alternatively, to the terminals of all users of the ST service; - the exclusion of users from the ST service in the subgroup of interest.
[0075] Thus, the identification device is configured to inform, or even alert, users of the ST service who are not identified as fraudsters, of the presence of fraudsters within the user group to which they belong. Once groups of fraudsters are detected, the process also allows other users of the ST service to be informed. As an alternative or in addition, the identification device is configured to deny access to the ST service to the group of fraudsters. This step makes it possible to identify suspicious or abusive uses of the service in near real-time and to stop them as quickly as possible.
[0076] Other types of action can of course be implemented depending on the application context and the type of interest group sought. For example, it is entirely possible to implement, within the framework of this technique, the sending of an informational message to users of a transactional service, identified as target consumers of a given product, for marketing targeting purposes.
[0077] In the particular embodiment described here, the identification process is triggered upon receipt of a request to detect a group of interest, transmitted on an ad hoc basis by the ST service supervisor. This implementation allows the supervisor to identify groups of interest in real time (or near real time) from one or more reference groups previously identified and selected by the supervisor. Indeed, before transmitting the request, the supervisor may have previously conducted a field investigation and detected a new type of fraud or misuse of the service based on the collected data. This makes it possible to construct a query for detecting groups of interest relative not to a single reference group, but to several user reference groups.This makes it possible to identify if there are other user groups with similar characteristics, and possibly to discover if a new type of fraud is currently being committed and detectable from the data collected by the SG server.
[0078] As an alternative or complementary measure, this request can be generated automatically and periodically over time, in order to trigger the algorithm regularly, for example once a day, once a week, or once a month, as needed. The triggering frequency is defined in advance and known to the identification device. Such a implementation This system allows for the regular monitoring of new interest groups and continuous adaptation to changes in their structure. It also enables adaptation to fraudsters' circumvention methods. Suspicious users and transactions are identified beforehand by a software process and escalated to a fraud expert for classification (fraudulent / suspicious / legitimate). Confirmed cases, or those classified as such, can then constitute the reference groups to be considered in the algorithm. The supervisor saves the identified reference groups in the database and assigns an identifier to each group type. This allows for the identification of interest groups for multiple types of reference groups simultaneously, thus optimizing the search for them.
[0079] Furthermore, in the particular embodiment described here, the feature association step (step E0) is implemented for each user of the ST service and for each interaction between users. Alternatively, such a feature association step is implemented for only some of the users of the ST service, for example, when a preliminary sorting is performed on all users of the ST service (sorting performed manually by the service supervisor or sorting performed by the device according to a predefined sorting criterion). This alternative approach, when applicable, speeds up the calculations.
[0080] Fig. 4 shows the simplified structure of an identification device 40 implementing the identification method according to the invention (for example the particular embodiment described above in relation to Fig. 3).
[0081] According to a particular embodiment, this identification device is a computer and / or electronic device that is implemented in the management server itself. Alternatively, this identification device is an independent piece of equipment connected to the management server SG.
[0082] This identification device 40 comprises a random access memory 43 (for example, RAM), a processing unit 41, equipped for example with a processor, and controlled by a computer program stored in a read-only memory 42 (for example, ROM or a hard drive). At initialization, the code instructions of the computer program are, for example, loaded into the random access memory 43 before being executed by the processor of the processing unit 41. The processing unit 41 receives as input a request REQ to detect one (or more) group(s) of interest within a set of users of a given transactional service. The processor of the processing unit 41 processes the content of the request and identifies the group of interest based on the pruning principle detailed above and generates as output information INF representative of the group of identified users of interest, an ALT alert message to user terminals and / or an EVC eviction request for users in the group of interest, according to the computer program's instructions.
[0083] This [Fig. 4] illustrates only one particular way, among several possible ways, of implementing the algorithm detailed above, in relation to [Fig. 3]. Indeed, the technique of the invention can be implemented interchangeably: - on a reprogrammable computing machine (a PC, a DSP processor, or a microcontroller) executing a program comprising a sequence of instructions, or - on a dedicated computing machine (for example a set of logic gates such as an FPGA or an ASIC, or any other hardware module).
[0084] In the case where the invention is implemented on a reprogrammable computing machine, the corresponding program (i.e. the sequence of instructions) may be stored in a removable storage medium (such as, for example, a floppy disk, a CD-ROM or a DVD-ROM) or not, this storage medium being readable partially or totally by a computer or a processor.
Claims
Demands
1. A method for identifying at least one subset of users, referred to as a group of interest, within a set of users of a transactional service via terminals (UT1-UT-n), said method being implemented in an identification device and comprising the steps of: - receiving a detection request for said at least one group of interest relative to at least one reference group of users previously identified among the users of said set, - deleting (E2) at least one user and / or an interaction, within the framework of said transactional service, between two users of said set, delivering a pruned set of users, said deleting taking into account at least one deletion rule and the result of a comparison of at least one characteristic associated with said at least one user and / or with said at least one interaction with at least one characteristic representative of said reference group,- identification (E3) of said at least one interest group among the users of said pruned set.
2. Identification method according to claim 1 comprising a preliminary step, implemented for at least one user of a plurality of users of said set and / or for at least one interaction between two users of said plurality of users, of associating (EO) at least one feature to said at least one user and / or to said at least one interaction, delivering a plurality of features associated with said at least one user and / or to said at least one interaction.
3. Identification method according to claim 1 or claim 2 wherein the deletion comprises: • for at least one user and at least one interaction between two users of a plurality of users of said set: • comparison (E1) of at least one feature associated with said at least one user or with said at least one interaction with at least one representative characteristic of said reference group, providing a proximity distance, • when the proximity distance is greater at a predetermined threshold, selection (El) of said at least one user or of said at least one interaction and addition to a subgroup of no interest, • removal of said set, according to said at least one removal rule, of the users and interactions selected in the subgroup of no interest, delivering said pruned set.
4. An identification method according to claim 3, comprising the assignment (E4) of a label to at least one user of said identified interest group, forming a subgroup of interest, said label being representative of the implementation of an action belonging to the group comprising: - the sending of an alert or information message to the terminals of users of said set or subgroup of non-interest; - the exclusion of said transactional service from users of the subgroup of interest.
5. An identification method according to any one of claims 3 and 4, comprising the creation of a graph representing said set of users, said graph comprising a plurality of nodes corresponding to said users of said transactional service and a plurality of arcs corresponding to the interactions between said users, and where: - said comparison (El) includes, for at least one node and at least one arc of said graph: • a measurement of a distance between a feature vector associated with said at least one node and / or said at least one arc and a feature vector associated with said reference group, providing a proximity distance, • a comparison of said proximity distance delivered with a predetermined threshold, delivering a positive comparison result when said proximity distance is greater than said threshold, - said deletion (E2) includes a deletion according to said at least one deletion rule, of said graph, of said at least one node and / or of said at least one edge when said comparison result is positive, delivering a pruned graph, - said identification (E3) includes the determination of at least one subgraph of interest in said pruned graph.
6. Identification method according to any one of claims 1 to 5, wherein the features associated with the users and interactions of said set belong to the group comprising: - topological features, - intrinsic features of said service, - features calculated by learning.
7. Identification method according to claim 5, wherein said at least one suppression rule belongs to the group comprising: - a so-called node rule, suppressing all users of said set present in the subgroup of no interest, - a so-called arc rule, suppressing all interactions of said set present in the subgroup of no interest, - a so-called simple rule, suppressing all users and interactions of said set present in the subgroup of no interest, - a so-called arc majority rule, suppressing all users and interactions of said set present in the subgroup of no interest except the nodes having a ratio of neighbors absent from the subgroup of no interest greater than a predefined threshold.
8. A device for identifying at least one subset of users, referred to as a group of interest, within a set of users of a transactional service via terminals, said identification device comprising at least one processor configured to:
9.
10. - receive a request to detect said at least one group of interest relative to at least one reference group of users previously identified among the users of said set, - to remove at least one user and / or interaction, within the framework of said transactional service, between two users of said set, delivering a pruned set of users, said removal taking into account at least one removal rule and the result of a comparison of at least one characteristic associated with said at least one user and / or said at least one interaction with at least one characteristic representative of said reference group, - identify said at least one interest group among the users of said pruned set. Product computer program, comprising program code instructions for implementing the method according to at least one of claims 1 to 7, when said program is executed on a computer. Computer-readable and non-transient storage medium, storing a computer program product according to claim 9.
Citation Information
Patent Citations
Fraudulent customer group identification method and device, terminal equipment and computer storage medium
CN113592517A
Fraud community discovery method and device, equipment and storage medium
CN114077709A