Abnormal user determination method and device, medium and program product
By dynamically determining the minimum support threshold and frequent subsequence matching in financial activities, the problem of low accuracy in abnormal user identification in the existing technology is solved, and more efficient and flexible abnormal user identification is achieved.
Patent Information
- Application Number
- CN202510688091.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
AI Technical Summary
In the prior art, when determining whether user behavior data meets abnormal user identification rules, there is a problem of low accuracy, resulting in inaccurate identification of abnormal users.
By dynamically determining the current minimum support threshold based on the current number of behavioral sequences in the database to be analyzed formed by users participating in financial activities and the attributes of the target financial activities, combined with the recursive relationship of frequent elements, frequent subsequences are mined and matched with pre-determined abnormal behavior sequences to identify abnormal users.
The accuracy and flexibility of abnormal user identification are improved, ensuring that frequent subsequences can accurately represent the behavioral sequences in the database to be analyzed, reducing the amount of data, and improving the accuracy and adaptability of abnormal user identification.
Smart Images

Figure CN120596530A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to an abnormal user identification method, device, medium and program product. Background Art
[0002] In the financial sector, certain activities held by financial institutions are often attacked by abnormal users. Therefore, it is very important to identify abnormal users based on their behavioral data.
[0003] In related technologies, it is possible to determine whether the user's behavior data meets pre-set abnormal user identification rules. If so, it means that the user is an abnormal user.
[0004] However, since the amount of user behavior data is generally large, it is difficult to accurately determine whether the user behavior data meets the abnormal user identification rules, which leads to low accuracy in identifying abnormal users. Summary of the Invention
[0005] The present invention provides a method, device, medium and program product for determining abnormal users, so as to solve the technical problem of low accuracy in determining abnormal users existing in the related art.
[0006] According to one aspect of the present invention, a method for determining an abnormal user is provided, the method comprising:
[0007] Determining a current minimum support threshold based on the current number of behavior sequences in the database to be analyzed formed by the user's participation in the target financial activity and the attributes of the target financial activity;
[0008] Determine a first frequent element in the database to be analyzed, and for each first frequent element, recursively determine respective second frequent elements from the projected database; wherein a frequent element is an element whose number of occurrences in the database to be analyzed or the projected database is greater than the current minimum support threshold, and the projected database includes a behavior subsequence in the database to be analyzed prefixed by the first frequent element, or a behavior subsequence in the projected database in a previous recursive process prefixed by the second frequent element in a previous recursive process;
[0009] Arrange each of the first frequent elements and its corresponding second frequent elements having a recursive relationship in chronological order to obtain a frequent subsequence corresponding to the database to be analyzed;
[0010] For each of the frequent subsequences, if the predetermined abnormal behavior sequence is a subsequence of the frequent subsequence, the user corresponding to the frequent subsequence is determined to be an abnormal user.
[0011] According to another aspect of the present invention, there is provided an abnormal user determination device, the device comprising:
[0012] A first determination module is configured to determine a current minimum support threshold based on the current number of behavior sequences in the database to be analyzed formed by users participating in a target financial activity and the attributes of the target financial activity;
[0013] a second determining module, configured to determine a first frequent element in the database to be analyzed, and recursively determine respective second frequent elements from a projected database for each first frequent element; wherein a frequent element is an element whose number of occurrences in the database to be analyzed or the projected database is greater than the current minimum support threshold, and the projected database includes a behavior subsequence in the database to be analyzed prefixed by the first frequent element, or a behavior subsequence in the projected database in a previous recursive process prefixed by the second frequent element in a previous recursive process;
[0014] a third determining module, configured to arrange each of the first frequent elements and its corresponding second frequent elements having a recursive relationship in chronological order to obtain a frequent subsequence corresponding to the database to be analyzed;
[0015] The fourth determining module is configured to, for each of the frequent subsequences, determine that a user corresponding to the frequent subsequence is an abnormal user if a predetermined abnormal behavior sequence is a subsequence of the frequent subsequence.
[0016] According to another aspect of the present invention, an electronic device is provided, comprising:
[0017] at least one processor; and
[0018] a memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the abnormal user determination method according to any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and the computer program is configured to enable a processor to implement the abnormal user determination method according to any embodiment of the present invention when executed.
[0021] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the abnormal user determination method according to any embodiment of the present invention is implemented.
[0022] The technical solution of the embodiment of the present invention, on the one hand, can obtain the frequent subsequences corresponding to the database to be analyzed based on the first frequent element in the database to be analyzed and its corresponding second frequent element with a recursive relationship, and then determine the abnormal users based on each frequent subsequence and the abnormal behavior sequence. Since the frequent subsequences can accurately represent the behavior sequences in the database to be analyzed, and the data volume of the frequent subsequences is necessarily smaller than the data volume of the database to be analyzed, it is possible to accurately judge whether the abnormal behavior sequence is a subsequence of the frequent subsequence, thereby improving the accuracy of the determined abnormal users; on the other hand, the method can determine the current minimum support threshold based on the current number of behavior sequences and the attributes of the target financial activity, so that the minimum support threshold when determining each frequent element matches the database to be analyzed and the target financial activity, thereby improving the flexibility and adaptability of the abnormal user determination method.
[0023] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0025] Figure 1 This is a flow chart of a method for determining abnormal users provided by an embodiment of the present invention;
[0026] Figure 2 is a schematic diagram of segmenting behavior data in an embodiment of the present invention;
[0027] Figure 3 is a schematic diagram of a target computer node in a processing cluster according to an embodiment of the present invention;
[0028] Figure 4 is a flow chart of another abnormal user identification method provided by an embodiment of the present invention;
[0029] Figure 5 1 is a schematic structural diagram of an abnormal user identification device provided by an embodiment of the present invention;
[0030] Figure 6 The figure is a schematic structural diagram of an electronic device for implementing the abnormal user determination method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of data in the embodiments of the present invention comply with the relevant provisions of national laws and regulations.
[0033] To facilitate subsequent understanding, the relevant terms involved in this embodiment are explained below.
[0034] Sequence Pattern Mining: Sequence pattern mining is the process of finding all frequent sequences in a database that meet a given minimum support threshold. These frequent sequences can reveal underlying patterns and regularities in the data.
[0035] Frequent Sequences: In a database, if the support (i.e., frequency of occurrence) of a sequence exceeds a user-defined minimum support threshold, then the sequence is called a frequent sequence. In sequential pattern mining, the goal is to find all frequent sequences.
[0036] Support: Support refers to the frequency with which a sequence appears in a database. If the ratio of the number of times sequence A appears in the database to the total number of sequences in the database exceeds the minimum support threshold set by the user, then sequence A is called a frequent sequence.
[0037] Projected Database: A core concept in sequential pattern mining is the projected database. At each step in the algorithm, a new projected database is generated based on the frequent sequences found in the previous step. This database contains only subsequences prefixed by a specific frequent sequence, thus reducing the search space and improving mining efficiency.
[0038] Prefix: In the sequence pattern mining algorithm, the prefix refers to the front part of the sequence. Each recursive mining is based on the prefix obtained in the previous mining.
[0039] Suffix: In contrast to the prefix, the suffix refers to the remaining part of the sequence. In the sequential pattern mining algorithm, it is checked whether the suffix of each frequent sequence is also frequent. If not, pruning can be performed to avoid unnecessary calculations.
[0040] Apriori Property: The Apriori property states that if all subsequences of a sequence are frequent, then the sequence itself must also be frequent. This property is often used in sequential pattern mining algorithms to reduce the number of candidate sequences.
[0041] Minimum Support Threshold: The specified minimum support threshold is used to determine whether a sequence is a frequent sequence. A sequence is considered frequent only when its support exceeds this threshold.
[0042] Figure 1 This is a flow chart of an abnormal user identification method provided by an embodiment of the present invention. This embodiment is applicable to scenarios where abnormal users participating in financial activities are identified. The method can be executed by an abnormal user identification device, which can be implemented in the form of hardware and / or software. The abnormal user identification device can be configured in an electronic device, such as a computer device. Figure 1 As shown, the method includes the following steps 101 to 104.
[0043] Step 101: Determine a current minimum support threshold based on the current number of behavior sequences in the database to be analyzed formed by users participating in the target financial activity and the attributes of the target financial activity.
[0044] In this embodiment, the target financial activity refers to a financial activity whose participating users need to be analyzed for abnormalities. Financial activities in this embodiment may include financing activities, investment activities, financial market transactions, risk management, asset management, and payment and settlement. More specifically, the financial activity in this embodiment may be a financial marketing activity conducted by a financial institution. For example, the target financial activity in this embodiment may be a marketing activity that a user participates in through a financial institution's online channel, such as a mobile banking app.
[0045] In this embodiment, users refer to at least a portion of users participating in the target financial activity. During their participation in the target financial activity, user behaviors form behavior sequences. Behavior sequences represent the mapping between user behaviors and time. In this embodiment, a behavior sequence may include at least one event (also referred to as an item set). Each event may include at least one element (also referred to as an item). For ease of description, in this embodiment, the collection of user behavior sequences is referred to as the database to be analyzed.
[0046] In this embodiment, a pattern mining algorithm based on prefix projection in sequential pattern mining is used to identify abnormal users. This algorithm requires specifying a minimum support threshold. In related sequential pattern mining techniques, the minimum support threshold is typically a fixed value. However, in financial activity scenarios under big data access models, a fixed minimum support threshold can easily lead to inaccurate frequent subsequences subsequently identified, and thus, low accuracy in identifying abnormal users.
[0047] In one implementation, in order to improve the flexibility of the abnormal user identification process and its adaptability to the financial field, the current minimum support threshold can be determined based on the current number of behavior sequences in the database to be analyzed and the attributes of the target financial activity.
[0048] In this embodiment, the current number of behavior sequences in the database to be analyzed refers to the number of behavior sequences included in the database to be analyzed. The attributes of financial activities are used to represent information such as the importance and influence of financial activities.
[0049] In another implementation, the current minimum support threshold may be determined based on the length of the behavior sequence in the database to be analyzed and the attributes of the target financial activity.
[0050] The implementation process of step 101 may include steps 1011 to 1012.
[0051] Step 1011: Determine the minimum support threshold corresponding to the current number of behavior sequences according to the current number of behavior sequences and the mapping relationship between the number of behavior sequences and the minimum support threshold.
[0052] In the mapping relationship between the number of behavior sequences and the minimum support threshold, the number of behavior sequences is proportional to the minimum support threshold.
[0053] In the above mapping relationship, the more behavior sequences there are, the larger the minimum support threshold is, and the fewer behavior sequences there are, the smaller the minimum support threshold is. The purpose of this setting is: the more behavior sequences there are, the more likely an element is to appear in the sequence. If the minimum support threshold is set too low, it will lead to misjudgment (i.e., false positive), that is, elements that are not frequent elements are judged as frequent elements, resulting in a decrease in the accuracy of the determined frequent elements; the fewer behavior sequences there are, the fewer times an element appears in the sequence. If the minimum support threshold is set too high, it will lead to missed judgment, that is, elements that are frequent elements are judged as not frequent elements, which will also reduce the accuracy of the determined frequent elements.
[0054] Step 1012: Determine the current minimum support threshold according to the weight coefficient corresponding to the attribute of the target financial activity and the minimum support threshold corresponding to the current number of behavior sequences.
[0055] Furthermore, after determining the minimum support threshold corresponding to the current number of behavior sequences, this embodiment also considers the impact of the attributes of the target financial activity on the minimum support threshold. This embodiment provides a correspondence between financial activity attributes and weight coefficients. The minimum support threshold corresponding to the current number of behavior sequences is adjusted based on the weight coefficients corresponding to the attributes of the target financial activity to obtain the current minimum support threshold. This embodiment does not impose any restrictions on the size of the weight coefficient corresponding to the attributes of the target financial activity; it can be a number greater than 0 and less than or equal to 1, or a number greater than 1.
[0056] Optionally, the current minimum support threshold may be determined as a value obtained by rounding up the product of the weight coefficient corresponding to the attribute of the target financial activity and the minimum support threshold corresponding to the current number of behavior sequences.
[0057] For example, if the attribute of a target financial activity represents that the activity has a relatively large influence, then the weight coefficient corresponding to the attribute of the target financial activity can be a number greater than 0 and less than or equal to 1, that is, the minimum support threshold corresponding to the current number of behavior sequences is reduced to avoid missed judgments.
[0058] Based on the implementation process of step 1011 and step 1012, it is possible to quickly and accurately determine the current minimum support threshold according to the current number of behavior sequences and the attributes of the target financial activity.
[0059] Optionally, before step 101 , the abnormal user determination method provided in this embodiment further includes the following steps 101 a to 101 c.
[0060] Step 101a: Obtain a dataset of users participating in target financial activities.
[0061] The data set includes multiple behavioral data arranged in chronological order.
[0062] In this embodiment, the difference between the behavior data and the behavior sequence is that the duration corresponding to the behavior data is greater than the duration corresponding to the behavior sequence.
[0063] Step 101b: Split each behavior data according to the data segmentation strategy to obtain a behavior sequence.
[0064] In order to avoid low accuracy of subsequent operations due to excessive data volume, in step 101b, each behavior data can be segmented according to a data segmentation strategy to obtain a behavior sequence.
[0065] Optionally, the data segmentation strategy in this embodiment is used to indicate a time window for data segmentation, for example, data segmentation is performed according to time periods such as days, weeks, and quarters.
[0066] Optionally, the data segmentation strategy in this embodiment is used to indicate the amount of data to be segmented, for example, data segmentation is performed according to the last ten or last hundred transactions.
[0067] Step 101c: Obtain a database to be analyzed based on the behavior sequence.
[0068] The database to be analyzed includes at least one behavior sequence corresponding to the same time period.
[0069] After obtaining the behavior sequences, the behavior sequences corresponding to the same time period are determined as the database to be analyzed.
[0070] Optionally, to improve data accuracy, after step 101b, the behavior sequence can be preprocessed to remove abnormal values, irrelevant field data, etc., to obtain a processed behavior sequence. Correspondingly, in step 101c, the database to be analyzed is obtained based on the preprocessed behavior sequence.
[0071] Figure 2 Schematic diagram of segmenting behavior data in an embodiment of the present invention. Figure 2 As shown in Figure 1, assume that there are five behavioral data: behavioral data 1, behavioral data 2, behavioral data 3, behavioral data 4, and behavioral data 5. The duration of these behavioral data is T. Each behavioral data is divided according to the time window t1. Assuming that T includes two t1s, the two databases to be analyzed 1 and 2 to be analyzed are as follows: Figure 2 shown.
[0072] The above steps 101a to 101c can achieve the acquisition of the best behavior sequence of a certain time period or a certain length, which facilitates the subsequent determination of frequent subsequences, further ensures the accuracy of the abnormal user determination, and increases the flexibility of the abnormal user determination process.
[0073] Step 102: Determine the first frequent elements in the database to be analyzed, and for each first frequent element, recursively determine each second frequent element from the projected database.
[0074] Frequent elements are elements whose occurrence counts in the database to be analyzed or the projected database are greater than the current minimum support threshold. The projected database includes the behavior subsequence prefixed by the first frequent element in the database to be analyzed, or the behavior subsequence prefixed by the second frequent element in the previous recursive process in the projected database from the previous recursive process.
[0075] It should be noted that when determining the number of occurrences of an element in the database to be analyzed or the projected database, if the element appears multiple times in the same behavior sequence, it is only considered to have appeared once. For example, suppose the database to be analyzed has two behavior sequences: Behavior Sequence 1: ABC, Behavior Sequence 2: AAC. Then, the number of occurrences of element A in the database to be analyzed is 2, the number of occurrences of element B in the database to be analyzed is 1, and the number of occurrences of element C in the database to be analyzed is 2. Assuming the current minimum support threshold is 1, then elements A and C are the first frequent elements. The number of occurrences of an element in the database to be analyzed or the projected database can also be called the support of the element.
[0076] In the first recursive process, at least one first frequent element is determined from the database to be analyzed. Then, for each first frequent element, each second frequent element is recursively determined from the projected database. In this recursive process, the projected database includes a subsequence of behaviors in the database to be analyzed that begins with the first frequent element. This projected database can also be referred to as a projection database of first frequent elements.
[0077] The projection database is determined in the following way: for each behavior sequence in the database to be analyzed, if it does not contain the first frequent element, the sequence becomes empty; if it contains the first frequent element, the position of the first frequent element is determined; if the first frequent element is in the item set, the first frequent element is represented by the placeholder "_"; if the first frequent element is not in the item set, the first frequent element is deleted; at the same time, regardless of whether the first frequent element is in the item set, all elements before the first frequent element are deleted.
[0078] In the second recursive process, the second frequent elements are determined from the projected database of the first frequent element. This determination process is similar to that for the first frequent element. For each second frequent element, the projected database for that second frequent element is determined from the projected database of the first frequent element. This determination process is similar to that for the first frequent element and will not be repeated here.
[0079] In the third recursive process, the second frequent element is determined from the projected database of the second frequent element determined in the second recursive process, that is, from the subsequence of the projected database of the first frequent element (the projected database in the previous recursive process) prefixed with the second frequent element in the previous recursive process. For each second frequent element, the projected database of the second frequent element determined in the second recursive process is then determined. This process is repeated in this way until all second frequent elements have been determined.
[0080] The size of the projected database affects computer device memory usage and the efficiency of determining the second frequent element. If the projected database is too large, it may lead to insufficient memory. Optionally, to improve the efficiency of determining frequent subsequences, a distributed computing framework can be used to determine the first frequent element and the second frequent element in step 102: based on the load information of each computer node in the processing cluster, a target computer node is determined to participate in the determination of abnormal users; using the target computer node and a distributed algorithm, the first frequent element in the database to be analyzed is determined, and for each first frequent element, each second frequent element is recursively determined from the projected database.
[0081] When determining target computer nodes for determining abnormal users based on load information of each computer node in the processing cluster, n computer nodes with loads less than a preset threshold or with smaller loads may be determined as target computer nodes.
[0082] Figure 3 FIG is a schematic diagram of a target computer node in a processing cluster according to an embodiment of the present invention. Figure 3 As shown, assuming that there are m computer nodes in the processing cluster, and the load of three computer nodes is less than a preset threshold, these three computer nodes are determined as target computer nodes 31. m is an integer greater than or equal to 3. These three target computer nodes 31 can respectively determine a portion of the second frequent elements.
[0083] In the process of determining the first and second frequent elements, we employ pattern mining using prefix projection. This technique divides the database to be analyzed into smaller projection databases, and then recursively mines the second frequent elements on these projection databases. This method significantly reduces the search space and storage requirements, resulting in high mining efficiency.
[0084] Step 103: Arrange each first frequent element and its corresponding second frequent element having a recursive relationship in chronological order to obtain a frequent subsequence corresponding to the database to be analyzed.
[0085] In this embodiment, the second frequent element with a recursive relationship refers to a second frequent element with a dependency relationship. That is, subsequent second frequent elements are determined based on the second frequent elements in the previous recursive process. For example, assume that there are two first frequent elements in the database to be analyzed: A1 and A2. In the second recursive process, three second frequent elements are determined for A1: A11, A12, and A13. In the third recursive process, one second frequent element A111 is determined for A11, one second frequent element A121 is determined for A12, and one second frequent element A131 is determined for A13. Assuming that the recursive termination condition is met at this point, in the above example, the second frequent elements with a recursive relationship are: A11 and A111, A12 and A121, and A13 and A131. The frequent subsequences derived based on the first frequent element A1 are: A1A11A111, A1A12A121, and A1A13A131. The process of obtaining a frequent subsequence based on the second frequent element A2 is similar to the above process and will not be described again here.
[0086] Step 104: For each frequent subsequence, if the predetermined abnormal behavior sequence is a subsequence of the frequent subsequence, the user corresponding to the frequent subsequence is determined to be an abnormal user.
[0087] In this embodiment, the abnormal behavior sequence can be determined in advance based on the abnormal behavior data. The abnormal behavior sequence in this embodiment can be determined based on an artificial intelligence model, or based on expert rules, or based on other pattern recognition algorithms. For example, the abnormal behavior sequence in this embodiment can be<a,a,c,c,d> wait.
[0088] In this embodiment, if a first sequence is a sequence obtained by deleting zero or more elements from a second sequence without changing the order of elements, then the first sequence is a subsequence of the second sequence.
[0089] Optionally, in order to maintain real-time performance, the abnormal behavior sequence in this embodiment may be continuously updated at a preset frequency.
[0090] Optionally, in order to further improve accuracy, the number of events included in the abnormal behavior sequence in this embodiment may be less than a preset event number threshold.
[0091] Optionally, the correspondence between frequent subsequences and users is pre-determined. After identifying abnormal users, a table of user behavior labels for a specific time period can be obtained, such as whether there were abnormal behaviors in the past day, week, or month. Normal behavior is marked as 1, and abnormal behavior is marked as 0. This user behavior label data can be associated with the user information table to provide label support for subsequent data asset development for wide tables, indicators, risk control, and other areas, thereby increasing the diversity of analytical data.
[0092] The abnormal user determination method provided in this embodiment can, on the one hand, obtain the frequent subsequences corresponding to the database to be analyzed based on the first frequent element in the database to be analyzed and its corresponding second frequent element with a recursive relationship, and then determine the abnormal users based on each frequent subsequence and the abnormal behavior sequence. Since the frequent subsequences can accurately represent the behavior sequences in the database to be analyzed, and the data volume of the frequent subsequences is necessarily smaller than the data volume of the database to be analyzed, it is possible to accurately determine whether the abnormal behavior sequence is a subsequence of the frequent subsequence, thereby improving the accuracy of the determined abnormal users; on the other hand, the method can determine the current minimum support threshold based on the current number of behavior sequences and the attributes of the target financial activity, so that the minimum support threshold when determining each frequent element matches the database to be analyzed and the target financial activity, thereby improving the flexibility and adaptability of the abnormal user determination method.
[0093] Figure 4 This is a flow chart of another abnormal user determination method provided by an embodiment of the present invention. Figure 1 Based on the embodiment shown and various optional implementations, the process of determining the first frequent element and the second frequent element is described in detail. Figure 4 As shown, the abnormal user determination method provided by this embodiment includes the following steps 401 to 406.
[0094] Step 401: Determine a current minimum support threshold based on the current number of behavior sequences in the database to be analyzed formed by users participating in the target financial activity and the attributes of the target financial activity.
[0095] The implementation process and technical principles of step 401 are similar to those of step 101 and will not be repeated here.
[0096] This embodiment describes in detail the process of determining the first frequent element and each second frequent element in step 102 through steps 402 to 404. For ease of description, the second frequent element in this embodiment includes the current second frequent element determined in the current recursive process and / or the next second frequent element to be determined in the next recursive process.
[0097] Step 402: Determine the first frequent elements in the database to be analyzed, and for each first frequent element, determine a projection database of the first frequent element from the database to be analyzed.
[0098] The process of determining the projection database of the first frequent element is shown in step 102 and will not be described in detail here.
[0099] If the number of first frequent elements is R, then in step 402, a projection database of R first frequent elements may be determined.
[0100] Optionally, after determining the first frequent element in the database to be analyzed, and before determining a projected database of the first frequent element from the database to be analyzed for each first frequent element, the method further includes the following steps: determining infrequent elements in the database to be analyzed; and deleting the infrequent elements from the behavior sequence of the database to be analyzed to obtain a pruned database to be analyzed. Correspondingly, determining the projected database of the first frequent element from the database to be analyzed includes determining the projected database of the first frequent element from the pruned database to be analyzed. Infrequent elements herein refer to elements whose number of occurrences in the database to be analyzed is less than or equal to a current minimum support threshold.
[0101] Step 403: If the current second frequent element can be determined from the projected database of the first frequent element and the recursive termination condition is not satisfied, then the projected database of the current second frequent element is determined from the projected database of the first frequent element.
[0102] During the actual implementation of this method, there may be a situation where the current second frequent element cannot be determined from the projection database of a first frequent element. That is, the number of occurrences of all elements in the projection database of the first frequent element is less than the current minimum support threshold. In this case, the first frequent element can be determined as a frequent subsequence of the database to be analyzed.
[0103] Optionally, the recursion termination condition may be that a preset number of recursions is met, or a preset recursion duration is met.
[0104] The process of determining the projected database of the current second frequent element from the projected database of the first frequent element can be as follows: for each behavioral subsequence of the projected database of the first frequent element, if it does not contain the current second frequent element, the sequence becomes empty; if it contains the current second frequent element, the position of the first current second frequent element is determined; if the first current second frequent element is in the item set, the first current second frequent element is represented by a placeholder "_"; if the first current second frequent element is not in the item set, the first current second frequent element is deleted; at the same time, regardless of whether the first current second frequent element is in the item set, all elements before the first current second frequent element are deleted.
[0105] Optionally, after determining the current second frequent element from the projected database of the first frequent element and before determining the current projected database of the second frequent element from the projected database of the first frequent element, the abnormal user identification method provided in this embodiment further includes: determining infrequent elements in the projected database of the first frequent element; and deleting infrequent elements from the behavioral subsequence of the projected database of the first frequent element to obtain a pruned projected database of the first frequent element. Correspondingly, determining the current projected database of the second frequent element from the projected database of the first frequent element includes: determining the current projected database of the second frequent element from the pruned projected database of the first frequent element. Infrequent elements herein refer to elements whose number of occurrences in the projected database of the first frequent element is less than or equal to the current minimum support threshold.
[0106] Step 404: If the next second frequent element can be determined from the projected database of the current second frequent element and the recursive termination condition is not met, then the projected database of the next second frequent element is determined from the projected database of the current second frequent element, and the next second frequent element is used as the new current second frequent element. The process then returns to this step and continues until the next second frequent element cannot be determined or the recursive termination condition is met.
[0107] Among them, the frequent elements are elements whose number of occurrences in the database to be analyzed or the projected database is greater than the current minimum support threshold. The projected database includes the behavior subsequence prefixed by the first frequent element in the database to be analyzed, or the behavior subsequence prefixed by the second frequent element in the previous recursive process in the projected database in the previous recursive process.
[0108] The implementation process of step 404 is similar to that of step 403 , which realizes recursively determining each second frequent element from the projection database.
[0109] It should be noted that in each recursive process, after determining the next second-frequent element, the next infrequent element in the projected database of the current second-frequent element is also determined. The next infrequent element is then deleted from the behavioral subsequence of the projected database of the current second-frequent element to obtain a pruned projected database of the current second-frequent element. Correspondingly, the projected database of the next second-frequent element is determined from the pruned projected database of the current second-frequent element.
[0110] The above-mentioned method of pruning the database to be analyzed or the projected database can reduce the search space and improve the efficiency of identifying abnormal users.
[0111] In one embodiment, the recursive termination condition is that the number of a first frequent element and its corresponding second frequent element having a recursive relationship is greater than a preset sequence length threshold. Accordingly, the method further includes the following steps: determining a preset sequence length threshold based on the length of the abnormal behavior sequence. Exemplarily, the length of the abnormal behavior sequence can be used as the preset sequence length threshold. Alternatively, a weighted value of the length of the preset sequence length threshold can be used as the preset sequence length threshold.
[0112] This implementation can facilitate comparison of a frequent subsequence consisting of a first frequent element and its corresponding second frequent element having a recursive relationship with an abnormal behavior sequence, thereby improving efficiency.
[0113] Step 405: Arrange each first frequent element and its corresponding second frequent element with a recursive relationship in chronological order to obtain a frequent subsequence corresponding to the database to be analyzed.
[0114] Step 406: For each frequent subsequence, if the predetermined abnormal behavior sequence is a subsequence of the frequent subsequence, the user corresponding to the frequent subsequence is determined to be an abnormal user.
[0115] The implementation process and technical principles of step 405 and step 103, and step 406 and step 104 are similar, and will not be repeated here.
[0116] The abnormal user identification method provided by this embodiment determines the first frequent element in the database to be analyzed, and for each first frequent element, determines a projection database of the first frequent element from the database to be analyzed. If the current second frequent element can be determined from the projection database of the first frequent element and the recursive termination condition is not met, the projection database of the current second frequent element is determined from the projection database of the first frequent element. If the next second frequent element can be determined from the projection database of the current second frequent element and the recursive termination condition is not met, the projection database of the next second frequent element is determined from the projection database of the current second frequent element, and the next second frequent element is used as the new current second frequent element. The method returns to this step and continues until the next second frequent element cannot be determined or the recursive termination condition is met. This method implements the partitioning of the database to be analyzed into smaller projection databases using the prefix projection technique, and then recursively mines second frequent elements from these projection databases. This method significantly reduces the search space and storage requirements, improves mining efficiency, and thus improves the efficiency of identifying abnormal users.
[0117] Figure 5 Schematic diagram of a structure of an abnormal user identification device provided by an embodiment of the present invention. The device is set in a computer device. Figure 5 As shown, the abnormal user determination device provided by this embodiment includes the following modules: a first determination module 51 , a second determination module 52 , a third determination module 53 and a fourth determination module 54 .
[0118] The first determining module 51 is configured to determine a current minimum support threshold according to the current number of behavior sequences in the database to be analyzed formed by users participating in a target financial activity and the attributes of the target financial activity.
[0119] The second determining module 52 is configured to determine the first frequent elements in the database to be analyzed, and recursively determine each second frequent element from the projected database for each first frequent element.
[0120] A frequent element is an element whose number of occurrences in the database to be analyzed or the projected database is greater than the current minimum support threshold. The projected database includes a behavior subsequence in the database to be analyzed prefixed with the first frequent element, or a behavior subsequence in the projected database in the previous recursive process prefixed with the second frequent element in the previous recursive process.
[0121] The third determining module 53 is configured to arrange each of the first frequent elements and its corresponding second frequent elements having a recursive relationship in chronological order to obtain a frequent subsequence corresponding to the database to be analyzed.
[0122] The fourth determining module 54 is configured to, for each of the frequent subsequences, determine that the user corresponding to the frequent subsequence is an abnormal user if the predetermined abnormal behavior sequence is a subsequence of the frequent subsequence.
[0123] In one embodiment, the device further includes: an acquisition module, a data segmentation module, and a fifth determination module.
[0124] An acquisition module is configured to acquire a dataset of users' participation in target financial activities. The dataset includes multiple behavioral data points arranged in chronological order. A data segmentation module is configured to segment each behavioral data point according to a data segmentation strategy to obtain behavioral sequences. A fifth determination module is configured to obtain the database to be analyzed based on the behavioral sequences. The database to be analyzed includes at least one behavioral sequence corresponding to the same time period.
[0125] In one embodiment, the first determination module 51 is specifically used to: determine the minimum support threshold corresponding to the current number of behavior sequences based on the current number of behavior sequences and the mapping relationship between the number of behavior sequences and the minimum support threshold, wherein the number of behavior sequences in the mapping relationship between the number of behavior sequences and the minimum support threshold is proportional to the minimum support threshold; determine the current minimum support threshold based on the weight coefficient corresponding to the attribute of the target financial activity and the minimum support threshold corresponding to the current number of behavior sequences.
[0126] In one embodiment, the second determination module 52 is specifically configured to: determine a target computer node that participates in determining abnormal users based on load information of each computer node in the processing cluster; determine the first frequent element in the database to be analyzed using the target computer node and a distributed algorithm; and recursively determine each second frequent element from the projected database for each first frequent element.
[0127] In one embodiment, the second frequent element includes the current second frequent element and / or the next second frequent element. The second determination module 52 is specifically configured to: determine the first frequent elements in the database to be analyzed, and for each first frequent element, determine a projection database of the first frequent element from the database to be analyzed; if the current second frequent element can be determined from the projection database of the first frequent element and the recursive termination condition is not satisfied, determine the projection database of the current second frequent element from the projection database of the first frequent element; if the next second frequent element can be determined from the projection database of the current second frequent element and the recursive termination condition is not satisfied, determine the projection database of the next second frequent element from the projection database of the current second frequent element, set the next second frequent element as the new current second frequent element, and return to this step until the next second frequent element cannot be determined or the recursive termination condition is satisfied.
[0128] In one embodiment, the device further includes: a sixth determining module and a deleting module.
[0129] The sixth determination module is configured to determine infrequent elements in the projected database of the first frequent element. The deletion module is configured to delete the infrequent elements from the behavioral subsequence of the projected database of the first frequent element to obtain a pruned projected database of the first frequent element. Regarding determining the projected database of the current second frequent element from the projected database of the first frequent element, the second determination module 52 is specifically configured to determine the projected database of the current second frequent element from the pruned projected database of the first frequent element.
[0130] In one embodiment, the recursive termination condition is that the number of the first frequent element and its corresponding second frequent element having a recursive relationship is greater than a preset sequence length threshold. The device also includes a seventh determination module for determining the preset sequence length threshold based on the length of the abnormal behavior sequence.
[0131] The abnormal user determination device provided in the embodiment of the present invention can execute the abnormal user determination method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0132] Figure 61 is a schematic diagram of the structure of an electronic device that implements the abnormal user determination method of an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0133] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0134] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0135] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the abnormal user determination method.
[0136] In some embodiments, the abnormal user determination method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the abnormal user determination method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the abnormal user determination method in any other appropriate manner (e.g., by means of firmware).
[0137] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0138] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0139] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0141] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0142] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0143] An embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the abnormal user determination method provided by any embodiment of the present invention is implemented.
[0144] The computer program product may be implemented in a computer program code for performing the operations of the present invention written in one or more programming languages, or a combination thereof, including object-oriented programming languages and conventional procedural programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0145] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0146] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for identifying abnormal users, characterized in that: The method comprises: Determining a current minimum support threshold based on the current number of behavior sequences in the database to be analyzed formed by the user's participation in the target financial activity and the attributes of the target financial activity; Determine a first frequent element in the database to be analyzed, and for each first frequent element, recursively determine respective second frequent elements from the projected database; wherein a frequent element is an element whose number of occurrences in the database to be analyzed or the projected database is greater than the current minimum support threshold, and the projected database includes a behavior subsequence in the database to be analyzed prefixed by the first frequent element, or a behavior subsequence in the projected database in a previous recursive process prefixed by the second frequent element in a previous recursive process; Arrange each of the first frequent elements and its corresponding second frequent elements having a recursive relationship in chronological order to obtain a frequent subsequence corresponding to the database to be analyzed; For each of the frequent subsequences, if the predetermined abnormal behavior sequence is a subsequence of the frequent subsequence, the user corresponding to the frequent subsequence is determined to be an abnormal user.
2. The method according to claim 1, characterized in that Before determining the current minimum support threshold based on the current number of behavior sequences in the database to be analyzed formed by the user's participation in the target financial activity and the attributes of the target financial activity, the method further includes: Obtaining a data set of users participating in a target financial activity; wherein the data set includes a plurality of behavioral data arranged in chronological order; Splitting each of the behavior data according to the data segmentation strategy to obtain the behavior sequence; The database to be analyzed is obtained according to the behavior sequence; wherein the database to be analyzed includes at least one behavior sequence corresponding to the same time period.
3. The method according to claim 1, characterized in that Determining the current minimum support threshold based on the current number of behavior sequences in the database to be analyzed formed by the user's participation in the target financial activity and the attributes of the target financial activity includes: Determining the minimum support threshold corresponding to the current number of behavior sequences according to the current number of behavior sequences and the mapping relationship between the number of behavior sequences and the minimum support threshold; wherein, in the mapping relationship between the number of behavior sequences and the minimum support threshold, the number of behavior sequences is proportional to the minimum support threshold; The current minimum support threshold is determined according to the weight coefficient corresponding to the attribute of the target financial activity and the minimum support threshold corresponding to the current number of the behavior sequences.
4. The method according to any one of claims 1 to 3, characterized in that The determining of the first frequent element in the database to be analyzed, and recursively determining each second frequent element from the projected database for each first frequent element, includes: Determining target computer nodes that participate in identifying abnormal users based on load information of each computer node in the processing cluster; The first frequent elements in the database to be analyzed are determined by using the target computer node and a distributed algorithm, and for each of the first frequent elements, respective second frequent elements are recursively determined from a projected database.
5. The method according to any one of claims 1 to 3, characterized in that The second frequent elements include the current second frequent element and / or the next second frequent element; The determining of the first frequent element in the database to be analyzed, and recursively determining each second frequent element from the projected database for each first frequent element, includes: determining a first frequent element in the database to be analyzed, and determining, for each first frequent element, a projection database of the first frequent element from the database to be analyzed; If the current second frequent element can be determined from the projection database of the first frequent element and the recursive termination condition is not satisfied, then determining the projection database of the current second frequent element from the projection database of the first frequent element; If the next second frequent element can be determined from the projected database of the current second frequent element and the recursive termination condition is not met, the projected database of the next second frequent element is determined from the projected database of the current second frequent element, and the next second frequent element is used as the new current second frequent element. The process then returns to this step and continues until the next second frequent element cannot be determined or the recursive termination condition is met.
6. The method according to claim 5, characterized in that After determining the current second frequent element from the projection database of the first frequent elements and before determining the current second frequent element projection database from the projection database of the first frequent elements, the method further includes: Determine infrequent elements in a projection database of the first frequent elements; Deleting the infrequent elements from the behavioral subsequence of the projection database of the first frequent element to obtain a pruned projection database of the first frequent element; The determining the projection database of the current second frequent element from the projection database of the first frequent element includes: The projection database of the current second frequent element is determined from the projection database of the pruned first frequent element.
7. The method according to claim 5, characterized in that The recursive termination condition is that the number of the first frequent elements and their corresponding second frequent elements with a recursive relationship is greater than a preset sequence length threshold; The method further comprises: The preset sequence length threshold is determined according to the length of the abnormal behavior sequence.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the abnormal user determination method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is configured to enable a processor to implement the abnormal user determination method according to any one of claims 1 to 7 when the computer program is executed.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the abnormal user determination method according to any one of claims 1 to 7.