Abnormal user determination method, determination apparatus, electronic device, and storage medium
By processing the initial asset information of target users to generate standard asset information, and using clustering to identify abnormal users, the problem of high error rate in assessment information caused by manual reading is solved, and the accuracy of identifying abnormal users is improved.
Patent Information
- Application Number
- CN202310368488.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-07
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2043-04-07
AI Technical Summary
In existing technologies, when banks and other enterprises identify abnormal users by manually reading asset statements, there is a high error rate in the assessment information, resulting in inaccurate identification results for abnormal users.
By processing the initial asset information of target users, standard asset information is generated. Then, multidimensional scaling analysis and hierarchical clustering are used to cluster abnormal and non-abnormal users. Combining the results of the first and second clustering, the abnormal results of users are determined, and suspicious users are further classified through readability indicators.
It reduces data extraction errors caused by manual operation, improves the accuracy of identifying abnormal users, and ensures the accuracy and reliability of clustering results.
Smart Images

Figure CN116644323B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of big data, and more particularly to an abnormal user determination method and device, electronic equipment and storage medium. BACKGROUND
[0002] In the related art, banks and other enterprises determine whether a user has transaction risk and whether the user is an abnormal user by disclosing asset information. Business personnel obtain evaluation information by reading asset reports and other disclosed asset information, and upload the evaluation information to a computer system, which determines whether the user is abnormal based on the evaluation information.
[0003] However, asset reports include multiple index categories and multiple evaluation criteria, as well as obscure textual description languages, resulting in a high error rate of evaluation information obtained manually and inaccurate determination results of abnormal users. SUMMARY
[0004] In view of the above problems, the present disclosure provides an abnormal user determination method and device, electronic equipment and storage medium.
[0005] According to a first aspect of the present disclosure, an abnormal user determination method is provided, comprising:
[0006] processing initial asset information of M target users to obtain standard asset information corresponding to the M target users, the initial asset information including asset reports of the target users, and M≥2;
[0007] clustering N abnormal class users and the M target users according to the standard asset information to obtain a first clustering result, N≥M≥2;
[0008] clustering N non-abnormal class users and the M target users according to the standard asset information to obtain a second clustering result; and
[0009] determining an abnormal result of the M target users according to the first clustering result and the second clustering result, the abnormal result including abnormal users, non-abnormal users and suspicious users.
[0010] According to an embodiment of the present disclosure, the first clustering result includes M first sub-clustering results, and the second clustering result includes M second sub-clustering results.
[0011] determining whether the M target users are abnormal according to the first clustering result and the second clustering result, comprising:
[0012] In a case where the n th first sub-clustering result represents that the n th target user belongs to the abnormal class and the n th second sub-clustering result represents that the n th target user does not belong to the non-abnormal class, the n th target user is determined as an abnormal user, M ≥ n ≥ 2;
[0013] In a case where the n th first sub-clustering result represents that the n th target user does not belong to the abnormal class and the n th second sub-clustering result represents that the n th target user belongs to the non-abnormal class, the n th target user is determined as a non-abnormal user.
[0014] In a case where the n th first sub-clustering result represents that the n th target user belongs to the abnormal class and the n th second sub-clustering result represents that the n th target user belongs to the non-abnormal class, or the n th first sub-clustering result represents that the n th target user does not belong to the abnormal class and the n th second sub-clustering result represents that the n th target user does not belong to the non-abnormal class, the n th target user is determined as a suspicious user.
[0015] According to an embodiment of the present disclosure, after determining the abnormal results of the M target users according to the first clustering result and the second clustering result, the method further comprises:
[0016] In a case where the m th target user is determined as a suspicious user, a readability index value is calculated according to the initial asset information of the m th target user, the readability index value being used to represent the readability of the initial asset information, M ≥ m ≥ 2.
[0017] In a case where the readability index value is greater than an index threshold value, the m th target user is determined as an abnormal user.
[0018] In a case where the readability index value is less than or equal to the index threshold value, the m th target user is determined as a non-abnormal user.
[0019] According to an embodiment of the present disclosure, in a case where the m th target user is determined as a suspicious user, a readability index value is calculated according to the initial asset information of the m th target user, comprising:
[0020] The total number of words and the total number of complex words in the initial asset information are calculated, the complex words representing words with text features higher than a preset text feature threshold value.
[0021] According to the total number of words and the total number of lines of the initial asset information, the average number of words per line of text is calculated.
[0022] The ratio of the total number of complex words to the total number of words is calculated.
[0023] According to the ratio and the average number of words per line of text, the readability index value is calculated.
[0024] According to an embodiment of the present disclosure, the N abnormal users and the M target users are clustered according to the standard asset information to obtain a first clustering result, including:
[0025] According to the standard asset information, the N abnormal users and the M target users are clustered by a multidimensional scaling analysis method and / or a hierarchical clustering method to obtain a first clustering result, and the first clustering result includes M first sub-clustering results.
[0026] According to an embodiment of the present disclosure, the N abnormal users and the M target users are clustered according to the standard asset information by a multidimensional scaling analysis method and / or a hierarchical clustering method to obtain a first clustering result, including:
[0027] According to the standard asset information, the distances between the N abnormal users and the M target users are calculated to obtain an initial distance matrix;
[0028] Based on the multidimensional scaling analysis method, the initial distance matrix is iterated to obtain a final distance matrix under a preset number of iterations;
[0029] According to the final distance matrix, a two-dimensional distance structure diagram is generated; and
[0030] According to the two-dimensional distance structure diagram, the M first sub-clustering results are determined.
[0031] According to an embodiment of the present disclosure, the N abnormal users and the M target users are clustered according to the standard asset information by a multidimensional scaling analysis method and / or a hierarchical clustering method to obtain a first clustering result, further including:
[0032] Based on the hierarchical clustering method, the M target users and the N abnormal users are clustered by using the standard asset information until the M target users and the N abnormal users are merged into one class;
[0033] M clustering level information corresponding to the M target users is determined; and
[0034] According to the M clustering level information, M first sub-clustering results corresponding to the M target users are determined.
[0035] According to an embodiment of the present disclosure, the first sub-clustering result is used to represent whether the target user belongs to an abnormal class;
[0036] After obtaining the first clustering result, further including:
[0037] Based on the M first sub-clustering results, L target users belonging to an abnormal class and (M-L) target users not belonging to an abnormal class are screened out from the M target users, and M≥L≥1; and
[0038] According to the standard asset information, the N abnormal class users and the (M-L) target users are clustered again, and the first sub-clustering result of the (M-L) target users is updated.
[0039] According to an embodiment of the present disclosure, the initial asset information includes P transaction characteristics and first parameter values of each transaction characteristic, and the standard asset information includes Q standard factors and second parameter values of each standard factor, P≥Q≥1.
[0040] The initial asset information of the M target users is processed to obtain the standard asset information corresponding to the M target users, including:
[0041] According to the correlation coefficient between the P transaction characteristics, R transaction characteristics without an association relationship are selected from the P transaction characteristics, P≥R≥Q≥1.
[0042] A conversion relationship between the R transaction characteristics and each standard factor is obtained, wherein the standard factor is related to at least one transaction characteristic; and
[0043] According to the conversion relationship and the first parameter values of the R transaction characteristics, the second parameter values of each standard factor in the Q standard factors are determined.
[0044] According to an embodiment of the present disclosure, the standard factor is determined after factor analysis on multiple transaction characteristics, and the standard factor includes a comprehensive factor, a revenue and profit factor, a fee factor, and a value change and income factor.
[0045] A second aspect of the present disclosure provides an abnormal user determination device, including:
[0046] The processing module is configured to process initial asset information of M target users to obtain standard asset information corresponding to the M target users, and the initial asset information includes asset statements of the target users, M≥2.
[0047] The first clustering module is configured to cluster the N abnormal class users and the M target users according to the standard asset information to obtain a first clustering result, N≥M≥2.
[0048] The second clustering module is configured to cluster the N non-abnormal class users and the M target users according to the standard asset information to obtain a second clustering result; and
[0049] The determination module is configured to determine an abnormal result of the M target users according to the first clustering result and the second clustering result, and the abnormal result includes abnormal users, non-abnormal users, and suspicious users.
[0050] The third aspect of the present disclosure provides an electronic device, comprising: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the above-mentioned abnormal user determination method.
[0051] The fourth aspect of the present disclosure further provides a computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the above-mentioned abnormal user determination method.
[0052] The fifth aspect of the present disclosure further provides a computer program product comprising a computer program that, when executed by a processor, implements the above-mentioned abnormal user determination method.
[0053] Embodiments of the present disclosure enable a computer to automatically extract asset information of a target user by processing initial asset information into standard asset information, without the need for manual reading of initial asset information, and converting initial asset information into standard asset information, which can reduce data extraction errors caused by manual operation, and improve the accuracy of determining abnormal users from a data perspective. In addition, by processing initial asset information into standard asset information, the clustered information is unified into standard asset information at the data level, avoiding inaccurate abnormal user determination results due to inconsistent asset indicators and different asset indicator division standards.
[0054] In addition, clustering the target user with abnormal class users and non-abnormal class users respectively ensures the accuracy of clustering from both abnormal class features and non-abnormal class features, which can improve the accuracy of abnormal user determination results. BRIEF DESCRIPTION OF DRAWINGS
[0055] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description of embodiments of the present disclosure taken in conjunction with the accompanying drawings, in which:
[0056] Figure 1 An application scenario of the abnormal user determination method according to an embodiment of the present disclosure is schematically shown;
[0057] Figure 2 A flowchart of the abnormal user determination method according to an embodiment of the present disclosure is schematically shown;
[0058] Figure 3 A flowchart of the method for determining abnormal users according to the first clustering result and the second clustering result according to an embodiment of the present disclosure is schematically shown;
[0059] Figure 4A An application scenario of determining abnormal users according to a specific embodiment of the present disclosure is schematically shown;
[0060] Figure 4BAn application scenario of determining a suspicious user as an abnormal user and a non-abnormal user is schematically shown according to an embodiment of the present disclosure;
[0061] Figure 5 A flowchart of a readability index calculation method is schematically shown according to an embodiment of the present disclosure;
[0062] Figure 6A A two-dimensional distance structure diagram of 29 target users determined based on a multidimensional scaling analysis method is schematically shown according to an embodiment of the present disclosure;
[0063] Figure 6B A scatter plot between actual distances and fitted distances of 29 target users is schematically shown according to an embodiment of the present disclosure;
[0064] Figure 7 A flowchart of a first sub-clustering result updating method is schematically shown according to an embodiment of the present disclosure;
[0065] Figure 8 A structural block diagram of an abnormal user determination apparatus is schematically shown according to an embodiment of the present disclosure; and
[0066] Figure 9 A block diagram of an electronic device adapted to an abnormal user determination method is schematically shown according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0067] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. It is to be understood, however, that the description is merely exemplary and is intended to provide a thorough understanding of the present disclosure. The following description, given together with the accompanying drawings, is intended to provide a thorough understanding of the present disclosure. However, it is apparent that one or more embodiments can be implemented without the specific details, as is apparent to those skilled in the art.
[0068] The terms used herein are merely used to describe specific embodiments, and are not intended to limit the present disclosure. The terms "include", "comprise" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0069] All terms used herein, including technical and scientific terms, have meanings commonly understood by one of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted in a manner consistent with the context of the present specification, and should not be interpreted in an idealized or overly formal manner.
[0070] In the case of using expressions similar to "at least one of A, B, and C, etc.", it is generally to be interpreted in the meaning as the person skilled in the art usually understands the expression (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having A alone, a system having B alone, a system having C alone, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having A, B, and C, etc.).
[0071] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application of data (such as including but not limited to user personal information) comply with relevant legal regulations, necessary security measures are taken, and do not violate public order and good customs.
[0072] Public asset information such as asset statements Asset statements not only include multiple index categories and multiple evaluation criteria, but also include obscure textual description languages.
[0073] On the one hand, the asset indicators in the asset statement are related to the industry to which the enterprise belongs. For example, the asset statement of the financial industry includes quarterly indicators, annual indicators, etc.; the asset statement of the construction industry includes annual indicators, engineering indicators, etc. In related technologies, bank personnel extract evaluation information from asset statements based on work experience by reading asset statements of multiple industries. Then, the extracted evaluation information is input into a computer device, and the computer device determines whether a user has a transaction risk based on the above evaluation information.
[0074] However, there are differences in work experience, understanding ability, reading ability, etc. among business personnel, which causes differences in the input information extracted by business personnel, and a high error rate of the selected evaluation information, further affecting the accuracy of the determination result of abnormal users.
[0075] On the other hand, the asset statement can be used to predict the subsequent development of the user. Therefore, in the case of unsatisfactory performance, the asset statement has a complex form and an obscure textual description language, which makes it difficult for bank personnel to accurately understand and extract evaluation information from the asset statement, resulting in low accuracy of the determination result of abnormal users.
[0076] In addition, in related technologies, the evaluation information in the asset statement is generally directly used to determine whether a user is abnormal; or the evaluation information is simply screened, and whether a user is abnormal is determined based on the screened evaluation information. The evaluation information includes asset indicators in the asset statement. However, different asset indicators have different degrees of influence on different users, and the same indicators also have different degrees of influence on users in different industries. In the case of directly using asset indicators in the asset statement or using simply screened asset indicators to predict whether a user is abnormal, the determination result of abnormal users is not accurate.
[0077] Embodiments of the present disclosure provide an abnormal user determination method, comprising: processing initial asset information of M target users to obtain standard asset information corresponding to the M target users, the initial asset information comprising asset statements of the target users, and M≥2; clustering N abnormal class users and the M target users according to the standard asset information to obtain a first clustering result, N≥M≥2; clustering N non-abnormal class users and the M target users according to the standard asset information to obtain a second clustering result; and determining an abnormal result of the M target users according to the first clustering result and the second clustering result, the abnormal result comprising abnormal users, non-abnormal users and suspicious users.
[0078] Figure 1 An application scenario of the abnormal user determination method according to an embodiment of the present disclosure is schematically shown.
[0079] As shown in Figure 1 the application scenario 100 according to this embodiment can include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104 and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0080] A user can use at least one of the first terminal device 101, the second terminal device 102 and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102 and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0081] For example, the user can use the first terminal device 101, the second terminal device 102 and the third terminal device 103 to log in to the web browser application and obtain asset information of target users through the server 105.
[0082] The first terminal device 101, the second terminal device 102 and the third terminal device 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers and desktop computers, etc.
[0083] The server 105 can be a server providing various services, for example, a background management server (for example only) providing support for a website browsed by a user using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server can analyze and process received user requests and the like, and feed back the processing results (for example, obtained initial asset information) to the terminal device.
[0084] For example, the server 105 can also receive initial asset information of at least one target user sent from the first terminal device 101, the second terminal device 102, and the third terminal device 103. The server 105 processes the initial asset information of the at least one target user to obtain corresponding standard asset information, and determines an abnormal result of the at least one target user based on the standard asset information through clustering.
[0085] Alternatively, any one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 sends a request for obtaining initial asset information of at least one target user to the server 105, and receives the initial asset information fed back from the server 105. Any one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 can process the initial asset information of the at least one target user to obtain corresponding standard asset information, and determine an abnormal result of the at least one target user based on the standard asset information through clustering.
[0086] It should be noted that the abnormal user determination method provided in the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the abnormal user determination apparatus provided in the embodiments of the present disclosure can generally be arranged in the server 105. The abnormal user determination method provided in the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the abnormal user determination apparatus provided in the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0087] The abnormal user determination method provided by the embodiments of the present disclosure can also be executed by any one of the first terminal device 101, the second terminal device 102, and the third terminal device 103. Correspondingly, the abnormal user determination apparatus provided by the embodiments of the present disclosure can be generally arranged in the first terminal device 101, the second terminal device 102, and the third terminal device 103. The abnormal user determination method provided by the embodiments of the present disclosure can also be executed by other terminal devices different from the first terminal device 101, the second terminal device 102, and the third terminal device 103 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the abnormal user determination apparatus provided by the embodiments of the present disclosure can also be arranged in other terminal devices different from the first terminal device 101, the second terminal device 102, and the third terminal device 103 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0088] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above scenario is only illustrative. Any number of terminal devices, networks, and servers can be provided according to implementation needs.
[0089] The abnormal user determination method of the embodiments of the present disclosure will be described in detail below based on the scenario described above. Figure 1 Figures 2-7 The abnormal user determination method of the embodiments of the present disclosure will be described in detail below based on the scenario described above.
[0090] Figure 2 A flowchart of the abnormal user determination method according to the embodiments of the present disclosure is schematically shown.
[0091] As shown in Figure 2 , the method 200 includes operations S210-S240.
[0092] In operation S210, initial asset information of M target users is processed to obtain standard asset information corresponding to the M target users, the initial asset information including asset statements of the target users, and M≥2.
[0093] According to embodiments of the present disclosure, the target users include enterprise users in multiple industries, for example, large enterprise users and small and medium-sized enterprise users in the construction industry, and large enterprise users and small and medium-sized enterprise users in the financial industry. The target users also include individual business users and other micro enterprises.
[0094] According to embodiments of the present disclosure, the initial asset information includes annual statements, quarterly statements, public account tweets containing asset information, newspapers and magazines, etc. The initial asset information is obtained with the permission of the user or through a public channel permitted by the user, such as the official website of the user, financial newspapers, etc.
[0095] According to the disclosed embodiments, the initial asset information is raw text information without processing, and the standard asset information is text information or digital information processed by a computer based on extraction rules. For example, the initial asset information is an annual asset report of an enterprise, which includes a description text of the development of the enterprise in a certain year and asset indicators. The processed standard asset data includes a plurality of asset indicators, converted standard indicators, or evaluation indicators corresponding to the text.
[0096] In operation S220, the N abnormal users and the M target users are clustered according to the standard asset information, to obtain a first clustering result, N≥M≥2.
[0097] According to the embodiments of the present disclosure, the abnormal users and the non-abnormal users can be determined according to whether a default behavior occurs. For example, the abnormal users include enterprise users who have been determined to have a default behavior, and the non-abnormal users include enterprise users who do not have a default behavior.
[0098] According to the embodiments of the present disclosure, the abnormal users and the non-abnormal users can also be determined according to the default type, the default times, the default time, the default amount, etc. For example, the abnormal users include enterprise users whose default times exceed 3 times, and the non-abnormal users include enterprise users who do not have a default behavior or whose default times do not exceed 3 times.
[0099] According to the embodiments of the present disclosure, the abnormal users and the non-abnormal users can also be determined according to specific rules within the bank.
[0100] According to the embodiments of the present disclosure, after the initial asset information of the M target users is processed into the standard asset information, the N abnormal users and the M target users can be clustered based on the standard asset information of the N abnormal users and the standard asset information of the M target users, to determine a first clustering result. The first clustering result includes M first sub-clustering results corresponding to the M target users. The first sub-clustering result represents whether the target user belongs to the abnormal user.
[0101] According to the embodiments of the present disclosure, when the clustering operation is performed, the number of abnormal users is greater than the number of target users to be determined.
[0102] In operation S230, the N non-abnormal users and the M target users are clustered according to the standard asset information, to obtain a second clustering result.
[0103] According to an embodiment of the present disclosure, after the initial asset information of the M target users is processed into the standard asset information, the N non-anomalous-class users and the M target users can be clustered based on the standard asset information of the N non-anomalous-class users and the standard asset information of the M target users, to determine a second clustering result. The second clustering result includes M second sub-clustering results corresponding to the M target users. The second sub-clustering result indicates whether the target user belongs to the non-anomalous-class users.
[0104] According to an embodiment of the present disclosure, the N non-anomalous-class users and the M target users can be clustered first, and then the N anomalous-class users and the M target users can be clustered. Conversely, the N anomalous-class users and the M target users can be clustered first, and then the N non-anomalous-class users and the M target users can be clustered. Alternatively, the two clustering operations can be performed simultaneously.
[0105] According to an embodiment of the present disclosure, sample imbalance can affect the learning effect of the model, and therefore, when clustering, the N anomalous-class users and the N non-anomalous-class users are selected to ensure that the anomalous-class clustering result and the non-anomalous-class clustering result are not affected by sample imbalance.
[0106] In operation S240, the anomaly result of the M target users is determined according to the first clustering result and the second clustering result. The anomaly result includes anomalous users, non-anomalous users, and suspicious users.
[0107] According to an embodiment of the present disclosure, after the target users are clustered using the N anomalous-class users and the N non-anomalous-class users, the anomaly result of each target user can be determined according to the first clustering result and the second clustering result.
[0108] According to an embodiment of the present disclosure, for a target user belonging to an anomalous user, the standard asset information of the user is similar to the standard asset information of the anomalous-class users, and is not similar to the standard asset information of the non-anomalous-class users. For a target user belonging to a non-anomalous user, the standard asset information of the user is similar to the standard asset information of the non-anomalous-class users, and is not similar to the standard asset information of the anomalous-class users.
[0109] A suspicious user indicates a target user whose standard asset information is similar to the standard asset information of the anomalous-class users and the non-anomalous-class users, and a target user whose standard asset information is not similar to the standard asset information of the anomalous-class users and the non-anomalous-class users.
[0110] Embodiments of the present disclosure can automatically extract asset information of target users by processing initial asset information into standard asset information, without manually reading the initial asset information, and convert the initial asset information into standard asset information, thereby reducing data extraction errors caused by manual operation, improving the accuracy of determining abnormal users from a data perspective. In addition, by processing the initial asset information into standard asset information, the clustered information is unified into standard asset information at the data level, avoiding inaccurate abnormal user determination results caused by inconsistent asset indicators and different asset indicator division standards.
[0111] In addition, clustering the target users with the abnormal class users and the non-abnormal class users respectively ensures the accuracy of clustering from both the abnormal class features and the non-abnormal class features, thereby improving the accuracy of the abnormal user determination result.
[0112] According to embodiments of the present disclosure, the first clustering result includes M first sub-clustering results, and the second clustering result includes M second sub-clustering results.
[0113] According to the first clustering result and the second clustering result, determining whether the M target users are abnormal includes: for the nth target user, in a case where the nth first sub-clustering result indicates that the nth target user belongs to the abnormal class and the nth second sub-clustering result indicates that the nth target user does not belong to the non-abnormal class, determining that the nth target user is an abnormal user, M≥n≥2.
[0114] In a case where the nth first sub-clustering result indicates that the nth target user does not belong to the abnormal class and the nth second sub-clustering result indicates that the nth target user belongs to the non-abnormal class, determining that the nth target user is a non-abnormal user.
[0115] In a case where the nth first sub-clustering result indicates that the nth target user belongs to the abnormal class and the nth second sub-clustering result indicates that the nth target user belongs to the non-abnormal class, or the nth first sub-clustering result indicates that the nth target user does not belong to the abnormal class and the nth second sub-clustering result indicates that the nth target user does not belong to the non-abnormal class, determining that the nth target user is a suspicious user.
[0116] According to embodiments of the present disclosure, after clustering the N abnormal class users and the M target users, a first sub-clustering result is obtained for each target user. After clustering the N non-abnormal class users and the M target users, a second sub-clustering result is obtained for each target user.
[0117] The first sub-cluster result indicates whether the target user belongs to the anomalous class, including those belonging to the same class as anomalous users and those not belonging to the same class. The second sub-cluster result indicates whether the target user belongs to the non-anomalous class, including those belonging to the same class as non-anomalous users and those not belonging to the same class.
[0118] According to embodiments of this disclosure, the same clustering method or different clustering methods can be used in obtaining the first clustering result and the second clustering result.
[0119] Figure 3 The flowchart illustrating a method for determining abnormal users based on a first clustering result and a second clustering result according to an embodiment of the present disclosure is shown.
[0120] like Figure 3 As shown, the method 300 for determining abnormal users based on the first clustering result and the second clustering result in this embodiment includes operations S301 to S305, which can be used as a specific embodiment of operation S240.
[0121] In operation S301, the user is checked to determine if they belong to the same category as the abnormal user. Specifically, clustering can be performed on N abnormal users and M target users to obtain a first sub-clustering result for at least one target user. Operation S301 then evaluates the first sub-clustering result. Regardless of whether the first sub-clustering result is "belongs to the same category as the abnormal user" or "does not belong to the same category as the abnormal user," the process proceeds to operation S302.
[0122] In operation S302, the user is checked to determine if they belong to the same category as non-abnormal users. Specifically, clustering can be performed on N non-abnormal users and M target users to obtain a second sub-clustering result for at least one target user. Operation S302 then evaluates the second sub-clustering result. Since the first sub-clustering result was evaluated before executing operation S302, the first and second sub-clustering results are combined to determine whether the target user is abnormal.
[0123] For example, if it is determined that the target user and the abnormal user belong to the same category, and the target user and the non-abnormal user do not belong to the same category, then proceed to operation S303.
[0124] If it is determined that the target user and the abnormal user belong to the same category, and the target user and the non-abnormal user belong to the same category, proceed to operation S304.
[0125] If it is determined that the target user and the abnormal user do not belong to the same category, and the target user and the non-abnormal user do not belong to the same category, proceed to operation S304.
[0126] If it is determined that the target user and the abnormal user do not belong to the same category, and the target user and the non-abnormal user belong to the same category, proceed to operation S305.
[0127] In operation S303, the target user is identified as an abnormal user.
[0128] In operation S304, the target user is identified as a suspicious user.
[0129] In operation S303, the target user is identified as a non-abnormal user.
[0130] According to embodiments of this disclosure, since a target user cannot be both an abnormal user and an abnormal user at the same time, if it cannot be determined whether the target user is an abnormal user based on the first sub-clustering result and the second sub-clustering result, the target user is identified as a suspicious user so that subsequent operations can be performed on the suspicious user.
[0131] Figure 4A This illustration depicts an application scenario for identifying abnormal users according to a specific embodiment of the present disclosure.
[0132] like Figure 4A As shown in the example, application scenario 400A demonstrates the process of determining whether multiple target users are abnormal users.
[0133] According to an embodiment of this disclosure, after obtaining the initial asset information 401 of the target user, the initial asset information 401 of the target user is processed to obtain the standard asset information 403 of the target user.
[0134] Input the standard asset information 402 of abnormal users and the standard asset information 403 of the target user into the clustering model. After clustering, the first clustering result 405 can be determined. Input the standard asset information 404 of non-abnormal users and the standard asset information 403 of the target user into the clustering model. After clustering, the second clustering result 406 can be determined.
[0135] Based on the first clustering result 405 and the second clustering result 406, it can be determined that the target user belongs to the abnormal user category 407, or the target user belongs to the non-abnormal user category 408, or the target user belongs to the suspicious user category 409.
[0136] According to an embodiment of the present disclosure, when only the target user is clustered with the abnormal user, the clustering process only focuses on the abnormal feature, and the user with no obvious abnormal feature cannot be accurately determined. Similarly, when only the target user is clustered with the non-abnormal user, the user cannot be accurately determined. The model that focuses on both the abnormal feature and the non-abnormal feature can be used to determine whether the target user belongs to the abnormal user, but in the training process, the model needs to spend a high cost to generate a training set and adjust parameters, resulting in high model training cost and high abnormal user determination cost.
[0137] Embodiments of the present disclosure cluster the target user with the abnormal user and the non-abnormal user respectively, without adjusting and improving the clustering model, to ensure simultaneous clustering of the abnormal feature and the non-abnormal feature, and improve the accuracy of determining the abnormal user.
[0138] According to an embodiment of the present disclosure, after determining the abnormal result of the M target users according to the first clustering result and the second clustering result, the method further includes:
[0139] In a case where the mth target user is determined as a suspicious user, a readability index value is calculated according to the initial asset information of the mth target user, the readability index value is used to represent the readability of the initial asset information, and M>m≥2;
[0140] In a case where the readability index value is greater than an index threshold value, the mth target user is determined as an abnormal user.
[0141] In a case where the readability index value is less than or equal to the index threshold value, the mth target user is determined as a non-abnormal user.
[0142] According to an embodiment of the present disclosure, after clustering the target user with the abnormal user and the non-abnormal user respectively, the target user is similar to the abnormal user and the non-abnormal user, or the target user is not similar to the abnormal user and the non-abnormal user. The readability of the initial asset information can reflect whether the user violates the rules.
[0143] Therefore, after the target user is determined as a suspicious user, the suspicious user is divided again according to the readability of the initial asset information of the suspicious user to determine whether the suspicious user is an abnormal user.
[0144] According to an embodiment of the present disclosure, the readability index value represents the readability of the initial asset information, and the readability index value can be determined by the text feature of the initial asset information.
[0145] According to embodiments of this disclosure, in linguistics, the Fog Index, or Gunning Fog Index, refers to a readability test for English documents, used to estimate the number of years of schooling required for a person to understand an English document. For example, a Fog Index of 12 requires a US high school graduation, which represents 12 years of schooling (6 years of elementary school, 3 years of middle school, and 3 years of high school).
[0146] According to embodiments of this disclosure, a preset threshold is used to determine whether the initial asset information, used to characterize user asset features, is easy to read. The threshold can be obtained by pre-analyzing the initial asset information of multiple anomalous users and / or multiple non-anomalous users. Multiple anomalous users can be the same as or different from the N anomalous users used for clustering. Similarly, multiple non-anomalous users can be the same as or different from the N non-anomalous users used for clustering.
[0147] For example, the initial asset information of suspicious user A includes a large number of data tables, and the text description is clear and simple. The initial asset information of suspicious user B does not include data tables, and the text description paragraphs are too long and complex. Therefore, the readability index value of suspicious user B is higher than that of suspicious user A.
[0148] According to embodiments of this disclosure, suspicious users can be further classified into abnormal users or non-abnormal users based on indicator thresholds.
[0149] In the embodiments of this disclosure, for suspicious users, by calculating the readability index value of the initial asset information, suspicious users are further divided into abnormal users and non-abnormal users from the perspective of whether the initial asset information is easy to understand, so as to improve the accuracy of identifying abnormal users.
[0150] Figure 4B This illustration depicts an application scenario where a suspicious user is identified as an abnormal user or a non-abnormal user according to a specific embodiment of this disclosure.
[0151] like Figure 4B As shown in the example, application scenario 400B demonstrates the process of determining whether a user is an abnormal user.
[0152] According to an embodiment of this disclosure, when a target user is identified as a suspicious user 409, initial asset information 410 of the suspicious user is obtained. Based on the initial asset information 410 of the suspicious user, a readability index value 411 of the initial asset information is calculated.
[0153] After obtaining the indicator threshold 412 from the database or the training library, the readability indicator value 411 and the indicator threshold 412 are compared, and the suspicious user 409 is further determined as the abnormal user 407 or the suspicious user 409 is further determined as the non-abnormal user 408.
[0154] According to an embodiment of the present disclosure, one or more suspicious users can be determined simultaneously according to the first clustering result and the second clustering result. Thus, in the process of further dividing the suspicious users, the readability indicator value of one or more suspicious users can also be calculated simultaneously, and whether one or more suspicious users are abnormal users is determined based on the readability indicator value of one or more suspicious users.
[0155] According to an embodiment of the present disclosure, in the case where the mth target user is determined as a suspicious user, the readability indicator value is calculated according to the initial asset information of the mth target user, including the following steps.
[0156] The total number of words and the total number of complex words in the initial asset information are calculated, and the complex word represents a word with a text feature higher than a preset text feature threshold.
[0157] According to the total number of words and the total number of lines of the initial asset information, the average number of words per line of text is calculated.
[0158] The ratio of the total number of complex words to the total number of words is calculated.
[0159] According to the ratio and the average number of words per line of text, the readability indicator value is calculated.
[0160] According to an embodiment of the present disclosure, the complex word represents a word with a text feature higher than a preset text feature value. For the initial asset information in English form, the text feature includes the number of syllables. The preset text feature value is 3 syllables. For example, we is a simple word because we has only one syllable. Difficulty is a complex word because Difficulty has 4 syllables.
[0161] According to an embodiment of the present disclosure, for the initial asset information in Chinese form, the text feature includes the number of strokes. The preset text feature value is the average number of strokes of Chinese. For example, the total number of strokes of the complex word is divided by the number of characters in the complex word to obtain the average number of strokes per character in the complex word. In the case where the average number of strokes per character is greater than the average number of strokes of Chinese, the word is a complex word.
[0162] According to an embodiment of the present disclosure, the formula for calculating the readability indicator value satisfies:
[0163]
[0164] wherein μ represents a parameter, and Fog represents the readability indicator value.
[0165] According to an embodiment of the present disclosure, in a case where the initial asset information of the suspicious user includes a Chinese form and an English form, a readability index value of the Chinese form and a readability index value of the English form are respectively calculated. According to an average value of the readability index values of the Chinese form and the English form, the suspicious user is determined as an abnormal user or a non-abnormal user.
[0166] Figure 5 A flowchart of a readability index calculation method according to an embodiment of the present disclosure is schematically shown.
[0167] As shown in Figure 5 The readability index calculation method 500 of this embodiment shows a process of determining a readability index value according to initial asset information.
[0168] According to an embodiment of the present disclosure, after obtaining the initial asset information 501 of the suspicious user, the total number of complex words 502, the total number of words 503 and the total number of lines 504 in the initial asset information 501 are determined, the ratio 505 between the total number of complex words 502 and the total number of words 503 is determined, and the average number of words per line of text 506 is determined according to the ratio between the total number of words 503 and the total number of lines. According to the sum of the ratio 505 and the average number of words per line of text 506, the readability index of the initial asset information is determined.
[0169] According to an embodiment of the present disclosure, according to the standard asset information, the N abnormal class users and the M target users are clustered to obtain a first clustering result, including the following steps.
[0170] According to the standard asset information, the N abnormal class users and the M target users are clustered by a multidimensional scaling analysis method and / or a hierarchical clustering method to obtain a first clustering result, and the first clustering result includes M first sub-clustering results.
[0171] According to an embodiment of the present disclosure, according to the standard asset information, the N non-abnormal class users and the M target users are clustered by a multidimensional scaling analysis method and / or a hierarchical clustering method to obtain a second clustering result, and the first clustering result includes M second sub-clustering results.
[0172] According to an embodiment of the present disclosure, the target users can be clustered with the abnormal class users and the non-abnormal class users respectively by the multidimensional scaling analysis method to obtain the first clustering result and the second clustering result. The target users can also be clustered with the abnormal class users and the non-abnormal class users respectively by the hierarchical clustering method to obtain the first clustering result and the second clustering result. The target users can be clustered with the abnormal class users by the multidimensional scaling analysis method to obtain the first clustering result, and the target users can be clustered with the non-abnormal class users by the hierarchical clustering method to obtain the second clustering result. Alternatively, the target users can be clustered with the abnormal class users by the hierarchical clustering method to obtain the first clustering result, and the target users can be clustered with the non-abnormal class users by the multidimensional scaling analysis method to obtain the second clustering result.
[0173] According to an embodiment of the present disclosure, the target users can be clustered with the abnormal class users by the multidimensional scaling analysis method to obtain a multidimensional scaling first clustering result, and the target users can be clustered with the abnormal class users again by the hierarchical clustering method to obtain a hierarchical clustering first clustering result. Then, the multidimensional scaling first clustering result can be verified by the hierarchical clustering first clustering result to determine the first clustering result. Similarly, the second clustering result can also be determined by the multidimensional scaling analysis method and the hierarchical clustering method.
[0174] According to an embodiment of the present disclosure, the process of clustering the target users and the abnormal class users by the hierarchical analysis method or the hierarchical clustering method will be described below by taking the process of obtaining the first clustering result as an example.
[0175] According to an embodiment of the present disclosure, according to the standard asset information, the N abnormal class users and the M target users can be clustered by the multidimensional scaling analysis method and / or the hierarchical clustering method to obtain the first clustering result, including:
[0176] According to the standard asset information, the distances between the N abnormal class users and the M target users are calculated to obtain an initial distance matrix.
[0177] Based on the multidimensional scaling analysis method, the initial distance matrix is iterated to obtain a final distance matrix under the condition that a preset number of iterations is performed.
[0178] According to the final distance matrix, a two-dimensional distance structure diagram is generated; and
[0179] According to the two-dimensional distance structure diagram, M first sub-clustering results are determined.
[0180] According to an embodiment of the present disclosure, according to the standard asset information converted based on the initial asset information, the distances between the N abnormal class users and the M target users can be calculated to form an initial distance matrix.
[0181] According to an embodiment of the present disclosure, the distance between the target user and the target user, the distance between the target user and the abnormal class user, and the distance between the abnormal class user and the abnormal class user can be calculated according to the Euclidean distance.
[0182] According to an embodiment of the present disclosure, the standard asset information includes a plurality of standard factors, and when calculating the distance between the nth abnormal class user and the mth target user, the distance between the nth abnormal class user and the mth target user is calculated for each standard factor, and the distance between the nth abnormal class user and the mth target user is obtained by synthesizing a plurality of standard factors.
[0183] According to an embodiment of the present disclosure, the multidimensional scaling method is a method for rearranging samples in an effective way, and the final purpose is to obtain a structure closest to the observed distance. In essence, the multidimensional scaling method can move the standard factor to the required dimension space, and at the same time, it can check whether the new structure is reasonable.
[0184] According to an embodiment of the present disclosure, the initial distance matrix is iterated by the steepest descent method, and the preset number of iterations is 30 times. After 30 iterations, the convergence criterion is 0.00100, the minimum S-stress is 0.00500, and the Tiestore is 406.
[0185] According to an embodiment of the present disclosure, the stress index of the final distance matrix obtained in the 30th iteration is 0.10109, and the relative standard deviation RSQ is 0.98154, indicating that the fitting degree is good.
[0186] According to the final distance matrix, a two-dimensional distance structure diagram is generated, and according to the two-dimensional distance structure diagram, M first sub-clustering results are determined.
[0187] Specifically, in the two-dimensional clustering structure diagram, the center point of the abnormal class user can be determined, and the target user with a distance from the center point less than or equal to a preset distance is determined as “belonging to the same class as the abnormal class user”, and the target user with a distance from the center point greater than the preset distance is determined as “not belonging to the same class as the abnormal class user”.
[0188] According to an embodiment of the present disclosure, the two-dimensional distance structure diagram only shows the relative distance between the target user and the abnormal class user. The abscissa and the ordinate only represent the relative distance between the target user and the abnormal class user, and do not represent the actual parameters. The relationship between the target user and the abnormal class user can be fully and intuitively reproduced, and on this basis, the classification of samples can be realized according to the distance between the object points.
[0189] Figure 6AA two-dimensional distance structure diagram of 29 target users determined based on a multidimensional scaling analysis method is shown schematically.
[0190] As shown in Figure 6A , the two-dimensional distance structure diagram 600A shows the relative distances between the 29 target users, and the target users located near the origin are more similar.
[0191] According to an embodiment of the present disclosure, the accuracy between the first clustering results can also be determined according to a scatter plot between the actual distances and the fitted distances.
[0192] Figure 6B A scatter plot between the actual distances and the fitted distances of the 29 target users according to an embodiment of the present disclosure is shown schematically.
[0193] As shown in Figure 6B , the scatter plot Figure 6B The abscissa "gap" represents the fitted distance, and the ordinate "distance" represents the actual distance. The scatter plot is basically near the 45° straight line, indicating that the effect of the multidimensional scaling analysis method is relatively ideal.
[0194] According to an embodiment of the present disclosure, the clustering based on the multidimensional scaling analysis method can map high-dimensional standard asset data (Q standard factors) including multiple factors to a low-dimensional image, while reducing the dimension of the data, the original relationship between the users is preserved to the greatest extent, and the clustering efficiency and accuracy can be improved.
[0195] According to an embodiment of the present disclosure, based on the standard asset information, the N abnormal class users and the M target users are clustered by the multidimensional scaling analysis method and / or the hierarchical clustering method to obtain the first clustering result, and the method further comprises:
[0196] Based on the hierarchical clustering method, the M target users and the N abnormal class users are clustered by using the standard asset information until the M target users and the N abnormal class users are merged into one class;
[0197] Determine M clustering level information corresponding to the M target users; and
[0198] According to the M clustering level information, determine M first sub-clustering results corresponding to the M target users.
[0199] According to an embodiment of the present disclosure, when starting clustering, the M target users and the N abnormal class users are taken as a separate class to obtain M+N classes, and in each round of clustering process, the two classes with the closest distance are merged; then, the distance between the new class and other classes is recalculated, and the two classes with the closest distance are merged again... until the M target users and the N abnormal class users are all merged into one class.
[0200] In the hierarchical clustering process, the distance between two classes is calculated according to the square Euclidean distance.
[0201] According to an embodiment of the present disclosure, the clustering level information includes the number of levels at which the target user starts to merge with other classes, the other classes including one abnormal class user as a class, and the classes after the merging has occurred.
[0202] According to an embodiment of the present disclosure, determining the M first sub-clustering results corresponding to the M target users according to the M clustering level information includes: taking the number of levels at which the N abnormal class users are merged into one class as a reference level; and selecting M1 target users with a level number less than the reference level and M2 target users with a level number greater than the reference level from the M target users according to the M clustering level information, where M = M1 + M2.
[0203] The first sub-clustering result of the M1 target users is determined as "belonging to the same class as the abnormal class user", and the first sub-clustering result of the M2 target users is determined as "not belonging to the same class as the abnormal class user".
[0204] According to an embodiment of the present disclosure, the N abnormal class users are taken as reference users, and when the N abnormal class users are merged into the same class, it indicates that the target users with similar characteristics have been merged into a class with the abnormal class users before this level. Therefore, by taking the level at which the abnormal class users are merged as a reference, the accuracy of determining whether the target user is an abnormal user can be improved.
[0205] According to an embodiment of the present disclosure, the target users and the abnormal class users are clustered by the hierarchical clustering method, which can improve the clustering efficiency and reduce the clustering cost.
[0206] According to an embodiment of the present disclosure, after obtaining the first clustering result, the method further includes:
[0207] Based on the M first sub-clustering results, L target users belonging to the abnormal class and (M-L) target users not belonging to the abnormal class are screened from the M target users, and M > L > 1; and
[0208] According to the standard asset information, the N abnormal class users and the (M-L) target users are clustered again, and the first sub-clustering result of the (M-L) target users is updated.
[0209] After obtaining the second clustering result, the method further includes:
[0210] Based on the M second sub-clustering results, L' target users belonging to the non-abnormal class and (M-L') target users not belonging to the non-abnormal class are screened from the M target users, and M > L' > 1; and
[0211] According to the standard asset information, the M target users and the N abnormal class users are clustered again, and a second sub-clustering result of the (M-L') target users is updated.
[0212] According to an embodiment of the present disclosure, the single clustering may miss part of the target users, and thus the M target users and the N abnormal class users are clustered multiple times, and the M target users and the N non-abnormal class users are clustered multiple times.
[0213] The following will be described by taking Figure 7 The first sub-clustering result updating method of the disclosed embodiment will be described in detail. The second sub-clustering result updating method is similar to the first sub-clustering result updating method, and thus will not be described herein.
[0214] Figure 7 A flowchart of the first clustering result updating method according to an embodiment of the present disclosure is schematically shown.
[0215] As shown in Figure 7 The M first sub-clustering results before updating include: the 1st first sub-clustering result 701_1, …, the Lth first sub-clustering result 701_L, the (L+1)th first sub-clustering result 701_L+1, …, the Mth first sub-clustering result 701_M.
[0216] The updated first clustering result includes: the M first sub-clustering results, such as the 1st first sub-clustering result 704_1, …, the Lth first sub-clustering result 704_L, the (L+1)th first sub-clustering result 704_L+1, …, the Mth first sub-clustering result 704_M.
[0217] In the case that the 1st first sub-clustering result 701_1, …, the Lth first sub-clustering result 701_L represent that the 1st target user, …, the Lth target user belong to the abnormal class, the 1st first sub-clustering result 701_1, …, the Lth first sub-clustering result 701_L are taken as the updated 1st first sub-clustering result 704_1, …, the Lth first sub-clustering result 704_L.
[0218] In the case that the (L+1)th first sub-clustering result 701_L+1, …, the Mth first sub-clustering result 701_M represent that the (L+1)th target user, …, the Mth target user do not belong to the abnormal class, according to the standard asset information 702_L+1 of the (L+1)th target user, …, the standard asset information 702_M of the Mth target user, and the standard asset information 703 of the N abnormal class users, the updated (L+1)th first sub-clustering result 704_L+1, …, the Mth first sub-clustering result 704_M are obtained through T times of clustering. The clustering times T can be determined according to actual conditions.
[0219] Embodiments of the present disclosure can improve clustering accuracy by multiple clustering, so as to reduce the number of determined suspicious users and reduce subsequent workload.
[0220] According to embodiments of the present disclosure, the initial asset information includes P transaction features and first parameter values for each transaction feature, the standard asset information includes Q standard factors and second parameter values for each standard factor, and P > Q > 1.
[0221] According to embodiments of the present disclosure, the initial asset information includes P transaction features, and each transaction feature corresponds to an asset index and is used to represent the asset transaction status of the target user. There may be differences between transaction features in different industries, and training different models using transaction features in different industries will increase training costs and application costs. In addition, there may be correlations between the P transaction features. When transaction features with correlations are input into a model, the learning effect and prediction effect of the model will be affected
[0222] According to embodiments of the present disclosure, the initial asset information of the M target users is processed to obtain standard asset information corresponding to the M target users, including:
[0223] According to the correlation coefficients between the P transaction features, R transaction features without correlation are selected from the P transaction features, and P > R > Q > 1.
[0224] A conversion relationship between the R transaction features and each standard factor is obtained, wherein the standard factor is related to at least one transaction feature; and
[0225] According to the conversion relationship and the first parameter values of the R transaction features, the second parameter values of each standard factor in the Q standard factors are determined.
[0226] According to embodiments of the present disclosure, the correlation coefficients between the P transaction features can be calculated according to the Pearson correlation coefficients, and R transaction features without correlation are selected from the P transaction features.
[0227] According to embodiments of the present disclosure, each standard factor is related to at least one transaction feature, and the conversion includes conversion of at least one transaction feature into a conversion coefficient of the standard factor. According to at least one conversion coefficient, at least one transaction feature is linearly superimposed to obtain the standard factor.
[0228] For example, the transaction characteristics R1, R2, R3 and R4 are related to the standard factor A. After obtaining the conversion coefficients λ1, λ2, λ3 and λ4 between the transaction characteristics R1, R2, R3 and R4 and the standard factor A, the first parameter values of the transaction characteristics and the conversion coefficients are linearly superimposed to obtain the second parameter values of the standard factor, such as: A = λ1*R1 + λ2*R2 + λ3*R3 + λ4*R4.
[0229] According to an embodiment of the present disclosure, the standard factor is determined after factor analysis on a plurality of transaction characteristics, and the standard factor includes a comprehensive factor, a profit factor, a cost factor and a value change profit factor.
[0230] As a specific embodiment, the initial asset information includes 24 transaction characteristics x1-x24, such as total operating revenue x1, operating revenue x2, total operating cost x3, operating cost x4, operating tax and additional x5, sales expense x6, management expense x7, financial expense x8, R&D expense x9, fair value change profit x10, investment profit x11, investment profit x12 to joint venture and joint venture, operating profit x13, non-operating income x14, non-operating expense x15, total profit x16, income tax expense x17, net profit x18, net profit attributable to the parent company x19, minority shareholder interest x20, basic earnings per share x21, other comprehensive income x22, total comprehensive income x23, and total comprehensive income attributable to the parent company x24.
[0231] By calculating the correlation coefficients between the 24 transaction characteristics, a correlation matrix is generated. According to the correlation matrix, it is found that the correlation coefficients between the total operating revenue x1, the operating revenue x2, the total operating cost x3 and the operating cost x4 are close to 1, and since there is a correlation relationship between the total operating revenue x1, the operating revenue x2, the total operating cost x3 and the operating cost x4, the correlation coefficient between the total profit x16 and the operating profit x13 is close to 1, and thus 20 transaction characteristics without correlation relationship are selected from the 24 transaction characteristics, such as operating cost x4, operating tax and additional x5, sales expense x6, management expense x7, financial expense x8, R&D expense x9, fair value change profit x10, investment profit x11, investment profit x12 to joint venture and joint venture, operating profit x13, non-operating income x14, non-operating expense x15, income tax expense x17, net profit x18, net profit attributable to the parent company x19, minority shareholder interest x20, basic earnings per share x21, other comprehensive income x22, total comprehensive income x23, and total comprehensive income attributable to the parent company x24.
[0232] After determining the four standard factors according to the factor analysis method, the above-mentioned 20 transaction characteristics are converted into four standard factors, such as a comprehensive factor, a profit factor, a cost factor, and a value change profit factor.
[0233] Embodiments of the present disclosure can greatly reduce the data volume, clustering difficulty and clustering time, and improve the clustering efficiency by converting the transaction characteristics into standard factors.
[0234] According to embodiments of the present disclosure, the process of determining the standard factors by using the factor analysis method includes the following steps.
[0235] According to the correlation coefficients of the R transaction characteristics, a correlation matrix including the R transaction characteristics is generated. The R transaction characteristics are calculated by using the principal component analysis method, and Q principal components are extracted. As a specific embodiment of the present disclosure, still taking the above-mentioned 20 transaction characteristics as an example, four principal components are extracted by using the principal component analysis method, the characteristic values are all greater than 1, and the cumulative contribution value of the first four principal components reaches 91.831%,
[0236] Then, the common factor variance ratio is calculated according to the first parameter values of the 20 transaction characteristics. Among the 20 common factor variance ratios, only the common factor variance ratio of the other comprehensive income x22 is greater than 0.5 and less than 0.7, the common factor variance ratios of 15 transaction characteristics are greater than 0.9, and the common factor variance ratios of four transaction characteristics are greater than 0.7 and less than 0.9. It is shown that the above-mentioned four principal components can better reflect the information of the transaction characteristics.
[0237] In addition, in the factor analysis process, the factor loading matrix before rotation and the factor loading matrix after rotation are calculated respectively. Among them, the Kaiser normalization maximum variance method is used for orthogonal rotation.
[0238] In the factor loading matrix before rotation, for the first principal component, the loadings of 17 transaction characteristics are greater than 0.5, such as operating costs x4, business tax and additional x5, sales expenses x6, management expenses x7, research and development expenses x9, investment income x11, investment income x12 to joint ventures and joint ventures, operating profit x13, non-operating income x14, non-operating expenses x15, income tax expense x17, net profit x18, net profit attributable to the parent company x19, minority shareholder interest x20, basic earnings per share x21, total comprehensive income x23, and comprehensive income attributable to the parent company x24. It is shown that most of the transaction characteristics have more loadings on the first principal component.
[0239] For the second principal component, the transaction characteristics basic earnings per share x21, financial expenses x8, and research and development expenses x9 have greater loadings on the second principal component.
[0240] For the third principal component, the management expense x7 has a large load on the third principal component.
[0241] For the fourth principal component, the fair value change income x10 has a large load on the fourth principal component.
[0242] Therefore, it is illustrated that the first principal component reflects the comprehensive situation, the second principal component reflects the income and expense, the third principal component reflects the management expense x7, and the fourth principal component reflects the fair value change income x10.
[0243] In the rotated factor loading matrix, for the first principal component, the controllable transaction characteristics include the minority shareholder profit and loss x20, the business tax and additional expenses x5, the financial expense x8, the investment income of the joint venture and the joint venture x12, the income tax expense x17, the non-operating expenses x15, the investment income x11, the non-operating income x14, the three operating profits x13, the operating cost x4, the eight total comprehensive income x23, and the five net profit x18.
[0244] For the second principal component, the controllable transaction characteristics include the income tax expense x17, the three operating profits x13, the basic earnings per share x21, the total comprehensive income attributable to the parent company x24, the net profit attributable to the parent company x19, the seven other comprehensive income x22, the eight total comprehensive income x23, and the five net profit x18.
[0245] For the third principal component, the controllable transaction characteristics include the non-operating income x14, the operating cost x4, the sales expense x6, the management expense x7, and the research and development expense x9.
[0246] For the fourth principal component, the controllable transaction characteristics include the fair value change income x10.
[0247] Therefore, the first principal component is determined as the first standard factor, reflecting the situation of each aspect of the profit-related financial indicators, and is called a comprehensive factor; the second principal component is determined as the second standard factor, reflecting the income and profit of the listed company, and is called an income and profit factor. The third principal component is determined as the third standard factor, reflecting the cost and expense, and is called an expense factor. The fourth principal component is determined as the fourth standard factor, reflecting the fair value change income, and is called a value change income factor.
[0248] Figure 8 A structural block diagram of an abnormal user determination apparatus according to an embodiment of the present disclosure is schematically shown.
[0249] As Figure 8 shown, the abnormal user determination apparatus 800 of this embodiment includes a processing module 810, a first clustering module 820, a second clustering module 830, and a determination module 840.
[0250] The processing module 810 is configured to process initial asset information of M target users to obtain standard asset information corresponding to the M target users, where the initial asset information includes asset statements of the target users, and M is greater than or equal to 2. In an embodiment, the processing module 810 can be configured to perform operation S210 described above, and details are not described herein again.
[0251] The first clustering module 820 is configured to cluster N abnormal-class users and the M target users according to the standard asset information to obtain a first clustering result, where N is greater than or equal to M and M is greater than or equal to 2. In an embodiment, the first clustering module 820 can be configured to perform operation S220 described above, and details are not described herein again.
[0252] The second clustering module 830 is configured to cluster N non-abnormal-class users and the M target users according to the standard asset information to obtain a second clustering result. In an embodiment, the second clustering module 830 can be configured to perform operation S230 described above, and details are not described herein again.
[0253] The determining module 840 is configured to determine abnormal results of the M target users according to the first clustering result and the second clustering result, where the abnormal results include abnormal users, non-abnormal users, and suspicious users. In an embodiment, the determining module 840 can be configured to perform operation S240 described above, and details are not described herein again.
[0254] According to an embodiment of the present disclosure, the determining module 840 includes a first determining unit, a second determining unit, and a third determining unit.
[0255] The first determining unit is configured to determine, for the nth target user, that the nth target user is an abnormal user in a case where the nth first sub-clustering result indicates that the nth target user belongs to an abnormal class and the nth second sub-clustering result indicates that the nth target user does not belong to a non-abnormal class, where M is greater than or equal to n and n is greater than or equal to 2.
[0256] The second determining unit is configured to determine, for the nth target user, that the nth target user is a non-abnormal user in a case where the nth first sub-clustering result indicates that the nth target user does not belong to an abnormal class and the nth second sub-clustering result indicates that the nth target user belongs to a non-abnormal class.
[0257] The third determining unit is configured to determine, for the nth target user, that the nth target user is a suspicious user in a case where the nth first sub-clustering result indicates that the nth target user belongs to an abnormal class and the nth second sub-clustering result indicates that the nth target user belongs to a non-abnormal class, or the nth first sub-clustering result indicates that the nth target user does not belong to an abnormal class and the nth second sub-clustering result indicates that the nth target user does not belong to a non-abnormal class.
[0258] According to an embodiment of the present disclosure, the abnormal user determining apparatus 800 further comprises a readability index calculating module. The readability index calculating module comprises a calculating unit, a fourth determining unit and a fifth determining unit.
[0259] The calculating unit is configured to, in a case where the mth target user is determined as a suspicious user, calculate a readability index value according to the initial asset information of the mth target user, the readability index value being used to represent the readability of the initial asset information, M≥m≥2.
[0260] The fourth determining unit is configured to, in a case where the readability index value is greater than an index threshold value, determine the mth target user as an abnormal user.
[0261] The fifth determining unit is configured to, in a case where the readability index value is less than or equal to the index threshold value, determine the mth target user as a non-abnormal user.
[0262] According to an embodiment of the present disclosure, the calculating unit comprises a first calculating sub-unit, a second calculating sub-unit, a third calculating sub-unit and a fourth calculating sub-unit.
[0263] The first calculating sub-unit is configured to calculate a total number of words in the initial asset information and a total number of complex words, the complex words representing words with text features higher than a preset text feature threshold value.
[0264] The second calculating sub-unit is configured to calculate an average number of words per line of text according to the total number of words and a total number of lines of the initial asset information.
[0265] The third calculating sub-unit is configured to calculate a ratio of the total number of complex words to the total number of words.
[0266] The fourth calculating sub-unit is configured to calculate the readability index value according to the ratio and the average number of words per line of text.
[0267] According to an embodiment of the present disclosure, the first clustering module 820 comprises a clustering sub-module configured to cluster the N abnormal class users and the M target users according to the standard asset information by a multidimensional scaling analysis method and / or a hierarchical clustering method to obtain a first clustering result, the first clustering result comprising M first sub-clustering results.
[0268] According to an embodiment of the present disclosure, the clustering sub-module comprises a first clustering unit, a second clustering unit, a third clustering unit and a fourth clustering unit.
[0269] The first clustering unit is configured to calculate distances between the N abnormal class users and the M target users according to the standard asset information to obtain an initial distance matrix.
[0270] The second clustering unit is configured to perform iteration on the initial distance matrix based on the multidimensional scaling analysis method, and obtain a final distance matrix in a case where a preset number of iterations is performed.
[0271] The third clustering unit is configured to generate a two-dimensional distance structure diagram according to the final distance matrix.
[0272] The fourth clustering unit is configured to determine M first sub-clustering results according to the two-dimensional distance structure diagram.
[0273] According to an embodiment of the present disclosure, the clustering sub-module further includes a fifth clustering unit, a sixth clustering unit, and a seventh clustering unit.
[0274] The fifth clustering unit is configured to cluster the M target users and the N abnormal-class users based on a hierarchical clustering method and using the standard asset information, until the M target users and the N abnormal-class users are merged into one class.
[0275] The sixth clustering unit is configured to determine M clustering hierarchy information corresponding to the M target users.
[0276] The seventh clustering unit is configured to determine M first sub-clustering results corresponding to the M target users according to the M clustering hierarchy information.
[0277] According to an embodiment of the present disclosure, the clustering sub-module further includes a first updating unit and a second updating unit. After obtaining the first clustering result,
[0278] The first updating unit is configured to filter L target users belonging to the abnormal class and (M-L) target users not belonging to the abnormal class from the M target users based on the M first sub-clustering results, where M≥L≥1.
[0279] The second updating unit is configured to re-cluster the N abnormal-class users and the (M-L) target users according to the standard asset information, and update the first sub-clustering results of the (M-L) target users.
[0280] According to an embodiment of the present disclosure, the processing module 810 includes a first processing unit, a second processing unit, and a third processing unit.
[0281] The first processing unit is configured to filter R transaction features without an association relationship from P transaction features according to correlation coefficients between the P transaction features, where P≥R≥Q≥1.
[0282] The second processing unit is configured to obtain a conversion relationship between the R transaction features and each standard factor, where the standard factor is related to at least one transaction feature.
[0283] The third processing unit is configured to determine a second parameter value of each standard factor in the Q standard factors according to the conversion relationship and a first parameter value of the R transaction features.
[0284] According to an embodiment of the present disclosure, the standard factors are determined after factor analysis on the plurality of transaction features, and the standard factors include a comprehensive factor, a yield profit factor, a cost factor, and a value change yield factor.
[0285] According to an embodiment of the present disclosure, any of the processing module 810, the first clustering module 820, the second clustering module 830, and the determining module 840 can be combined in one module, or any of them can be split into multiple modules. Alternatively, at least part of the function of one or more of these modules can be combined with at least part of the function of the other modules, and implemented in one module.
[0286] According to an embodiment of the present disclosure, at least one of the processing module 810, the first clustering module 820, the second clustering module 830, and the determining module 840 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner that can be integrated or packaged by a circuit, etc. hardware or firmware, or any one of software, hardware and firmware or any appropriate combination of several of them. Alternatively, at least one of the processing module 810, the first clustering module 820, the second clustering module 830, and the determining module 840 can be at least partially implemented as a computer program module which can perform corresponding functions when executed.
[0287] Figure 9 A block diagram of an electronic device suitable for the abnormal user determination method according to an embodiment of the present disclosure is schematically shown.
[0288] As Figure 9 shown, the electronic device 900 according to an embodiment of the present disclosure includes a processor 901 which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 902 or loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a special-purpose microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 901 can also include an on-board memory for cache use. The processor 901 can include a single processing unit or multiple processing units for performing different actions of the method processes according to embodiments of the present disclosure.
[0289] In the RAM 903, various programs and data required for the operation of the electronic device 900 are stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other via the bus 904. The processor 901 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs can also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 can also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.
[0290] According to an embodiment of the present disclosure, the electronic device 900 can further include an input / output (I / O) interface 905, which is also connected to the bus 904. The electronic device 900 can further include one or more of the following components connected to the input / output I / O interface 905: an input part 906 including a keyboard, a mouse, etc.; an output part 907 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 908 including a hard disk, etc.; and a communication part 909 including a network interface card such as a LAN card, a modem, etc. The communication part 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as necessary. A removable medium 911 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 910 as necessary, so that a computer program read out therefrom is installed in the storage part 908 as necessary.
[0291] The present disclosure also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0292] According to an embodiment of the present disclosure, the computer readable storage medium can be a nonvolatile computer readable storage medium, for example, can include, but is not limited to, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or Flash memory), a portable compact disc read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer readable storage medium can include one or more memories such as the ROM 902 and / or the RAM 903 described above and / or one or more memory other than the ROM 902 and the RAM 903.
[0293] Embodiments of the present disclosure also include a computer program product that includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the above-mentioned methods provided by the embodiments of the present disclosure.
[0294] The above-mentioned functions defined in the system / device of the embodiments of the present disclosure are performed when the computer program is executed by the processor 901. According to an embodiment of the present disclosure, the above-mentioned system, device, module, unit, etc. can be implemented by computer program modules.
[0295] In one embodiment, the computer program can rely on a tangible storage medium such as an optical storage device, a magnetic storage device, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of a signal on a network medium and installed and downloaded through the communication part 909 and / or installed from the detachable medium 911. The program codes contained in the computer program can be transmitted by any appropriate network medium, including but not limited to wireless, wired, etc., or any appropriate combination thereof.
[0296] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909 and / or installed from the detachable medium 911. When the computer program is executed by the processor 901, the above-mentioned functions defined in the system of the embodiments of the present disclosure are performed. According to an embodiment of the present disclosure, the above-mentioned system, device, apparatus, module, unit, etc. can be implemented by computer program modules.
[0297] According to embodiments of the present disclosure, program code of the computer programs provided by embodiments of the present disclosure can be written in any combination of one or more programming languages, and specifically, these computer programs can be implemented using a high-level procedural and / or object-oriented programming language, and / or an assembly / machine language. The programming language includes, but is not limited to, a programming language such as Java, C++, Python, "C" language, or a similar programming language. The program code can be executed entirely on a user computing device, partially on a user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case involving a remote computing device, the remote computing device can be connected to the user computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, connected through the Internet by using an Internet service provider).
[0298] The flow diagrams and the block diagrams in the drawings are illustrations of possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0299] Those skilled in the art can understand that the features described in various embodiments of the present disclosure and / or claims can be combined or / and integrated, even if such combinations or integrations are not explicitly described in the present disclosure. In particular, the features described in various embodiments of the present disclosure and / or claims can be combined and / or integrated in various combinations, without departing from the spirit and teachings of the present disclosure. All such combinations and / or integrations are within the scope of the present disclosure.
[0300] The specific embodiments described above are intended to be illustrative of the present disclosure and its best mode known to the inventors at the time of filing the application. It is understood that the present disclosure is not limited to the specific details described herein, and that various modifications, equivalents, and alternatives to the embodiments described herein can be implemented without departing from the spirit and scope of the disclosure.
Claims
1. A method for determining abnormal users, comprising: processing initial asset information of M target users to obtain standard asset information corresponding to the M target users, wherein the initial asset information comprises asset statements of the target users, and M≥2; clustering N abnormal users and the M target users according to the standard asset information to obtain a first clustering result, wherein N≥M≥2; clustering N non-abnormal users and the M target users according to the standard asset information to obtain a second clustering result; determining abnormal results of the M target users according to the first clustering result and the second clustering result, wherein the abnormal results comprise abnormal users, non-abnormal users and suspicious users, the first clustering result comprises M first sub-clustering results, and the second clustering result comprises M second sub-clustering results; in a case that an nth first sub-clustering result indicates that an nth target user belongs to an abnormal class and an nth second sub-clustering result indicates that the nth target user belongs to a non-abnormal class, or the nth first sub-clustering result indicates that the nth target user does not belong to the abnormal class and the nth second sub-clustering result indicates that the nth target user does not belong to the non-abnormal class, determining that the nth target user is a suspicious user, M≥n≥2; in a case that an mth target user is determined to be a suspicious user, calculating a readability index value according to initial asset information of the mth target user, wherein the readability index value is used to indicate readability of the initial asset information, M≥m≥2; in a case that the readability index value is greater than an index threshold, determining that the mth target user is an abnormal user; and in a case that the readability index value is less than or equal to the index threshold, determining that the mth target user is a non-abnormal user. in a case that an nth target user belongs to the abnormal class and does not belong to the non-abnormal class, determining that the nth target user is an abnormal user; in a case that the nth target user does not belong to the abnormal class and belongs to the non-abnormal class, determining that the nth target user is a non-abnormal user.
2. The method of claim 1, wherein, in a case that an mth target user is determined to be a suspicious user, calculating a readability index value according to initial asset information of the mth target user, comprising: calculating a total number of words and a total number of complex words in the initial asset information, wherein the complex words represent words with text features higher than a preset text feature threshold; calculating an average number of words per line of text according to the total number of words and a total number of lines of the initial asset information; 3. The method of claim 1, wherein, calculating a ratio of the total number of complex words to the total number of words; and According to the ratio and the average number of words per text, the readability index value is calculated.
4. The method of claim 1, wherein, The clustering of the N abnormal class users and the M target users according to the standard asset information comprises: According to the standard asset information, the N abnormal class users and the M target users are clustered by a multidimensional scaling analysis method and / or a hierarchical clustering method to obtain a first clustering result, and the first clustering result comprises M first sub-clustering results.
5. The method of claim 4, wherein, The clustering of the N abnormal class users and the M target users according to the standard asset information comprises: According to the standard asset information, the distance between the N abnormal class users and the M target users is calculated to obtain an initial distance matrix; Based on the multidimensional scaling analysis method, the initial distance matrix is iterated to obtain a final distance matrix under a preset iteration number; According to the final distance matrix, a two-dimensional distance structure diagram is generated; and According to the two-dimensional distance structure diagram, the M first sub-clustering results are determined.
6. The method of claim 4, wherein, The clustering of the N abnormal class users and the M target users according to the standard asset information further comprises: Based on the hierarchical clustering method, the M target users and the N abnormal class users are clustered by using the standard asset information until the M target users and the N abnormal class users are merged into one class; M clustering level information corresponding to the M target users is determined; and According to the M clustering level information, the M first sub-clustering results corresponding to the M target users are determined.
7. The method of claim 4, wherein, The first sub-clustering result is used to represent whether the target user belongs to an abnormal class; After obtaining the first clustering result, the method further comprises: Based on the M first sub-clustering results, L target users belonging to an abnormal class and (M-L) target users not belonging to an abnormal class are screened from the M target users, and M≥L≥1; and According to the standard asset information, the N abnormal class users and the (M-L) target users are clustered again to update the first sub-clustering result of the (M-L) target users.
8. The method of claim 1, wherein, The initial asset information comprises P transaction characteristics and a first parameter value for each transaction characteristic, and the standard asset information comprises Q standard factors and a second parameter value for each standard factor, and P≥Q≥1; The processing of the initial asset information of the M target users to obtain the standard asset information corresponding to the M target users comprises: According to the correlation coefficient between the P transaction characteristics, R transaction characteristics without an association relationship are screened from the P transaction characteristics, and P≥R≥Q≥1; A conversion relationship between the R transaction characteristics and each standard factor is obtained, wherein the standard factor is related to at least one transaction characteristic; and According to the conversion relationship and the first parameter value of the R transaction characteristics, a second parameter value of each of the Q standard factors is determined.
9. The method of claim 8, wherein, The standard factors are determined after factor analysis on a plurality of transaction characteristics, and the standard factors include a comprehensive factor, a yield profit factor, a fee factor, and a value change yield factor. 10.An abnormal user determining apparatus, comprising: a processing module configured to process initial asset information of M target users to obtain standard asset information corresponding to the M target users, the initial asset information comprising asset statements of the target users, and M≥2; a first clustering module configured to cluster N abnormal class users and the M target users according to the standard asset information to obtain a first clustering result, and N≥M≥2; a second clustering module configured to cluster N non-abnormal class users and the M target users according to the standard asset information to obtain a second clustering result; and a determining module configured to determine an abnormal result of the M target users according to the first clustering result and the second clustering result, the abnormal result comprising abnormal users, non-abnormal users and suspicious users, the first clustering result comprising M first sub-clustering results, and the second clustering result comprising M second sub-clustering results; the determining module is further configured to: determine the nth target user as a suspicious user when an nth first sub-clustering result indicates that the nth target user belongs to the abnormal class and an nth second sub-clustering result indicates that the nth target user belongs to the non-abnormal class, or when the nth first sub-clustering result indicates that the nth target user does not belong to the abnormal class and the nth second sub-clustering result indicates that the nth target user does not belong to the non-abnormal class, M≥n≥2; calculate a readability index value according to the initial asset information of the mth target user when the mth target user is determined as a suspicious user, the readability index value being used to indicate the readability of the initial asset information, M≥m≥2; determine the mth target user as an abnormal user when the readability index value is greater than an index threshold, or determine the mth target user as a non-abnormal user when the readability index value is less than or equal to the index threshold. 11.An electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein the one or more programs, when executed by the one or more processors, enable the one or more processors to perform the method according to any one of claims 1-9. 12.A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1-9. 13.A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-9.
Citation Information
Patent Citations
Abnormality detection device for diagnosis and treatment behaviors of physician, computer equipment and storage medium
CN113990514A