A target user identification method and apparatus, a terminal device, and a storage medium

By using a decision tree prediction model and the RSA algorithm for data alignment and encryption in the telecom fraud identification model, the problems of single data source and leakage of sensitive information are solved, thereby improving the identification accuracy and model building efficiency.

CN115659261BActive Publication Date: 2026-02-13GUANGZHOU HANTELE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211111838.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-13
Publication Date
2026-02-13
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

Existing telecom fraud identification models have low accuracy, rely on a single data source, pose risks in data exchange, and cannot fully utilize computing power.

Method used

The decision tree prediction model is used to obtain users' personal information from different dimensions. The CART decision tree prediction model is used to extract data from data systems from different dimensions. The RSA algorithm is used for data alignment and encryption to avoid leakage of sensitive information, reduce labor costs, and improve model building efficiency.

Benefits of technology

It improved the accuracy of the telecom fraud identification model, reduced the risk of sensitive information leakage, saved manpower costs, and improved the efficiency of model building.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115659261B_ABST
    Figure CN115659261B_ABST
Patent Text Reader

Abstract

The application discloses a target user identification method and device, a terminal device, and a storage medium. After obtaining a first personal information of a user in different dimensions, the first personal information is input into a pre-constructed decision tree prediction model to determine whether the user is a target category user. The application increases the data source in the target user identification process by obtaining the first personal information in different dimensions, avoids the technical problem that the decision tree prediction model has a low identification accuracy due to too single data source, and stores the node information of the decision tree prediction model in different dimensional data systems, so that sensitive information from different data systems in the decision tree prediction model can be stored in the original data system, the risk of sensitive information leakage is reduced, a special person does not need to be arranged to be responsible for data security and secrecy, human cost is saved, and the efficiency of the decision tree prediction model in construction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of artificial intelligence, and in particular to a target user identification method and device, a terminal device, and a storage medium. BACKGROUND

[0002] In the prior art, a telecommunications fraud identification model is generally used to identify fraudsters. The existing telecommunications fraud identification model is mostly based on telecommunications operator data and is constructed by combining various traditional machine learning algorithms or deep learning algorithms. Even though the introduction of a distributed system has solved the problems of data storage and model calculation, there is still the technical problem of low identification accuracy.

[0003] To sum up, how to improve the identification accuracy of the telecommunications fraud identification model has become a technical problem that needs to be solved at present. SUMMARY

[0004] Embodiments of the present application provide a target user identification method and device, a terminal device, and a storage medium, which solve the technical problem of low identification accuracy of the telecommunications fraud identification model in the prior art.

[0005] In a first aspect, embodiments of the present application provide a target user identification method, comprising the following steps:

[0006] obtaining first personal information of a user to be identified in different dimensions;

[0007] inputting the first personal information into a pre-constructed decision tree prediction model, so that the decision tree prediction model extracts first data corresponding to each node from the first personal information, determines an accessible path according to the first data and classification conditions of the nodes; wherein node information of each node is stored in a data system in the corresponding dimension, and the node information includes classification conditions of the node and a next accessible node, and the node is obtained by cutting according to personal information of a user in different dimensions;

[0008] determining whether the user is a target category user according to a leaf node on the accessible path.

[0009] In a second aspect, embodiments of the present application provide a target user identification device, comprising:

[0010] a personal information acquisition module configured to obtain first personal information of a user to be identified in different dimensions;

[0011] a user identification module, configured to input the first personal information into a pre-constructed decision tree prediction model, so that the decision tree prediction model extracts first data corresponding to each node from the first personal information, and determines an accessible path according to the first data and classification conditions of the nodes;

[0012] a category determination module, configured to determine whether the user is a target category user according to leaf nodes on the accessible path.

[0013] In a third aspect, an embodiment of the present application provides a terminal device, which comprises a processor and a memory.

[0014] The memory is configured to store a computer program and transmit the computer program to the processor.

[0015] The processor is configured to execute the target user identification method according to instructions in the computer program.

[0016] In a fourth aspect, an embodiment of the present application provides a storage medium storing computer executable instructions, which, when executed by a computer processor, are configured to execute the target user identification method.

[0017] According to the above, after obtaining the first personal information of the user in different dimensions, the first personal information is input into a pre-constructed decision tree prediction model to determine whether the user is a target category user. According to the embodiment of the present application, the first personal information in different dimensions is obtained to increase the data source in the target user identification process, so as to avoid the technical problem that the decision tree prediction model has a low recognition accuracy due to too single data source. In addition, in the embodiment of the present application, the node information of the decision tree prediction model is stored in the data system in different dimensions, so that the sensitive information from different data systems in the decision tree prediction model can be stored in the original data system, the risk of sensitive information leakage is reduced, and the data of the decision tree prediction model does not need to be saved to a specified data center, so that a person in charge of data security and secrecy does not need to be arranged, the human cost is saved, and the efficiency of the decision tree prediction model in construction is improved. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 FIG. 1 is a flowchart of a target user identification method according to an embodiment of the present application.

[0019] Figure 2A structural schematic diagram of a decision tree prediction model provided for an embodiment of the present application.

[0020] Figure 3 A flowchart of constructing a decision tree prediction model provided for an embodiment of the present application.

[0021] Figure 4 A schematic diagram of aligning user personal information provided for an embodiment of the present application.

[0022] Figure 5 A structural schematic diagram of constructing a decision tree prediction model provided for an embodiment of the present application.

[0023] Figure 6 A structural schematic diagram of a target user identification device provided for an embodiment of the present application.

[0024] Figure 7 A structural schematic diagram of a terminal device provided for an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following description and drawings are illustrative of specific embodiments of the application and are not intended to be limiting thereof Any variations of these specific embodiments can be considered to be within the scope of the application. Individual components and functions can be optional, and the order of operations can be varied. Parts and features of some embodiments can be included or substituted in or for parts and features of other embodiments. The scope of the embodiments of the application encompasses the entire scope of the claims, and all available equivalents of the claims. In this document, the term "application" can be used to refer to one or more embodiments of the application, and the terms "application" and "embodiments" can be used interchangeably. The term "application" can be used to refer to the entire application, or to individual features or groups of features of the application. Unless specifically stated otherwise, the use of terms such as first and second, etc., does not imply any actual relationship or order between structures, operations, or materials. The use of the term "including" or "containing" to describe the elements of a process, method, or device, or the like, is intended to be non-exclusive, such that a process, method, or device, or the like that "includes" or "contains" a list of elements is not necessarily limited to those elements, but can include other elements not expressly listed or inherent to such process, method, or device. The various embodiments are described in a progressive fashion, each embodiment highlighting a different aspect of the embodiments disclosed. The same or similar parts and features of the embodiments are cross-referenced among the various embodiments. The structure, products, etc. disclosed in the embodiments are described in a relatively simple manner, as they correspond to the parts disclosed in the embodiments, and the relevant parts are cross-referenced in the method part.

[0026] Telecommunications fraud refers to the criminal behavior of creating false information and setting up a fraud to implement fraud through telephone, network and short message. The traditional telecommunications fraud identification model mostly uses network communication features and combines classification and discrimination model algorithm to identify fraudulent groups, generally including the following steps:

[0027] (a) Obtain abnormal communication behavior characteristics of users through network crawler and other technical means, such as users who sell a large number of mobile cards, email, bank cards or identity card information on network forums.

[0028] (b) Extract abnormal communication behavior of users such as one card multi-card users, number calling behavior active, abnormal time period and pseudo base station communication from telecommunications communication data.

[0029] (c) Feature extraction is performed on user data, and user data is labeled to build a data set. According to the actual situation of data set and feature selection, decision tree, Boosting or other classification and discrimination model is selected to build, and telecommunications fraud identification model is obtained.

[0030] (d) For data without annotation, K-means can be used to build clustering model to obtain telecommunications fraud identification model.

[0031] (e) The communication data in the process of telecommunications fraud may have certain time sequence event characteristics, such as a large number of short time communication, suddenly entering long time communication state with a certain user, and having certain communication behavior contact with the user in the subsequent period. Such data can use deep learning algorithm model such as long short-term memory artificial neural network (LSTM) to complete the identification of telecommunications fraud.

[0032] At present, due to the large amount of recorded data of operators, various communication behaviors, account behaviors and residence information of users are recorded in the server and deleted regularly, so in the process of building telecommunications fraud identification model, data collection, conversion, storage and model building need to be completed in a very short time. At the same time, with the increasing identification dimension of telecommunications fraud groups, the data collection accuracy has been improved to a certain extent, and the traditional offline single machine identification model can no longer meet the needs of related work, so various telecommunications fraud identification models based on distributed computing architecture design have emerged. By combining hdfs, sqoop, hive, spark Mllib and other component platforms in hadoop ecosystem, a series of work from data storage to model building can be efficiently completed, which has become the mainstream processing method of telecommunications fraud work in various cities and operators.

[0033] However, the existing telecommunications fraud identification model still has the following defects;

[0034] (1) The data source is too singular, and the data dimensions for identifying telecom fraud are insufficient.

[0035] Currently, the construction of telecom fraud identification models is generally carried out independently within each institution. Therefore, the data sampled during the construction process has strong industry characteristics, such as communication data from operators and user transaction data from banks. Even if there is some information exchange between departments, it is far from sufficient for the construction of telecom fraud identification models.

[0036] (2) Data exchange carries significant risks.

[0037] Led by relevant ministries, there may be some cross-enterprise and inter-departmental cooperation, such as telecom operators, banks, third-party payment platforms, and public security technology departments, to complete the construction of telecom fraud detection models. However, cross-departmental cooperation also brings about various highly sensitive data security issues, including how data is transmitted, how it is stored, and how to design data sensitivity classification and isolation.

[0038] (3) The computing performance cannot be fully utilized.

[0039] Given the sensitive nature of the data, operators may centralize the data in a dedicated working group to handle model building when constructing telecom fraud detection models. In cross-enterprise and cross-departmental collaborations, data is typically imported into a designated data center, with specific personnel responsible for data security and confidentiality. However, both of these approaches severely limit overall operational capacity, including slow model design and construction due to limited personnel and insufficient computing performance due to limitations in the number of platform servers.

[0040] In summary, to improve the accuracy of telecommunications fraud detection models, this invention provides a target user identification method, such as... Figure 1 As shown, Figure 1 This is a flowchart illustrating a target user identification method provided in an embodiment of the present invention. The target user identification method provided in this embodiment can be executed by a terminal device, which can be implemented through software and / or hardware. The terminal device can consist of two or more physical entities, or it can consist of a single physical entity. For example, the terminal device can be a computer, a host computer, a tablet, or other devices. The method includes the following steps:

[0041] Step 101: Obtain the first personal information of the user to be identified from different dimensions.

[0042] In the embodiment, first, the first personal information of the user to be identified in different dimensions is acquired. For example, when the identification of the user of telecom fraud is needed, the first personal information of the user to be identified can be the age of the user, the number of credit cards opened by the user, the number of long-term low-balance accounts, the number of calls per month, the average call duration, the number of SMS sent, the roaming situation, whether there is a criminal record, the entry and exit situation, and the like. It can be understood that in the embodiment, the first personal information can be set according to the actual situation, and the specific content of the first personal information is not limited in the embodiment.

[0043] In step 102, the first personal information is input into the decision tree prediction model constructed in advance, so that the decision tree prediction model extracts the first data corresponding to each node from the first personal information, and determines the reachable path according to the first data and the classification condition of each node; wherein the node information of each node is stored in the data system of the corresponding dimension, and the node information includes the classification condition of the node and the next reachable node, and the node is obtained by cutting the personal information of the user in the data system of different dimensions.

[0044] After the first personal information of the user to be identified is acquired, the first personal information is input into the decision tree prediction model constructed in advance to determine whether the user to be identified is a target category user. Specifically, in the embodiment, the decision tree prediction model adopts a CART decision tree prediction model, and the nodes in the CART decision tree prediction model include branch nodes and leaf nodes. Each branch node includes two child nodes, and the leaf node refers to a node without child nodes, also known as a terminal node. In the embodiment, each node in the CART decision tree prediction model is obtained by cutting the personal information of the user in the data system of different dimensions. Different dimensions of the data system refer to data systems storing different personal information of the user. For example, in the process of identifying telecom fraud personnel, the data systems of different dimensions can select the data system of a bank, the data system of a communication operator, and the data system of the public security. The personal information of the user stored in the data system of the bank includes the number of credit cards opened by the user and the number of long-term low-balance accounts, and the like. The personal information of the user stored in the data system of the communication operator includes the number of calls per month, the average call duration, the number of SMS sent, and the roaming situation, and the like. The personal information of the user stored in the data system of the public security includes whether there is a criminal record and whether there is an entry and exit situation, and the like.

[0045] In the process of constructing the CART decision tree prediction model, each node can be obtained by cutting the personal information of the user in the data system of different dimensions. In an embodiment, each node can be obtained by cutting the personal information by calculating the Gini coefficient. For example, as shown in FIG. 1, the Gini coefficient of the node A is calculated by the following formula: Figure 2As shown, in the CART decision tree prediction model, the classification condition for the root node 1 at the first level is whether the number of invalid SIM cards is greater than X. The classification conditions for branch nodes 2 and 3 at the second level are whether there is a criminal record and whether the average call duration is greater than Y minutes, respectively. The terminal nodes on each branch are leaf nodes, which include two categories: telecom fraudsters and non-telecom fraudsters. Furthermore, it should be further explained that in this embodiment, the node information of each node is stored in the corresponding dimensional data system. The node information includes the node's classification conditions and the next reachable node. For example, for... Figure 2 The root node 1 in the first level stores the node information in the bank's data system. The node information of branch node 2 in the second level is stored in the public security data system. The node information of branch node 3 in the second level is stored in the telecommunications operator's data system. The node information includes the node's classification conditions and the next reachable node. When a node is a branch node, there are two next reachable nodes.

[0046] When the first piece of personal information is input into a pre-built decision tree prediction model, the model extracts the first data corresponding to each node from the first piece of personal information and determines the reachable path based on the first data and the classification conditions of each node. For example, when the first piece of personal information is input into a pre-built decision tree prediction model, the model extracts the first data corresponding to each node and determines the reachable path. Figure 2 After the CART decision tree prediction model is shown, the CART decision tree prediction model first obtains the classification conditions of the root node 1 and the next reachable node from the bank's data system. Then, it extracts the number of invalid cards of the user from the first personal information. If the number of invalid cards is greater than X, then according to the determination that the next reachable node is the second-level branch node 2, it continues to obtain the classification conditions of branch node 2 and the next node from the public security data system. It continues to obtain whether the user has a criminal record from the first personal information and judges whether it meets the classification conditions of branch node 2. This process continues until the leaf node on the reachable path is reached, thereby confirming the reachable path corresponding to the user to be identified.

[0047] Step 103: Determine whether the user is a target category user based on the leaf nodes on the reachable path.

[0048] Finally, based on the leaf nodes on the reachable path, it is determined whether the user belongs to the target category. For example, if a leaf node on the reachable path represents a user suspected of being a telecom fraudster, then the user to be identified is a telecom fraudster. If a leaf node on the reachable path represents a user who is not a telecom fraudster, then the user to be identified is not a telecom fraudster.

[0049] The above, the embodiment of the application inputs the first personal information into the decision tree prediction model constructed in advance to determine whether the user is a target category user after obtaining the first personal information of the user in different dimensions. The embodiment of the application increases the data source in the target user identification process by obtaining the first personal information in different dimensions, avoids the technical problem that the decision tree prediction model has low identification accuracy due to too single data source, and in addition, the node information of the decision tree prediction model in the embodiment of the application is stored in different dimensional data systems, so that sensitive information from different data systems in the decision tree prediction model can be stored in the original data system, reducing the risk of sensitive information leakage, and at the same time, the data of the decision tree prediction model does not need to be saved to the designated data center, and special personnel do not need to be arranged to be responsible for data security work, which saves the cost of manpower, and also improves the efficiency of the decision tree prediction model when constructing.

[0050] In one embodiment, as shown in Figure 3 The decision tree prediction model is constructed in advance, including the following steps:

[0051] Step 201, align the personal information in the data systems in different dimensions, and determine the target users common to the data systems in different dimensions. The personal information of different users is stored in the data systems in different dimensions in advance.

[0052] First, in this embodiment, the user groups of the data systems in different dimensions do not completely overlap because the personal information of different users is stored in the data systems in different dimensions in advance, so it is necessary to align the personal information in the data systems in different dimensions, so as to determine the target users common to the data systems in different dimensions. For example, the data systems of banks, public security and communication operators respectively store the personal information of users with ids (1, 2, 3, 4, 5), (2, 3, 4, 5, 6) and (4, 5, 6, 7, 8), and data alignment is to let the data systems of banks, public security and communication operators know that only the personal information of users with ids 4 and 5 needs to participate in the construction of the decision tree prediction model.

[0053] On the basis of the above embodiment, the personal information in the data systems in different dimensions is aligned in step 201 to determine the target users common to the data systems in different dimensions, including:

[0054] Step 2011, align the personal information in the data systems in different dimensions based on the RSA algorithm to determine the target users common to the data systems in different dimensions.

[0055] In the embodiment, the data alignment can be achieved by using the RSA encryption algorithm with homomorphic encryption characteristics. The homomorphic encryption refers to that the processing of the data subjected to the homomorphic encryption obtains an output, and the decryption of the output is the same as the output obtained by processing the original data without encryption by using the same method. For example, for any data a and b, when new data [a] and [b] are generated by using an encryption method, the original data a and b and the encrypted data [a] and [b] have a certain mathematical relationship, and the encryption method can be considered as a homomorphic encryption method. For data a and b, if [a]+[b]=[a+b], it is a homomorphic encryption algorithm supporting addition; if [c*a]=c*[a], where c is any real number, it is a homomorphic encryption algorithm supporting multiplication. For data a and b, if [a]*[b]=[a*b], and the result of multiplying the original data a and b after encryption is the same as the result of multiplying the original data a and b and then encrypting, it is a homomorphic encryption algorithm supporting multiplication, and the RSA algorithm is a typical homomorphic encryption algorithm supporting multiplication.

[0056] The existing RSA algorithm is often processed by using a random number salt, but the result obtained after each encryption is different due to using the same public key, and is not suitable for the data alignment process in the embodiment. Therefore, in this stage, the original RSA algorithm is used to process the personal information in the data systems of different dimensions, and the signature data subjected to the hash processing is compared to align the data. Specifically, as shown in FIG. 11, the method comprises the following steps: Figure 4

[0057] Step 20111, in the personal information in the data systems of different dimensions, the identity information of each user is obtained, and the identity information is subjected to hash calculation to obtain hash identity information.

[0058] Firstly, before the personal information in the data systems of different dimensions is subjected to hash calculation, the identity information of each user is obtained in the personal information in the data systems of different dimensions. The identity information of each user can be an identity card number of each user or a unique identifier labeled later. Then, the identity information in different data systems is subjected to hash calculation to obtain hash identity information. The purpose of the hash calculation is to format the data, which facilitates data transmission and improves the cracking difficulty.

[0059] Step 20112, in the data systems of different dimensions, a target dimension data system having a first private key is determined.

[0060] ​Afterwards, in the different dimension data systems, a target dimension data system with the first private key is determined, and the first private key is sent to the target dimension data system, and the first public key is sent to each dimension data system.

[0061] Step 20113, in the remaining data systems except the target dimension data system, multiply the hash identity information in each data system by the random number encrypted by the first public key to obtain the first encrypted personal information, and send the first encrypted personal information to the target dimension data system.

[0062] Afterwards, in the remaining data systems except the target dimension data system, a random number (for example, r) is selected, which is encrypted by the first public key to obtain r_sec, and multiplied by the hash identity information Hb to obtain the first encrypted personal information r_sec*Hb, and then the first encrypted personal information r_sec*Hb is sent to the target dimension data system.

[0063] Step 20114, in the target dimension data system, the first encrypted personal information is decrypted using the first private key to obtain the first decrypted data, and the first decrypted data is sent to the corresponding dimension data system to divide the first decrypted data by the random number in the corresponding dimension data system to obtain the first signature data.

[0064] After each data system sends the first encrypted personal information r_sec*Hb to the target dimension data system, although the target dimension data system has the first private key, due to the existence of the random number r, the decryption result is not the original content, so that the target dimension data system cannot learn the hash identity information of the remaining data systems. When the target dimension data system decrypts the first encrypted personal information r_sec*Hb, due to the multiplication homomorphism property of the RSA algorithm itself, encrypting (a*b) is equal to (encrypting a)*(encrypting b). In this embodiment, the encryption formula of the RSA algorithm is:

[0065] Ciphertext=(plaintext^e)%n

[0066] plaintext=(ciphertext^d)%n

[0067] Where e is the first public key, n is the product of two large prime numbers, and d is the first private key.

[0068] As can be seen from the above formula, the encryption and decryption process can actually be interchanged, so decrypting the first encrypted personal information r_sec*Hb can actually be seen from another angle as encrypting the decryption of the first encrypted personal information r_sec*Hb using the first private key, so it also has the multiplication homomorphism property.

[0069] After the first encrypted personal information r_sec*Hb is decrypted to obtain the first decrypted data in the target dimension data system, the first decrypted data can be regarded as a signature, and then the first decrypted data is sent to the corresponding dimension data system, so that the corresponding dimension data system divides the first decrypted data by the random number r generated in advance after receiving the first decrypted data, and the first signature data sig b is obtained.

[0070] Step 20115, the hash identity information in the target dimension data system is decrypted using the first private key to obtain the second signature data.

[0071] Meanwhile, in the target dimension data system, the hash identity information is decrypted using the first private key to obtain the second signature data sig a.

[0072] Step 20116, according to each first signature data and second signature data, the alignment of the identity information is performed to determine the target user common to the data systems in different dimensions.

[0073] Finally, according to the second signature data sig a in the target dimension data system and the first signature data sig b in the data system in other dimensions, the alignment of the identity information is performed to determine the target user common to the data systems in different dimensions.

[0074] On the basis of the above embodiment, the alignment of the identity information is performed according to the first signature data and the second signature data in step 20116 to determine the target user common to the data systems in different dimensions, including:

[0075] Step 201161, the first hash value of each first signature data and the second hash value of the second signature data are calculated.

[0076] In this embodiment, the second hash value of the second signature data sig a is calculated in the target dimension data system, and the first hash value of the first signature data sig b is calculated in the data system in the remaining dimension.

[0077] Step 201162, each first hash value and second hash value are compared to determine the target user common to the data systems in different dimensions.

[0078] Then, the data systems in the remaining dimensions respectively send the first hash value to the target dimension data system, and the target dimension data system compares the first hash value with each second hash value to determine the intersection of the hash identity information in each dimension data system, thereby determining the target user common to the data systems in different dimensions, and completing the alignment of the data. Then, the intersection of the hash identity information in the target dimension data system is sent to the remaining data systems for synchronization, and the specific process is as shown in Figure 4 .

[0079] Step 202, determining target personal information corresponding to the target user in the different dimension data system, calculating Gini coefficient according to the target personal information, and dividing the target personal information according to the Gini coefficient to obtain each node of the decision tree prediction model and node information of each node.

[0080] After determining the target user common to the different dimension data systems, target personal information corresponding to the target user in the different dimension data systems is determined respectively. Then, Gini coefficient is calculated according to the corresponding target personal information in the different dimension data systems, and the target personal information is divided according to the calculated Gini coefficient, so as to obtain each node of the decision tree prediction model and node information of each node. Specifically, in the embodiment, the CART decision tree prediction model completes the construction of the whole tree by calculating the Gini coefficient of each node branch. For a sample, the calculation formula of the Gini coefficient is When performing binary classification, the calculation formula can be simplified as Gini = 2p(1-p) = 2pq. Wherein, classm represents that the model has m classifications, and in the embodiment, class = 1 is "telecommunication fraud personnel" and class = 2 is "non-telecommunication fraud personnel". P represents the probability of belonging to a certain classification under a certain division.

[0081] And in each data system, the Gini coefficient calculated by dividing all the target personal information is calculated by the following formula:

[0082]

[0083] Wherein, D represents the number of target personal information, A represents the division method, D1 and D2 represent the number of target personal information in each category according to different division methods, and Gini(D1) and Gini(D2) represent the Gini coefficient of target personal information in D1 and D2, that is, Gini = 2p(1-p) in the above.

[0084] For example, taking the construction of the decision tree prediction model corresponding to the data system of the public security department as an example, it is assumed that the following target personal information exists in the data system of the public security department:

[0085] (1) There are 10 people in the age group of teenagers, 3 people belong to classification a, and 7 people belong to classification b.

[0086] (2) There are 20 people in the age group of middle-aged people, 15 people belong to classification a, and 5 people belong to classification b.

[0087] (3) There are 10 people in the age group of old people, 6 people belong to classification a, and 4 people belong to classification b.

[0088] For example, all target personal information is divided into two parts D1 and D2 by cutting into youth and non-youth. There are 10 youths and 30 non-youths, i.e. D1 = 10, D2 = 30.

[0089] Gini(D1) = 2 * p * (1-p), where p is the probability of classification a in D1. Among the 10 youths, 3 are classified as a, so p = 3 / 10, 1-p = 7 / 10, thus:

[0090] Gini(D1) = 2 * p * (1-p) = 2 * 1 / 10 * 7 / 10

[0091] Gini(D2) Similarly, there are 30 people in D2 (20 middle-aged and 10 old), of which 21 (15 middle-aged and 6 old) belong to classification a, so p = 21 / 30, 1-p = 9 / 30, thus:

[0092] Gini(D2) = 2 * p * (1-p) = 2 * 21 / 30 * 9 / 30

[0093] Gini(D, A = youth) = D1 / D * Gini(D1) + D2 / D * Gini(D2)

[0094] = (10 / 40 * 2 * 1 / 10 * 7 / 10) + (30 / 40 * 2 * 21 / 30 * 9 / 30) = 0.42

[0095] The calculation of the Gini coefficient for middle-aged and old people is similar to that for young people:

[0096]

[0097]

[0098] After calculating the Gini coefficients of the three age segmentations respectively, the Gini coefficient obtained by the first segmentation method of young and non-young is the lowest, so the best segmentation method for the age field is to divide the teenagers into one group and the other ages into another group. Similar calculations are performed on other fields in the target personal information to obtain the best segmentation points on other fields. Then compare the Gini coefficients of the best segmentation of each field, and the segmentation method corresponding to the global minimum Gini coefficient is the optimal segmentation method, so as to divide the target personal information in the public security data system into left and right segments, and perform the same processing on the target personal information corresponding to the left segment and the target personal information corresponding to the right segment. Finally, the decision tree prediction model corresponding to the target personal information in the public security data system is obtained, that is, the construction process of the decision tree prediction model of a single data source is completed. In this embodiment, since the target personal information in different dimensional data systems is not the same, after calculating the Gini coefficients of different dimensional data systems, a global comparison is also needed to determine the segmentation method with the lowest Gini coefficient.

[0099] On the basis of the above embodiment, in step 202, the Gini coefficient is calculated according to the target personal information, and the target personal information is segmented according to the Gini coefficient to obtain each node of the decision tree prediction model and the node information of each node, including:

[0100] Step 2021, obtaining the classification information of the target user from the first target personal information of the specified first dimensional data system, and the classification information includes information about whether different users are target category users.

[0101] In this embodiment, first, the classification information of the target user is obtained from the first target personal information of the specified first dimensional data system, and the classification information includes information about whether different users are target category users. It should be noted that, for example, when a communication operator, a bank and a public security cooperate to construct a decision tree prediction model, the classification information may exist only in one data system, such as the classification information of whether a user is a telecom fraudster, which exists only in the public security data system. Therefore, in this embodiment, the classification information of different users needs to be obtained from the first target personal information of the specified first dimensional data system.

[0102] Step 2022, encrypting the classification information to obtain encrypted classification information, and distributing the encrypted classification information to the remaining second dimensional data systems.

[0103] For privacy protection or other issues (e.g. case review process needs to be kept secret, etc.), the public security department may not want other departments to know whether a user is a telecom fraud personnel when building a model, and the data systems of the rest of the departments do not have the classification information of the user. Therefore, after obtaining the classification information of the target user in the first dimension data system, the classification information needs to be encrypted in the first dimension data system to obtain encrypted classification information, and then the encrypted classification information is distributed to the second dimension data systems of the rest of the departments, such as the data systems of the bank and the communication operator, so that the second dimension data systems calculate the Gini coefficient based on the encrypted classification information.

[0104] Step 2023, in the second dimension data system, obtain a second classification label from the second target personal information, encrypt the second classification label to obtain an encrypted classification label, and the second target personal information in each second dimension data system includes a plurality of second classification labels.

[0105] Then, the second classification label is obtained from the second target personal information of each second dimension data system. The second classification label is data summarized and sorted by each data system according to business logic, which can be classification data (such as the above-mentioned old, middle-aged, and young), or specific numerical data (such as 10 years old, 11 years old, 20 years old, etc.). In actual processing, a plurality of second classification labels corresponding to the user's personal information can be generated according to specific business experience and feature engineering before modeling, for example, the second classification labels of the bank include: whether it is one card with multiple cards, whether there are a large number of blank accounts, the number of credit card opening, and the number of long-term low-balance accounts, etc., the second classification labels of the communication operator include: the number of monthly calls, the average call duration, the number of SMS sending, and the roaming situation, etc. The definition and division of the second classification label are obtained according to the business experience of each department, and the labels mastered by different departments are different. In this embodiment, the specific content of the second classification label is not limited. After obtaining the second classification label of each second dimension data system, the second classification label is encrypted to obtain an encrypted second classification label, so as to avoid other data systems from knowing the specific content of the second classification label.

[0106] Step 2024, in the first dimension data system, the Gini coefficient is calculated according to the classification information, in the second dimension data system, the Gini coefficient is calculated according to the encrypted classification information and the encrypted classification label, and the target personal information is divided according to the Gini coefficient to obtain each node of the decision tree prediction model.

[0107] Then, in the first dimension data system, the Gini coefficient of the corresponding target personal information when divided under different classification labels can be calculated according to the classification information; in the second dimension data system, the Gini coefficient of the corresponding target personal information when divided under different classification labels can be calculated according to the encrypted classification information and the encrypted classification label. Finally, the target personal information in the data system of different dimensions is divided according to the smallest Gini coefficient, so as to obtain each node of the decision tree prediction model.

[0108] On the basis of the above embodiment, the Gini coefficient is calculated in the first dimension data system according to the classification information in step 2024, and the Gini coefficient is calculated in the second dimension data system according to the encrypted classification information and the encrypted classification label, and the target personal information is divided according to the Gini coefficient, so as to obtain each node of the decision tree prediction model, including:

[0109] Step 20241, in the first dimension data system, the first Gini coefficient corresponding to different first classification labels in the first target personal information is calculated according to the first classification label and the classification information, and the first target personal information also includes the first classification label.

[0110] In this embodiment, the first target personal information saved in the first dimension data system also includes the first classification label corresponding to the first target personal information. First, in the first dimension data system, the first Gini coefficient corresponding to different first classification labels is calculated according to the first classification label and the classification information.

[0111] Step 20242, in each second dimension data system, the second Gini coefficient corresponding to each encrypted classification label in the second target personal information is calculated according to the encrypted classification label and the encrypted classification information.

[0112] At the same time, in each second dimension second data system, the second Gini coefficient corresponding to each encrypted classification label in the second target personal information is calculated according to the encrypted classification label and the encrypted classification information.

[0113] Exemplarily, after the personal information alignment, the target users commonly owned in different dimensional data systems are [id=1, 2, 3, 4...8, 9, 10] in total 10. When calculating to a branch node of the decision tree prediction model, the data system of the public security department owns data of users [id=1, 2, 3, 4, 5, 6] at the node, so a vector [1, 1, 1, 1, 1, 1, 0, 0, 0, 0] can be generated based on the aligned target users and the users owned at the current node, indicating whether the personal information of the corresponding target user is available, and the vector is split by the different classification information. Assuming that 135 is type a and 246 is type b, the split vectors are type a: [1, 0, 1, 0, 1, 0, 0, 0, 0, 0] and type b: [0, 1, 0, 1, 0, 1, 0, 0, 0, 0], and the public security data system sends the two vectors to the data systems of other departments respectively after encrypting the 1s and 0s in the two vectors using the first public key, indicating that there are two classifications at present, and the users under each classification are represented by the corresponding vector.

[0114] The data systems of each department generate corresponding label vectors and perform encryption processing for each classification label according to the personal information of the local target users. For example, the label vector of the 10 people in the data system of a certain department is defined as [1, 1, 0, 0, 0, 0, 0, 0, 0, 0] for youth as 1 and non-youth as 0, and the classification vectors transmitted by the public security department are a: [1, 0, 1, 0, 1, 0, 0, 0, 0, 0] and b: [0, 1, 0, 1, 0, 1, 0, 0, 0, 0]. Multiplying the label vector and the classification vector can obtain the vector of the youth users under classification a as [1, 0, 0, 0, 0, 0, 0, 0, 0, 0] and the vector of the youth users under classification b as [0, 1, 0, 0, 0, 0, 0, 0, 0, 0]. Then the Gini coefficient when dividing at the current node by youth / non-youth can be obtained according to the above Gini coefficient calculation formula. Since there are a total of 6 people at the node, 3 people are type a and 3 people are type b, of which 1 is youth and 2 are non-youth in type a, and 1 is youth and 2 are non-youth in type b, then:

[0115] Gini(D) = youth / (youth+non-youth)*Gini(youth)+non-youth / (youth+non-youth)*Gini(non-youth) = 2 / 6*2*1 / 3*(1-1 / 3)+4 / 6*2*1 / 3*(1-1 / 3) = 4 / 9.

[0116] The Gini coefficients of other classification labels are calculated in the same way, which will not be described in detail in this embodiment. Since the encryption algorithm has the homomorphic encryption feature [a*b]=[a]*[b], it is very convenient to perform relevant calculation and processing on the encrypted 1 and 0. In addition, in the process of calculating the Gini coefficient, the vector is formed by using the encrypted [0] and [1] instead of the original 0 and 1, so as to ensure the security of information in the transmission process.

[0117] Step 20243, selecting the first target Gini coefficient with the minimum value in the first Gini coefficients, and selecting the second target Gini coefficient with the minimum value in the second Gini coefficients.

[0118] After the first Gini coefficients corresponding to each classification label are calculated in the data system of the first dimension, the first target Gini coefficient with the minimum value is selected from all the first Gini coefficients. Similarly, in each data system of the second dimension, the second target Gini coefficient with the minimum value is selected. For example, in the data system of a communication operator, the second Gini coefficients of whether the network age is greater than 5 years, whether the network age is greater than 10 years, whether a plurality of mobile phone cards are owned, and whether the monthly short message sending volume is greater than a certain value are obtained. Since the second Gini coefficients are in plaintext, the comparison of the second Gini coefficients can be performed locally in the data system of the communication operator to determine the second target Gini coefficient with the minimum value, and the classification label corresponding to the second target Gini coefficient is taken as the optimal split mode of the data system of the communication operator.

[0119] Step 20243, comparing the first target Gini coefficient and the plurality of second target Gini coefficients to determine the third target Gini coefficient with the minimum value.

[0120] After the optimal split mode of each dimension is calculated in the data system, a global comparison needs to be performed to determine which split mode is the best split mode. For example, in the data system of a public security department, whether a person has a criminal record is calculated as the current optimal split mode A, in the data system of a bank, whether a person has a large number of blank bank cards is calculated as the optimal split mode B, and in the data system of a communication operator, whether a person has a plurality of mobile phone cards is calculated as the optimal split mode C. At this time, it is necessary to compare and determine which one of the three split modes is the globally optimal split mode.

[0121] In an embodiment, encryption protection also needs to be performed in the comparison process of the first target Gini coefficient and the second target Gini coefficient. In an embodiment, the Paillier algorithm based on homomorphic addition and subtraction encryption can be used to encrypt and protect the first Gini coefficient and the second Gini coefficient. For example, the comparison process of the Gini coefficients encrypted based on the Paillier algorithm is as follows:

[0122] In the data system of the bank, a pair of public key and private key is generated, the data system of the communication operator encrypts the second target Gini coefficient by the public key respectively to obtain the second encrypted Gini coefficient [a2], and the data system of the public security encrypts the first target Gini coefficient by the public key to obtain the first encrypted Gini coefficient [a1]. Then, the first encrypted Gini coefficient [a1] is sent to the data system of the communication operator, or the second encrypted Gini coefficient [a2] is sent to the data system of the public security. Since neither party has the private key, the original data cannot be obtained. According to the addition (or subtraction) method of homomorphic encryption, [a2]-[a1]=[a2-a1], the result of subtraction is sent to the data system of the bank which has the private key. After decryption, the bank's data system can only obtain the result of subtraction, and cannot reversely deduce the real content of [a2] and [a1]. At this time, the data system of the communication operator and the data system of the public security only publish which target Gini coefficient is smaller, and do not publish the calculation result, so that the other party cannot obtain the real result of the Gini coefficient.

[0123] Then, the data system of the party with the smaller target Gini coefficient (assuming the public security) repeats the above comparison process with the data system of the bank. This time, the public key generated by the third party (i.e. the data system of the communication operator) is used for encryption. After calculating the result of subtraction of the two target Gini coefficients, the result of subtraction is sent to the data system of the communication operator for decryption and publication of the comparison result, and finally the smallest Gini coefficient in the three-dimensional data system is obtained.

[0124] Step 20244, the target classification label corresponding to the third target Gini coefficient is used as the optimal split mode of the root node, and the classification condition of the root node is determined according to the optimal split mode.

[0125] Then, the target classification label corresponding to the third target Gini coefficient can be used as the optimal split mode of the root node, and the classification condition of the root node is determined according to the optimal split mode.

[0126] Step 20245, the optimal split mode and classification condition of the nodes corresponding to different branches under the root node are determined until the leaf nodes under each branch are determined, and a constructed decision tree prediction model is obtained.

[0127] Afterwards, for each node under each branch of the root node, the global minimum Gini coefficient of the node can be recalculated according to the target personal information of the target user in the data system of different dimensions after the split, and the corresponding classification label is obtained according to the global minimum Gini coefficient to split the current node, and the classification condition of each node is determined, until the leaf node under each branch is determined, and a constructed decision tree prediction model is obtained. The specific process can refer to the above process or the process of constructing a decision tree prediction model according to the Gini coefficient in the prior art, which will not be described in detail in this embodiment.

[0128] In step 203, the node information of the node is stored in the data system of the corresponding dimension according to the source of the target personal information used by each node when splitting.

[0129] In this embodiment, after the optimal splitting mode of each node is calculated, the node information of the node is also recorded on the corresponding node. The node information includes the classification condition of the node and the next reachable node. For example, the node information includes eight parameters {ind, dept, label, val, gi, p, l, s}, which indicates that the splitting of the ind-th node of the decision tree prediction model is performed on the data system dept, the classification condition is the two-splitting of the data system dept on the classification label label according to whether the data value is greater than a certain threshold val, the Gini coefficient obtained by the splitting is gi, and in addition, the number p of the parent node of the current node, the number l of the left branch child node and the number s of the right branch child node are also recorded. Afterwards, the node information of the node can be stored in the corresponding data system according to the source of the target personal information used by each node when splitting, so that the corresponding node information can be obtained from different data systems in the subsequent process of identifying the target user. For example, the classification label used by the optimal splitting of a certain node is the number of waste cards, and the target personal information used by the node when splitting is obtained from the data system of the bank, so the node information of the node is stored in the data system of the bank. The overall process is as shown in Figure 5 .

[0130] As described above, in the process of constructing the decision tree prediction model, when the data exchange is performed in the data systems of different dimensions, the data is encrypted by the homomorphic encryption method before the data exchange, so that the data systems of different dimensions can still complete the related calculation work under the premise that the data cannot be decrypted, and at the same time, the sensitive data will not be leaked, and the safety of the sensitive data is ensured.

[0131] As shown in Figure 6 , a structure diagram of a target user identification device provided by the embodiment of the present application is shown in Figure 6 .

[0132] The personal information acquisition module 301 is configured to acquire first personal information of different dimensions of a user to be identified.

[0133] The user identification module 302 is configured to input the first personal information into a pre-constructed decision tree prediction model, so that the decision tree prediction model extracts first data corresponding to each node from the first personal information, and determines an accessible path according to the first data and classification conditions of the nodes.

[0134] The category determination module 303 is configured to determine whether the user is a target category user according to a leaf node on the accessible path.

[0135] On the basis of the above-mentioned embodiments, the information alignment module, the information segmentation module and the information storage module are further included.

[0136] The information alignment module is configured to align the personal information in the data systems of different dimensions, and determine target users common in the data systems of different dimensions.

[0137] The information segmentation module is configured to determine target personal information corresponding to the target users in the data systems of different dimensions, calculate a Gini coefficient according to the target personal information, segment the target personal information according to the Gini coefficient, and obtain each node of the decision tree prediction model and node information of each node.

[0138] The information storage module is configured to store the node information of each node in the data system of a corresponding dimension according to a source of target personal information used by each node when being segmented.

[0139] On the basis of the above-mentioned embodiments, the information alignment module is specifically configured to align the personal information in the data systems of different dimensions based on RSA algorithm, and determine target users common in the data systems of different dimensions.

[0140] On the basis of the above-mentioned embodiments, the information alignment module includes a hash encryption submodule, a system determination submodule, an information encryption submodule, a first decryption submodule, a second decryption submodule and a target user determination submodule.

[0141] The hash encryption submodule is configured to acquire identity information of each user from the personal information in the data systems of different dimensions, and perform hash calculation on the identity information to obtain hash identity information.

[0142] The system determination sub-module is configured to determine a target dimensional data system having a first private key in the different dimensional data systems;

[0143] The information encryption sub-module is configured to multiply the hash identity information in each of the remaining data systems other than the target dimensional data system by a random number encrypted using a first public key to obtain first encrypted personal information, and send the first encrypted personal information to the target dimensional data system;

[0144] The first decryption sub-module is configured to decrypt the first encrypted personal information using the first private key in the target dimensional data system to obtain first decrypted data, and send the first decrypted data to a corresponding dimensional data system to divide the first decrypted data by the random number in the corresponding dimensional data system to obtain first signature data;

[0145] The second decryption sub-module is configured to decrypt the hash identity information in the target dimensional data system using the first private key to obtain second signature data;

[0146] The target user determination sub-module is configured to align the identity information according to each of the first signature data and the second signature data, and determine a target user common to the different dimensional data systems.

[0147] In the above embodiment, the target user determination sub-module is specifically configured to calculate a first hash value of each of the first signature data and a second hash value of the second signature data, and compare each of the first hash value and the second hash value to determine the target user common to the different dimensional data systems.

[0148] In the above embodiment, the user identification module includes a classification information acquisition sub-module, a classification information encryption sub-module, a classification label encryption sub-module, and an information segmentation sub-module.

[0149] The classification information acquisition sub-module is configured to acquire classification information of the target user from first target personal information of a first dimensional data system, and the classification information includes information about whether the different users are target category users.

[0150] The classification information encryption sub-module is configured to encrypt the classification information to obtain encrypted classification information, and distribute the encrypted classification information to the remaining second dimensional data systems.

[0151] The classification label encryption submodule is configured to obtain a second classification label from second target personal information in the second dimension data system, encrypt the second classification label to obtain an encrypted classification label, and each of the second target personal information in the second dimension data system includes a plurality of the second classification labels.

[0152] The information segmentation submodule is configured to calculate a Gini coefficient according to the classification information in the first dimension data system, calculate a Gini coefficient according to the encrypted classification information and the encrypted classification label in the second dimension data system, and segment the target personal information according to the Gini coefficient to obtain each node of the decision tree prediction model.

[0153] Based on the above embodiment, the information segmentation submodule is specifically configured to calculate a first Gini coefficient corresponding to different first classification labels in the first target personal information according to the first classification label and the classification information in the first dimension data system, and the first target personal information further includes the first classification label; calculate a second Gini coefficient corresponding to each of the encrypted classification labels in the second target personal information according to the encrypted classification label and the encrypted classification information in each of the second dimension data system; select a first target Gini coefficient with the smallest value from the first Gini coefficients, and select a second target Gini coefficient with the smallest value from the second Gini coefficients; compare the first target Gini coefficient with a plurality of the second target Gini coefficients to determine a third target Gini coefficient with the smallest value; take a target classification label corresponding to the third target Gini coefficient as an optimal segmentation manner of a root node, determine a classification condition of the root node according to the optimal segmentation manner, determine optimal segmentation manners and classification conditions of nodes corresponding to different branches under the root node until leaf nodes under each branch are determined, and obtain a constructed decision tree prediction model.

[0154] The embodiment also provides a terminal device, as shown in a terminal device 40, the terminal device includes a processor 400 and a memory 401; Figure 7 The memory 401 is configured to store a computer program 402 and transmit the computer program 402 to the processor;

[0155] The memory 401 is configured to store a computer program 402 and transmit the computer program 402 to the processor;

[0156] The processor 400 is configured to execute the steps in the above-mentioned target user identification method embodiment according to instructions in the computer program 402.

[0157] For example, the computer program 402 can be divided into one or more modules / units, which are stored in the memory 401 and executed by the processor 400 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 402 in the terminal device 40.

[0158] The terminal device 40 can be a desktop computer, a notebook computer, a palm computer, a cloud server and the like. The terminal device 40 can include, but is not limited to, the processor 400 and the memory 401. Those skilled in the art can understand that the terminal device 40 can include more or less components, or combine certain components, or different components, for example, the terminal device 40 can also include an input / output device, a network access device, a bus and the like. Figure 7 The terminal device 40 is only an example and does not constitute a limitation on the terminal device 40, and can include more or less components than the illustration, or combine certain components, or different components, for example, the terminal device 40 can also include an input / output device, a network access device, a bus and the like.

[0159] The processor 400 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0160] The memory 401 can be an internal storage unit of the terminal device 40, such as a hard disk or a memory of the terminal device 40. The memory 401 can also be an external storage device of the terminal device 40, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card and the like equipped on the terminal device 40. Further, the memory 401 can include both the internal storage unit and the external storage device of the terminal device 40. The memory 401 is used to store the computer program and other programs and data required by the terminal device 40. The memory 401 can also be used to temporarily store data that has been output or will be output.

[0161] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0162] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic, and the division of the units is merely a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0163] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.

[0164] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware, or in the form of software functional units.

[0165] When the integrated unit is realized in the form of software functional units and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application essentially or the part that makes a contribution to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various computer program storage media.

[0166] The embodiment of the present application also provides a storage medium comprising computer executable instructions, which are used for executing a target user identification method when executed by a computer processor, and the method comprises the following steps:

[0167] acquiring first personal information of different dimensions of a user to be identified;

[0168] inputting the first personal information into a pre-constructed decision tree prediction model, so that the decision tree prediction model extracts first data corresponding to each node from the first personal information, determines an accessible path according to the first data and classification conditions of the nodes; node information of each node is stored in a data system of a corresponding dimension, and the node information comprises a classification condition of the node and a next accessible node, and the node is obtained by cutting according to personal information of a user in different dimension data systems;

[0169] determining whether the user is a target category user according to a leaf node on the accessible path.

[0170] It should be noted that the above are only preferred embodiments of the present application and the technical principles applied. Those skilled in the art will understand that the embodiments of the present application are not limited to the specific embodiments described herein, and those skilled in the art can make various obvious changes, readjustments and substitutions without departing from the scope of the embodiments of the present application. Therefore, although the embodiments of the present application have been described in more detail through the above embodiments, the embodiments of the present application are not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the embodiments of the present application, and the scope of the embodiments of the present application is determined by the scope of the appended claims.

Claims

1. A method for identifying target users, characterized in that, Includes the following steps: Obtain primary personal information of the user to be identified from different dimensions; The first personal information is input into a pre-built decision tree prediction model, so that the decision tree prediction model extracts the first data corresponding to each node from the first personal information, and determines the reachable path based on the first data and the classification conditions of each node; The node information of each node is stored in the corresponding dimension of the data system. The node information includes the classification conditions of the node and the next reachable node. The node is obtained by segmenting the user's personal information in the data system of different dimensions. The segmentation of the node based on the user's personal information in the data system of different dimensions includes: calculating the Gini coefficient based on the target personal information, and selecting the segmentation method corresponding to the smallest Gini coefficient to segment the target personal information to obtain the node information of the decision tree prediction model. Based on the leaf nodes on the reachable path, determine whether the user is a target category user.

2. The target user identification method according to claim 1, characterized in that, The decision tree prediction model is pre-built, including: Align personal information in the data systems of different dimensions to identify the target users common to the data systems of different dimensions, wherein personal information of different users is pre-stored in the data systems of different dimensions; In the data system of different dimensions, target personal information corresponding to the target user is determined, the Gini coefficient is calculated based on the target personal information, and the target personal information is segmented based on the Gini coefficient to obtain each node of the decision tree prediction model and the node information of each node; Based on the source of the target personal information used by each node during the segmentation, the node information of the node is stored in the data system of the corresponding dimension.

3. The target user identification method according to claim 2, characterized in that, The process of aligning personal information in the data systems of different dimensions to determine the common target users in the data systems of different dimensions includes: Based on the RSA algorithm, personal information in the data systems of different dimensions is aligned to identify the common target users in the data systems of different dimensions.

4. The target user identification method according to claim 3, characterized in that, The method of aligning personal information in the data systems of different dimensions based on the RSA algorithm to determine the common target users in the data systems of different dimensions includes: In the personal information of the data system of different dimensions, the identity information of each user is obtained, and the identity information is hashed to obtain hash identity information; In the data systems of different dimensions, identify the target dimension data system that possesses the first private key; In the data systems other than the target dimension data system, the hash identity information in each data system is multiplied by a random number encrypted with the first public key to obtain the first encrypted personal information, and the first encrypted personal information is sent to the target dimension data system. The first encrypted personal information is decrypted using the first private key in the target dimension data system to obtain first decrypted data. The first decrypted data is then sent to the corresponding dimension data system, where the first decrypted data is divided by the random number to obtain first signature data. The hash identity information in the target dimension data system is decrypted using the first private key to obtain the second signature data; Based on each of the first signature data and the second signature data, the identity information is aligned to determine the target users common to the data systems of different dimensions.

5. The target user identification method according to claim 4, characterized in that, The step of aligning identity information based on the first signature data and the second signature data to determine the target user common to the different dimensions of the data system includes: Calculate the first hash value for each of the first signature data and the second hash value for each of the second signature data; By comparing each of the first hash value and the second hash value, the target users common to the data systems of different dimensions are determined.

6. The target user identification method according to claim 2, characterized in that, The step of calculating the Gini coefficient based on the target personal information, and segmenting the target personal information based on the Gini coefficient to obtain each node of the decision tree prediction model and the node information of each node includes: From the first target personal information of the data system of the specified first dimension, obtain the classification information of the target user, the classification information including whether the different users belong to the target category user; The classification information is encrypted to obtain encrypted classification information, which is then distributed to the remaining data systems of the second dimension. In the second dimension data system, a second category label is obtained from the second target personal information, and the second category label is encrypted to obtain an encrypted category label. Each second target personal information in the second dimension data system includes multiple second category labels. In the first-dimensional data system, the Gini coefficient is calculated based on the classification information. In the second-dimensional data system, the Gini coefficient is calculated based on the encrypted classification information and the encrypted classification label. The target personal information is segmented based on the Gini coefficient to obtain the nodes of the decision tree prediction model.

7. The target user identification method according to claim 6, characterized in that, The calculation of the Gini coefficient based on the classification information in the first-dimensional data system, and the calculation of the Gini coefficient based on the encrypted classification information and the encrypted classification label in the second-dimensional data system, followed by segmentation of the target personal information based on the Gini coefficient to obtain the nodes of the decision tree prediction model, including: In the data system of the first dimension, based on the first classification label and the classification information, the first Gini coefficient corresponding to different first classification labels in the first target personal information is calculated, and the first target personal information also includes the first classification label; In each of the second-dimensional data systems, a second Gini coefficient corresponding to each of the encrypted classification labels is calculated in the second target personal information based on the encrypted classification labels and the encrypted classification information; Select the first target Gini coefficient with the smallest value from the first Gini coefficient, and select the second target Gini coefficient with the smallest value from the second Gini coefficient; The first target Gini coefficient is compared with multiple second target Gini coefficients to determine the third target Gini coefficient with the smallest value; The target classification label corresponding to the third target Gini coefficient is taken as the optimal segmentation method of the root node, and the classification condition of the root node is determined according to the optimal segmentation method. The optimal splitting method and classification conditions for the nodes corresponding to different branches under the root node are determined until the leaf nodes under each branch are determined, thus obtaining the constructed decision tree prediction model.

8. A target user identification device, characterized in that, include: The personal information acquisition module is used to acquire different dimensions of the primary personal information of the user to be identified. The user identification module is used to input the first personal information into a pre-built decision tree prediction model, so that the decision tree prediction model extracts the first data corresponding to each node from the first personal information, and determines the reachable path based on the first data and the classification conditions of each node. The node information of each node is stored in the corresponding dimension of the data system. The node information includes the classification conditions of the node and the next reachable node. The node is obtained by segmenting the user's personal information in the data system of different dimensions. The user identification module is specifically used to: calculate the Gini coefficient based on the target personal information, select the segmentation method corresponding to the smallest Gini coefficient to segment the target personal information to obtain the node information of the decision tree prediction model. The category determination module is used to determine whether the user is a target category user based on the leaf nodes on the reachable path.

9. A terminal device, characterized in that, The terminal device includes a processor and a memory; The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to execute a target user identification method as described in any one of claims 1-7 according to instructions in the computer program.

10. A storage medium for storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform a target user identification method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Black product user recognition method, TEE node and computer readable storage medium

    CN113837303A

  • Method and device for constructing decision tree

    US20220230071A1