A data prediction method, device and readable storage medium
By de-identifying the target data and combining it with various data prediction techniques, the problems of low accuracy and information leakage risk of manual identity prediction have been solved, achieving more efficient and secure identity prediction and service optimization.
Patent Information
- Application Number
- CN202210454694.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-27
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2042-04-27
AI Technical Summary
In the existing technology, the prediction method of manually judging the identity information of the object has low accuracy and high cost, and there is a risk of identity information leakage.
Multiple data prediction techniques are used to process the de-identified data. The de-identification process enhances data security, and the combination of multiple data prediction techniques determines the identity category of the object, thereby improving the accuracy of the prediction.
By de-identifying information, the risk of identity leakage can be reduced, the accuracy and efficiency of data prediction can be improved, targeted services and information push can be achieved, and the user experience can be enhanced.
Smart Images

Figure CN117034181B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a data prediction method, device and readable storage medium. BACKGROUND
[0002] At present, there are many businesses that need to judge the identity of an object to realize personalized services and improve the accuracy of services. At present, the identity information of an object is generally predicted by a person according to experience. The accuracy of this prediction method is low, and the labor cost is high. In addition, this prediction method is prone to cause user identity information leakage in the prediction process, and has a high risk of identity information leakage. SUMMARY
[0003] The embodiments of the present application provide a data prediction method, device and readable storage medium, which can improve the accuracy of data prediction and reduce the risk of identity information leakage.
[0004] In a first aspect, the present application provides a data prediction method, comprising:
[0005] obtaining target data to be predicted, the target data being used to indicate an identity category of an object;
[0006] performing desensitization processing on the target data to obtain target desensitization data;
[0007] respectively processing the target desensitization data by using N kinds of data prediction technologies to obtain N identity prediction sets, each identity prediction set including a probability that the object belongs to M predicted identity categories, N being a positive integer greater than or equal to 2, and M being a positive integer;
[0008] determining a target identity category of the object based on the N identity prediction sets.
[0009] In a second aspect, the present application provides a data prediction device, comprising:
[0010] a data acquisition unit configured to acquire target data to be predicted, the target data being used to indicate an identity category of an object;
[0011] a data desensitization unit configured to perform desensitization processing on the target data to obtain target desensitization data;
[0012] a data processing unit configured to respectively process the target desensitization data by using N kinds of data prediction technologies to obtain N identity prediction sets, each identity prediction set including a probability that the object belongs to M predicted identity categories, N being a positive integer greater than or equal to 2, and M being a positive integer;
[0013] an identity determination unit configured to determine a target identity category of the object based on the N identity prediction sets.
[0014] In a third aspect, the present application provides a computer device, comprising: a processor, a memory, and a network interface;
[0015] The processor is connected to the memory and the network interface, wherein the network interface is configured to provide a data communication function, the memory is configured to store a computer program, and the processor is configured to invoke the computer program to enable the computer device comprising the processor to perform the data prediction method.
[0016] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program adapted to be loaded and executed by a processor to enable a computer device comprising the processor to perform the data prediction method.
[0017] In a fifth aspect, the present application provides a computer program product or a computer program, which comprises computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the data prediction method provided in various optional manners in the first aspect of the present application.
[0018] In the embodiments of the present application, the target data to be predicted is desensitized to improve data security and reduce the risk of object identity information leakage. Since the target desensitized data is processed using multiple data prediction technologies, each data prediction technology processes the target desensitized data in a different way, and thus the results obtained are different. By combining multiple data prediction technologies to determine the identity category of the object, identity prediction can be performed from multiple dimensions, and the accuracy of data prediction can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0020] Figure 1 is an architecture diagram of a data prediction system provided by the embodiments of the present application;
[0021] Figure 2 is an application scenario diagram of a data prediction method provided by the embodiments of the present application;
[0022] Figure 3is a flowchart of a data prediction method provided by an embodiment of the present application;
[0023] Figure 4 is a flowchart of a model training method provided by an embodiment of the present application;
[0024] Figure 5 is a schematic diagram of average pooling provided by an embodiment of the present application;
[0025] Figure 6 is a flowchart of a training model provided by an embodiment of the present application;
[0026] Figure 7 is a schematic diagram of processing desensitization data based on a model provided by an embodiment of the present application;
[0027] Figure 8 is another schematic diagram of processing desensitization data based on a model provided by an embodiment of the present application;
[0028] Figure 9 is still another schematic diagram of processing desensitization data based on a model provided by an embodiment of the present application;
[0029] Figure 10a is a model effect comparison schematic diagram provided by an embodiment of the present application;
[0030] Figure 10b is a business effect comparison schematic diagram provided by an embodiment of the present application;
[0031] Figure 11 is a flowchart of another data prediction method provided by an embodiment of the present application;
[0032] Figure 12 is a flowchart of still another data prediction method provided by an embodiment of the present application;
[0033] Figure 13 is a component structure schematic diagram of a data prediction device provided by an embodiment of the present application;
[0034] Figure 14 is a component structure schematic diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0035] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0036] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software level technology. Artificial intelligence basic technology generally includes, such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning and other fields.
[0037] Among them, machine learning (Machine Learning, ML) is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. Its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rule-based learning.
[0038] The data related to user information (for example, target data) involved in the embodiments of the present application are all data authorized by the user. The present application relates to machine learning technology in the field of artificial intelligence. Optionally, for example, the target desensitization data can be processed using machine learning technology to obtain N identity prediction sets, so as to determine the target identity category of the object (i.e. the user) based on the N identity prediction sets. Or, the target data can also be desensitized using machine learning technology to obtain target desensitized data, etc. The technical scheme of the present application can be applied to the scene of predicting the target data of the user and determining the identity category of the user. By determining the identity category of the object, targeted advertising can be implemented, and the click-through rate of the user can be improved. Or, by determining the identity category of the object, the financial loan capacity of the object can be predicted, so that targeted loans can be issued in the scene of buying a house, buying a car, etc., reducing the loss of the loan institution. The technical scheme of the present application can also be applied to other scenes that need to predict the identity category of the object, which is not limited by the present application. By desensitizing the target data of the object, the data security can be improved, and the risk of identity information leakage of the object can be reduced. By using various data prediction technologies to process the target desensitized data, the identity category of the object can be determined, which can predict the identity from multiple dimensions and improve the accuracy of data prediction.
[0039] See Figure 1 , Figure 1 is a schematic diagram of the architecture of a data prediction system provided by the embodiments of the present application, such as Figure 1As shown, the computer device can interact with the terminal device, and the number of terminal devices can be one or at least two. For example, when the number of terminal devices is multiple, the terminal devices can include Figure 1 For example, the computer device 102 can obtain target data to be predicted. Further, the computer device 102 can perform desensitization processing on the target data to obtain target desensitization data. Further, the computer device 102 can process the target desensitization data by using N kinds of data prediction technologies respectively to obtain N identity prediction sets, and determine a target identity category of the object based on the N identity prediction sets. Optionally, the computer device 102 can send the target identity category of the object to the terminal device 101a, so that the terminal device 101a performs corresponding business processing based on the target identity category of the object. By performing desensitization processing on the target data of the object, the data security can be improved, and the risk of identity information leakage of the object can be reduced. By using multiple data prediction technologies to process the target desensitization data and determine the identity category of the object, identity prediction can be performed from multiple dimensions, and the accuracy of data prediction can be improved.
[0040] It can be understood that the computer device mentioned in the embodiments of the present application includes but is not limited to a terminal device or a server. In other words, the computer device can be a server or a terminal device, or a system composed of a server and a terminal device. The terminal device mentioned above can be an electronic device, including but not limited to a mobile phone, a tablet computer, a desktop computer, a notebook computer, a palm computer, a vehicle-mounted device, a smart voice interaction device, an augmented reality / virtual reality (AR / VR) device, a head-mounted display, a wearable device, a smart speaker, a smart home appliance, a flying device, a digital camera, a camera, and other mobile internet devices (MID) with network access capability. The server mentioned above can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, vehicle-road cooperation, content delivery networks (CDN), and big data and artificial intelligence platforms, etc. basic cloud computing services.
[0041] Further, please refer to Figure 2 , Figure 2 is an application scenario diagram of a data prediction method provided by the embodiments of the present application. As Figure 2As shown, the computer device 20 can obtain target data 21 to be predicted, for example, the target data 21 includes "name: Zhang San, interest: sports, permanent residence: XX city". Further, the computer device 20 can perform desensitization processing on the target data 21 to obtain target desensitization data 22, for example, the target desensitization data 22 is "name: Zhang*, interest: **, permanent residence: **". Further, the computer device 20 can process the target desensitization data 22 by using N kinds of data prediction techniques respectively to obtain N identity prediction sets. For example, N is equal to 3, that is, the target desensitization data 22 is processed by using 3 target model corresponding data prediction techniques to obtain 3 identity prediction sets. Each identity prediction set includes the probability that the object belongs to M (for example, 4) predicted identity categories, for example, the first identity prediction set includes "first identity category 0.3, second identity category 0.6, third identity category 0.55, fourth identity category 0.35"; the second identity prediction set includes "first identity category 0.28, second identity category 0.58, third identity category 0.62, fourth identity category 0.38"; and the third identity prediction set includes "first identity category 0.4, second identity category 0.68, third identity category 0.5, fourth identity category 0.4", so as to determine the target identity category of the object based on the three identity prediction sets, for example, the target identity category of the object is the second identity category. Wherein, the first identity category, the second identity category, the third identity category and the fourth identity category are four different identity categories.
[0042] Further, please refer to Figure 3 , Figure 3 is a flow diagram of a data prediction method provided by an embodiment of the present application; as Figure 3 shown, the data prediction method can be applied to a computer device, and the data prediction method includes but is not limited to the following steps:
[0043] S101, obtaining target data to be predicted.
[0044] In the embodiments of the present application, the computer device can obtain the target data to be predicted from the terminal device, or obtain the target data to be predicted from the local storage, or obtain the target data to be predicted from a third-party terminal, and the embodiments of the present application do not limit this. The target data can be used to indicate the identity category of the object, for example, the target data can include the identity information of the object, for example, the target data can include the interest and hobby of the object, the permanent residence, the zip code, the work, the type of the application installed on the terminal of the object, the application use time period, the traffic use condition corresponding to the application, and the like. Since the work state of the object is different, the type of the application installed on the terminal held by the object is also different, for example, the work state of the object with the learning application installed on the terminal can be a student, and the work state of the object with the Tencent meeting installed on the terminal can include different categories such as the first identity category, the second identity category, and the third identity category. Alternatively, the first identity category, the second identity category, and the third identity category can be determined according to the position of the user. Therefore, by combining the target data of this type to judge the target identity category of the object, for example, the work state of the object, the identity category of the object can be determined, and then targeted services can be provided based on the target identity type of the object, the user experience is improved, and resources are saved.
[0045] Alternatively, when the computer device detects the starting instruction for the target application, the target data to be predicted is obtained. The target application can refer to a pre-set application, or can also refer to an application with certain functions. For example, the target application can refer to an application with an information pushing function, or the target application can also be associated with a shopping application, and when the purchase instruction for the target product in the target application is detected, the shopping application associated with the target application can be jumped to, so as to facilitate the user to quickly purchase the target product and improve the user experience.
[0046] Alternatively, the computer device can pre-obtain the target data to be predicted, store the target identity category of the object determined based on the target data to be predicted, and when the starting instruction for the target application is detected, the target identity category of the object can be quickly determined, so as to quickly push the information based on the target identity category, improve the data pushing efficiency, and improve the user experience. Further, the computer device can obtain the target data to be predicted every target time period, and determine the target identity category of the object based on the obtained target data, so as to update the stored target identity category of the object. Since the identity category of the object will be updated, by obtaining the target data of the object every target time period to determine the target identity category of the object, the identity category of the object can be updated, so as to adjust the content of the information pushing, improve the accuracy of the information pushing, and thus improve the click rate of the user and increase the user experience.
[0047] Optionally, the computer device can further acquire association data of the association object, determine the identity category of the association object based on the association data, and push information to the object based on the identity category of the association object. The association object can be a user having an association relationship with the object, for example, can include but is not limited to a friend, a parent, a child, a spouse, etc. The acquisition of the association data needs to obtain authorization of the user. The association data can be used to indicate the identity category of the association object. By determining the identity category of the association object, information related to the identity category of the association object can be pushed to the object, thereby improving the click rate of the user.
[0048] Optionally, the computer device can further process the target data, for example, perform feature processing on the target data. Optionally, the computer device can construct portrait features for the object, which can include but are not limited to user basic attributes, device basic attributes, network connection attributes, etc. Further, the computer device can construct business vertical type features based on business characteristics. The vertical type features can include the click rate and conversion rate of the user on a specific type of advertisement. Further, the computer device can aggregate portrait features and business features of different time spans in combination with the time dimension. In subsequent processing of the target data, the features of the user can be spliced to obtain spliced features, and then the spliced features are desensitized and predicted. By performing feature processing on the target data, the target feature vector of the object can be combined with the time dimension, and the features of the object can be more complete.
[0049] S102, desensitizing the target data to obtain target desensitized data.
[0050] In the embodiments of the present application, since the target data is used to predict the identity category of the object, there is a risk of leakage of the identity information of the object. Therefore, the target data is desensitized before subsequent processing, which can reduce the risk of data leakage and improve the security of the data.
[0051] Optionally, the computer device can perform desensitization processing on the target data based on preset desensitization rules to obtain target sensitive data. Specifically, the computer device can obtain key characters in the target data and replace the key characters with preset characters; and determine the target data with the replaced preset characters as target desensitized data. The key characters can be determined according to the type of the target data. For example, if the type of the target data is a name type, the key characters can be characters other than the family name, i.e., the key characters are the given name, for example, the target data is Zhang San, and the key characters are 'San'. If the type of the target data is an address type, the key characters can include characters of districts or streets, for example, the target data is 'Shenzhen Nanshan District XX Street XX Science and Technology Park', and the target characters can be 'Nanshan District XX Street XX Science and Technology Park', 'XX Street XX Science and Technology Park', 'XX Science and Technology Park', etc. The preset characters are used to replace the key characters, and the preset characters can include '*', '#', '?' and the like. By replacing the key characters with the preset characters, the target desensitized data can be obtained. For example, the target data is of the address type and is 'Shenzhen Nanshan District XX Street XX Science and Technology Park', the key characters in the target data are 'XX Street XX Science and Technology Park', and the target desensitized data obtained by replacing the key characters with the preset characters is 'Shenzhen Nanshan District ***********', etc.
[0052] Optionally, the computer device can perform desensitization processing on the target data based on preset desensitization rules to obtain target sensitive data. Specifically, the computer device can obtain key characters in the target data and replace the key characters with preset characters; and determine the target data with the replaced preset characters as target desensitized data. The key characters can be determined according to the type of the target data. For example, if the type of the target data is a name type, the key characters can be characters other than the family name, i.e., the key characters are the given name, for example, the target data is Zhang San, and the key characters are 'San'. If the type of the target data is an address type, the key characters can include characters of districts or streets, for example, the target data is 'Shenzhen Nanshan District XX Street XX Science and Technology Park', and the target characters can be 'Nanshan District XX Street XX Science and Technology Park', 'XX Street XX Science and Technology Park', 'XX Science and Technology Park', etc. The preset characters are used to replace the key characters, and the preset characters can include '*', '#', '?' and the like. By replacing the key characters with the preset characters, the target desensitized data can be obtained. For example, the target data is of the address type and is 'Shenzhen Nanshan District XX Street XX Science and Technology Park', the key characters in the target data are 'XX Street XX Science and Technology Park', and the target desensitized data obtained by replacing the key characters with the preset characters is 'Shenzhen Nanshan District ***********', etc.
[0053] S103, respectively using N data prediction technologies to process the target desensitized data to obtain N identity prediction sets.
[0054] In the embodiments of the present application, the computer device can use multiple data prediction technologies to process the target de-sensitized data respectively, and obtain multiple identity prediction sets. Each identity prediction set includes the probability that the object belongs to M predicted identity categories, N is a positive integer greater than or equal to 2, and M is a positive integer. Assuming that N is equal to 3 and M is equal to 4, the target de-sensitized data is processed by using N data prediction technologies respectively, and three identity prediction sets are obtained, for example, the first identity prediction set includes “first identity category 0.3, second identity category 0.6, third identity category 0.55, and fourth identity category 0.35”; the second identity prediction set includes “first identity category 0.28, second identity category 0.62, third identity category 0.58, and fourth identity category 0.38”; and the third identity prediction set includes “first identity category 0.4, second identity category 0.68, third identity category 0.5, and fourth identity category 0.4”. Each identity prediction set contains M predicted identity categories, that is, the probability that the object belongs to M predicted identity categories can be determined by using each data prediction technology to process the target de-sensitized data, thereby obtaining N identity prediction sets.
[0055] Alternatively, the target de-sensitized data includes one of dense features, sparse features, and sparse features and dense features, so the computer device can use N data prediction technologies to process the dense features and / or sparse features respectively, and obtain N identity prediction sets. The dense features can refer to the number of non-zero values in the features being greater than a threshold, that is, most of the features in the dense features correspond to non-zero values, and only a small part of the features correspond to zero. The sparse features can refer to the number of non-zero values in the features being less than or equal to a threshold, that is, most of the features in the sparse features correspond to zero, and only a small part of the features correspond to non-zero values.
[0056] Alternatively, for example, the N identity prediction sets include at least two of the first identity prediction set, the second identity prediction set, and the third identity prediction set; since the target de-sensitized data includes sparse features and dense features, the computer device can determine the N identity prediction sets in combination with at least two of the following modes:
[0057] The first mode is that the computer device can perform feature compression on the sparse features, perform feature splicing on the compressed features and the dense features, obtain first spliced features, and determine the first identity prediction set based on the first spliced features.
[0058] The second mode is that the computer device can perform feature compression on the sparse features and the dense features, perform feature splicing on the compressed features, obtain second spliced features, and determine the second identity prediction set based on the second spliced features.
[0059] In the third mode, the computer device can perform feature compression on the sparse features, perform feature splicing on the compressed features and the dense features to obtain third spliced features, perform weight processing on the third spliced features based on an attention mechanism to determine a third identity prediction set.
[0060] In the first mode, since the target data contains sparse features, that is, most feature values in the target data are 0 and a small amount of feature values are non-0, the prediction effect is poor when directly using the sparse features for prediction. Therefore, the sparse features can be compressed to compress the high-dimensional sparse features into low-dimensional dense features for prediction, which can improve the prediction effect. In the second mode, by directly compressing the sparse features and the dense features, the data processing efficiency can be improved. In the third mode, since the sparse features are compressed, the prediction effect is better, and the features are processed by weight using the attention mechanism, which allows each feature to be trained with other features. According to the weight of the feature, the importance of each feature and other features is determined, and the weight of the more important combination is higher, which ultimately improves the prediction effect.
[0061] Optionally, the computer device can also determine the identity prediction set based on any one of the above three modes. Thus, the target identity category of the object is determined based on the identity prediction set. For example, the identity prediction set includes the probability that the object belongs to M predicted identity categories, and the computer device can obtain the predicted identity category corresponding to the maximum probability, and determine the predicted identity category corresponding to the maximum probability as the target identity category of the object. Alternatively, if there are multiple probabilities greater than a target threshold in the probabilities of the M predicted identity categories, the predicted identity categories corresponding to the multiple probabilities can also be determined as the target identity category of the object. By outputting the target identity category of the object and the corresponding probability, subsequent further judgment of the target identity category of the object based on artificial can be performed, thereby improving the identity prediction accuracy.
[0062] Optionally, the computer device can use N target model corresponding data prediction technology to process the target desensitization data, and obtain N identity prediction sets. Wherein, the N target models can include but are not limited to Logistic regression model, support vector machine (support vector machines, SVM), convolutional neural network (Convolutional Neural Network, CNN), long short term memory network (Long Short Term Memroy, LSTM), deep model (Deep&Cross, DCN), probability neural network (Product-based Neural Network, PNN), deep recommendation model (Automatic Feature Interaction Learning via Self-Attentive Neural Networks, AutoInt), etc.
[0063] Optionally, the computer device can perform feature compression on the sparse features based on the Deep&Cross model, perform feature splicing on the compressed features and the dense features to obtain first splicing features, and determine the first identity prediction set based on the first splicing features. Optionally, the computer device can perform feature compression on the sparse features and the dense features using the PNN model, perform feature splicing on the compressed features to obtain second splicing features, and determine the second identity prediction set based on the second splicing features. Optionally, the computer device can perform feature compression on the sparse features using the AutoInt model, perform feature splicing on the compressed features and the dense features to obtain third splicing features, and perform weight processing on the third splicing features based on an attention mechanism to determine the third identity prediction set.
[0064] In the embodiments of the present application, because the model structures of the N target models are different, the ways of processing features by each target model are different, and therefore the identity category results of the objects obtained by processing the target desensitization data based on each target model can be different, which can more completely reflect the identity category of the objects. By processing the target desensitization data by combining multiple target models, the target desensitization data can be predicted from multiple dimensions, thereby improving the accuracy of data processing. Optionally, before processing the target desensitization data using the N target models, the N target models can be pre-trained, and the N target models after training are saved. When the target data to be predicted is obtained in the subsequent desensitization processing of the target data to obtain the target desensitization data, the N target models can be used to process the target desensitization data. The process of training the N target models can be referred to in the description of the method in the corresponding embodiments. Figure 4 The method in the corresponding embodiments will not be described in detail here.
[0065] S104, determining a target identity category of the object based on the N identity prediction sets.
[0066] In the embodiments of the present application, since the probability that the object belongs to the M predicted identity categories is determined based on each data prediction technology, the N identity prediction sets are obtained, and thus the target identity category of the object can be determined based on the N identity prediction sets. The object can refer to a user who needs to be predicted for identity. Alternatively, if the target identity category is a user work category, the user work category is used to indicate the position of the working population in the work unit, i.e., the identity category of the user can be determined based on the position of the user. The target identity category of the object can include but is not limited to a first identity category, a second identity category, a third identity category, a fourth identity category, and the like. Alternatively, the identity category of the user can be pre-set or pre-divided. Alternatively, the target identity category can include a user lifestyle category, and the target identity category of the object can include but is not limited to a first lifestyle category, a second lifestyle category, a third lifestyle category, and the like. The user lifestyle category can be used to reflect the situation of the user in work or life, and by determining the lifestyle category of the user, targeted services such as targeted advertisement pushing can be realized, resources can be saved, and user experience can be improved.
[0067] Alternatively, the identity category of the user can also be determined based on the work category of the user, for example, the identity category of the user corresponding to the first work category is the first identity category, the identity category of the user corresponding to the second work category is the second identity category, the identity category of the user corresponding to the third work category is the third identity category, and the like. The work category of the user can include but is not limited to social services, culture and education, scientific research, art and creation, calculation and mathematics, and the like. Since there are differences in the fields corresponding to each work category, the types of information that the users of each work category pay attention to are different. By determining the work category of the user, targeted information recommendation can be realized, user experience can be improved, and user click rate can be improved.
[0068] Alternatively, the computer device can determine the predicted identity categories belonging to the same category from each identity prediction set, and determine the target identity category of the object based on the predicted identity categories belonging to the same category. Specifically, the computer device can determine the predicted identity categories belonging to the same category from the N identity prediction sets, and the probabilities of the predicted identity categories belonging to the same category; determine the probability of each predicted identity category in the M predicted identity categories based on the probabilities of the predicted identity categories belonging to the same category in the N identity prediction sets, to obtain the total probability of each predicted identity category; determine the maximum probability from the total probability of the M predicted identity categories; and determine the predicted identity category corresponding to the maximum probability as the target identity category of the object.
[0069] Optionally, the computer device can determine the probability of each predicted identity category in M predicted identity categories based on the average probability of the same type of predicted identity categories in N identity prediction sets, obtain the total probability of each predicted identity category, and thus determine the target identity category of the object.
[0070] For example, N is 3, and the N identity prediction sets include a first identity prediction set, a second identity prediction set, and a third identity prediction set. For example, the first identity prediction set is "first identity category 0.3, second identity category 0.6, third identity category 0.55, fourth identity category 0.35"; the second identity prediction set is "first identity category 0.28, second identity category 0.58, third identity category 0.62, fourth identity category 0.38"; and the third identity prediction set is "first identity category 0.4, second identity category 0.68, third identity category 0.5, fourth identity category 0.4." By counting the probabilities of the same type of predicted identity categories, the computer device can determine that the total probability of the predicted identity category being the "first identity category" is (0.3+0.28+0.4) / 3=0.327; the total probability of the predicted identity category being the "second identity category" is (0.6+0.58+0.68) / 3=0.62; the total probability of the predicted identity category being the "third identity category" is (0.55+0.62+0.5) / 3=0.56; the total probability of the predicted identity category being the "fourth identity category" is (0.35+0.38+0.4) / 3=0.377. The computer device can then determine the maximum probability, i.e., 0.62, from the probabilities of the four predicted identity categories, and determine the predicted identity category "second identity category" corresponding to the maximum probability as the target identity category of the object.
[0071] Optionally, if there are multiple total probabilities in the total probabilities of each predicted identity category in the M predicted identity categories that are greater than the target threshold, the computer device can determine the predicted identity category corresponding to the multiple total probabilities as the target identity category of the object, and then manually judge the multiple target identity categories to determine which target identity category the object belongs to.
[0072] Optionally, if the probability of the target predicted identity category existing in each identity prediction set is greater than the probability of other predicted identity categories in the M predicted identity categories, and the target predicted identity category in each identity prediction set is the same, then the target predicted identity category is determined as the target identity category of the object.
[0073] For example, the N identity prediction sets include a first identity prediction set, a second identity prediction set, and a third identity prediction set; the first identity prediction set, the second identity prediction set, and the third identity prediction set each include a first identity category, a second identity category, and a third identity category. And the probability of the second identity category in the first identity prediction set is greater than the probabilities of other predicted identity categories in the first identity prediction set; the probability of the second identity category in the second identity prediction set is greater than the probabilities of other predicted identity categories in the second identity prediction set; and the probability of the second identity category in the third identity prediction set is greater than the probabilities of other predicted identity categories in the third identity prediction set. In other words, the probability of the second identity category is the greatest in the three identity prediction sets, and the second identity category is the target predicted identity category, so the second identity category can be directly determined as the target identity category of the object.
[0074] Since the probability of the second identity category is the greatest probability in each identity prediction set, even if the average of the total probabilities of each predicted category in the N identity prediction sets is calculated, the average of the total probability of the second identity category is greater than the average of the total probabilities of other predicted categories. Therefore, by directly determining the second identity category as the target identity category of the object, without calculating and determining the total probabilities of other predicted identity categories, the data processing efficiency can be saved.
[0075] In the embodiments of the present application, by performing desensitization processing on the target data to be predicted, the data security can be improved, and the risk of object identity information leakage can be reduced. Since the target desensitized data is processed by using multiple data prediction technologies, each data prediction technology processes the target desensitized data in a different way, and therefore the results obtained by processing are different. By combining multiple data prediction technologies to determine the identity category of the object, identity prediction can be performed from multiple dimensions, and the accuracy of data prediction can be improved.
[0076] Optionally, please refer to Figure 4 , Figure 4 is a flowchart of a model training method provided by an embodiment of the present application. The model training method can be applied to a computer device; as shown in Figure 4 , the model training method includes but is not limited to the following steps:
[0077] S201, obtaining sample data to be predicted.
[0078] In the embodiments of the present application, the computer device can obtain the sample data to be predicted from the terminal device, can obtain the sample data to be predicted from the local storage, or can obtain the sample data to be predicted from a third-party terminal, and the embodiments of the present application do not limit this. The sample data is used to indicate the sample identity category of the sample object. For example, the sample data can include the interest and hobby of the object, the permanent residence, the zip code, the work, the type of the application installed on the terminal of the object, the application use time period, the traffic use condition corresponding to the application, and the like.
[0079] Optionally, the computer device can obtain at least one initial sample data, filter the at least one initial sample data based on an abnormal rule to obtain sample filtered data, the abnormal rule including at least one abnormal behavior, filter the sample filtered data based on a distribution abnormal theorem to obtain the sample data to be predicted, and the distribution abnormal theorem is used to filter the data based on probability theory. That is, after obtaining the at least one initial sample data, it can be detected whether the initial sample data includes at least one abnormal behavior, if the initial sample data includes the abnormal behavior, the initial sample data can be filtered. For example, if it is detected that the frequency of starting an application by the object in the sleep time period is greater than a preset number of times, the initial sample data corresponding to the object is filtered. Alternatively, if it is detected that the time length of operating an application by the object exceeds 24 hours, the initial sample data corresponding to the object is filtered.
[0080] Further, filtering the sample filtered data based on the distribution abnormal theorem can be to use the Léonard criterion to judge the abnormal value, so as to filter the sample filtered data. The Léonard criterion refers to first assuming that a group of detection data only contains random errors, performing calculation and processing to obtain a standard deviation, determining an interval according to a certain probability, considering that the errors exceeding the interval are not random errors but gross errors, and the data containing the errors should be removed.
[0081] In a specific implementation, the computer device can obtain seed users with labels based on artificial labeling and business logic, i.e., seed users with identity category labels. The seed users can be explicitly labeled positive and negative samples. For example, a batch of seed users can be roughly recalled, and then filtered based on artificial screening, such as screening out users that obviously do not meet the rules, such as users who use a learning application to attend class during class time and are labeled as the first identity category. The users of the first identity category can be employees or other identity categories except students. Further, the computer device can verify the seed users based on business logic, such as the first identity category not playing game applications for too much time, etc. For example, a user labeled as the first identity category is obtained, and frequently plays game applications during working hours, which indicates that the seed user is abnormal and can be filtered, etc. Further, a filtered seed user basic portrait can be obtained. The basic portrait can include some behavior data of the user in some applications, such as whether the user's terminal installs a mobile phone manager, whether the user uses the mobile phone manager to disturb and block the function, answers the assistant function, etc., so as to further filter the seed user. Further, an abnormal type index of the seed user can be calculated, i.e., an abnormal user type evaluation index. In a real business scenario, there can be false users and computer-controlled terminal devices. In order to eliminate the influence of non-real users on modeling analysis, an abnormal type index can be set based on business experience, such as obvious abnormalities in the user's traffic usage in a certain category of applications, the time distribution of traffic generation, etc., such as continuously detecting that the user continuously operates the terminal device during the sleep time period within a week. Further, the sample filtered data is filtered based on the distribution anomaly theorem. By filtering the abnormal seed users, normal seed users are obtained, and the normal seed users can be stored in a distributed file system (The Hadoop Distributed File System, HDFS) for quick access in subsequent processes.
[0082] Optionally, the computer device can perform feature processing on the stored seed users, for example, offline feature processing. Specifically, the computer device can construct portrait features, for example, rich portrait features can be constructed based on user historical behavior data, and the portrait features can include but are not limited to: user basic attributes, device basic attributes, network connection attributes, etc. For example: user basic attributes can reflect user identity information, device basic attributes (mobile phone brand: Huawei), network connection attributes (Wi-Fi connection times this week is 10 times). Further, the computer device can construct business vertical type features based on business characteristics. Vertical type features can include user click rate and conversion rate for a specific type of advertisement. Further, the computer device can also aggregate portrait features and business features of different time spans in combination with time dimensions. For example, the computer device can calculate the aggregated portrait of the user in the past half year, the past 3 months, the past 1 month, and the past 1 week. The aggregation method can use any one or more of summation, median, and standard deviation. As shown in Figure 5 Figure 5 is a schematic diagram of average pooling provided by an embodiment of the present application, Figure 5 The numbers in the figure represent features. By performing average pooling (i.e., averaging) on features with a large amount of data, the features can be converted into aggregated features with a small amount of data. Figure 5 In Figure 5 , by performing average pooling on the values (1, 2, 3, 0) in the upper left corner of the four cells in the left 4*4 grid, the value (1.5) in the upper left corner of the right 2*2 grid can be converted, thereby converting 4*4 data into 2*2 data, which can reduce the amount of subsequent calculation.
[0083] Further, the features in the sample data can also be normalized and discretized. Discretization includes the following methods: One-Hot Encoding, for example, for user interest features, One-Hot Encoding becomes: (1, 0) for liking sports and (0, 1) for not liking sports. Count Encoding, for example, for user WiFi POI (point of interest) features, Count Encoding is used to identify the degree of interest of the user to the POI. For example, a user went to the "food-Chinese food-Cantonese food" POI 3 times this week. Consolidation Encoding, for example, multiple values under certain category variables can be summarized into the same information. For example, the Android system version feature includes "4.2", "4.4" and "5.0" three versions, which can be summarized as "low version Android system". Experiments show that the Consolidation Encoding processing method can bring greater positive benefits than directly one-hot "Android system version" feature. Finally, the computer device can merge the processed features and store them offline in the HDFS system for quick access by subsequent processes. For each user, the data input to the model can be a Y*1 numerical vector, Y is a positive integer, representing the dimension of the feature, such as (1, 0, 31, 4, 0.2, 9.3, 8.8, …, 0, 0, 1, 2, 34), the numerical values represent various features of the user, and a user has a Y-dimensional feature.
[0084] Optionally, the computer device can train multiple models using the sample data and select N target models therefrom. Specifically, the computer device can input the sample data to be predicted into K models for training, determine index parameters of the K models, the index parameters being used to reflect the classification performance of the models, K being a positive integer greater than N; and determine N target models from the K models based on the index parameters of the K models. When the N target models are determined, the sample desensitization data can be processed using data prediction techniques corresponding to the N target models respectively to obtain N sample prediction sets.
[0085] The index parameter of the model can be AUC (area under curve), which is a model evaluation index in the field of machine learning. The AUC is the area under the ROC curve. By calculating the AUC value of each model in the K models, the best N target models can be selected for parameter optimization. Parameter optimization refers to grid optimization of the hyperparameters of the selected model, so as to expect the evaluation index AUC of the model to be improved. The larger the AUC value, the more likely it is that the classification algorithm in the current model will place the positive samples in front of the negative samples, resulting in a better classification result.
[0086] Optionally, the computer device can randomly divide the sample set for feature processing, that is, divide the sample data into training sample data and test sample data. For example, the sample data can be divided according to the time window to which the sample data belongs, and the sample data of earlier time is taken as the training sample data and the sample data of later time is taken as the verification sample data, wherein the proportion of the training sample data and the verification sample data can be 5:1. Further, based on the default parameters, a plurality of models (such as K models) can be trained in parallel, and the better models (such as N target models) are selected from the plurality of models. The K models can include but are not limited to: Logistic regression model, SVM model, CNN model, LSTM model, Deep&Cross model, PNN model, AutoInt model, etc. Further, after retraining the model based on the parameters, the stability of the model effect is verified on multiple verification sample data, which is convenient for subsequent use.
[0087] S202, desensitizing the sample data to obtain sample desensitization data.
[0088] In the embodiments of the present application, since there is a risk of object identity information leakage when using sample data to predict the identity category of the object, the sample data is desensitized before subsequent step processing, which can reduce the risk of data leakage and improve data security.
[0089] Optionally, the computer device can desensitize the sample data in the following manner: obtaining the original sample sensitive probability of the sample data; adding noise to the sample data to obtain sample noise data, determining the target sample sensitive probability of the sample noise data based on the sample noise data and the original sample sensitive probability; if the target sample sensitive probability is in the target probability interval, the sample noise data is determined as the sample desensitization data.
[0090] Specifically, the computer device can perform K-anonymity processing on the sample data to obtain an equivalence group M1 after anonymity processing, which can be as shown in Table 1:
[0091] Table 1
[0092]
[0093] Furthermore, the computer device can extract the set of sensitive attributes S in the equivalence group M1, and extract and analyze the sample sensitivity probability α of each sensitive attribute in the set S. The set of sensitive attributes S may include sensitive data in the equivalence group M1, or may include all data in the equivalence group M1. That is to say, in the embodiment of the present application, only the sensitive data in the sample data may be processed, or the entire sample data may be processed, and the embodiment of the present application does not limit this. Furthermore, the computer device can enter the sample sensitivity probability α and add a small amount of noise to the data of the set S. The calculation formula is shown in formula (1-1):
[0094] α p =α+Lap<ΔS / ε> (1-1)
[0095] Among them, the given sample sensitive data set S = {S1, S2, ..., S n}, α p The value range belongs to the interval (0,1], α p is the target sample sensitive probability value of the required new sensitive attribute, that is, the sensitive probability value α of the sample data after desensitization, and its value range belongs to the interval (0,1], α is the original sample sensitive probability value, and Lap<ΔS / ε> is the trace random noise parameter.
[0096] Optionally, the computer device may determine the noise parameter in the following manner: The computer device may construct a vector group that obeys the Laplace distribution for the S set, as shown in Formula (1-2):
[0097] Δf=max||f(D1)-f(D2)||1 (1-2)
[0098] Where Δf represents the sensitivity of the function, D is the data set, and f: D→R d is a function, the max dataset has the largest difference, and ||·||1 represents the Manhattan distance. Laplace mechanism, given a dataset D, there is a function f: D→R d , if the random mechanism M satisfies the following formula (1-3), then K provides ε-differential privacy:
[0099] M(D)=f(D)+Lap(Δf / ε) (1-3)
[0100] Where Δf is the sensitivity, Lap(Δf / ε) is the random noise following the Laplace distribution, and the noise size depends on the sensitivity Δf and differential privacy ε.
[0101] Further, through the above formula (1-2), the sensitive parameter conforming to the alpha sensitive probability is generated, the sensitive parameter is substituted into formula (1-3), and the computer device can determine the random noise subject to Laplace distribution, that is, determine the noise. By substituting the determined noise into the above formula (1-1), the target sample sensitive probability of the data added with the noise can be calculated; if the target sample sensitive probability satisfies the target interval, for example, alpha p belongs to the interval (0, 1], the noise parameter is recorded in the equivalence group M1, and thus the equivalence group M2 is obtained, that is, the sample desensitization data is obtained.
[0102] Through the desensitization processing of the data, even if the illegal terminal device obtains the desensitization processed data, it is also difficult to decrypt the data, so that the original data cannot be obtained, and thus the risk of identity information leakage can be reduced and the data security can be improved.
[0103] In S203, the sample desensitization data is processed by using N kinds of data prediction technologies respectively, and N sample prediction sets are obtained.
[0104] In the embodiment of the application, the computer device can process the sample desensitization data by using the data prediction technology corresponding to the N target models respectively, and obtain N sample prediction sets. Each sample prediction set includes the probability that the sample object belongs to M kinds of prediction identity categories.
[0105] Optionally, the number of sample data to be predicted is multiple, and the computer device can divide the sample data into a training set and a test set, train the model using the training set, and test the model using the test set to improve the accuracy of model prediction. Specifically, the computer device can divide the multiple sample desensitization data to determine the training set and the test set, and the data amount of the training set is greater than that of the test set; the N target models are trained using the training set respectively, and N trained target models are obtained; the test set is processed based on the N trained target models respectively, and N test probabilities are obtained; and N sample prediction sets are determined based on the N test probabilities.
[0106] Optionally, the computer device can read the low-order feature matrix and the high-order feature matrix in the sample de-sensitization data, and splice the low-order feature matrix and the high-order feature matrix. The low-order feature matrix can include but is not limited to name, interest, and the like, and the high-order feature matrix can include but is not limited to time. Illustratively, the computer device can evenly divide the training set into 3 parts, use the "leave-one-out method" to train N target models, and use the trained N target models to predict the remaining one part of data and the test set. Then, the 3 parts of predicted training data are combined to obtain a new training data; the 3 parts of predicted test set are combined using the mean method to obtain a new test data. On the new training data and the test data, the output results of the N target models are obtained, and the final result, i.e., the sample identity category to which the sample data belongs, can be obtained by averaging and pooling the output results of the N target models.
[0107] As Figure 6 shown, Figure 6 is a flowchart of training a model provided by an embodiment of the present application, wherein the training set is divided into 4 parts to obtain 4 training sets and 4 validation sets, the first target model in the N target models is trained using the 4 training sets respectively, and the first target model after training is verified using the 4 validation sets respectively to obtain 4 probabilities x1 to x4, i.e., the probability of the sample category to which the sample object belongs, and S 1 train is obtained by splicing x1 to x4. Further, the first target model can be tested using the 4 test sets to obtain 4 probabilities c1 to c4, and S 1 test For other target models in the N target models, the training and testing can also be performed in the above-mentioned manner to obtain S train and S test corresponding to each target model. The output results of the N target models are obtained by using S train and S test , and the final result, i.e., the sample identity category to which the sample data belongs, can be obtained by averaging and pooling the output results of the N target models.
[0108] S204, determining the target sample identity category of the sample object based on the N sample prediction sets.
[0109] In the embodiment of the present application, the manner of determining the target sample identity category of the sample object based on the N sample prediction sets can refer to the manner of determining the target identity category of the object based on the N identity prediction sets in the corresponding step S104, which will not be described herein again. Figure 3
[0110] S205 , obtaining sample identity labels of the sample data, and training to obtain N target models based on the sample identity labels and target sample identity categories.
[0111] In an embodiment of the present application, a computer device obtains sample identity labels from sample data and trains N target models based on the sample identity labels and target sample identity categories. The sample identity labels refer to the actual identity categories of the sample data. When training the target models, the sample identity labels of the sample data can be predetermined, which is equivalent to knowing the actual identity categories of the sample data. By processing the sample data using the model, a model output is obtained, namely, the sample identity category to which the sample object belongs, where the sample identity category belongs to one or more of the M predicted identity categories. The purpose of training the model is to ensure that the sample identity category to which the sample object belongs and the sample identity label of the sample data are as consistent as possible. If the sample identity category and sample identity label corresponding to a predetermined number of sample data in a plurality of sample data are consistent, the current model can be saved for subsequent use. If the sample identity category and sample identity label corresponding to a predetermined number of sample data in a plurality of sample data are consistent, the model can be further trained, and the model parameters within the model can be adjusted so that after processing the sample data based on the model, the sample identity category and sample identity label output by the model are as consistent as possible.
[0112] Optionally, the N target models may include Deep&Cross models, such as Figure 7 As shown, Figure 7 This is a schematic diagram of processing desensitized data based on a model provided in an embodiment of the present application. Figure 7 The model includes an embedding and stacking layer, a cross network layer, a deep network layer, and a link output layer. By inputting the sample desensitized data into the embedding and stacking layer, the sparse features in the sample desensitized data can be compressed, and the compressed features can be spliced with the dense features to obtain the first spliced features. Furthermore, by inputting the first spliced features into the cross network layer and the deep network layer for processing respectively, the first cross feature and the first deep feature are obtained. Among them, based on the cross network layer, bounded prediction cross features can be effectively learned; the deep network layer can capture highly nonlinear interactions. Then, the first cross feature and the first deep feature are spliced and input into the link output layer, which can realize the merging of the first cross feature and the first deep feature, and the merged features are predicted to obtain the model prediction result, that is, the sample identity category to which the sample object belongs.
[0113] Wherein, since the core architecture of the Deep&Cross model includes two parts, discrete feature embedding and high-order cross feature. Among them, the discrete feature embedding method is inspired by the idea of Word2Vec. The original problem solved by Word2Vec is that the one-hot representation of words is too sparse, and the vector form representation between different words has no connection. Finally, an upper ten thousand-dimensional word one-hot representation is embedded into a few hundred-dimensional dense vector. The combination of high-order cross features often brings positive business effects, such as: “USA” and “Thanksgiving”, “China” and “Chinese New Year” such associated features. By designing this cross network, explicitly apply feature cross on each layer, effectively learn bounded prediction cross features, and do not need manual feature engineering or exhaustive search. Secondly, the cross network is simple and effective, through design, the highest polynomial degree of each layer is determined by the layer depth. The network is composed of all cross terms, and their coefficients are different. And, the cross network is memory efficient and easy to implement. In addition, the cross network has less than an order of magnitude of parameter quantity compared with DNN on LogLoss. Optionally, Relu function can be used as the activation function when training the Deep&Cross model, and Dropout is added.
[0114] Optionally, the N target models can include a PNN model, as shown in Figure 8 Figure 8 is another schematic diagram for processing desensitization data based on a model provided by the embodiments of the present application, Figure 8 The model in the above formula (1) includes an embedding layer, a product layer, a first hidden layer, and a second hidden layer. If the sample desensitization data contains N field characteristics and the one-hot vector is X, an embedding vector is generated for each field. By inputting the sample desensitization data into the embedding layer, the model can learn the embedding representation of each field from each field to obtain first embedding features. The first-order features and the second-order cross features of the first embedding features are spliced through the product layer to obtain second spliced features. The second spliced features are fully learned by the first hidden layer to obtain high-order combination features to obtain first hidden features. The first hidden features are fully learned by the second hidden layer to obtain high-order combination features to obtain prediction probability, and finally the model prediction result, i.e., the sample identity category to which the sample object belongs, is obtained.
[0115] Since the PNN model uses a pair-wise connected product layer to perform vector products between pairs of embedded vectors, the results are used as input for subsequent processing. The PNN model designs a product layer to combine features, including inner product and outer product operations, to increase the depth of feature combination crossover. For the inner product form of PNN, the result of multiplying two vectors is a scalar, which can be directly "spliced" into a large vector as input to the MLP. For the outer product form of PNN, multiplying two vectors is equivalent to matrix multiplication of a column vector and a row vector, resulting in a matrix. Each matrix is directly spliced into a large vector as input to the MLP. Alternatively, for the hidden layer, a three-layer 200-400-100 structure can be used; the relu function is used as the activation function; and the Dropout is increased.
[0116] Alternatively, the N target models can include an AutoInt model, as shown in Figure 9 Figure 9 is another schematic diagram for processing desensitized data based on a model provided by an embodiment of the present application, Figure 9 The model in the above formula (1) includes an input layer, an embedding layer, an interaction layer, and an output layer. The sample desensitized data is input into the input layer, and the input data is transmitted from the input layer to the embedding layer to obtain second embedding features. The embedding layer can map discrete features and continuous features of the input data into an embedding vector of the same length. The discrete features are directly looked up in the embedding table, and the multi-value discrete features use average pooling. The continuous features are equivalent to multiplying a Dense layer without bias. Further, the second embedding features are processed by the interaction layer, which can be stacked with multiple layers to realize high-order cross of features and obtain interaction features. Since the key to feature combination is to know which feature combination has strong representation ability, the interaction layer is actually equivalent to feature selection in manual feature engineering. At the same time, based on the self-attention mechanism, each field feature and other field features are considered to be attentioned, and the importance of the combination of the field feature and other field features is determined according to the weight of the attention. The higher the importance of the combination, the higher the weight given. Finally, the weighted sum pooling is generated as the result of the combination of the field feature and all other field features. The interaction features are processed by the output layer to obtain a prediction probability, and finally a model prediction result, i.e., a sample identity category to which a sample object belongs, is obtained.
[0117] Since the AutoInt model can realize the method of automatically performing high-order cross by finding a feature, it can not only make up for the weakness of MLP in capturing multiplicative feature combinations, but also better explain which feature combinations are more effective, so using the model can improve the data prediction efficiency. By proposing a method for learning high-dimensional feature cross, the interpretability of the model output result can be improved; based on the self-attentive neural network, a new method is proposed to automatically learn high-dimensional feature cross, which effectively improves the prediction accuracy. Optionally, the batch size of the AutoInt model can be 1024, and the embedding dimension d can be 16; further, the Adam optimizer can be used, and the dropout parameter is set to 0.5.
[0118] Optionally, by predicting a large amount of sample data, the sample identity category to which the sample object belongs can be determined, so as to carry out targeted advertising for the sample object, and then the model parameters can be adjusted according to the parameters such as the advertising click rate and the advertising conversion rate of the sample object, thereby improving the prediction effect of the model.
[0119] Optionally, the model process can be further solidified, and the trained target model can be solidified, and the training, verification, alarm and solidification are performed offline at regular intervals. The model can also be tested offline. The data set is trained based on the model, and after parameter optimization, the trained model is solidified based on the Saver() method of TensorFlow, generating four files: a checkpoint text file for recording the path information list of the model file; a model.ckpt.data file for recording network weight information; and a model.ckpt.index.data file and an.index file are binary files for saving variable weight information in the model. Since the trained model is solidified, the technical solution of the present application has strong reusability. Further, the user identity type to which the positive sample belongs can be changed, such as "user lifestyle", and then the server accumulates the corresponding log data, and finally the same feature splicing, feature processing and model training method is used to determine the result.
[0120] Optionally, in an embodiment of the present application, the effectiveness of the model can be evaluated based on the online traffic of the A / B Test, and the evaluation indicators may include but are not limited to the ad click-through rate and ad conversion rate. Specifically, after training the model, the computer device can use the trained model to detect the target data, determine the target identity category of the object, and push information to the object based on the target identity category of the object; obtain the ad click-through rate and / or ad click-through rate of the object for the pushed information within a preset time period; if the ad click-through rate and / or ad click-through rate are less than the expected threshold, it is determined that the current loss of the model is greater than the loss threshold, and the model is continued to be trained until the target identity category of the object is determined based on the model, and the corresponding ad click-through rate and / or ad click-through rate after pushing information to the object based on the target identity category is greater than or equal to the expected threshold.
[0121] Since this method can determine the effectiveness of the model, it can also be used to select N target models from multiple models. Among them, A / B Testing refers to developing two plans (such as two pages) for the same goal, having some users use Plan A and other users use Plan B, and recording user usage to see which plan is more in line with the design. After determining the user's target identity category by using the model, information can be pushed to the user based on the user's identity category, such as targeted advertising. The user's ad click-through rate and ad conversion rate for the pushed information can be obtained over a period of time, so that the plan can be adjusted, such as the model.
[0122] like Figure 10a As shown, Figure 10a : This is a schematic diagram of a model effect comparison provided in an embodiment of the present application, which includes three solutions for predicting the user's working status, namely, a manually formulated strong rule solution, a non-deep learning solution, and the technical solution of the present application. Among them, the AUC value corresponding to the offline manually formulated strong rule solution is 0.59; the AUC value corresponding to the offline non-deep learning solution is 0.65; the AUC value corresponding to the offline technical solution of the present application is 0.8. In addition, the AUC value corresponding to the online manually formulated strong rule solution is 0.57, the AUC value corresponding to the online non-deep learning solution is 0.6; the AUC value corresponding to the online technical solution of the present application is 0.75. From the perspective of offline AUC effect, the AUC of the technical solution of the present application is significantly improved compared with other technical solutions (manually formulated strong rule solution and non-deep learning solution). From the perspective of offline AUC efficiency, the AUC of the technical solution of the present application is also significantly improved compared with other technical solutions, which shows that the model effect of the technical solution of the present application is better.
[0123] Furthermore, if Figure 10b As shown, Figure 10bis a business effect comparison diagram provided in the embodiment of the application, wherein the advertisement click rate of the artificial strong rule scheme is 0.8%, the advertisement click rate of the non-deep learning scheme is 1.8%, and the advertisement click rate of the technical scheme of the application is 2.1%; the advertisement conversion rate of the artificial strong rule scheme is 0.1%, the advertisement conversion rate of the non-deep learning scheme is 0.7%, and the advertisement conversion rate of the technical scheme of the application is 1.3%. From the advertisement click rate, the advertisement click rate of the technical scheme of the application is obviously improved compared with other technical schemes. From the advertisement conversion rate, the advertisement conversion rate of the technical scheme of the application is also obviously improved compared with other technical schemes. It can be seen that the technical scheme of the application is better than other technical schemes.
[0124] In the embodiment of the application, a large amount of sample data is obtained to train the target model, so that the target model can predict the input sample data, and the parameters in the target model can be continuously adjusted to improve the accuracy of model prediction.
[0125] Optionally, please refer to Figure 11 , Figure 11 is a flow diagram of another data prediction method provided by the embodiment of the application. The data prediction method can be applied to a computer device, and the technical scheme of the application is described from three stages of sample data acquisition, model training and model use; as shown in Figure 11 , the data prediction method includes but is not limited to the following steps:
[0126] S301, sample data preparation.
[0127] Among them, the sample data preparation is mainly used for obtaining sample data, filtering the sample data based on rules, and filtering the abnormal data in the sample data. The specific sample data preparation method corresponds to step S201 in Figure 4 , which will not be repeated here.
[0128] S302, feature processing of sample data.
[0129] Among them, the feature processing mode of sample data corresponds to step S201 in Figure 4 , which will not be repeated here.
[0130] S303, multi-path model selection.
[0131] Among them, the multi-path model selection is based on the filtered sample data to train K models, and select N target models from the K models. The specific way of selecting N target models corresponds to step S201 in Figure 4 , which will not be repeated here.
[0132] S304, sample data desensitization.
[0133] The manner of sample data desensitization corresponds to Figure 4 In step S202, details are not repeated here.
[0134] In step S305, the selected model is trained based on the desensitized sample data, and the trained model is saved.
[0135] Because N target models are selected, the N target models can be trained using the desensitized sample data, and the N trained target models are saved. The specific implementation of training the model in step S305 can refer to the implementation of steps S203-S204, and details are not repeated here.
[0136] In step S306, target data to be predicted is obtained.
[0137] In step S307, the target data is desensitized.
[0138] In step S308, the desensitized target data is processed based on the trained model to determine the target identity category of the object.
[0139] The specific implementation of processing the target data in steps S306-S308 can refer to the implementation of processing the target data in steps S101-S104, and details are not repeated here.
[0140] In the embodiments of the present application, the target data to be predicted is desensitized, which can improve data security and reduce the risk of object identity information leakage. The target desensitized data is processed using N data prediction technologies to obtain N identity prediction sets. Because the target desensitized data is processed using multiple data prediction technologies, each data prediction technology processes the target desensitized data in a different manner, and the results obtained by processing are different. By combining multiple data prediction technologies to determine the identity category of the object, identity prediction can be performed from multiple dimensions, and the accuracy of data prediction can be improved.
[0141] Optionally, please refer to Figure 12 , Figure 12 is a flowchart of another data prediction method provided by the embodiments of the present application. As shown in Figure 12 , the data prediction method includes but is not limited to the following steps:
[0142] In step S401, a seed user with a label is obtained.
[0143] The seed user with a label can be a positive or negative sample with an explicit label.
[0144] In step S402, a basic portrait of the seed user is obtained.
[0145] The basic portrait can include some behavior data of the seed user in some applications, such as whether a certain type of function in some applications is started. For example, whether the spam blocking function, the answering assistant function and the like in the mobile phone manager are started.
[0146] S403, calculate an abnormal user type evaluation index.
[0147] The abnormal user type evaluation index can be used to determine whether there is a false user manipulating the terminal device, for example, whether it is an abnormal user can be determined according to the time of the user operating the application, the continuous duration and the like.
[0148] S404, filter abnormal seed users based on a distribution anomaly theorem.
[0149] The distribution anomaly theorem is used to determine the abnormal values in the sample data based on the distribution anomaly theorem, thereby filtering the abnormal seed users.
[0150] S405, determine whether the number of seed users meets the standard.
[0151] If yes, that is, the number of seed users meets the standard, indicating that the sample training data is sufficient, then step S406 is performed; if no, that is, the number of seed users does not meet the standard, indicating that the sample training data is not sufficient, then step S401 is performed until the number of seed users meets the standard. Steps S401-S405 are an offline data preparation process, and the specific implementation manner can refer to step S201 in Figure 4 .
[0152] S406, construct object features.
[0153] The object features can include but are not limited to portrait features of the object and business vertical type features.
[0154] S407, construct aggregated features in combination with a time dimension, perform feature processing on the aggregated features, and obtain processed features.
[0155] The computer device can obtain the aggregated features by aggregating portrait features and business features of different time spans. The feature processing on the aggregated features can include normalization feature processing and discretization feature processing, etc.
[0156] S408, merge the processed features and store them offline in HDFS.
[0157] By merging the processed features, multi-dimensional features of each user can be obtained, for example, a Y-dimensional feature of a user.
[0158] S409, solidify feature processing logic, perform automatic calculation offline at a fixed time, and store the offline calculation results in HDFS.
[0159] Among them, by solidifying the logic of feature processing, the data can be processed based on the processing logic later. Steps S406 to S409 are offline feature processing processes. The specific implementation method can be referred to Figure 4 Step S201.
[0160] S410, randomly divide the feature-processed sample set into a training set and a test set.
[0161] By dividing the sample set into a training set and a test set, the model can be trained using the training set to improve its accuracy. The model can then be tested using the test set to determine whether the model's detection results are consistent with the true sample values. This can then determine the model's accuracy, and the model can then be selected and adjusted based on its accuracy.
[0162] S411: Based on default parameters, multiple models are trained and N models are selected from the multiple models.
[0163] Among them, the default parameters may refer to the initial parameters in the model, and training the model is actually to adjust the initial parameters in the model multiple times so that the results obtained by processing the model input data based on the adjusted parameters are as similar as possible to the true values of the samples. After training multiple models, N models can be selected from multiple models based on the accuracy of the model predictions, such as the number of times the model prediction results are the same as the true values of the samples, and the number of times the results predicted by the N models are the same as the true values of the samples is greater than the number of times the prediction results of other models are the same as the true values of the samples. Alternatively, the AUC value of each model can be calculated, and N models can be selected from multiple models based on the AUC value of the model. Steps S410 to S411 are the model selection process, and the specific implementation method can be referred to. Figure 4 Step S201.
[0164] S412: Read in low-order features and high-order features, and concatenate them by column.
[0165] The low-order features and high-order features are the object features stored in HDFS. A user can include both high-order and low-order features. By combining the high-order and low-order features, a more complete picture of the user can be obtained. Alternatively, the low-order and high-order features can be combined in other ways.
[0166] S413, performing K-anonymity processing on the concatenated features to obtain an equivalent group M1.
[0167] S414: Extract the sensitive attribute set S in M1.
[0168] S415 , extracting and analyzing the sensitivity probability α of each sensitive attribute in the set S.
[0169] S416, add noise in the set S.
[0170] S417, construct a vector set conforming to laplace distribution for each sensitive group S.
[0171] S418, generate sensitive values conforming to the sensitive probability a and record the data into the equivalent group M2, to obtain the sample data after desensitization.
[0172] Wherein, the steps S412-S418 are the data desensitization process, and the specific implementation manner can refer to the step S202 in the Figure 4 .
[0173] S419, divide the training data into Z parts evenly, and use the leave-one-out method to train N target models.
[0174] Wherein, Z is a positive integer, and the leave-one-out method refers to dividing a large data set into q small data sets, wherein q-1 are used as the training set, and the remaining one is used as the test set, then the next one is selected as the test set, and the remaining q-1 are used as the training set. By using the leave-one-out method, as much effective information as possible can be obtained from limited data, so that the sample can be learned from multiple angles to avoid falling into local extreme value.
[0175] S420, use the trained N target models to predict the remaining one part of the training data and the test data.
[0176] S421, combine the Z parts of the predicted training data to obtain new training data.
[0177] S422, combine the Z parts of the predicted test data using the mean method to obtain new test data.
[0178] S423, based on the new training data and the new test data, determine the N target models, and solidify the N models.
[0179] S424, solidification model training process.
[0180] Wherein, the steps S419-S424 are the process of training the model based on sensitive data, and the specific implementation manner can refer to the steps S203-S204 in the Figure 4 . Wherein, after the solidification model training process, the model can also be trained offline at regular intervals, the model is verified, an alarm is given, and the model after offline processing is solidified. For example, when any one link in the training process is abnormal, an alarm can be given to prompt the relevant personnel to handle it.
[0181] Optionally, the processes of steps S419 to S424 can also be referred to as a Stacking ensemble learning framework, which refers to fusing multiple classification models through a meta-classifier. The secondary classifier outputs a prediction result after being trained based on a training set, and then the meta-classifier trains the output according to the secondary classifier. That is, the Stacking ensemble learning framework can refer to predicting data through N target models, and then performing average pooling on the prediction results of the N target models to obtain a final prediction result.
[0182] Optionally, after the target model is trained, the target identity category of the object can be determined based on the target model. For details, refer to the related description of the above embodiments, which will not be repeated here.
[0183] In the embodiments of the present application, by performing desensitization processing on the target data to be predicted, the data security can be improved, and the risk of object identity information leakage can be reduced. The target desensitization data is processed by using N kinds of data prediction technologies respectively, N identity prediction sets are obtained. Since the target desensitization data is processed by using multiple data prediction technologies, each data prediction technology processes the target desensitization data in a different way, and the results obtained by processing are different. By combining multiple data prediction technologies to determine the identity category of the object, identity prediction can be performed from multiple dimensions, and the accuracy of data prediction can be improved.
[0184] The method of the embodiments of the present application is introduced above, and the device of the embodiments of the present application is introduced below.
[0185] Referring to Figure 13 , Figure 13 is a schematic diagram of the composition structure of a data prediction device provided by the embodiments of the present application. The above data prediction device can be a computer program (including program code) running in a terminal device; the data prediction device can be used to execute the corresponding steps in the data prediction method provided by the embodiments of the present application. For example, the data processing device 130 includes:
[0186] The data acquisition unit 1301 is configured to acquire target data to be predicted, the target data being used to indicate the identity category of an object;
[0187] The data desensitization unit 1302 is configured to perform desensitization processing on the target data to obtain target desensitization data;
[0188] The data processing unit 1303 is configured to process the target desensitization data by using N kinds of data prediction technologies respectively, and obtain N identity prediction sets, each identity prediction set including the probability that the object belongs to M predicted identity categories, N being a positive integer greater than or equal to 2, and M being a positive integer;
[0189] The identity determination unit 1304 is configured to determine a target identity category of the object based on the N identity prediction sets.
[0190] Optionally, the target desensitization data comprises dense features and / or sparse features; and the data processing unit 1303 is specifically configured to process the target desensitization data by using N data prediction technologies respectively, to obtain the N identity prediction sets.
[0191] Optionally, the N identity prediction sets comprise at least two of a first identity prediction set, a second identity prediction set and a third identity prediction set; and the data processing unit 1303 is specifically configured to:
[0192] perform feature compression on the sparse features, perform feature splicing on the compressed features and the dense features to obtain first spliced features, and determine the first identity prediction set based on the first spliced features; and / or,
[0193] perform feature compression on the sparse features and the dense features, perform feature splicing on the compressed features to obtain second spliced features, and determine the second identity prediction set based on the second spliced features; and / or,
[0194] perform feature compression on the sparse features, perform feature splicing on the compressed features and the dense features to obtain third spliced features, and perform weight processing on the third spliced features based on an attention mechanism to determine the third identity prediction set.
[0195] Optionally, the identity determination unit 1304 is specifically configured to:
[0196] determine, from the N identity prediction sets, a prediction identity category belonging to the same category and a probability of the prediction identity category of the same category;
[0197] determine, based on the probability of the prediction identity category of the same category in the N identity prediction sets, a probability of each prediction identity category in M prediction identity categories, to obtain a total probability of each prediction identity category;
[0198] determine a maximum probability from the total probability of the M prediction identity categories;
[0199] determine a prediction identity category corresponding to the maximum probability as the target identity category of the object.
[0200] Optionally, the data desensitization unit 1302 is specifically configured to:
[0201] obtain an original sensitive probability of the target data;
[0202] add noise to the target data to obtain noise data, and determine a target sensitive probability of the noise data based on the noise data and the original sensitive probability.
[0203] If the target sensitive probability is in the target probability interval, the noise data is determined as the target desensitization data.
[0204] Optionally, the data prediction apparatus 130 further comprises a model training unit 1305 configured to:
[0205] Obtain sample data to be predicted, the sample data being used to indicate a sample identity category of a sample object;
[0206] Perform desensitization processing on the sample data to obtain sample desensitization data;
[0207] Process the sample desensitization data by using N kinds of data prediction techniques respectively to obtain N sample prediction sets, each sample prediction set including probabilities of the sample object belonging to M predicted identity categories;
[0208] Determine a target sample identity category of the sample object based on the N sample prediction sets;
[0209] Obtain a sample identity label of the sample data, and train N target models based on the sample identity label and the target sample identity category;
[0210] The data processing unit 1303 is specifically configured to:
[0211] Process the target desensitization data by using data prediction techniques corresponding to the N target models respectively to obtain N identity prediction sets.
[0212] Optionally, the model training unit 1305 is specifically configured to:
[0213] Obtain at least one initial sample data, filter the at least one initial sample data based on an abnormal rule to obtain sample filtered data, the abnormal rule including at least one abnormal behavior;
[0214] Filter the sample filtered data based on a distribution abnormality theorem to obtain the sample data to be predicted, the distribution abnormality theorem being used to filter data based on probability theory.
[0215] Optionally, the model training unit 1305 is specifically configured to:
[0216] Input the sample data to be predicted into K models for training to determine index parameters of the K models, the index parameters being used to reflect classification performance of the models, K being a positive integer greater than N;
[0217] Determine N target models from the K models based on the index parameters of the K models;
[0218] The N target models correspondingly predict the sample desensitization data by using data prediction technology to obtain N sample prediction sets.
[0219] It should be noted that, Figure 13 The contents not mentioned in the corresponding embodiments can be referred to the description of the method embodiments, which will not be repeated here.
[0220] In the embodiments of the present application, by desensitizing the target data to be predicted, the data security can be improved and the risk of object identity information leakage can be reduced. Since multiple data prediction technologies are used to process the target desensitization data, each data prediction technology processes the target desensitization data in different ways, and therefore the results obtained by processing are different. By combining multiple data prediction technologies to determine the identity category of the object, identity prediction can be performed from multiple dimensions, and the accuracy of data prediction can be improved.
[0221] Referring to Figure 14 , Figure 14 is a schematic diagram of a computer device provided by an embodiment of the present application. As Figure 14 shown, the computer device 140 can include a processor 1401, a memory 1402, and a network interface 1403. The processor 1401 is connected to the memory 1402 and the network interface 1403, for example, the processor 1401 can be connected to the memory 1402 and the network interface 1403 through a bus. The computer device can be a terminal device or a server.
[0222] The processor 1401 is configured to support the data processing device to perform the corresponding functions in the above-mentioned data processing method. The processor 1401 can be a central processing unit (CPU), a network processor (NP), a hardware chip or any combination thereof. The hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The PLD can be a complex programmable logic device (CPLD), a field programmable logic gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0223] The memory 1402 is configured to store program codes and the like. The memory 1402 can include a volatile memory (VM), for example, a random access memory (RAM); the memory 1402 can also include a non-volatile memory (NVM), for example, a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); the memory 1402 can also include a combination of the above-mentioned memories. In the embodiments of the present application, the memory 1402 is configured to store programs of website security detection, interactive traffic data and the like.
[0224] The network interface 1403 is configured to provide network communication functions.
[0225] The processor 1401 can invoke the program codes to perform the following operations:
[0226] Obtain target data to be predicted, the target data being configured to indicate an identity category of an object;
[0227] Perform desensitization processing on the target data to obtain target desensitization data;
[0228] Respectively perform processing on the target desensitization data by using N kinds of data prediction technologies to obtain N identity prediction sets, each identity prediction set including probabilities of the object belonging to M predicted identity categories, N being a positive integer greater than or equal to 2, and M being a positive integer;
[0229] Determine a target identity category of the object based on the N identity prediction sets.
[0230] It should be understood that the computer device 140 described in the embodiments of the present application can perform the description of the above-mentioned data processing method in the foregoing Figure 3 、 Figure 4 、 Figure 11 and Figure 12 corresponding embodiments, and can also perform the description of the above-mentioned data processing apparatus in the foregoing Figure 13 corresponding embodiments, which will not be described here again. In addition, the description of the beneficial effects of using the same method will not be described again.
[0231] The embodiments of the present application further provide a computer readable storage medium storing a computer program, the computer program comprising program instructions which, when executed by a computer, cause the computer to perform the method of the foregoing embodiments. The computer can be part of the computer device mentioned above. For example, the computer can be the processor 1401 mentioned above. As an example, the program instructions can be deployed to execute on one computer device, or on multiple computer devices located in one place, or on multiple computer devices distributed in multiple places and interconnected through a communication network, which can constitute a blockchain network.
[0232] The embodiments of the present application further provide a computer program product or computer program comprising computer instructions stored in a computer readable storage medium. A processor of a computer device can read the computer instructions from the computer readable storage medium, and the processor can execute the computer instructions to cause the computer device to perform the steps performed in the embodiments of the methods described above.
[0233] A person of ordinary skill in the art can understand that all or part of the processes in the methods of the embodiments described above can be completed by a computer program instructing relevant hardware. The program can be stored in a computer readable storage medium, and when the program is executed, the program can include the processes of the embodiments of the methods described above. The storage medium can be a disk, an optical disk, a Read-Only Memory (ROM) or a Random Access Memory (RAM), etc.
[0234] The above only describes the preferred embodiments of the present application, and of course cannot limit the scope of the rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope of the present application.
Claims
1. A data prediction method characterized by, The method comprises the following steps: obtaining target data to be predicted, the target data being used to indicate an identity category of an object; performing desensitization processing on the target data to obtain target desensitization data; the desensitization processing comprises: obtaining an original sensitive probability of the target data; adding noise to the target data to obtain noise data, and determining a target sensitive probability of the noise data based on the noise data and the original sensitive probability; if the target sensitive probability is in a target probability interval, the noise data is determined as the target desensitization data; respectively using N kinds of data prediction technologies to process the target desensitization data to obtain N identity prediction sets, each identity prediction set comprising a probability that the object belongs to M predicted identity categories, N being a positive integer greater than or equal to 2, and M being a positive integer; the target desensitization data comprises dense features and / or sparse features, and different data prediction technologies have different processing modes for the dense features and / or the sparse features; determining a target identity category of the object based on the N identity prediction sets.
2. The method of claim 1, wherein, The method comprises the following steps: respectively using N kinds of data prediction technologies to process the dense features and / or the sparse features to obtain N identity prediction sets.
3. The method of claim 2, wherein, The N identity prediction sets comprise at least two of a first identity prediction set, a second identity prediction set and a third identity prediction set. The method comprises the following steps: performing feature compression on the sparse features, performing feature splicing on the compressed features and the dense features to obtain first spliced features, and determining the first identity prediction set based on the first spliced features; and / or performing feature compression on the sparse features and the dense features, performing feature splicing on the compressed features to obtain second spliced features, and determining the second identity prediction set based on the second spliced features; and / or performing feature compression on the sparse features, performing feature splicing on the compressed features and the dense features to obtain third spliced features, and performing weight processing on the third spliced features based on an attention mechanism to determine the third identity prediction set.
4. The method according to any one of claims 1 to 3, characterized in that, The method comprises the following steps: determining, from the N identity prediction sets, predicted identity categories of the same category and probabilities of the predicted identity categories of the same category; determining, based on the probabilities of the predicted identity categories of the same category in the N identity prediction sets, a probability of each predicted identity category of the M predicted identity categories to obtain a total probability of each predicted identity category; determining a maximum probability from the total probabilities of the M predicted identity categories; determining a predicted identity category corresponding to the maximum probability as the target identity category of the object.
5. The method of claim 1, wherein, The method further comprises the following steps: obtaining sample data to be predicted, the sample data being used to indicate a sample identity category of a sample object; performing desensitization processing on the sample data to obtain sample desensitization data; respectively adopting N kinds of data prediction technologies to process the sample desensitization data, to obtain N sample prediction sets, each sample prediction set including the probability of the sample object belonging to M kinds of predicted identity categories; determining the target sample identity category of the sample object based on the N sample prediction sets; obtaining a sample identity label of the sample data, and training N target models based on the sample identity label and the target sample identity category; respectively adopting N kinds of data prediction technologies to process the sample desensitization data, to obtain N sample prediction sets, each sample prediction set including the probability of the sample object belonging to M kinds of predicted identity categories; respectively adopting N kinds of data prediction technologies to process the sample desensitization data, to obtain N sample prediction sets, each sample prediction set including the probability of the sample object belonging to M kinds of predicted identity categories.
6. The method of claim 5, wherein, The method further comprises: inputting the sample data to be predicted into K models for training, and determining index parameters of the K models, the index parameters being used to reflect the classification performance of the models, K being a positive integer greater than N; determining N target models from the K models based on the index parameters of the K models; 7. The method of claim 5, wherein, respectively adopting N kinds of data prediction technologies to process the sample desensitization data, to obtain N sample prediction sets, each sample prediction set including the probability of the sample object belonging to M kinds of predicted identity categories. Comprise: a processor, a memory and a network interface; The processor is connected with the memory and the network interface, wherein the network interface is used to provide data communication function, the memory is used to store program code, and the processor is used to call the program code, so that the computer equipment executes the method of any one of claims 1-7. The computer readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by the processor, so that the computer equipment with the processor executes the method of any one of claims 1-7.
8. A computer device, comprising: Comprise: a data acquisition unit, configured to acquire target data to be predicted, the target data being used to indicate an identity category of an object; a data desensitization unit, configured to perform desensitization processing on the target data to obtain target desensitization data; 9. A computer-readable storage medium, characterized in that, The desensitization processing comprises: acquiring an original sensitive probability of the target data; adding noise to the target data to obtain noise data, and determining a target sensitive probability of the noise data based on the noise data and the original sensitive probability; if the target sensitive probability is in a target probability interval, the noise data is determined as the target desensitization data; 10. A data prediction apparatus, characterized by comprising: A data processing unit is configured to process the target de-identification data by using N data prediction techniques respectively, to obtain N identity prediction sets, each of which includes probabilities of the object belonging to M predicted identity categories, N is a positive integer greater than or equal to 2, and M is a positive integer; the target de-identification data includes dense features and / or sparse features, and different data prediction techniques have different processing manners for the dense features and / or the sparse features; An identity determination unit is configured to determine a target identity category of the object based on the N identity prediction sets.
11. A computer program product, characterised in that, The computer program product includes computer instructions, which, when executed, implement the method of any one of claims 1-7.
Citation Information
Patent Citations
Data processing method and device for realizing privacy protection
CN111523146A
Object attribute identification method and device, storage medium and computer equipment
CN113569111A
Predicting object identity using an ensemble of predictors
US8484225B1