User data processing method, device, computer equipment and storage medium
By filtering and clustering user data and training classification models, the problem of low user data classification accuracy in the existing technology is solved, and higher classification accuracy and targeted information push are achieved.
Patent Information
- Application Number
- CN202010612623.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-30
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-06-30
AI Technical Summary
The prior art is difficult to accurately judge the behavioral characteristics of users' sparse behavior, resulting in a low accuracy rate of user data classification results.
By obtaining multiple user data, filtering out the target user data, determining the similarity degree based on the characteristic data and weights of each dimension, clustering, obtaining categories, training classification models, classifying the remaining user data, and finally pushing information.
It improves the accuracy of user data classification, avoids the problem of insufficient computing and resources caused by full data clustering, and achieves more accurate and targeted information push.
Smart Images

Figure CN111667022B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a user data processing method, apparatus, computer equipment and storage medium. Background Art
[0002] With the development of computer technology, people increasingly rely on the Internet to obtain information from all aspects. In order to provide products or services to users in a timely manner and avoid providing useless products or services as much as possible, the target population that receives products or services is usually determined based on the classification labels of user data.
[0003] However, the existing classification of user data is mainly based on statistical description of user image. This method makes it difficult to accurately judge the sparse behavioral characteristics of users, and further makes it difficult to judge the category to which the user data truly belongs, resulting in a relatively low accuracy rate in the classification results of the obtained user data. Summary of the invention
[0004] Based on this, it is necessary to provide a user data processing method, apparatus, computer device and storage medium that can improve the accuracy of user data classification in response to the above technical problems.
[0005] A user data processing method, the method comprising:
[0006] Acquire multiple user data; each of the user data includes feature data of multiple dimensions;
[0007] Filtering out a plurality of target user data from the user data;
[0008] Determine the similarity between the target user data according to the feature data of each dimension and the corresponding weight of each dimension;
[0009] Clustering the target user data based on the similarity between each target user data to obtain the category to which each target user data belongs;
[0010] Training a classification model according to the target user data and the category to which the target user data belongs;
[0011] Classify the user data remaining after screening by the classification model obtained through training, and obtain the category to which each user data remaining after screening belongs;
[0012] Information is pushed according to the category to which each of the user data belongs.
[0013] A user data processing device, the device comprising:
[0014] An acquisition module, used to acquire multiple user data; each of the user data includes feature data of multiple dimensions;
[0015] A first screening module, used to screen out a plurality of target user data from the user data;
[0016] A determination module, configured to determine the similarity between the target user data according to the feature data of each dimension and the corresponding weight of each dimension;
[0017] A clustering module, used for clustering the target user data based on the similarity between the target user data to obtain the category to which the target user data belongs;
[0018] A training module, used for training a classification model according to the target user data and the category to which the target user data belongs;
[0019] A second screening module is used to classify the user data remaining after screening by using the classification model obtained through training, so as to obtain the category to which each user data remaining after screening belongs;
[0020] The application module is used to push information according to the category to which each user data belongs.
[0021] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0022] Acquire multiple user data; each of the user data includes feature data of multiple dimensions;
[0023] Filtering out a plurality of target user data from the user data;
[0024] Determine the similarity between the target user data according to the feature data of each dimension and the corresponding weight of each dimension;
[0025] Clustering the target user data based on the similarity between each target user data to obtain the category to which each target user data belongs;
[0026] Training a classification model according to the target user data and the category to which the target user data belongs;
[0027] Classify the user data remaining after screening by the classification model obtained through training, and obtain the category to which each user data remaining after screening belongs;
[0028] Information is pushed according to the category to which each of the user data belongs.
[0029] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0030] Acquire multiple user data; each of the user data includes feature data of multiple dimensions;
[0031] Filtering out a plurality of target user data from the user data;
[0032] Determine the similarity between the target user data according to the feature data of each dimension and the corresponding weight of each dimension;
[0033] Clustering the target user data based on the similarity between each target user data to obtain the category to which each target user data belongs;
[0034] Training a classification model according to the target user data and the category to which the target user data belongs;
[0035] Classify the user data remaining after screening by the classification model obtained through training, and obtain the category to which each user data remaining after screening belongs;
[0036] Information is pushed according to the category to which each of the user data belongs.
[0037] After obtaining a large amount of multi-dimensional user data, the above-mentioned user data processing method, device, computer equipment and storage medium screen out a small amount of user data as a target for clustering and classification, obtain target user data with categories, and then use the target user data with categories to train the classification model, and then use the trained classification model to classify the remaining user data after screening. In this way, on the one hand, when clustering, different weight parameters are set for feature data of different dimensions, which can well process discrete user data in a targeted manner, and then more accurately cluster user data according to the importance of feature data of different dimensions, thereby improving user classification accuracy; and the classification model trained based on this part of target user data with categories can also accurately and effectively classify the remaining user data after screening; on the other hand, clustering only part of the user data can also avoid the amount of calculation and possible insufficient computing resources caused by clustering the entire amount of data; in addition, after obtaining the category to which the user data belongs, information push can be more accurate and targeted. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 An application environment diagram of a user data processing method in one embodiment;
[0039] Figure 2 is a flowchart of a method for processing user data in one embodiment;
[0040] Figure 3 A flowchart of user data processing in one embodiment;
[0041] Figure 4 is a schematic diagram of a clustering effect after clustering target user data in one embodiment;
[0042] Figure 5 It is a schematic diagram of the classification effect after classifying the remaining user data after screening in one embodiment;
[0043] Figure 6 A schematic diagram of the structure of a classification model in one embodiment;
[0044] Figure 7 is a schematic diagram of the structure of an attention network structure in one embodiment;
[0045] Figure 8 is a structural block diagram of a user data processing device in one embodiment;
[0046] Fig. 9 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0048] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0049] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, mechatronics, etc. Artificial intelligence software technologies include natural language processing technology and machine learning.
[0050] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.
[0051] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing.
[0052] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, which can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The backend services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the rapid development and application of the Internet industry, in the future, each item may have its own identification mark, which needs to be transmitted to the backend system for logical processing. Data of different levels will be processed separately. All kinds of industry data need strong system backing support, which can only be achieved through cloud computing.
[0053] The main research directions of cloud security include: 1. Cloud computing security, which mainly studies how to ensure the security of the cloud itself and various applications on the cloud, including cloud computer system security, secure storage and isolation of user data, user access authentication, information transmission security, network attack protection, compliance auditing, etc.; 2. Cloudification of security infrastructure, which mainly studies how to use cloud computing to build and integrate security infrastructure resources and optimize security protection mechanisms, including building a large-scale security event, information collection and processing platform through cloud computing technology, realizing the collection and correlation analysis of massive information, and improving the control ability of security events and risk control capabilities of the entire network; 3. Cloud security services, which mainly studies various security services provided to users based on cloud computing platforms, such as antivirus services.
[0054] The following will explain the user data processing method provided in the embodiment of the present application based on machine learning and cloud technology of artificial intelligence technology.
[0055] The user data processing method provided in this application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network, and the number of terminals 102 is multiple. Specifically, each terminal 102 can upload user data to the server 104, and each user data includes feature data of multiple dimensions. The server 104 obtains multiple user data; then screens out multiple target user data from these user data; determines the similarity between each target user data according to the feature data of each dimension of each target user data and the corresponding weight of each dimension; then clusters the target user data based on the similarity between each target user data to obtain the category to which each target user data belongs; then trains a classification model according to the target user data and the category to which the target user data belongs; then classifies the remaining user data after screening through the classification model obtained by training, and obtains the category to which each user data remaining after screening belongs; and pushes information according to the category to which each user data belongs. In another embodiment, the terminal 102 can also upload user data to the server 104 through the application running thereon.
[0056] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices, and the server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and this application does not limit this.
[0057] In one embodiment, Figure 2 As shown, a user data processing method is provided, and the user data processing method is applied to a computer device (eg Figure 1 Taking the server in the example as an example, the user data processing method includes the following steps:
[0058] Step 202, obtaining multiple user data; each user data includes feature data of multiple dimensions.
[0059] Among them, user data is data reflecting user characteristics. User data includes user action data. User action data is data reflecting user action characteristics. User action data includes social behavior data, browsing behavior data, or payment behavior data. Social behavior data includes social conversation data, social message publishing data, or social message comment information. Browsing behavior data includes news browsing data, audio and video browsing data, or product browsing data. Payment behavior data includes consumption behavior data or transfer behavior data. User data may also include user basic data. User basic data is data reflecting user basic attributes. User basic data includes gender, age, or education level.
[0060] It can be understood that in the present application, a plurality includes two and more than two. In a specific embodiment, the plurality of user data may be a data set of millions, tens of millions or even hundreds of millions.
[0061] Dimensions can refer to the field to which feature data belongs, and can also be called feature domains. Dimensions include age dimension, gender dimension, video dimension, or image and text dimension.
[0062] In one embodiment, the dimensional division can be customized according to actual needs. The multiple dimensions obtained by dividing in one dimensional division method can be a dimension obtained by dividing in another dimensional division method. For example, in method A, the age dimension, gender dimension and regional dimension are obtained, and in method B, the basic information dimension is obtained. Then, it can be considered that the age dimension, gender dimension and regional dimension can be sub-dimensions of the basic information dimension. A dimension obtained by dividing in one dimensional division method can also be a plurality of dimensions obtained by dividing in another dimensional division method. For example, in method A, the video dimension is obtained, and in method B, the sports video dimension and the entertainment video dimension are obtained. Then, it can be considered that the sports video dimension and the entertainment video dimension can be sub-dimensions of the video dimension.
[0063] It is understandable that, usually, each user will only take actions on the scattered content that he or she is interested in. Then, the actions generated based on these scattered contents will also generate some scattered user data accordingly, and the computer device will obtain the discrete user data of each user. In the prior art, it is usually not possible to perform user classification based on these high-dimensional and discrete user data.
[0064] In one embodiment, the user data acquired by the computer device in step 202 may be original user data, and the user data needs to be filtered before subsequent processing. For example, the computer device acquires M-dimensional user data, filters out N-dimensional data from the M-dimensional data, obtains N-dimensional user data, and then performs subsequent processing; wherein M>N.
[0065] In one embodiment, the user data acquired by the computer device in step 202 may be pre-processed user data, and the dimensions of the pre-processed user data including feature data are specific dimensions, such as dimensions with higher weights. For example, the computer device acquires user data of N dimensions that have been screened out.
[0066] Step 204: Filter out a plurality of target user data from the user data.
[0067] The target user data refers to the user data selected from a large amount of user data and processed as a target. The target user data may also be referred to as seed user data.
[0068] Specifically, the seed user data may be user data of a seed user. The computer device may filter out user data of a plurality of seed users from the user data as target user data.
[0069] It should be noted that the seed user data is not necessarily the user data of the seed user.
[0070] In one embodiment, step 204 includes: performing data cleaning on the user data; performing multi-level division on the cleaned user data to obtain multiple user data subsets; and randomly selecting user data from each of the minimum number of user data subsets as target user data.
[0071] Specifically, the computer device can clean the user data, such as deleting duplicate information and correcting existing erroneous information, etc. to convert dirty data into data that meets data quality requirements. Since the amount of user data is large, the computer device can initially divide the user data in a simple division method, such as dividing by age group, and then further divide the divided user data to achieve multi-level division of user data, and obtain multiple user data subsets, each of which has a small number of user data subsets, and then randomly select user data from each of the minimum number of user data subsets as target user data.
[0072] For example, see Figure 3 , which shows a flowchart of user data processing in an embodiment. It can be seen that the computer device can first collect the original user data, then perform data cleaning on the collected original user data, and then filter out a part of the user data after the data cleaning as seed user data, and use the remaining user data after the screening as non-seed user data. Among them, the amount of seed user data is much smaller than the amount of non-seed user data, which can reduce the amount of computation for clustering the seed user data.
[0073] In this embodiment, the user data is first divided into multiple levels and split into smaller subsets, and then the user data is randomly selected from the subsets. Compared with directly randomly selecting from a large amount of data, various representative user data can be selected more comprehensively, avoiding screening out user data that is too similar.
[0074] Step 206 : Determine the similarity between the target user data according to the feature data of each dimension and the corresponding weight of each dimension.
[0075] It can be understood that user data includes feature data of multiple dimensions. At present, no matter in what specific scenario the user data is clustered, the feature data of each dimension is usually processed indiscriminately. This may make the clustering result unsuitable for the current application scenario, and thus fail to distinguish the value of users in specific scenarios. For example, in a video promotion scenario, the importance of the user's video browsing behavior data to user data clustering is higher than the importance of the user's social conversation data. In this embodiment, different weight parameters are set for feature data of different dimensions, so that more attention can be paid to more important features in the current application scenario during clustering, and the obtained clustering results are more suitable for the current application scenario.
[0076] Specifically, the computer device may calculate the first similarity between the feature data of any two target user data in each dimension of the target user data, and then fuse the first similarity in each dimension according to the weight of each dimension to obtain the second similarity as the similarity between the two target user data. The first similarity in each dimension is fused according to the weight of each dimension to obtain the second similarity, and specifically, the first similarity in each dimension is weighted and summed according to the weight of the corresponding dimension to obtain the second similarity.
[0077] For example, assuming that each target user data includes feature data of three dimensions (dimension 1, dimension 2, and dimension 3), then for target user data A and target user data B, the first similarity 1 of dimension 1, the first similarity 2 of dimension 2, and the first similarity 3 of dimension 3 are calculated respectively, and then the first similarity of each dimension is merged according to the weight of each dimension to obtain the second similarity. Assuming the weight of dimension 1 is 1, the weight of dimension 2 is 2, and the weight of dimension 3 is 3, the first similarity 1, the first similarity 2, and the first similarity 3 are weighted and summed according to the weights of the corresponding dimensions to obtain the first similarity 1*weight 1+first similarity 2*weight 2+first similarity 3*weight 3, that is, the second similarity, which is the similarity between A and B.
[0078] The higher the similarity between target user data, the higher the similarity between users corresponding to the target user data, that is, the more similar the users are. User data can be vectorized, feature data can be numericalized, the similarity between feature data can be quantified by the distance between numerical values, and the similarity between user data can be quantified by the distance between vectors.
[0079] In one embodiment, step 206 includes: vectorizing each target user data to obtain a feature vector of each target user data; the feature vector includes feature values of multiple dimensions; each feature value corresponds to feature data of one dimension; obtaining a corresponding weight for each dimension; for the feature vectors of any two target user data, calculating the distance between the feature values of each dimension according to the weight of each dimension to obtain the similarity between any two target user data.
[0080] Vectorization refers to expressing data in other forms in mathematical form. For example, the text form of "XXX" is expressed in mathematical form as "[0 0 0 1 0 0 0 0 0 0 0...]". In this case, "[0 0 0 1 0 0 0 00 00...]" is the result of vectorizing "XXX", that is, the vector of "XXX". It can be understood that there is no limitation on what kind of vector representation data in other forms can be converted into, as long as data in other forms can be expressed mathematically.
[0081] Each vector element of the feature vector represents the feature value corresponding to the feature data of one dimension. For example, assuming that the user data of the target user A includes feature data of four dimensions, 1, 2, 3, and 4, the feature vector obtained by vectorizing the user data of user A is (X1, X2, X3, X4). Then, X1 is the feature value corresponding to the feature data of dimension 1, X2 is the feature value corresponding to the feature data of dimension 2, X3 is the feature value corresponding to the feature data of dimension 3, and X4 is the feature value corresponding to the feature data of dimension 4.
[0082] Specifically, the computer device can convert discrete user data into continuous feature vectors for representation by performing an Embedding operation on the user data.
[0083] In one embodiment, the corresponding weight of each dimension can be set based on historical experience data, or can be obtained by learning existing classified user data through a machine learning model, or can be obtained through other methods, which are not limited in the embodiments of the present application.
[0084] In a specific embodiment, the computer device can measure the similarity between user data by weighted Euclidean distance. The weighted Euclidean distance calculation formula is as follows:
[0085]
[0086] Where X = {x 1 ,x 2 ,x 3 ,…,x k ,…,x n} is a feature vector corresponding to the target user data, Y = {y 1 ,y 2 ,y 3 ,…,y k ,…,y n} is the feature vector corresponding to another target user data, n is the dimension of the number of features Euc Distance(X,Y) is the weighted Euclidean distance between X and Y, reflecting the similarity between X and Y. k is the weight of the feature data of the kth dimension.
[0087] It is understood that in other embodiments, the computer device may also use other distances to measure the similarity between user data, such as Manhattan distance, Chebyshev distance, and earth-moving distance, etc. When other distances are used to measure the similarity between user data, weight parameters are also introduced in a weighted manner.
[0088] In the above embodiment, when clustering user data, different weight parameters are set for feature data of different dimensions, so that discrete user data can be processed in a targeted manner, and user data can be clustered more accurately according to the importance of feature data of different dimensions, thereby improving the classification accuracy of user data and improving the accuracy of user value division.
[0089] In one embodiment, step 206 includes: obtaining the corresponding weight of each dimension; for each target user data, retaining a preset number of feature data of the dimension with the largest weight, and obtaining target user sub-data for clustering; for the feature vectors of any two target user sub-data, calculating the distance of the feature values of each dimension according to the weight of each dimension, and obtaining the similarity between any two target user sub-data.
[0090] It can be understood that when the user data obtained in step 202 is the original user data, when calculating the similarity between the user data, the feature data of all dimensions can be taken into account, that is, the similarity between any two target user data is calculated according to the steps in the previous embodiment. When the user data obtained in step 202 is the user data of a specific dimension, the similarity between any two target user data is calculated according to the steps in the previous embodiment, then the feature data of the specific dimension is taken into account, which can reduce the amount of calculation while meeting the requirements for calculating the similarity of user data in the current scenario.
[0091] In this embodiment, when the user data obtained in step 202 is the original user data, the original user data may be further dimensionally screened, a preset number of feature data with the largest weight in the dimension may be retained, target user sub-data for clustering may be obtained, and then the target user sub-data may be clustered. At this time, it may be considered that the similarity between any two target user sub-data may represent the similarity between the corresponding two target user data, and further represent the similarity between the corresponding users.
[0092] Step 208: cluster the target user data based on the similarity between the target user data to obtain the category to which the target user data belongs.
[0093] Specifically, after calculating the similarity between the target user data, the computer device can cluster the target user data according to the similarity to obtain multiple clusters. Each cluster corresponds to a category. The cluster generated by clustering is a set of target user data. The similarity between target user data in the same cluster is high, and the similarity between target user data in different clusters is low.
[0094] Continue to refer Figure 3 After filtering out the seed user data, the computer device can cluster the seed user data to obtain multiple clusters. Each cluster corresponds to a category. It can be understood that the classification of user data can be used to group users, thereby obtaining a user grouping result.
[0095] refer to Figure 4 , which shows a schematic diagram of the clustering effect after clustering the target user data in one embodiment. The classification of user data can be used to group users by value, such as Figure 4As shown, in this embodiment, the computer device clusters the user data using the feature data of the three dimensions with the largest weights, and the three-dimensional coordinates of the figure represent the dimensions of different feature data. The computer device divides the users into eight groups through clustering: important value users, important retention users, important development users, important retention users, general value users, general retention users, general development users, and general retention users.
[0096] Step 210: training a classification model according to the target user data and the category to which the target user data belongs.
[0097] Specifically, after obtaining the category to which each target user data belongs, the computer device can use the category to which each target user data belongs as the corresponding training label of each target user data, and then train the classification model in a supervised manner based on the target user data and its corresponding training label.
[0098] After the classification model is trained, it can be used to classify user data. The categories obtained by the classification model from classifying the user data include the categories to which the target user data belongs.
[0099] In one embodiment, when the target user data is used to train the classification model, feature data of all dimensions may be used as training samples, or feature data of some dimensions may be used as training samples. However, no matter what feature data is used, the category to which the user data belongs does not change.
[0100] In one embodiment, step 210 includes: for each target user data, retaining a preset number of feature data of the dimension with the largest weight, to obtain target user sub-data for training the classification model; and training the classification model according to the target user sub-data and the category to which the target user sub-data belongs.
[0101] Assume that the weight is δ = {δ 1 ,δ 2 ,δ 3 ,δ 4 ,δ 5 ,…,δ n}, the preset number is 3 and the top three weights are δ 3 ,δ 5 and δ 1 , then the feature data of the dimensions corresponding to the three weights are included, and the feature data of these dimensions are vectorized as the input of the classification model. Specifically, the computer device can convert the discrete user data into a continuous feature vector for representation by performing an Embedding operation on the user data, so that the feature vector can be directly input into the classification model for processing.
[0102] In this way, using user data from some more important dimensions for model training can not only reduce the amount of data and computation required for model training, but also ensure the effectiveness of model training.
[0103] Step 212: classify the user data remaining after screening by using the trained classification model to obtain the category to which each user data remaining after screening belongs.
[0104] Specifically, the computer device may process the remaining user data after screening according to the data form of the training samples during the classification model training, and then classify the remaining user data after screening through the trained classification model to obtain the category to which each remaining user data after screening belongs. The data form here includes the dimension of the feature data and the representation method of the feature data.
[0105] In one embodiment, step 212 includes: for each user data remaining after screening, retaining a preset number of feature data of the dimension with the largest weight, and obtaining user sub-data for using the classification model; classifying the user sub-data through the trained classification model to obtain the category to which each user sub-data belongs.
[0106] refer to Figure 5 , which shows a schematic diagram of the classification effect of classifying the remaining user data after screening in one embodiment. Figure 5 , the classification model uses the feature data of the three dimensions with the largest weights of the seed user data as training samples during training, and the seed user data is clustered using the feature data of the three dimensions with the largest weights. When the classification model is used for non-seed user data, the feature data of the three dimensions with the largest weights also needs to be used. The three-dimensional coordinates of the figure represent the dimensions of different feature data. The classification results output by the classification model also include the eight groups obtained by clustering the seed user data: important value users, important retention users, important development users, important retention users, general value users, general retention users, general development users, and general retention users.
[0107] Continue to refer Figure 3 After the computer device clusters the seed user data to obtain the seed user data with categories, the seed user data with categories can be used to train the classification model. After the classification model training is completed, the classification model is used to classify the remaining non-seed user data to obtain the categories described by each non-seed user data. It can be understood that the classification of user data can be used to group users, thereby obtaining the user grouping result.
[0108] Step 214: Push information according to the category to which each user data belongs.
[0109] Specifically, after obtaining the category to which each user data belongs, the computer device can push targeted information according to the category to which each user data belongs. Pushing information according to the category to which each user data belongs can specifically push information to the user corresponding to each user data according to the category to which each user data belongs. Different data can be pushed to users belonging to different categories.
[0110] The pushed information may be products, news, videos, resources or audio, etc.
[0111] In one embodiment, step 214 includes: acquiring push information corresponding to each category; and pushing push information corresponding to the category to which each user data belongs to a user terminal corresponding to each user data.
[0112] Among them, push information is information pushed to users in specific application scenarios, such as video messages in video promotion scenarios, or product messages in product promotion scenarios, etc. By classifying users, different business processes are performed based on different user value groups, such as using different red envelope rewards.
[0113] It is understandable that high-dimensional data distance calculation will cause memory explosion, insufficient resources, and serious time consumption. Whether it is Euclidean distance, Manhattan distance, Chebyshev distance, etc., there are similar situations. When facing a huge number of high-dimensional user data sets, it is impossible to load all the data into memory for pairwise calculations like traditional clustering algorithms K-Means, DBSCA, and hierarchical clustering. In this application, a divide-and-conquer processing method for big data is cleverly proposed. By screening a small amount of seed user data, the seed user data is clustered to obtain seed user data with categories, and then the seed user data with categories is used to supervise the classification model to obtain a more accurate classification model. The classification model obtained by training is then used to classify a large amount of remaining non-seed user data.
[0114] Moreover, on the one hand, the present application sets different weight parameters for different features, and quantifies the similarity between user data, i.e., user distance, in a weighted manner, so as to classify user data in specific scenarios well, and thus better distinguish the value of users in specific scenarios; and the weight parameters are obtained by learning a large amount of user data based on the machine learning model, which is objective and accurate. On the other hand, the classification model is designed so that the classification model can learn the user's feature data more accurately, and the classification model after learning can accurately classify user data, so as to identify the user's value well.
[0115] After obtaining the categories of user data, we can obtain the user's value grouping, and then perform different processing according to the user's value grouping. For example, we can carry out different marketing activities according to the user's value grouping, use different red envelope rewards for different users, high-value users will receive higher rewards, low-value users will be pushed different new products, and greater discounts will be given to activate users' purchasing power, and other product solutions.
[0116] After obtaining a large amount of multi-dimensional user data, the above user data processing method selects a small amount of user data as the target for clustering and classification, obtains the target user data with categories, and then uses the target user data with categories to train the classification model, and then uses the trained classification model to classify the remaining user data after screening. In this way, on the one hand, when clustering, different weight parameters are set for feature data of different dimensions, which can well process discrete user data in a targeted manner, and then more accurately cluster user data according to the importance of feature data of different dimensions, thereby improving user classification accuracy; and the classification model trained based on this part of the target user data with categories can also accurately and effectively classify the remaining user data after screening; on the other hand, clustering only part of the user data can also avoid the amount of calculation and possible insufficient computing resources caused by clustering the entire amount of data; in addition, after obtaining the category to which the user data belongs, information push can be more accurate and targeted.
[0117] In one embodiment, obtaining the corresponding weight of each dimension includes: obtaining a weight vector output by a trained sorting model; wherein the sorting model is obtained through supervised training of user data samples with training labels, and the influence of feature data of multiple dimensions in the user sample data is sorted during the supervised training; the weight vector includes the weights of the feature data of multiple dimensions.
[0118] Among them, the sorting model can be a tree model, SVM (Support Vector Machine) model, LR (Logistic Regression) model, neural network model and other types of models, or the sorting model can also be a model combining multiple types of algorithms such as tree model, SVM (Support Vector Machine) model, LR (Logistic Regression) model, neural network model, etc., or the sorting model can also be a SVM model, LR model or neural network model combined with SHAP (SHa pley Additive exPlanations, model interpretability), Permutation and other interpretability algorithms to obtain, the embodiment of the present application is not limited to this.
[0119] Specifically, when the sorting model is a tree model, the tree model may include a GBDT (Gradient Boosting Decision Tree) model, a LightGBM (Light Gradient Boosting Machine) model, or a combination of LightGBM and GBDT. The LightGBM model uses a histogram algorithm for feature selection, converting multiple continuous values into a preset number of discrete values in a histogram, with high computational efficiency. The LightGBM model abandons the layer-by-layer growth strategy and adopts a leaf-by-leaf growth strategy. When the number of splits is the same, it can reduce unnecessary searches and splits, thereby improving the accuracy of the model.
[0120] The weight vector output by the ranking model can be specifically δ={δ 1 ,δ 2 ,δ 3 ,δ 4 ,δ 5 ,…,δ n}. Among them, δ 1 It can represent the weight corresponding to the feature data of dimension 1, δ 2 It can represent the weight corresponding to the feature data of dimension 2, δ 3 It can represent the weight corresponding to the feature data of dimension 3, and so on, where n is the number of dimensions.
[0121] In one embodiment, the training steps of the sorting model include: obtaining user data samples and training labels corresponding to the user data samples; the user data samples include feature data of multiple dimensions; predicting the user data samples through the sorting model to obtain prediction results; optimizing the sorting model based on the difference between the prediction results and the training labels; sorting the degree of influence of the feature data of multiple dimensions in the user sample data on the prediction through the sorting model, and outputting the sorting results.
[0122] Among them, the training labels of the user data samples used by the training sorting model can be related to the importance of the user data samples. It can be understood that setting weights for feature data of different dimensions is essentially to sort the influence of feature data of different dimensions on the classification of user data. The feature data of a dimension that has a greater influence on the importance classification of user data has a higher weight corresponding to the feature data of that dimension.
[0123] Specifically, the computer device can divide the user data samples into two categories, one category is considered to be important user data samples as positive samples, and the other category is considered to be unimportant user data samples as negative samples. The sorting model is supervised by positive samples and negative samples. The output of the sorting model includes two types of data, one is the classification prediction of the user data samples, and the other is the importance ranking of the feature data of each different dimension. In the continuous iterative training, the sorting model learns to pay attention to the feature data of the dimensions that are more useful for improving the accuracy of the classification results, and then can achieve the ranking of the degree of influence of the feature data of each different dimension on the classification of user data. After training, the sorting model will output a stable sorting result. The sorting result includes the weights corresponding to the feature data of each dimension, that is, the weight vector.
[0124] In one embodiment, the user data processing method also includes: obtaining the vector element with the largest value and the vector element with the smallest value in the weight vector; determining the difference between the vector element with the largest value and the vector element with the smallest value; and normalizing each vector element in the weight vector according to the vector element with the smallest value and the difference.
[0125] Specifically, the ranking model output weight vector is δ = {δ 1 ,δ 2 ,δ 3 ,δ 4 ,δ 5 ,…,δ n}, obtain the vector element δ_max with the largest value and the vector element δ_min with the smallest value in the weight vector, determine the difference δ_max-δ_min between the vector element with the largest value and the vector element with the smallest value, and normalize each vector element in the weight vector according to the vector element with the smallest value and the difference:
[0126] Normal_δ=(δ-δ min ) / (δ_max-δ_min) (2)
[0127] Furthermore, the computer device can use the normalized weights to calculate the similarity of the user data. Normalization is a dimensionless processing method that converts the absolute value of the physical system value into a relative value relationship, which can simplify the calculation and reduce the value.
[0128] In the above embodiment, a large amount of user data is learned according to the machine learning model to obtain weight parameters corresponding to the feature data of each dimension, which is objective, reliable and highly accurate.
[0129] In one embodiment, the classification model includes multiple classification substructures; each classification substructure includes an attention network structure; the processing steps of the attention network structure include: assigning weights to the feature vectors input into the attention network structure respectively to obtain a key vector, a request vector and a value vector; processing the key vector, the request vector and the value vector through the attention network structure to obtain a processing result.
[0130] Specifically, the present application creatively designs a model structure of a classification model, which includes multiple classification substructures, each of which includes an attention network structure. Among them, the attention network structure is a network structure based on the attention mechanism. The attention mechanism is a way of building a model based on the dependency relationship between the hidden states of the encoder and the decoder. Multiple attention network structures can be used to capture feature data in different representation spaces respectively. Each attention network structure will obtain feature data in a representation space after processing the data. In this way, the attention network structure can make the classification model pay more attention to important features, learn more useful information for the classification purpose of the classification model, and focus on the feature information of certain dimensions in the multidimensional features.
[0131] The input of the classification model is the feature vector (User Feature Embedding) obtained by vectorizing the discrete user data through the Embedding operation. When processing discrete user data, the Embedding operation can be performed only on the feature data of a preset number of dimensions with higher weights. For example, the Embedding operation is performed only on the feature data of the three dimensions with the highest weights to obtain a three-dimensional feature vector.
[0132] After the feature vector (User Feature Embedding) obtained by vectorizing discrete user data is input into the classification model, it is processed in turn through each classification substructure of the classification model to obtain a new feature vector which is input into the next classification substructure for processing until the output layer of the classification model outputs the classification result.
[0133] The output of the classification model is the classification result of the user data. For example, if there are 8 classification categories, the output of the classification model can be the corresponding probabilities of (0, 1, 2, 3, 4, 5, 6, 7), and the category with the highest probability is the category to which the user data belongs.
[0134] For example, see Figure 6, which shows a schematic diagram of the structure of a classification model in an embodiment. As can be seen from the figure, the classification model (SA-NET) includes multiple classification substructures, each of which includes an attention network structure (SA, SP-Attention). Adjacent classification substructures can be directly connected or transitionally connected through a pooling layer. The classification model finally outputs the classification result through a regression layer. Among them, the pooling operation of the pooling layer is such as maximum pooling (MaxPooling) or global maximum pooling (Global Max Pooling). The regression layer is such as a Softmax layer.
[0135] refer to Figure 7 , which shows a schematic diagram of the structure of an attention network structure (SA, SP-Attention) in one embodiment. As can be seen from the figure, the feature vector z input to the attention network structure is assigned weights respectively, and the key vector K(z)=w is obtained. 2 z 2 , request vector Q(z) = w 1 z 1 Sum vector V(z)=w 3 z 3 , z=z 1 =z 2 =z 3 Among them, w 1 、w 2 and w 3 It is the weight parameter that the classification model needs to learn during training.
[0136] The output of the attention network structure is shown as follows:
[0137]
[0138] Among them, K T is the transpose of K, Represents the Huffman distance between Q and K.
[0139] In one embodiment, the feature vector z input to the attention network structure is weighted respectively to obtain the key vector K(z)=w 2 z 2 , request vector Q(z) = w 1 z 1 Sum vector V(z)=w 3 z 3 When 1 、z 2 and z + It can also be feature data of different dimensions of z.
[0140] In one embodiment, a key vector, a request vector, and a value vector are processed through an attention network structure to obtain a processing result, including: mapping the key vector, the request vector, and the value vector into more than one set of key vectors, request vectors, and value vectors through a nonlinear activation function layer of the attention network structure; processing more than one set of key vectors, request vectors, and value vectors respectively through multiple attention mechanism layers of the attention network structure to obtain intermediate results; and processing the intermediate results in sequence through a concatenation layer and a convolution layer of the attention network structure to obtain a processing result.
[0141] Continue to refer Figure 7 , then we get the key vector K(z)=w 2 z 2 , request vector Q(z) = w 1 z 1 Sum vector V(z)=w 3 z 3 Afterwards, the key vector, request vector, and value vector are mapped into more than one set of key vectors, request vectors, and value vectors through the nonlinear activation function layer of the attention network structure, such as the prelu activation function layer. For example, the mapping is into h groups, and each set of key vectors, request vectors, and value vectors are the same, that is, K(z) = w 2 z 2 Q(z) = w 1 z 1 and V(z)=w 3 z 3 Then, through multiple attention mechanism layers of the attention network structure, such as Scaled Dot-Product Attention, more than one set of key vectors, request vectors, and value vectors are processed respectively to obtain intermediate results; and then the concatenation layer (Concatenate) and convolution layer (Convolution) of the attention network structure process the intermediate results in turn to obtain the processing result SA (Q, K, V).
[0142] In one embodiment, the classification substructure also includes a convolutional network structure and a batch normalization network structure; the processing steps of the classification substructure include: performing a convolution operation on the data input into the classification substructure through the convolutional network structure, and outputting the convolution operation result to the batch normalization network structure; performing distribution adjustment on the convolution operation result through the batch normalization network structure, and outputting the adjustment result to the attention network structure.
[0143] The convolutional network structure is a network structure used to perform convolution operations on data. The number of convolutional layers in the convolutional network structures included in different classification substructures may be the same or different.
[0144] Batch Normalization (BN) is used to keep the input of each layer of the network structure in the same distribution during the model training process. Batch Normalization can speed up the training of the model and improve the generalization ability of the model.
[0145] Continue to refer Figure 6 , we can see that the classification network includes 13 classification substructures. The 5th classification substructure and the 6th classification substructure, as well as the 10th classification substructure and the 11th classification substructure are connected through Max Pooling transition, and the other classification substructures are directly connected. The 13th classification substructure outputs the classification result after the global maximum pooling operation (GlobalMax Pooling), inner product operation (Inner Product) and normalization operation (Softmax).
[0146] In a specific scenario, the convolution layer of the first classification substructure includes 66 3x3 convolution kernels; the convolution layers of the second, third, and fourth classification substructures include 128 3x3 convolution kernels; the convolution layers of the fifth, sixth, seventh, eighth, and ninth classification substructures include 192 3x3 convolution kernels; the convolution layers of the tenth and eleventh classification substructures include 288 3x3 convolution kernels; the convolution layer of the twelfth classification substructure includes 355 3x3 convolution kernels; the convolution layer of the thirteenth classification substructure includes 432 3x3 convolution kernels. Of course, the number and size of the convolution kernels of the convolution layers of the classification substructures can also be other situations.
[0147] It should be understood that, although the various steps in the flow chart of the above-described embodiment are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flow chart of the above-described embodiment may include a plurality of steps or a plurality of stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.
[0148] In one embodiment, Figure 8As shown, a user data processing device is provided. The user data processing device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The user data processing device specifically includes: an acquisition module 801, a first screening module 802, a determination module 803, a clustering module 804, a training module 805, a second screening module 806 and an application module 807.
[0149] in,
[0150] The acquisition module 801 is used to acquire multiple user data; each user data includes feature data of multiple dimensions;
[0151] A first screening module 802 is used to screen out a plurality of target user data from the user data;
[0152] A determination module 803, configured to determine the similarity between each target user data according to the feature data of each target user data in each dimension and the corresponding weight of each dimension;
[0153] A clustering module 804 is used to cluster the target user data based on the similarity between the target user data to obtain the category to which each target user data belongs;
[0154] A training module 805 is used to train a classification model according to target user data and the category to which the target user data belongs;
[0155] The second screening module 806 is used to classify the user data remaining after screening by using the trained classification model to obtain the category to which each user data remaining after screening belongs;
[0156] The application module 807 is used to push information according to the category to which each user data belongs.
[0157] In one embodiment, the determination module 803 is also used to vectorize each target user data to obtain a feature vector of each target user data; the feature vector includes feature values of multiple dimensions; obtain the corresponding weight of each dimension; for the feature vectors of any two target user data, calculate the distance of the feature values of each dimension according to the weight of each dimension to obtain the similarity between any two target user data.
[0158] In one embodiment, the determination module 803 is also used to obtain the corresponding weight of each dimension; for each target user data, retain a preset number of feature data of the dimension with the largest weight to obtain the target user sub-data for clustering; for the feature vectors of any two target user sub-data, calculate the distance of the feature values of each dimension according to the weight of each dimension to obtain the similarity between any two target user sub-data.
[0159] In one embodiment, the determination module 803 is also used to obtain the weight vector output by the trained sorting model; wherein the sorting model is obtained through supervised training of user data samples with training labels, and the influence of feature data of multiple dimensions in the user sample data is sorted during the supervised training process; the weight vector includes the weights of feature data of multiple dimensions.
[0160] In one embodiment, the training module 805 is also used to obtain user data samples and training labels corresponding to the user data samples; the user data samples include feature data of multiple dimensions; the user data samples are predicted by a sorting model to obtain prediction results; the sorting model is optimized according to the difference between the prediction results and the training labels; the degree of influence of the feature data of multiple dimensions in the user sample data on the prediction is sorted by the sorting model, and the sorting results are output.
[0161] In one embodiment, the determination module 803 is also used to obtain the vector element with the largest value and the vector element with the smallest value in the weight vector; determine the difference between the vector element with the largest value and the vector element with the smallest value; and normalize each vector element in the weight vector according to the vector element with the smallest value and the difference.
[0162] In one embodiment, the training module 805 is further used to retain a preset number of feature data of the dimension with the largest weight for each target user data, and obtain target user sub-data for training the classification model; train the classification model according to the target user sub-data and the category to which the target user sub-data belongs. The second screening module 806 is further used to retain a preset number of feature data of the dimension with the largest weight for each user data remaining after screening, and obtain user sub-data for using the classification model; classify the user sub-data through the classification model obtained by training, and obtain the category to which each user sub-data belongs.
[0163] In one embodiment, the classification model includes multiple classification substructures; each classification substructure includes an attention network structure. The second screening module 806 is also used to assign weights to the feature vectors input into the attention network structure to obtain a key vector, a request vector and a value vector; the key vector, the request vector and the value vector are processed by the attention network structure to obtain a processing result.
[0164] In one embodiment, the second screening module 806 is also used to map the key vector, request vector and value vector into more than one set of key vectors, request vectors and value vectors through the nonlinear activation function layer of the attention network structure; process more than one set of key vectors, request vectors and value vectors respectively through multiple attention mechanism layers of the attention network structure to obtain intermediate results; and process the intermediate results in turn through the splicing layer and convolution layer of the attention network structure to obtain processing results.
[0165] In one embodiment, the classification substructure further includes a convolutional network structure and a batch normalization network structure. The second screening module 806 is further used to perform a convolution operation on the data input into the classification substructure through the convolutional network structure, and output the convolution operation result to the batch normalization network structure; perform distribution adjustment on the convolution operation result through the batch normalization network structure, and output the adjustment result to the attention network structure.
[0166] In one embodiment, the first screening module 802 is further used to clean the user data; divide the cleaned user data into multiple levels to obtain multiple user data subsets; and randomly select user data from each minimum number of user data subsets as target user data.
[0167] In one embodiment, the application module 807 is further used to obtain push information corresponding to each category; and push push information corresponding to the category to which each user data belongs to to the user terminal corresponding to each user data.
[0168] After acquiring a large amount of multi-dimensional user data, the above user data processing device selects a small amount of user data as a target for clustering and classification, obtains target user data with categories, and then uses the target user data with categories to train a classification model, and then uses the trained classification model to classify the remaining user data after screening. In this way, on the one hand, when clustering, different weight parameters are set for feature data of different dimensions, which can be used to carry out targeted processing of discrete user data, and then more accurately cluster user data according to the importance of feature data of different dimensions, thereby improving user classification accuracy; and the classification model trained based on this part of target user data with categories can also accurately and effectively classify the remaining user data after screening; on the other hand, clustering only part of the user data can also avoid the amount of calculation and possible insufficient computing resources caused by clustering the entire amount of data; in addition, after obtaining the category to which the user data belongs, information push can be more accurate and targeted.
[0169] For the specific definition of the user data processing device, please refer to the definition of the user data processing method above, which will not be repeated here. Each module in the above user data processing device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0170] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig. 9As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store user data, etc. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a user data processing method executed by a server is implemented.
[0171] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0172] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0173] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0174] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the above-mentioned method embodiments.
[0175] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0176] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0177] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for processing user data, It is characterized in that The method comprises: Acquire multiple user data; each of the user data includes feature data of multiple dimensions; Filtering out a plurality of target user data from the user data; Get the corresponding weight of each dimension; For each target user data, retain a preset number of feature data of dimensions with the largest weight, and obtain target user sub-data for clustering; For any two feature vectors of the target user sub-data, the feature values of each dimension are calculated according to the weight of each dimension to obtain the similarity between any two target user sub-data; Clustering the target user data based on the similarity between each target user data to obtain the category to which each target user data belongs; Training a classification model according to the target user data and the category to which the target user data belongs; Classify the user data remaining after screening by the classification model obtained through training, and obtain the category to which each user data remaining after screening belongs; Information is pushed according to the category to which each of the user data belongs.
2. The method according to claim 1, It is characterized in that The obtaining of the corresponding weight of each dimension includes: Get the weight vector output by the trained sorting model; The ranking model is obtained through supervised training of user data samples with training labels, and the influence of feature data of multiple dimensions in the user sample data is ranked during the supervised training process; the weight vector includes the weights of the feature data of the multiple dimensions.
3. The method according to claim 2, It is characterized in that The training steps of the sorting model include: Obtaining a user data sample and a training label corresponding to the user data sample; the user data sample includes feature data of multiple dimensions; Predicting the user data sample by using the sorting model to obtain a prediction result; Optimizing the ranking model according to the difference between the prediction result and the training label; The ranking model is used to sort the degree of influence of feature data of multiple dimensions in the user sample data on the prediction, and the ranking result is output.
4. The method according to claim 2, It is characterized in that The method further comprises: Obtaining the vector element with the largest value and the vector element with the smallest value in the weight vector; Determine the difference between the vector element with the largest value and the vector element with the smallest value; Normalizing each vector element in the weight vector according to the vector element with the smallest value and the difference.
5. The method according to claim 1, It is characterized in that The training of the classification model according to the target user data and the category to which the target user data belongs includes: For each target user data, retain a preset number of feature data of dimensions with the largest weight, and obtain target user sub-data for training a classification model; Training a classification model according to the target user sub-data and the category to which the target user sub-data belongs; The classification model obtained through training classifies the user data remaining after screening to obtain the category to which each user data remaining after screening belongs, including: For each user data remaining after screening, retain the preset number of feature data of the dimension with the largest weight, and obtain user sub-data for using the classification model; The user sub-data are classified by the classification model obtained through training to obtain the category to which each user sub-data belongs.
6. The method according to claim 1, It is characterized in that The classification model includes multiple classification substructures; each classification substructure includes an attention network structure; The processing steps of the attention network structure include: Assigning weights to the feature vectors input into the attention network structure respectively to obtain a key vector, a request vector, and a value vector; The key vector, the request vector and the value vector are processed by the attention network structure to obtain a processing result.
7. The method according to claim 6, It is characterized in that The processing of the key vector, the request vector, and the value vector by the attention network structure to obtain a processing result includes: Mapping the key vector, the request vector, and the value vector into more than one set of key vectors, request vectors, and value vectors through a nonlinear activation function layer of the attention network structure; Processing more than one set of key vectors, request vectors, and value vectors respectively through multiple attention mechanism layers of the attention network structure to obtain intermediate results; The intermediate results are processed in sequence through the concatenation layer and the convolution layer of the attention network structure to obtain the processing result.
8. The method according to claim 6, It is characterized in that The classification substructure also includes a convolutional network structure and a batch normalization network structure; The processing steps of the classification substructure include: Performing a convolution operation on the data input into the classification substructure through the convolution network structure, and outputting the convolution operation result to the batch normalization network structure; The convolution operation result is distributedly adjusted through the batch normalization network structure, and the adjusted result is output to the attention network structure.
9. The method according to claim 1, It is characterized in that The step of filtering out a plurality of target user data from the user data comprises: Performing data cleaning on the user data; Divide the cleaned user data into multiple levels to obtain multiple user data subsets; User data is randomly selected from each minimum number of user data subsets as target user data.
10. The method according to claim 1, It is characterized in that The information push according to the category to which each user data belongs includes: Get the corresponding push information of each category; Push information corresponding to the category to which each of the user data belongs is pushed to the user terminal corresponding to each of the user data.
11. A user data processing device, It is characterized in that The device comprises: An acquisition module, used to acquire multiple user data; each of the user data includes feature data of multiple dimensions; A first screening module, used to screen out a plurality of target user data from the user data; A determination module is used to obtain the corresponding weight of each dimension; for each target user data, retain a preset number of feature data of the dimension with the largest weight, and obtain the target user sub-data for clustering; for the feature vectors of any two target user sub-data, calculate the distance of the feature values of each dimension according to the weight of each dimension, and obtain the similarity between any two target user sub-data; A clustering module, used for clustering the target user data based on the similarity between the target user data to obtain the category to which the target user data belongs; A training module, used for training a classification model according to the target user data and the category to which the target user data belongs; A second screening module is used to classify the user data remaining after screening by using the classification model obtained through training, so as to obtain the category to which each user data remaining after screening belongs; The application module is used to push information according to the category to which each user data belongs.
12. The device according to claim 11, It is characterized in that The determination module is also used to obtain the weight vector output by the trained sorting model; wherein the sorting model is obtained through supervised training of user data samples with training labels, and the influence of feature data of multiple dimensions in the user sample data is sorted during the supervised training process; the weight vector includes the weights of the feature data of the multiple dimensions.
13. The device according to claim 12, It is characterized in that The determination module is also used to obtain user data samples and training labels corresponding to the user data samples; the user data samples include feature data of multiple dimensions; the user data samples are predicted by the sorting model to obtain prediction results; the sorting model is optimized according to the difference between the prediction results and the training labels; the degree of influence of the feature data of multiple dimensions in the user sample data on the prediction is sorted by the sorting model, and the sorting results are output.
14. The device according to claim 12, It is characterized in that The determination module is also used to obtain the vector element with the largest value and the vector element with the smallest value in the weight vector; determine the difference between the vector element with the largest value and the vector element with the smallest value; and normalize each vector element in the weight vector according to the vector element with the smallest value and the difference.
15. The device according to claim 11, It is characterized in that The training module is further used to retain a preset number of feature data of the dimension with the largest weight for each target user data, and obtain target user sub-data for training the classification model; and train the classification model according to the target user sub-data and the category to which the target user sub-data belongs; The second screening module is also used to retain the preset number of feature data of the dimension with the largest weight for each user data remaining after screening, and obtain user sub-data for using the classification model; classify the user sub-data through the classification model obtained by training to obtain the category to which each user sub-data belongs.
16. The device according to claim 11, It is characterized in that The classification model includes multiple classification substructures; each classification substructure includes an attention network structure; The second screening module is also used to assign weights to the feature vectors input into the attention network structure respectively to obtain a key vector, a request vector and a value vector; the key vector, the request vector and the value vector are processed through the attention network structure to obtain a processing result.
17. The device according to claim 16, It is characterized in that The second screening module is further used to map the key vector, the request vector and the value vector into more than one set of key vectors, request vectors and value vectors through a nonlinear activation function layer of the attention network structure; Through the multiple attention mechanism layers of the attention network structure, more than one set of key vectors, request vectors and value vectors are processed respectively to obtain intermediate results; the intermediate results are processed in turn through the concatenation layer and the convolution layer of the attention network structure to obtain the processing results.
18. The device according to claim 16, It is characterized in that The classification substructure also includes a convolutional network structure and a batch normalization network structure; The second screening module is also used to perform a convolution operation on the data input into the classification substructure through the convolution network structure, and output the convolution operation result to the batch normalization network structure; perform distribution adjustment on the convolution operation result through the batch normalization network structure, and output the adjustment result to the attention network structure.
19. The device according to claim 11, It is characterized in that The first screening module is also used to clean the user data; divide the cleaned user data into multiple levels to obtain multiple user data subsets; and randomly select user data from each minimum number of user data subsets as target user data.
20. The device according to claim 11, It is characterized in that The application module is also used to obtain push information corresponding to each category; and push push information corresponding to the category to which each user data belongs to to the user terminal corresponding to each user data.
21. A computer device comprising a memory and a processor, wherein the memory stores a computer program. It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.
22. A computer-readable storage medium storing a computer program, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
23. A computer program product comprising computer instructions, It is characterized in that When the computer instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Similar crowd expansion method and device and electronic equipment
CN109903086A