A data processing method, device, computer and readable storage medium
By training a target prediction model, based on candidate behavioral indicators of social applications and user historical data, the probability of users performing business operations is predicted, which solves the problem that self-service cannot handle customer complaints and improves the accuracy and efficiency of data processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2021-03-26
- Publication Date
- 2026-05-22
AI Technical Summary
When users encounter application problems, the existing self-service technology cannot effectively solve them, resulting in a large number of application complaints being pushed to human customer service, increasing the burden on human processing and reducing efficiency.
By acquiring candidate behavioral metrics from social applications and historical behavioral data from sample users, a target prediction model is trained to predict the probability of users performing business operations. Based on the prediction results, a decision is made on whether to push the data to business personnel or send a default reply, thereby reducing unnecessary manual intervention.
It improved the accuracy and efficiency of data processing, reduced the workload of human customer service, and ensured the timely handling of customer complaints.
Smart Images

Figure CN115130711B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a data processing method, apparatus, computer, and readable storage medium. Background Technology
[0002] With the development of the internet, the number of users engaging in app transactions is increasing. When making app transactions, users may encounter problems and, unsure how to resolve them, will seek help from app customer service, resulting in an app customer complaint. When an app customer complaint is generated, the computer system typically first checks if the problem can be resolved using self-service tools, such as user authentication (e.g., facial recognition). If authentication is successful, self-service is provided; if authentication fails, the computer system can perform multiple authentication attempts to instruct the user to resolve the problem independently. However, in processing app customer complaints, self-service may not be able to resolve the issue, thus requiring the complaint to be forwarded to human customer service. If all user complaints were forwarded to human customer service, it would generate a massive amount of data requiring human processing. This not only places a huge workload on human customer service but also leads to low efficiency due to the sheer volume of data. Therefore, often, an app customer complaint is only forwarded to human customer service after multiple attempts have been made, resulting in delayed processing and reduced data processing efficiency. Summary of the Invention
[0003] This application provides a data processing method, apparatus, computer, and readable storage medium, which can improve the accuracy and efficiency of data processing.
[0004] One embodiment of this application provides a data processing method, the method comprising:
[0005] Obtain M candidate behavioral metrics associated with social applications, and obtain historical behavioral data of N sample users under the M candidate behavioral metrics; N is a positive integer, M is a positive integer; social applications include the target business;
[0006] Obtain behavioral execution tags for N sample users targeting the target business. Based on the behavioral execution tags and historical behavioral data of the N sample users, determine the execution impact of M candidate behavioral indicators on the target business.
[0007] Based on the impact of M candidate behavioral indicators on the execution of the target business, the target behavioral indicator is determined from the M candidate behavioral indicators.
[0008] In historical behavior data, target behavior data of N sample users under target behavior indicators are obtained. The initial prediction model is trained using the behavior execution labels and target behavior data corresponding to the N sample users to obtain the target prediction model. The target prediction model is used to predict the probability of the target user executing the target business.
[0009] One embodiment of this application provides a data processing method, the method comprising:
[0010] The process involves acquiring business inquiry information submitted by target users within a social application, and obtaining predictable behavioral data of these users under target behavioral metrics. The social application includes the target business. The target behavioral metrics are determined from M candidate behavioral metrics based on their respective impact on the target business. These M candidate behavioral metrics are associated with the social application. The impact of each of the M candidate behavioral metrics on the target business is determined based on the historical behavioral data of N sample users under the M candidate behavioral metrics, and the behavioral execution tags of these N sample users for the target business. M and N are positive integers.
[0011] The behavioral data to be predicted is input into the target prediction model. In the target prediction model, the target user's target execution probability for the target business is predicted based on the behavioral data to be predicted. The target execution probability is used to represent the probability that the target user will execute the target business.
[0012] The method also includes:
[0013] If the target execution probability is greater than or equal to the business execution threshold, then business consultation information is sent to the business device associated with the business personnel so that the business device can display the business consultation information to the business personnel.
[0014] If the target execution probability is less than the business execution threshold, then the default reply information corresponding to the business consultation information is obtained and sent to the user device associated with the target user.
[0015] One embodiment of this application provides a data processing apparatus, the apparatus comprising:
[0016] The metrics acquisition module is used to acquire M candidate behavioral metrics associated with social applications and to acquire historical behavioral data of N sample users under the M candidate behavioral metrics; N is a positive integer, M is a positive integer; social applications include target businesses;
[0017] The tag acquisition module is used to acquire the action tags of N sample users for the target business.
[0018] The impact determination module is used to determine the impact of M candidate behavioral indicators on the target business based on the behavioral execution tags and historical behavioral data of N sample users.
[0019] The indicator selection module is used to determine the target behavior indicator from the M candidate behavior indicators based on their respective impact on the target business.
[0020] The data acquisition module is used to acquire target behavior data of N sample users under the target behavior index from historical behavior data;
[0021] The model generation module is used to train the initial prediction model using the behavior execution labels and target behavior data corresponding to N sample users to obtain the target prediction model; the target prediction model is used to predict the probability of the target user executing the target business.
[0022] The behavior execution label includes positive sample behavior execution labels and negative sample behavior execution labels;
[0023] This tag acquisition module is specifically used for:
[0024] Obtain the historical execution counts of N sample users for the target business. Assign a positive sample behavior execution label to the first sample user among the N sample users based on the historical execution counts, and assign a negative sample behavior execution label to the second sample user among the N sample users based on the historical execution counts. The first sample user is the sample user whose historical execution count for the target business is greater than or equal to the number of times the threshold is applied. The second sample user is the sample user whose historical execution count for the target business is less than the number of times the threshold is applied.
[0025] Among them, the behavior execution label includes positive sample behavior execution label and negative sample behavior execution label; M candidate behavior indicators include the i-th candidate behavior indicator; i is a positive integer, i is less than or equal to M; the number of historical behavior data of N sample users under the i-th candidate behavior indicator is N;
[0026] The impact determination module includes:
[0027] The data clustering unit is used to cluster the N historical behavior data corresponding to the i-th candidate behavior indicator to obtain h behavior data sets; h is a positive integer, h is less than or equal to N; the h behavior data sets include the j-th behavior data set; j is a positive integer; j is less than or equal to h;
[0028] The quantity acquisition unit is used to acquire the number of positive samples of historical behavior data with positive behavior execution labels in the j-th behavior data set, and to acquire the number of negative samples of historical behavior data with negative behavior execution labels in the j-th behavior data set.
[0029] The information determination unit is used to determine the set information content of the j-th behavior data set based on the number of positive samples corresponding to the j-th behavior data set and the number of negative samples corresponding to the j-th behavior data set.
[0030] The impact determination unit is used to sum the set information corresponding to each of the h behavioral data sets to obtain the execution impact of the i-th candidate behavioral indicator on the target business.
[0031] The information determining unit includes:
[0032] The total number determination subunit is used to determine the total number of positive samples based on the number of positive samples corresponding to each of the h behavioral data sets, and to determine the total number of negative samples based on the number of negative samples corresponding to each of the h behavioral data sets.
[0033] The weight determination subunit is used to obtain the positive sample ratio between the number of positive samples corresponding to the j-th behavior data set and the total number of positive samples, and to obtain the negative sample ratio between the number of negative samples corresponding to the j-th behavior data set and the total number of negative samples, and to determine the set weight of the j-th behavior data set based on the positive sample ratio and the negative sample ratio.
[0034] The weighted processing subunit is used to obtain the difference between the proportion of positive samples and the proportion of negative samples. Based on the set weight, the difference between the proportion of positive samples and the proportion of negative samples is weighted to obtain the set information of the j-th row data set.
[0035] The indicator selection module includes:
[0036] The combination acquisition unit is used to obtain a first candidate behavioral indicator from M candidate behavioral indicators, combine the first candidate behavioral indicators into L combined behavioral indicators, and obtain the combined impact degree of each of the L combined behavioral indicators on the target business. The first candidate behavioral indicator refers to the candidate behavioral indicator whose execution impact degree belongs to the first impact degree range. Each combined behavioral indicator is obtained by combining at least two candidate behavioral indicators. The combined impact degree is used to represent the execution impact degree of the at least two candidate behavioral indicators included in the corresponding combined behavioral indicator on the target business. L is a positive integer.
[0037] The indicator determination unit is used to obtain the second candidate behavior indicator from M candidate behavior indicators, and to determine the combined behavior indicators whose combined influence degree belongs to the second influence degree range from L combined behavior indicators, as well as the second candidate behavior indicators, as the target behavior indicators; the second candidate behavior indicator refers to the candidate behavior indicator whose execution influence degree belongs to the second influence degree range.
[0038] The data acquisition module includes:
[0039] The data lookup unit is used to find the target behavior data of N sample users under the target behavior indicator in historical behavior data;
[0040] The missing data acquisition unit is used to determine the data missing rate of N sample users based on the ratio between the missing sample users and the N sample users if there are missing sample users for the target behavior indicator among the N sample users.
[0041] The distribution acquisition unit is used to acquire the behavioral data distribution of target behavioral data for N sample users (excluding missing sample users) under the target behavioral indicator if the data missing rate is less than the data filling threshold.
[0042] The data determination unit is used to determine the target behavior data of missing sample users under the target behavior indicator based on the distribution of behavior data, and to obtain the target behavior data of N sample users under the target behavior indicator.
[0043] The model generation module includes:
[0044] The type acquisition unit is used to acquire the indicator type of the target behavior indicator;
[0045] The label encoding unit is used to encode the target behavior data corresponding to N sample users respectively if the indicator type is an ordered indicator type, generate N target behavior features, input the N target behavior features into the initial prediction model for prediction, obtain the sample execution probability corresponding to the N target behavior features respectively, and train the initial prediction model based on the sample execution probability corresponding to the N target behavior features and the error value between the behavior execution labels of the N sample users to obtain the target prediction model.
[0046] One-hot encoding unit is used to perform one-hot encoding on the target behavior data corresponding to N sample users if the indicator type is unordered, generating N target behavior features. The N target behavior features are then input into the initial prediction model for prediction to obtain the sample execution probability corresponding to each of the N target behavior features. Based on the sample execution probabilities corresponding to the N target behavior features and the error value between the behavior execution labels of the N sample users, the initial prediction model is trained to obtain the target prediction model.
[0047] The model generation module includes:
[0048] The model creation unit is used to create S initial prediction models; the number of decision layers in each of the S initial prediction models is different; the S initial prediction models include initial prediction model d; S is a positive integer; initial prediction model d is an initial prediction model containing d decision layers; d is a positive integer.
[0049] The model training unit is used to train the initial prediction model d using the target behavior data and behavior execution labels corresponding to N sample users, so as to obtain the training prediction model corresponding to the initial prediction model d.
[0050] The accuracy acquisition unit is used to acquire S training prediction models obtained from S initial prediction models, acquire the model prediction accuracy corresponding to each of the S training prediction models, and determine the training prediction model with the highest model prediction accuracy as the target prediction model.
[0051] The model training unit includes:
[0052] The sample partitioning subunit is used to divide the target behavior data and behavior execution tags corresponding to N sample users into k sample sets; k is a positive integer, k is less than or equal to N, and the sample set includes the target behavior data and behavior execution tags corresponding to Z sample users; Z is less than or equal to N;
[0053] The verification acquisition subunit is used to acquire a verification sample set from k sample sets, determine the target behavior data included in the verification sample set as verification behavior data, and determine the behavior execution labels included in the verification sample set as verification behavior execution labels;
[0054] The training acquisition sub-unit is used to determine the sample set other than the validation sample set from the k sample sets as the training sample set, the target behavior data included in the training sample set as the training behavior data, and the behavior execution labels included in the training sample set as the training behavior execution labels.
[0055] The model training subunit is used to train the initial prediction model d based on the training behavior data and the training behavior execution labels, so as to obtain the training prediction model corresponding to the initial prediction model d.
[0056] This accuracy acquisition unit includes:
[0057] The verification prediction subunit is used to input verification behavior data into the training prediction model d for prediction, and obtain the verification execution probability.
[0058] The accuracy determination subunit is used to determine the model prediction accuracy of the training prediction model corresponding to the initial prediction model d based on the verification error between the verification execution probability and the verification behavior execution label, until the model prediction accuracy corresponding to S training prediction models is obtained respectively.
[0059] The device also includes:
[0060] The time series acquisition module is used to acquire the time series features of target behavior data of N sample users under the target behavior index, and adjust the target prediction model based on the time series features to obtain the adjusted target prediction model; the time series features are used to represent the features composed of the generation time series of the target behavior data;
[0061] The accuracy acquisition module is used to obtain the first model prediction accuracy of the target prediction model and the second model prediction accuracy of the adjusted target prediction model.
[0062] The accuracy comparison module is used to delete the adjusted target prediction model if the prediction accuracy of the first model is greater than or equal to the prediction accuracy of the second model.
[0063] The accuracy comparison module is also used to determine the adjusted target prediction model as the model for predicting the probability of the target user executing the target business if the prediction accuracy of the first model is less than that of the second model.
[0064] One embodiment of this application provides a data processing apparatus, the apparatus comprising:
[0065] The information acquisition module is used to acquire business consultation information submitted by target users in social applications, and to acquire the target users' predictable behavior data under target behavior indicators. The social application includes the target business. The target behavior indicators are determined from M candidate behavior indicators based on their respective impact on the target business. The M candidate behavior indicators are associated with the social application. The impact of each of the M candidate behavior indicators on the target business is determined based on the historical behavior data of N sample users under the M candidate behavior indicators, and the behavior execution labels of the N sample users for the target business. M is a positive integer, and N is a positive integer.
[0066] The probability prediction module is used to input the behavioral data to be predicted into the target prediction model. In the target prediction model, the target user's target execution probability for the target business is predicted based on the behavioral data to be predicted. The target execution probability is used to represent the probability that the target user will execute the target business.
[0067] The device also includes:
[0068] The business push module is used to send business consultation information to the business device associated with the business personnel if the target execution probability is greater than or equal to the business execution threshold, so that the business device can display the business consultation information to the business personnel.
[0069] The default reply module is used to obtain the default reply information corresponding to the business consultation information and send the default reply information to the user device associated with the target user if the target execution probability is less than the business execution threshold.
[0070] One embodiment of this application provides a computer device, including a processor, a memory, and an input / output interface;
[0071] The processor is connected to a memory and an input / output interface, respectively. The input / output interface is used to receive and output data, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device containing the processor executes the data processing method in one aspect of the embodiments of this application.
[0072] One aspect of this application provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor, so that a computer device having the processor performs the data processing method of one aspect of this application.
[0073] One aspect of this application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional embodiments of this application.
[0074] Implementing the embodiments of this application will have the following beneficial effects:
[0075] In this embodiment, M candidate behavioral indicators associated with a social application are obtained, and historical behavioral data of N sample users under the M candidate behavioral indicators are obtained; N is a positive integer, and M is a positive integer; the social application includes a target business; behavioral execution tags for the target business are obtained for each of the N sample users; based on the behavioral execution tags and historical behavioral data of the N sample users, the execution impact of each of the M candidate behavioral indicators on the target business is determined; based on the execution impact of each of the M candidate behavioral indicators on the target business, the target behavioral indicator is determined from the M candidate behavioral indicators; target behavioral data of the N sample users under the target behavioral indicator is obtained from the historical behavioral data; the initial prediction model is trained using the behavioral execution tags and target behavioral data of the N sample users to obtain a target prediction model; the target prediction model is used to predict the execution probability of the target user executing the target business. Through the above process, the execution probability of the target user executing the target business is predicted, enabling push notifications and other processing to the target user when there is a high probability that the target user will execute the target business, thereby improving the efficiency of data processing. The types of candidate behavioral metrics associated with social applications are generally numerous, meaning that a single user generates a large amount of data within that application. By filtering candidate behavioral metrics, the amount of data used for model training and prediction can be reduced, thereby improving data processing efficiency. Furthermore, the target behavioral metrics used for model training are determined based on the execution impact of each candidate behavioral metric on the target business. This execution impact represents the degree to which the corresponding candidate behavioral metric influences the probability of a sample user performing the target business, thus improving the accuracy of predicting the execution probability of the target business. Attached Figure Description
[0076] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0077] Figure 1 This is a network interaction architecture diagram for data processing provided in an embodiment of this application;
[0078] Figure 2 This is a schematic diagram of a model processing scenario provided in an embodiment of this application;
[0079] Figure 3 This is a flowchart of a data processing method provided in an embodiment of this application;
[0080] Figure 4This is a flowchart illustrating a data processing method provided in an embodiment of this application.
[0081] Figure 5 This application provides a schematic diagram of a scenario for obtaining candidate behavior indicators.
[0082] Figure 6 This is a schematic diagram of a scenario for creating combined behavioral indicators provided in an embodiment of this application;
[0083] Figure 7 This is a schematic diagram of a model prediction scenario provided in an embodiment of this application;
[0084] Figure 8 This is a flowchart of a model prediction method provided in an embodiment of this application;
[0085] Figure 9 This is a diagram illustrating the architecture of a data processing device provided in an embodiment of this application.
[0086] Figure 10 This is a schematic diagram of a data processing device provided in an embodiment of this application;
[0087] Figure 11 This is a schematic diagram of another data processing device provided in an embodiment of this application;
[0088] Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0089] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0090] Optionally, this application may employ machine learning techniques in the field of artificial intelligence to train a target prediction model, and may use the target prediction model to predict the probability of a target user executing a target business.
[0091] Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine capable of reacting in a manner similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities. For example, this application involves the automated screening of candidate behavioral indicators and the training of a target prediction model based on samples corresponding to the screened target behavioral indicators. Furthermore, when it is necessary to predict the probability of a target user executing a target business, the data generated by the target user under the target behavioral indicators can be input into the target prediction model for prediction, yielding the probability of the target user executing the target business. All of the above processes can be considered to be implemented based on artificial intelligence.
[0092] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning. This application may use any single AI technology individually, or it may combine various AI technologies randomly, such as using machine learning / deep learning alone, or combining natural language processing with machine learning / deep learning, etc., without any restrictions. The use of AI technologies improves the efficiency of data processing in this application.
[0093] Deep learning (DL) is a new research direction in the field of machine learning (ML). Deep learning learns the inherent patterns and hierarchical representations of sample data. The information gained during this learning process greatly aids in the interpretation of data such as text, images, and sound. By applying deep learning to the target behavior data generated by sample users under target behavior indicators, a target prediction model can be obtained, which can be used in this application to predict the execution probability of target users performing target business. Furthermore, based on the prediction results of the target prediction model, error feedback adjustments can be made, enabling the target prediction model to possess analytical and learning capabilities, continuously optimizing and updating itself, much like a human. Deep learning is a complex machine learning algorithm that has achieved results in speech and image recognition far exceeding previous related technologies. Deep learning typically includes techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and formulaic learning.
[0094] With the research and advancement of artificial intelligence technology, it has been studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role, such as the business processing field in this application.
[0095] The solutions provided in this application involve technologies such as machine learning in the field of artificial intelligence, and are specifically illustrated through the following embodiments:
[0096] In the embodiments of this application, please refer to Figure 1 , Figure 1 This is a network interaction architecture diagram for data processing provided in an embodiment of this application. This embodiment can be implemented using a computer device. Wherein, as... Figure 1As shown, computer device 101 can interact with user devices such as user devices 102a, 102b, and 102c. Computer device 101 can be a server or terminal device associated with a social application, a system associated with a social application, or a device used for training models, etc., without limitation. Computer device 101 can acquire M candidate behavioral indicators associated with the social application. These M candidate behavioral indicators refer to the types of data that may be generated in the social application. These M candidate behavioral indicators can be stored in computer device 101 or in a blockchain network, without limitation, where M is a positive integer. Furthermore, computer device 101 can acquire historical behavioral data of N sample users under the M candidate behavioral indicators. This historical behavioral data can be stored in computer device 101, in each user device, in a blockchain network, or based on cloud storage technology, etc., where N is a positive integer. For example, the historical behavior data is stored in various user devices. Computer device 101 can obtain the historical behavior data of N sample users under M candidate behavior indicators from the user devices corresponding to the N sample users respectively. Assuming that sample user 1 corresponds to user device 102a and sample user 2 corresponds to user device 102b, computer device 101 can obtain the historical behavior data of sample user 1 under M candidate behavior indicators from user device 102a; computer device 101 can obtain the historical behavior data of sample user 2 under M candidate behavior indicators from user device 102b, and so on. In this embodiment, positive integers refer to integers greater than 0, such as 1, 2, 3, etc.
[0097] Furthermore, the computer device 101 can determine the execution impact of M candidate behavioral indicators on the target service based on the acquired historical behavioral data. A higher execution impact indicates a greater influence of the corresponding candidate behavioral indicator on whether the user executes the target service. In other words, a higher execution impact means that when the data corresponding to that candidate behavioral indicator changes, the probability of the user executing the target service is likely to change accordingly. Therefore, the computer device can determine the target behavioral indicator to be used for model training based on the execution impact of the M candidate behavioral indicators, thereby reducing the amount of data involved in model training, improving the efficiency of model training, and simultaneously improving the efficiency of model prediction. The computer device 101 can then train the model based on the target behavioral data corresponding to N sample users to generate a target prediction model. Optionally, when the computer device needs to predict the probability of a target user executing a target service, it can acquire the target user's behavioral data under the target behavioral indicators, predict the behavioral data using a target prediction model, and obtain the probability that the target user will execute the target service, i.e., the target execution probability. Based on this target execution probability, the service processing result for the target user is determined, including service push results or service hold results, etc. On the one hand, the acquired candidate behavioral indicators are screened to reduce the amount of data required for model training and prediction; on the other hand, based on the execution impact corresponding to each of the M candidate behavioral indicators, the M candidate behavioral indicators are screened to obtain the target behavioral indicators with a strong correlation to the execution probability of the target service, so that the reduction in the amount of data processed will not reduce the accuracy of model training and prediction. The computer device used to train the target prediction model and the computer device used to use the target prediction model can be the same device or different devices; there is no restriction on this.
[0098] For details, please see Figure 2 , Figure 2 This is a schematic diagram of a model processing scenario provided in an embodiment of this application. For example... Figure 2As shown, the computer device can acquire M candidate behavioral indicators 202 associated with the social application 201, including candidate behavioral indicator 1, candidate behavioral indicator 2, ... and candidate behavioral indicator M. The computer device acquires historical behavioral data 203 of N sample users under the M candidate behavioral indicators. Assuming that the N sample users include sample user 1, sample user 2, ... and sample user N, the computer device acquires N historical behavioral data of the N sample users under candidate behavioral indicator 1, acquires N historical behavioral data of the N sample users under candidate behavioral indicator 2, ..., acquires N historical behavioral data of the N sample users under candidate behavioral indicator M. In other words, the computer device acquires the historical behavior data 1 of sample user 1 under candidate behavior indicator 1, acquires the historical behavior data 2 of sample user 2 under candidate behavior indicator 1, ..., acquires the historical behavior data N of sample user N under candidate behavior indicator 1; acquires the historical behavior data 1 of sample user 1 under candidate behavior indicator 2, acquires the historical behavior data 2 of sample user 2 under candidate behavior indicator 2, ..., acquires the historical behavior data N of sample user N under candidate behavior indicator 2; ...; acquires the historical behavior data 1 of sample user 1 under candidate behavior indicator M, acquires the historical behavior data 2 of sample user 2 under candidate behavior indicator M, ..., acquires the historical behavior data N of sample user N under candidate behavior indicator M.
[0099] Furthermore, the computer device can acquire behavioral execution labels for N sample users targeting the target service. These behavioral execution labels include positive and negative sample behavioral execution labels, where the probability of a sample user corresponding to a positive behavioral execution label executing the target service is greater than the probability of a sample user corresponding to a negative behavioral execution label executing the target service. Based on the behavioral execution labels corresponding to the N sample users and N historical behavioral data points corresponding to each candidate behavioral indicator, the computer device determines the execution impact of M candidate behavioral indicators targeting the target service. Specifically, based on the behavioral execution labels corresponding to the N sample users and N historical behavioral data points corresponding to candidate behavioral indicator 1, the computer device determines the execution impact 1 of candidate behavioral indicator 1 targeting the target service; based on the behavioral execution labels corresponding to the N sample users and N historical behavioral data points corresponding to candidate behavioral indicator 2, the computer device determines the execution impact 2 of candidate behavioral indicator 2 targeting the target service; ...; based on the behavioral execution labels corresponding to the N sample users and N historical behavioral data points corresponding to candidate behavioral indicator M, the computer device determines the execution impact M of candidate behavioral indicator M targeting the target service. The computer equipment determines the target behavior index based on the execution impact of M candidate behavior indices 202 on the target business. Target behavior data of N sample users under the target behavior index is obtained and used as samples for training the model. Based on the behavior execution labels of the N sample users for the target business and the corresponding target behavior data of the N sample users, the initial prediction model is trained to obtain the target prediction model.
[0100] It is understood that the computer equipment or user equipment mentioned in the embodiments of this application includes, but is not limited to, terminal equipment or servers. In other words, the computer equipment or user equipment can be a server or a terminal device, or a system composed of servers and terminal devices. The terminal device mentioned above can be an electronic device, including but not limited to mobile phones, tablets, desktop computers, laptops, handheld computers, in-vehicle devices, augmented reality / virtual reality (AR / VR) devices, head-mounted displays, smart TVs, wearable devices, smart speakers, digital cameras, webcams, and other mobile internet devices (MIDs) with network access capabilities, or terminal devices in scenarios such as trains, ships, and flights. The server mentioned above can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, vehicle-to-everything (V2X) communication, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0101] Optionally, the data involved in the embodiments of this application (such as M candidate behavioral indicators, historical behavioral data of N sample users under M candidate behavioral indicators, or behavioral execution tags of N sample users for the target business, etc.) can be stored in computer devices, or can be stored based on cloud storage technology, or can be stored in a blockchain network, without any restrictions.
[0102] Further, please see Figure 3 , Figure 3 This is a flowchart of a data processing method provided in an embodiment of this application. Figure 3 As shown, in Figure 3 In the described method embodiment, the data processing procedure includes the following steps:
[0103] Step S301: Obtain M candidate behavioral indicators associated with social applications, and obtain historical behavioral data of N sample users under the M candidate behavioral indicators.
[0104] In this embodiment, the computer device acquires M candidate behavioral indicators associated with a social application. The social application refers to an application that enables data interaction between different devices or between the application and the user, such as an instant messaging application, a game application, a shopping application, or a video application, etc., without limitation. The social application can be an application (APP), a website, or a webpage, etc., where M is a positive integer. The M candidate behavioral indicators refer to indicators that can generate data within the social application; simply put, they can be considered as parameter indicators used in the social application to represent user behavior or personal information. For example, taking an instant messaging system that enables application transactions as an example, the instant messaging system can include transaction information, interception information, complaint information, basic user information, user registration information, user device information, communication interaction information, and user policy information, etc. Each piece of information included in the instant messaging system represents an indicator type, and each indicator type includes at least one candidate behavioral indicator. The candidate behavioral indicators corresponding to all indicator types included in the instant messaging system constitute the M candidate behavioral indicators associated with the social application. In other words, the M candidate behavioral indicators may be numerous. For example, the transaction information may include candidate behavioral indicators such as transaction scenario, number of transactions, transaction amount, historical transaction behavior, transaction frequency, and transaction time. The interception information may include… that is, the M candidate behavioral indicators may include, but are not limited to, transaction scenario, number of transactions, transaction amount, historical transaction behavior, transaction frequency, and transaction time. For example, taking this social application as a game application, the M candidate behavioral indicators associated with the game application may include, but are not limited to, game account, game items, game instance scenario, number of game instances, number of players in game instances, game transaction data, and game character level. Furthermore, the computer device can obtain historical behavioral data of N sample users under the M candidate behavioral indicators. That is, the computer device can obtain historical behavioral data of each of the N sample users under the M candidate behavioral indicators. Optionally, the historical behavioral data of the sample users under the M candidate behavioral indicators can be recorded as sample behavioral data. That is, the computer device obtains sample behavioral data of N sample users under the M candidate behavioral indicators, and the sample behavioral data corresponding to one sample user includes M historical behavioral data.
[0105] Optionally, if the data involved in this step is stored in a blockchain network, the computer device can obtain the application identifier of the social application, search for the application block associated with the application identifier in the blockchain network based on the application identifier, and obtain M candidate behavioral indicators associated with the social application from the application block. Further, the computer device can obtain user information blocks corresponding to N sample users from the blockchain network, and obtain historical behavioral data of the N sample users under the M candidate behavioral indicators from the user information blocks corresponding to the N sample users.
[0106] Among them, the N sample users are users associated with the social application, that is, users who have used the social application.
[0107] Step S302: Obtain the behavior execution tags of N sample users for the target business. Based on the behavior execution tags and historical behavior data of the N sample users, determine the execution impact of M candidate behavior indicators for the target business.
[0108] In this embodiment, the computer device can acquire behavioral execution tags for N sample users targeting a specific business. These behavioral execution tags include positive and negative sample behavioral execution tags, representing the historical execution count of the corresponding sample user for the target business. Further, based on the behavioral execution tags corresponding to the N sample users, the computer device determines the positive and negative sample ratios of the historical behavioral data of the N sample users under M candidate behavioral indicators. Based on these ratios, the execution impact of each of the M candidate behavioral indicators on the target business is determined. In this way, the distribution of the historical behavioral data of the N sample users under the M candidate behavioral indicators can be obtained when the historical execution count of the target business by the sample users in their historical use of the social application satisfies the positive sample behavioral execution tag. This execution impact can represent the degree of influence of the corresponding candidate behavioral indicator on the business processing results of the sample users for the target business.
[0109] Step S303: Based on the impact of the M candidate behavioral indicators on the execution of the target business, determine the target behavioral indicator from the M candidate behavioral indicators.
[0110] In this embodiment, the computer device can determine the target behavior indicator from the M candidate behavior indicators based on their respective impact on the execution of the target business. Specifically, the computer device can directly determine u target behavior indicators from the M candidate behavior indicators based on their execution impact, where the execution impact of u target behavior indicators is greater than the execution impact of all other candidate behavior indicators in the M candidate behavior indicators, where u is a positive integer; or, based on the respective execution impact of the M candidate behavior indicators on the target business, the M candidate behavior indicators can be divided into v indicator groups, each corresponding to an impact range. Based on the v indicator groups, the M candidate behavior indicators can be combined to determine the target behavior indicator, where v is a positive integer.
[0111] Step S304: In the historical behavior data, obtain the target behavior data of N sample users under the target behavior index. Train the initial prediction model using the behavior execution labels and target behavior data corresponding to the N sample users respectively to obtain the target prediction model.
[0112] In this embodiment, a computer device can obtain target behavior data for N sample users under a target behavior indicator from historical behavior data of N sample users under M candidate behavior indicators. This includes target behavior data corresponding to each sample user. Feature transformation is performed on the target behavior data of the N sample users under the target behavior indicator to obtain target behavior features corresponding to each of the N sample users. An initial prediction model is trained using the behavior execution labels and target behavior features corresponding to each of the N sample users to obtain a target prediction model. For example, taking a single sample user as an example, the computer device can adjust the initial prediction model based on the error between the behavior execution label corresponding to that sample user and the prediction results of the initial prediction model for the target behavior features, thus obtaining the target prediction model.
[0113] Further, please see Figure 4 , Figure 4 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 4 As shown, in Figure 4 In the described method embodiment, the data processing procedure includes the following steps:
[0114] Step S401: Obtain M candidate behavioral indicators associated with social applications.
[0115] In this embodiment, the computer device can acquire M candidate behavioral indicators associated with a social application. These M candidate behavioral indicators can be preset, obtained by detecting the social application to identify the M candidate behavioral indicators that may generate data within it, or by traversing the application code of the social application to determine the M candidate behavioral indicators associated with it. No limitation is imposed here. Optionally, the social application can be an application that implements human-computer interaction, such as an instant messaging application, a game application, a shopping application, or a video application. The above are merely examples of possible types of social applications and do not limit other types of applications. The candidate behavioral indicators associated with different social applications are not necessarily the same.
[0116] For example, please see Figure 5 , Figure 5 This is a schematic diagram illustrating a scenario for obtaining candidate behavior indicators provided in an embodiment of this application. For example... Figure 5 As shown, assuming the social application is one that enables transactions, it can include transaction information, interception information, complaint information, basic user information, user registration information, user device information, communication and interaction information, and user policy information, etc. The transaction information may include, but is not limited to, transaction scenarios, number of transactions, transaction amounts, historical transaction behavior, transaction frequency, and transaction time. Optionally, the transaction scenario refers to the transaction method used by the user when conducting transactions within the application, such as money transfers, red envelopes, or commercial payments. The interception information indicates data generated when an application transaction encounters an anomaly, such as the transaction being intercepted by the device corresponding to the social application, or the transaction being interrupted due to network problems. This includes, but is not limited to, the number of interceptions, the intercepted amount, and large-amount interceptions. The data includes: multi-strategy blocking, blocking reasons, blocking days, and blocking time; complaint information refers to data generated when users file customer service complaints, including but not limited to behavior at the time of the first complaint, historical complaint behavior (such as the number of historical complaints, the time of historical complaints, and the targets of historical complaints), and complaint text information (such as the reason for the complaint); basic user information refers to the user's personal information, such as age, gender, and activity level; user registration information is used to represent the data registered by the user in the social application, such as the number of user data cards bound, the time of user data card binding, and the functional information registered by the user in the social application; user device information refers to the device information associated with the user when using the social application, such as the device system; user strategy information is used to represent transaction blocking strategies, such as account abnormality blocking or amount abnormality blocking. The above are just some examples of candidate behavioral indicators. Computer devices can determine the acquired behavioral indicators into M candidate behavioral indicators associated with the social application, including but not limited to transaction scenarios, transaction times, ... and age.
[0117] For example, taking this social application as a game application, the M types of candidate behavioral indicators associated with the game application may include, but are not limited to, game account, game items, game instance scenes, number of game instances, number of players in game instances, game transaction data, and game character level.
[0118] For example, taking this social application as a shopping application, the M candidate behavioral indicators associated with the shopping application may include, but are not limited to, the number of shopping trips, shopping frequency, shopping data volume, and shopping categories.
[0119] Optionally, the social application can be any other application, and the M candidate behavioral indicators associated with the social application can also be other candidate behavioral indicators. There are no restrictions here, and it can also be considered that the M candidate behavioral indicators are determined based on the corresponding social application.
[0120] Step S402: Obtain historical behavior data of N sample users under M candidate behavior indicators.
[0121] In the embodiments of this application, the computer device can acquire the historical behavior data of each sample user under M candidate behavior indicators. That is, each sample user corresponds to M historical behavior data, and each candidate behavior indicator corresponds to N historical behavior data. In other words, it can be considered that a sample user generates one historical behavior data under one candidate behavior indicator.
[0122] Step S403: Obtain the behavior execution tags of N sample users for the target business.
[0123] In this embodiment, the behavior execution label includes positive sample behavior execution labels and negative sample behavior execution labels. The sample user corresponding to the positive sample behavior execution label can be considered a positive sample used for training the model, and the sample user corresponding to the negative sample behavior execution label can be considered a negative sample used for training the model. The target service refers to the object that the trained model needs to predict. For example, if the model needs to be used to predict the probability of a target user filing an application customer complaint, then the target service is an application customer complaint; or if the model needs to be used to predict the probability of a target user entering a target game instance, then the target service is a target game instance, etc.
[0124] Specifically, the computer device can determine the behavior execution tags of sample users who have performed the target service as positive sample behavior execution tags, and the behavior execution tags of sample users who have not performed the target service as negative sample behavior execution tags. Optionally, the computer device can also assign behavior execution tags to N sample users based on their historical execution counts and a threshold for the target service. Specifically, the computer device can obtain the historical execution counts of the target service for each of the N sample users, assign a positive sample behavior execution tag to the first sample user among the N sample users based on the historical execution counts, and assign a negative sample behavior execution tag to the second sample user among the N sample users based on the historical execution counts; the first sample user is the sample user whose historical execution counts for the target service are greater than or equal to the threshold; the second sample user is the sample user whose historical execution counts for the target service are less than the threshold. For example, assuming the threshold is 3, positive sample behavior execution tags can be assigned to sample users whose historical execution counts are greater than or equal to 3, and negative sample behavior execution tags can be assigned to sample users whose historical execution counts are less than 3. The threshold for the number of executions can be preset, determined based on historical application data of the social application, or determined based on the characteristics of the target business of the social application. By using this threshold, N sample users are divided into positive and negative sample users. This allows users with a higher historical execution frequency for the target business to be identified as positive sample users, thereby improving the accuracy of model training. For example, taking application customer service complaints as the target business, computer devices typically only push an application customer service complaint to human customer service when the number of executions is greater than or equal to the threshold. By setting this threshold as the basis for dividing N sample users into positive and negative sample users, the positive and negative behavior execution labels assigned to each sample user can more accurately represent the positive and negative sample attributes of each sample user during model training. For example, taking the social application as a game application and the target business as a copy of the game application, if a user executes the copy of the game application more than or equal to a certain number of times based on historical behavior data, the user can be considered to have a higher probability of executing the copy of the game application. This user can be recorded as a positive sample user, thereby making the distinction between positive and negative sample users more accurate and improving the accuracy of model training.
[0125] Optionally, the computer device can also acquire at least two theoretical probability labels, obtain the historical execution counts of N sample users for the target business, and determine the behavioral indicator labels corresponding to the N sample users based on the theoretical probability labels corresponding to the historical execution counts of the N sample users. For example, the at least two theoretical execution labels include 100%, 80%, 50%, 20%, and 0. The theoretical probability label for sample users whose historical execution counts fall within the first range is 100%, for those within the second range is 80%, for those within the third range is 50%, for those within the fourth range is 20%, and for those within the fifth range is 0. The computer device can acquire the ranges corresponding to the N historical execution counts, determine the theoretical probability labels corresponding to the N historical execution counts, and determine the behavioral execution labels for the N sample users corresponding to the N historical execution counts based on these theoretical probability labels. Optionally, the ranges can also be divided in other ways, which are not limited here.
[0126] Step S404: Based on the behavior execution tags and historical behavior data corresponding to N sample users, determine the execution impact of M candidate behavior indicators on the target business.
[0127] In this embodiment, the behavior execution label includes positive sample behavior execution labels and negative sample behavior execution labels; M candidate behavior indicators include the i-th candidate behavior indicator; i is a positive integer, i is less than or equal to M; the number of historical behavior data of N sample users under the i-th candidate behavior indicator is N. Taking the i-th candidate behavior indicator as an example, the computer device can cluster the N historical behavior data corresponding to the i-th candidate behavior indicator to obtain h behavior data sets; h is a positive integer, h is less than or equal to N; the h behavior data sets include the j-th behavior data set; j is a positive integer; j is less than or equal to h. Optionally, the computer device can cluster the N historical behavior data corresponding to the i-th candidate behavior indicator based on any clustering method. For example, the computer device can cluster the N historical behavior data corresponding to the i-th candidate behavior indicator based on the h indicator data intervals corresponding to the i-th candidate behavior indicator to obtain h behavior data sets; or, the computer device can directly cluster the N historical behavior data corresponding to the i-th candidate behavior indicator, that is, divide the historical behavior data with similar values (the similarity between historical behavior data is greater than the data similarity threshold) into one category to obtain h behavior data sets; or, the computer device can use a clustering algorithm to cluster the N historical behavior data corresponding to the i-th candidate behavior indicator to obtain h behavior data sets, etc., without any restrictions.
[0128] For example, let's take the example of a computer device being able to cluster N historical behavioral data corresponding to the i-th candidate behavioral indicator based on h indicator data intervals corresponding to the i-th candidate behavioral indicator, thus obtaining a set of h behavioral data. For example, assuming the i-th candidate behavioral indicator is age, and the h corresponding indicator data ranges include 0-15 years old, 16-25 years old, 26-35 years old, and over 35 years old (this indicator data range is only one possible division), the computer device can divide historical behavioral data belonging to ages 0-15 years old into one behavioral data set, historical behavioral data belonging to ages 16-25 years old into one behavioral data set, historical behavioral data belonging to ages 26-35 years old into one behavioral data set, historical behavioral data belonging to ages 35 years old into one behavioral data set, and so on. Similarly, assuming the i-th candidate behavioral indicator is a transaction venue, and the h corresponding indicator data ranges include transfers, red envelopes, and commercial payments, the computer device can divide historical behavioral data indicating transactions via transfers into one behavioral data set, historical behavioral data indicating transactions via red envelopes into one behavioral data set, historical behavioral data indicating transactions via commercial payments into one behavioral data set, and so on. The indicator data ranges mentioned above can also be determined or updated based on the corresponding candidate behavioral indicators, and are not limited to the above-mentioned indicator data range division methods.
[0129] Furthermore, the computer device can obtain the number of positive samples of historical behavior data with positive behavior execution labels in the j-th behavior data set, and the number of negative samples of historical behavior data with negative behavior execution labels in the j-th behavior data set. For example, suppose the j-th behavior data set includes f historical behavior data, which are data generated by f sample users under the i-th candidate behavior indicator. There is a one-to-one correspondence between the f historical behavior data and the f sample users, and the f sample users belong to N sample users, where f is a positive integer less than or equal to N. The computer device can obtain the f sample users corresponding to the f historical behavior data in the j-th behavior data set, obtain the behavior execution labels of the f sample users for the target business, count the number of first sample users with positive behavior execution labels among the f sample users, and determine the number of first sample users as the number of positive samples; count the number of second sample users with negative behavior execution labels among the f sample users, and determine the number of second sample users as the number of negative samples. Furthermore, the computer device can determine the set information content of the j-th behavior data set based on the number of positive samples and the number of negative samples corresponding to the j-th behavior data set; and sum the set information content corresponding to the h behavior data sets respectively to obtain the execution impact of the i-th candidate behavior indicator on the target business.
[0130] Specifically, when determining the set information content of the j-th behavior data set based on the number of positive samples and the number of negative samples corresponding to the j-th behavior data set, the computer device can determine the total number of positive samples based on the number of positive samples corresponding to each of the h behavior data sets, and determine the total number of negative samples based on the number of negative samples corresponding to each of the h behavior data sets. In other words, the sum of the number of positive samples corresponding to each of the h behavior data sets is determined as the total number of positive samples, and the sum of the number of negative samples corresponding to each of the h behavior data sets is determined as the total number of negative samples. Further, the computer device can obtain the positive sample ratio between the number of positive samples corresponding to the j-th behavior data set and the total number of positive samples, and obtain the negative sample ratio between the number of negative samples corresponding to the j-th behavior data set and the total number of negative samples, and determine the set weight of the j-th behavior data set based on the positive and negative sample ratios. Similarly, the set weights corresponding to each of the h behavior data sets can be obtained. Optionally, the set weights corresponding to each of the h behavior data sets can also be preset. Alternatively, the computer device can determine the set weights corresponding to each of the h behavioral data sets based on the number of historical behavioral data included in each of the h behavioral data sets. For example, the set weights can be determined based on the proportion of the number of historical behavioral data included in each of the h behavioral data sets in the N historical behavioral data corresponding to the i-th candidate behavioral indicator. This ensures that the set weights of behavioral data sets with strong relevance to the social application are larger, so that the execution impact determined based on these set weights can better represent the characteristics of the social application's audience (i.e., users), thereby improving the accuracy of model training and prediction. For example, if the social application is game application A, and the i-th candidate behavioral indicator is age, and the users of game application A are relatively young, mostly between 15 and 35 years old, then the set weights of the behavioral data sets corresponding to the 15-35 age group can be reset to the first set weight. However, there are fewer users of game application A between 0 and 15 years old, and due to age restrictions, the usage time of users between 0 and 15 years old is limited. Therefore, the set weights of the behavioral data sets corresponding to the 0-15 age group can be reset to the second set weight, where the first set weight is greater than the second set weight. Optionally, the computer device can also determine the set weights corresponding to the h behavioral data sets based on other weight determination methods. Alternatively, the set weights corresponding to the h behavioral data sets can also be determined based on the social application; this is not limited here. For example, assuming the set weights are determined based on the number of positive and negative samples corresponding to the j-th behavioral data set, the determination process of these set weights can be seen in formula ①:
[0131]
[0132] Among them, WOE jThe set weight used to represent the data set of the j-th row, py j This is used to represent the proportion of positive samples corresponding to the j-th behavior data set. This proportion can be the ratio between the number of positive samples corresponding to the j-th behavior data set and the total number of positive samples, where y... j Let y represent the number of positive samples corresponding to the j-th row of the data set, and y represent the total number of positive samples. That is, py j =y j / y;pt j This is used to represent the proportion of negative samples corresponding to the j-th behavior data set. This proportion can be the ratio between the number of negative samples corresponding to the j-th behavior data set and the total number of negative samples, where t... j Let represent the number of negative samples corresponding to the j-th row of data, and t represent the total number of negative samples, i.e., pt. j =t j / t. Here, the ratio between the proportion of positive samples and the proportion of negative samples corresponding to the j-th row data set is transformed to obtain "(y j / t j Therefore, the computer device can also obtain the ratio between the number of positive samples and the number of negative samples corresponding to the j-th behavior data set, to obtain the set size ratio, and obtain the ratio between the total number of positive samples and the total number of negative samples, to obtain the total size ratio. Based on the ratio between the set size ratio and the total size ratio, the set weight of the j-th behavior data set is determined. Here, ln() refers to the natural logarithm function. Optionally, a normalization algorithm can be used to normalize the ratio between the positive and negative sample proportions corresponding to the j-th behavior data set to obtain the set weight of the j-th behavior data set. Similarly, using the method of determining the set weight of the j-th behavior data set, the set weights corresponding to the h behavior data sets can be obtained respectively.
[0133] Furthermore, the computer device can obtain the difference between the proportion of positive samples and the proportion of negative samples, and perform weighted processing on the difference between the proportion of positive samples and the proportion of negative samples based on the set weights to obtain the set information content of the j-th row data set. The formula for generating the set information content of the j-th row data set can be found in Formula ②:
[0134] IV j =ln(py h -pt j )×WOE j ②
[0135] Among them, IV jThe set information content is used to represent the j-th behavior data set, where "X" represents a multiplication operation. Similarly, based on formulas ① and ② above, the set information content corresponding to each of the h behavior data sets can be obtained. The set information content corresponding to each of the h behavior data sets is summed to obtain the execution impact of the i-th candidate behavior indicator on the target business.
[0136] Among them, based on the method of determining the execution impact of the i-th candidate behavior indicator, the execution impact of the M candidate behavior indicators on the target business is obtained respectively.
[0137] Step S405: Based on the impact of the M candidate behavioral indicators on the execution of the target business, determine the target behavioral indicator from the M candidate behavioral indicators.
[0138] In this embodiment of the application, the computer device can sort the M candidate behavioral indicators based on the execution impact corresponding to each of the M candidate behavioral indicators, and determine u target behavioral indicators from the sorted M candidate behavioral indicators.
[0139] Optionally, the computer equipment can also obtain a first candidate behavioral indicator from M candidate behavioral indicators, combine the first candidate behavioral indicators into L combined behavioral indicators, and obtain the combined impact degree of each of the L combined behavioral indicators on the target business; wherein, the first candidate behavioral indicator refers to the candidate behavioral indicator whose execution impact degree falls within the first impact degree range; each combined behavioral indicator is obtained by combining at least two candidate behavioral indicators; the combined impact degree is used to represent the execution impact degree of the at least two candidate behavioral indicators included in the corresponding combined behavioral indicator on the target business; L is a positive integer. A second candidate behavioral indicator is obtained from M candidate behavioral indicators, and the combined behavioral indicators whose combined impact degree falls within the second impact degree range from the L combined behavioral indicators, along with the second candidate behavioral indicator, are determined as the target behavioral indicator; the second candidate behavioral indicator refers to the candidate behavioral indicator whose execution impact degree falls within the second impact degree range.
[0140] Furthermore, when acquiring combined behavioral indicators, the computer device can divide the M candidate behavioral indicators into v indicator groups based on their respective execution impact degrees. Each indicator group corresponds to an impact degree range. Based on the v indicator groups, the M candidate behavioral indicators are combined to determine the target behavioral indicator, where v is a positive integer. For example, assuming v is 3, the v indicator groups correspond to a first impact degree range, a second impact degree range, and a third impact degree range, respectively. The execution impact degree corresponding to the second impact degree range > the execution impact degree corresponding to the first impact degree range > the execution impact degree corresponding to the third impact degree range. The computer device can delete candidate behavioral indicators whose execution impact degree belongs to the third impact degree range, retain the candidate behavioral indicators whose execution impact degree belongs to the second impact degree range, and designate the candidate behavioral indicators whose execution impact degree belongs to the second impact degree range as the second candidate behavioral indicator. The first candidate behavioral indicator whose execution impact falls within the first impact range is obtained. These first candidate behavioral indicators are then combined into L combined behavioral indicators. Each combined behavioral indicator can consist of at least two candidate behavioral indicators. Different combined behavioral indicators may contain the same candidate behavioral indicators. For example, candidate behavioral indicators 1, 2, and 3 can form "a combined behavioral indicator obtained by combining candidate behavioral indicators 1 and 2," "a combined behavioral indicator obtained by combining candidate behavioral indicators 1 and 3," "a combined behavioral indicator obtained by combining candidate behavioral indicators 2 and 3," and "a combined behavioral indicator obtained by combining candidate behavioral indicators 1, 2, and 3," etc. An upper limit can be set on the number of at least two candidate behavioral indicators included in a combined behavioral indicator; that is, the number of at least two candidate behavioral indicators included in a combined behavioral indicator is less than or equal to the upper limit. The computer equipment can obtain the combined impact of the combined behavioral indicators. The method for determining this combined impact is the same as the method for determining the execution impact of candidate behavioral indicators described above. The quantity counted when determining the combined impact is the number of at least two candidate behavioral indicators included in the combined impact. Among the L combined behavioral indicators, the combined behavioral indicators whose combined influence falls within the second influence range, and the second candidate behavioral indicators whose execution influence falls within the second influence range, are identified as target behavioral indicators.
[0141] For example, please see Figure 6 , Figure 6 This is a schematic diagram of a scenario for creating combined behavioral indicators provided in an embodiment of this application. For example... Figure 6As shown, assuming the combined behavioral indicator can be composed of a first candidate behavioral indicator and a second candidate behavioral indicator, a combined indicator tree is created based on the first candidate behavioral indicator and the second candidate behavioral indicator. The combined indicator tree includes a first candidate layer 6011 and a second candidate layer 6012. The first candidate layer 6011 is composed of nodes corresponding to h1 behavioral data sets corresponding to the first candidate behavioral indicator, and the second candidate layer 6012 is composed of nodes corresponding to h2 behavioral data sets corresponding to the second candidate behavioral indicator. The process of dividing the N historical behavioral data corresponding to the first candidate behavioral indicator into h1 behavioral data sets and the process of dividing the N historical behavioral data corresponding to the second candidate behavioral indicator into h2 behavioral data sets can be referred to in step S404 above, which is the process of dividing the N historical behavioral data corresponding to the i-th candidate behavioral indicator into h behavioral data sets. The first candidate behavioral indicator and the second candidate behavioral indicator belong to M types of candidate behavioral indicators. The computer device can divide the N historical behavioral data corresponding to the first candidate behavioral indicator and the N historical behavioral data corresponding to the second candidate behavioral indicator into (h1*h2) combined data sets based on the h1 indicator data ranges corresponding to the first candidate behavioral indicator and the h2 indicator data ranges corresponding to the second candidate behavioral indicator. It then obtains the number of positive and negative samples for each combined data set, determines the total number of positive samples based on the number of positive samples for each combined data set, and determines the total number of negative samples based on the number of negative samples for each combined data set. Based on the number of positive and total positive samples for each combined data set, the computer device can determine the proportion of positive samples for each combined data set; the proportion of negative samples for each combined data set; and the combined weight for each combined data set based on the proportion of positive and negative samples. Finally, based on the combined weight, the proportion of positive samples, and the proportion of negative samples for each combined data set, the combined information content for each combined data set is determined. The combined information content corresponding to each combined data set is summed to obtain the combined impact of the combined behavioral index on the target business.
[0142] Step S406: In the historical behavior data, obtain the target behavior data of N sample users under the target behavior index. Train the initial prediction model using the behavior execution labels and target behavior data corresponding to the N sample users respectively to obtain the target prediction model.
[0143] In this embodiment, the computer device can search for target behavior data of N sample users under a target behavior indicator in historical behavior data. If there are missing sample users among the N sample users for the target behavior indicator, the data missing rate of the N sample users is determined based on the ratio between the missing sample users and the N sample users. If the data missing rate is less than the data imputation threshold, the behavior data distribution of the target behavior data of the sample users (excluding the missing sample users) under the target behavior indicator is obtained. Based on the behavior data distribution, the target behavior data of the missing sample users under the target behavior indicator is determined, thus obtaining the target behavior data of the N sample users under the target behavior indicator. For example, assuming the behavior data distribution is a uniform distribution, the computer device can determine the mean of the historical behavior data of the N sample users (excluding the missing sample users) under the target behavior indicator as the target behavior data of the missing sample users under the target behavior indicator. Optionally, the computer device can also delete the target behavior indicator when the data missing rate is greater than or equal to the data imputation threshold. Alternatively, the computer device can determine the default imputation value as the target behavior data of the missing sample user under the target behavior indicator; or, the computer device can perform interpolation processing on the historical behavior data of N sample users (excluding the missing sample user) under the target behavior indicator to determine the target behavior data of the missing sample user under the target behavior indicator, such as random interpolation, multiple interpolation, hot plateau interpolation, Lagrange interpolation, or Newton interpolation; or, the computer device can directly input the historical behavior data of N sample users (excluding the missing sample user) under the target behavior indicator into the missing imputation model for prediction to determine the target behavior data of the missing sample user under the target behavior indicator. This missing imputation model can be a regression model, a Bayesian model, a random forest model, or a decision tree model, etc., without restriction.
[0144] Furthermore, the computer device can acquire the indicator type of the target behavior indicator; if the indicator type is an ordered indicator type, then the target behavior data corresponding to N sample users are labeled and encoded to generate N target behavior features. The N target behavior features are then input into the initial prediction model for prediction to obtain the sample execution probability corresponding to each of the N target behavior features. Based on the sample execution probabilities corresponding to the N target behavior features and the error value between the behavior execution labels of the N sample users, the initial prediction model is trained to obtain the target prediction model; if the indicator type is an unordered indicator type, then the target behavior data corresponding to N sample users are one-hot encoded to generate N target behavior features. The N target behavior features are then input into the initial prediction model for prediction to obtain the sample execution probability corresponding to each of the N target behavior features. Based on the sample execution probabilities corresponding to the N target behavior features and the error value between the behavior execution labels of the N sample users, the initial prediction model is trained to obtain the target prediction model.
[0145] Optionally, if the indicator type is text, then taking one target behavior data corresponding to the target behavior indicator as an example, the computer device can perform word segmentation processing on the target behavior data corresponding to the target behavior indicator to obtain at least two text segmentation word groups. Based on the at least two text segmentation word groups, the word frequency (tf) of each text segmentation word group in the target behavior data is determined; an inverse file frequency set is obtained, and the inverse file frequency (idf) of each text segmentation word group in the target behavior data is obtained from the inverse file frequency set. According to the word frequency (tf) and inverse file frequency (idf) of each text segmentation word group, the word importance of each text segmentation word group is determined. The word frequency (tf) represents the proportion of the corresponding text segmentation word group in the target behavior data. For example, if there are 10 text segmentation word groups, and 3 of them contain the content "account abnormality," then the word frequency of the "account abnormality" text segmentation word group is 0.3. The inverse file frequency (idf) represents the difference of each text segmentation word group in different target behavior data. One method for determining the inverse file frequency, taking a target text segmentation phrase as an example, is to obtain the number of associated texts containing the target text segmentation phrase in the N target behavior data corresponding to the candidate behavior indicator. The inverse file frequency of the target text segmentation phrase can then be log(N / number of associated texts), indicating that the target text segmentation phrase belongs to at least two of the aforementioned text segmentation phrases. Optionally, the inverse file frequency corresponding to each text segmentation phrase can also be determined through other methods. Based on the product of the phrase frequency and the inverse file frequency corresponding to each text segmentation phrase, the phrase importance of each text segmentation phrase is determined. Keyword groups are determined based on the phrase importance of each text segmentation phrase, and feature transformation is performed on the keyword groups to obtain the target behavior features corresponding to the target behavior data. Similarly, the target behavior data of N sample users under the target behavior indicator and the corresponding target behavior features can be obtained. N target behavior features are input into the initial prediction model for prediction, and the sample execution probabilities corresponding to the N target behavior features are obtained. Based on the sample execution probabilities corresponding to the N target behavior features and the error values between the behavior execution labels of the N sample users, the initial prediction model is trained to obtain the target prediction model.
[0146] Furthermore, the computer device can create S initial prediction models; the number of decision layers in each of the S initial prediction models is different; the S initial prediction models include an initial prediction model d; S is a positive integer; the initial prediction model d is an initial prediction model containing d decision layers; d is a positive integer. Using the target behavior data and behavior execution labels corresponding to N sample users, the initial prediction model d is trained to obtain the training prediction model corresponding to the initial prediction model d; the S training prediction models obtained from the S initial prediction models are obtained, and the model prediction accuracy corresponding to each of the S training prediction models is obtained; the training prediction model with the highest model prediction accuracy is determined as the target prediction model. The initial prediction model can be an XGBoost model, which is an optimized distributed gradient boosting library with advantages such as strong classifier concepts, built-in rules for handling missing values, regularization, and parallel processing. The XGBoost model includes General Parameters, Booster Parameters, and Learning Parameters types. The General Parameters type guides the overall modeling direction of the initial prediction model, the Booster Parameters type indicates the growth method of each tree (classification or regression) in each iteration, and the Learning Parameters type indicates the model optimization method and weights within the current framework.
[0147] Taking the training process of the initial prediction model d as an example, the computer device can divide the target behavior data and behavior execution labels corresponding to N sample users into k sample sets; k is a positive integer, k is less than or equal to N, and the sample set includes the target behavior data and behavior execution labels corresponding to Z sample users; Z is less than or equal to N. A validation sample set is obtained from the k sample sets. The target behavior data included in the validation sample set is determined as validation behavior data, and the behavior execution labels included in the validation sample set are determined as validation behavior execution labels. The sample set other than the validation sample set in the k sample sets is determined as the training sample set. The target behavior data included in the training sample set is determined as training behavior data, and the behavior execution labels included in the training sample set are determined as training behavior execution labels. Based on the training behavior data and training behavior execution labels, the initial prediction model d is trained to obtain the training prediction model corresponding to the initial prediction model d. In this approach, when the computer device obtains the model prediction accuracy corresponding to each of the S training prediction models, it can input the verification behavior data into the training prediction model d for prediction to obtain the verification execution probability. Based on the verification error between the verification execution probability and the verification behavior execution label, the model prediction accuracy of the training prediction model corresponding to the initial prediction model d is determined, until the model prediction accuracy corresponding to each of the S training prediction models is obtained.
[0148] Optionally, the k sample sets can correspond to k different set partitioning methods. Under the g-th set partitioning method, the g-th sample set among the k sample sets is determined as the validation sample set, and the sample sets other than the validation sample set among the k sample sets are determined as the training sample set. The initial prediction model d is trained based on the training sample set to obtain the training prediction model under the g-th set partitioning method. The training prediction model under the g-th set partitioning method is then validated based on the validation sample set to obtain the model prediction accuracy of the training prediction model under the g-th set partitioning method. Similarly, the computer device can obtain the model prediction accuracy of the training prediction models corresponding to the k set partitioning methods respectively. Based on the model prediction accuracy of the training prediction models corresponding to the k set partitioning methods respectively, the final model prediction accuracy corresponding to the training prediction model containing d decision layers is determined. For example, the average of the model prediction accuracies of the training prediction models corresponding to the k set partitioning methods is taken as the final model prediction accuracy corresponding to the training prediction model containing d decision layers. Similarly, the prediction accuracy of the training prediction model corresponding to each initial prediction model is obtained. The number of decision layers included in the training prediction model with the highest prediction accuracy is determined as the prediction model depth, and the initial prediction model corresponding to the prediction model depth is determined as the initial prediction model to be updated. Based on the sample execution probabilities corresponding to N target behavior features and the error values between the behavior execution labels of N sample users, the initial prediction model to be updated is trained to obtain the target prediction model.
[0149] For example, assuming k is 10, the first sample set is determined as the validation sample set, and the second to tenth sample sets are determined as the training sample sets. The initial prediction model d is trained based on the training sample sets to obtain the first training prediction model corresponding to the initial prediction model d. The first training prediction model corresponding to the initial prediction model d is then validated based on the validation sample set to obtain the first model prediction accuracy corresponding to the initial prediction model d. The second sample set is then determined as the validation sample set, and the first sample set and the third to tenth sample sets are determined as the training sample sets. The initial prediction model d is trained based on the training sample sets to obtain... The second training prediction model corresponding to the initial prediction model d is validated based on the validation sample set to obtain the prediction accuracy of the second model corresponding to the initial prediction model d; ...; the 10th sample set is determined as the validation sample set, and the 1st to 9th sample sets are determined as the training sample sets. The initial prediction model d is trained based on the training sample sets to obtain the 10th training prediction model corresponding to the initial prediction model d. The 10th training prediction model corresponding to the initial prediction model d is validated based on the validation sample set to obtain the prediction accuracy of the 10th model corresponding to the initial prediction model d. The final prediction accuracy of the initial prediction model d is determined by considering the prediction accuracies from the 1st to the 10th model corresponding to the initial prediction model d.
[0150] Step S407: Obtain the temporal features of target behavior data of N sample users under the target behavior index, and adjust the target prediction model based on the temporal features to obtain the adjusted target prediction model.
[0151] In this embodiment, the computer device can acquire the temporal features of target behavior data of N sample users under the target behavior index, and adjust the target prediction model based on the temporal features to obtain the adjusted target prediction model. The temporal features are used to represent the features composed of the generation time series of the target behavior data, and can optionally be obtained by algorithms such as Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), or Seq2Seq. Seq2Seq is a variant of a recurrent neural network, including an encoder and a decoder. The computer device can acquire the first model prediction accuracy of the target prediction model and the second model prediction accuracy of the adjusted target prediction model. If the first model prediction accuracy is greater than or equal to the second model prediction accuracy, the adjusted target prediction model is deleted; if the first model prediction accuracy is less than the second model prediction accuracy, the adjusted target prediction model is determined as the model for predicting the execution probability of the target user performing the target service. This step can be optional.
[0152] Optionally, when training the model in step S406, the computer device can further screen the target behavior indicators, such as by using the feature importance algorithm or shapley value (SHAP value) algorithm in XGBoost to screen the target behavior indicators, so as to further reduce the amount of data that needs to be processed during model training or prediction. The SHAP value is used to represent the influence of features in each sample. In this application, the influence of each target behavior indicator on the prediction result of the model can be considered, and the positive or negative nature of the influence can be expressed.
[0153] In this embodiment, a computer device can acquire M candidate behavioral indicators associated with a social application, and acquire historical behavioral data of N sample users under the M candidate behavioral indicators; N is a positive integer, and M is a positive integer; the social application includes a target business; acquire behavioral execution tags for the target business for each of the N sample users, and determine the execution impact of the M candidate behavioral indicators on the target business based on the behavioral execution tags and historical behavioral data of the N sample users; determine the target behavioral indicator from the M candidate behavioral indicators based on the execution impact of the M candidate behavioral indicators on the target business; acquire target behavioral data of the N sample users under the target behavioral indicator from the historical behavioral data, and train an initial prediction model using the behavioral execution tags and target behavioral data of the N sample users to obtain a target prediction model; the target prediction model is used to predict the execution probability of a target user executing the target business. Through the above process, the prediction of the execution probability of a target user executing the target business is realized, enabling push notifications and other processing to be performed on the target user when there is a high probability that the target user will execute the target business, thereby improving the efficiency of data processing. The types of candidate behavioral metrics associated with social applications are generally numerous, meaning that a single user generates a large amount of data within that application. By filtering candidate behavioral metrics, the amount of data used for model training and prediction can be reduced, thereby improving data processing efficiency. Furthermore, the target behavioral metrics used for model training are determined based on the execution impact of each candidate behavioral metric on the target business. This execution impact represents the degree to which the corresponding candidate behavioral metric influences the probability of a sample user performing the target business, thus improving the accuracy of predicting the execution probability of the target business.
[0154] Furthermore, the computer device can determine the target execution probability of a target user performing a target service based on a target prediction model, and then determine the service processing result based on this target execution probability. For example, please see... Figure 7 , Figure 7 This is a schematic diagram of a model prediction scenario provided in an embodiment of this application, such as... Figure 7 As shown, when target user 701 is conducting application transaction 702, the application transaction 702 is intercepted 703. The computer device can obtain the predicted behavior data of target user 701 under the target behavior indicator, input the predicted behavior data into the target prediction model for prediction, and determine the business processing result. If the business processing result is a business push result, then target user 701 is pushed to human customer service, that is, area 704 indicates human customer service at this time; if the business processing result is a business hold result, then target user 701 is pushed to self-service customer service, that is, area 704 indicates self-service customer service at this time.
[0155] For details, please see Figure 8 , Figure 8 This is a flowchart of a model prediction method provided in an embodiment of this application. Figure 8 As shown, in Figure 8 In the described method embodiment, the model prediction process includes the following steps:
[0156] Step S801: Obtain business consultation information submitted by the target user in the social application, and obtain the target user's predictable behavior data under the target behavior indicators.
[0157] In this embodiment, business consultation information submitted by target users is obtained in a social application, and the target user's predictable behavior data under target behavior indicators is obtained; the social application includes the target business; the target behavior indicators are determined from M candidate behavior indicators based on the execution impact of M candidate behavior indicators on the target business; the M candidate behavior indicators are associated with the social application; the execution impact of the M candidate behavior indicators on the target business is determined based on the historical behavior data of N sample users under the M candidate behavior indicators, and the behavior execution tags of N sample users on the target business; M is a positive integer, and N is a positive integer.
[0158] Step S802: Input the behavioral data to be predicted into the target prediction model. In the target prediction model, predict the target user's target execution probability for the target business based on the behavioral data to be predicted.
[0159] In this embodiment of the application, the behavioral data to be predicted is input into the target prediction model. In the target prediction model, the target execution probability of the target user for the target business is predicted based on the behavioral data to be predicted. The target execution probability is used to represent the probability that the target user will execute the target business.
[0160] Furthermore, if the target execution probability is greater than or equal to the business execution threshold, the business processing result for the target user regarding the target business is determined to be the business push result; if the target execution probability is less than the business execution threshold, the business processing result for the target user regarding the target business is determined to be the business hold result. When the target business processing result is the business push result, the computer device can obtain the target object associated with the target business and provide services to the target user based on the target object.
[0161] For example, if the target execution probability is greater than or equal to the business execution threshold, a business consultation message is sent to the business device associated with the business personnel, so that the business device displays the business consultation message to the business personnel; that is, the target object is the business personnel. If the target execution probability is less than the business execution threshold, the default reply message corresponding to the business consultation message is obtained and sent to the user device associated with the target user. Alternatively, assuming the social application is a game application and the target business is a target game instance, if the target execution probability is greater than or equal to the business execution threshold, the computer device obtains the target object associated with the target game instance, such as game items, and adds the target object to the virtual user of the target user in the game application; or, it displays object switching information, and in response to the confirmation operation for the object switching information, adds the target object to the virtual user of the target user in the game application.
[0162] Optionally, the computer device can adjust the parameters of the target prediction model based on feedback information from providing services to the target user, so as to continuously optimize the target prediction model.
[0163] Further, please see Figure 9 , Figure 9 This is a diagram illustrating the architecture of a data processing device provided in an embodiment of this application. Figure 9As shown, the data processing architecture includes an application layer, model selection, feature engineering, and a basic feature layer. The basic feature layer includes M candidate behavioral indicators, such as those found in transaction information, interception information, complaint information, complaint link information, complaint text information, and user basic information. Feature engineering processes the historical behavioral data of N sample users under the M candidate behavioral indicators. This includes missing value handling (identifying missing sample user behavioral data), data encoding (converting target behavioral data into target behavioral features), text processing (converting features of text-type target behavioral indicators), feature binning (clustering the N historical behavioral data corresponding to each of the M candidate behavioral indicators to obtain h sets of behavioral data corresponding to each of the M candidate behavioral indicators), correlation analysis (determining the execution impact of each of the M candidate behavioral indicators), and feature pre-selection (determining the target behavioral indicator). The model is selected for training. Specifically, M candidate behavioral indicators are added to tree-like features to determine the second candidate behavioral indicator and derived indicators (i.e., combined behavioral indicators that meet the second influence range) based on the execution influence corresponding to each of the M candidate behavioral indicators, thus obtaining the target behavioral indicator. The model is trained based on the target behavioral data of N sample users under the target behavioral indicator to obtain the target prediction model. Optionally, the computer device can add time-series features to the M candidate behavioral indicators to adjust the target prediction model and detect the model gain before and after adjustment. If model gain is achieved, the adjusted target prediction model is determined as the model used to predict the probability of the target user executing the target business; if model gain is not achieved, the unadjusted target prediction model is determined as the model used to predict the probability of the target user executing the target business. This application layer can obtain the target user's behavior data to be predicted under the target behavioral indicator and, based on the target prediction model 9011, determine the business processing result corresponding to the behavior data to be predicted. This business processing result includes business push results and business hold results. Based on the business processing results, the computer device determines the target object 9012 associated with the target business. Optionally, under the business push result, the target object 9012 associated with the target business is the target push object, and under the business hold result, the target object 9012 associated with the target business is the target hold object. For example, in the above "transaction -> interception -> complaint, etc." scenario, under the business push result, the target object 9012 associated with the target business is a human customer service representative, and under the business hold result, the target object 9012 associated with the target business is a self-service customer service representative, etc.The computer device can adjust the parameters of the target prediction model 9011 based on the feedback information 9013 of the service provided to the target user, and obtain the adjusted target prediction model 9014.
[0164] Further, please see Figure 10 , Figure 10 This is a schematic diagram of a data processing apparatus provided in an embodiment of this application. The data processing apparatus can be a computer program (including program code, etc.) running on a computer device; for example, the data processing apparatus can be application software. The apparatus can be used to execute corresponding steps in the methods provided in the embodiments of this application. Figure 10 As shown, the data processing device 1000 can be used for Figure 3 Specifically, the computer device in the corresponding embodiment may include: an indicator acquisition module 11, a label acquisition module 12, an influence determination module 13, an indicator selection module 14, a data acquisition module 15, and a model generation module 16.
[0165] The indicator acquisition module 11 is used to acquire M candidate behavioral indicators associated with social applications and to acquire historical behavioral data of N sample users under the M candidate behavioral indicators; N is a positive integer and M is a positive integer; the social application includes the target business;
[0166] The tag acquisition module 12 is used to acquire the action tags of N sample users for the target business.
[0167] The impact determination module 13 is used to determine the impact of M candidate behavioral indicators on the target business based on the behavioral execution tags and historical behavioral data corresponding to N sample users.
[0168] The indicator selection module 14 is used to determine the target behavior indicator from the M candidate behavior indicators based on the execution impact of the target business.
[0169] Data acquisition module 15 is used to acquire target behavior data of N sample users under the target behavior index from historical behavior data;
[0170] The model generation module 16 is used to train the initial prediction model using the behavior execution labels and target behavior data corresponding to N sample users to obtain the target prediction model; the target prediction model is used to predict the execution probability of the target user executing the target business.
[0171] The behavior execution label includes positive sample behavior execution labels and negative sample behavior execution labels;
[0172] This tag acquisition module 12 is specifically used for:
[0173] Obtain the historical execution counts of N sample users for the target business. Assign a positive sample behavior execution label to the first sample user among the N sample users based on the historical execution counts, and assign a negative sample behavior execution label to the second sample user among the N sample users based on the historical execution counts. The first sample user is the sample user whose historical execution count for the target business is greater than or equal to the number of times the threshold is applied. The second sample user is the sample user whose historical execution count for the target business is less than the number of times the threshold is applied.
[0174] Among them, the behavior execution label includes positive sample behavior execution label and negative sample behavior execution label; M candidate behavior indicators include the i-th candidate behavior indicator; i is a positive integer, i is less than or equal to M; the number of historical behavior data of N sample users under the i-th candidate behavior indicator is N;
[0175] The impact determination module 13 includes:
[0176] Data clustering unit 131 is used to cluster the N historical behavior data corresponding to the i-th candidate behavior indicator to obtain h behavior data sets; h is a positive integer, h is less than or equal to N; the h behavior data sets include the j-th behavior data set; j is a positive integer; j is less than or equal to h;
[0177] The quantity acquisition unit 132 is used to acquire the number of positive samples of historical behavior data with positive behavior execution labels in the j-th behavior data set, and to acquire the number of negative samples of historical behavior data with negative behavior execution labels in the j-th behavior data set.
[0178] The information determination unit 133 is used to determine the set information content of the j-th behavior data set based on the number of positive samples corresponding to the j-th behavior data set and the number of negative samples corresponding to the j-th behavior data set.
[0179] The impact determination unit 134 is used to sum the set information corresponding to each of the h behavioral data sets to obtain the execution impact of the i-th candidate behavioral indicator on the target business.
[0180] The information determining unit 133 includes:
[0181] The total number determination subunit 1331 is used to determine the total number of positive samples based on the number of positive samples corresponding to each of the h behavioral data sets, and to determine the total number of negative samples based on the number of negative samples corresponding to each of the h behavioral data sets.
[0182] The weight determination subunit 1332 is used to obtain the positive sample ratio between the number of positive samples corresponding to the j-th behavior data set and the total number of positive samples, obtain the negative sample ratio between the number of negative samples corresponding to the j-th behavior data set and the total number of negative samples, and determine the set weight of the j-th behavior data set based on the positive sample ratio and the negative sample ratio.
[0183] The weighted processing subunit 1333 is used to obtain the difference between the proportion of positive samples and the proportion of negative samples, and to perform weighted processing on the difference between the proportion of positive samples and the proportion of negative samples based on the set weight to obtain the set information of the j-th row data set.
[0184] The indicator selection module 14 includes:
[0185] The combination acquisition unit 141 is used to obtain a first candidate behavior indicator from M candidate behavior indicators, combine the first candidate behavior indicator into L combined behavior indicators, and obtain the combined impact degree of each of the L combined behavior indicators on the target business; the first candidate behavior indicator refers to the candidate behavior indicator whose execution impact degree belongs to the first impact degree range; each combined behavior indicator is obtained by combining at least two candidate behavior indicators; the combined impact degree is used to represent the execution impact degree of the at least two candidate behavior indicators included in the corresponding combined behavior indicator on the target business; L is a positive integer;
[0186] The indicator determination unit 142 is used to obtain the second candidate behavior indicator from M candidate behavior indicators, and to determine the combined behavior indicators whose combined influence degree belongs to the second influence degree range among L combined behavior indicators, as well as the second candidate behavior indicators, as the target behavior indicators; the second candidate behavior indicator refers to the candidate behavior indicator whose execution influence degree belongs to the second influence degree range.
[0187] The data acquisition module 15 includes:
[0188] The data lookup unit 151 is used to find the target behavior data of N sample users under the target behavior index in the historical behavior data;
[0189] The missing data acquisition unit 152 is used to determine the data missing rate of N sample users based on the ratio between the missing sample users and the N sample users if there are missing sample users for the target behavior indicator among the N sample users.
[0190] Distribution acquisition unit 153 is used to acquire the behavioral data distribution of target behavioral data under the target behavioral index for sample users other than missing sample users among N sample users if the data missing rate is less than the data filling threshold.
[0191] The data determination unit 154 is used to determine the target behavior data of missing sample users under the target behavior indicator based on the distribution of behavior data, and to obtain the target behavior data of N sample users under the target behavior indicator.
[0192] The model generation module 16 includes:
[0193] Type acquisition unit 161 is used to acquire the indicator type of the target behavior indicator;
[0194] The label encoding unit 162 is used to encode the target behavior data corresponding to N sample users respectively if the indicator type is an ordered indicator type, generate N target behavior features, input the N target behavior features into the initial prediction model for prediction, obtain the sample execution probability corresponding to the N target behavior features respectively, and train the initial prediction model based on the sample execution probability corresponding to the N target behavior features and the error value between the behavior execution labels of the N sample users to obtain the target prediction model.
[0195] One-hot encoding unit 163 is used to perform one-hot encoding on the target behavior data corresponding to N sample users if the indicator type is unordered, generate N target behavior features, input the N target behavior features into the initial prediction model for prediction, obtain the sample execution probability corresponding to the N target behavior features, and train the initial prediction model based on the sample execution probabilities corresponding to the N target behavior features and the error value between the behavior execution labels of the N sample users to obtain the target prediction model.
[0196] The model generation module 16 includes:
[0197] Model creation unit 164 is used to create S initial prediction models; the number of decision layers contained in the S initial prediction models are different; the S initial prediction models include initial prediction model d; S is a positive integer; initial prediction model d is an initial prediction model containing d decision layers; d is a positive integer.
[0198] The model training unit 165 is used to train the initial prediction model d using the target behavior data and behavior execution labels corresponding to N sample users, so as to obtain the training prediction model corresponding to the initial prediction model d.
[0199] The accuracy acquisition unit 166 is used to acquire S training prediction models obtained by training S initial prediction models respectively, acquire the model prediction accuracy corresponding to each of the S training prediction models, and determine the training prediction model with the highest model prediction accuracy as the target prediction model.
[0200] The model training unit 165 includes:
[0201] The sample partitioning subunit 1651 is used to divide the target behavior data and behavior execution tags corresponding to N sample users into k sample sets; k is a positive integer, k is less than or equal to N, and the sample set includes the target behavior data and behavior execution tags corresponding to Z sample users; Z is less than or equal to N.
[0202] The verification acquisition subunit 1652 is used to acquire a verification sample set from k sample sets, determine the target behavior data included in the verification sample set as verification behavior data, and determine the behavior execution labels included in the verification sample set as verification behavior execution labels;
[0203] The training acquisition subunit 1653 is used to determine the sample set other than the validation sample set in the k sample sets as the training sample set, the target behavior data included in the training sample set as the training behavior data, and the behavior execution labels included in the training sample set as the training behavior execution labels.
[0204] The model training subunit 1654 is used to train the initial prediction model d based on the training behavior data and the training behavior execution label, so as to obtain the training prediction model corresponding to the initial prediction model d.
[0205] The accuracy acquisition unit 166 includes:
[0206] The verification prediction subunit 1661 is used to input the verification behavior data into the training prediction model d for prediction to obtain the verification execution probability.
[0207] The accuracy determination subunit 1662 is used to determine the model prediction accuracy of the training prediction model corresponding to the initial prediction model d based on the verification error between the verification execution probability and the verification behavior execution label, until the model prediction accuracy corresponding to S training prediction models is obtained respectively.
[0208] The device 1000 also includes:
[0209] The time series acquisition module 17 is used to acquire the time series features of target behavior data of N sample users under the target behavior index, and adjust the target prediction model based on the time series features to obtain the adjusted target prediction model; the time series features are used to represent the features composed of the generation time series of the target behavior data;
[0210] The accuracy acquisition module 18 is used to acquire the first model prediction accuracy of the target prediction model and the second model prediction accuracy of the adjusted target prediction model.
[0211] The accuracy comparison module 19 is used to delete the adjusted target prediction model if the prediction accuracy of the first model is greater than or equal to the prediction accuracy of the second model.
[0212] The accuracy comparison module 20 is also used to determine the adjusted target prediction model as the model for predicting the probability of the target user executing the target business if the prediction accuracy of the first model is less than that of the second model.
[0213] This application provides a data processing apparatus that can acquire M candidate behavioral indicators associated with a social application, and acquire historical behavioral data of N sample users under the M candidate behavioral indicators; N is a positive integer, M is a positive integer; the social application includes a target business; acquire behavioral execution tags for the target business for each of the N sample users, and determine the execution impact of the M candidate behavioral indicators on the target business based on the behavioral execution tags and historical behavioral data of the N sample users; determine the target behavioral indicator from the M candidate behavioral indicators based on the execution impact of the M candidate behavioral indicators on the target business; acquire target behavioral data of the N sample users under the target behavioral indicator from the historical behavioral data, and train an initial prediction model using the behavioral execution tags and target behavioral data of the N sample users to obtain a target prediction model; the target prediction model is used to predict the execution probability of the target user executing the target business. Through the above process, the execution probability of the target user executing the target business is predicted, enabling push notifications and other processing to the target user when there is a high probability that the target user will execute the target business, thereby improving the efficiency of data processing. The types of candidate behavioral metrics associated with social applications are generally numerous, meaning that a single user generates a large amount of data within that application. By filtering candidate behavioral metrics, the amount of data used for model training and prediction can be reduced, thereby improving data processing efficiency. Furthermore, the target behavioral metrics used for model training are determined based on the execution impact of each candidate behavioral metric on the target business. This execution impact represents the degree to which the corresponding candidate behavioral metric influences the probability of a sample user performing the target business, thus improving the accuracy of predicting the execution probability of the target business.
[0214] Further, please see Figure 11 , Figure 11 This is a schematic diagram of another data processing apparatus provided in an embodiment of this application. The data processing apparatus can be a computer program (including program code, etc.) running on a computer device; for example, the data processing apparatus can be application software. The apparatus can be used to execute corresponding steps in the methods provided in the embodiments of this application. Figure 11 As shown, the data processing device 1100 can be used for Figure 8 Specifically, the computer device in the corresponding embodiment may include: an information acquisition module 31 and a probability prediction module 32.
[0215] Information acquisition module 31 is used to acquire business consultation information submitted by target users in social applications, and to acquire the target user's predictable behavior data under target behavior indicators; the social application includes the target business; the target behavior indicators are determined from the M candidate behavior indicators based on the execution impact of each of the M candidate behavior indicators on the target business; the M candidate behavior indicators are associated with the social application; the execution impact of each of the M candidate behavior indicators on the target business is determined based on the historical behavior data of N sample users under the M candidate behavior indicators, and the behavior execution labels of the N sample users on the target business; M is a positive integer, and N is a positive integer;
[0216] The probability prediction module 32 is used to input the behavioral data to be predicted into the target prediction model. In the target prediction model, the target execution probability of the target user for the target business is predicted based on the behavioral data to be predicted. The target execution probability is used to represent the probability that the target user will execute the target business.
[0217] The device 1100 also includes:
[0218] The business push module 33 is used to send business consultation information to the business device associated with the business personnel if the target execution probability is greater than or equal to the business execution threshold, so that the business device can display the business consultation information to the business personnel.
[0219] The default reply module 34 is used to obtain the default reply information corresponding to the business consultation information and send the default reply information to the user device associated with the target user if the target execution probability is less than the business execution threshold.
[0220] See Figure 12 , Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 12 As shown, the computer device in this embodiment may include one or more processors 1201, a memory 1202, and an input / output interface 1203. The processor 1201, memory 1202, and input / output interface 1203 are connected via a bus 1204. The memory 1202 stores a computer program, which includes program instructions. The input / output interface 1203 receives and outputs data, such as for data interaction between the computer device and a user device. The processor 1201 executes the program instructions stored in the memory 1202.
[0221] When the processor 1201 is running on a device used for model training, it can perform the following operations:
[0222] Obtain M candidate behavioral metrics associated with social applications, and obtain historical behavioral data of N sample users under the M candidate behavioral metrics; N is a positive integer, M is a positive integer; social applications include the target business;
[0223] Obtain behavioral execution tags for N sample users targeting the target business. Based on the behavioral execution tags and historical behavioral data of the N sample users, determine the execution impact of M candidate behavioral indicators on the target business.
[0224] Based on the impact of M candidate behavioral indicators on the execution of the target business, the target behavioral indicator is determined from the M candidate behavioral indicators.
[0225] In historical behavior data, target behavior data of N sample users under target behavior indicators are obtained. The initial prediction model is trained using the behavior execution labels and target behavior data corresponding to the N sample users to obtain the target prediction model. The target prediction model is used to predict the probability of the target user executing the target business.
[0226] Alternatively, when the processor 1201 is running in a device used for model prediction, it can perform the following operations:
[0227] The process involves acquiring business inquiry information submitted by target users within a social application, and obtaining predictable behavioral data of these users under target behavioral metrics. The social application includes the target business. The target behavioral metrics are determined from M candidate behavioral metrics based on their respective impact on the target business. These M candidate behavioral metrics are associated with the social application. The impact of each of the M candidate behavioral metrics on the target business is determined based on the historical behavioral data of N sample users under the M candidate behavioral metrics, and the behavioral execution tags of these N sample users for the target business. M and N are positive integers.
[0228] The behavioral data to be predicted is input into the target prediction model. In the target prediction model, the target user's target execution probability for the target business is predicted based on the behavioral data to be predicted. The target execution probability is used to represent the probability that the target user will execute the target business.
[0229] In some feasible implementations, the processor 1201 may be a central processing unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0230] The memory 1202 may include read-only memory and random access memory, and provides instructions and data to the processor 1201 and input / output interface 1203. A portion of the memory 1202 may also include non-volatile random access memory. For example, the memory 1202 may also store device type information.
[0231] In practice, the computer device can perform actions such as these through its built-in functional modules. Figure 3 or Figure 8 For details on the implementation methods provided for each step, please refer to [the relevant documentation / document / etc.]. Figure 3 or Figure 8 The implementation methods provided for each step are not elaborated here.
[0232] This application provides a computer device including a processor, an input / output interface, and a memory. The processor retrieves a computer program from the memory and executes it. Figure 3 The steps of the method shown are for model training, or for executing the... Figure 8 Each step of the method shown is used for model prediction. This embodiment of the application enables the prediction of the execution probability of a target user performing a target service, allowing for targeted push notifications and other processing when there is a high probability that the target user will perform the target service, thus improving data processing efficiency. Since there are generally many types of candidate behavioral indicators associated with social applications, meaning a single user generates a large amount of data within the application, filtering candidate behavioral indicators reduces the amount of data used for model training and prediction, thereby improving data processing efficiency. Furthermore, the target behavioral indicators used for model training are determined based on the execution impact of each candidate behavioral indicator on the target service. This execution impact represents the degree to which the corresponding candidate behavioral indicator influences the probability of a sample user performing the target service, improving the accuracy of predicting the execution probability of the target service.
[0233] This application also provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor. Figure 3 or Figure 8 For details on the data processing methods provided in each step, please refer to the document. Figure 3 or Figure 8 The implementation methods provided for each step are not repeated here. Furthermore, the beneficial effects of using the same method are also not repeated. For technical details not disclosed in the computer-readable storage medium embodiments involved in this application, please refer to the description of the method embodiments of this application. As an example, a computer program may be deployed to execute on a single computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed across multiple locations and interconnected via a communication network.
[0234] The computer-readable storage medium can be the data processing apparatus provided in any of the foregoing embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0235] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform... Figure 3 or Figure 8The methods provided among the various optional approaches enable the prediction of the execution probability of a target user performing a target service. This allows for targeted push notifications and other processing when there is a high probability that a target user will perform the target service, improving data processing efficiency. Furthermore, social applications typically generate a large number of candidate behavioral metrics, meaning a single user in such an application generates a significant amount of data. By filtering candidate behavioral metrics, the amount of data used for model training and prediction is reduced, thereby improving data processing efficiency. Additionally, the target behavioral metrics used for model training are determined based on the execution impact of each candidate behavioral metric on the target service. This execution impact represents the degree to which the corresponding candidate behavioral metric influences the probability of a sample user performing the target service, improving the accuracy of predicting the execution probability of the target service.
[0236] The terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0237] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0238] The methods and related apparatuses provided in this application are described with reference to the method flowcharts and / or structural diagrams provided in this application. Specifically, each block of the method flowchart and / or structural diagram, as well as combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to create a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the process. Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 A process or multiple processes and / or structures illustrate the steps of the functions specified in one or more boxes.
[0239] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A data processing method, characterized in that, The method includes: Obtain M candidate behavioral indicators associated with social applications, and obtain historical behavioral data of N sample users under the M candidate behavioral indicators; N is a positive integer, and M is a positive integer; the social applications include the target business; Obtain the behavioral execution tags of the N sample users for the target business. Based on the behavioral execution tags of the N sample users, determine the positive and negative sample ratios of the historical behavioral data of the N sample users under the M candidate behavioral indicators. Based on the positive and negative sample ratios, determine the execution impact of the M candidate behavioral indicators for the target business. The behavioral execution tags include positive sample behavioral execution tags and negative sample behavioral execution tags. A first candidate behavior indicator is obtained from the M candidate behavior indicators. The first candidate behavior indicator is combined into L combined behavior indicators, and the combined impact of the L combined behavior indicators on the target business is obtained. The first candidate behavior indicator refers to the candidate behavior indicator whose execution impact falls within the first impact range. Each combined behavior indicator is obtained by combining at least two candidate behavior indicators. The combined impact is used to represent the execution impact of the at least two candidate behavior indicators included in the corresponding combined behavior indicator on the target business. L is a positive integer. A second candidate behavior indicator is obtained from the M candidate behavior indicators. The combined behavior indicators whose combined influence falls within the second influence range from the L combined behavior indicators, along with the second candidate behavior indicator, are determined as the target behavior indicator. The second candidate behavior indicator refers to the candidate behavior indicator whose execution influence falls within the second influence range. In the historical behavior data, target behavior data of the N sample users under the target behavior index are obtained. The initial prediction model is trained using the behavior execution tags corresponding to the N sample users and the target behavior data to obtain the target prediction model. The target prediction model is used to predict the execution probability of the target user executing the target business.
2. The method as described in claim 1, characterized in that, The behavior execution labels include positive sample behavior execution labels and negative sample behavior execution labels; The step of obtaining the behavior execution tags of the N sample users for the target service includes: Obtain the historical execution counts of the N sample users for the target service, assign the positive sample behavior execution label to the first sample user among the N sample users based on the historical execution counts, and assign the negative sample behavior execution label to the second sample user among the N sample users based on the historical execution counts; the first sample user is a sample user whose historical execution count for the target service is greater than or equal to a count threshold; the second sample user is a sample user whose historical execution count for the target service is less than the count threshold.
3. The method as described in claim 1, characterized in that, The M candidate behavioral indicators include the i-th candidate behavioral indicator; i is a positive integer, i is less than or equal to M; the number of historical behavioral data of the N sample users under the i-th candidate behavioral indicator is N; Based on the behavioral execution tags corresponding to the N sample users, the positive and negative sample ratios of the historical behavioral data of the N sample users under the M candidate behavioral indicators are determined. Based on these positive and negative sample ratios, the execution impact of the M candidate behavioral indicators on the target business is determined, including: Cluster the N historical behavior data corresponding to the i-th candidate behavior indicator to obtain h behavior data sets; h is a positive integer, h is less than or equal to N; the h behavior data sets include the j-th behavior data set; j is a positive integer; j is less than or equal to h; Obtain the number of positive samples of historical behavior data in the j-th behavior data set whose behavior execution label is the positive sample behavior execution label; obtain the number of negative samples of historical behavior data in the j-th behavior data set whose behavior execution label is the negative sample behavior execution label. Based on the number of positive samples corresponding to the j-th behavior data set and the number of negative samples corresponding to the j-th behavior data set, the set information content of the j-th behavior data set is determined. The set information corresponding to each of the h behavioral data sets is summed to obtain the execution impact of the i-th candidate behavioral indicator on the target business.
4. The method as described in claim 3, characterized in that, The determination of the set information content of the j-th behavior data set based on the number of positive samples and the number of negative samples corresponding to the j-th behavior data set includes: The total number of positive samples is determined based on the number of positive samples corresponding to each of the h behavioral data sets, and the total number of negative samples is determined based on the number of negative samples corresponding to each of the h behavioral data sets. Obtain the positive sample ratio between the number of positive samples corresponding to the j-th behavior data set and the total number of positive samples; obtain the negative sample ratio between the number of negative samples corresponding to the j-th behavior data set and the total number of negative samples; and determine the set weight of the j-th behavior data set based on the positive sample ratio and the negative sample ratio. Obtain the difference between the positive sample ratio and the negative sample ratio, and perform weighted processing on the difference between the positive sample ratio and the negative sample ratio based on the set weight to obtain the set information of the j-th behavior data set.
5. The method as described in claim 1, characterized in that, The step of obtaining the target behavior data of the N sample users under the target behavior indicator from the historical behavior data includes: In the historical behavior data, find the target behavior data of the N sample users under the target behavior indicator; If there are missing sample users among the N sample users for the target behavior indicator, then the data missing rate of the N sample users is determined based on the ratio between the missing sample users and the N sample users. If the data missing rate is less than the data filling threshold, then obtain the behavioral data distribution of the target behavioral data of the sample users other than the missing sample users among the N sample users under the target behavioral indicator. Based on the distribution of the behavioral data, the target behavioral data of the missing sample users under the target behavioral indicator is determined, and the target behavioral data of the N sample users under the target behavioral indicator is obtained.
6. The method as described in claim 1, characterized in that, The step of training the initial prediction model using the behavior execution labels corresponding to the N sample users and the target behavior data to obtain the target prediction model includes: Obtain the indicator type of the target behavior indicator; If the indicator type is an ordered indicator type, then the target behavior data corresponding to the N sample users are labeled and encoded to generate N target behavior features. The N target behavior features are input into the initial prediction model for prediction to obtain the sample execution probability corresponding to the N target behavior features. Based on the sample execution probabilities corresponding to the N target behavior features and the error value between the behavior execution labels of the N sample users, the initial prediction model is trained to obtain the target prediction model. If the indicator type is an unordered indicator type, then one-hot encoding is performed on the target behavior data corresponding to the N sample users to generate N target behavior features. The N target behavior features are input into the initial prediction model for prediction to obtain the sample execution probability corresponding to the N target behavior features. Based on the sample execution probabilities corresponding to the N target behavior features and the error value between the behavior execution labels of the N sample users, the initial prediction model is trained to obtain the target prediction model.
7. The method as described in claim 1, characterized in that, The step of training the initial prediction model using the behavior execution labels corresponding to the N sample users and the target behavior data to obtain the target prediction model includes: Create S initial prediction models; the number of decision layers in each of the S initial prediction models is different; the S initial prediction models include an initial prediction model d; S is a positive integer; the initial prediction model d is an initial prediction model containing d decision layers; d is a positive integer. The initial prediction model d is trained using the target behavior data and behavior execution labels corresponding to the N sample users to obtain the trained prediction model corresponding to the initial prediction model d. Obtain S training prediction models obtained from the S initial prediction models respectively, obtain the model prediction accuracy corresponding to the S training prediction models respectively, and determine the training prediction model with the highest model prediction accuracy as the target prediction model.
8. The method as described in claim 7, characterized in that, The step of training the initial prediction model d using the target behavior data and behavior execution labels corresponding to the N sample users to obtain the trained prediction model corresponding to the initial prediction model d includes: The target behavior data and behavior execution tags corresponding to the N sample users are divided into k sample sets; k is a positive integer, k is less than or equal to N, and each sample set includes the target behavior data and behavior execution tags corresponding to Z sample users; Z is less than or equal to N. A verification sample set is obtained from the k sample sets. The target behavior data included in the verification sample set is determined as the verification behavior data, and the behavior execution labels included in the verification sample set are determined as the verification behavior execution labels. The set of samples other than the validation set in the k sample sets is determined as the training sample set, the target behavior data included in the training sample set is determined as the training behavior data, and the behavior execution labels included in the training sample set are determined as the training behavior execution labels. The initial prediction model d is trained based on the training behavior data and the training behavior execution label to obtain the training prediction model corresponding to the initial prediction model d. The step of obtaining the model prediction accuracy corresponding to the S trained prediction models includes: The verification behavior data is input into the trained prediction model d for prediction to obtain the verification execution probability; Based on the verification error between the verification execution probability and the verification behavior execution label, the model prediction accuracy of the training prediction model corresponding to the initial prediction model d is determined until the model prediction accuracy corresponding to the S training prediction models is obtained.
9. The method as described in claim 1, characterized in that, The method further includes: The time-series features of the target behavior data of the N sample users under the target behavior index are obtained, and the target prediction model is adjusted based on the time-series features to obtain the adjusted target prediction model; the time-series features are used to represent the features composed of the generation time series of the target behavior data. Obtain the first model prediction accuracy of the target prediction model, and obtain the second model prediction accuracy of the adjusted target prediction model; If the prediction accuracy of the first model is greater than or equal to the prediction accuracy of the second model, then the adjusted target prediction model is deleted. If the prediction accuracy of the first model is less than that of the second model, then the adjusted target prediction model is determined as the model for predicting the probability of the target user executing the target service.
10. A data processing method, characterized in that, The method includes: The system acquires business consultation information submitted by target users through a social application, and obtains the target users' predictable behavior data under target behavior indicators. The social application includes the target business. The target behavior indicators are determined from the M candidate behavior indicators based on their respective impact on the target business. The M candidate behavior indicators are associated with the social application. The impact of the M candidate behavior indicators on the target business is determined based on the positive and negative sample ratios corresponding to the historical behavior data of N sample users under the M candidate behavior indicators. M is a positive integer, and N is a positive integer. The positive and negative sample ratios are determined based on the behavior execution labels corresponding to the N sample users, and the behavior execution labels include positive sample behavior execution labels and negative sample behavior execution labels. The behavioral data to be predicted is input into the target prediction model, in which the target prediction model predicts the target user's target execution probability for the target service based on the behavioral data to be predicted; the target execution probability is used to represent the probability that the target user will execute the target service; The target behavior indicator includes a second candidate behavior indicator from the M candidate behavior indicators, and a combined behavior indicator from the L combined behavior indicators whose combined influence falls within the second influence range. The second candidate behavior indicator refers to a candidate behavior indicator whose execution influence falls within the second influence range. The L combined behavior indicators are formed by combining the first candidate behavior indicators from the M candidate behavior indicators, where the first candidate behavior indicator refers to a candidate behavior indicator whose execution influence falls within the first influence range. The combined influence is used to represent the execution influence of at least two candidate behavior indicators included in the corresponding combined behavior indicator on the target business. L is a positive integer.
11. The method as described in claim 10, characterized in that, The method further includes: If the target execution probability is greater than or equal to the business execution threshold, the business consultation information is sent to the business device associated with the business personnel, so that the business device displays the business consultation information to the business personnel. If the target execution probability is less than the business execution threshold, then the default reply information corresponding to the business consultation information is obtained, and the default reply information is sent to the user device associated with the target user.
12. A data processing apparatus, characterized in that, The device includes: The indicator acquisition module is used to acquire M candidate behavioral indicators associated with the social application and to acquire historical behavioral data of N sample users under the M candidate behavioral indicators; N is a positive integer and M is a positive integer; the social application includes the target business; The tag acquisition module is used to acquire the behavior execution tags of the N sample users for the target service; the behavior execution tags include positive sample behavior execution tags and negative sample behavior execution tags; The impact determination module is used to determine the positive and negative sample ratios of the historical behavior data of the N sample users under the M candidate behavior indicators based on the behavior execution tags corresponding to the N sample users respectively, and to determine the execution impact of the M candidate behavior indicators on the target business according to the positive and negative sample ratios. The indicator selection module is used to obtain a first candidate behavior indicator from the M candidate behavior indicators, combine the first candidate behavior indicator into L combined behavior indicators, and obtain the combined impact degree of the L combined behavior indicators on the target business respectively; the first candidate behavior indicator refers to the candidate behavior indicator whose execution impact degree belongs to the first impact degree range; each combined behavior indicator is obtained by combining at least two candidate behavior indicators; the combined impact degree is used to represent the execution impact degree of the at least two candidate behavior indicators included in the corresponding combined behavior indicator on the target business; L is a positive integer; The indicator selection module is further configured to obtain a second candidate behavior indicator from the M candidate behavior indicators, and to determine the combined behavior indicators whose combined influence degree belongs to the second influence degree range among the L combined behavior indicators, as well as the second candidate behavior indicator, as the target behavior indicator; the second candidate behavior indicator refers to the candidate behavior indicator whose execution influence degree belongs to the second influence degree range. The model generation module is used to obtain target behavior data of the N sample users under the target behavior index from the historical behavior data, and to train the initial prediction model using the behavior execution labels corresponding to the N sample users and the target behavior data to obtain the target prediction model; the target prediction model is used to predict the execution probability of the target user executing the target business.
13. A computer device, characterized in that, Includes processor, memory, and input / output interfaces; The processor is connected to the memory and the input / output interface respectively, wherein the input / output interface is used to receive data and output data, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device executes the method according to any one of claims 1-9, or executes the method according to any one of claims 10-11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-9, or the method of any one of claims 10-11.