User activity determination method, apparatus, and electronic device
By acquiring user interaction data, constructing feature information, and using clustering and classification models to generate activity segmentation criteria, the problem of inaccurate user activity segmentation is solved, achieving highly accurate and business-applicable activity determination.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2023-03-15
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies have poor accuracy in segmenting user activity and are not conducive to business operations. Existing methods rely on historical experience or single pieces of information from machine learning, which leads to inaccurate segmentation.
By acquiring user interaction data, constructing feature information, and utilizing pre-defined clustering algorithms and classification models, an interpretable and business-specific activity classification criterion is generated, including feature selection, clustering, and hyperplane optimization, combined with a global search algorithm to optimize activity classification.
It improves the accuracy and business applicability of user activity segmentation, and the generated activity segmentation criteria are stable and effective in different time periods, supporting the continuous development of the business.
Smart Images

Figure CN116361549B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, specifically relating to a method, apparatus, and electronic device for determining user activity. Background Technology
[0002] In information feed recommendations, users with different activity levels typically exhibit different behavioral habits, leading to significant differences in recommendation strategies. Therefore, accurately segmenting user activity is crucial for determining the recommendation strategy. However, current technologies often rely on historical experience to segment user activity, but this method suffers from poor accuracy. Alternatively, machine learning can be used, but this method depends solely on user information, which is detrimental to business operations. In other words, current technologies for segmenting user activity are both inaccurate and unfavorable for business development. Summary of the Invention
[0003] The purpose of this application is to provide a method, apparatus, and electronic device for determining user activity, which can solve the problem that existing methods for classifying user activity are inaccurate and detrimental to business operations.
[0004] In a first aspect, embodiments of this application provide a method for determining user activity, the method comprising:
[0005] Obtain primary feature information that reflects the activity level of any user;
[0006] The first user group is clustered according to the first feature information and the preset clustering algorithm to obtain at least one user cluster, wherein each user cluster corresponds to an activity level.
[0007] Based on the at least one user cluster and a preset classification model, multiple activity level division criteria are determined, wherein each of the activity level division criteria involves constraints on the first feature information.
[0008] The activity level of the target user is determined based on the target activity level criteria among the multiple activity level criteria.
[0009] Secondly, embodiments of this application provide a user activity determination device, the method comprising:
[0010] The acquisition module is used to acquire primary feature information that reflects the activity level of any user.
[0011] The clustering module is used to cluster the first user group according to the first feature information and the preset clustering algorithm to obtain at least one user cluster, wherein one user cluster corresponds to an activity level.
[0012] The first determining module is used to determine multiple activity level division criteria based on the at least one user cluster and a preset classification model, wherein each of the activity level division criteria involves constraints on the first feature information.
[0013] The second determining module is used to determine the activity level of the target user based on the target activity level division criteria among the multiple activity level division criteria.
[0014] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions, when executed by the processor, implementing the steps of the method described in the first aspect.
[0015] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0016] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0017] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0018] In this embodiment, the electronic device selects a first feature information reflecting the activity level of any user based on expert experience. Then, it clusters a first user group according to a preset clustering algorithm and the first feature information, obtaining user clusters with different activity levels. Based on these user clusters and a preset classification model, multiple activity level classification criteria are determined, each involving constraints on the first feature information. In other words, this embodiment effectively generates classification criteria from the preset clustering algorithm, ensuring they are communicable with business needs and highly interpretable. This results in highly accurate and business-oriented activity level classification criteria. Consequently, the activity level of target users determined based on the target activity level classification criteria is more accurate and more conducive to business operations. Attached Figure Description
[0019] Figure 1 This is a flowchart of the user activity determination method provided in the embodiments of this application;
[0020] Figure 2 This is one of the schematic diagrams of the hyperplanes corresponding to the two user clusters with different activity levels provided in the embodiments of this application;
[0021] Figure 3 This is a second schematic diagram of the hyperplanes corresponding to the two user clusters with different activity levels provided in the embodiments of this application;
[0022] Figure 4 This is a schematic diagram of the user activity determination device provided in the embodiments of this application;
[0023] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;
[0024] Figure 6 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0027] The user activity determination method provided in this application embodiment can be executed by the user activity determination device provided in this application embodiment, or by an electronic device integrating the user activity determination device, wherein the user activity determination device can be implemented in hardware or software.
[0028] The method for determining user activity provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0029] Please refer to Figure 1 This is a flowchart illustrating a method for determining user activity according to an embodiment of this application. User activity is used to characterize the frequency of interaction between a user and an information stream. This method can be applied to electronic devices, such as mobile phones, tablets, and laptops. Figure 1As shown, the method may include steps 1100 to 1400, which will be described in detail below.
[0030] Step 1100: Obtain first feature information to reflect the activity level of any user.
[0031] The first feature information is used to reflect the activity level of any user. It is understood that the first feature information used to reflect the activity level of any user refers to feature information highly correlated with the activity level of that user. Typically, this first feature information can be selected based on expert experience. In an optional embodiment, step 1100, obtaining the first feature information reflecting the activity level of any user, may further include the following steps 1110 to 1140:
[0032] Step 1110: Obtain the interaction data of the first user group.
[0033] Specifically, based on expert experience and understanding of the business, we can identify and represent the interactive behaviors between users and information feeds. For example, user interaction behaviors with information feeds include exposure, clicks, browsing time, likes, comments, favorites, and shares.
[0034] The first user group is a group of users who interact with the information flow. This embodiment does not limit the number of users included in the first user group.
[0035] Specifically, the electronic device can collect interaction data from each user in the first user group within the statistics window. The statistics window can be selected based on business experience, and can be 1 day, 3 days, 7 days, 30 days, 90 days, or 180 days. For example, the information stream is content provided by application A, and the first user group includes user 1, user 2, and user 3. Interaction behaviors include exposure, clicks, browsing time, likes, comments, favorites, and shares. The statistics for user 1 within one day are as follows: 5 exposures, 20 clicks, 25 minutes browsing time, 10 likes, 10 comments, 8 favorites, and 15 shares. The interaction data for user 2 is as follows: 6 exposures, 18 clicks, 36 minutes browsing time, 8 likes, 12 comments, 10 favorites, and 20 shares. User 3's interaction data: exposure "7 times", clicks "17 times", browsing time "49 minutes", likes "9 times", comments "8 times", favorites "6 times", and shares "7 times".
[0036] Step 1120: Obtain multiple second feature information based on the interaction data.
[0037] Specifically, electronic devices can perform feature cross-referencing and feature construction on interactive data to generate new feature information. For example, cross-referencing exposure volume and browsing duration in interactive data (browsing duration / exposure volume) can generate a new feature information: exposure contribution duration. Another example is exposure volume in interactive data. If it's necessary to track whether a user browses the information feed daily, features can be constructed based on the user's exposure behavior to statistically determine daily exposure and thus obtain a new feature information: the user's active days. Based on this, the second feature information includes exposure volume, click volume, browsing duration, number of likes, number of comments, number of favorites, number of shares, exposure contribution duration, and the user's active days.
[0038] Step 1130: Obtain the sparsity corresponding to each of the multiple second feature information, and the similarity between any two of the multiple second feature information.
[0039] In step 1130, the second feature information f i The corresponding sparsity is sparse(f i It satisfies the following formula:
[0040]
[0041] In step 1130, any two second feature information f i ,f j The similarity r(f) between them i ,f j It satisfies the following formula:
[0042]
[0043] Among them, Cov(f i ,f j ) represents the second feature information f i Second feature information f j The covariance between them, Var(f i ) represents the second feature information f i The variance, Var(f) j ) represents the second feature information f j The variance of the second feature is also considered. Furthermore, a higher similarity indicates a greater correlation between the two second features, resulting in poorer discriminative power.
[0044] Step 1140: Determine the first feature information to reflect the activity level of any user based on the sparsity corresponding to the multiple second feature information and the similarity between any two second feature information among the multiple second feature information.
[0045] In step 1140, the electronic device needs to perform feature filtering on the multiple second feature information based on the sparsity corresponding to the multiple second feature information and the similarity between any two of the multiple second feature information, and retain only the second feature information that is highly related to user activity as the first feature information.
[0046] Specifically, the electronic device first performs initial feature filtering on the second feature information based on sparsity, for example, retaining second feature information with sparsity less than the sparsity threshold alpha; then, it performs secondary feature filtering on the retained second feature information based on relevance, for example, if the relevance of two second feature information is greater than the relevance threshold beta, then these two second feature information are retained. The second feature information retained through secondary feature filtering is the second feature information that is highly related to user activity, and is used as the first feature information.
[0047] Based on steps 1110 to 1140 above, it is possible to select feature information that is highly correlated with user activity based on expert experience, thereby making the segmentation of user activity based on this feature information more accurate.
[0048] Step 1200: Cluster the first user group according to the first feature information and the preset clustering algorithm to obtain at least one user cluster, wherein each user cluster corresponds to an activity level.
[0049] In this embodiment, to avoid the influence of different units on the clustering effect, the electronic device needs to first normalize the first feature information, and then cluster the first user group according to the normalized first feature information and a preset clustering algorithm to obtain user clusters with different activity levels. The normalization algorithm may include one of z-score normalization, max-min normalization, or standard deviation normalization.
[0050] In an optional embodiment, step 1200, which clusters the first user group based on the first feature information and a preset clustering algorithm to obtain at least one user cluster, may further include the following steps 1210 to 1230:
[0051] Step 1210: Determine the cluster center of the first user group based on the feature values of each user in the first user group corresponding to the first feature information.
[0052] In step 1210, the first user group can be simply divided by electronic devices based on business experience and the quantiles of the feature values of the first feature information corresponding to each user in the first user group, and the center point of the first user group can be calculated as the cluster center of the preset clustering algorithm. This cluster center is the initial cluster center of the preset clustering algorithm.
[0053] Step 1220: Determine the similarity between each user and the cluster center based on the feature values of each user in the first user group for the first feature information.
[0054] In step 1220, for any user in the first user group, the electronic device can calculate the similarity between the feature value of the first feature information corresponding to that user and the feature value of the first feature information corresponding to the cluster center, based on a similarity calculation function, and use this similarity as the similarity between the user and the cluster center. The similarity calculation function includes one of cosine similarity, Euclidean distance, and Minkowski distance.
[0055] Step 1230: Based on the preset clustering algorithm and the similarity between each user and the cluster center, obtain the at least one user cluster.
[0056] In step 1230, the electronic device can process the similarity between each user and the cluster center based on a preset clustering algorithm, such as the k-means clustering algorithm, to obtain user clusters corresponding to different activity levels, such as high-activity user clusters, medium-activity user clusters, and low-activity user clusters. That is, the first user group is divided into three different activity levels—high activity, medium activity, and low activity—using a preset clustering algorithm.
[0057] Based on steps 1210 to 1230 above, it can cluster the first user group according to the preset clustering algorithm and the first feature information. Since the first feature information is highly correlated with user activity, the user activity obtained by the preset clustering algorithm is more accurate.
[0058] Step 1300: Based on the at least one user cluster and the preset classification model, determine multiple activity level division criteria, wherein each of the activity level division criteria involves constraints on the first feature information.
[0059] In this embodiment, although the preset clustering algorithm can determine the user's activity level by calculating the distance between the user and the cluster center, this method is not conducive to business development and interpretation. Therefore, it is necessary to generate activity level classification criteria based on the clustering results. The generated activity level classification criteria are highly interpretable and have strong business characteristics.
[0060] In an optional embodiment, step 1300, which determines multiple activity level division criteria based on the at least one user cluster and a preset classification model, may further include: determining first hyperplanes corresponding to two user clusters with adjacent activity levels based on the preset classification model; optimizing the first hyperplanes corresponding to the two user clusters with adjacent activity levels based on a first optimization objective to obtain target hyperplanes corresponding to the two user clusters with adjacent activity levels; and determining multiple activity level division criteria based on the target hyperplanes corresponding to the two user clusters with adjacent activity levels.
[0061] Specifically, electronic devices use a pre-defined classification model, such as an SVM model, to find hyperplanes that can effectively distinguish between two levels of user activity and have the largest distance between them. Based on these two hyperplanes, the activity classification criteria that distinguish between these two levels of activity can be easily mapped. The following explanation uses the activity classification criterion that distinguishes between high and medium activity levels as an example:
[0062] Step 1: Assign identification information to each user in the high-activity user cluster and each user in the medium-activity user cluster. This identification information is used to distinguish whether the user's activity level is high or medium.
[0063] Step 2: Assign positive classification (y=1) to users in the high-activity user cluster and negative classification (y=-1) to users in the medium-activity user cluster, referring to... Figure 2 Using a soft-margin SVM model, we find the two initial hyperplanes with the largest margins: wx + b = 1 (the initial hyperplane corresponding to the cluster of highly active users) and wx + b = -1 (the initial hyperplane corresponding to the cluster of moderately active users). These two initial hyperplanes can effectively distinguish between highly active and moderately active users. However, these two initial hyperplanes are easily affected by a small number of outliers, making the classification criteria somewhat random and biased. (Refer to...) Figure 3 To reduce the impact of outlier data and make the partitioning criteria more robust, slack variables are added to give the partitioning hyperplane a certain degree of fault tolerance, ensuring the robustness of the partitioning criteria. The first hyperplane with added slack variables is then used to obtain the first hyperplane w corresponding to the high-activity user clusters. T +=-1, the first hyperplane w corresponding to the cluster of users with medium activity level T + = 1.
[0064] Step 3: Optimize the first hyperplanes corresponding to the high-activity user cluster and the medium-activity user cluster according to the constructed first optimization objective, and obtain the target hyperplanes corresponding to the high-activity user cluster and the medium-activity user cluster respectively.
[0065] The first optimization objective, as constructed, satisfies the following formula:
[0066]
[0067] Here, C is a hyperparameter representing the penalty coefficient for misclassification. It can be designed according to business needs, and a larger C indicates a greater penalty for misclassification, while a smaller C indicates a smaller penalty. In other words, the choice of hyperparameter C affects the rationality and robustness of activity level segmentation. The first optimization objective is to ensure that, given a greater than preset interval between the first hyperplanes corresponding to two adjacent user clusters, the number of misclassifications of user activity by the first hyperplanes corresponding to these two adjacent user clusters is less than or equal to a preset number. The preset interval and preset number can be designed, with the preset interval typically being a larger value and the preset number typically a smaller value. Therefore, the first optimization objective can also be described as ensuring that, while the interval between the first hyperplanes corresponding to two adjacent user clusters is sufficiently large, the number of misclassifications of user activity by the first hyperplanes corresponding to these two adjacent user clusters is sufficiently small.
[0068] Step 4: Based on the target hyperplanes corresponding to the high-activity user clusters and the medium-activity user clusters, determine the activity classification criteria that distinguish between high-activity and medium-activity users.
[0069]
[0070] Here, each value of the x vector is the value of each first feature information. In other words, any hyperplane between the two target hyperplanes can effectively distinguish between highly active and moderately active groups.
[0071] Step 1400: Determine the activity level of the target user based on the target activity level criteria among the multiple activity level criteria.
[0072] The target user can be any user whose activity level needs to be determined. It can be a user in the first user group or a user who is not in the first user group. This embodiment does not limit this.
[0073] In this embodiment, since there are multiple activity level classification criteria, the electronic device can determine one of these criteria as the target activity level classification criterion, and then determine the activity level of the target user based on the target activity level classification criterion.
[0074] According to embodiments of this application, the electronic device selects first feature information reflecting the activity level of any user based on expert experience, and clusters a first user group according to a preset clustering algorithm and the first feature information to obtain user clusters with different activity levels. Then, based on these user clusters and a preset classification model, multiple activity level classification criteria are determined, each involving constraints on the first feature information. That is, through embodiments of this application, the clustering results of the preset clustering algorithm can be effectively used to generate classification criteria that are communicable with business needs and highly interpretable. This results in highly accurate and business-oriented activity level classification criteria. Consequently, the activity level of target users determined based on the target activity level classification criteria is more accurate and more conducive to business operations.
[0075] In one embodiment, multiple activity segmentation criteria are obtained according to the above embodiments. However, not all activity segmentation criteria can stably and accurately segment user activity over time. Therefore, it is necessary to find a target activity segmentation criterion from the multiple activity segmentation criteria, so that the target activity segmentation criterion can be stably effective over time. Based on this, before performing step 1400 above to determine the activity of the target user according to the target activity segmentation criterion among the multiple activity segmentation criteria, the user activity determination method of this application embodiment further includes the following steps 2100 to 2300:
[0076] Step 2100: Construct the second and third optimization objectives.
[0077] In this embodiment, to better meet the needs of business development, the user activity levels segmented based on the activity level segmentation criteria should satisfy two aspects. One aspect is local stability, which means that within a short time interval, the user's state and the overall user distribution should be relatively similar, with minimal changes in state. The other aspect is the observability of state changes over a long time window, which means that within a longer time interval, the user's activity level should exhibit relatively significant changes.
[0078] The second optimization objective is to minimize the change in user activity between any two adjacent time intervals within the first time interval. In other words, the second optimization objective reflects local stability; the smaller the second optimization objective, the greater the local stability of the activity classification criterion. Typically, the local stability of user activity can be represented from three aspects: distribution stability (related to the distribution stability function), minimization of extreme transition probabilities (related to the first transition probability function), and stability of transition probabilities (related to the second transition probability function). In an optional embodiment, constructing the second optimization objective in step 2100 may further include the following steps 2110a to 2140a:
[0079] Step 2110a: Construct a distribution stability function based on the divergence function of the user's activity distribution in any two adjacent first time intervals within the first time interval.
[0080] Understandably, for a relatively stable activity level classification system (e.g., an activity level classification criterion that includes high and medium activity levels, or medium and low activity levels), the changes in the user activity distribution within similar time windows should be extremely small. Here, the changes in user activity distribution within adjacent windows are used to characterize the distribution stability of the activity level classification criterion. KL divergence is an indicator of the similarity between two distributions, and the distribution stability function satisfies:
[0081]
[0082] Where z represents the activity level, such as high activity, medium activity, and low activity, and p t (z) represents the distribution of user activity at time t, where the smaller the distribution stability function, the higher the local stability.
[0083] Step 2120a: Construct the first transition probability function that reflects the first event.
[0084] The first event is the event in which a user's activity level shifts between the highest and lowest activity levels. For example, in an activity level classification system that distinguishes between high, medium, and low activity levels, the shift from high activity to low activity and vice versa is called the first event, or an extreme transition event.
[0085] Understandably, user activity levels should change gradually over a short period. A good activity level classification system should minimize the probability of extreme transition events, i.e., minimize the probability of extreme transitions. For an activity level classification system that distinguishes between high, medium, and low activity levels, a user's state changing from high activity to low activity, and vice versa, is called an extreme transition event, denoted as event p1.
[0086] The first transition probability function, reflecting the first event, satisfies:
[0087]
[0088] The smaller the first transition probability function, i.e. the smaller the extreme transition probability, the higher the local stability.
[0089] Step 2130a: Construct a second transition probability function that reflects the second and third events.
[0090] The second event is an event in which the user's activity level shifts between any two adjacent activity levels, and the third event is an event in which the user's activity level remains unchanged.
[0091] Understandably, other non-extreme transition events can be categorized into two types: one where the user's activity level remains unchanged, such as a user's activity level changing from high to high, denoted as event p2; and another where the user's activity level changes between two adjacent activity levels, such as a user's activity level changing from high to medium or vice versa, denoted as event p3. The second transition probability function, reflecting the second and third events, satisfies:
[0092] min op3=E(p3)-E(p2) (7)
[0093] The smaller the second transition probability function, the higher the local stability.
[0094] Step 2140a: Construct the second optimization objective based on the distribution stability function, the first transition probability function, and the second transition probability function.
[0095] In step 2140a, the first optimization objective loss1, representing local stability, satisfies:
[0096]
[0097] The third optimization objective is to maximize the change in user activity between any two adjacent time intervals within the second time period. It's understandable that one purpose of dividing user activity into segments is to observe long-term changes in user activity; therefore, it's desirable for user activity to show significant changes over a longer time window. In other words, the third optimization objective reflects the observability of state changes over a long time window; the larger the third optimization objective, the greater the change in user activity over the long time window. Here, the long time window interval is denoted as T. L Typically, modeling the observability of state changes over a long window can be approached from two aspects: the dissimilarity of the distribution (related to the distribution dissimilarity function) and maximizing the transition probabilities of adjacent states (related to the fourth transition probability function). In an optional embodiment, this step of constructing the third optimization objective may further include the following steps 2110b to 2120b:
[0098] Step 2110b: Construct a distribution difference function based on the divergence function of the user's activity distribution within any two adjacent second time windows in the second time period.
[0099] In this step 2110b, at a distance T L The distributions of activity levels of two users should be as different as possible, where the distribution difference function satisfies:
[0100]
[0101] The larger the distributional dissimilarity function, the more significant the change in user activity over a long time window.
[0102] Step 2120b: Construct the fourth transition probability function that reflects the fourth event.
[0103] The fourth event is an event in which a user's activity level shifts between any two adjacent activity levels.
[0104] In this step 2120b, at a distance T L The user's state should change as much as possible between two adjacent activity levels; this event is denoted as p4. The fourth transition probability function satisfies:
[0105] max op5=E(p4) (10)
[0106] The larger the fourth transition probability function, the more significant the change in user activity over a long window.
[0107] Step 2130b: Construct the third optimization objective based on the distribution difference function and the fourth transition probability function.
[0108] In step 2130b, the second optimization objective loss2, representing the observability of state changes over a long time window, satisfies:
[0109]
[0110] Step 2200: Construct the fitness function of the global search algorithm based on the second optimization objective and the third optimization objective.
[0111] In this embodiment, the global search algorithm is an algorithm that determines the target activity partitioning criterion from multiple activity partitioning criteria. The global search algorithm may include one of the following: Particle Swarm Optimization (PSO), Simplex Algorithm, Ant Colony Algorithm, and Genetic Algorithm.
[0112] In an optional embodiment, taking the PSO algorithm as an example of a global search algorithm, the fitness function of the global search algorithm can be constructed as follows:
[0113] Minloss = loss1 - loss2 (12)
[0114] Step 2300: Determine the target activity partitioning criterion based on the fitness function and multiple activity partitioning criteria.
[0115] In this embodiment, taking the fitness function as the above formula as an example, step 2300 is specifically implemented as follows:
[0116] Step 1: Initialize a group of N particles. Each particle has two parameters: its position x and its velocity v. The initialization of the particle positions must satisfy constraints, such as the high school activity level classification criteria. Low to medium activity level classification criteria Then the particle position x of the i-th particle i and particle velocity v i satisfy:
[0117]
[0118] Step 2: Calculate the fitness function value loss of the particle, and denote the particle's own historical best position as pbest and the population's best position as gbest.
[0119] Step 3: For each particle, compare the fitness function value (loss) with the particle's own optimal solution (pbest). If it is less than the optimal solution, then use the current particle position (x) as the benchmark. iReplace the individual's historical best position; compare each particle's fitness function value (loss) with the ensemble's historical best solution (gbest); if it is less than the best solution, use the current position (x). i Replace the group in the optimal position;
[0120] Step 4: Calculate and continuously update the particle position and velocity using the particle update formula until the preset number of iterations is reached or the swarm's historical best solution (gbest) converges. At this point, the particle position x corresponding to the swarm's historical best solution (gbest) is determined. i This refers to the target activity level classification standard.
[0121]
[0122] According to the embodiments of this application, user activity is divided using the target activity division criteria determined by the global search algorithm. This optimal activity division criterion can not only satisfy the reasonableness of the division within a single time period, but also the reasonableness of the division on the time axis.
[0123] The user activity determination method provided in this application can be executed by a user activity determination device. This application uses the example of a user activity determination device executing the user activity determination method to illustrate the user activity determination device provided in this application.
[0124] Corresponding to the above embodiments, see [link to relevant documentation]. Figure 4 This application embodiment also provides a user activity determination device 400, which includes an acquisition module 410, a clustering module 420, a first determination module 430, and a second determination module 440.
[0125] The acquisition module 410 is used to acquire first feature information that reflects the activity level of any user.
[0126] Clustering module 420 is used to cluster the first user group according to the first feature information and a preset clustering algorithm to obtain at least one user cluster, wherein one user cluster corresponds to an activity level;
[0127] The first determining module 430 is used to determine multiple activity division criteria based on the at least one user cluster and a preset classification model, wherein each of the activity division criteria involves constraints on the first feature information.
[0128] The second determining module 440 is used to determine the activity level of the target user based on the target activity level division criteria among the multiple activity level division criteria.
[0129] In one embodiment, the acquisition module 410 is specifically configured to: acquire interaction data of the first user group; obtain multiple second feature information based on the interaction data; acquire the sparsity corresponding to each of the multiple second feature information, and the similarity between any two of the multiple second feature information; and determine first feature information to reflect the activity level of any user based on the sparsity corresponding to each of the multiple second feature information and the similarity between any two of the multiple second feature information.
[0130] In one embodiment, the clustering module 420 is specifically configured to: determine the cluster center of the first user group based on the feature values of each user in the first user group corresponding to the first feature information; determine the similarity between each user and the cluster center based on the feature values of each user in the first user group corresponding to the first feature information; and obtain the at least one user cluster based on the preset clustering algorithm and the similarity between each user and the cluster center.
[0131] In one embodiment, the first determining module 430 is specifically used to: determine the first hyperplane corresponding to two user clusters with adjacent activity levels according to the preset classification model; optimize the first hyperplane corresponding to two user clusters with adjacent activity levels according to the first optimization objective to obtain the target hyperplane corresponding to two user clusters with adjacent activity levels; and determine multiple activity level division criteria according to the target hyperplane corresponding to two user clusters with adjacent activity levels.
[0132] The first optimization objective is to ensure that, when the interval between the first hyperplanes corresponding to two user clusters with adjacent activity levels is greater than a preset interval, the number of misclassifications of user activity by the first hyperplanes corresponding to the two user clusters with adjacent activity levels is less than or equal to a preset number.
[0133] In one embodiment, the apparatus further includes a third determining module (not shown in the figure), configured to: construct a second optimization objective and a third optimization objective; wherein the second optimization objective is to minimize the change in user activity between any two adjacent first time periods within a first time period, and the third optimization objective is to maximize the change in user activity between any two adjacent second time periods within a second time period; construct a fitness function for a global search algorithm based on the second optimization objective and the third optimization objective; and determine a target activity partitioning criterion based on the fitness function and multiple activity partitioning criteria.
[0134] In one embodiment, the third determining module is specifically configured to: construct a distribution stability function based on the divergence function of the user's activity distribution in any two adjacent first time intervals within the first time interval; construct a first transition probability function reflecting a first event, wherein the first event is an event in which the user's activity transitions between the highest and lowest activity levels; construct a second transition probability function reflecting a second and a third event, wherein the second event is an event in which the user's activity transitions between any two adjacent activity levels, and the third event is an event in which the user's activity remains unchanged; and construct a second optimization objective based on the distribution stability function, the first transition probability function, and the second transition probability function.
[0135] In one embodiment, the third determining module is specifically used to: construct a distribution difference function based on the divergence function of the distribution of user activity in any two adjacent second time periods within the second time period; construct a fourth transition probability function reflecting a fourth event, wherein the fourth event is an event in which user activity shifts between any two adjacent activity levels; and construct the third optimization objective based on the distribution difference function and the fourth transition probability function.
[0136] According to embodiments of this application, the electronic device selects first feature information reflecting the activity level of any user based on expert experience, and clusters a first user group according to a preset clustering algorithm and the first feature information to obtain user clusters with different activity levels. Then, based on these user clusters and a preset classification model, multiple activity level classification criteria are determined, each involving constraints on the first feature information. That is, through embodiments of this application, the clustering results of the preset clustering algorithm can be effectively used to generate classification criteria that are communicable with business needs and highly interpretable. This results in highly accurate and business-oriented activity level classification criteria. Consequently, the activity level of target users determined based on the target activity level classification criteria is more accurate and more conducive to business operations.
[0137] The user activity determination device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0138] The user activity determination device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0139] The user activity determination device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0140] Optionally, such as Figure 5 As shown, this application embodiment also provides an electronic device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they implement the various steps of the above-described user activity determination method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0141] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0142] Figure 6 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0143] The electronic device 1000 includes, but is not limited to, components such as: radio frequency unit 1001, network module 1002, audio output unit 1003, input unit 1004, sensor 1005, display unit 1006, user input unit 1007, interface unit 1008, memory 1009, and processor 1010.
[0144] Those skilled in the art will understand that the electronic device 1000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0145] The processor 1010 is configured to: acquire first feature information reflecting the activity level of any user; cluster a first user group according to the first feature information and a preset clustering algorithm to obtain at least one user cluster, wherein each user cluster corresponds to a certain level of activity; determine multiple activity level classification criteria according to the at least one user cluster and a preset classification model, wherein each activity level classification criterion involves constraints on the first feature information; and determine the activity level of a target user according to a target activity level classification criterion among the multiple activity level classification criteria.
[0146] Optionally, the processor 1010 is configured to acquire interaction data of the first user group; obtain multiple second feature information based on the interaction data; acquire the sparsity corresponding to each of the multiple second feature information, and the similarity between any two of the multiple second feature information; and determine first feature information reflecting the activity level of any user based on the sparsity corresponding to each of the multiple second feature information and the similarity between any two of the multiple second feature information.
[0147] Optionally, the processor 1010 is configured to determine the cluster center of the first user group based on the feature values of each user in the first user group corresponding to the first feature information; determine the similarity between each user and the cluster center based on the feature values of each user in the first user group corresponding to the first feature information; and obtain the at least one user cluster based on the preset clustering algorithm and the similarity between each user and the cluster center.
[0148] Optionally, the processor 1010 is configured to: determine, according to the preset classification model, the first hyperplanes corresponding to two user clusters with adjacent activity levels; optimize the first hyperplanes corresponding to the two user clusters with adjacent activity levels according to a first optimization objective to obtain target hyperplanes corresponding to the two user clusters with adjacent activity levels; and determine multiple activity level division criteria based on the target hyperplanes corresponding to the two user clusters with adjacent activity levels. The first optimization objective is to ensure that, when the interval between the first hyperplanes corresponding to two user clusters with adjacent activity levels is greater than a preset interval, the number of misclassifications of user activity by the first hyperplanes corresponding to the two user clusters with adjacent activity levels is less than or equal to a preset number.
[0149] Optionally, the processor 1010 is configured to construct a second optimization objective and a third optimization objective; wherein, the second optimization objective is to minimize the change in user activity between any two adjacent first time periods within a first time period, and the third optimization objective is to maximize the change in user activity between any two adjacent second time periods within a second time period; based on the second optimization objective and the third optimization objective, a fitness function for a global search algorithm is constructed; and based on the fitness function and multiple activity partitioning criteria, a target activity partitioning criterion is determined.
[0150] Optionally, the processor 1010 is configured to: construct a distribution stability function based on the divergence function of the user's activity distribution in any two adjacent first time intervals within the first time interval; construct a first transition probability function reflecting a first event, wherein the first event is an event in which the user's activity transitions between the highest and lowest activity levels; construct a second transition probability function reflecting a second and a third event, wherein the second event is an event in which the user's activity transitions between any two adjacent activity levels, and the third event is an event in which the user's activity remains unchanged; and construct a second optimization objective based on the distribution stability function, the first transition probability function, and the second transition probability function.
[0151] Optionally, the processor 1010 is configured to construct a distribution difference function based on the divergence function of the user's activity distribution in any two adjacent second time periods within the second time period; construct a fourth transition probability function reflecting a fourth event, wherein the fourth event is an event in which the user's activity shifts between any two adjacent activity levels; and construct the third optimization objective based on the distribution difference function and the fourth transition probability function.
[0152] According to embodiments of this application, the electronic device selects first feature information reflecting the activity level of any user based on expert experience, and clusters a first user group according to a preset clustering algorithm and the first feature information to obtain user clusters with different activity levels. Then, based on these user clusters and a preset classification model, multiple activity level classification criteria are determined, each involving constraints on the first feature information. That is, through embodiments of this application, the clustering results of the preset clustering algorithm can be effectively used to generate classification criteria that are communicable with business needs and highly interpretable. This results in highly accurate and business-oriented activity level classification criteria. Consequently, the activity level of target users determined based on the target activity level classification criteria is more accurate and more conducive to business operations.
[0153] It should be understood that, in this embodiment, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0154] The memory 1009 can be used to store software programs and various data. The memory 1009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1009 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0155] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor 1010.
[0156] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described user activity determination method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0157] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0158] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described user activity determination method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0159] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0160] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the user activity determination method embodiment described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0161] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0163] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for determining user activity, characterized in that, The method includes: Obtain primary feature information that reflects the activity level of any user; The first user group is clustered according to the first feature information and the preset clustering algorithm to obtain at least one user cluster, wherein each user cluster corresponds to an activity level. Based on the at least one user cluster and a preset classification model, multiple activity level division criteria are determined, wherein each of the activity level division criteria involves constraints on the first feature information. The activity level of the target user is determined based on the target activity level criteria among the multiple activity level criteria. The step of determining multiple activity level classification criteria based on the at least one user cluster and a preset classification model includes: determining the first hyperplane corresponding to two user clusters with adjacent activity levels according to the preset classification model; optimizing the first hyperplane corresponding to the two user clusters with adjacent activity levels according to a first optimization objective to obtain target hyperplanes corresponding to the two user clusters with adjacent activity levels; determining multiple activity level classification criteria based on the target hyperplanes corresponding to the two user clusters with adjacent activity levels; the first optimization objective is that, when the interval between the first hyperplanes corresponding to the two user clusters with adjacent activity levels is greater than a preset interval, the number of misclassifications of user activity by the first hyperplanes corresponding to the two user clusters with adjacent activity levels is less than or equal to a preset number.
2. The method according to claim 1, characterized in that, The acquisition of the first feature information reflecting the activity level of any user includes: Obtain the interaction data of the first user group; Based on the interaction data, multiple second feature information is obtained; Obtain the sparsity corresponding to each of the multiple second feature information, and the similarity between any two of the multiple second feature information; Based on the sparsity corresponding to the multiple second feature information and the similarity between any two second feature information among the multiple second feature information, a first feature information for reflecting the activity of any user is determined.
3. The method according to claim 1, characterized in that, The step of clustering the first user group according to the first feature information and a preset clustering algorithm to obtain at least one user cluster includes: Based on the feature values of each user in the first user group corresponding to the first feature information, the cluster center of the first user group is determined. Based on the feature values of each user in the first user group corresponding to the first feature information, the similarity between each user and the cluster center is determined. Based on the preset clustering algorithm and the similarity between each user and the cluster center, at least one user cluster is obtained.
4. The method according to claim 1, characterized in that, Before determining the activity level of a target user based on the target activity level criteria among the multiple activity level criteria, the method further includes: Construct a second optimization objective and a third optimization objective; wherein, the second optimization objective is to minimize the change in user activity between any two adjacent time periods within the first time period, and the third optimization objective is to maximize the change in user activity between any two adjacent time periods within the second time period. Based on the second optimization objective and the third optimization objective, a fitness function for the global search algorithm is constructed. The target activity partitioning criterion is determined based on the fitness function and multiple activity partitioning criteria.
5. The method according to claim 4, characterized in that, The construction of the second optimization objective includes: Based on the divergence function of the user activity distribution in any two adjacent time intervals within the first time interval, construct the distribution stability function; Construct a first transition probability function that reflects the first event, where the first event is the event in which the user's activity level transitions between the highest and lowest activity levels; Construct a second transition probability function that reflects the second event and the third event, wherein the second event is the event in which the user's activity level shifts between any two adjacent activity levels, and the third event is the event in which the user's activity level remains unchanged; The second optimization objective is constructed based on the distribution stability function, the first transition probability function, and the second transition probability function.
6. The method according to claim 4, characterized in that, The construction of the third optimization objective includes: Construct a distribution difference function based on the divergence function of the distribution of user activity in any two adjacent second time periods within the second time period; Construct a fourth transition probability function that reflects the fourth event, where the fourth event is the event in which a user's activity level shifts between any two adjacent activity levels; The third optimization objective is constructed based on the distribution difference function and the fourth transition probability function.
7. A user activity determination device, characterized in that, The device includes: The acquisition module is used to acquire primary feature information that reflects the activity level of any user. The clustering module is used to cluster the first user group according to the first feature information and the preset clustering algorithm to obtain at least one user cluster, wherein one user cluster corresponds to an activity level. The first determining module is used to determine multiple activity division criteria based on the at least one user cluster and a preset classification model, wherein each of the activity division criteria involves constraints on the first feature information. The second determining module is used to determine the activity level of the target user based on the target activity level division criteria among the multiple activity level division criteria. Specifically, the first determining module is used to determine, according to the preset classification model, the first hyperplanes corresponding to two user clusters with adjacent activity levels; optimize the first hyperplanes corresponding to two user clusters with adjacent activity levels according to a first optimization objective to obtain target hyperplanes corresponding to two user clusters with adjacent activity levels; and determine multiple activity level division criteria based on the target hyperplanes corresponding to two user clusters with adjacent activity levels. The first optimization objective is to ensure that, when the interval between the first hyperplanes corresponding to two user clusters with adjacent activity levels is greater than a preset interval, the number of misclassifications of user activity by the first hyperplanes corresponding to the two user clusters with adjacent activity levels is less than or equal to a preset number.
8. The apparatus according to claim 7, characterized in that, The acquisition module is specifically used for: Obtain the interaction data of the first user group; Based on the interaction data, multiple second feature information is obtained; Obtain the sparsity corresponding to each of the multiple second feature information, and the similarity between any two of the multiple second feature information; Based on the sparsity corresponding to the multiple second feature information and the similarity between any two second feature information among the multiple second feature information, a first feature information for reflecting the activity of any user is determined.
9. The apparatus according to claim 7, characterized in that, The clustering module is specifically used for: Based on the feature values of each user in the first user group corresponding to the first feature information, the cluster center of the first user group is determined. Based on the feature values of each user in the first user group corresponding to the first feature information, the similarity between each user and the cluster center is determined. Based on the preset clustering algorithm and the similarity between each user and the cluster center, at least one user cluster is obtained.
10. The apparatus according to claim 7, characterized in that, The device further includes a third determining module, used for: Construct a second optimization objective and a third optimization objective; wherein, the second optimization objective is to minimize the change in user activity between any two adjacent time periods within the first time period, and the third optimization objective is to maximize the change in user activity between any two adjacent time periods within the second time period. Based on the second optimization objective and the third optimization objective, a fitness function for the global search algorithm is constructed. The target activity partitioning criterion is determined based on the fitness function and multiple activity partitioning criteria.
11. The apparatus according to claim 10, characterized in that, The third determining module is specifically used for: Based on the divergence function of the user activity distribution in any two adjacent time intervals within the first time interval, construct the distribution stability function; Construct a first transition probability function that reflects the first event, where the first event is the event in which the user's activity level transitions between the highest and lowest activity levels; Construct a second transition probability function that reflects the second event and the third event, wherein the second event is the event in which the user's activity level shifts between any two adjacent activity levels, and the third event is the event in which the user's activity level remains unchanged; The second optimization objective is constructed based on the distribution stability function, the first transition probability function, and the second transition probability function.
12. The apparatus according to claim 10, characterized in that, The third determining module is specifically used for: Construct a distribution difference function based on the divergence function of the distribution of user activity in any two adjacent second time periods within the second time period; Construct a fourth transition probability function that reflects the fourth event, where the fourth event is the event in which a user's activity level shifts between any two adjacent activity levels; The third optimization objective is constructed based on the distribution difference function and the fourth transition probability function.
13. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the user activity determination method as described in any one of claims 1-6.
14. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the user activity determination method as described in any one of claims 1-6.
Citation Information
Patent Citations
Determining active persona of user device
CN106104510A
Method for discovering active user cluster in network community, terminal device and storage medium
CN107749033A