Recommended information determination methods, devices, electronic equipment, and non-volatile storage media

By selecting cluster centers based on the distance and density between data points in big data user profiles, the problem of poor clustering effect caused by human selection in the traditional K-Means algorithm is solved, and more accurate user recommendation information and higher user stickiness are achieved.

CN117312672BActive Publication Date: 2025-11-14CHINA TELECOM CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311309613.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-10
Publication Date
2025-11-14
Estimated Expiration
2043-10-10

AI Technical Summary

Technical Problem

Existing technologies, when using big data to recommend information based on user profiles, rely on manually selecting user data cluster centers, resulting in poor clustering effects and consequently, inaccurate user recommendation information.

Method used

By acquiring the target user dataset, the target density and average density of each target user's data are determined. Based on the distance and density between data points, the cluster centers of the traditional K-Means algorithm are selected, and clustering is performed to obtain a preset number of clusters. Recommendation information is then determined based on user characteristics.

Benefits of technology

It effectively avoids the problem of poor clustering results caused by improper human selection of cluster centers in the traditional K-Means algorithm, improves the accuracy and matching degree of user recommendation information, and enhances user stickiness and marketing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117312672B_ABST
    Figure CN117312672B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, electronic device, and non-volatile storage medium for determining recommendation information. The method includes: acquiring a target user dataset; determining the target density and the average density of each target user dataset based on user attribute parameters; determining a predetermined number of cluster centers among the data points corresponding to the target user data based on the target density and average density; clustering the target user data according to the cluster centers to obtain a predetermined number of clusters; determining the user characteristics of the target user data within the clusters; and determining the recommendation information corresponding to the clusters based on the user characteristics. This application solves the technical problem of poor clustering results and inaccurate user recommendation information caused by the manual selection of user data cluster centers in related technologies when using big data to recommend information based on user profiles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data mining technology, and more specifically, to a method, apparatus, electronic device, and non-volatile storage medium for determining recommendation information. Background Technology

[0002] Currently, the communications market has reached saturation, and the development model of continuously increasing user numbers to increase profits is no longer applicable. It has become crucial to utilize data mining technology to create user profiles from multiple perspectives based on big data in order to increase user stickiness.

[0003] However, when using big data to recommend information based on user profiles, related technologies often rely on manually selecting user data cluster centers, resulting in poor clustering effects and consequently, inaccurate user recommendation information.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides a method, apparatus, electronic device, and non-volatile storage medium for determining recommendation information, in order to at least solve the technical problem that the related technologies, when using big data to recommend information based on user profiles, employ a manual selection of user data cluster centers, resulting in poor clustering effects and thus poor accuracy of user recommendation information.

[0006] According to one aspect of the embodiments of this application, a method for determining recommendation information is provided, comprising: acquiring a target user dataset, wherein the target user dataset contains multiple target user data, and each target user data includes multiple user attribute parameters; determining a target density and an average density corresponding to each target user data based on the user attribute parameters, wherein the target density is used to characterize the density of the distribution of data points corresponding to other target user data around a data point corresponding to a target user data in the target user dataset; determining a preset number of cluster centers among the data points corresponding to the target user data based on the target density and the average density, and clustering the target user data based on the cluster centers to obtain a preset number of clusters; determining user features of the target user data in the clusters, and determining recommendation information corresponding to the clusters based on the user features.

[0007] Optionally, determining the target density of each target user data and the average density of the target user dataset based on user attribute parameters includes: determining the target distance between any pairwise data points of the target user data in the target user dataset based on user attribute parameters, and determining the average distance of the target user dataset based on the target distance; determining the target density of the target user data based on the target distance and the average distance; and calculating the average value of the target density of the target user data in the target user dataset to obtain the average density of the target user dataset.

[0008] Optionally, determining the target density of the target user data based on the target distance and the average distance includes: adding the remaining target user data in the target user data set whose target distance to the data point corresponding to the first user data is no greater than the average distance to the first user data in the first data set, wherein the first user data is any target user data in the target user data set; counting the number of target user data in the first data set to obtain a first quantity; and calculating the target density of the first user data based on the target distance between the first user data and each target user data in the first data set, and the first quantity.

[0009] Optionally, determining a preset number of cluster centers among the data points corresponding to the target user data based on the target density and average density includes: adding target user data with a target density not less than the average density in the target user dataset to the second dataset; determining the data points corresponding to the target user data with the highest target density in the second dataset as cluster centers and adding the cluster centers to the set of center points; determining new cluster centers among the data points corresponding to the target user data in the second dataset based on the first distance from each data point corresponding to the target user data in the second dataset to the set of center points, and adding them to the set of center points, until the number of cluster centers in the set of center points equals the preset number.

[0010] Optionally, determining new cluster centers among the data points corresponding to the target user data in the second dataset includes: determining a first distance corresponding to the target user data in the second dataset, wherein the first distance is the minimum target distance among the target distances from the data points corresponding to the target user data in the second dataset to each cluster center in the set of center points; determining the data points corresponding to the target user data with the largest first distance in the second dataset as new cluster centers, and adding the cluster centers to the set of center points.

[0011] Optionally, determining the recommendation information corresponding to the cluster based on user characteristics includes: randomly selecting a preset proportion of target user data in the cluster for feature extraction to obtain user characteristics, wherein the user characteristics are used to characterize at least one of the data traffic usage, voice call usage, and wireless network traffic usage of the user group corresponding to the cluster, and the usage level is used to characterize the amount / duration of data used for traffic / calls; determining the recommendation information corresponding to the user characteristics, and sending the recommendation information to the mobile terminal device corresponding to the target user data.

[0012] Optionally, obtaining the target user dataset includes: obtaining the original user dataset, wherein the original user dataset includes at least one of the following: user account information, communication records; obtaining the target attributes of the original user data in the original user dataset, wherein the target attributes include at least one of the following: average monthly outgoing call duration, average monthly number of calls, average monthly number of inter-network calls, average monthly data traffic, average monthly Wi-Fi traffic, average monthly number of point-to-point SMS / MMS messages sent; and correcting outliers and / or missing values ​​in the target attributes of the obtained original user data using corresponding default values ​​to obtain the target user dataset.

[0013] According to another aspect of the embodiments of this application, a recommendation information determination apparatus is also provided, comprising: a user data acquisition module, configured to acquire a target user dataset, wherein the target user dataset contains multiple target user data, and each target user data includes multiple user attribute parameters; a target density determination module, configured to determine the target density of each target user data and the average density corresponding to the target user dataset based on the user attribute parameters, wherein the target density is used to characterize the density of the distribution of data points corresponding to other target user data around a data point corresponding to a target user data in the target user dataset; a cluster center determination module, configured to determine a preset number of cluster centers among the data points corresponding to the target user data based on the target density and the average density, and to cluster the target user data based on the cluster centers to obtain a preset number of clusters; and a recommendation information determination module, configured to determine the user characteristics of the target user data in the clusters, and to determine the recommendation information corresponding to the clusters based on the user characteristics.

[0014] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the program executes a recommendation information determination method during runtime.

[0015] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, the non-volatile storage medium including a stored computer program, wherein the device where the non-volatile storage medium is located executes a recommendation information determination method by running the computer program.

[0016] In this embodiment, a target user dataset is obtained, which contains multiple target user data points, each of which includes multiple user attribute parameters. Based on the user attribute parameters, the target density and the average density of the target user dataset are determined. The target density characterizes the density of the distribution of data points of other target user data points around a data point in the target user dataset. Based on the target density and average density, a preset number of cluster centers are determined among the data points of the target user data points. The target user data is then clustered based on these cluster centers to obtain a preset number of clusters. The user characteristics of the target user data in each cluster are determined, and the recommendation information corresponding to the cluster is determined based on these user characteristics. By selecting cluster centers based on the distance and density between data points in the dataset, the traditional K-Means algorithm's clustering effect is avoided from being significantly affected by improper human selection of cluster centers. This solves the technical problem of poor clustering effect and inaccurate user recommendation information caused by manually selecting user data cluster centers when using big data to recommend information based on user profiles. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0018] Figure 1 This is a hardware structure block diagram of a computer terminal (or electronic device) for implementing a method for determining recommendation information, according to an embodiment of this application.

[0019] Figure 2 This is a schematic diagram of a method for determining recommendation information according to an embodiment of this application;

[0020] Figure 3 This is a schematic diagram of a method for improving user stickiness based on data mining, according to an embodiment of this application.

[0021] Figure 4 This is a schematic diagram of a recommendation information determining device provided in an embodiment of this application. Detailed Implementation

[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0024] With the development of internet technology, the world has entered the era of big data. Every industry is constantly generating massive amounts of data, leading to the rapid development of data mining technology. Major telecommunications operators carry various types of user data, including basic user information, call data, mobile internet data, and various consumption data, all characterized by their massive volume, continuity, and stability.

[0025] Currently, the telecommunications market has reached saturation, and the development model of continuously increasing user numbers to boost profits is no longer applicable. Traditional marketing and service customer retention methods primarily rely on decision-makers' experience and market feedback from competitors' similar products, lacking precision and targeting, resulting in relatively poor effectiveness. Therefore, leveraging data mining techniques to create multi-faceted user profiles based on big data and pushing relevant recommendations to increase user engagement has become crucial.

[0026] However, when using big data to recommend information based on user profiles, most related technologies employ the traditional K-Means algorithm for clustering. This may result in poor clustering performance due to improper human selection of cluster centers, leading to poor accuracy of user recommendations, poor matching with actual customer needs, and difficulty in improving customer retention rates through recommendations.

[0027] To address the aforementioned issues, this application provides relevant solutions, which are detailed below.

[0028] According to an embodiment of this application, a method embodiment for determining recommendation information is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0029] The methods and embodiments provided in this application can be executed on mobile terminals, computer terminals, or similar computing devices. Figure 1 A hardware block diagram of a computer terminal (or electronic device) for implementing a method for determining recommendation information is shown. Figure 1 As shown, the computer terminal 10 (or electronic device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0030] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or electronic device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0031] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the recommendation information determination method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned recommendation information determination method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0032] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0033] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or electronic device).

[0034] Under the above operating environment, this application provides a method for determining recommendation information. Figure 2 This is a schematic diagram of a method for determining recommendation information according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:

[0035] Step S202: Obtain the target user dataset, wherein the target user dataset contains multiple target user data, and each target user data includes multiple user attribute parameters;

[0036] Step S204: Based on the user attribute parameters, determine the target density of each target user data and the average density of the target user dataset. The target density is used to characterize the density of the distribution of data points corresponding to other target user data around the data point corresponding to a target user data in the target user dataset.

[0037] Step S206: Based on the target density and average density, determine a preset number of cluster centers in the data points corresponding to the target user data, and cluster the target user data according to the cluster centers to obtain a preset number of clusters.

[0038] Step S208: Determine the user characteristics of the target user data in the cluster, and determine the recommendation information corresponding to the cluster based on the user characteristics.

[0039] By using the above steps to select cluster centers for the traditional K-Means algorithm based on the distance and density between data points in the dataset, the algorithm avoids the significant impact on clustering results caused by improper human selection of cluster centers. This solves the technical problem of poor clustering results and inaccurate user recommendation information caused by manually selecting user data cluster centers when using big data to recommend information based on user profiles.

[0040] The method for determining recommendation information in steps S202 to S208 of the embodiments of this application will be further described below.

[0041] Figure 3 This is a schematic diagram of a method for improving user stickiness based on data mining, according to an embodiment of this application. Figure 3 As shown, the method includes the following steps.

[0042] First, obtain the original user dataset D, which includes at least one of the following: user account information, communication records, user basic information, user call data, mobile internet data, various consumption data, etc. For example, the original user data can be obtained through the operator's central business support system (cBSS), business support system (BSS), and some local databases.

[0043] It should be noted that all relevant information and data involved in this application (including but not limited to various data and information in the original user dataset) are information and data authorized by the user or fully authorized by all parties. For example, if there is an interface between this system and the relevant user or organization, before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information after receiving consent from the aforementioned user or organization.

[0044] To extract user features from user data more accurately, after obtaining the original user dataset D, data preprocessing can be performed on the original user dataset D to obtain the target user dataset D′. The specific steps are as follows.

[0045] In some embodiments of this application, obtaining the target user dataset includes the following steps: obtaining the target attributes of the original user data in the original user dataset, wherein the target attributes include at least one of the following: average monthly outgoing call duration, average monthly number of callers, average monthly number of inter-network callers, average monthly data traffic, average monthly wireless network traffic, and average monthly number of point-to-point SMS / MMS messages sent; correcting the outliers and / or missing values ​​in the target attributes of the obtained original user data using the corresponding default values ​​to obtain the target user dataset.

[0046] Specifically, during preprocessing, unnecessary interfering data in the original user dataset needs to be filtered out while preserving its relevant characteristics as much as possible. This reduces the complexity of subsequent data processing. For example, based on business experience, user attribute parameters that best reflect user characteristics can be selected for subsequent data processing, eliminating data from other attributes. These target parameters include: average monthly outgoing call duration (h), average monthly number of calls, average monthly number of inter-network calls, average monthly data traffic (Gb), average monthly Wireless Local Area Network (WLAN) traffic (Gb) (i.e., the aforementioned average monthly wireless network traffic), and average monthly amount of point-to-point SMS / MMS messages sent. Additionally, associated user package name, network type, whether there is a contract phone, whether there is broadband under the same ID card, development channel, network access time, and average monthly billing amount are also considered.

[0047] After obtaining the target user dataset D′ through preprocessing, it is necessary to perform clustering of the target user data in the target user dataset D′ with a predetermined number of s categories, based on the user attribute parameters of the target user data. s}, and the cluster C for each category. i All target user data points included in

[0048] The clustering process in this application will be further explained below. Based on the traditional K-Means algorithm, this application introduces the concept of distance density to obtain the improved DKM (Distentivity K-Means) clustering algorithm. The specific steps of this algorithm are as follows.

[0049] First, based on the user attribute parameters, determine the target density of each target user data and the average density of the target user dataset. The target density is used to characterize the density of the distribution of data points of other target user data around the data point corresponding to a target user data in the target user dataset.

[0050] In some embodiments of this application, determining the target density of each target user data and the average density of the target user dataset based on user attribute parameters includes the following steps: determining the target distance between any pairwise data points corresponding to the target user data in the target user dataset based on user attribute parameters, and determining the average distance of the target user dataset based on the target distance; determining the target density of the target user data based on the target distance and the average distance; calculating the average value of the target density of the target user data in the target user dataset to obtain the average density of the target user dataset.

[0051] Specifically, first calculate the target distance Dis between any pair of data points corresponding to each target user data in the target user dataset D′. For example, the target distance Dis can be Euclidean distance, as shown in the following formula.

[0052]

[0053] Where p,o are k-dimensional data points (i.e., data points corresponding to the target user data), k is the number of user attribute parameters corresponding to each target user data, and x i ,y i These are the user attribute parameters of p and o in dimension i, respectively.

[0054] Then, the average distance MDis between data points is calculated based on the target distance Dis, as shown in the following formula.

[0055]

[0056] Where x in the above formula i ,x j Let C be the data point corresponding to the target user data in the target user dataset D′, where n is the number of target user data points in the target user dataset D′, and C is the number of target user data points in the target user dataset D′. n 2 It is the number of pairs extracted from the data of n target users.

[0057] After obtaining the average distance MDis and target distance Dis corresponding to the target user dataset D′, the target density Den and average density MDe of the target user dataset D′ can be calculated based on the average distance MDis and target distance Dis. The specific steps are as follows.

[0058] In some embodiments of this application, determining the target density of target user data based on the target distance and the average distance includes the following steps: adding the remaining target user data in the target user data set whose target distance to the data point corresponding to the first user data is not greater than the average distance to the first user data in the target user data set, wherein the first user data is any target user data in the target user data set; counting the number of target user data in the first data set to obtain a first quantity; and calculating the target density of the first user data based on the target distance between the first user data and each target user data in the first data set, and the first quantity.

[0059] Specifically, first determine the first user data x i The most recent first quantity of m target user data x j (i.e., the first dataset mentioned above), for example, it can be determined that the first user data x i The target distance of the corresponding data point is no greater than the average distance of the other target user data x j Or, with the first user data x i The target distance of the corresponding data points is not greater than the average distance, and the difference between the target distances and the average distances is not less than a preset difference (e.g., 1) for the remaining target user data x. j As shown in the following formula.

[0060]

[0061] Where n is the number of target user data in the target user dataset D′, and u(x) is calculated as shown in the following formula.

[0062]

[0063] Then, based on the first user data x j Compared with the target user data x in the first dataset j Distance between targets Dis(x) i ,x j The target density Den(x) of the first user data is obtained by calculating the first quantity m. i The specific formula is as follows.

[0064]

[0065] Where, x j For x i One of the most recent m data points, where e is the natural constant.

[0066] After obtaining the target density Den, the average density MDe can be calculated based on the target density Den, as shown in the following formula.

[0067]

[0068] After obtaining the target density Den and the average density MDe, a predetermined number of cluster centers can be determined among the data points corresponding to the target user data based on the target density Den and the average density MDe. Then, the target user data is clustered according to these cluster centers to obtain a predetermined number s clusters C = {C1, C2, ..., C...}. s The specific steps are as follows.

[0069] In some embodiments of this application, determining a preset number of cluster centers among the data points corresponding to the target user data based on the target density and the average density includes: adding target user data with a target density not less than the average density in the target user dataset to a second dataset; determining the data points corresponding to the target user data with the highest target density in the second dataset as cluster centers and adding the cluster centers to the set of center points; determining new cluster centers among the data points corresponding to the target user data in the second dataset based on the first distance from the data points corresponding to each target user data in the second dataset to the set of center points, and adding them to the set of center points, until the number of cluster centers in the set of center points equals the preset number.

[0070] Specifically, this can be achieved by first comparing the target density Den of each target user data in the target user dataset D′. i The size of the average density MDe in Den i The target user data ≥MDen is added to the initially empty second dataset T; then, the target density Den is determined in the second dataset T. i The data point D′ corresponding to the largest target user data p The data point is then identified as the initial cluster center for the K-Means algorithm, and this cluster center is added to the set of center points F.

[0071] Then, select the data point D′ that is farthest from the set of center points F in the second dataset T. k As the next new cluster center, it is added to the set of center points. That is, based on the first distance from the data points corresponding to the target user data in the second dataset to the set of center points, the new cluster center is determined among the data points corresponding to the target user data in the second dataset. The specific steps are as follows.

[0072] In some embodiments of this application, determining new cluster centers among the data points corresponding to the target user data in the second dataset includes the following steps: determining a first distance corresponding to the target user data in the second dataset, wherein the first distance is the minimum target distance among the target distances from the data points corresponding to the target user data in the second dataset to each cluster center in the set of center points; determining the data point corresponding to the target user data with the largest first distance in the second dataset as the new cluster center, and adding the cluster center to the set of center points.

[0073] Specifically, the data points x corresponding to each target user data in the second dataset i The first distance D to the center point set F set The calculation method is shown in the following formula.

[0074] D set (x i F) = min(Dis(x) i ,x j ),x j ∈F)

[0075] As can be seen from the above formula, the first distance is the data point x corresponding to the target user data in the second dataset. i The minimum target distance among all target distances to the cluster centers in the centroid set F. Each time a new cluster center is determined, it is added to the centroid set F until the centroid set F contains a preset number of cluster centers, s. The preset number s can be determined empirically based on the specific user dataset.

[0076] After obtaining a preset number of s cluster centers, clustering can be performed based on the distances between the remaining ns target user data points in the target user dataset D′ and the s cluster centers. The target user data points in the target user dataset D′ are then sequentially assigned to the classes corresponding to the respective cluster centers, resulting in s clusters C = {C1, C2, ..., C...}. s}, and the cluster C for each category. i All target user data contained therein.

[0077] Finally, user analysis can be performed on the target user data in each cluster to obtain user characteristics. Based on these characteristics, corresponding customer retention measures can be implemented to improve user stickiness. For example, recommendation information corresponding to the user characteristics of the cluster can be sent to the customer's mobile terminal. The specific steps are as follows.

[0078] In some embodiments of this application, determining the recommendation information corresponding to a cluster based on user characteristics includes the following steps: randomly selecting a preset proportion of target user data in the cluster for feature extraction to obtain user characteristics, wherein the user characteristics are used to characterize at least one of the data traffic usage, voice call usage, and wireless network traffic usage of the user group corresponding to the cluster, and the usage level is used to characterize the amount of data used / duration of traffic / call usage; determining the recommendation information corresponding to the user characteristics, and sending the recommendation information to the mobile terminal device corresponding to the target user data.

[0079] Specifically, the clustering result C = {C1, C2, ..., C} can be analyzed. k Cluster C for each category in} i A predetermined percentage (e.g., 30%) of target user data is randomly selected for further analysis to extract user characteristics. Since the target user data has already been clustered into a predetermined number of clusters (s), the target user data in each cluster is similar and shares similar user characteristics. Therefore, only random sampling analysis is needed to obtain the individual characteristics of the users corresponding to each cluster. An example is provided below.

[0080] For example, analysis of target user data within a specific cluster reveals that these users consume a large amount of mobile data per month but almost no Wi-Fi (WLAN) data, and most make very few voice calls daily. This suggests that the users in this cluster are likely younger, heavily reliant on smart devices for internet access, such as watching videos and playing games. Furthermore, these users may communicate more with instant messaging software and less with traditional voice functions.

[0081] Furthermore, further analysis of user acquisition channels, whether broadband is available under the same ID card, and whether they own contract phones reveals that these users often join the network through e-commerce and campus marketing, are highly receptive to new things, and exhibit significant autonomy in choosing telecommunications products. Analyzing their average monthly billing amount also indicates that these users are not particularly concerned about their spending on telecommunications products. Therefore, based on these user characteristics, relevant recommendations for contract phone policies can be pushed to them to increase user stickiness. Alternatively, promotional information such as free video streaming software memberships with prepaid subscriptions can be pushed to retain users.

[0082] For example, analysis of target user data in another cluster revealed that users in this cluster consume approximately 15GB of data per month and have an average monthly outgoing call duration of 2-3 hours, which is considered moderate. Their average monthly WLAN traffic usage, however, is relatively high. This suggests that these users heavily rely on internet access via smart devices, such as for watching videos. However, they are also cost-conscious and will likely use WLAN more frequently to save money. Furthermore, they also make greater use of traditional voice services.

[0083] Therefore, these users are more likely to choose plans with more data and a certain amount of voice call time. For these users, you can push relevant information about preferential policies such as deposit-and-get-free credit to increase user stickiness. Alternatively, you can compare product types and push relevant information about plans with more data at the same price to meet their internet access needs in the absence of WLAN.

[0084] Alternatively, user characteristics obtained by analyzing the clustered target user data can be used to determine the user category corresponding to the cluster. For different user categories, corresponding recommendation information can be pushed. For example, for category 1 users, recommendation information related to upgrading from 4G to 5G, changing terminals, and adding integration can be pushed for business recommendation and marketing; for category 2 users, recommendation information related to adding contracts or benefits can be pushed to stabilize and retain users.

[0085] This application introduces the concept of distance-density to select cluster centers for the K-Means algorithm based on the distance and density between data points in the dataset. This avoids the shortcomings of related clustering algorithms, which often suffer from poor clustering results due to improper human selection of cluster centers, as well as the influence of isolated and discrete points on the clustering results. It also effectively improves the shortcomings of related clustering algorithms, which are prone to getting trapped in local optima. Based on the clustering results obtained from clustering user data, this application analyzes user characteristics and implements corresponding customer retention measures to enhance user stickiness, thereby improving marketing efficiency, retention rate and matching degree of existing users, and enhancing user perception and loyalty.

[0086] According to an embodiment of this application, an embodiment of a recommendation information determination device is also provided. Figure 4 This is a schematic diagram of a recommendation information determining device provided according to an embodiment of this application. Figure 4 As shown, the device includes:

[0087] User data acquisition module 40 is used to acquire target user dataset, wherein the target user dataset contains multiple target user data, and each target user data includes multiple user attribute parameters;

[0088] In some embodiments of this application, obtaining the target user dataset includes: obtaining an original user dataset, wherein the original user dataset includes at least one of the following: user account information, communication records; obtaining target attributes of the original user data in the original user dataset, wherein the target attributes include at least one of the following: average monthly outgoing call duration, average monthly number of calls, average monthly number of inter-network calls, average monthly data traffic, average monthly Wi-Fi traffic, average monthly number of point-to-point SMS / MMS messages sent; and correcting outliers and / or missing values ​​in the target attributes of the obtained original user data using corresponding default values ​​to obtain the target user dataset.

[0089] The target density determination module 42 is used to determine the target density of each target user data and the average density of the target user dataset based on the user attribute parameters. The target density is used to characterize the density of the distribution of data points corresponding to other target user data around the data point corresponding to a target user data in the target user dataset.

[0090] In some embodiments of this application, determining the target density of each target user data and the average density corresponding to the target user dataset based on user attribute parameters includes: determining the target distance between any pairwise data points corresponding to the target user data in the target user dataset based on user attribute parameters, and determining the average distance corresponding to the target user dataset based on the target distance; determining the target density of the target user data based on the target distance and the average distance; and calculating the average value of the target density of the target user data in the target user dataset to obtain the average density corresponding to the target user dataset.

[0091] In some embodiments of this application, determining the target density of target user data based on the target distance and the average distance includes: adding the remaining target user data in the target user data set whose target distance to the data point corresponding to the first user data is not greater than the average distance to the first user data in the target user data set, wherein the first user data is any target user data in the target user data set; counting the number of target user data in the first data set to obtain a first quantity; and calculating the target density of the first user data based on the target distance between the first user data and each target user data in the first data set, and the first quantity.

[0092] The cluster center determination module 44 is used to determine a preset number of cluster centers in the data points corresponding to the target user data based on the target density and average density, and to cluster the target user data based on the cluster centers to obtain a preset number of clusters.

[0093] In some embodiments of this application, determining a preset number of cluster centers among the data points corresponding to the target user data based on the target density and the average density includes: adding target user data with a target density not less than the average density in the target user dataset to a second dataset; determining the data points corresponding to the target user data with the highest target density in the second dataset as cluster centers and adding the cluster centers to the set of center points; determining new cluster centers among the data points corresponding to the target user data in the second dataset based on the first distance from the data points corresponding to each target user data in the second dataset to the set of center points, and adding them to the set of center points, until the number of cluster centers in the set of center points equals the preset number.

[0094] In some embodiments of this application, determining new cluster centers among the data points corresponding to the target user data in the second dataset includes: determining a first distance corresponding to the target user data in the second dataset, wherein the first distance is the minimum target distance among the target distances from the data points corresponding to the target user data in the second dataset to each cluster center in the set of center points; determining the data point corresponding to the target user data with the largest first distance in the second dataset as the new cluster center, and adding the cluster center to the set of center points.

[0095] The recommendation information determination module 46 is used to determine the user characteristics of the target user data in the cluster, and determine the recommendation information corresponding to the cluster based on the user characteristics.

[0096] In some embodiments of this application, determining the recommendation information corresponding to a cluster based on user characteristics includes: randomly selecting a preset proportion of target user data in the cluster for feature extraction to obtain user characteristics, wherein the user characteristics are used to characterize at least one of the data traffic usage, voice call usage, and wireless network traffic usage of the user group corresponding to the cluster, and the usage level is used to characterize the amount / duration of data used for traffic / calls; determining the recommendation information corresponding to the user characteristics, and sending the recommendation information to the mobile terminal device corresponding to the target user data.

[0097] It should be noted that the modules in the above-mentioned recommendation information determining device can be program modules (e.g., a set of program instructions to implement a certain function) or hardware modules. For the latter, they can be in the following forms, but are not limited to these: each of the above modules is in the form of a processor, or the functions of each of the above modules are implemented by a processor.

[0098] It should be noted that the recommendation information determination device provided in this embodiment can be used to perform... Figure 2The method for determining recommendation information shown above is also applicable to the embodiments of this application, and will not be repeated here.

[0099] This application embodiment also provides a non-volatile storage medium, which includes a stored computer program. The device containing the non-volatile storage medium executes the following recommendation information determination method by running the computer program: acquiring a target user dataset, wherein the target user dataset contains multiple target user data points, each of which includes multiple user attribute parameters; determining the target density of each target user data point and the average density corresponding to the target user dataset based on the user attribute parameters, wherein the target density characterizes the density of the distribution of data points corresponding to other target user data points around a data point corresponding to a target user data point in the target user dataset; determining a preset number of cluster centers among the data points corresponding to the target user data points based on the target density and average density, and clustering the target user data based on the cluster centers to obtain a preset number of clusters; determining the user characteristics of the target user data in the clusters, and determining the recommendation information corresponding to the clusters based on the user characteristics.

[0100] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0101] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0102] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0103] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0104] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0105] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0106] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for determining recommendation information, characterized in that, include: Obtain a target user dataset, wherein the target user dataset contains multiple target user data, and each target user data includes multiple user attribute parameters; Based on the user attribute parameters, determining the target density of each target user data point and the average density of the target user dataset includes: determining the target distance between any pairwise data points of the target user data points in the target user dataset based on the user attribute parameters, and determining the average distance of the target user dataset based on the target distance; determining the target density of the target user data based on the target distance and the average distance; calculating the average value of the target density of the target user data points in the target user dataset to obtain the average density of the target user dataset, wherein the target density is used to characterize the density of the distribution of data points of other target user data points around a data point of a target user data point in the target user dataset; Determining the target density of the target user data based on the target distance and the average distance includes: adding the remaining target user data in the target user dataset whose target distance to the data point corresponding to the first user data is not greater than the average distance to the first user data, wherein the first user data is any one of the target user data in the target user dataset; counting the number of target user data in the first dataset to obtain a first quantity; and calculating the target density of the first user data based on the target distance between the first user data and each of the target user data in the first dataset, and the first quantity. Based on the target density and the average density, a preset number of cluster centers are determined among the data points corresponding to the target user data, and the target user data is clustered based on the cluster centers to obtain the preset number of clusters; Determine the user characteristics of the target user data in the cluster, and determine the recommendation information corresponding to the cluster based on the user characteristics.

2. The method for determining recommendation information according to claim 1, characterized in that, Determining a preset number of cluster centers among the data points corresponding to the target user data based on the target density and the average density includes: Add the target user data in the target user dataset whose target density is not less than the average density to the second dataset; The data point corresponding to the target user data with the highest target density in the second dataset is determined as the cluster center, and the cluster center is added to the center point set; Based on the first distance from the data points corresponding to each target user data in the second dataset to the set of centroids, new cluster centers are determined among the data points corresponding to the target user data in the second dataset and added to the set of centroids until the number of cluster centers in the set of centroids equals the preset number.

3. The method for determining recommendation information according to claim 2, characterized in that, Determining new cluster centers among the data points corresponding to the target user data in the second dataset includes: Determine the first distance corresponding to the target user data in the second dataset, wherein the first distance is the minimum target distance among the target distances from the data points corresponding to the target user data in the second dataset to each of the cluster centers in the set of center points; The data point corresponding to the target user data with the largest first distance in the second dataset is determined as the new cluster center, and the cluster center is added to the set of center points.

4. The method for determining recommendation information according to claim 1, characterized in that, Based on the user characteristics, the recommendation information corresponding to the cluster includes: A preset proportion of the target user data is randomly selected from the cluster for feature extraction to obtain the user features. The user features are used to characterize at least one of the following: the degree of data traffic usage, the degree of voice call usage, and the degree of wireless network traffic usage of the user group corresponding to the cluster. The degree of usage is used to characterize the amount of data used for traffic / calls / usage duration. The recommended information corresponding to the user characteristics is determined, and the recommended information is sent to the mobile terminal device corresponding to the target user data.

5. The method for determining recommendation information according to claim 1, characterized in that, Obtaining the target user dataset includes: Obtain the original user dataset, wherein the original user dataset includes at least one of the following: user account information, communication records; Obtain the target attributes of the original user data in the original user dataset, wherein the target attributes include at least one of the following: average monthly outgoing call duration, average monthly number of callers, average monthly number of inter-network callers, average monthly data traffic, average monthly wireless network traffic, and average monthly number of point-to-point SMS / MMS messages sent. The outliers and / or missing values ​​in the target attributes of the obtained original user data are corrected using the corresponding default values ​​to obtain the target user dataset.

6. A device for determining recommendation information, characterized in that, include: The user data acquisition module is used to acquire a target user dataset, wherein the target user dataset contains multiple target user data, and each target user data includes multiple user attribute parameters; The target density determination module is used to determine the target density of each target user data and the average density of the target user dataset based on the user attribute parameters. This includes: determining the target distance between any pairwise data points corresponding to the target user data in the target user dataset based on the user attribute parameters, and determining the average distance of the target user dataset based on the target distance; determining the target density of the target user data based on the target distance and the average distance; and calculating the average value of the target densities of the target user data in the target user dataset to obtain the average density of the target user dataset. The target density is used to characterize the density of the distribution of data points corresponding to other target user data around a data point corresponding to one target user data in the target user dataset. Determining the target density of the target user data based on the target distance and the average distance includes: adding the remaining target user data in the target user dataset whose target distance to the data point corresponding to the first user data is not greater than the average distance to the first user data, wherein the first user data is any one of the target user data in the target user dataset; counting the number of target user data in the first dataset to obtain a first quantity; and calculating the target density of the first user data based on the target distance between the first user data and each of the target user data in the first dataset, and the first quantity. The cluster center determination module is used to determine a preset number of cluster centers in the data points corresponding to the target user data based on the target density and the average density, and to cluster the target user data based on the cluster centers to obtain the preset number of clusters; The recommendation information determination module is used to determine the user characteristics of the target user data in the cluster, and determine the recommendation information corresponding to the cluster based on the user characteristics.

7. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, performs the recommendation information determination method according to any one of claims 1 to 5.

8. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored computer program, wherein the device containing the non-volatile storage medium executes the recommendation information determination method according to any one of claims 1 to 5 by running the computer program.

Citation Information

Patent Citations

  • A group recommendation method and device based on customer characteristics

    CN109408562A

  • Differential K-means load clustering method based on center optimization

    CN112819299A