Content distribution optimization method and system based on user portrait
By collecting and standardizing multi-source telecommunications user data, constructing a profile tagging system and performing cluster analysis, a correlation table is generated between the user multi-dimensional profile set and the business content tag set. This solves the problem that traditional content distribution methods cannot accurately process multi-source user data, and achieves the accuracy and comprehensiveness of content distribution.
Patent Information
- Application Number
- CN202511305291.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-12-16
AI Technical Summary
Traditional content distribution methods cannot accurately process multi-source user data, making it difficult to meet the requirements for accurate evaluation and optimization. This results in one-sided and insufficiently targeted content distribution, failing to meet the needs for precise content distribution and performance optimization based on user profiles.
Multi-source telecommunications user data is collected, standardized, and then a user profile and tagging system is constructed. A classifier is generated by training a support vector machine to determine the multi-dimensional user profile set. The telecommunications service content is clustered and tagged to generate a service content tag set. The user multi-dimensional profile set and the service content tag set are matched and optimized to form a user profile-service content association table. A multi-level distribution strategy is formulated to distribute content and optimize the effect feedback.
It achieves accurate and comprehensive content distribution based on user profiles, meets the technical requirements for precise evaluation and optimization, and improves the targeting and effectiveness of content distribution.
Smart Images

Figure CN121146844A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and system for optimizing content distribution based on user profiles. Background Technology
[0002] In the field of telecommunications marketing, the effectiveness of content distribution is crucial for business promotion. Current technologies largely rely on traditional push algorithms for content distribution, which have proven effective in stable, closed information environments. However, with the diversification of user needs, traditional methods have revealed significant limitations when dealing with complex and multi-source user data. Due to the diverse sources and complex characteristics of user data, traditional methods cannot accurately match user needs with business content, resulting in fragmented and insufficiently targeted distribution data, failing to meet the requirements for precise content distribution and performance optimization based on user profiles. Summary of the Invention
[0003] This application provides a content distribution optimization method and system based on user profiles to solve the technical problem that traditional content distribution methods cannot accurately process multi-source user data and are difficult to meet the requirements for accurate evaluation and optimization.
[0004] The first aspect of this application provides a content distribution optimization method based on user profiles. The method includes: collecting multi-source telecommunications user data, including telecommunications service data, user consumption data, and third-party data; standardizing the multi-source telecommunications user data to obtain standard multi-source telecommunications user data; constructing a profile tagging system; evaluating and labeling the standard multi-source telecommunications user data based on the profile tagging system to determine a multi-dimensional user profile set; performing cluster analysis and content tagging on telecommunications service content according to the profile tagging system to generate a service content tag set; matching and optimizing the multi-dimensional user profile set with the service content tag set to determine a user profile-service content association table; parsing the distribution strategy based on the user profile-service content association table to determine a multi-level user content distribution strategy; and optimizing service content distribution and effect feedback through the multi-level user content distribution strategy.
[0005] A second aspect of this application provides a content distribution optimization system based on user profiles. The system includes: a standard user data acquisition module for collecting multi-source telecommunications user data, including telecommunications service data, user consumption data, and third-party data; standardizing the multi-source telecommunications user data to obtain standard multi-source telecommunications user data; a user multi-dimensional profile set acquisition module for constructing a profile tagging system; evaluating and labeling the standard multi-source telecommunications user data based on the profile tagging system to determine a user multi-dimensional profile set; a service content tag set acquisition module for performing cluster analysis and content tagging on telecommunications service content according to the profile tagging system to generate a service content tag set; and a multi-level distribution strategy acquisition module for matching and optimizing the user multi-dimensional profile set with the service content tag set to determine a user profile-service content association table; parsing the distribution strategy based on the user profile-service content association table to determine a multi-level user content distribution strategy; and optimizing service content distribution and effect feedback through the multi-level user content distribution strategy.
[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application collects and standardizes multi-source telecommunications user data, constructs a profile tagging system to evaluate and label it to determine a multi-dimensional user profile set. Simultaneously, it performs cluster analysis and tagging of telecommunications service content to generate a service content tag set. The two are then matched and optimized to determine a correlation table. Based on this, a multi-level user content distribution strategy is derived, and distribution and effect feedback are optimized. This achieves precise content distribution based on user profiles in telecommunications digital data processing, resulting in better content distribution effects. It achieves accurate matching between user profiles and service content, making content distribution more comprehensive and accurate, and meeting the technical requirements for precise evaluation and optimization. Attached Figure Description
[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 This is a flowchart illustrating the content distribution optimization method based on user profiles provided in this application embodiment.
[0009] Figure 2 This is a schematic diagram of the structure of the content distribution optimization system based on user profiles provided in the embodiments of this application.
[0010] Figure labeling: Standard user data acquisition module 1, user multi-dimensional profile acquisition module 2, business content tag acquisition module 3, multi-level distribution strategy acquisition module 4. Detailed Implementation
[0011] This application provides a content distribution optimization method and system based on user profiles to solve the technical problem that traditional content distribution methods cannot accurately process multi-source user data and are difficult to meet the requirements for accurate evaluation and optimization.
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0013] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.
[0014] Example 1, as Figure 1 As shown, a content distribution optimization method based on user profiles is provided, wherein the method includes: Step A100: Collect multi-source telecommunications user data, which includes telecommunications service data, user consumption data, and third-party data. Standardize the multi-source telecommunications user data to obtain standard multi-source telecommunications user data.
[0015] Specifically, the first step is to collect multi-source telecommunications user data, which involves gathering data related to telecommunications users from multiple channels. This includes: telecommunications service data (data generated during user engagement with basic and value-added telecommunications services, such as call duration, call frequency, SMS volume, total and time-of-use data usage, and service transaction records); user consumption data (data generated during telecommunications service consumption, including monthly communication package fees, value-added service subscription amounts, payment records, and consumption frequency); and third-party data (user-related data from partner organizations or platforms outside of telecommunications operators, such as browsing history on partner e-commerce platforms and telecommunications-related interactions on social media platforms). These data collectively constitute multi-source telecommunications user data, reflecting user characteristics from different dimensions.
[0016] Next, based on the data application standards, abnormal conditions containing missing, duplicate, and out-of-range data are set to identify multi-source abnormal user data. After cleaning with matching cleaning strategies, usable data is obtained. Then, standard data is obtained by standardizing according to the data application standards. The specific steps are explained in detail in A110-A140.
[0017] By processing multi-source telecommunications user data, matching user and service content tags, and formulating and optimizing multi-level distribution strategies, accurate content distribution based on user profiles in telecommunications digital data processing has been achieved, improving the content distribution effect.
[0018] Step A200: Construct a profile tagging system, evaluate and label the standard multi-source telecommunications user data based on the profile tagging system, and determine the user multi-dimensional profile set.
[0019] Optionally, firstly, user profile dimensions are divided into five dimensions, including basic information, based on telecommunications business objectives. Each dimension is used as a primary label and further subdivided into multiple levels to obtain a multi-level label set. Value ranges are determined according to business needs. Then, the primary labels, multi-level label sets, and value range sets are hierarchically associated and mapped. The specific steps are explained in detail in A210-A240.
[0020] Next, training data is extracted proportionally from standard multi-source telecommunications user data. Sample data is obtained by evaluating and labeling according to the profile labeling system. A classifier is generated by training and optimizing the support vector machine. The classification label is then evaluated on the standard data based on the classifier, and the user multidimensional profile set is determined. The specific steps are explained in detail in A250-A280.
[0021] Step A300: Perform cluster analysis and content tagging on the telecommunications service content according to the aforementioned profile tagging system to generate a service content tag set.
[0022] In this embodiment, the user profile tagging system is constructed by dividing user profile dimensions according to telecommunications service objectives, using each dimension as a primary profile tag, further subdividing it into multi-level profile tag sets, defining value ranges for multi-level profile tags according to service requirements, and then hierarchically mapping the primary tags, multi-level tag sets, and value range sets. Telecommunications service content refers to various service-related content provided by telecommunications operators.
[0023] In one embodiment of this application, the business content feature dimension is determined according to the profile tag system, the features are extracted to obtain the business content dimension feature set, and after obtaining multiple business content feature clusters through iterative clustering analysis, the business content tag set is generated by tagging according to the profile tag system. The specific steps are described in detail in A310-A340.
[0024] Step A400: Match and optimize the user multidimensional profile set with the business content tag set to determine the user profile-business content association table. Based on the user profile-business content association table, perform distribution strategy analysis to determine the multi-level user content distribution strategy. Then, optimize the business content distribution and effect feedback through the multi-level user content distribution strategy.
[0025] Specifically, firstly, principal component analysis and weight assignment are performed on the user profile information of each dimension to determine the weight factor of the profile dimension. The matching degree set of the user multidimensional profile set and the business content tag set is calculated according to the factor. Then, the user profile-business content association table is determined based on the matching degree set and the association filtering. The specific steps are explained in detail in A410-A430.
[0026] Next, value assessment is performed on each user profile in the user profile-business content association table to obtain a tiered value user set. A business content distribution effect target is preset, and a tiered value user set analysis strategy is performed based on the target to determine a multi-level user content distribution strategy. The specific steps are explained in detail in A440-A450.
[0027] Then, when distributing business content using a multi-level user content distribution strategy, differentiated approaches are matched based on the characteristics of the tiered value user groups: for high-value users, high-end business content that closely matches their profile is pushed and supplemented with dedicated service follow-up; for medium-value users, cost-effective business content is pushed in a targeted manner and incentive measures are added; for low-value users, basic business content is pushed in a low-frequency manner to reduce interference, ensuring that the content received by users at each level is adapted to their needs and value.
[0028] Optimizing the feedback process requires collecting key data after distribution, such as the business conversion rate of users at each level, content clicks, and complaint rates. This data should be compared and analyzed with the preset performance targets. If the conversion rate of a certain user level is lower than the target, issues such as the matching degree between the content and the user profile, and the timing of the push notifications need to be investigated. If the clicks are insufficient, the content presentation format may need to be adjusted. Then, the strategy should be iterated and optimized based on the analysis results, such as adjusting the tag matching rules, push frequency, or content type, forming a closed loop of distribution-feedback-optimization to continuously improve the accuracy and effectiveness of business content distribution.
[0029] Furthermore, step A100 in the method provided in this application embodiment includes: A110: According to the data application standard, set multi-source data anomaly conditions, which include missing data, duplicate data, and out-of-range data.
[0030] A120: Based on the aforementioned multi-source data anomaly conditions, perform anomaly identification on the multi-source telecommunications user data to obtain multi-source abnormal user data.
[0031] A130: Based on the aforementioned multi-source data anomaly conditions, an anomaly data cleaning strategy is matched, and the anomaly data cleaning strategy is used to clean the multi-source abnormal user data to obtain usable multi-source telecommunications user data.
[0032] A140: Based on the data application standard, the available multi-source telecommunications user data is standardized to obtain the standard multi-source telecommunications user data.
[0033] In this embodiment, the data application standard is used as a basis for guiding the processing of multi-source telecommunications user data, and is specifically set and modified by those skilled in the art according to actual business.
[0034] Specifically, firstly, according to the data application standards, setting multi-source data anomaly conditions requires clarifying the specific specifications in the standard regarding the completeness, uniqueness, and reasonableness of values for telecommunications user data. For missing data, based on the standard's requirement that core data fields must be fully recorded, key fields such as mobile phone numbers in basic user information, monthly data usage records in communication behavior data, and package fees in consumption data are identified. If these fields are missing or empty, they are considered missing data. For duplicate data, according to the standard's requirement that the same user's same business behavior data should exist uniquely, data such as value-added service subscription information and call records for the same user at the same time point are identified. If two or more records with completely identical content appear, they are considered duplicate data. For out-of-range data, based on the standard's definition of reasonable value ranges for each data field, such as the normal range for call duration being 0 to 24 hours, monthly data usage not exceeding three times the user's package limit, and consumption amounts being positive and conforming to conventional consumption logic, data values exceeding these preset ranges are considered out-of-range data. By combining the specific requirements of data application standards for data integrity, uniqueness, and reasonableness of values, the abnormal conditions of multi-source data can be set.
[0035] Subsequently, when conducting system screening of the collected multi-source telecommunications user data based on the aforementioned abnormal conditions of missing data, duplicate data, and out-of-range data, it is necessary to first establish a data association index based on user identifiers to aggregate the telecommunications service data, consumption data, and third-party data of the same user. For missing data, the completeness of core fields is verified one by one. For example, checking whether the user's monthly data usage record field has null values, and whether the package fee field in the consumption data has any missing information; if so, it is marked as missing data. For duplicate data, the service records of the same user at the same time dimension are compared, such as checking whether a user's value-added service subscription information within the same minute has multiple completely identical entries; if so, it is marked as duplicate data. For out-of-range data, the values of each data field are checked for range verification, such as checking whether the call duration field has a value exceeding 24 hours, and whether the data usage exceeds three times the package limit; if so, it is marked as out-of-range data. Through this categorized and field-based system screening, the identification of abnormal data in the multi-source telecommunications user data is completed. These data collectively constitute multi-source abnormal user data.
[0036] Next, corresponding cleaning strategies were applied to different types of abnormal data. For missing data traffic records, the average data traffic of the user over the previous three months was used to supplement them; for duplicate subscription information, the earliest record was retained and the remaining duplicates were deleted; for call durations exceeding the allowed range, they were verified and corrected to reasonable values. After cleaning, the remaining data became usable multi-source telecom user data.
[0037] Finally, the available data is standardized based on data application standards. This includes unifying data formats, such as converting traffic units from MB to GB and calculating consumption amounts to two decimal places, ultimately resulting in standardized, multi-source telecommunications user data with consistent format and reasonable values.
[0038] By setting abnormal conditions, identifying and cleaning abnormal data, and performing standardized processing, high-quality standard multi-source telecommunications user data was obtained, providing a reliable data foundation for building accurate user profiles in the future.
[0039] Furthermore, step A200 in the method provided in this application embodiment includes: A210: Based on telecommunications business objectives, user profile dimensions are defined, including basic information dimension, communication behavior dimension, consumption behavior dimension, interest preference dimension, and social relationship dimension.
[0040] A220: Take each of the user profile dimensions as a first-level profile tag, and then subdivide the first-level profile tags into multiple levels to obtain a multi-level profile tag set.
[0041] A230: Divide the multi-level profile tag set into value ranges according to business application requirements, and determine the multi-level profile tag value range set.
[0042] A240: The first-level portrait tag, the multi-level portrait tag set, and the multi-level portrait tag value range set are hierarchically associated and mapped to construct the portrait tag system.
[0043] Optionally, when constructing a user profile tagging system, the user profile dimensions are first determined based on the telecommunications business objectives. For example, if the business objective is to accurately push communication packages, the dimensions would be divided into basic information, communication behavior, consumption behavior, interest preferences, and social relationships. These dimensions reflect the relationship characteristics between users and telecommunications services from different perspectives.
[0044] Next, these five dimensions are used as first-level profile tags, and then further subdivided into multiple levels. Taking the communication behavior dimension as an example, as a first-level tag, it can be subdivided into second-level tags: call behavior, data usage behavior, and SMS behavior; call behavior can be further subdivided into third-level tags: call duration, call frequency, and caller / caller ratio. Specific examples are shown in Table 1, thus forming a multi-level profile tag set containing multiple levels of tags.
[0045] Next, the value ranges of the multi-level profile tag set are divided according to business application requirements. For example, the monthly consumption amount of the three-level tag under the consumption behavior dimension can be divided into value ranges of 0-50 yuan, 51-100 yuan, 101-200 yuan, and above 200 yuan, combined with the needs of package promotion. Other tags are set in the same way, by those skilled in the art according to actual business, to form a set of value ranges for multi-level profile tags.
[0046] Finally, the first-level profile tags, multi-level profile tag sets, and multi-level profile tag value range sets are hierarchically associated and mapped. For example, the first-level tag "social relationship dimension" is associated with the second-level tag "contact features," the contact features are associated with the third-level tags "number of contacts" and "types of high-frequency contacts," and the third-level tag is further associated with the multi-level profile tag value range sets, ultimately constructing a complete profile tag system.
[0047] By dividing dimensions, subdividing tags, dividing intervals, and associating mappings, a clear and hierarchical user profile tagging system is constructed, providing a standardized framework for the evaluation and labeling of user profiles.
[0048] Table 1: Multidimensional Profile Tag Subdivision Table
[0049] Furthermore, step A200 in the method provided in this application embodiment includes: A250: The standard multi-source telecommunications user data is proportionally extracted to obtain telecommunications user training data.
[0050] A260: Perform multi-dimensional evaluation and labeling on the telecommunications user training data according to the aforementioned profile labeling system to obtain telecommunications user sample data.
[0051] A270: Use a support vector machine to perform label classification training and cross-validation optimization on the telecommunications user sample data to generate a profile label classifier.
[0052] A280: Based on the profile label classifier, evaluate, classify, and label the standard multi-source telecommunications user data to determine the user multi-dimensional profile set.
[0053] Specifically, firstly, a proportional sampling method is used to extract data from standard multi-source telecommunications user data, typically employing stratified sampling. This ensures that the distribution of data across different user characteristic dimensions in the training set is consistent with the full dataset, avoiding data bias that could negatively impact subsequent model training. For example, if the standard multi-source telecommunications user data contains 1 million records covering users of different ages, consumption levels, and communication habits, 700,000 records are extracted at a 7:3 ratio as training data. The proportion of users in each age group and consumption range remains consistent with the full dataset, ensuring the representativeness of the training data.
[0054] Next, a multi-dimensional label radiation map is constructed based on the multi-level image label set of the image label system. The coordinate interval of the radiation map is divided according to the value interval set of its multi-level labels and labels are assigned to generate a multi-dimensional sample annotation map. Then, this map is used to perform multi-dimensional evaluation annotation on the training data of telecommunications users to obtain telecommunications user sample data. The specific steps are explained in detail in A261-A264.
[0055] Then, Support Vector Machines (SVMs) were used to train and fine-tune the label classification of the telecommunications user sample data. SVMs can effectively handle high-dimensional feature data and distinguish user features with different labels by finding the optimal classification hyperplane. During training, a 5-fold cross-validation method was used, dividing the sample data into 5 groups. Four groups were used as the training set and one group as the validation set in turn, and the model parameters, such as the penalty coefficient and kernel function parameters, were adjusted iteratively multiple times to improve the model's classification accuracy. For example, the initial model had an accuracy of 75% in identifying high-frequency traffic user labels. After cross-validation fine-tuning, the accuracy improved to 82%, enhancing the model's ability to classify complex user features and ultimately generating a stable and reliable user profile label classifier.
[0056] Subsequently, when evaluating and classifying standard multi-source telecommunications user data based on the user profile label classifier, a batch processing approach is adopted. The entire dataset is input into the classifier in batches for each user. The classifier automatically matches each user's labels at various levels within the user profile label system based on the classification logic developed during training. For example, based on data such as monthly spending, data usage periods, and call frequency, the classifier will label users with tags such as monthly spending of 101-200 yuan, high data usage at night, and high-frequency local calls. This ultimately forms a set containing multi-dimensional feature labels for each user, i.e., a multi-dimensional user profile set.
[0057] By proportionally extracting training data, training and optimizing the classifier using support vector machines, and labeling the entire dataset, a multi-dimensional user profile set that comprehensively reflects user characteristics was obtained, laying the foundation for accurate matching of user profiles with business content in the future.
[0058] Furthermore, step A260 in the method provided in this application embodiment includes: A261: Based on the multi-level image tag set of the image tag system, construct a multi-dimensional tag radiation map.
[0059] A262: Divide the multidimensional label radiation map into coordinate intervals according to the multi-level image label value interval set of the image label system to obtain the radiation map coordinate interval.
[0060] A263: Labels are assigned sequentially to the coordinate intervals of the radiation map to generate a multidimensional sample annotation map.
[0061] A264: The multidimensional sample annotation map is used to perform multidimensional evaluation annotation on the telecommunications user training data to obtain the telecommunications user sample data.
[0062] Specifically, firstly, based on the hierarchical structure of the tagging system, the first-level profile tags serve as the core dimension of the radial graph, extending outward from the center. Second-level tags serve as branch dimensions of the first-level tags, and third-level tags further refine the branches, forming a multi-dimensional radial distribution. For example, using five first-level tags such as basic information and communication behavior as the core axes, each axis extends outward to form second-level tag branches, and each second-level tag branch further extends to form third-level tags, constructing a radial graph structure containing multiple hierarchical dimensions.
[0063] Next, the multidimensional label radiation map is divided into coordinate intervals according to the multi-level image label value range set of the image label system, and the value range of each dimension is mapped to the coordinate scale of the radiation map. For example, the value range of call duration under the communication behavior dimension, such as 0-50 hours and 51-100 hours, corresponds to the coordinate intervals of 0-5 and 5-10 on the axis of that dimension in the radiation map. Each value range is accurately mapped to a specific coordinate range, forming a clear coordinate interval of the radiation map.
[0064] Then, labels are assigned sequentially to the coordinate intervals of the radiation map, and the multi-level labels corresponding to each interval are combined and clearly marked. For example, if a coordinate interval simultaneously covers labels such as 20-30 years old in the basic information dimension, monthly consumption of 51-100 yuan in the consumption behavior dimension, and video applications in the interest preference dimension, these labels are combined and assigned to that interval, so that each coordinate interval corresponds to a unique label combination, thereby generating a complete multi-dimensional sample annotation map.
[0065] Finally, when using multidimensional sample annotation maps to perform multidimensional evaluation annotation on the training data of telecommunications users, it is necessary to match the user features in the training data with the coordinate range of the annotation map. If the features of each dimension of the user data fall within a certain coordinate range, the corresponding label combination of that range is assigned, and finally, telecommunications user sample data with clear labels is obtained.
[0066] By constructing a multidimensional label radiation map, dividing coordinate intervals, assigning labels to generate labeled maps, and labeling training data, accurate telecommunications user sample data was obtained, providing a high-quality sample foundation for the subsequent training of the profile label classifier.
[0067] Furthermore, step A300 in the method provided in this application embodiment includes: A310: Determine the business content feature dimensions based on the aforementioned profile tagging system.
[0068] A320: Based on the service content feature dimension, feature extraction is performed on the telecommunications service content to obtain the service content dimension feature set.
[0069] A330: Perform iterative clustering analysis on the business content dimension feature set until the preset convergence condition is met to obtain multi-business content feature clusters.
[0070] A340: Tag the multi-service content feature clusters according to the portrait tagging system to generate the service content tag set.
[0071] Specifically, when determining the feature dimensions of business content based on the user profile tagging system, the first step is to use the multi-level tags of the user profile as a reference to ensure that the feature dimensions of the business content correspond to the user feature dimensions. For example, if the user profile tagging system includes monthly spending amount in the consumption behavior dimension and content preferences in the interest preference dimension, the corresponding feature dimensions of the business content would be set as business price range, business content type, applicable traffic range, etc., to ensure that the features of the business content accurately reflect the attributes that users may be interested in.
[0072] Then, when extracting features from telecommunications service content based on the defined service content feature dimensions, it is necessary to traverse all service content and extract key information. Assuming that telecommunications service content includes 200 different communication packages and 50 types of value-added services, for the service price range dimension, data such as the monthly fee for each package and the single-use charge for value-added services are extracted; for the service content type dimension, types such as video data packages, voice call packages, and social targeted data are distinguished; for the applicable data range dimension, the total data allowance and time period restrictions included in the package are recorded, ultimately forming a service content dimension feature set containing 200+50 service content items, each with 3-5 feature values.
[0073] Subsequently, when performing iterative clustering analysis on the business content dimension feature set, the K-means clustering algorithm can be used, with a preset number of clusters of 8-12. The convergence condition is that the change rate of cluster centers between two consecutive iterations is less than 0.5%. Taking the above 250 business content feature data as an example, 8 cluster centers are initially randomly selected. The Euclidean distance between each data point and each center is calculated and the data is classified. After updating the centers, the iteration is repeated. After 15 iterations, the convergence condition is met, and finally 8 business content feature clusters are obtained. The business content within each cluster has high similarity in features such as price, type, and traffic range, such as the low-priced video traffic package cluster and the mid-priced comprehensive package cluster.
[0074] Finally, the multi-service content feature clusters are tagged according to the profile tagging system, and the features of the multi-service content feature clusters are mapped to the value range of the profile tags. For example, the low-priced video traffic package cluster is priced between 10-30 yuan, mainly targeting users with monthly consumption of 51-100 yuan and video-related interests in the profile tags. Therefore, this cluster is assigned tags such as low price, video-oriented, and suitable for users with medium consumption, ultimately generating a set of business content tags containing all cluster tags.
[0075] By defining the business content dimensions that match user characteristics, extracting and clustering features, and associating them with profile tags for labeling, a set of business content tags that is highly compatible with the user profile system was generated, providing a standardized basis for the accurate matching of users and business content.
[0076] Furthermore, step A400 in the method provided in this application embodiment includes: A410: Perform principal component analysis and weight assignment on the information of each dimension in the user profile dimension to determine the profile dimension weight factor.
[0077] A420: Match the user multidimensional profile set with the business content tag set according to the profile dimension weight factor to obtain the user business content matching degree set.
[0078] A430: Based on the user business content matching degree set, associate and filter the user multidimensional profile set with the business content tag set to determine the user profile-business content association table.
[0079] In one embodiment, principal component analysis (PCA) and weight assignment are performed on the information of each dimension in the user profile. When calculating the variance contribution rate of each dimension in PCA, firstly, the raw data of the five dimensions of the user profile—basic information, communication behavior, consumption behavior, interests and preferences, and social relationships—are standardized to eliminate the influence of differences in units (such as the amount of consumption in yuan or the duration of calls in minutes) between different dimensions, so that the data of each dimension are on the same order of magnitude. Next, the correlation matrix between these five dimensions is calculated to reflect the degree of linear correlation between different dimensions.
[0080] Next, the eigenvalues and eigenvectors of the correlation matrix are calculated. The eigenvalues represent the total amount of original data information contained in the corresponding principal component, i.e., the variance, while the eigenvectors reflect the linear combination relationship between the original dimensions and the principal components. Then, all eigenvalues are arranged in descending order. The ratio of each eigenvalue to the sum of all eigenvalues is the variance contribution rate of the principal component corresponding to that eigenvalue, representing the proportion of the total variation in the original data that the principal component can explain. Finally, based on the original dimension combinations corresponding to each principal component and the magnitude of the variance contribution rate, it is determined which original dimensions have a higher proportion in explaining user characteristics, thereby selecting the core dimensions that have a more significant impact on user characteristics and reducing information redundancy.
[0081] For example, principal component analysis was performed on five user profile dimensions: basic information, communication behavior, consumption behavior, interests and preferences, and social relationships. The results showed that the variance contribution rate for consumption behavior was 30%, for interests and preferences 25%, for communication behavior 20%, for social relationships 15%, and for basic information 10%. Based on these variance contribution rates, corresponding weights were assigned to each dimension: consumption behavior 0.3, interests and preferences 0.25, communication behavior 0.2, social relationships 0.15, and basic information 0.1, forming profile dimension weighting factors that highlight the role of key dimensions in matching.
[0082] Next, the user multidimensional profile set and the business content tag set are matched and calculated sequentially to obtain the profile dimension content matching degree set. Then, the set is weighted according to the profile dimension weight factor to obtain the user business content matching degree set. The specific steps are explained in detail in A421-A422.
[0083] Finally, when associating and filtering user multidimensional profiles with business content tag sets based on the user business content matching score set, a matching score threshold, such as 0.6, needs to be set to retain association combinations with a matching score that reaches or exceeds this threshold. By filtering these high-matching-score combinations, the user profile-business content association table is finally determined, clearly presenting the appropriate business content corresponding to each user profile.
[0084] By determining dimensional weights through principal component analysis and filtering highly matched combinations, a user profile-business content association table that accurately reflects the user-business fit relationship is constructed, providing a direct and reliable basis for subsequent content distribution strategy formulation.
[0085] Furthermore, step A420 in the method provided in this application embodiment includes: A421: The user multidimensional profile set is matched and calculated sequentially with the business content tag set to obtain the profile dimension content matching degree set.
[0086] A422: The user business content matching degree set is obtained by weighting the content matching degree set of the profile dimension according to the profile dimension weight factor.
[0087] Optionally, when calculating the matching degree set of user multidimensional profiles with the business content tag set in sequence to obtain the profile dimension content matching degree set, it is necessary to perform an adaptation analysis on each dimension of the user profile and the corresponding dimension of the business content tag. Assume that the user multidimensional profile set contains profiles of 1000 users, each profile covering 5 dimensions: basic information, communication behavior, consumption behavior, interests and preferences, and social relationships; the business content tag set contains 50 business items, each business also corresponding to tags of these 5 dimensions.
[0088] During the matching calculation, for each user profile and each business item, the matching degree in 5 dimensions is calculated separately: For each dimension, the specific characteristics or range of user profile tags and business content tags are first identified; then the matching degree is determined by calculating the degree of overlap, adaptation ratio or feature matching degree of the two in that dimension: If it is a range-type tag, such as consumption amount or age, the proportion of the overlapping part of the range is calculated as the matching degree; if it is a feature-type tag, such as interest preference or social activity, a value between 0 and 1 is assigned according to the degree of feature fit. The more the features fit, the closer the matching degree is to 1; finally, each dimension obtains a matching degree value between 0 and 1.
[0089] For example, if the user's consumption behavior dimension tag is monthly spending of 101-200 yuan, and the corresponding business content dimension tag is target user monthly spending of 80-250 yuan, with an overlap of 80%, then the matching degree for this dimension is 0.8. If the user's interest preference dimension tag is frequent use of video applications, and the corresponding business content dimension tag is including video-targeted traffic, the matching degree is 0.9. In the basic information dimension, the matching degree between a user's age of 25 and the business's suitability for users aged 20-35 is 0.9; the matching degree between the communication behavior dimension of an average daily call time of 30 minutes and the business's inclusion of 50 minutes of free calls is 0.7; and the matching degree between the social relationship dimension of 10 frequently contacted individuals and the business's suitability for socially active users is 0.6. Through this dimension-level matching, each user and each business will form a profile dimension content matching degree set containing 5 values. 1000 users and 50 businesses will generate 50,000 such sets.
[0090] Next, the matching degree set of the profile dimensions is weighted according to the profile dimension weight factor. Combining the previously determined weights of each dimension, such as consumption behavior 0.3, interest preference 0.25, communication behavior 0.2, social relationship 0.15, and basic information 0.1, the matching degree of each dimension is multiplied by the corresponding weight and then summed. Taking a user's profile dimension content matching degree set of 0.9, 0.7, 0.8, 0.9, and 0.6 with a certain business as an example, the weighted calculation is: 0.9 × 0.1 (basic information) + 0.7 × 0.2 (communication behavior) + 0.8 × 0.3 (consumption behavior) + 0.9 × 0.25 (interest preference) + 0.6 × 0.15 (social relationship) = 0.09 + 0.14 + 0.24 + 0.225 + 0.09 = 0.785. This result is the user business content matching degree between the user and the business. Perform the same calculation on all 50,000 profile dimensions content matching score sets to finally obtain a user business content matching score set containing 50,000 matching score values.
[0091] By performing dimensional matching calculations and weighted summation, the multi-dimensional compatibility between users and businesses is transformed into a quantifiable overall matching degree, forming a set of user-business content matching degrees that can be directly used for association filtering, providing data support for accurately constructing user profile-business content association tables.
[0092] Furthermore, step A400 in the method provided in this application embodiment includes: A440: Evaluate the value of each user profile in the user profile-business content association table to obtain a tiered value user set.
[0093] A450: Preset business content distribution effect target, analyze the distribution strategy of the hierarchical value user set based on the business content distribution effect target, and determine the multi-level strategy for user content distribution.
[0094] In one embodiment, when evaluating the value of each user profile in the user profile-business content association table to obtain a tiered value user set, a multi-dimensional evaluation model needs to be constructed first. This model includes three core evaluation dimensions: consumption contribution dimension, business activity dimension, and potential conversion space dimension. The consumption contribution dimension uses the average monthly consumption amount over the past three months and the annual consumption growth rate as core indicators, with the average monthly consumption amount having a weight of 0.25 and the annual consumption growth rate having a weight of 0.15, comprehensively reflecting the user's current direct value. The business activity dimension uses the number of business transactions per month and the frequency of service inquiries as indicators, with weights of 0.15 and 0.1 respectively, reflecting the closeness of the interaction between the user and the business. The potential conversion space dimension uses the number of high-value services not yet processed and the upgrade rate of similar users as indicators, with weights of 0.2 and 0.15 respectively, predicting the user's future value enhancement potential.
[0095] After the model is built, each user profile in the association table is quantitatively scored, with a maximum score of 100. For example, a user's average monthly consumption of 200 yuan over the past 3 months scores 80, annual consumption growth rate of 10% scores 70, 2 business transactions per month scores 85, 1 consultation per month scores 75, no high-value business transactions scores 80, and the upgrade rate of similar users is 30% scores 70. Therefore, their comprehensive score is: 80×0.25+70×0.15+85×0.15+75×0.1+80×0.2+70×0.15=77.25 points. The user profiles in the association table are divided into three levels based on their scores: 80 points and above are high-value users, 50-79 points are medium-value users, and 49 points and below are low-value users, forming a tiered value user set.
[0096] When setting target performance for content distribution, it is essential to align with user segmentation. For high-value users, set targets of core business conversion rate ≥35%, average monthly business clicks ≥6, and user retention rate ≥95%. For mid-value users, set targets of core business conversion rate ≥25%, average monthly business clicks ≥4, and user retention rate ≥90%. For low-value users, set targets of basic business conversion rate ≥15%, average monthly business clicks ≥2, and user complaint rate ≤1.5%. Simultaneously, a uniform requirement of a ≥30% open rate for pushed content is required to ensure measurable distribution effectiveness.
[0097] When analyzing the distribution strategy for tiered user groups based on the aforementioned performance objectives, differentiated push methods and content are necessary. High-value users will receive personalized push notifications and exclusive services, with 1-2 high-end services tailored to their consumption habits and interests daily, such as premium packages or customized data packages. A dedicated customer service representative will follow up with a phone call within 48 hours of the push notification to increase conversion rates. Mid-value users will receive interest-targeted push notifications and incentive-based activities, with 3-4 cost-effective bundled services per week, such as limited-time discount packages or data / voice packages, supplemented with incentives like points redemption and referral rewards to enhance interaction. Low-value users will receive basic service push notifications and low-frequency outreach, with 1-2 entry-level services every 10 days, such as low-priced data packages or basic call packages, to avoid high-frequency interruptions that could cause annoyance. Through this tiered strategy design, a multi-level content distribution strategy covering all user groups of different values will be ultimately determined.
[0098] By constructing a multi-dimensional evaluation model to achieve precise user segmentation, setting differentiated effect goals, and analyzing adaptation strategies, a highly targeted multi-level strategy for user content distribution has been formed, which effectively improves the accuracy and conversion efficiency of business content distribution, while balancing the experience of users with different values.
[0099] In summary, the content distribution optimization method based on user profiles provided in this application has the following technical effects: This application collects and standardizes multi-source telecommunications user data to construct a user profile tagging system to determine a multi-dimensional user profile set. Based on this system, it performs cluster analysis and tagging of telecommunications service content to generate a service content tag set. It then matches and optimizes the user multi-dimensional profile set with the service content tag set to determine an association table. Based on this association table, it parses out a multi-level user content distribution strategy, thereby optimizing service content distribution and feedback. This achieves precise distribution of service content, making user profile-based content distribution more efficient and accurate. It achieves precise matching between user profiles and service content, making content distribution more comprehensive and accurate, and meeting the technical requirements for precise evaluation and optimization.
[0100] Example 2, as Figure 2As shown, based on the same inventive concept as in Embodiment 1 above, this application provides a content distribution optimization system based on user profiles, the system comprising: The standard user data acquisition module 1 is used to collect multi-source telecommunications user data, which includes telecommunications service data, user consumption data and third-party data. The multi-source telecommunications user data is standardized to obtain standard multi-source telecommunications user data.
[0101] User multidimensional profile acquisition module 2 is used to construct a profile tag system, evaluate and label the standard multi-source telecommunications user data based on the profile tag system, and determine the user multidimensional profile set.
[0102] The business content tag set acquisition module 3 is used to perform cluster analysis and content tagging on telecommunications business content according to the profile tag system, and generate a business content tag set.
[0103] The multi-level distribution strategy acquisition module 4 is used to match and optimize the user multi-dimensional profile set with the business content tag set, determine the user profile-business content association table, perform distribution strategy parsing based on the user profile-business content association table, determine the user content distribution multi-level strategy, and perform business content distribution and effect feedback optimization through the user content distribution multi-level strategy.
[0104] Furthermore, the standard user data acquisition module 1 is used to perform the following steps: According to data application standards, multi-source data anomaly conditions are set, including missing data, duplicate data, and out-of-range data. Anomalies are identified in the multi-source telecommunications user data according to these conditions to obtain multi-source abnormal user data. Based on these conditions, an anomaly cleaning strategy is matched, and the multi-source abnormal user data is cleaned using this strategy to obtain usable multi-source telecommunications user data. Finally, the usable multi-source telecommunications user data is standardized based on the data application standards to obtain standardized multi-source telecommunications user data.
[0105] Furthermore, the user multidimensional profile acquisition module 2 is used to perform the following steps: Based on telecommunications service objectives, user profile dimensions are defined, including basic information, communication behavior, consumption behavior, interest preferences, and social relationships. Each of these user profile dimensions is used as a first-level profile tag, and these first-level profile tags are then further subdivided into multiple levels to obtain a multi-level profile tag set. The multi-level profile tag set is then divided into value ranges according to business application requirements to determine the multi-level profile tag value range set. Finally, the first-level profile tags, the multi-level profile tag set, and the multi-level profile tag value range set are hierarchically mapped to construct the profile tag system.
[0106] Furthermore, the user multidimensional profile acquisition module 2 is used to perform the following steps: The standard multi-source telecommunications user data is proportionally extracted to obtain telecommunications user training data; the telecommunications user training data is multidimensionally evaluated and labeled according to the profile labeling system to obtain telecommunications user sample data; the telecommunications user sample data is trained and cross-validated using a support vector machine to generate a profile label classifier; the standard multi-source telecommunications user data is evaluated, classified, and labeled based on the profile label classifier to determine the user multidimensional profile set.
[0107] Furthermore, the user multidimensional profile acquisition module 2 is used to perform the following steps: Based on the multi-level image tag set of the image tag system, a multi-dimensional tag radiation map is constructed; the multi-dimensional tag radiation map is divided into coordinate intervals according to the value interval set of the multi-level image tags of the image tag system to obtain the radiation map coordinate interval; the coordinate intervals of the radiation map are sequentially labeled to generate a multi-dimensional sample annotation map; the multi-dimensional sample annotation map is used to perform multi-dimensional evaluation annotation on the telecommunications user training data to obtain the telecommunications user sample data.
[0108] Furthermore, the business content tag set acquisition module 3 is used to perform the following steps: Based on the profile tagging system, the service content feature dimensions are determined; features are extracted from the telecommunications service content based on the service content feature dimensions to obtain a service content dimension feature set; iterative clustering analysis is performed on the service content dimension feature set until a preset convergence condition is met to obtain a multi-service content feature cluster; the multi-service content feature cluster is tagged according to the profile tagging system to generate the service content tag set.
[0109] Furthermore, the multi-level distribution strategy acquisition module 4 is used to perform the following steps: Principal component analysis and weight assignment are performed on the information of each dimension in the user profile dimension to determine the profile dimension weight factor; the user multidimensional profile set and the business content tag set are matched and calculated according to the profile dimension weight factor to obtain the user business content matching degree set; based on the user business content matching degree set, the user multidimensional profile set and the business content tag set are associated and filtered to determine the user profile-business content association table.
[0110] Furthermore, the multi-level distribution strategy acquisition module 4 is used to perform the following steps: The user multidimensional profile set is matched and calculated sequentially with the business content tag set to obtain the profile dimension content matching degree set; the profile dimension content matching degree set is weighted and calculated according to the profile dimension weight factor to obtain the user business content matching degree set.
[0111] Furthermore, the multi-level distribution strategy acquisition module 4 is used to perform the following steps: Value assessment is performed on each user profile in the user profile-business content association table to obtain a tiered value user set; a business content distribution effect target is preset, and a distribution strategy analysis is performed on the tiered value user set based on the business content distribution effect target to determine a multi-level user content distribution strategy.
[0112] The content distribution optimization system based on user profiles provided in the embodiments of the present invention can execute the content distribution optimization method based on user profiles provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.
[0113] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.
[0114] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A content distribution optimization method based on user profiles, characterized in that, The method includes: Collect multi-source telecommunications user data, which includes telecommunications service data, user consumption data, and third-party data; and perform standardization processing on the multi-source telecommunications user data to obtain standard multi-source telecommunications user data. Construct a user profile labeling system, evaluate and label the standard multi-source telecommunications user data based on the user profile labeling system, and determine a multi-dimensional user profile set; Based on the aforementioned profile tagging system, cluster analysis and content tagging of telecommunications service content are performed to generate a service content tag set; The user multidimensional profile set is matched and optimized with the business content tag set to determine the user profile-business content association table. Based on the user profile-business content association table, the distribution strategy is parsed to determine the multi-level user content distribution strategy. The business content distribution and effect feedback are optimized through the multi-level user content distribution strategy.
2. The content distribution optimization method based on user profiles as described in claim 1, characterized in that, The obtained standard multi-source telecommunications user data includes: According to data application standards, multi-source data anomaly conditions are set, including missing data, duplicate data, and out-of-range data. Anomalies are identified in the multi-source telecommunications user data according to the aforementioned multi-source data anomaly conditions to obtain multi-source abnormal user data. Based on the aforementioned multi-source data anomaly conditions, an anomaly data cleaning strategy is matched, and the anomaly data cleaning strategy is used to clean the multi-source abnormal user data to obtain usable multi-source telecommunications user data. Based on the data application standard, the available multi-source telecommunications user data is standardized to obtain the standardized multi-source telecommunications user data.
3. The content distribution optimization method based on user profiles as described in claim 1, characterized in that, The construction of the profile tagging system includes: Based on telecommunications service objectives, user profile dimensions are defined, including basic information dimension, communication behavior dimension, consumption behavior dimension, interest and preference dimension, and social relationship dimension. Each user profile dimension is taken as a first-level profile tag, and the first-level profile tags are subdivided into multiple levels to obtain a multi-level profile tag set. The multi-level profile tag set is divided into value ranges according to business application requirements to determine the multi-level profile tag value range set. The first-level portrait tag, the multi-level portrait tag set, and the multi-level portrait tag value range set are hierarchically associated and mapped to construct the portrait tag system.
4. The content distribution optimization method based on user profiles as described in claim 1, characterized in that, The determination of the user multidimensional profile set includes: The standard multi-source telecommunications user data is proportionally extracted to obtain telecommunications user training data; The training data of telecommunications users is evaluated and labeled in multiple dimensions according to the aforementioned profile labeling system to obtain telecommunications user sample data. The support vector machine is used to train the label classification of the telecommunications user sample data and perform cross-validation optimization to generate a profile label classifier. The standard multi-source telecommunications user data is evaluated, classified, and labeled based on the profile label classifier to determine the user multidimensional profile set.
5. The content distribution optimization method based on user profiles as described in claim 4, characterized in that, The acquisition of telecommunications user sample data includes: Based on the multi-level image tag set of the aforementioned image tag system, a multi-dimensional tag radiation map is constructed. The multidimensional label radiation map is divided into coordinate intervals according to the multi-level image label value interval set of the image label system to obtain the radiation map coordinate interval; Labels are assigned sequentially to the coordinate intervals of the radiation map to generate a multidimensional sample annotation map; The multidimensional sample annotation map is used to perform multidimensional evaluation annotation on the telecommunications user training data to obtain the telecommunications user sample data.
6. The content distribution optimization method based on user profiles as described in claim 1, characterized in that, The generated business content tag set includes: Based on the aforementioned profile tagging system, determine the dimensions of business content features; Based on the service content feature dimensions, feature extraction is performed on the telecommunications service content to obtain a service content dimension feature set; Perform iterative clustering analysis on the business content dimension feature set until a preset convergence condition is met to obtain multi-business content feature clusters; The multi-service content feature clusters are tagged according to the aforementioned profile tagging system to generate the service content tag set.
7. The content distribution optimization method based on user profiles as described in claim 3, characterized in that, The process of determining the user profile-business content association table includes: Principal component analysis and weight assignment are performed on the information of each dimension in the user profile dimension to determine the profile dimension weight factor; The user multidimensional profile set and the business content tag set are matched and calculated according to the profile dimension weight factor to obtain the user business content matching degree set; Based on the user business content matching degree set, the user multidimensional profile set and the business content tag set are associated and filtered to determine the user profile-business content association table.
8. The content distribution optimization method based on user profiles as described in claim 7, characterized in that, The obtained user service content matching degree set includes: The user multidimensional profile set is matched and calculated sequentially with the business content tag set to obtain the profile dimension content matching degree set; The user business content matching degree set is obtained by weighting the content matching degree set of the profile dimension according to the profile dimension weight factor.
9. The content distribution optimization method based on user profiles as described in claim 1, characterized in that, The determination of the multi-level user content distribution strategy includes: The value of each user profile in the user profile-business content association table is evaluated to obtain a tiered value user set. A preset business content distribution effect target is set, and a distribution strategy is analyzed for the graded value user set based on the business content distribution effect target to determine a multi-level strategy for user content distribution.
10. A content distribution optimization system based on user profiles, characterized in that, The system is used to implement the content distribution optimization method based on user profiles as described in any one of claims 1-9, the system comprising: The standard user data acquisition module is used to collect multi-source telecommunications user data, which includes telecommunications service data, user consumption data and third-party data. The multi-source telecommunications user data is standardized to obtain standard multi-source telecommunications user data. The user multi-dimensional profile acquisition module is used to construct a profile tag system, evaluate and label the standard multi-source telecommunications user data based on the profile tag system, and determine the user multi-dimensional profile set. The business content tag set acquisition module is used to perform cluster analysis and content tagging on telecommunications business content according to the profile tag system, and generate a business content tag set. The multi-level distribution strategy acquisition module is used to match and optimize the user multi-dimensional profile set with the business content tag set, determine the user profile-business content association table, parse the distribution strategy based on the user profile-business content association table, determine the user content distribution multi-level strategy, and optimize the business content distribution and effect feedback through the user content distribution multi-level strategy.