User portrait construction method and device, equipment, medium and product

By preprocessing and optimizing multi-source heterogeneous data, and performing time-series-aware interest mining, user profile vectors are generated, solving the problems of one-sided and static user profiles and improving the accuracy and real-time performance of personalized recommendations.

CN121786729APending Publication Date: 2026-04-03CHINA MOBILE GROUP JIANGSU +1
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing user profiling methods rely on single-dimensional data, resulting in one-sided and static profiles that fail to fully reflect user characteristics. Furthermore, the lack of dynamic update mechanisms leads to insufficient accuracy and poor real-time performance in marketing recommendations.

Method used

By preprocessing multi-source heterogeneous data, static attribute vectors and dynamic behavior vectors are generated. Then, multi-objective optimized clustering analysis and time-aware interest mining mechanisms are used to fuse and generate user profile vectors.

Benefits of technology

By constructing comprehensive, granular, and dynamically evolving user profiles, the accuracy and real-time performance of personalized recommendation services are improved, thereby increasing user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786729A_ABST
    Figure CN121786729A_ABST
Patent Text Reader

Abstract

The invention discloses a user portrait construction method and device, equipment, a medium and a product. The method comprises the following steps: preprocessing multi-source heterogeneous data of a target user to obtain standardized data; the standardized data at least comprises static attribute data and dynamic behavior data; processing the static attribute data through a clustering analysis mechanism based on multi-objective optimization, and generating a basic attribute vector; processing the dynamic behavior data through an interest mining mechanism based on time sequence perception, and generating a behavior preference vector and an interest vector; and fusing the basic attribute vector, the behavior preference vector and the interest vector to generate a user portrait vector corresponding to the target user. According to the scheme, the static attribute data and the dynamic behavior data in the multi-source heterogeneous data are processed differentially, and the generated basic attribute vector, the behavior preference vector and the interest vector are fused into the unified user portrait vector, so that a comprehensive, fine-grained and dynamically evolved user portrait can be constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data mining technology, and in particular to a method, apparatus, device, medium and product for constructing user profiles. Background Technology

[0002] With the development of big data and precision marketing, telecom operators widely adopt user profiling technology to understand customer needs. Traditional user profiling methods have the following limitations: First, they often rely on single-dimensional data, such as coarse-grained grouping based solely on static demographic attributes or preference analysis based only on historical consumption records. This results in a one-sided and static profile that fails to fully reflect user characteristics. Second, even when integrating multiple types of data, there is a lack of effective fusion and dynamic updating mechanisms. For example, the processing of behavioral data fails to fully consider its timeliness, and the mining of user interests often relies on static topic models, making it difficult to capture the evolution of interests. This leads to a disconnect between the profile and actual needs, resulting in insufficient accuracy and poor real-time performance of profile-based marketing recommendations. Summary of the Invention

[0003] This invention provides a user profile construction method, apparatus, device, medium, and product to solve the problem that existing user profile construction methods have one-sided, static, and outdated profile features, resulting in insufficient accuracy of personalized marketing recommendations.

[0004] According to one aspect of the present invention, a user profile construction method is provided, the method comprising:

[0005] The multi-source heterogeneous data of the target users is preprocessed to obtain standardized data; the standardized data includes at least static attribute data and dynamic behavior data.

[0006] Static attribute data is processed using a clustering analysis mechanism based on multi-objective optimization to generate basic attribute vectors;

[0007] Dynamic behavioral data is processed through a time-aware interest mining mechanism to generate behavioral preference vectors and interest vectors.

[0008] The basic attribute vector, behavioral preference vector, and interest vector are fused to generate a user profile vector corresponding to the target user.

[0009] According to another aspect of the present invention, a user profile building apparatus is provided, the apparatus comprising:

[0010] The preprocessing module is used to preprocess the multi-source heterogeneous data of the target user to obtain standardized data; the standardized data includes at least static attribute data and dynamic behavior data.

[0011] The static processing module is used to process static attribute data through a clustering analysis mechanism based on multi-objective optimization to generate basic attribute vectors;

[0012] The dynamic processing module is used to process dynamic behavioral data through a time-aware interest mining mechanism to generate behavioral preference vectors and interest vectors.

[0013] The user profile generation module is used to fuse basic attribute vectors, behavioral preference vectors, and interest vectors to generate user profile vectors corresponding to the target user.

[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the user profile construction method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the user profile construction method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the user profile construction method described in any embodiment of the present invention.

[0020] The technical solution of this invention preprocesses multi-source heterogeneous data of target users to obtain standardized data. The standardized data includes at least static attribute data and dynamic behavior data. Static attribute data is processed using a multi-objective optimization-based clustering analysis mechanism to generate basic attribute vectors. Dynamic behavior data is processed using a time-series-aware interest mining mechanism to generate behavior preference vectors and interest vectors. The basic attribute vectors, behavior preference vectors, and interest vectors are then fused to generate a user profile vector corresponding to the target user. This solution differentiates the static attributes and dynamic behavior data in multi-source heterogeneous data and uses a multi-objective optimization-based clustering analysis mechanism and a time-series-aware interest mining mechanism respectively to generate three types of feature vectors: basic attributes, behavior preferences, and interests. These are ultimately fused into a unified user profile vector, effectively solving the shortcomings of traditional methods such as one-sided, static, and lagging profile updates. It can construct a comprehensive, fine-grained, and dynamically evolving user profile, thereby significantly improving the accuracy, real-time performance, and user satisfaction of subsequent personalized recommendation services.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a user profile construction method provided in Embodiment 1 of the present invention;

[0024] Figure 2 This is a flowchart of a user profile construction method provided in Embodiment 2 of the present invention;

[0025] Figure 3 This is a flowchart of a user profile construction method provided in Embodiment 3 of the present invention;

[0026] Figure 4 This is a schematic diagram of a user profile building device according to Embodiment 4 of the present invention;

[0027] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the user profile construction method of this invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Example 1

[0031] Figure 1 This is a flowchart of a user profile construction method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where fine-grained user profiles are constructed to support precision marketing. This method can be executed by a user profile construction device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown in the figure, the user profile construction method provided in this embodiment includes the following steps:

[0032] S110. Preprocess the multi-source heterogeneous data of the target user to obtain standardized data; the standardized data includes at least static attribute data and dynamic behavior data.

[0033] Multi-source heterogeneous data can refer to a collection of user data that comes from different data channels and has different data structures and formats. For example, it can include, but is not limited to, structured data (such as form attribute data), semi-structured data (such as log data), and unstructured data (such as text-based consultation records).

[0034] Preprocessing can refer to a series of preliminary processing operations performed on raw, multi-source, heterogeneous data. These operations may include data cleaning, feature transformation, and standardization, with the aim of improving data quality and unifying data format.

[0035] Standardized data refers to a set of user data that has been preprocessed and has a unified format, unified dimensions, regular structure, and high data purity. It is the input data for building user profiles.

[0036] Static attribute data refers to a type of data that represents the basic and stable attributes of a user. Its content does not easily change over time or with changes in user behavior. It is the core data for defining the basic identity and fixed attributes of a user. For example, it may include, but is not limited to, demographic information (age, gender) and basic business attributes (network duration, package level).

[0037] Dynamic behavioral data refers to a type of data that represents various behaviors of users at different points in time. Its content has obvious timeliness and volatility, and can reflect dynamic characteristics such as users' behavioral habits and interests. For example, it may include, but is not limited to, call records, internet behavior, consumption records, business change records, etc.

[0038] In this embodiment of the invention, multi-source heterogeneous raw data related to target users can be collected from various business systems of telecommunications operators. Then, through preprocessing operations such as data cleaning, anomaly handling, format alignment, and feature engineering, these raw data are transformed into high-quality, well-organized standardized data. Then, based on the frequency of data change and business meaning, the data is divided into static attribute data representing explicit user feedback and dynamic behavioral data representing implicit feedback. These two types of data will provide unified and reliable data support for the subsequent generation of various feature vectors for user profiles.

[0039] S120. Static attribute data is processed through a clustering analysis mechanism based on multi-objective optimization to generate basic attribute vectors.

[0040] Among them, the clustering analysis mechanism based on multi-objective optimization can be understood as an improved clustering method. Its multi-objective optimization is reflected in the calculation of user similarity. It not only minimizes the distance between users in the feature space, but also maximizes the discriminativeness or information content of the feature set used. That is, it ensures that the static features of users in the same cluster are highly overlapping, and also ensures that the feature differences between different clusters are significant, thereby achieving accurate division of user groups.

[0041] Basic attribute vectors can refer to feature vectors generated based on static attribute data that combine individual-specific attributes with group-common attributes. They can comprehensively and accurately represent the basic characteristics of users and provide underlying support for the overall construction of user profiles.

[0042] In this embodiment of the invention, an improved clustering algorithm can be used to process static attribute data. The core of this algorithm is that it does not use a traditional single distance metric for clustering, but instead uses a multi-objective optimization function to intelligently evaluate and assign differentiated weights to different features while calculating similarity, especially emphasizing features that can effectively distinguish user groups. Then, based on the optimized similarity, clustering analysis is performed to identify the target user cluster to which the target user belongs. Finally, the target user cluster is encoded to generate a cluster identifier vector, which is then fused and concatenated with the normalized static features to finally output a high-dimensional vector representing the basic attributes of the target user, namely the basic attribute vector.

[0043] S130. Process dynamic behavior data through a time-aware interest mining mechanism to generate behavior preference vectors and interest vectors.

[0044] Among them, the interest mining mechanism based on time sequence awareness can be understood as a behavior and interest analysis mechanism that integrates time dimension features. Its purpose is to not only focus on the content of the behavior / interest itself when mining user behavior preferences and interests, but also combine the time sequence attributes such as the time of occurrence, duration, and trend of change to achieve accurate capture of user dynamic behavior.

[0045] A behavioral preference vector can be a quantitative representation vector of a user's structured behavioral data (such as consumption frequency, package usage, service activation records, etc.). Each dimension of the vector corresponds to a specific behavioral feature or behavioral pattern, and its numerical value (weight or score) directly reflects the significance of the behavioral pattern in the user's current behavior and recent activity, and is used to quantitatively describe the user's behavioral habits and operational tendencies.

[0046] An interest vector can be a quantitative representation vector of a user's unstructured behavioral data (such as user comments, browsing content, interaction feedback, etc.). The dimension of this vector corresponds to the mined interest topic, and its numerical value (weight or rating) reflects the strength of the user's preference for a certain interest topic and its temporal attributes (such as long-term stable attention or short-term concentrated activity). It is used to quantitatively describe the user's interest focus and its dynamic evolution.

[0047] In this embodiment of the invention, structured behavioral data (such as quantifiable numerical sequences like consumption frequency and package usage) and unstructured behavioral data (such as semantic text information like search terms and browsing content titles) can be separated from dynamic behavioral data first. Then, these two types of behavioral data are processed differently: For structured behavioral data, a time-aware multi-dimensional weight adjustment mechanism can be used to quantify each behavioral feature. For example, recent behaviors are given higher weights, while the contribution of long-term behaviors decays over time, ultimately generating behavioral preference vectors that reflect the user's recent behavioral patterns and intensity preferences. For unstructured behavioral data, a multi-level topic model (such as the LDA model) can be used for analysis. This model not only identifies the distribution of user-focused interest topics from the text, but also divides interests into long-term and short-term interests by analyzing the frequency and persistence of interest topics on the timeline, ultimately generating an interest vector that reflects the user's interest preferences and their dynamic changes. Understandably, the two vectors mentioned above (behavioral preference vector + interest vector) together quantify the target user's dynamic and fine-grained preferences and tendencies in the two core dimensions of behavior and interest, constructing a complete, structured, and computable mathematical representation of the user's dynamic aspects, and providing a core data foundation for subsequent accurate personalized services.

[0048] S140. Merge the basic attribute vector, behavioral preference vector, and interest vector to generate a user profile vector corresponding to the target user.

[0049] Among them, the user profile vector can refer to a comprehensive quantitative carrier that integrates three types of features: basic user attributes, behavioral preferences, and potential interests. It can comprehensively and uniformly represent the overall characteristics of the target user and provide core data support for subsequent personalized service recommendations and other businesses.

[0050] In this embodiment of the invention, three feature vectors representing a user's static attributes, dynamic behavioral preferences, and semantic interests—namely, the basic attribute vector, the behavioral preference vector, and the interest vector—can be standardized to eliminate differences in their numerical scales. Subsequently, the three adjusted vectors are concatenated dimensionally in a predetermined order to generate a unified, high-dimensional user profile vector. Furthermore, this user profile vector can provide high-quality input for downstream tasks such as personalized recommendations and precision marketing, thereby enabling targeted personalized recommendations and precision marketing.

[0051] The technical solution of this invention preprocesses multi-source heterogeneous data of target users to obtain standardized data. The standardized data includes at least static attribute data and dynamic behavior data. Static attribute data is processed using a multi-objective optimization-based clustering analysis mechanism to generate basic attribute vectors. Dynamic behavior data is processed using a time-series-aware interest mining mechanism to generate behavior preference vectors and interest vectors. The basic attribute vectors, behavior preference vectors, and interest vectors are then fused to generate a user profile vector corresponding to the target user. This solution differentiates the static attributes and dynamic behavior data in multi-source heterogeneous data and uses a multi-objective optimization-based clustering analysis mechanism and a time-series-aware interest mining mechanism respectively to generate three types of feature vectors: basic attributes, behavior preferences, and interests. These are ultimately fused into a unified user profile vector, effectively solving the shortcomings of traditional methods such as one-sided, static, and lagging profile updates. It can construct a comprehensive, fine-grained, and dynamically evolving user profile, thereby significantly improving the accuracy, real-time performance, and user satisfaction of subsequent personalized recommendation services.

[0052] Example 2

[0053] Figure 2 This is a flowchart of a user profile construction method provided in Embodiment 2 of the present invention. It is further optimized and extended based on the above embodiments and can be combined with various optional technical solutions in the above embodiments. For example... Figure 2 As shown in the figure, the user profile construction method provided in this embodiment includes the following steps:

[0054] S210. Extract the corresponding multi-source heterogeneous data from the preset telecommunications system database according to the target user's identification information; wherein, the multi-source heterogeneous data includes user personal basic attribute data, circuit-switched domain signaling data, and packet-switched domain IP packet data.

[0055] The target user's identification information can refer to the feature information used to uniquely identify the target user, such as including but not limited to the user's mobile phone number, ID card number, user number, device IMEI code, etc., which is the core retrieval basis for various types of telecommunications data associated with the user.

[0056] A pre-built telecommunications system database can refer to an integrated data storage system that telecommunications operators build in advance to centrally store various types of user communication and basic information. It typically includes multiple functional partitions such as user basic information partitions, signaling data partitions, and IP packet data partitions, and has core capabilities such as data retrieval, storage, and verification.

[0057] User personal basic attribute data can refer to data used to describe the user's basic identity and business contract characteristics, such as static attribute information such as user name, age, gender, network access time, contracted package type, and business activation status.

[0058] Circuit-switched (CS) domain signaling data refers to the signaling information generated when users conduct communication activities such as voice calls and SMS sending and receiving in traditional circuit-switched networks. For example, it may include data such as call initiation time, call duration, called / called number, and communication network node, which can reflect the user's communication behavior characteristics.

[0059] Packet Switch (PS) domain IP packet data refers to data related to IP data packets generated by users in packet-switched networks such as mobile Internet. For example, it may include data such as data transmission time, traffic consumption, accessed network address, and data transmission protocol type, which can characterize the user's mobile Internet usage behavior.

[0060] In this embodiment of the invention, the identification information of the target user can be used as an index to extract the user's basic personal attribute data, circuit-switched domain signaling data and packet-switched domain IP packet data from different functional partitions of the preset telecommunications system database, and the above data can be used as the multi-source heterogeneous data corresponding to the target user.

[0061] S220. Perform data cleaning, feature transformation and standardization on multi-source heterogeneous data to obtain a regularized dataset.

[0062] Data cleaning refers to the process of identifying, correcting, or removing missing values, outliers, duplicate values, and dirty data from the original data. Its core objective is to improve data quality and ensure the integrity, accuracy, and consistency of the data.

[0063] Feature transformation refers to the process of converting raw data of different types and formats into numerical or vector features in a unified format that can participate in subsequent calculations through encoding, parsing, extraction and other means. It is a key step in solving the fusion analysis of heterogeneous data.

[0064] Standardization refers to preprocessing methods that eliminate differences in the dimensions and distributions of data features. Through operations such as normalization and standardization, features of different magnitudes and distributions are brought to the same analytical scale, making them meet the requirements of subsequent quantitative analysis or model algorithms.

[0065] In this embodiment of the invention, after acquiring multi-source heterogeneous data, data cleaning methods, including but not limited to missing value completion, outlier correction, duplicate value removal, and dirty data processing, can be used to eliminate noise and operations in the original data, thereby providing a high-quality and reliable data foundation for subsequent analysis. Then, feature transformations such as classification feature encoding, unstructured feature extraction, and format normalization are performed on the cleaned data to map complex information in the real world (such as categories, text, etc.) into numerical forms that subsequent model algorithms can directly understand and operate on, eliminating the heterogeneity barrier of the data. Finally, the data after feature transformation is normalized or Z-score standardized to eliminate the dimensional differences of different features and form a regularized dataset.

[0066] S230. Separate static attribute data and dynamic behavior data from the regularized dataset and use them as standardized data.

[0067] In this embodiment of the invention, static attribute data and dynamic behavior data can be separated from the regularized dataset based on the frequency of data change and business meaning, and then integrated into hierarchical standardized data, which provides a high-quality, multi-dimensional data foundation for the subsequent hierarchical generation of user profile vectors.

[0068] S240. Based on static attribute data, construct a global static feature set that includes the target user and other users.

[0069] Among them, other users refer to the sample user set selected to provide group reference and context for the target users in this user profile construction and analysis. Their static attribute data will be used to construct feature reference benchmarks, provide group feature comparison dimensions for the clustering and affiliation determination of the target users, and form the basis for realizing multi-user feature association analysis.

[0070] The global static feature set can refer to a collection of static attribute data of the target user and all other sample users. It is the core data input for subsequent multi-objective optimization clustering analysis.

[0071] In this embodiment of the invention, static attribute data of other sample users (such as all active users in the same region and at the same time) within a preset range can be retrieved. This data has been preprocessed (e.g., cleaned and standardized) to form a regular structured record. Then, the static attribute data of the target user and other users are aggregated to form a standardized global static feature set. This set has the characteristics of unified feature dimensions, standardized data format, and wide coverage of user groups, and is the core data input for subsequent multi-objective optimization clustering analysis. In one embodiment, each row in the global static feature set can represent a user, and each column can represent a static attribute feature.

[0072] S250. Perform cluster analysis on the global static feature set to determine the target user cluster to which the target user belongs; wherein, in the cluster analysis, a preset multi-objective optimization function is used to determine the similarity between users, and the preset multi-objective optimization function is constructed based on feature difference degree and information entropy.

[0073] In this context, a user cluster can refer to a group of users with high overall similarity after cluster analysis. Users within the same user cluster share significant commonalities in static attributes, while users in different user clusters exhibit significant differences in static attributes.

[0074] A predefined multi-objective optimization function refers to a predefined function used to simultaneously optimize multiple objectives in user clustering analysis. The core feature of this function is the synergistic optimization of the sub-objectives of "minimizing feature differences" and "regulating the information entropy of feature distribution." Feature difference can be a quantitative indicator measuring the degree of difference between two users on one or more static features. A higher value indicates a more significant difference in static features between users. Common calculation methods include, but are not limited to, Euclidean distance, cosine similarity, and Manhattan distance. Information entropy refers to the distribution entropy of user static features in the global static feature set. It is an indicator measuring the degree of disorder in feature distribution. A higher entropy value indicates a more uniform global distribution of the feature and a weaker uniqueness; a lower entropy value indicates a more concentrated global distribution of the feature and a stronger representativeness. Understandably, features with wider distribution and greater differences will be assigned higher weights, while more uniform features will receive lower weights. This weight adjustment method allows for more accurate similarity calculations, ensuring the completeness and accuracy of user profiles in detail.

[0075] In this embodiment of the invention, after obtaining a global static feature set that integrates the static attribute data of the target user and a large number of sample users, the comprehensive similarity between users can be calculated using a preset multi-objective optimization function that integrates feature difference and information entropy. Subsequently, based on this comprehensive similarity, a clustering analysis algorithm (such as K-means, hierarchical clustering, etc.) is executed to divide the user set into multiple user clusters with high internal similarity and large differences between each other. Finally, the target user cluster to which the target user belongs is determined according to the similarity between the static feature vector of the target user and the feature vector of the center of each user cluster.

[0076] It's important to understand that users' static attribute data typically doesn't change significantly over time. Therefore, cluster analysis can be performed in advance based on a large amount of users' static attribute data, and the clustering results can be saved. When it's necessary to build a profile for a target user, the similarity between the target user's static feature vector and the pre-stored feature vectors of each user cluster center can be used to quickly locate the target user cluster to which the target user belongs, thus accelerating the profile building process.

[0077] S260. Encode the target user cluster to obtain the cluster label vector.

[0078] Encoding can refer to the operation of converting abstract information in non-vector form, such as category identifiers and core features of user clusters, into standardized vector data that can be recognized and calculated by computers. This is used to establish a mapping relationship between the abstract features of user clusters and numerical vectors. For example, encoding methods can include, but are not limited to, one-hot encoding and embedding.

[0079] A cluster label vector can be a numerical vector obtained after performing encoding operations on a target user cluster. This vector can represent the group affiliation category of the target user and can also integrate the core attribute features of the user cluster. It is a digital label for the identity of the user group.

[0080] In this embodiment of the invention, encoding methods such as One-hot encoding can be used to convert the abstract target user cluster into a structured cluster label vector, thereby realizing the digitization and quantification of user group characteristics.

[0081] S270. The cluster label vector is fused with the normalized static attribute data to obtain the basic attribute vector.

[0082] In this embodiment of the invention, the cluster label vector and the normalized static attribute data can be integrated into a new, high-dimensional basic attribute vector by splicing or weighted fusion. This vector organically combines the common characteristics of the group with the specific details of the individual, together forming a complete digital expression of the user's basic attributes.

[0083] S280. Extract structured and unstructured behavioral data from dynamic behavioral data.

[0084] Structured behavioral data can refer to dynamic behavioral data with a fixed data format, clear field attributes, and direct quantification and statistics. Its data form is usually a database table, structured log, etc. For example, structured behavioral data in the telecommunications scenario may include, but is not limited to, users' monthly data usage quota, call duration statistics, recharge consumption amount and time, package change records, etc.

[0085] Unstructured behavioral data refers to dynamic behavioral data that lacks a unified fixed format, is mainly composed of natural text or unformatted data streams, and cannot be directly quantitatively analyzed. Its data form is usually text, voice (which needs to be converted to text), images (which need to be parsed), etc. For example, unstructured behavioral data in telecommunications scenarios may include, but is not limited to, dialogue text between users and customer service, service satisfaction evaluation messages, and freely entered text for business handling inquiries.

[0086] In this embodiment of the invention, in addition to conventional structured data (such as consumption frequency, package usage, etc.), unstructured data (such as user reviews, browsing content, etc.) can also provide rich information on user interests and preferences, providing important input data for building fine-grained dynamic user profiles. Based on this, before conducting dynamic user behavior analysis, structured and unstructured behavioral data can be extracted from the dynamic behavior data, thus providing differentiated data input for subsequently generating behavioral preference vectors and interest vectors respectively.

[0087] S290. Using a preset time weighting adjustment formula and / or a preset behavior weighting adjustment formula, weights are assigned to each behavior feature in the structured behavior data to generate a behavior preference vector.

[0088] The preset time weight adjustment formula can be a pre-defined mathematical function used to adjust the weight of behavioral features based on the proximity of their occurrence time. The core logic is that the closer the time, the higher the weight. Common forms can include exponential decay formulas, linear decay formulas, etc. Its function is to strengthen the influence of recent behavior on the user's current preferences and weaken the interference of long-term ineffective behavior.

[0089] The preset behavior weight adjustment formula can be a pre-defined mathematical function used to adjust the weight of behavior features based on their business importance and relevance. The core logic is that the higher the business relevance and the greater the impact on user needs, the higher the weight. Its function is to evaluate the changes in the activity of behavior features in different time periods in order to dynamically adjust the weight of each behavior feature.

[0090] Behavioral characteristics can refer to indicators or variables that can quantify and describe a user's specific behavioral patterns or habits. They are the basic data units that constitute a user behavior profile. For example, behavioral characteristics may include, but are not limited to: usage characteristics (such as data usage), transaction characteristics (such as the number of times a user changes their data plan), frequency characteristics (such as the number of times a user makes a call per week), etc.

[0091] In this embodiment of the invention, one or a combination of two of the preset time weight adjustment formula and preset behavior weight adjustment formula can be flexibly selected according to business needs. That is, the timeliness of recent behavior is highlighted by the time dimension, and the importance of behavior is distinguished by the business dimension. Dynamic weight calculation is performed on each behavioral feature in the structured behavioral data. Finally, the weighted behavioral features are integrated by dimension to generate a behavioral preference vector that can accurately map user behavior patterns.

[0092] S2100. Preprocess the unstructured behavioral data to obtain a keyword set.

[0093] Preprocessing refers to a series of normalization operations on unstructured behavioral data, which may include data filtering, cleaning, word segmentation, keyword extraction, etc. The purpose is to remove invalid information, extract core semantics, and transform messy unstructured behavioral data into a standardized set of keywords that can be used as input to the model.

[0094] A keyword set can refer to a collection of words or phrases extracted from unstructured behavioral data that can represent the core semantics of the original text. The relevance and accuracy of these words directly determine the effectiveness of subsequent interest mining.

[0095] In this embodiment of the invention, unstructured behavioral data can be preprocessed, including but not limited to word segmentation and stop word removal, part-of-speech tagging and named entity recognition, and noise reduction, in order to extract a set of keywords that represent user concerns.

[0096] S2110. Input the keyword set into the preset multi-level topic model to mine interest topics and obtain interest vectors; wherein, the preset multi-level topic model embeds an interest persistence detection mechanism.

[0097] The pre-defined multi-level topic model can refer to a hierarchical interest mining model based on semantic clustering. Unlike single-level topic models, it can achieve multi-level topic segmentation from bottom-level subdivided interest words to top-level core interest domains, simultaneously capturing both shallow and deep user interests and adapting to complex user interest representation needs. In one embodiment, the pre-defined multi-level topic model may include at least a Latent Dirichlet Allocation (LDA) model.

[0098] Interest topic mining refers to the process of transforming a set of keywords into hierarchical user interest topics by pre-setting a multi-level topic model. The core is to realize the transformation of interests from scattered words to systematic topics, so that user interests have quantifiable and categorizable characteristics.

[0099] An interest persistence detection mechanism can refer to a rule embedded in a pre-defined multi-level topic model to assess the stability of user interests. By analyzing the time distribution patterns and frequency changes of specific interest topics in the user's behavior sequence, it dynamically determines and quantifies whether the interest is a stable preference that the user has always held or a temporary attention that is only active within a specific time period. This mechanism is the key to realizing dynamic interest segmentation.

[0100] In this embodiment of the invention, the obtained keyword set can be input into a preset multi-level topic model embedded with an interest persistence detection mechanism for processing. This model not only mines the semantic topics behind the keyword text, but also automatically classifies them into long-term stable interests or short-term transient interests based on the time pattern of topic appearance through the embedded interest persistence detection algorithm. Finally, the system comprehensively encodes these analysis results into a multi-dimensional interest vector, which contains each interest topic and its corresponding intensity and persistence attributes, thereby completing a refined and dynamic digital model of user interest profiling.

[0101] S2120. Standardize the basic attribute vector, behavior preference vector, and interest vector.

[0102] Standardization processing can refer to a unified feature space calibration operation performed on vectors from different sources before vector fusion. The goal is to eliminate fusion bias caused by differences in feature dimensions, units, and numerical ranges among different vectors, so that each vector has a comparable and integrable feature base.

[0103] In this embodiment of the invention, the three key feature vectors generated for the target user above—basic attribute vector, behavioral preference vector, and interest vector—can be uniformly standardized to eliminate the imbalance caused by differences in dimensions, numerical ranges, or distributions between different feature vectors, ensuring their comparability during final fusion and preventing any single vector from dominating the final profile due to its excessively large numerical scale.

[0104] S2130. The standardized basic attribute vector, behavioral preference vector, and interest vector are concatenated to obtain the user profile vector.

[0105] In this embodiment of the invention, the three standardized feature vectors can be concatenated in a preset dimensional order to form a single, higher-dimensional user profile vector. This vector integrates multiple dimensions of information such as the target user's static attributes, dynamic behavioral preferences, and long-term and short-term interests, and is a fine-grained and complete feature expression of the target user.

[0106] Furthermore, based on the above embodiments of the invention, after generating the user profile vector, the method further includes:

[0107] Based on user profile vectors, at least one preset recommendation algorithm is used to generate personalized recommendation results for the target user; wherein, the preset recommendation algorithm includes at least a content recommendation algorithm based on vector matching and a collaborative filtering algorithm based on user similarity.

[0108] The pre-defined recommendation algorithm can refer to a recommendation strategy model or rule pre-embedded in the recommendation system. It can be selected or combined according to the actual business scenario requirements. Its core function is to establish the association between user profiles and candidate recommendation objects, outputting recommendation results that meet user needs. In one embodiment, the pre-defined recommendation algorithm may include at least a vector-matching-based content recommendation algorithm and a user similarity-based collaborative filtering algorithm. The vector-matching-based content recommendation algorithm is a type of recommendation algorithm that uses vector similarity calculation as its core. It measures the user's potential interest in candidate objects by comparing the similarity between the user profile vector and the feature vector of candidate objects, and makes recommendations accordingly. The user similarity-based collaborative filtering algorithm is a type of recommendation algorithm that relies on the correlation of group user behavior. It identifies user groups with similar characteristics to the target user and then recommends candidate objects favored by that group to the target user. Personalized recommendation results can refer to an ordered list of products, services, or content automatically generated by the recommendation system for the target user, conforming to their individual preferences and needs.

[0109] In this embodiment of the invention, after generating the user profile vector, at least one preset recommendation algorithm can be selected to execute the recommendation process according to actual business needs, and candidate recommendation results corresponding to each preset recommendation algorithm can be determined. Then, dynamic weights are assigned to the candidate recommendation results from different algorithm sources, and the comprehensive interest score of each candidate recommendation result relative to the target user is calculated. Finally, all candidate recommendation results are sorted according to the comprehensive interest score, and the top-ranked items are selected to generate the final personalized recommendation result, which is then pushed to the target user to achieve precision marketing.

[0110] Example 3

[0111] Figure 3 This is a flowchart of a user profile construction method provided in Embodiment 3 of the present invention. Based on the above embodiments, this embodiment provides an implementation method for user profile construction, which can achieve the fusion of multi-dimensional user feature vectors to generate a unified user profile and apply it to personalized recommendations. Figure 3 As shown, the user profile construction method provided in Embodiment 3 of the present invention specifically includes the following steps:

[0112] S310. Obtain multi-source heterogeneous data related to the target user from the telecommunications operator's business system.

[0113] In this embodiment of the invention, user profiling involves characterizing and analyzing users across various dimensions, which relies heavily on data support. The richer the data dimensions, the better the fusion effect, the more information can be mined, and the more comprehensive the user profile. This solution, based on actual business scenarios, obtains multi-source heterogeneous data related to target users from the telecom operator's business system. The raw data provided by the telecom operator's business system mainly includes three parts:

[0114] (1) User's basic personal attribute data, such as user's basic personal information, document information, address information, contact information, consumption information, package information and terminal information, etc.

[0115] (2) CS domain signaling data, such as user telephone call records, SMS sending records and other interaction records between the terminal and the network;

[0116] (3) PS domain IP packet data, such as the control plane and user plane data packet records when users use the network: control plane data such as authentication data and authentication data packets of AAA (Authentication Authorization Accounting), PDP (Packet Data Protocol) establishment, deletion, and update, etc., and user plane data mainly consists of user network usage data.

[0117] Next, the following valid data can be extracted from the above three types of multi-source heterogeneous data:

[0118] (1) Basic user information: mainly includes demographic dimensions that affect customers’ lifestyles and thus their purchase and use of telecommunications products, such as age, gender, occupation, etc.

[0119] (2) User status: This mainly refers to the customer status recorded in the China Unicom system;

[0120] (3) Telecommunications-related information: This mainly refers to other telecommunications-related information of customers in China Unicom.

[0121] S320. Preprocess the multi-source heterogeneous data to obtain standardized data.

[0122] Because telecommunications companies collect massive amounts of user information data, they inevitably encounter many problems during data collection, such as missing values ​​in variables and dirty data. These issues affect the accuracy and coverage of the user profile tags created. Therefore, data preprocessing is necessary before building user profiles to prepare for the subsequent construction process. Specifically, the preprocessing process includes:

[0123] (1) Handling missing values

[0124] The dataset contains some missing value variables, such as total user spending amount and user online time. Since the number of missing values ​​accounts for a very small proportion of the total data and the labels they belong to are all negative samples, the missing value handling method adopted in this embodiment is to directly delete the data rows with missing values.

[0125] (2) Dirty Data Processing

[0126] Dirty data refers to data in the source system that is outside the given scope or has no meaning for actual business operations. In the business scenario of telecommunications marketing, this embodiment defines dirty data as data for which no records are made for any indicator for a user for 30 consecutive days. This is because these users have not performed any meaningful operations, indicating that the user uses smart devices such as mobile phones too little, and their usage cannot be analyzed through mobile phone usage. Users with indicator records in other 30 days are retained.

[0127] (3) Numerical processing of categorical variables

[0128] The user dataset collected in this embodiment includes several feature variables, such as users' basic personal information, consumption information, and service activation information. Categorical variables include user gender, user age (whether the user is elderly), whether the user has joined a family network, whether they have activated 4G / 5G network service, whether they have activated telephone service, whether they have activated multi-line service, whether they have activated broadband, whether they have activated IPTV, contract signing method, whether they have activated e-billing service, and payment method. Since the original dataset already quantified the user gender and user age variables, this embodiment primarily focuses on quantifying the remaining categorical variables. Preliminary data analysis reveals that all categorical variables are nominal variables, and there is no order relationship between their values. Therefore, this invention employs One-hot encoding to quantify the categorical variables.

[0129] (4) Duplicate data processing

[0130] Duplicate values ​​may occur during data collection. This could be due to repeated startups of the platform program or problems during the data entry phase. This embodiment uses a merging method, which merges identical records for the same user into a single record by determining whether the user IDs are the same.

[0131] S330. Generate basic attribute vectors, behavioral preference vectors, and interest vectors based on standardized data.

[0132] In this embodiment of the invention, to achieve accurate personalized recommendations, the marketing recommendation system needs to mine user interests and build user profiles by modeling user attributes. User profiles need to reflect user preference information to support the operation of subsequent recommendation algorithms. The data used to build user profiles include static attribute data (such as demographic information) and dynamic behavioral data (such as structured and unstructured behavioral data). Feature extraction is performed on these data to construct labels for the user profile. In the field of marketing system recommendations, the user demographic information used to build user profiles mainly includes gender, age, purchased service packages, occupation, and location. This data represents explicit user feedback and can be directly obtained from user profiles. This method of obtaining user information is direct, easy to process, low in complexity, and highly transparent; however, it cannot dynamically update user features and lacks flexibility. Therefore, building user profiles cannot rely entirely on explicit user input.

[0133] Collecting dynamic user behavior data is an indispensable part of building user profiles. This data reflects users' consumption habits and serves as an information source for user modeling and profiling. Dynamic behavior data is essentially implicit user feedback. The backend server obtains user access information through system backend logs without interfering with users' normal access and use of the recommendation platform, analyzes user behavioral characteristics, and uncovers user interests.

[0134] The specific generation process of the three eigenvectors is described below:

[0135] (1) Generate basic attribute vectors based on static attribute data

[0136] When the amount of user information in a system is large, it is necessary to cluster the user information to discover the implicit information among a large number of users. In this embodiment of the invention, clustering algorithms can be used to classify static attribute data, thereby finding the inherent patterns in the data. Faced with a large amount of data, clustering algorithms can divide the data into different classes or clusters. The data between each class is different, but the data within each class is not significantly different, showing similarity. Before the recommendation algorithm finds similar users to the target user, clustering algorithms can be used to group similar users in the system into the same cluster. When finding the set of similar users to the target user, only the cluster where the target user is located is considered. That is, the similarity between other users and the target user is calculated only in the cluster where the target user is located. This avoids calculating the similarity of all users in the system, narrowing the scope of finding similar users, thereby improving the speed and efficiency of recommendation.

[0137] Since users' dynamic behavior data changes over time, the data may change significantly after clustering, rendering the previous clustering results inapplicable and requiring frequent calculations. In contrast, users' static attribute data changes less over time and requires little adjustment after clustering is established, thus achieving the goal of reusing calculation results.

[0138] The clustering process is as follows:

[0139] S1. Randomly select a sample as the first cluster center point C1;

[0140] S2. Calculate the shortest distance between each sample and the existing cluster centers; the larger the distance value, the greater the probability of being selected as a cluster center; finally, use the roulette wheel method to select the next cluster center C. N ;

[0141] S3. Repeat step S2 until the Kth cluster center C is selected. K ;

[0142] S4. Assign each sample point to the cluster represented by the nearest cluster center.

[0143] S5. Use the center point of all sample points in each cluster to represent the center point of the cluster;

[0144] S6. Repeat steps S4 and S5 until the cluster center point remains unchanged, or the set number of iterations is reached, or the set fault tolerance range is reached.

[0145] Furthermore, in the construction of user profiles, to make the calculation of similarity between users more accurate, this embodiment of the invention introduces optimizations in two aspects: "information entropy" and "feature difference degree". This ensures that when calculating user similarity, not only the weights of each feature are considered, but also the distribution of user data is better taken into account, making the similarity obtained by the model more representative.

[0146] The calculation of user similarity is viewed as a multi-objective optimization problem. Instead of simply using a single similarity metric, it considers both similarity and information entropy to ensure the diversity and importance of features. The optimized objective function L can be expressed as:

[0147]

[0148] In the formula, and Let be the feature values ​​of user a and user b on the i-th feature; Feature similarity can be calculated using methods such as Euclidean distance and cosine similarity. The dynamic weight of the i-th feature is automatically adjusted based on the distribution of the feature among all users (such as information entropy). The more uneven the distribution and the higher the discriminative power of the feature, the greater the weight. The joint information entropy of all features is used to measure the overall uniformity (uncertainty) of the distribution of all feature values. The larger the entropy value, the more uniform the feature distribution and the richer the information. This term is introduced to optimize feature weights and avoid similarity calculations from relying too much on a few concentrated features. and is an adjustment coefficient used to control the importance of information entropy and feature difference in the overall objective, and can be adjusted according to the actual application scenario; n is the total number of features.

[0149] Using the multi-objective optimization-based clustering analysis mechanism described above, the weights of each feature are automatically adjusted based on information entropy. The advantage of this is that features with wider distribution and greater diversity are assigned higher weights, while more uniform features receive lower weights. This adjustment allows for more accurate similarity calculations, ensuring the completeness and accuracy of user profiles in detail.

[0150] To use the clustering results for building user profiles, the clustering information needs to be converted into a numerical vector representation. The process of generating basic attribute vectors is as follows:

[0151] ① Cluster tag vector generation: For each user, a cluster tag vector is generated based on the cluster they belong to; then, it is encoded using one-hot encoding to obtain the cluster tag vector.

[0152] ② Basic attribute vector generation: The cluster label vector is fused with the normalized static attribute data to obtain the basic attribute vector.

[0153] (2) Generate behavioral preference vectors based on structured behavioral data

[0154] Because structured user behavior data contains a wealth of information about user habits and preferences, such as data plan usage, internet traffic, and consumption frequency, if a user frequently uses a certain data plan or regularly purchases a certain type of value-added service, these behaviors can be numerically recorded as components in a feature vector so that algorithms can identify the user's behavioral preferences. By vectorizing this behavioral data, complex behavioral features can be transformed into numerical vector representations, making them data that subsequent profile building algorithms can process. For example, structured behavioral data may include:

[0155] 1. Package Usage: This includes the type of package selected by the user, data usage, and call duration. This type of data reflects the user's preference for different types of packages, such as whether the user prefers a large data package or a package that supports voice calls.

[0156] 2. Internet Traffic and Application Preferences: This data shows a user's internet traffic usage at different times, the types of applications accessed, and the frequency of use. This data can reveal a user's online habits, such as whether they prefer using video streaming, social media, or shopping applications.

[0157] 3. Consumption Records: This includes the user's recharge history, value-added service purchase history, and consumption frequency on the platform. This information can be used to analyze the user's consumption habits and purchasing power, and to predict the services the user may be interested in.

[0158] 4. Call and SMS records: This includes the user's call frequency and SMS volume. This type of data not only reflects the user's daily communication habits but also helps determine whether the user has a need for certain services (such as international calls and SMS packages).

[0159] 5. Device and Network Usage Preferences: This includes the type of device used by the user (e.g., phone brand, operating system version), network connection type (e.g., 4G, 5G, Wi-Fi), and their switching patterns. This information can help predict the frequency of device upgrades and the user's sensitivity to network speed.

[0160] 6. Location Information and Movement Trajectory: User location information and movement trajectory at different time periods, such as frequently visited locations and travel patterns. This information can help infer user lifestyle habits (such as whether they travel frequently for business or their daily activity areas) and their demand for services in different regions.

[0161] The system collects implicit user feedback, namely, users' historical behavioral information, including their data plan usage and internet traffic. To more accurately reflect changes in user interests, this invention introduces a time-aware weight adjustment mechanism based on traditional user behavior modeling methods. By dynamically considering the temporal distribution of user behavior and combining it with the behavioral characteristics of users in different time periods, different time weights are assigned to the behavioral data. Specifically, the preset time weight adjustment formula is expressed as follows:

[0162]

[0163] In the formula, This represents the time weight of term w in document d, i.e., the timeliness weight of a certain behavioral feature in the user behavior sequence; The frequency of term w in document d at time t, that is, the frequency of a certain behavioral feature at a specific point in time; The maximum frequency of all terms in document d; This is the time decay coefficient, a constant greater than 0, used to control the rate at which the weight decays over time; the larger the value, the faster the decay. is the difference between the current time and the time when the behavior occurred.

[0164] In addition, in view of the dynamic changes in user behavior, the embodiments of the present invention introduce an adaptive behavior change weight mechanism, which dynamically adjusts the weights of each behavior label by evaluating the activity changes of behavior characteristics in different time periods. The preset behavior weight adjustment formula is expressed as follows:

[0165]

[0166] In the formula, is the adaptive behavior weight of the term w at time t; N is the total number of documents, that is, the total number of user behavior records or the total number of independent events; is the number of documents containing the term w, that is, the number of records in which this behavior occurs. The smaller this value is, the more unique this behavior is, and the larger the logarithmic part value in the weight is; is the adaptive time decay coefficient, which can be flexibly adjusted to adapt to the decay characteristics of different behaviors.

[0167] Through this improved time-series aware behavior weight mechanism, the system can more accurately capture the spatio-temporal dynamic changes of user behavior, effectively reflect the latest trends of user interests, and extract the information of these changes as dynamic labels in the user profile.

[0168] In summary, the preset time weight adjustment formula and / or the preset behavior weight adjustment formula can be used to assign weights to each behavior feature in the structured behavior data to generate corresponding behavior preference vectors.

[0169] (3) Generating an interest vector based on unstructured behavior data

[0170] When constructing a user profile, in addition to relying on structured behavior data, unstructured behavior data (such as text data) can also provide rich user interest and preference information for the system. By deeply mining this data, more targeted interest vectors can be generated to provide decision support for personalized recommendation. Taking text data as an example, the specific generation process of the interest vector is described below:

[0171] 1. Text preprocessing

[0172] ① Word segmentation and stop word removal: Perform word segmentation on the text content and remove irrelevant stop words such as "of" and "is", and retain useful keywords. ]>

[0173] ② Part-of-speech tagging and named entity recognition: Tag the part-of-speech of the keywords and identify specific entities such as products and locations. For example, "5G package" can be recognized as a product entity to provide support for subsequent label generation.

[0174] ③ Noise removal: Remove redundant sentences and other noisy content to ensure the semantic clarity of the text.

[0175] After preprocessing, the text content is transformed into a high-quality keyword set, which facilitates subsequent topic mining.

[0176] 2. Multi-level topic modeling based on dynamic tracking of interest preferences

[0177] To better reflect changes in users' interests over different periods, this invention adds a dynamic interest preference tracking mechanism to the traditional LDA model, constructing a multi-level topic model structure to capture subtle changes in users' interests. This method not only focuses on users' topic preferences but also automatically stratifies users based on behavioral changes, identifying long-term stable interests and short-term fluctuating interests. Specifically, users' interest topics can be divided into two layers: "long-term interests" and "short-term interests," and an interest persistence detection mechanism is introduced to dynamically identify user behavior patterns and automatically adjust the topic layering structure.

[0178] Long-term interest layer: This layer focuses on the thematic preferences that users maintain over a long period of time, such as a user's long-term interest in topics like "short videos" and "data plans." By setting a thematic persistence threshold (such as themes that appear repeatedly in multiple behaviors), the system automatically identifies and marks the user's long-term interest themes.

[0179] Short-term interest layer: This layer captures recent peaks in user behavior, such as a sudden and frequent search for "cloud gaming" or "discount offers," which the system marks as short-term interests. Topics in the short-term interest layer gradually diminish over time, ensuring the real-time nature of the user profile.

[0180] The interest persistence detection mechanism can determine whether a topic belongs to long-term or short-term interest by calculating the "persistence frequency" and "decay frequency" of each topic. The specific formula is as follows:

[0181]

[0182] In the formula, The probability or score of a user's sustained interest in topic i is represented by T. The higher this value, the more likely the topic is to appear continuously in the user's behavior and be a long-term interest. T represents the total number of time periods, such as the total number of days or weeks observed. The frequency of topic i at time t, i.e., the number of times a user exhibits behavior related to that topic within the time period t; N is the total number of identified topics; This represents the total frequency of all topics at time t.

[0183] The interest mining process based on this optimization method is as follows:

[0184] S1. Topic Extraction and Interest Persistence Detection: The system first extracts topics from the text data using the LDA model, and then uses the frequency of topic occurrence and the duration of topic persistence to determine the topic category.

[0185] S2. Dynamic Updates of Interest Tags: The system updates interest tags periodically based on changes in user behavior. Short-term interest tags are updated quickly with new behavioral changes, while long-term interest tags remain stable.

[0186] S3. Fusion of Long-Term and Short-Term Interest Tags: The system will fuse the generated long-term and short-term interest tags to form a complete interest vector. Among them, short-term interest tags have a higher weight, reflecting the user's immediate needs; long-term interest tags have a lower weight, but provide the user's stable interest direction.

[0187] After the above processing, the improved LDA multi-level topic model can output the following final result:

[0188] ① Long-term interest tags: such as "short videos" and "social media", which represent the user's consistent interests over a long period of time;

[0189] ② Short-term interest tags: such as "5G packages" and "cloud gaming" and other recently popular interest topics, representing the user's current interest needs;

[0190] ③ Interest Vector: A complete user interest profile that integrates long-term and short-term interest tags, accurately capturing changes in user interests.

[0191] By introducing a multi-level topic model and a dynamic interest tracking mechanism, the system can make tiered recommendations based on long-term and short-term interest tags in user profiles. Short-term interest tags prioritize recommending currently trending content, while long-term interest tags provide supplementary content that aligns with the user's long-term habits. For example, when a user shows short-term interests such as "5G plans," "short videos," or "data packages" in their profile, the system will prioritize recommending this type of content, while long-term interest tags such as "video entertainment" will serve as secondary options.

[0192] S340. Generate user profile vectors corresponding to the target user based on basic attribute vectors, behavioral preference vectors, and interest vectors.

[0193] In this embodiment of the invention, the generated basic attribute vector, behavior preference vector, and interest vector can be standardized respectively, and then concatenated along the feature dimension to form the final high-dimensional unified user profile vector. Each dimension of the vector represents a feature dimension, and the value on each dimension represents the user's weight or score on the corresponding feature.

[0194] S350: Based on user profile vectors, generate personalized recommendation results for target users.

[0195] In this embodiment of the invention, the personalized recommendation process is as follows:

[0196] S1. Initial screening of recommended content

[0197] Based on the basic quantity vector of the target users, content that matches the basic characteristics of the users is initially filtered from the recommendation content library. This process can filter out content that is obviously irrelevant to the users, making the initial scope of recommendations more consistent with the basic characteristics of the user profile.

[0198] S2, Multidimensional Comprehensive Vector Calculation

[0199] After initial screening of candidate content, the system utilizes behavioral preference vectors and interest vectors, such as combining consumption preference tags, historical behavior tags, and interest tags, to calculate a user's multi-dimensional composite vector. This vector is then weighted and summed based on the weights of various tags in the behavioral preference vector and interest vector to form the final multi-dimensional composite vector. The calculation formula is as follows:

[0200]

[0201] In the formula, The recommended content i is a comprehensive interest score for the user; is the weight of the j-th type of label. The weights of different types of labels (such as basic attributes, historical behavior, short-term interests, etc.) can be dynamically adjusted according to their importance; M is the total number of label or feature categories. The degree of matching between a user's preference on tag j and candidate content i can be calculated using cosine similarity.

[0202] In this way, tags and feature information from different sources can be comprehensively considered to generate a user's overall interest score for each candidate content.

[0203] S3, Dynamic Timing Adjustment

[0204] Based on the generated multi-dimensional composite vector, the system also dynamically adjusts the comprehensive interest score by incorporating the time weight of user behavior. That is, recently occurring behaviors are given higher weights, while earlier behaviors or long-term interest tags are given relatively lower weights. This mechanism ensures that recommended content reflects the user's latest needs, enhancing the real-time nature and timeliness of recommendations.

[0205] S4, Collaborative Filtering and Similar User Recommendation

[0206] The system uses clustering algorithms to perform collaborative filtering recommendations based on the profiles of similar users. For example, if multiple similar users frequently purchase similar value-added services after selecting a certain package, the system will add these value-added services as recommendations to the current user's recommendation list to improve the diversity and coverage of recommendations.

[0207] S5. Sorting and Outputting Recommendation Results

[0208] All candidate content is sorted according to interest scores and displayed to the user based on its final priority. Higher-priority content appears at the top of the recommendation page, while less desirable content appears as an alternative in the recommendation list. The final recommendation results satisfy both the user's short-term needs and long-term interests, ensuring the relevance and usability of the content.

[0209] The technical solutions of the embodiments of the present invention have at least the following beneficial effects:

[0210] (1) Multi-dimensional data integration improves user profile accuracy: This invention integrates explicit and implicit user feedback to achieve multi-dimensional data integration and comprehensively depict user characteristics. Compared with traditional methods that rely on a single data source, this solution significantly improves the accuracy and completeness of user profiles.

[0211] (2) Dynamic weight optimization improves similarity calculation accuracy: A dynamic weight adjustment mechanism for information entropy and difference is introduced, making similarity calculation more flexible. Unlike traditional methods, this scheme can automatically adjust feature weights to highlight distinctive features, thereby improving the accuracy of similarity calculation.

[0212] (3) Enhanced Real-Time Recommendation Through Dynamic Interest Tracking: A multi-level topic model is adopted, combining long-term and short-term interests to achieve dynamic interest tracking. Compared with existing static topic models, this solution can capture changes in user interests in real time, making recommended content more closely aligned with the user's current needs.

[0213] Example 4

[0214] Figure 4 This is a schematic diagram of a user profile building device provided in Embodiment 4 of the present invention. Figure 4 As shown, the device includes:

[0215] Preprocessing module 41 is used to preprocess the multi-source heterogeneous data of the target user to obtain standardized data; the standardized data includes at least static attribute data and dynamic behavior data;

[0216] The static processing module 42 is used to process static attribute data through a clustering analysis mechanism based on multi-objective optimization to generate basic attribute vectors.

[0217] The dynamic processing module 43 is used to process dynamic behavior data through a time-aware interest mining mechanism to generate behavior preference vectors and interest vectors.

[0218] The user profile generation module 44 is used to fuse basic attribute vectors, behavioral preference vectors, and interest vectors to generate user profile vectors corresponding to the target user.

[0219] Furthermore, based on the above embodiments of the invention, the preprocessing module 41 is specifically used for:

[0220] Based on the target user's identification information, corresponding multi-source heterogeneous data is extracted from the preset telecommunications system database; among which, multi-source heterogeneous data includes user personal basic attribute data, circuit-switched domain signaling data, and packet-switched domain IP packet data;

[0221] Data cleaning, feature transformation, and standardization are performed on multi-source heterogeneous data to obtain a regularized dataset;

[0222] Static attribute data and dynamic behavior data are separated from the regularized dataset and used as standardized data.

[0223] Furthermore, based on the above embodiments of the invention, the static processing module 42 is specifically used for:

[0224] Based on static attribute data, construct a global static feature set that includes the target user and other users;

[0225] Cluster analysis is performed on the global static feature set to determine the target user cluster to which the target user belongs; in the cluster analysis, a pre-set multi-objective optimization function is used to determine the similarity between users, and the pre-set multi-objective optimization function is constructed based on feature difference degree and information entropy;

[0226] Encode the target user cluster to obtain the cluster label vector;

[0227] The cluster label vector is fused with the normalized static attribute data to obtain the basic attribute vector.

[0228] Furthermore, based on the above embodiments of the invention, the dynamic processing module 43 is specifically used for:

[0229] Extracting structured and unstructured behavioral data from dynamic behavioral data;

[0230] By using a preset time weighting formula and / or a preset behavior weighting formula, weights are assigned to each behavior feature in the structured behavior data to generate a behavior preference vector;

[0231] Unstructured behavioral data is preprocessed to obtain a keyword set;

[0232] The keyword set is input into a pre-defined multi-level topic model to mine interest topics and obtain interest vectors; the pre-defined multi-level topic model embeds an interest persistence detection mechanism.

[0233] Furthermore, based on the above embodiments of the invention, the user profile generation module 44 is specifically used for:

[0234] Standardize the basic attribute vector, behavior preference vector, and interest vector.

[0235] The standardized basic attribute vector, behavioral preference vector, and interest vector are concatenated to obtain the user profile vector.

[0236] Furthermore, based on the above embodiments of the invention, the device also includes a recommendation module, specifically used for:

[0237] After generating user profile vectors, personalized recommendation results for the target user are generated based on the user profile vectors using at least one preset recommendation algorithm; wherein the preset recommendation algorithm includes at least a content recommendation algorithm based on vector matching and a collaborative filtering algorithm based on user similarity.

[0238] The user profile building apparatus provided in this embodiment of the invention can execute the user profile building method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0239] Example 5

[0240] Figure 5 A schematic diagram of an electronic device 50 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0241] like Figure 5As shown, the electronic device 50 includes at least one processor 51 and a memory, such as a read-only memory (ROM) 52 and a random access memory (RAM) 53, communicatively connected to the at least one processor 51. The memory stores computer programs executable by the at least one processor. The processor 51 can perform various appropriate actions and processes based on the computer program stored in the ROM 52 or loaded from storage unit 58 into the RAM 53. The RAM 53 can also store various programs and data required for the operation of the electronic device 50. The processor 51, ROM 52, and RAM 53 are interconnected via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.

[0242] Multiple components in electronic device 50 are connected to I / O interface 55, including: input unit 56, such as keyboard, mouse, etc.; output unit 57, such as various types of monitors, speakers, etc.; storage unit 58, such as disk, optical disk, etc.; and communication unit 59, such as network card, modem, wireless transceiver, etc. Communication unit 59 allows electronic device 50 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0243] Processor 51 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 51 performs the various methods and processes described above, such as user profile building methods.

[0244] In some embodiments, the user profile building method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 58. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 50 via ROM 52 and / or communication unit 59. When the computer program is loaded into RAM 53 and executed by processor 51, one or more steps of the user profile building method described above may be performed. Alternatively, in other embodiments, processor 51 may be configured to execute the user profile building method by any other suitable means (e.g., by means of firmware).

[0245] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0246] In some embodiments, the user profile building method may be implemented as a computer program, which is implicitly included in a computer program product. When executed by a processor, the computer program implements the user profile building method of the present invention. The computer program product can be understood as a software product that primarily implements its solution through a computer program. The computer program used to implement the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer program causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer program may be executed entirely on a machine, partially on a machine, partially on a remote machine as a standalone software package, or entirely on a remote machine or server.

[0247] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0248] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0249] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0250] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0251] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0252] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for constructing user profiles, characterized in that, The method includes: The multi-source heterogeneous data of the target user is preprocessed to obtain standardized data; the standardized data includes at least static attribute data and dynamic behavior data. The static attribute data is processed by a clustering analysis mechanism based on multi-objective optimization to generate basic attribute vectors; The dynamic behavior data is processed by a time-aware interest mining mechanism to generate behavior preference vectors and interest vectors. The basic attribute vector, the behavioral preference vector, and the interest vector are fused to generate a user profile vector corresponding to the target user.

2. The method according to claim 1, characterized in that, The preprocessing of multi-source heterogeneous data from the target user to obtain standardized data includes: Based on the identification information of the target user, the corresponding multi-source heterogeneous data is extracted from the preset telecommunications system database; wherein, the multi-source heterogeneous data includes user personal basic attribute data, circuit-switched domain signaling data, and packet-switched domain IP packet data; The multi-source heterogeneous data is cleaned, feature transformed, and standardized to obtain a regularized dataset; The static attribute data and the dynamic behavior data are separated from the regularized dataset and used as the standardized data.

3. The method according to claim 1, characterized in that, The process of processing the static attribute data through a clustering analysis mechanism based on multi-objective optimization to generate basic attribute vectors includes: Based on the static attribute data, a global static feature set containing the target user and other users is constructed; Cluster analysis is performed on the global static feature set to determine the target user cluster to which the target user belongs; wherein, in the cluster analysis, a preset multi-objective optimization function is used to determine the similarity between users, and the preset multi-objective optimization function is constructed based on feature difference degree and information entropy; The target user cluster is encoded to obtain a cluster label vector; The cluster label vector is fused with the normalized static attribute data to obtain the basic attribute vector.

4. The method according to claim 1, characterized in that, The process of processing the dynamic behavior data through a time-aware interest mining mechanism to generate behavior preference vectors and interest vectors includes: Structured and unstructured behavioral data are extracted from the dynamic behavioral data; The structured behavioral data is weighted by using a preset time weighting formula and / or a preset behavior weighting formula to generate the behavior preference vector. The unstructured behavioral data is preprocessed to obtain a keyword set; The keyword set is input into a preset multi-level topic model for interest topic mining to obtain the interest vector; wherein, the preset multi-level topic model embeds an interest persistence detection mechanism.

5. The method according to claim 1, characterized in that, The step of fusing the basic attribute vector, the behavioral preference vector, and the interest vector to generate a user profile vector corresponding to the target user includes: The basic attribute vector, the behavioral preference vector, and the interest vector are standardized. The standardized basic attribute vector, the behavioral preference vector, and the interest vector are concatenated to obtain the user profile vector.

6. The method according to claim 1, characterized in that, After generating the user profile vector, the process also includes: Based on the user profile vector, at least one preset recommendation algorithm is used to generate personalized recommendation results for the target user; wherein, the preset recommendation algorithm includes at least a content recommendation algorithm based on vector matching and a collaborative filtering algorithm based on user similarity.

7. A user profile building device, characterized in that, The device includes: The preprocessing module is used to preprocess the multi-source heterogeneous data of the target user to obtain standardized data; the standardized data includes at least static attribute data and dynamic behavior data. The static processing module is used to process the static attribute data through a clustering analysis mechanism based on multi-objective optimization to generate basic attribute vectors; The dynamic processing module is used to process the dynamic behavior data through a time-aware interest mining mechanism to generate behavior preference vectors and interest vectors. The user profile generation module is used to fuse the basic attribute vector, the behavioral preference vector, and the interest vector to generate a user profile vector corresponding to the target user.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the user profile construction method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the user profile construction method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the user profile construction method according to any one of claims 1-6.

Citation Information

Cited By

  • User portrait construction method and system based on multi-source heterogeneous data fusion

    CN121980062A