Broadband user portrait construction method and device, medium and product
By processing multi-source data using homomorphic encryption technology, the problem of balancing privacy and security with service availability in broadband user profiling is solved. It enables data cleaning, feature extraction, and fusion under encryption, constructing accurate user profiles and supporting personalized services and network optimization.
Patent Information
- Application Number
- CN202511166703.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies struggle to achieve secure retrieval and efficient querying while protecting privacy when constructing broadband user profiles, making it difficult to balance the security and availability of user data.
Homomorphic encryption technology is used to process multi-source data, including encryption, cleaning, feature extraction and fusion, to build broadband user profiles, ensuring that the data is always in an encrypted state, supporting operations and queries in the encrypted state, and avoiding the decryption process.
It achieves the goal of protecting user privacy while ensuring data availability and query efficiency, building accurate broadband user profiles, and supporting personalized service recommendations and network optimization.
Smart Images

Figure CN121000366A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a broadband user portrait construction method and device, medium and product. BACKGROUND
[0002] With the expansion of the broadband user scale, the privacy protection and safe use of user basic information have become the core demand of portrait construction. Although the existing technology can encrypt user information, it has obvious limitations: the encrypted data often needs to be decrypted to be queried, and it is difficult to realize safe retrieval under the premise of protecting privacy; at the same time, the existing query mechanism is easy to sacrifice the effectiveness and efficiency of the query when processing encrypted information, and cannot accurately retrieve specific user information from encrypted data to support multi-scene data analysis and decision-making. This leads to the difficulty in balancing the security and usability of user data, and restricts the reliable application of broadband user portrait technology. SUMMARY
[0003] At least one embodiment of the present application provides a broadband user portrait construction method, device, medium and product, which solves the problem that the existing technology cannot guarantee the privacy security and actual business usability of user data at the same time.
[0004] In order to solve the above technical problems, the present application is implemented as follows:
[0005] In a first aspect, the embodiments of the present application provide a broadband user portrait construction method, comprising:
[0006] obtaining homomorphic encrypted multi-source data;
[0007] cleaning and preprocessing the multi-source data to obtain cleaned multi-source data;
[0008] extracting and fusing features from the cleaned multi-source data, determining fused comprehensive features, and constructing a broadband user portrait based on the comprehensive features.
[0009] Optionally, obtaining homomorphic encrypted multi-source data comprises:
[0010] encrypting flow data in a deep packet inspection (DPI) original data table according to a preset additive homomorphic encryption model, and encrypting measurement report (MR) data according to a preset fully homomorphic encryption model;
[0011] obtaining homomorphic encrypted multi-source data; the multi-source data at least includes encrypted DPI original flow data and encrypted MR data.
[0012] Optionally, cleaning and preprocessing the multi-source data to obtain cleaned multi-source data comprises:
[0013] transforming the encrypted multi-source data into a unified format to obtain first data in a unified format;
[0014] performing a deduplication operation on the first data to obtain second data after deduplication;
[0015] performing missing value filling on the second data based on time series analysis or a preset random forest regression algorithm to obtain third data after filling;
[0016] identifying abnormal values in the third data based on a preset business rule or a preset anomaly detection model to obtain fourth data after removing abnormal values;
[0017] performing a preset standardization process on the fourth data to obtain cleaned multi-source data.
[0018] Optionally, feature extraction and fusion are performed on the cleaned multi-source data to determine fused comprehensive features, including:
[0019] obtaining user access application records and traffic records in DPI data from the cleaned multi-source data;
[0020] determining total traffic consumption of each type of application in a preset time period from the traffic records in the DPI data, and taking the total traffic consumption as traffic feature data;
[0021] storing the user access application records and the traffic feature data in a preset user network behavior data table;
[0022] processing location information and geographic information data of MR data in the cleaned multi-source data using a preset clustering algorithm to obtain location pattern features of users;
[0023] storing the location pattern features in a preset user geographic location information table;
[0024] determining fused comprehensive features from the user network behavior data table and the user geographic location information table.
[0025] Optionally, determining fused comprehensive features from the user network behavior data table and the user geographic location information table includes:
[0026] encoding features in the user network behavior data table and the user geographic location information table to obtain encoded features;
[0027] The preset function in the feature fusion network is used to fuse the coded network behavior features and geographic location features to determine comprehensive features after fusion.
[0028] Optionally, after constructing the broadband user portrait based on the comprehensive features, the method further includes:
[0029] Based on the constructed broadband user portrait and the corresponding labeled data of the user in different scenarios, a scenario recognition model is determined.
[0030] According to the obtained geographic information and network behavior information of the to-be-recognized scenario, an encryption feature vector of the to-be-recognized scenario is determined.
[0031] The encryption feature vector is input into the scenario recognition model to determine the personalized service content of the recommended user.
[0032] Optionally, after constructing the broadband user portrait based on the comprehensive features, the method further includes:
[0033] An incremental learning model is constructed; the incremental learning model updates model parameters using online stochastic gradient descent in the training process, and adjusts the adaptive learning rate through a preset optimizer.
[0034] When the update frequency or the preset user behavior is reached, the incremental learning model is updated based on a sliding window mechanism.
[0035] The broadband user portrait is updated through the incremental learning model.
[0036] In a second aspect, the embodiments of the present application further provide a broadband user portrait construction device, which includes:
[0037] A first acquisition module is configured to acquire homomorphic encryption multi-source data.
[0038] A second acquisition module is configured to clean and preprocess the multi-source data to obtain cleaned multi-source data.
[0039] A first processing module is configured to extract and fuse features of the cleaned multi-source data to determine comprehensive features after fusion, and construct a broadband user portrait based on the comprehensive features.
[0040] In a third aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method according to any one of the first aspect are implemented.
[0041] In a fourth aspect, the embodiments of the present application further provide a computer program product comprising computer instructions for implementing the steps of the method according to any one of the first aspect when executed by a processor.
[0042] Compared with the prior art, the wideband user portrait construction method, device, medium and product provided by the embodiments of the present application respectively use homomorphic encryption for multi-source data, ensure that the data is always in an encrypted state during storage and transmission, and guarantee privacy security from the source. Meanwhile, the homomorphic encryption supports specific operations in an encrypted state, and reserves data availability for subsequent processing. The encrypted data is directly cleaned and preprocessed based on the homomorphic encryption algorithm to obtain cleaned multi-source data, and the data cleaning can be completed without decryption, which avoids privacy leakage and guarantees data quality to support business analysis. The cleaned multi-source data is subjected to feature extraction and fusion, the integrated features after fusion are determined, and a wideband user portrait is constructed based on the integrated features. The entire process does not require decryption of data, and the feature integrity and precision required for portrait construction are met, thereby achieving a balance between privacy protection and business application. BRIEF DESCRIPTION OF DRAWINGS
[0043] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are intended to depict only preferred embodiments of the application, and therefore should not be considered to narrow the scope of the present application. Rather, the entire disclosure including the full description and the drawings are to be considered to define the scope of the application. Moreover, the same reference numerals in different drawings represent the same or similar components.
[0044] Figure 1 A flowchart of a wideband user portrait construction method provided by the embodiments of the present application;
[0045] Figure 2 A flowchart of a wideband user portrait construction method provided by the embodiments of the present application;
[0046] Figure 3 A structure diagram of a wideband user portrait construction device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0047] The terms "first", "second", and the like in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second" are generally a class, not limited to the number of objects, for example, the first object can be one or more. In addition, "or" in this application means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, scenario one: including A and not including B; scenario two: including B and not including A; scenario three: including A and B. The character " / " generally represents that the objects before and after are in an "or" relationship.
[0048] The term "indication" in this application can be a direct indication (or explicit indication) or an indirect indication (or implicit indication). Among them, the direct indication can be understood as that the sender explicitly informs the receiver of the specific information, the operation to be performed or the request result, etc. in the sent indication; the indirect indication can be understood as that the receiver determines the corresponding information according to the indication sent by the sender, or judges and determines the operation to be performed or the request result according to the judgment result.
[0049] As described in the background, in the prior art, most solutions rely on offline processing mode, which causes delay in updating data analysis and user portrait, cannot reflect the latest behavior changes of users in real time, and affects the performance and effect of real-time tasks or functions. Secondly, the prior art solution often ignores the importance of data privacy protection when processing and analyzing user data, lacks effective mechanisms to protect the security of user data during analysis. Finally, although some solutions try to use machine learning methods to improve the accuracy of user portrait, they do not consider the frequent update of data source, large-scale data set and complex model, etc., which leads to the need to retrain the model, increases the complexity and maintenance cost of the system, and may introduce more errors and problems. To solve at least one of the above problems, the embodiments of the present application provide a broadband user portrait construction method, device, medium and product, which can reduce or avoid the above situations, improve the efficiency and accuracy of user portrait construction, and pay more attention to the privacy protection of user data.
[0050] Please refer to Figure 1 The broadband user portrait construction method provided by the embodiments of the present application comprises:
[0051] Step 11, obtaining the homomorphic encryption of the multi-source data.
[0052] In the embodiments of the present application, step 11 realizes data privacy protection from the source while retaining the operability of the data. The multi-source data range includes but is not limited to: covering various types of data related to broadband users, such as DPI (Deep Packet Inspection) traffic data (including access records, traffic consumption), MR (Measurement Report) data (including signal strength, location information), user basic information (such as ID, package type), KPI performance data (such as network delay, bandwidth), etc. According to the data type and operation requirements, different homomorphic encryption models are adopted, such as AHE (Additive Homomorphic Encryption) for scenarios that require frequent addition operations on traffic data, supporting summation in the encrypted state, such as calculating total traffic; PHE (Partially Homomorphic Encryption) for basic information that requires multiplication operations, such as user packages, to meet specific analysis requirements, such as package matching; FHE (Fully Homomorphic Encryption) for behavior data that requires complex feature extraction, supporting algorithm execution in the encrypted state, such as time series analysis. The present application ensures that the data exists in an encrypted form throughout the entire process of collection, transmission, and storage, eliminating the risk of plaintext leakage; at the same time, the operability of homomorphic encryption retains the data value for subsequent processing, avoiding the loss of business usability due to encryption.
[0053] Step 12: cleaning and preprocessing the multi-source data to obtain cleaned multi-source data.
[0054] It should be noted that, based on the characteristics of the homomorphic encryption algorithm, the cleaning operation is directly performed on the encrypted data without decryption. The preprocessing operations include but are not limited to: deduplication: identifying duplicate records through encrypted hash value comparison and removing redundant data; missing value completion: using the additive property of AHE to calculate encrypted mean, interpolation algorithm of FHE, etc., to fill in missing data in the encrypted state; outlier processing: identifying outliers through statistical analysis in the encrypted domain and marking or correcting them in the encrypted state; standardization: performing linear transformation on encrypted features to map them to a unified encrypted range, eliminating dimension differences. Data denoising is completed without exposing the plaintext, ensuring that the data input into the subsequent steps is both secure and reliable, solving the contradiction between encryption and usability.
[0055] Step 13: feature extraction and fusion of the cleaned multi-source data, determining the integrated features after fusion, and constructing a broadband user portrait based on the integrated features.
[0056] In the embodiments of the present application, based on the operations supported by homomorphic encryption, key features are extracted from encrypted data, such as application usage frequency, traffic consumption patterns, etc. from encrypted DPI data; spatial features such as frequently visited locations and movement trajectories are extracted from encrypted MR data; network experience indicators are extracted from encrypted KPI data. In the encrypted state, multiple source features are fused into comprehensive features to form a comprehensive user feature system. Based on the fused encrypted comprehensive features, a user portrait is generated containing dimensions such as user behavior habits, network demand, service preferences, etc. to support subsequent personalized service recommendation, network optimization and other business scenarios. The entire process of the present application completes feature integration and portrait construction in the encrypted domain, which not only guarantees that user privacy is not leaked, but also ensures the accuracy and business guidance of the portrait, and finally realizes the balance between security and usability.
[0057] The three-step process of the present application constructs a full-link security processing mechanism of encrypted collection, encrypted cleaning and encrypted fusion through the application of homomorphic encryption technology, which not only eliminates the risk of data plaintext exposure, but also retains the core value of data in business analysis, thereby solving the problem that privacy security and business usability are difficult to balance in the prior art.
[0058] Optionally, step 11 described above comprises:
[0059] According to the preset additive homomorphic encryption model, the traffic data in the deep packet inspection (DPI) original data table is encrypted, and according to the preset full homomorphic encryption model, the measurement report (MR) data is encrypted.
[0060] Obtain the homomorphic encrypted multi-source data; the multi-source data at least includes the encrypted DPI original traffic data and the encrypted MR data.
[0061] It should be noted that the deep packet inspection (DPI) original traffic data in the multi-source data can be monitored in real time by deep packet inspection (DPI) technology to accurately capture and record online activity data of users, thereby providing rich original data sources for constructing user fingerprint portraits. The specific implementation includes but is not limited to: data packet capture, data packet classification, feature extraction and other sub-steps.
[0062] In the process of data packet capture, DPI devices are deployed at key nodes of the network, and specific capture rules are used to intercept and store all data packets passing through the network in real time. This process can be represented by the following formula:
[0063] N captrued =f(N total )
[0064]
[0065] where f(x) represents the set capture rule function, N captrued represents the number of packets captured by the DPI device within a given time window, N total represents the total number of packets passing through the node, C r represents the packet capture rate. By optimizing the configuration and capture rules of the DPI device, the goal is to make C r as close to 1 as possible, i.e., to capture as comprehensively as possible the packets passing through the network.
[0066] In the process of packet classification, packet classification is based on the header information and payload content of the packet, using pattern recognition and classification algorithms to analyze the type, destination IP, source IP, protocol type, application type, etc. Key information of the captured packet and classify, such as video traffic, social media traffic, email traffic, etc., providing a basis for subsequent feature extraction and user behavior analysis.
[0067] In the process of feature extraction, i.e., extracting key features from each classified packet for constructing user fingerprint portraits. This includes application recognition, session duration, access frequency, etc. Feature extraction can be done by the following formula:
[0068] F extracted =f(D header ,D payload )
[0069] where f(x) represents the feature extraction function, F extracted represents the extracted feature set, D header and D payload represent the header information and payload content of the packet, respectively. The design of the feature extraction function f(x) needs to consider the representativeness and discriminability of the features, in order to accurately reflect the user's online behavior.
[0070] In the embodiments of the present application, through the homomorphic encryption technology and the private information retrieval (PIR) mechanism, secure query processing is realized on the encrypted user base information table without decrypting the data, to protect the privacy and security of user data. In the data collection phase, according to the preset additive homomorphic encryption model, the traffic data in the deep packet inspection (DPI) original data table is encrypted, and according to the preset fully homomorphic encryption model, the measurement report (MR) data is encrypted. The present application uses homomorphic encryption technology to independently encrypt each item of data (e.g., user ID, age, gender, subscription package, etc.) in the user behavior data and location data table model before storing it into the database, to ensure that the data still maintains operational flexibility in encrypted form.
[0071] Specifically, the preset additive homomorphic encryption model AHE is used to encrypt the flow data in the DPI original data table. The preset additive homomorphic encryption model allows addition operation on the encrypted flow data, so that the total flow usage of a user can be calculated without decrypting the data. The specific steps are as follows: (1) generating a key pair: generating two large prime numbers p and q, ensuring that their lengths are sufficient (64 bits) to meet security requirements. (2) calculating n = p * q and λ = lcm(p-1, q-1), where lcm(x, y) represents the least common multiple of x and y. (3) selecting a random integer g such that g is in the multiplicative group modulo n 2 . (4) calculating μ = (L(g λ mod(n 2 )) -1 mod(n)) where The public key is (n, g), and the private key is (λ, μ). (5) encrypting the flow data. For each user's flow data m, use the public key (n, g) to encrypt the formula c = g m *r n mod(n 2 ) for encryption operation. Where m satisfies the relationship m < n, and r is a random value satisfying 0 < r < n. (6) addition operation. Given two encrypted flow data c1 and c2, the encrypted result c total of the total flow of a user can be obtained in the encrypted form: c total = c1 * c2 mod(n 2 ). (7) decryption (only when needed). Use the private key (λ, μ) to decrypt the encrypted total flow c total , and get the plaintext result m total of the total flow of a user:
[0072] The partial homomorphic encryption model PHE is used to encrypt and protect the user basic information table containing basic information such as user subscription packages. This encryption method supports multiplication operation on encrypted data, which can process certain specific data analysis requirements while ensuring data security.
[0073] The full homomorphic encryption model FHE is used to encrypt and store the user behavior feature table obtained through DPI data and other source analysis. This encryption method supports directly executing complex feature extraction algorithms (such as time series analysis, frequent item set mining, etc.) on encrypted data, without the need to decrypt to update user behavior features.
[0074] Through the above homomorphic encryption, the homomorphic encrypted multi-source data is obtained; the multi-source data at least includes the encrypted DPI original flow data and the encrypted MR data.
[0075] Optionally, the step 12 described above comprises:
[0076] Converting the encrypted multi-source data into a unified format to obtain first data in a unified format;
[0077] Performing a deduplication operation on the first data to obtain second data after deduplication;
[0078] Based on time series analysis or a preset random forest regression algorithm, filling missing values in the second data to obtain third data after filling;
[0079] Based on a preset business rule or a preset anomaly detection model, identifying abnormal values in the third data to obtain fourth data after removing abnormal values;
[0080] Performing a preset standardization process on the fourth data to obtain cleaned multi-source data.
[0081] In the embodiments of the present application, the step 12 of performing cleaning and preprocessing on the multi-source data using a homomorphic encryption processing algorithm is a core link for optimizing data quality in an encrypted state. In combination with automatic cleaning logic and homomorphic encryption characteristics, the specific steps are explained as follows: For encrypted multi-source data such as DPI traffic data, MR location data, and user basic information, different encrypted data such as ciphertext generated by different encryption algorithms and encrypted fields with different structures are converted into a unified encrypted data format such as a unified ciphertext coding rule and a field naming specification by using a homomorphic encryption compatible format conversion rule, first data in a unified format is obtained, format differences of encrypted data caused by different sources are eliminated, a consistent processing basis is provided for subsequent cleaning operations, and cross-source data can be analyzed collaboratively.
[0082] Optionally, a unique encrypted hash value is generated for the first data in a unified format by using SHA256 hash technology. Based on ciphertext content calculation, repeated encrypted records are quickly identified by comparing the encrypted hash values, and redundant data is removed. Here, SHA256 (Secure Hash Algorithm 256-bit) is a cryptographic hash function belonging to the SHA-2 family of standard algorithms. It converts input data of any length into a fixed length (256 bits or 32 bytes) hash value, which is usually represented as a 64-bit hexadecimal string.
[0083] The deduplication formula is represented as: R unique = R total - R duplicate ; where R unique represents the number of unique data records, R total represents the total number of data records, and R duplicateThe number of repeated data records is represented. The repeated information is accurately removed without decrypting the data, the data redundancy is reduced, and the accuracy of subsequent analysis is ensured.
[0084] Based on time series analysis or preset random forest regression algorithm, the second data is filled with missing values to obtain the third data after filling, that is, for the missing values in the encrypted data, the operation characteristics supported by homomorphic encryption are used to perform time series analysis such as ARIMA model on encrypted data with time dependence such as time series flow records, mine the time regularity of encrypted data, and predict and fill in the missing encrypted values; the random forest regression model can be used for non-time series encrypted data, and the missing encrypted values are predicted by training the model based on the correlation of encrypted features. While protecting data privacy, the data gap is completed through algorithm logic, avoiding the influence of missing values on the integrity of subsequent feature extraction.
[0085] Further, in order to guarantee the accuracy and efficiency in the data processing flow, rule-based and machine learning-based methods are used to identify and correct error values and abnormal values in the data: the rule-based method is to check the data rationality in the encrypted domain according to the preset business rules, such as the normal range of encrypted traffic value and the continuity of encrypted timestamp. For example, according to the historical data analysis of different applications or services, the normal range of traffic consumption is determined. For each traffic record, the system automatically checks whether its traffic value is within the normal range of the application or service. If the traffic value of a record exceeds the normal range, the system will mark it as abnormal and further manual review or automatic correction processing will be performed. The user network activity record should maintain the continuity of the timestamp to avoid time jumps or repetition. By comparing the timestamps of consecutive records, the system can automatically check the continuity of the time sequence. If time jumps or repeated timestamps are detected, it may indicate that there is an error in the data collection or recording process.
[0086] Based on the machine learning method, the local outlier factor (LOF) model can be used to train an anomaly detection model in an encrypted state, learn the normal distribution characteristics of encrypted data, and identify encrypted abnormal values that deviate from the distribution. Specifically, a part of the verified data without anomaly from the historical data of different applications or services is selected as the training set. The local outlier factor LOF is used to train the anomaly detection model on these training data to learn the normal distribution characteristics of the data. After the model training is completed, the real-time collected data is input into the trained anomaly detection model, and the learned data distribution characteristics are used to identify whether there are abnormal values in the data that may represent data collection errors or user abnormal behavior.
[0087] Further, a denoising algorithm based on statistical characteristics is introduced to filter the identified abnormal data using the size, transmission frequency and other characteristics of the data packet. The anomaly detection formula can be represented as: Snormal = {s i | μ - kδ ≤ s i ≤ μ + kδ}; wherein, S normal is the data set identified as normal, s i is a certain statistical feature value of a single data packet, μ and δ are the average value and standard deviation of the feature in the entire data set, respectively, and k is a constant that controls the threshold looseness. In this way, the error or abnormal data is accurately identified and removed in the encrypted state, ensuring data quality.
[0088] Further, the fourth data is subjected to preset standardization processing to obtain the cleaned multi-source data. The data standardization processing makes the data from different sources comparable, which is particularly important when multiple data sources are involved. For each feature f i , the standardized value f i is calculated as follows: Through this process, all feature values will be scaled between 0 and 1, thereby eliminating the influence of different magnitudes and units of measurement, providing standardized input data.
[0089] Step 12 of the present application realizes full-process data optimization in the encrypted state through the combination of homomorphic encryption technology and automated cleaning algorithm, which not only guarantees the privacy and security of user data, but also improves data quality through operations such as deduplication, completion, denoising, and standardization, providing a reliable encrypted data source for subsequent feature extraction and portrait construction.
[0090] Optionally, in step 13, feature extraction and fusion are performed on the cleaned multi-source data to determine the integrated features after fusion, including:
[0091] According to the DPI data in the cleaned multi-source data, user access application records and traffic records in the DPI data are obtained;
[0092] According to the traffic records in the DPI data, the total traffic consumption of each type of application in a preset time period is determined, and the total traffic consumption is taken as traffic feature data;
[0093] The user access application records and the traffic feature data are stored in a preset user network behavior data table;
[0094] According to the location information and geographic information data of the MR data in the cleaned multi-source data, a preset clustering algorithm is used to process the location information and the geographic information data to obtain the location pattern features of the user;
[0095] The location pattern features are stored in a preset user geographic location information table;
[0096] According to the user network behavior data table and the user geographic location information table, a fused comprehensive feature is determined.
[0097] It should be noted that the DPI data in the cleaned multi-source data has undergone the data preprocessing process in step 12, that is, the data preprocessing is performed on the data from different sources by using the development data cleaning script, including cleaning and denoising of the DPI raw data and standardization processing, and consistency adjustment of the MR measurement data, the user basic information and the network quality monitoring data.
[0098] (1) For the DPI raw data table, the hash value of each data record is calculated, and the hash mapping technology is used to quickly identify and remove duplicate data records, so as to ensure the accuracy of the analysis.
[0099] (2) Check whether the signal strength value in the MR measurement data table is within a reasonable range; if not, correct it according to the neighboring data or mark it for manual review.
[0100] (3) Use time series analysis method or linear interpolation method to predict and fill in the missing values in the user basic information table and the network quality monitoring table.
[0101] (4) Standardize the format of the data in all tables to ensure that the data can be interoperated between different tables. For example, unify the timestamp format, standardize the representation of geographic location information, and ensure that all data conform to the same standard.
[0102] (5) Use the local outlier factor (LOF) algorithm to learn the distribution pattern of the data set, and identify and remove abnormal data or noise data in the data set that do not belong to any main group.
[0103] In the embodiments of the present application, the key features are extracted from the preprocessed data and stored in the corresponding data table. For example, after the DPI data is preprocessed, the user network behavior features and the location pattern features in the geographic information data are obtained, the user network behavior features are stored in the user network behavior data table, and the location pattern features are stored in the user geographic location information table. The user network behavior data table and the user geographic location information table are used to reflect the user's network usage habits and physical world activity patterns. Specifically, the records of user access to applications are filtered from the DPI raw data, the frequent item set mining (FIM) algorithm is used to identify the types of applications frequently accessed by the user within a certain time period, the occurrence frequency of each application type set is formed, and the application type set with an occurrence frequency higher than a threshold value θ is taken as an application usage pattern. Let I = {i1, i2,..., i n} be the set of all application types, T = {t1, t2,..., t n} is a set of application types whose frequency F is higher than a predetermined threshold θ. The frequency calculation formula is: where N represents the total number of access time points, and C represents the number of occurrences of a certain application type set in the access record. According to the flow record in the DPI original data, the total flow consumption of each type of application in a specified time period is calculated, and the flow usage amount is extracted as a flow usage amount feature. Let L = {l1, l2,..., li} represent the total flow usage amount of the user in a certain time period, and li n i represents the flow usage amount of the i-th application. The calculation formula of the flow usage amount is: where flow i (t) represents the flow generated by the user at time point t for the i-th application. In combination with the location information in the MR data and the geographic information data, the Kmeans clustering algorithm is used to analyze the frequency of the user at different times and locations, identify the frequently visited places of the user, and form a location pattern feature. In combination with the time stamp and location information in the MR data and the geographic information data, data fusion technologies such as Kalman filtering are used to predict and update the user state, estimate the continuous movement trajectory of the user, and record in the user geographic location information table.
[0104] For example, the multi-source data includes DPI user plane data, MR data, high-precision building map, CAD building file, KPI performance data, complaint data, word-of-mouth data, and package data, etc. The present application can construct a comprehensive feature set that comprehensively reflects the user behavior and preferences by integrating information from multiple data sources such as DPI user plane data, MR data, high-precision building map, CAD building file, KPI performance data, complaint data, word-of-mouth data, and package data, etc. This strategy covers the whole process from preprocessing, feature extraction to feature fusion, ensuring the quality and consistency of the data, and at the same time capturing and fusing the complex cross-source data relationships through advanced deep learning technology.
[0105] Further, according to the user network behavior data table and the user geographic location information table, a fused comprehensive feature is determined.
[0106] Optionally, according to the user network behavior data table and the user geographic location information table, a fused comprehensive feature is determined, comprising:
[0107] encoding the features in the user network behavior data table and the user geographic location information table to obtain encoded features;
[0108] using a preset function in the feature fusion network to fuse the encoded network behavior features and geographic location features to determine a fused comprehensive feature; wherein the comprehensive feature is decoded and stored in a preset comprehensive feature data table.
[0109] In the embodiments of the present application, based on the preset deep learning model, the high-dimensional features in the user network behavior data table and the user geographic location information table are encoded, fused and decoded by using the feature fusion network automatic encoder in the preset deep learning model, and then integrated into the comprehensive feature data table, so as to integrate key features from different sources, build a more comprehensive and accurate broadband user fingerprint portrait, and provide strong data support for personalized services.
[0110] Specifically, the features in the user network behavior data table and the user geographic location information table are encoded to reduce the feature dimension and capture key information. The encoding process can be represented as: z = f encode (x); wherein x is the input feature, z is the encoded feature, f encode (x) is the encoding function of the automatic encoder.
[0111] The encoded features are fused by using the feature fusion network (FFN). The fusion process can be represented as: z fusion = f fusion (z1, z2); wherein z1 is the network behavior feature, z2 is the geographic location feature, f fusion is an RNN deep network designed according to the type of the feature and the fusion requirement, and z fusion is the output fused feature.
[0112] The fused feature z fusion is decoded and stored in the comprehensive feature data table. The decoding process can be represented as: x ’ = f decode (z fusion ); wherein x ’ is the decoded comprehensive feature, f decode (x) is the decoding function.
[0113] Optionally, the process of constructing the broadband user portrait based on the comprehensive feature is to realize the accurate description of the user behavior, demand, preference and other portrait dimensions through the aggregation, mapping and dynamic optimization of multi-dimensional features under the premise of full-process homomorphic encryption protection. Specifically:
[0114] First, the dimension analysis of the comprehensive feature is performed. The comprehensive feature is a high-dimensional feature set obtained by encrypting, cleaning, extracting and fusing multi-source data (such as DPI traffic, MR location, and package information), which covers three core dimensions:
[0115] Network behavior feature: such as application usage frequency (obtained by frequent item set mining), traffic consumption mode (such as traffic proportion of video or social application), access time period distribution, etc.;
[0116] Spatial and scene features: such as frequently visited places (Kmeans clustering results), moving track, scene-related behaviors (such as traffic peak during "at home" period), etc.
[0117] Service and preference features: such as package matching degree, network quality sensitivity (based on KPI data), service demand reflected by complaints or word-of-mouth, etc.
[0118] Further, mapping and aggregation of the portrait dimensions. Through feature association analysis in the encrypted state, complex algorithms supported by homomorphic encryption are used to map the comprehensive features to specific dimensions of the user portrait. For example, based on the encrypted user ID, package type and other features, "user group labels" such as "family user / enterprise user" are associated; through the aggregation of network behavior features, "high-frequency application preference" and "traffic use peak period" labels are generated; combined with spatial features and behavior features, "commuting period network dependence" and "high demand for home entertainment" are mapped; based on KPI data and complaint information, "network delay sensitivity" and "package cost performance concern" service preference labels are generated.
[0119] The entire process is completed in the data encryption state, and feature aggregation and label generation are achieved through homomorphic encryption supported operations without decryption: using clustering and association rule algorithms supported by full homomorphic encryption (FHE), the association rules of behavior-scene-demand are mined in encrypted features; combined with a dynamic incremental learning model, new data is absorbed in real time, and feature weights are updated through a sliding window mechanism. The generated user portrait has privacy protection and business usability. This whole process is encrypted, and the label is the mapping result of the encrypted feature, supporting label query and matching in the encrypted state.
[0120] The scheme of the present application can be directly used for personalized service recommendation and support network optimization decision. The process is designed in a closed loop of encrypted features-portrait dimensions-dynamic optimization, which not only ensures that user data is not leaked throughout the process, but also realizes the accuracy and business guidance of the portrait through the deep association of multi-dimensional features, thereby solving the core problem of balancing privacy security and usability.
[0121] Optionally, in the data analysis stage, the present application develops a query processing mechanism based on homomorphic encryption, which aims to safely retrieve specific user information from the encrypted user basic information table without sacrificing the effectiveness and efficiency of the query function, i.e. considering supporting querying different types of data tables to meet data analysis and decision support in different scenarios, such as DPI user plane data table, MR measurement data table, high-precision building map and CAD building file, KPI performance data table, complaint data table and word-of-mouth data table, package data table, etc., while ensuring the privacy and security of the data.
[0122] (1) For each data table, define the supported query operations and the corresponding homomorphic encryption operation conversion rules.
[0123] (2) Develop a query conversion engine to convert standard query requests into homomorphic operation sequences suitable for encrypted data.
[0124] (3) The result processing unit is responsible for completing the aggregation, sorting, and other operations on the encrypted query results in the encrypted state.
[0125] (4) Only in the final user interface, i.e., under the premise of meeting the security policy, the results are decrypted and displayed to the user.
[0126] Specific applications of table models include: DPI user plane data table: query the traffic usage or access frequency of a specific user. For example, perform a homomorphic addition operation on the encrypted traffic data field to calculate the total traffic usage of a user within a period of time. MR measurement data table: query involves the signal quality of a user within a specific time period. For example, by performing a comparison operation through a homomorphic algorithm, records with signal quality below a certain threshold can be found. High-precision building map and CAD building file: support location-based queries, such as identifying the user's stay time at a specific geographic location. KPI performance data table: query network events that meet certain performance indicator conditions, such as identifying cases where network latency exceeds a threshold. Complaint data table and reputation data table: by performing a homomorphic addition operation on encrypted rating data, calculate the total score or average score to analyze user satisfaction with a specific service or product. Package data table: by performing a matching query on encrypted package information, identify the type of service the user subscribes to.
[0127] Optionally, in step 13 described above, after constructing the broadband user portrait based on the comprehensive features, the method of the present application further comprises:
[0128] Based on the constructed broadband user portrait and the corresponding labeled data of the user in different scenarios, a scenario recognition model is determined;
[0129] According to the obtained geographic information and network behavior information of the to-be-recognized scenario, an encrypted feature vector of the to-be-recognized scenario is determined;
[0130] The encrypted feature vector is input into the scenario recognition model to determine the personalized service content recommended for the user.
[0131] In the embodiments of the present application, based on wideband user portrait and labeled data, a scene recognition model is determined, to construct a completed wideband user portrait (including comprehensive features such as user network behavior, location preference) as the basis, combined with the labeled data of the user in different scenes, the scene recognition model M is trained using the SVM algorithm. The scene recognition model M learns the association rule between portrait features and scene labels, and has the ability to predict the current scene of the user based on new data, providing a scene basis for personalized recommendation.
[0132] Real-time data of the user to be identified is collected, including geographic information: GPS coordinates, base station ID; network behavior information: access application record, traffic data, connected WiFi type, etc., to construct a feature vector The mathematical expression is: Wherein, x i represents the i-th feature extracted from the data. The feature vectors are preprocessed, including standardization and homomorphic encryption technology, to form an encrypted feature vector Based on the labeled data including the activity records of the user in different places and the corresponding scene labels (such as "at home", "at the office"), the scene recognition model is constructed using the SVM algorithm. Such encryption processing ensures that the real-time data does not leak privacy during transmission and model input, and standardization ensures the consistency of the feature vector and the model training data, avoiding the influence of dimension difference on prediction accuracy.
[0133] The encrypted feature vector is input into the scene recognition model, which performs operations in an encrypted state and outputs a predicted scene; combined with the predicted scene and the wideband user portrait of the user, the personalized algorithm is used to recommend adaptive content. The present application completes scene prediction and recommendation based on encrypted data throughout, protecting the real-time location and behavior privacy of the user, and through the linkage of scene and portrait, ensuring that the service recommendation is highly matched with the current needs of the user.
[0134] In a specific embodiment, the scene intelligent recognition technology based on geographic information analyzes the location data and broadband usage behavior of the user, and provides personalized service or content recommendation according to the location of the user, the process being as follows:
[0135] Geographic location information and network behavior data are collected, including user ID, GPS coordinates, base station ID, timestamp, access application record and traffic data, etc., to construct a feature vector The mathematical expression is: Wherein, x i represents the i-th feature extracted from the data. The feature vectors are preprocessed, including standardization and homomorphic encryption technology, to form an encrypted feature vector Based on the labeled data including users' activity records at different locations and their corresponding scenario labels (e.g., "at home", "at office"), a scenario recognition model M is constructed using SVM algorithm. The encrypted feature vector is input into the scenario recognition model M, and the predicted scenario S is output. Based on the recognized scenario and the user's broadband usage profile, personalized algorithm is used to recommend services or content to the user. For example, if the model recognizes that the user is at home S = "at home", then video content or information services related to home entertainment are recommended.
[0136] where the algorithm formula is outlined as follows: Feature vector construction: Encrypted feature vector: Scenario prediction model:
[0137] In the scenario intelligent recognition based on geographic information, the detailed list of scenarios and labels is crucial for training an efficient and accurate machine learning model. The present application provides the following scenarios and their corresponding labels, which cover the range of users' daily activities and can be used to improve the quality and accuracy of personalized services, meeting users' immediate needs and preferences.
[0138] (1) Features of being at home: connecting to home WiFi, activities during night time period, accessing home entertainment related applications, etc.
[0139] (2) Features of being at office: connecting to enterprise WiFi, activities during work time period, using office software and services, etc.
[0140] (3) Features of being on public transportation: fast moving speed, connecting to public transportation WiFi (e.g., subway, bus WiFi), accessing transportation or map applications, etc.
[0141] (4) Features of being at a coffee shop or restaurant: connecting to coffee shop or restaurant WiFi, accessing food, booking platforms, etc.
[0142] (5) Features of being at a shopping mall: multiple location changes within the shopping mall, accessing shopping related applications, using payment applications, etc.
[0143] (6) Features of being at a gym: connecting to gym WiFi, using fitness tracking applications, long stay at the gym's geographic location, etc.
[0144] (7) Features of being at a library or learning place: connecting to library WiFi, accessing learning resources, long stationary position, etc.
[0145] (8) Features of being on a trip: high speed movement, accessing travel and booking applications, using navigation, etc.
[0146] (9) In entertainment venues (such as cinemas, amusement parks): features such as staying in a specific geographic location, visiting entertainment-related applications, etc.
[0147] (10) In outdoor activities (such as parks, suburbs): features such as activities in large open areas, use of health monitoring or outdoor sports applications, etc.
[0148] Optionally, in step 13 described above, after constructing the broadband user portrait based on the comprehensive features, the method of the application further comprises:
[0149] Constructing an incremental learning model; the incremental learning model updates model parameters using online stochastic gradient descent in the training process, and adjusts the adaptive learning rate through a preset optimizer;
[0150] When the update frequency or the preset user behavior is reached, updating the incremental learning model based on a sliding window mechanism;
[0151] Updating the broadband user portrait through the incremental learning model.
[0152] In the embodiments of the application, the incremental learning model is designed to take the fused comprehensive features as input, and the model updates parameters in real time using online stochastic gradient descent (SGD) in the model training process. The formula for model updating is as follows: where w t is the model parameter at time point t, η is the learning rate, L is the loss function, y t is the true label, is the predicted value of the model under the current parameters, is the gradient of the loss function. Through continuous iteration, the model can gradually adapt to new data. At the same time, a preset optimizer (such as Adam) is introduced for adaptive learning rate adjustment, which dynamically optimizes the learning rate η through first-order moment and second-order moment estimation, ensuring that the model quickly converges when new data flows in. The model of the application supports continuous learning on encrypted data, can absorb new features (such as the user's recent new application preferences) without retraining, provides algorithm support for dynamic updating of the portrait, and at the same time avoids the low efficiency caused by full-volume retraining.
[0153] Further, time windows W i and update frequencies F i are set for different types of data. When the update frequency (such as triggering daily checks) or the user behavior is significantly changed (such as a sudden change in traffic pattern), the sliding window moves forward by a preset step S i , and only the latest encrypted data within the window is used to update the model parameters. For example, if the user's recent video traffic proportion suddenly increases, the window length W iAccelerate the learning of new behaviors with a model. By dynamically adjusting the window, the model focuses on recent behavior characteristics of the user, avoiding interference from outdated data, such as a significant difference between the user's preferences in the past six months and the present, ensuring the model's sensitivity to changes in user behavior.
[0154] The incremental learning model is based on updated parameters, and the integrated features are re-aggregated in an encrypted state. For example, the "high-frequency live application access" feature in the new window is included in the portrait, the feature weight is adjusted, such as increasing the "video traffic" feature weight and reducing the "old social application" weight, and the label system of the user portrait is updated, such as adjusting from "social preference" to "video entertainment preference". After updating, evaluate the accuracy of the portrait through the validation set, such as whether the recommendation click rate is improved, to ensure the update effect. The present application realizes real-time iteration of the portrait, so that the portrait always reflects the user's latest behavior and demand, such as the proportion of office software usage features in the portrait gradually increasing with incremental learning after the user changes jobs, and all operations are based on encrypted data to ensure privacy and security.
[0155] In a specific embodiment, in order to comprehensively utilize multi-source data including DPI user plane data, MR data, high-precision building maps, CAD building files, KPI performance data, complaint data, reputation data, and package data, an incremental learning model is constructed as follows: Before designing the incremental learning model, the following problems are considered to ensure that the model can adapt to data streams from different data sources and improve its generalization ability: Heterogeneity processing: Due to the obvious differences in data characteristics (such as data format, update frequency, and data quality) of different data sources, data characteristic analysis is needed to identify and understand the characteristics and potential value of each type of data. Time series analysis: Analyzing the time dependence and periodicity of data with obvious time series characteristics, such as MR data and KPI performance data, helps to reasonably process time features in model design and capture the time dynamics of data. Spatial data processing: Use spatial data analysis methods such as geographic information system (GIS) to process and analyze spatial data such as high-precision building maps and CAD building files, extract spatial features such as building height, number of floors, and building area, and network performance indicators such as signal strength and network delay.
[0156] During the design and implementation of the incremental learning model, the present application adopts a series of strategies to ensure that the model can continuously learn from new data while retaining old knowledge and improving overall performance. The following are the key steps in model design: (1) Feature vector construction. The feature vector is composed of features extracted from each data source, represented as X = [x1, x2, …, x n ], where x i represents a certain feature value. (2) Model update mechanism. Online stochastic gradient descent (SGD) is used to adapt to the continuous input of data streams. The formula for model updating is as follows: where wt is the model parameter at time point t, η is the learning rate, L is the loss function, y t is the true label, is the prediction value of the model under the current parameters, is the gradient of the loss function. Through continuous iteration, the model can gradually adapt to new data.(3) Addressing model forgetting. For the problem of model forgetting, the experience replay mechanism can be used, that is, periodically randomly select samples from historical data to join the training set, or use regularization methods such as Elastic Net to balance the performance of the model on new and old data.(4) Adaptive learning rate adjustment. To adapt to the variability and complexity of data flow, an adaptive learning rate adjustment strategy is adopted, such as the Adam optimizer, whose update formula is as follows: where, and are the bias correction of the first and second moment estimates of the gradient, ∈ is a very small number to prevent division by zero.(5) Performance monitoring and adjustment. By regularly evaluating the performance of the model on the validation set, such as accuracy, recall rate and other indicators, the learning progress and effect of the model are monitored; according to the evaluation results, the learning rate, regularization parameter and other parameters are adjusted to ensure that the model can continuously improve performance.
[0157] Further, in order to effectively utilize multi-source data including DPI user plane data, MR data, high-precision building map, CAD building file, KPI performance data, complaint data, reputation data, package data, etc., this step designs an algorithm framework based on sliding window mechanism. This framework optimizes data processing and model updating by defining the time window and update frequency of different data types, as follows: initialization: set the initial time window W i and update frequency F i for each data type D i . Data stream processing: regularly check and monitor according to the corresponding F i , and update the model as needed. Window update: when the update frequency F i of the data type D i is reached, slide the window forward according to the step S i , and use the latest data in the window for model training or updating. Dynamic adjustment: based on model performance feedback and data flow changes, dynamically adjust the time window W i , update frequency F i and step S i to optimize model performance.
[0158] For example, for DPI user plane data, define W DPI = 1 day, F DPI = every day, if the user behavior changes in the last week, dynamically reduce the time window WDPI and step size S DPI :W DPI_new =W DPI_old / 2;S DPI_new =S DPI_old / 2;
[0159] To ensure that user profiles reflect the latest changes in user behavior and preferences, this application introduces an incremental learning update mechanism. Specifically, when new data is received, this step determines the measures to update the user profile based on the update frequency and importance score of this data. Assume the initial user profile is P0, based on the initial user feature vector X0 = {x1, x2, ..., x...}. n} generated, where x i This represents the features extracted from the data source. When a new dataset D is received... new First, based on the data freshness and its potential impact on user behavior, dataset D is analyzed. new Importance scoring is performed. Next, based on the data type and importance, an appropriate incremental learning method is selected to update the model; for example, for DPI and KPI performance data, an online learning algorithm is used. The mathematical expression is as follows: X new =f(D new ;X0);P new =g(X) new ;P0); where X new P new These are the updated feature vector and the user profile, respectively, and the function f(D) new ;X0) is the feature vector update function, function g(X new P0) is the user profile update function. After the update, the model's performance is evaluated using a validation dataset, such as accuracy and recall, to ensure that the updated user profile is accurate and timely, thereby improving the level of personalized services and user experience.
[0160] Optionally, this application integrates information from multiple data sources, including DPI user face data, MR data, high-precision building maps, CAD building documents, KPI performance data, complaint data, reputation data, and package data, and utilizes techniques such as feature alignment and feature fusion to achieve more efficient cross-domain feature fusion. Feature alignment ensures that similar or identical features from different data sources have a consistent representation after fusion. This step is achieved through the following strategies:
[0161] Standardize the data format: For all data sources, timestamps, numerical and categorical data are standardized to a unified format to ensure data consistency.
[0162] Encoding and mapping, for categorical features, use one-hot encoding to convert them into numerical features that machine learning models can understand. In addition, for features that express the same meaning in different data sources, mapping rules are used to unify them to eliminate potential ambiguity.
[0163] Based on feature alignment, feature fusion integrates features from multiple domains into a unified feature space through the following four main strategies:
[0164] Simple fusion: directly concatenate the pre-processed and aligned feature vectors horizontally to form a comprehensive feature set. This method is simple and fast, but may face the problem of high feature dimension.
[0165] Advanced fusion: apply dimension reduction techniques such as PCA (Principal Component Analysis) and tSNE (tdistributed Stochastic Neighbor Embedding) to effectively reduce feature dimension while preserving key information as much as possible.
[0166] Deep learning: use multi-view learning, ensemble learning and other deep learning structures to learn the low-dimensional compressed representation of data. These models can capture complex non-linear relationships between features and extract high-level features with strong representation ability.
[0167] Multi-view learning: considering that different data sources may provide different perspectives on user behavior, multi-view learning techniques are used to optimize the integration of these perspectives in a shared feature space, revealing a more comprehensive user portrait.
[0168] Ensemble learning techniques: through model stacking or weighted averaging, etc. Ensemble learning techniques integrate the advantages of different deep learning models or dimension reduction methods to further improve the effectiveness of feature fusion and the generalization ability of the model.
[0169] Through a series of feature alignment and fusion strategies, the application can effectively integrate cross-domain data and build a more comprehensive and accurate user portrait, providing enterprises and service providers with more comprehensive and in-depth user insights, enabling them to make more accurate and effective decisions in product development, service optimization, market strategy formulation, etc.
[0170] Step 13 of the present application fuses the key components of the user portrait determined by the feature and the possible application scenarios, including: the key components of the user portrait are: from DPI data and MR data, including the user's online frequency, the most frequently visited website or application, data usage, signal quality, etc., which reflect the user's online behavior habits and mobile network use experience. Combined with high-precision building maps and CAD building files, the spatial distribution characteristics of the user's residence are provided, such as the type of building commonly entered, geographic location preferences, etc., which help to understand the user's living and working environment. Through KPI performance data, the user's network access speed, delay, etc. are analyzed to evaluate the user's network service experience quality. Integrate complaint data and word-of-mouth data, extract user satisfaction, complaint reasons, preferences and needs, etc. through text analysis, to provide a basis for improving service quality and user satisfaction. Combined with package data, analyze the user's consumption preferences, such as preferred service types, consumption levels, etc., and the user's behavior patterns obtained through all data analysis, such as active time period, preferred content, etc.
[0171] Optionally, the application scenarios of the user portrait are: using the user's network behavior characteristics and preference information to recommend content or services that the user may be interested in, improving the user experience. By analyzing the user's geographic location information and network performance experience, data support is provided for network optimization and base station layout to improve the user's network service experience. Combined with service feedback and satisfaction information, the user's service satisfaction is analyzed in depth, and the user's complaints and service quality are targeted to be improved. According to the multi-dimensional characteristics in the user portrait, the user group is more finely subdivided, and more targeted market strategies and service provision are realized.
[0172] Referring to Figure 2 The overall flowchart of the broadband user portrait construction method is shown. The broadband user portrait construction method of the present application comprises:
[0173] DPI process real-time monitoring and capture; develop an automatic data cleaning script; use the automatic data cleaning script to pre-process the data, then extract key features and fuse features from the pre-processed data, obtain high-dimensional fused features, and use the high-dimensional fused features to construct the portrait. Wherein, the present application uses homomorphic encryption technology throughout the process to ensure that the encrypted data can be processed without decryption.
[0174] Optionally, the constructed portrait can be applied to specific scenarios to realize scenario intelligent recognition based on geographic information and recommend personalized content to the current scenario.
[0175] Optionally, the application can also set up a dynamic incremental learning driven user portrait mechanism update. Specifically, an incremental learning model can be designed, a data window for model training is defined, and finally the user portrait is updated in real time through the trained incremental learning model.
[0176] Optionally, the portrait constructed by the application can realize cross-domain fusion analysis.
[0177] In summary, the scheme of the application introduces a series of innovative technical means to address the many shortcomings in the prior art, aiming to provide obvious technical advantages and solutions. These include homomorphic encryption technology, incremental learning algorithm, optimized data processing flow, cross-domain data fusion, and user portrait visualization and management, which work together to improve the accuracy, real-time performance of user portrait construction, and the level of privacy protection of user data.
[0178] First, by using homomorphic encryption technology, the application allows analysis and processing of encrypted data while ensuring user data privacy and security, thereby addressing the shortcomings of existing technologies in user data privacy protection.
[0179] Second, the incremental learning mechanism introduced enables real-time updating of user portraits, effectively reflecting immediate changes in user behavior, and addressing the portrait update delay problem caused by the reliance of traditional technologies on batch processing and data accumulation.
[0180] In addition, by optimizing the data processing flow, such as implementing efficient data sampling, compression, and outlier processing, the application not only improves data processing efficiency and reduces the demand for computing resources, but also reduces the overall burden on the system.
[0181] Finally, the ability of cross-domain data fusion analysis makes the user portrait more comprehensive and in-depth, and can reflect the user's behavior characteristics in different platforms and scenarios, overcoming the limitations of existing technologies that are limited to single data source analysis. By providing visual display of user portraits and allowing users to participate in portrait information adjustment, the application not only increases user trust and satisfaction with services, but also optimizes the personalized service experience.
[0182] In summary, by combining advanced technologies such as homomorphic encryption and incremental learning, the application effectively overcomes the problems in existing technical solutions, such as user portrait update lag, insufficient data privacy protection, and low resource utilization, and provides a practical solution for improving the real-time performance and accuracy of user portrait construction, enhancing data processing efficiency, and significantly improving user privacy security. These technical advantages provide a new solution for the construction and application of broadband user fingerprint portraits.
[0183] The above describes various methods of the embodiments of the application. The following will further provide a device for implementing the above methods.
[0184] Reference is made to Figure 3 The embodiment of the application further provides a wideband user portrait construction device, comprising:
[0185] The first obtaining module 31 is configured to obtain homomorphic encryption multi-source data.
[0186] The second obtaining module 32 is configured to clean and preprocess the multi-source data to obtain cleaned multi-source data.
[0187] The first processing module 33 is configured to extract and fuse features of the cleaned multi-source data, determine integrated features after fusion, and construct a wideband user portrait based on the integrated features.
[0188] Optionally, the first obtaining module 31 comprises:
[0189] The first processing unit is configured to encrypt traffic data in a deep packet inspection (DPI) original data table according to a preset additive homomorphic encryption model, and encrypt measurement report (MR) data according to a preset full homomorphic encryption model.
[0190] The first obtaining unit is configured to obtain homomorphic encryption multi-source data; the multi-source data at least comprises encrypted DPI original traffic data and encrypted MR data.
[0191] Optionally, the second obtaining module 32 comprises:
[0192] The second obtaining unit is configured to convert the encrypted multi-source data into a unified format to obtain first data in the unified format.
[0193] The third obtaining unit is configured to perform a de-duplication operation on the first data to obtain second data after de-duplication.
[0194] The fourth obtaining unit is configured to fill in missing values in the second data based on time series analysis or a preset random forest regression algorithm to obtain third data after filling in missing values.
[0195] The fifth obtaining unit is configured to identify abnormal values in the third data based on a preset business rule or a preset anomaly detection model to obtain fourth data after removing abnormal values.
[0196] The sixth obtaining unit is configured to perform a preset standardization processing on the fourth data to obtain cleaned multi-source data.
[0197] Optionally, the first processing module 33 comprises:
[0198] A seventh obtaining unit is configured to obtain user application access records and traffic records in DPI data from the cleaned multi-source data according to the DPI data in the cleaned multi-source data;
[0199] A second processing unit is configured to determine total traffic consumption of each type of application in a preset time period according to the traffic records in the DPI data, and take the total traffic consumption as traffic feature data;
[0200] A third processing unit is configured to store the user application access records and the traffic feature data in a preset user network behavior data table;
[0201] A fourth processing unit is configured to process location information and geographic information data of MR data in the cleaned multi-source data according to the location information and the geographic information data, obtain location mode features of a user by using a preset clustering algorithm;
[0202] A fifth processing unit is configured to store the location mode features in a preset user geographic location information table;
[0203] A first determining unit is configured to determine fused comprehensive features according to the user network behavior data table and the user geographic location information table.
[0204] Optionally, the first determining unit is specifically configured to:
[0205] encode features in the user network behavior data table and the user geographic location information table to obtain encoded features;
[0206] fuse the encoded network behavior features and the geographic location features by using a preset function in a feature fusion network to determine fused comprehensive features; and store the comprehensive features after decoding in a preset comprehensive feature data table.
[0207] Optionally, the device further includes:
[0208] A first determining module is configured to determine a scene recognition model based on a constructed wideband user portrait and corresponding labeled data of the user in different scenes.
[0209] A second determining module is configured to determine an encrypted feature vector of a to-be-recognized scene according to obtained geographic information and network behavior information of the to-be-recognized scene.
[0210] A third determining module is configured to input the encrypted feature vector into the scene recognition model to determine personalized service content recommended for the user.
[0211] Optionally, the device further includes:
[0212] The construction module is configured to construct an incremental learning model; the incremental learning model updates model parameters by using an online stochastic gradient descent in a training process, and adjusts a self-adaptive learning rate by using a preset optimizer;
[0213] The second processing module is configured to update the incremental learning model based on a sliding window mechanism when a preset user behavior or an update frequency is reached.
[0214] The third processing module is configured to update the wideband user portrait by using the incremental learning model.
[0215] It should be noted that the device in this embodiment corresponds to the wideband user portrait construction method described above, and the implementation modes in the above embodiments are applicable to the embodiments of the device and can achieve the same technical effects. The device described above provided in the embodiments of the present application can implement all the method steps achieved by the method embodiments and can achieve the same technical effects. Therefore, the same parts and beneficial effects in the method embodiments will not be described in detail.
[0216] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement various processes of the wideband user portrait construction embodiments described above and can achieve the same technical effects. To avoid repetition, the same parts will not be described in detail. The computer readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0217] The embodiments of the present application also provide a computer program product, which includes computer instructions. The computer instructions are executed by a processor to implement various processes of the wideband user portrait construction embodiments described above and can achieve the same technical effects. To avoid repetition, the same parts will not be described in detail.
[0218] It should be noted that in the technical solutions of the present disclosure, the collection, collection, update, analysis, processing, use, transmission, storage and other aspects of user personal information comply with relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken to prevent illegal access to user personal information data and to maintain user personal information security and network security.
[0219] It should be noted that, in the present document, the terms "comprises / comprising" or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.
[0220] From the above description of the embodiments, it can be clear to those skilled in the art that the above-mentioned example methods can be realized by means of software plus a necessary general hardware platform, and of course can also be realized by hardware, but in many cases the former is a better embodiment. Based on such an understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk), and includes a plurality of instructions for causing a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
[0221] The embodiments of the present application are described above in combination with the accompanying drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are merely illustrative and not restrictive, and those of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.
Claims
1. A method for constructing broadband user profiles, characterized in that, include: Obtain homomorphically encrypted multi-source data; The multi-source data is cleaned and preprocessed to obtain cleaned multi-source data; Feature extraction and fusion are performed on the cleaned multi-source data to determine the integrated features, and a broadband user profile is constructed based on the integrated features.
2. The method according to claim 1, characterized in that, Obtain homomorphically encrypted multi-source data, including: Based on the preset additive homomorphic encryption model, the traffic data in the DPI raw data table of deep packet inspection is encrypted, and based on the preset fully homomorphic encryption model, the measurement report MR data is encrypted. Obtain homomorphically encrypted multi-source data; the multi-source data includes at least encrypted DPI raw traffic data and encrypted MR data.
3. The method according to claim 1, characterized in that, The multi-source data is cleaned and preprocessed to obtain cleaned multi-source data, including: The encrypted multi-source data is converted into a unified format to obtain the first data after the unified format. Perform a deduplication operation on the first data to obtain the deduplicated second data; Based on time series analysis or a preset random forest regression algorithm, missing values are imputed in the second data to obtain the imputed third data. Based on preset business rules or preset anomaly detection models, identify abnormal values in the third data and obtain fourth data after removing the abnormal values. The fourth data is subjected to preset standardization processing to obtain cleaned multi-source data.
4. The method according to claim 1, characterized in that, Feature extraction and fusion are performed on the cleaned multi-source data to determine the fused comprehensive features, including: Based on the DPI data in the cleaned multi-source data, obtain user access application records and traffic records in the DPI data; Based on the traffic records in the DPI data, determine the total traffic consumption of each type of application within a preset time period, and use the total traffic consumption as traffic feature data. The user's application access records and the traffic characteristic data are stored in a preset user network behavior data table; Based on the location information and geographic information data of the MR data in the cleaned multi-source data, a preset clustering algorithm is used to process the location information and geographic information data to obtain the user's location pattern features. The location pattern features are stored in a preset user geographic location information table; Based on the user network behavior data table and the user geographic location information table, the integrated features after fusion are determined.
5. The method according to claim 4, characterized in that, Based on the user network behavior data table and the user geographic location information table, the integrated features after fusion are determined, including: Encode the features in the user network behavior data table and the user geographic location information table to obtain encoded features; Using a preset function in the feature fusion network, the encoded network behavior features and geographic location features are fused to determine the fused comprehensive features; wherein, the comprehensive features are decoded and stored in a preset comprehensive feature data table.
6. The method according to claim 1, characterized in that, After constructing a broadband user profile based on the comprehensive features, the method further includes: Based on the constructed broadband user profile and the corresponding labeled data of users in different scenarios, a scene recognition model is determined. Based on the acquired geographic information and network behavior information of the scene to be identified, the encrypted feature vector of the scene to be identified is determined; The encrypted feature vector is input into the scene recognition model to determine the personalized service content recommended to the user.
7. The method according to claim 1, characterized in that, After constructing a broadband user profile based on the comprehensive features, the method further includes: An incremental learning model is constructed; during the training process, the incremental learning model updates the model parameters using online stochastic gradient descent and adjusts the learning rate adaptively using a preset optimizer; When the update frequency is reached or a preset user behavior is achieved, the incremental learning model is updated based on the sliding window mechanism. The broadband user profile is updated through the incremental learning model.
8. A broadband user profile construction device, characterized in that, include: The first acquisition module is used to acquire homomorphically encrypted multi-source data; The second acquisition module is used to perform cleaning and preprocessing on the multi-source data to acquire the cleaned multi-source data. The first processing module is used to extract and fuse features from the cleaned multi-source data, determine the fused comprehensive features, and construct a broadband user profile based on the comprehensive features.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, Includes computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 7.