High-precision crowd portrait construction system based on multi-source data mining
Through a high-precision crowd portrait construction system based on multi-source data mining, the problems of insufficient data collection limitations and dynamic adjustment capabilities in the existing technology are solved, and accurate, meticulous and dynamically adaptive crowd portrait construction is achieved.
Patent Information
- Application Number
- CN202510208928.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing population portrait construction technology has data collection limitations, and it is impossible to obtain changes in users' behavior details and interest preferences in different scenarios, resulting in the portrait being too rough and lacking dynamic adjustment and business collaboration capabilities.
Design a high-precision crowd portrait construction system based on multi-source data mining. Through the multi-source data acquisition module, data is collected from Internet platforms, sensors and internal databases of the enterprise, combined with data preprocessing, data mining and analysis modules, a dynamic adaptive portrait model is built, and visually presented through the application display module.
It realizes accurate acquisition of users' multi-dimensional in-depth information, builds a detailed and comprehensive crowd portrait, and has dynamic adaptability, can reflect changes in crowd characteristics in real time, and meets business needs flexibility.
Smart Images

Figure CN120144829A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of crowd profiling, and in particular to a high-precision crowd profiling construction system based on multi-source data mining. Background Art
[0002] In the current digital business and data analysis fields, accurate crowd portraits are crucial for the decision-making of enterprises and institutions. However, the existing crowd portrait construction technology has many defects. On the one hand, most systems have limitations in data collection, relying on only a few data sources, and the collected data is mostly surface information and lacks depth. It is impossible to obtain the behavioral details and interest preference changes of users in different scenarios, resulting in the constructed crowd portraits being too rough and difficult to reflect the true picture of users. On the other hand, in the portrait construction and application links, existing technologies often lack dynamic adjustment and business collaboration capabilities. In the process of portrait construction, the label system and weight calculation method are solidified and cannot adapt to the dynamic changes of business scenarios.
[0003] In view of the above situation, we proposed a high-precision crowd portrait construction system based on multi-source data mining. This system deeply integrates various types of data; the dynamic adaptive mechanism of its portrait construction module can adjust the portrait at any time according to business changes, which can effectively solve the problems existing in the existing technology and provide more accurate and efficient crowd portrait construction and application solutions for various industries. Summary of the invention
[0004] In view of the deficiencies in the prior art, the present invention provides a high-precision population portrait construction system based on multi-source data mining, thereby solving the technical problems mentioned in the background technology.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0006] A high-precision population portrait construction system based on multi-source data mining, including a multi-source data acquisition module, a data preprocessing module, a data mining and analysis module, a portrait construction module and an application display module;
[0007] The multi-source data acquisition module collects data from the Internet platform, sensors and internal databases of enterprises to provide a rich data foundation for portrait construction. In the data collection of the Internet platform, the open platform API interface is used to obtain multi-dimensional information of users; the e-commerce platform uses its API to obtain purchase record data, and when obtaining purchase records, the detailed attributes of the product can be obtained by associating the product information table with the order number; the search engine obtains search-related data through the log analysis interface and web crawler technology;
[0008] The data preprocessing module cleans, removes noise, standardizes and integrates the raw data;
[0009] The data mining and analysis module extracts valuable information and features from the preprocessed data;
[0010] The portrait construction module constructs a population portrait based on the results of data mining and analysis;
[0011] The application display module visually presents the constructed population portrait to users and provides query and analysis functions.
[0012] In a possible implementation manner, when the multi-source data acquisition module acquires data from different data sources, it has respective adapted data acquisition methods and transmission processes for Internet platforms, sensors, and enterprise internal databases.
[0013] In a possible implementation manner, the data preprocessing module includes a data cleaning component, a data standardization component, and a data integration component. The data cleaning component uses a hash algorithm to remove duplicate data; for missing values in numerical data, the mean filling method is adopted, and the formula is where x i is the i-th known age value, and n is the number of known age values; for missing values in categorical data, the mode filling method is adopted; the data standardization component performs unified format conversion on the data. For numerical data, the Z-score standardization method is adopted, and the formula is where x is the original data value, μ is the mean of the data, σ is the standard deviation of the data, and x new is the standardized data value; for categorical data, unique encoding is adopted; the data integration component integrates data from different data sources.
[0014] In a possible implementation manner, the data mining and analysis module includes an association rule mining sub-module, a clustering analysis sub-module, and a classification algorithm sub-module. The association rule mining sub-module uses the Apriori algorithm and the FP-Growth algorithm to mine frequent item sets and association rules in the data; the clustering analysis sub-module uses the K-Means algorithm and the DBSCAN algorithm to cluster the data; the classification algorithm sub-module uses the decision tree algorithm and the random forest algorithm to construct a classification model.
[0015] In a possible implementation manner, the portrait construction module includes designing a label system, calculating label weights, and generating a portrait model. The label system design combines business requirements and data analysis to construct a comprehensive and detailed label system and hierarchy. The label weight calculation determines the label weights. For labels reflecting behavior frequency, the weight calculation formula is where count user is the number of times this user has performed actions on a specific commodity, and count totalIt is the total number of behaviors of all users towards the product, and the portrait model generates population portraits by applying rules and machine learning algorithms.
[0016] In a possible implementation manner, the application display module visually presents the constructed population portrait to users, and provides query and analysis functions. Through the visual interface design, bar charts, pie charts, heat maps, and node charts are used to display portrait information, and a query interface is designed to facilitate user queries and integrate with the business system.
[0017] Beneficial effects compared with the prior art:
[0018] 1. In this solution, the multi-source data collection module has a powerful data collection ability. It not only covers a wide range of data sources such as Internet platforms, sensors, and enterprise internal databases, but also far exceeds the conventional level in terms of the depth of data collection. For different platforms, by using the APIs of their open platforms, it can accurately obtain multi-dimensional in-depth information of users. This data fusion in terms of depth and breadth provides a unique data basis for constructing an extremely detailed and comprehensive population portrait.
[0019] 2. In this solution, in the portrait construction module, the label system design component can adjust the label structure in real time according to business requirements and data analysis, and the label weight calculation component can also flexibly determine the label weights according to data characteristics and business scenario changes. For example, during an e-commerce promotion event, the weight of the consumption behavior label will change dynamically according to the consumption data during the event, so that the population portrait generated by the portrait model generation component can reflect the dynamic changes of population characteristics in real time and accurately. This dynamic adaptive ability enables the population portrait to always maintain a high degree of accuracy and meet the continuously changing needs of the business. Brief Description of the Drawings
[0020] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement it according to the content of the specification, the following will describe in detail with reference to the preferred embodiments of the present invention and the accompanying drawings.
[0021] Figure 1 It is the system framework diagram of the high-precision population portrait construction based on multi-source data mining of the present invention. Detailed Embodiments
[0022] The preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. However, the present invention can be implemented in various different forms, so the present invention is not limited to the embodiments described below. In addition, in order to describe the present invention more clearly, components not connected to the invention will be omitted from the drawings;
[0023] The technical solutions in the embodiments of the present application are to solve the problems in the above background technology, and the general idea is as follows:
[0024] Example: This example introduces a high-precision population portrait construction system based on multi-source data mining, including a multi-source data collection module, a data preprocessing module, a data mining and analysis module, a portrait construction module, and an application display module;
[0025] I. Multi-source data collection module
[0026] This module is responsible for collecting data related to the target population from diverse data sources, which cover Internet platforms, sensors, and enterprise internal databases, etc., aiming to comprehensively obtain multi-dimensional information such as population characteristics, behaviors, and preferences, providing a rich data foundation for subsequent data processing and population portrait construction. For example, Internet platform data reflects the social behaviors and interests of the population; sensor data can obtain the real-time dynamics of the population in specific scenarios; enterprise internal data focuses on the behavioral data of customers within the scope of enterprise business. By integrating these multi-source data, a more three-dimensional and accurate population portrait can be constructed.
[0027] 1. Internet platform data collection
[0028] Social media platform: Taking Weibo as an example, through the API interface provided by the Weibo Open Platform, the basic information of users (such as nickname, gender, location), posted content (Weibo text, pictures, videos), following list, fan list, and interaction data (likes, comments, reposts), etc. can be obtained. When calling the API interface, the OAuth 2.0 authentication protocol needs to be followed to ensure the legality and security of data access. In terms of data collection frequency, it can be set according to business requirements. For example, for the user groups that are focused on, the latest posted Weibo content of each user can be collected once an hour; for ordinary users, their basic information and interaction data can be collected once a day.
[0029] E-commerce platform: Taking Taobao as an example, using the API of the Taobao Open Platform, the purchase records of users (purchased product names, prices, purchase times, purchase quantities), product favorite lists, browsing histories, and user evaluations, etc. can be obtained. When obtaining purchase records, the detailed attributes of products (such as brand, model, category) can be obtained by associating the product information table through the order number. To reduce the pressure on the platform server, the paging query method can be adopted, and a certain number of order data are obtained in each request. For example, 100 order records are obtained in each request.
[0030] Search Engine: Through the log analysis interface of the search engine (if available), the user's search keywords, search time, and click situation of search results can be obtained. For some search engines that do not provide a direct log analysis interface, web crawler technology can be used to simulate user search behavior and record search keywords and corresponding search result pages. However, it should be noted to comply with the search engine's robots protocol to avoid IP bans caused by excessive crawling. For example, the access frequency of the crawler can be set to send a search request every 10 seconds.
[0031] 2. Sensor Data Acquisition
[0032] Traffic Sensors: In an intelligent transportation system, devices such as geomagnetic sensors and cameras deployed on roads can collect data on vehicle driving speed, traffic flow, vehicle types, as well as pedestrian traffic flow and walking trajectories. The geomagnetic sensor senses the magnetic field changes generated when a vehicle passes by and transmits the signal to the data acquisition terminal. The acquisition terminal summarizes the data at regular time intervals (such as every 5 minutes) and uploads it to the data center. Cameras, on the other hand, use image recognition technology to identify vehicle and pedestrian information and transmit the recognition results to the data center after encoding.
[0033] Passenger Flow Sensors: In commercial places such as shopping malls and supermarkets, infrared sensors or Wi-Fi probes are deployed to collect passenger flow data. Infrared sensors count the passenger flow by detecting changes in the infrared signals of the human body, while Wi-Fi probes detect the Wi-Fi signals of devices such as mobile phones, analyze the frequency and time of the appearance of device MAC addresses, and thus count the passenger flow and customer stay time. The data acquisition device transmits the collected data to the data aggregation node through a wireless communication module (such as Bluetooth, ZigBee, or 4G), and then the aggregation node uploads it to the data center.
[0034] 3. Enterprise Internal Data Acquisition
[0035] Customer Relationship Management System (CRM): Extract data such as the basic information of customers (name, contact information, age, occupation, income, etc.), sales records (sales order amount, sales time, sales products), and customer service records (consultation content, complaint content, handling results) from the enterprise's CRM system. Through ETL (Extract, Transform, Load) tools, the data in the CRM system is extracted to the data warehouse at a predetermined time period (such as every early morning). During the extraction process, the data can be cleaned and transformed, such as unifying the date format and removing duplicate customer records.
[0036] Enterprise Business Database: For manufacturing enterprises, their business databases store product production data (production batches, output, pass rates), raw material procurement data (procurement suppliers, procurement quantities, procurement prices), etc. By writing SQL query statements, relevant data related to the target population can be extracted from the business database. For example, if one wants to analyze the customer characteristics of those who purchase a certain type of product, the sales record table and the customer information table can be associated through the product ID to extract relevant customer information.
[0037] II. Data Preprocessing Module
[0038] This module processes the collected raw data through cleaning, denoising, standardization, and integration to improve data quality, making it meet the requirements of subsequent data mining and analysis, and laying a solid data foundation for constructing an accurate population portrait. It mainly consists of a data cleaning component, a data standardization component, and a data integration component. The data cleaning component is responsible for identifying and processing error values, missing values, and duplicate values in the data; the data standardization component performs unified format conversion on data of different magnitudes and types; the data integration component integrates data from different data sources. High-quality data is the key to constructing an accurate population portrait.
[0039] 1. Data Cleaning Component
[0040] Removing Duplicate Data: The hash algorithm is used to process data records, generating a unique hash value for each record. The generated hash values are stored in a hash table. When processing new data records, calculate their hash values and compare them with the values in the hash table. If the hash values are the same, further compare the detailed content of the records. After confirming duplication, delete them. For example, for the user purchase records on an e-commerce platform, there may be duplicate order records due to network latency and other reasons. These duplicate records can be quickly identified and deleted through the hash algorithm.
[0041] Handling Missing Values:
[0042] Numeric Data: For missing values in numeric data, the mean filling method is adopted. Suppose there is a set of user age data with some missing values. First, calculate the mean of the existing age values in this set of data where x i is the i-th known age value and n is the number of known age values. Then fill the missing age values with this mean.
[0043] Categorical Data: For missing values in categorical data, the mode filling method is adopted. For example, in the user occupation data, if there are missing values, count the category that appears most frequently among the existing occupation categories and use it as the filling value for the missing values.
[0044] Correcting Error Data:
[0045] Data type error: By writing a data type check script, traverse each data field in the dataset to check whether its data type meets the expectations. For example, if the user's age field is incorrectly recorded as a string type, it can be converted to a numerical type through a data type conversion function.
[0046] Data range error: Set a reasonable range for the data. For example, the user's age should be between 0 and 120 years old. For data outside the range, correct it through manual review or according to the relevance of the data. For example, if a user's age is recorded as 200 years old, the true age can be inferred and corrected by querying other relevant information of the user (such as registration time, purchase behavior, etc.).
[0047] 2. Data normalization component
[0048] Normalization of numerical data: Adopt the Z-score normalization method, and the formula is where x is the original data value, μ is the mean of the data, σ is the standard deviation of the data, and x new is the normalized data value. For example, for the user's income data, the income levels of different users vary greatly. Through Z-score normalization, it can be converted into data with the same magnitude, which is convenient for subsequent data analysis and model training.
[0049] Encoding of categorical data: For categorical data, adopt the one-hot encoding method. Taking the user's gender as an example, it is encoded into two dimensions, [1,0] represents male, and [0,1] represents female. For categorical data with multiple categories, such as user occupations (teachers, doctors, engineers, etc.), assuming there are n categories, then n dimensions will be generated after encoding, each dimension corresponds to a category, and only the dimension value corresponding to this category is 1, and the other dimension values are 0.
[0050] 3. Data integration component
[0051] Unify data format: Use a data conversion tool (such as Apache NiFi) to write corresponding conversion rules according to the data formats of different data sources. For example, convert data in CSV format to Parquet format. Parquet format has better compression performance and query efficiency and is suitable for large-scale data storage and processing. For JSON format data, by parsing the JSON structure, convert it into a table form suitable for subsequent processing.
[0052] Data Association and Integration: By establishing association relationships between data, such as user IDs, device IDs, etc., data from different data sources is integrated. For example, the user purchase records on an e-commerce platform and the user interest data on a social media platform are associated through the user ID and integrated into a single dataset. During the integration process, data inconsistency issues may be encountered. For example, the user ID in the e-commerce platform is of integer type, while the user ID in the social media platform is of string type, and data type conversion and unification are required.
[0053] III. Data Mining and Analysis Module
[0054] Extract valuable information and features from the preprocessed data, such as discovering association rules in the data, performing clustering analysis to divide different population groups, using classification algorithms to predict the behaviors and attributes of the population, etc., to provide data support for constructing a population portrait. This module includes an association rule mining sub-module, a clustering analysis sub-module, and a classification algorithm sub-module. The association rule mining sub-module uses algorithms such as Apriori algorithm and FP-Growth algorithm to mine frequent item sets and association rules in the data; the clustering analysis sub-module uses algorithms such as K-Means algorithm and DBSCAN algorithm to cluster the data; the classification algorithm sub-module uses algorithms such as decision tree algorithm and random forest algorithm to build classification models.
[0055] 1. Association Rule Mining Sub-module
[0056] Apriori Algorithm:
[0057] Frequent Item Set Generation: Taking the product purchase data on an e-commerce platform as an example, first set a support threshold, such as 0.01, indicating that at least 1% of the transactions contain a certain item set. Scan the dataset, count the occurrence times of each individual product, and filter out the products that meet the support threshold to form frequent 1-item sets. Then combine the frequent 1-item sets to generate candidate 2-item sets, scan the dataset again, count the support of the candidate 2-item sets, and filter out the 2-item sets that meet the support threshold to form frequent 2-item sets. And so on, continuously generating higher-order frequent item sets. The support calculation formula is where X is the item set, σ(X) represents the number of transactions containing the item set X, and N is the total number of transactions.
[0058] Association Rule Generation: Generate association rules from the frequent item sets. Set a confidence threshold, such as 0.8, indicating that in the transactions containing the antecedent, at least 80% of the transactions also contain the consequent. For the frequent item set X∪Y, generate the association rule X→Y, and the confidence calculation formula is For example, if the support of the frequent item set {milk, bread} is 0.05 and the support of {milk} is 0.1, then the confidence of the association rule {milk}→{bread} is 0.5.
[0059] FP-Growth Algorithm:
[0060] Construct the FP-Tree: First, scan the dataset once to count the support of each item, and filter out the frequent items according to the support threshold. Then scan the dataset again, sort the frequent items in each transaction in descending order of support, and insert them into the FP-Tree. During the insertion process, if the node already exists, increment its count; if not, create a new node. For example, for the transaction {milk, bread, egg}, if milk, bread, and egg are all frequent items, and the support of milk is the highest, bread is the second, and egg is the lowest, then insert the milk node first, then insert the bread node under the milk node, and finally insert the egg node under the bread node.
[0061] Mine frequent item sets: Start from the leaf nodes of the FP-Tree and generate frequent item sets by backtracking. For example, starting from a certain leaf node, backtrack up to the root node, and the nodes on the path form a frequent item set. By continuously traversing the leaf nodes, all frequent item sets can be mined.
[0062] 2. Clustering Analysis Sub-module
[0063] K-Means Algorithm:
[0064] Initialize the cluster centers: Randomly select K data points as the initial cluster centers. For example, for the user consumption behavior data (including dimensions such as consumption amount and consumption frequency), if the users are to be divided into 5 categories (K = 5), then randomly select the consumption behavior data of 5 users from the dataset as the initial cluster centers.
[0065] Assign data points to clusters: Calculate the distance from each data point to each cluster center using the Euclidean distance formula where x and y are two data points, x i and y i are the i-th attribute values of data points x and y respectively, and n is the number of attributes. Assign each data point to the cluster where the nearest cluster center is located.
[0066] Update the cluster centers: Recalculate the mean of the data points within each cluster as the new cluster center. For example, for all the user consumption behavior data within a certain cluster, calculate the mean of the consumption amount and consumption frequency, and use this mean as the new cluster center. Repeat the above steps until the cluster centers no longer change or the maximum number of iterations is reached.
[0067] DBSCAN Algorithm:
[0068] Define core points, border points, and noise points: Set the neighborhood radius and the minimum number of points MinPts. For example, for the user location data, if Set it to 100 meters and MinPts to 5, which means that within a circular area with a certain user location as the center and a radius of 100 meters, if there are at least 5 users, then this user is a core point. Among the points in the neighborhood of the core point, those with the number of users in their own neighborhood less than MinPts are border points, and the others are noise points.
[0069] Clustering process: Starting from a core point, continuously expand the points in its neighborhood to form a cluster. If the neighborhood of a certain core point contains other core points, then merge the neighborhoods of these core points and continue to expand until it can no longer be expanded. By continuously traversing the points in the dataset, all clusters and noise points are identified.
[0070] 3. Classification algorithm sub-module
[0071] Decision tree algorithm:
[0072] Construct a decision tree: Taking the prediction of whether a user will purchase a certain product as an example, construct a decision tree based on attributes such as the user's age, income, and consumption habits. Select the attribute with the largest information gain as the splitting attribute of the root node. The information gain calculation formula is IG(D,A) = H(D) - H(D|A), where D is the dataset, A is the attribute, H(D) is the information entropy of the dataset D, and H(D|A) is the conditional entropy after partitioning the dataset D on the attribute A. The information entropy calculation formula is where p i is the probability that belongs to the i-th class in the dataset. For example, if the information gain of the age attribute is the largest, then use age as the splitting attribute of the root node, and partition the dataset into different subsets according to different value ranges of age.
[0073] Generate decision rules: Each path from the root node to the leaf node of the decision tree forms a decision rule. For example, if a path is age > 30 years old and income > 5000 yuan, then the decision rule corresponding to this path is: Users with age > 30 years old and income > 5000 yuan tend to purchase this product.
[0074] Random forest algorithm:
[0075] Construct a set of decision trees: Randomly draw multiple sample subsets from the original dataset with replacement. For each sample subset, randomly select a part of the attributes to construct a decision tree. For example, from the original dataset containing 10,000 user data, draw 100 times with replacement, and each time draw 8,000 user data as a sample subset. For each sample subset, randomly select 6 attributes from 10 user attributes to construct a decision tree.
[0076] Prediction and integration: For classification problems, the voting method is adopted, that is, multiple decision trees perform classification predictions on a data point, and the category with the most votes is used as the final prediction result; for regression problems, the averaging method is adopted, and the prediction results of multiple decision trees are averaged to obtain the final predicted value.
[0077] IV. Portrait Construction Module
[0078] This module constructs a population portrait based on the data mining and analysis results, including designing a label system, calculating label weights, and generating a portrait model, transforming the data into an intuitive and understandable description of population characteristics, and providing support for decision-making in various fields. It consists of a label system design component, a label weight calculation component, a portrait model generation component, and a portrait storage and update component. The label system design component defines and manages the label structure of the population portrait; the label weight calculation component determines the importance of each label according to data characteristics and business requirements; the portrait model generation component generates a population portrait using rules or machine learning algorithms.
[0079] 1. Label system design: Combine business requirements and data analysis to construct a comprehensive and detailed label system. In the field of marketing, demographic labels are subdivided into age ranges (such as 18 - 25 years old, 26 - 35 years old, etc.), gender, marital status, education level, etc.; interest labels cover entertainment interests (movies, music, games), life interests (food, fitness, travel), cultural interests (reading, art exhibitions), etc.; consumer behavior labels include consumption ability levels (low, medium, high), consumption frequencies (daily, weekly, monthly, etc.), consumption preference types (brand preference, category preference), etc. Design a reasonable label hierarchy, such as interest labels are divided into first-level labels (entertainment, life, culture), second-level labels (movies, fitness, reading), and even third-level labels (action movies, aerobic exercise, science fiction novels), which is convenient for management and use, and facilitates different granularity screening and combination during portrait analysis.
[0080] 2. Label weight calculation: For labels reflecting behavior frequencies, such as the number of views or purchases of a certain type of product by e-commerce platform users, the weight calculation formula is count user is the number of times this user has performed an action on a specific product, and count total is the total number of actions on this product by all users. For example, if a user has purchased a certain brand of sports shoes 5 times and all users have purchased 1000 times in total, the weight of the user's "purchased a certain brand of sports shoes" label is When evaluating the credit risk of users in the field of financial risk control, the analytic hierarchy process (AHP) is used to determine the weights. Invite financial experts to pairwise compare the relative importance of different labels (credit record, income stability, debt situation) to construct a judgment matrix. For example, if an expert believes that the credit record is slightly more important than the income stability, the value at the corresponding position in the judgment matrix is 3, and the value at the position of the inverse matrix is The eigenvector of the judgment matrix is calculated to obtain the weight of each label. Assume that the weight of credit record is 0.5, the weight of income stability is 0.3, and the weight of debt situation is 0.2.
[0081] 3. Portrait model generation: formulate clear rules to generate crowd portraits. If the user is between 22 and 30 years old, has a monthly income of more than 8,000 yuan, follows more than 10 fashion brand accounts on social media, and has purchased fashion clothing more than 3 times in the past three months, he / she is marked as a "young fashion high-spending group". The rules can be optimized and adjusted based on business experience and data analysis. Generate portraits using cluster analysis or classification algorithm results. Taking cluster analysis as an example, use the K-Means algorithm to divide users into 5 clusters and analyze the characteristics of users in each cluster. If the average age of users in a cluster is 35 years old, the main occupation is middle-level management personnel in the enterprise, the average monthly consumption is 5,000-8,000 yuan, and they are interested in electronic products and travel, the cluster is defined as a "middle-aged workplace elite consumer group" and given corresponding labels and feature descriptions.
[0082] 5. Application display module
[0083] The constructed crowd portraits are displayed intuitively to users, providing convenient query and analysis functions to help users make decisions based on the portrait information.
[0084] 1. Visual interface design: Use a variety of charts to display crowd portraits. Use a bar chart to display the distribution of the number of people in different age groups, with the horizontal axis being the age range and the vertical axis being the number of people; use a pie chart to display the proportion of people with different interests and hobbies, with each sector area representing an interest and hobby category and its proportion; use a heat map to display the consumption popularity of people in different regions, the darker the color, the higher the popularity. For user social relationships, use a node graph to display, with nodes representing users and edges representing relationships (follow, friends, etc.). Design a simple and easy-to-use query interface, where users can enter keywords (tag names, user attribute values), select query conditions (age range, consumption amount range), and other queries. For example, if you enter "users aged between 25 and 35 years old, and whose consumption amount is greater than 5,000 yuan", the query engine will search the portrait database according to the conditions and display user portrait information that meets the conditions.
[0085] 2. Integration with business systems: Integrate with the marketing automation system to formulate personalized marketing strategies based on the population portrait. For the portrait of the "young, fashionable and high-consumption population", the marketing system automatically pushes new product advertisements of fashion brands, invitations to high-end fashion events, etc. The population portrait data is transmitted to the marketing automation system through an interface. The marketing system screens the target customer group according to the portrait tags, formulates a marketing activity plan and tracks the effects. In the risk assessment system of financial institutions, the population portrait data serves as an important basis for evaluating users' credit risk and fraud risk. If the user portrait shows tags such as frequent job changes, high debt, and bad credit records, the risk assessment system raises their risk level and takes corresponding risk control measures in business processes such as loan approval and credit card issuance, such as reducing the loan amount, raising the loan interest rate or rejecting the application, and realizes data sharing and real-time interaction between the population portrait system and the financial risk control system through a data interface.
[0086] Finally, it should be noted that: Obviously, the above embodiments are merely examples for clearly illustrating the present invention and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to exhaustively list all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.
Claims
1. A high-precision crowd portrait construction system based on multi-source data mining, characterized in that: It includes multi-source data acquisition module, data preprocessing module, data mining and analysis module, image building module and application display module; The multi-source data acquisition module collects data from the Internet platform, sensors and internal databases of enterprises to provide a rich data foundation for portrait construction. In the data collection of the Internet platform, the open platform API interface is used to obtain multi-dimensional information of users; the e-commerce platform uses its API to obtain purchase record data, and when obtaining purchase records, the detailed attributes of the product can be obtained by associating the product information table with the order number; the search engine obtains search-related data through the log analysis interface and web crawler technology; The data preprocessing module cleans, removes noise, standardizes and integrates the raw data; The data mining and analysis module extracts valuable information and features from the preprocessed data; The portrait building module builds a crowd portrait based on the data mining and analysis results; The application display module intuitively presents the constructed crowd portrait to the user and provides query and analysis functions.
2. A high-precision crowd portrait construction system based on multi-source data mining as claimed in claim 1, characterized in that: When the multi-source data acquisition module collects data from different data sources, it has its own adapted data acquisition methods and transmission processes for the Internet platform, sensors and internal databases of enterprises.
3. A high-precision crowd portrait construction system based on multi-source data mining as claimed in claim 1, characterized in that: The data preprocessing module includes a data cleaning component, a data standardization component and a data integration component. The data cleaning component uses a hash algorithm to remove duplicate data; the missing values of numerical data are filled with the mean value, and the formula is where x i is the i-th known age value, and n is the number of known age values; the missing values of categorical data are filled by the mode method; the data standardization component converts the data into a unified format, and the numerical data is standardized by the Z-score method, and the formula is Where x is the original data value, μ is the mean of the data, σ is the standard deviation of the data, and x new is the standardized data value; The categorized data is uniquely coded; the data integration component integrates data from different data sources.
4. A high-precision crowd portrait construction system based on multi-source data mining as claimed in claim 1, characterized in that: The data mining and analysis module includes an association rule mining submodule, a cluster analysis submodule and a classification algorithm submodule. The association rule mining submodule uses the Apriori algorithm and the FP-Growth algorithm to mine frequent item sets and association rules in the data; the cluster analysis submodule uses the K-Means algorithm and the DBSCAN algorithm to cluster the data; and the classification algorithm submodule uses the decision tree algorithm and the random forest algorithm to build a classification model.
5. A high-precision crowd portrait construction system based on multi-source data mining as claimed in claim 1, characterized in that: The portrait construction module includes designing a label system, calculating label weights, and generating a portrait model. The label system design combines business needs and data analysis to build a comprehensive and detailed label system and hierarchy. The label weight calculation determines the label weight. For labels that reflect behavior frequency, the weight calculation formula is: where count user is the number of times the user has acted on a specific product, count total It is the total number of behaviors of all users on the product. The portrait model generates crowd portraits using rules and machine learning algorithms.
6. A high-precision crowd portrait construction system based on multi-source data mining as claimed in claim 1, characterized in that: The application display module intuitively presents the constructed population portrait to the user and provides query and analysis functions. Through visual interface design, it uses bar charts, pie charts, heat maps, and node maps to display portrait information. The query interface is designed to facilitate user queries and is integrated with the business system.
Citation Information
Cited By
Matrix type data layering intelligent management system
CN121070987A