Method and apparatus for information push
By combining data from multiple sources to align user identifiers and analyze movement trajectories, customer profiles are constructed, solving the problems of limited data dimensions and insufficient real-time data in the mining of potential new customers in business districts. This results in more accurate user profiles and information pushes, thereby improving marketing effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JINGDONG CITY BEIJING DIGITS TECH CO LTD
- Filing Date
- 2022-04-12
- Publication Date
- 2026-07-21
AI Technical Summary
When identifying potential new customers, business districts often suffer from issues such as limited data dimensions, lack of real-time customer profiling, and inaccurate matching of target new customers, leading to reduced marketing effectiveness.
By combining data from multiple sources, user identifier alignment, movement trajectory analysis, and spatial clustering are performed to construct customer profiles. Similarity is extracted using social network features to identify target users and push information.
It enables more accurate and real-time user profiling, improves the success rate of information delivery, and expands the marketing reach of the business district.
Smart Images

Figure CN114741595B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method and apparatus for information push. Background Technology
[0002] In current business district marketing scenarios (e.g., advertising, product recommendations), marketing primarily focuses on two aspects: repeat purchases from existing customers and activating potential new customers. Currently, potential new customer acquisition often relies on long-term and short-term user behavior data from the local business district for audience probing.
[0003] In the process of realizing this invention, the inventors discovered at least the following problems in the prior art:
[0004] Business district operators rely solely on long-term and short-term user behavior data from the local area for demographic exploration, lacking detailed understanding of external customer traffic and internal shop spending patterns. Furthermore, due to concerns about commercial interests and personal data privacy, business districts cannot effectively utilize data from external organizations, resulting in a single data dimension and hindering effective marketing. Customer consumption characteristics in customer profiles are influenced by market sentiment, with significant variations in characteristic values across different time periods. Existing business district insight methods do not consider the real-time nature of user data, leading to inaccurate customer profiles and failing to meet the real-time dynamic query requirements of business analysis. When matching target new customers based on existing visitors, current technologies often rely on expanding similar groups based on single user characteristics, resulting in inaccurate target new customers and impacting marketing effectiveness.
[0005] In summary, current practices in identifying potential new customers in business districts suffer from limitations such as limited data dimensions, lack of real-time customer profiling, and inaccurate matching of target customers, thus reducing the effectiveness of marketing efforts. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a method and apparatus for information push, which can aggregate data from multiple sources to characterize user profiles and push information, making up for the shortcomings of the single dimension of user data features in business district operation scenarios. It can better, more in real time and more accurately characterize user groups, more accurately identify potential customers in the business district, and facilitate improving the success rate of information push.
[0007] To achieve the above objectives, according to one aspect of the present invention, an information push method is provided.
[0008] A method for pushing information includes:
[0009] Obtain the set of user identifiers to be aligned, and perform identifier alignment processing between the set of user identifiers to be aligned and the set of user identifiers in the business district to obtain aligned user identifiers;
[0010] The movement trajectory data of the aligned user identifiers is obtained based on the business data, and the dense source areas of the visiting users are determined based on the movement trajectory data;
[0011] Based on the characteristic data of visiting users, calculate the customer profile of visiting users from each dense source area;
[0012] The customer profile is compared with the historical user profile of the business district to identify target users, and information is pushed to the target users.
[0013] Optionally, the identifier alignment process includes:
[0014] Processing can be done by taking the intersection, taking the union, or performing a left outer join or a right outer join.
[0015] Optionally, the process of aligning the user identifier set to be aligned with the business district user identifier set includes:
[0016] The set of user identifiers to be aligned is sent to a third-party arbitration institution so that the third-party arbitration institution can perform identifier alignment processing based on the set of user identifiers to be aligned and the set of user identifiers in the business district.
[0017] Optionally, before performing identifier alignment processing between the user identifier set to be aligned and the business district user identifier set, the method further includes:
[0018] The processing rules for obtaining user identifiers pre-agreed with the business district include the data format and encryption method of the user identifiers;
[0019] The processing rules are used to process the set of user identifiers to be aligned, and the processed set of user identifiers to be aligned is used as the set of user identifiers to be aligned.
[0020] Optionally, determining the dense source areas of visiting users based on the movement trajectory data includes:
[0021] The residence and workplace of each visiting user are determined based on the aforementioned movement trajectory data;
[0022] A spatial clustering algorithm was used to cluster residential and workplace locations, resulting in multiple clustered regions.
[0023] From the multiple cluster regions, select a specified number of cluster regions where the visiting users are most densely distributed in space as the dense source of visiting users.
[0024] Optionally, comparing the customer profile with the historical user profiles of the business district to determine target users includes:
[0025] The structured profile data in the customer profile and the historical user profile of the business district are categorized and encoded, and the profile representation vector is obtained by dimensionality reduction.
[0026] Social network representations are extracted from the unstructured profile data in the customer profile and the historical user profile of the business district, respectively, to obtain network representation vectors;
[0027] By concatenating the profile representation vector and the network representation vector together, feature vectors for the customer profile and the historical user profile of the business district are obtained respectively.
[0028] Based on the similarity between feature vectors, the similarity between the customer profile and the historical user profile of the business district is compared, and the visiting users corresponding to the customer profile with a similarity that meets a preset threshold are identified as target users.
[0029] Optionally, pushing information to the target user includes:
[0030] The first dense source location of the target user is obtained and returned to the business district so that the business district can push information based on the first dense source location.
[0031] According to another aspect of the present invention, an information push device is provided.
[0032] An information push device, comprising:
[0033] The user identifier alignment module is used to obtain a set of user identifiers to be aligned and perform identifier alignment processing between the set of user identifiers to be aligned and the set of user identifiers in the business district to obtain aligned user identifiers.
[0034] The user origin determination module is used to obtain the movement trajectory data of the aligned user identifier based on business data, and determine the dense origin of the visiting user based on the movement trajectory data;
[0035] The customer profile calculation module is used to calculate the customer profile of visitors from each dense source area based on the characteristic data of the visitors.
[0036] The target user identification module is used to compare the similarity between the customer profile and the historical user profile of the business district to identify target users and push information to the target users.
[0037] According to another aspect of the present invention, an electronic device for information push is provided.
[0038] An electronic device for information push includes: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement the information push method provided in the embodiments of the present invention.
[0039] According to another aspect of the present invention, a computer-readable medium is provided.
[0040] A computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the information push method provided in the embodiments of the present invention.
[0041] One embodiment of the above invention has the following advantages or beneficial effects: by acquiring a set of user identifiers to be aligned and performing identifier alignment processing with a set of user identifiers in the business district to obtain aligned user identifiers; by acquiring the movement trajectory data of the aligned user identifiers based on business data and determining the dense source areas of visiting users based on the movement trajectory data; by calculating the customer profile of visiting users in each dense source area based on the feature data of visiting users; by comparing the similarity between the customer profile and the historical user profile of the business district to determine the target users and pushing information to the target users, it is possible to aggregate data from multiple sources to characterize user profiles and push information, making up for the deficiency of the single dimension of user data features in the business district operation scenario, realizing the full mining of the source areas of users who have not visited the business district, thereby better pushing information such as advertisements in the scenario of attracting new customers in the business district; based on the spatiotemporal attributes of user trajectory information, the precise dense source areas of users are obtained, thereby enabling better, more real-time and more accurate characterization of user customer profiles, more accurate determination of potential customers in the business district, and facilitating the improvement of the success rate of information push.
[0042] The further effects of the aforementioned unconventional alternative methods will be explained below in conjunction with specific implementation methods. Attached Figure Description
[0043] The accompanying drawings are provided to better understand the invention and are not intended to unduly limit the scope of the invention. Wherein:
[0044] Figure 1 This is a schematic diagram illustrating the main steps of an information push method according to an embodiment of the present invention;
[0045] Figure 2 This is a schematic diagram of the implementation process of an embodiment of the present invention;
[0046] Figure 3 This is a schematic diagram illustrating the customer profile construction and new customer acquisition process according to an embodiment of the present invention;
[0047] Figure 4This is a schematic diagram of the main modules of an information push device according to an embodiment of the present invention;
[0048] Figure 5 This is an exemplary system architecture diagram in which embodiments of the present invention can be applied;
[0049] Figure 6 This is a schematic diagram of the structure of a computer system suitable for implementing terminal devices or servers of the present invention. Detailed Implementation
[0050] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0051] In the current marketing landscape of commercial districts, marketing primarily focuses on two aspects: repeat purchases from existing customers and activating potential new customers. However, local data suffers from a severe lack of data and feature dimensions, and is consistently limited to data generated within the current commercial district. For example, commercial district operators may want to understand the origin clusters and demographic profiles of potential customers who haven't visited the district before, and conduct targeted analysis and traffic generation from these origins to attract new customers to various business sectors. To fully expand the marketing reach of a commercial district, it's necessary to thoroughly explore the demographic characteristics of customers outside the district, analyze various data features and travel patterns for comprehensive comparison, identify the focal points of existing customer traffic, and find potential customer sources, providing data support for the new customer insight section of the business insights platform. However, commercial district operators lack access to demographic and consumption data of customers outside the district, and even specific consumption data from shops within the district is difficult to obtain due to trade secrets. This creates significant obstacles to attracting and marketing to target new customers outside the district. If commercial districts could access online customer traffic data, they could analyze target new customers to attract new customers to various business sectors, thereby increasing revenue.
[0052] The main shortcomings of current business districts in acquiring new customers are as follows:
[0053] 1. Limited Feature Dimensions: The data of the business district operator is only based on long-term and short-term user behavior data in the business district for audience exploration. There is a lack of detailed understanding of the situation of the customer flow outside the business district and the consumption situation of the stores within the business district. Furthermore, due to the protection of commercial interests and the security of personal data privacy, the business district cannot effectively use data from other external institutions. As a result, the data dimensions are limited and cannot be effectively used for marketing.
[0054] 2. Customer profiling lacks real-time accuracy: Consumer spending characteristics in customer profiles are influenced by market sentiment, and feature values often vary significantly over time. Existing business district insight methods do not consider the real-time nature of user data, resulting in inaccurate customer profiles that cannot meet the real-time dynamic query requirements of business analysis.
[0055] 3. Inaccurate matching of target new customers: When matching target new customers based on existing visitors, existing technologies only use the TF-IDF method to expand similar groups based on a single user characteristic attribute. The target new customers obtained are often inaccurate, thus affecting the marketing effect.
[0056] To overcome the deficiencies in the prior art, the present invention mainly employs the following technical means:
[0057] 1. Solve the problem of insufficient data dimensions for business district operators. By combining data from one or more external partners, while protecting the privacy of personal user data and preventing the leakage of business secrets, it can securely realize cross-domain data statistics and queries from multiple parties, supplement the data characteristics of users inside and outside the business district, discover the gathering places of target new customers, and achieve the goal of assisting commercial marketing to expand the radiation area of the business district through targeted traffic attraction methods such as advertising, traffic shuttle buses, and business format optimization.
[0058] 2. Based on the spatiotemporal attributes of user trajectory information, spatial clustering is performed through different long-term dwell time periods of users to obtain the precise origin of users; combined with the dynamic features among the user's multi-faceted characteristics, customer profiles of visiting users and target new customers are constructed according to short-term data in hours and minutes and long-term data in days and weeks, reflecting customer characteristics in real time, and recommending target new customers in real time based on visiting users; the business district collaborates with social platforms, using social graph networks, and through LINE's graph embedding method, extracts the representations reflecting first-order and second-order similarity of user social graphs as features to improve the matching accuracy of target new customers.
[0059] In the description of the embodiments of the present invention, the technical terms involved and their definitions are as follows.
[0060] 1. OPTICS Algorithm: The full name of the OPTICS algorithm is Ordering points to identify the clustering structure. Its goal is to cluster spatial data according to density distribution. Its idea is very similar to DBSCAN, but DBSCAN is more difficult to select input parameters, meaning it is quite sensitive to input parameters. Unlike DBSCAN, OPTICS allows the density-based clustering structure to present a special order. This order corresponds to a clustering structure that contains information about each level of clustering and is easy to analyze. Slight changes in neighborhood radii do not affect the clustering results. The OPTICS algorithm has many advantages:
[0061] 1) It does not require specifying the number of clusters in advance and can discover clusters of any shape;
[0062] 2) It is not sensitive to outliers and can automatically identify outliers during the clustering process;
[0063] 3) The clustering results do not depend on the order in which the nodes are traversed;
[0064] 4) It is not sensitive to the neighborhood radius parameter, and the clustering results are more stable.
[0065] 2. TF-IDF: TF-IDF (Term Frequency-Inverse Document Frequency) is a statistical analysis method for keywords used to evaluate the importance of a word to a document set or corpus. The importance of a word is directly proportional to the frequency of its occurrence in an article and inversely proportional to its frequency in the corpus. This calculation method effectively avoids the influence of common words on keywords, improving the relevance between keywords and articles. TF refers to the total number of times a word appears in an article. This indicator is usually normalized and defined as TF = (number of times the word appears in a document / total number of words in the document), which prevents the results from being biased towards excessively long documents (the same word usually has a higher frequency in long documents than in short documents). IDF is the inverse document frequency; the fewer documents containing a word, the larger the IDF value, indicating that the word has strong discriminative power. IDF = loge(total number of documents in the corpus / number of documents containing the word + 1), with the +1 to avoid a denominator of 0. TF-IDF = TF x IDF; the larger the TF-IDF value, the greater the importance of the feature word to the text.
[0066] 3. Cosine Similarity: Cosine similarity assesses the similarity between two vectors by calculating the cosine of the angle between them. The vectors are plotted in vector space according to their coordinates, the angle between them is calculated, and the corresponding cosine value is used to characterize the similarity between the two vectors. The smaller the angle, the closer the cosine value is to 1, and the more similar their directions are, the more similar they are.
[0067] 4. LINE: LINE (Larg-scale Information Network Embedding), proposed by Jian Tang et al. in 2015, is a node embedding algorithm applicable to large networks with arbitrary edge types. It achieves network embedding by considering first-order proximity (local structure) and second-order proximity (global structure). Compared to previous Graph / Network Embedding methods, LINE has the following advantages: 1. It is applicable to any type of network. "Arbitrary type" in this paper mainly refers to the arbitrary weights and directions of edges: directed, undirected, weighted, and unweighted. (LINE does not consider heterogeneous networks with different node and edge types, which has certain limitations. This led to later research on heterogeneous and homogeneous networks); 2. LINE proposes an edge-sampling algorithm to improve and optimize the objective function, thus overcoming the limitations of traditional stochastic gradient descent.
[0068] 5. Homomorphic Encryption: Homomorphic encryption (HE) is a type of encryption method with special natural properties. This concept was first proposed by Rivest et al. in the 1970s. Compared with general encryption algorithms, homomorphic encryption can not only perform basic encryption operations, but also perform various computational functions between ciphertexts. That is, computation before decryption is equivalent to decryption before computation.
[0069] Figure 1 This is a schematic diagram illustrating the main steps of an information push method according to an embodiment of the present invention. In this embodiment, the entity executing the information push method is a partner of the business district. This partner can be a mobile app, an online shopping website, an offline shopping platform, an online or offline tourism product sales platform, etc., mainly depending on the business needs of the business district and the actual information that the partner can provide. Figure 1As shown, the information push method of this embodiment mainly includes the following steps S101 to S104.
[0070] Step S101: Obtain the set of user identifiers to be aligned, and perform identifier alignment processing between the set of user identifiers to be aligned and the set of user identifiers in the business district to obtain aligned user identifiers;
[0071] Step S102: Obtain the movement trajectory data of the aligned user identifiers based on the business data, and determine the dense source areas of the visiting users based on the movement trajectory data;
[0072] Step S103: Based on the characteristic data of visiting users, calculate the customer profile of visiting users in each dense source area;
[0073] Step S104: Compare the similarity between the customer profile and the historical user profile of the business district to determine the target users, and push information to the target users.
[0074] According to one embodiment of the present invention, the identifier alignment process includes: intersection processing, union processing, left outer join processing, or right outer join processing. The identifier (ID) alignment method is not fixed, and multiple methods can be used for security considerations, but the core is to find the intersection. If actual business needs require, methods such as union, left outer join, or right outer join can also be used.
[0075] According to another embodiment of the present invention, the identification alignment process between the user identification set to be aligned and the business district user identification set includes: sending the user identification set to be aligned to a third-party arbitration institution, so that the third-party arbitration institution can perform identification alignment processing based on the user identification set to be aligned and the business district user identification set. The third-party arbitration institution is an independent, neutral, and secure third-party node. However, in specific implementations, besides aggregating user identifications to a third-party arbitration institution for identification alignment, if the merchant and the business district mutually trust each other, the partner can directly send the user identification set to be aligned to the business district operator. After the business district completes the identification alignment, it then sends it to the partner, or vice versa, depending on which method meets the business and security compliance requirements. That is, the third-party arbitration institution can be a neutral, secure third-party node different from the merchant and the business district operator, or it can be either the merchant or the business district operator that meets the business and security compliance requirements.
[0076] According to another embodiment of the present invention, before performing identifier alignment processing on the user identifier set to be aligned with the business district user identifier set, the method further includes: obtaining processing rules for user identifiers pre-agreed with the business district, the processing rules including the data format and encryption method of the user identifiers; processing the user identifier set to be aligned using the processing rules, and using the processed user identifier set to be aligned as the user identifier set to be aligned. Before performing identifier alignment processing, the user identifier set to be aligned and the business district user identifier set need to be uniformly processed. Specifically, the business district party and the partner define a uniform data format for user identifiers in order to perform identifier ID alignment. Here, the ID is a unique identity identifier. Depending on different marketing scenarios, the ID type for alignment can be different. In addition, a uniform ID encryption method must also be defined. The ID cannot be transmitted in plaintext; it must be encrypted, and the encryption method must be unified. The encryption method can take multiple forms, such as MD5 or hash encryption, or a combination of multiple encryption methods, to prevent the encrypted ID from being decrypted, thereby better ensuring data security and privacy.
[0077] After identifier alignment, for aligned user identifiers, partners can query the aligned user's mobile trajectory information over the past year based on their own business data (such as mobile signaling data). Signaling data is data collected by mobile operators at a fixed frequency regarding the current time and location of mobile devices; it can be obtained after security processing and authorization by the mobile operator. Additionally, assuming the partner is an e-commerce platform, they can query the user's mobile trajectory information based on the shipping address and other details in the aligned user's order data, and so on. Different methods can be used to obtain the aligned user's mobile trajectory information for different business scenarios.
[0078] According to another embodiment of the present invention, determining the dense source areas of visiting users based on the mobile trajectory data includes: determining the residence and workplace of each visiting user based on the mobile trajectory data; using a spatial clustering algorithm to cluster the residence and workplace to obtain multiple clustering regions; and selecting a specified number of clustering regions from the multiple clustering regions where the visiting users are most densely distributed in space as the dense source areas of the visiting users. In one embodiment of the present invention, assuming that the mobile trajectory information of aligned users is obtained through mobile signaling data, firstly, after obtaining the trajectory information of aligned users, the partner determines the residence and workplace of the users based on the trajectory information, wherein the determination logic is: the high-frequency dwelling point of the user between 9:00 and 17:00 on weekdays is the workplace, and the high-frequency dwelling point of the user between 10:00 and 8:00 is the residence (dwelling point: the location where the sample trajectory information stays for a long time). Then, the calculation of the residence and workplace of all users is completed, and the spatial clustering algorithm is used to cluster the residence and workplace to obtain clustering regions. Finally, the clusters with the highest density of visitors in the space (e.g., 10) are selected from the clustering results and displayed on the map, which are the top ten densest sources of visitors to the business district.
[0079] After identifying the densely populated sources of visitors, customer profiles will be calculated for each densely populated source based on the visitor's characteristic data. The dimensions of the customer profiles include: sample visit volume over a fixed time period, the temporal distribution of visitor appearances (e.g., counting the number of samples visiting at different times), component analysis, and social graph analysis. Component analysis can include structured features such as gender, age, presence of children, spending power, and shopping preferences. Social graph analysis primarily utilizes relevant data from social media platforms, extracting representations reflecting first-order and second-order similarity based on the relationships between existing visitors and their relatives, friends, and colleagues using the LINE algorithm in GraphEmbedding. This information serves as the basis for later matching the existing visitor profiles with the target new customer profiles. Meanwhile, in the processing of structured features, different approaches are taken for the time attributes of user tags: For static features, their representations are stored in the content library for easy later use; for dynamic features, such as user preferences and interests that change frequently over time, long-term representation vectors can be calculated on a daily or weekly basis, and short-term representation vectors can be calculated on an hourly or minutely basis. The short-term and long-term representation vectors are then weighted and used to calculate the matching degree with target new customers, thereby achieving real-time user recommendations. Here, customer profiling does not determine an individual's specific profile or tags.
[0080] According to another embodiment of the present invention, comparing the customer profile with the historical user profile of the business district to determine the target user includes:
[0081] The structured profile data in the customer profile and the historical user profile of the business district are categorized and encoded, and the profile representation vector is obtained by dimensionality reduction.
[0082] Social network representations are extracted from the unstructured profile data in the customer profile and the historical user profile of the business district, respectively, to obtain network representation vectors;
[0083] By concatenating the profile representation vector and the network representation vector together, feature vectors for the customer profile and the historical user profile of the business district are obtained respectively.
[0084] Based on the similarity between feature vectors, the similarity between the customer profile and the historical user profile of the business district is compared, and the visiting users corresponding to the customer profile with a similarity that meets a preset threshold are identified as target users.
[0085] The historical user profile of the business district is created by extracting features from users who have visited the district. When comparing the similarity between this customer profile and the historical user profile of the business district, a similar user mining algorithm is used. This algorithm works as follows: First, for structured profile data (e.g., age, gender, spending power), categorical encoding is performed (e.g., male is labeled 0001, female is labeled 1000), and dimensionality reduction is used to obtain a more concise profile representation vector. Then, for unstructured relationships (e.g., social graphs), the LINE algorithm of Graph Embedding is used to extract social network representations, obtaining network representation vectors, which are then concatenated with the profile representation vectors. Finally, the similarity between samples is calculated based on the concatenated sample representation vectors. Samples exceeding a certain threshold are considered part of the target new customer group. Here, a sample represents a single user.
[0086] According to another embodiment of the present invention, pushing information to the target user includes: obtaining a first dense source location of the target user and returning the first dense source location to the business district, so that the business district pushes information based on the first dense source location. After determining the target user, since the target user's movement trajectory is not limited to a certain business district, the dense source location of the target user can also be calculated according to the method for determining the dense source location of the user in the aforementioned step S102, so as to better determine the scope and content of information push for the business district.
[0087] Furthermore, to ensure data security and privacy, the comparison of customer profiles with historical user profiles within the business district to identify target users can be conducted through a third-party arbitration institution. Additionally, after determining the primary source of target users, the data can be aggregated by the third-party arbitration institution and then sent to the business district operator for subsequent marketing planning. This collaboration can involve one or more partners, depending on the marketing content requirements. When constructing customer profiles, user characteristics originate from the business district's own tag features, e-commerce tag features, and social graph features from social platforms, etc. The feature vector data from these various parties cannot be aggregated to calculate cosine similarity on a single platform. Therefore, homomorphic encryption technology is used. The results calculated from the different feature vectors of the three platforms are encrypted using a public key and sent to the third-party arbitration institution. The third-party arbitration institution decrypts the data using its private key and then sends the cosine similarity result back to the business district.
[0088] Figure 2 This is a schematic diagram illustrating the implementation process of an embodiment of the present invention. For example... Figure 2 As shown, this illustrates how, in one embodiment of the present invention, information is pushed based on business partners, collaborators, and third-party arbitration institutions. It mainly includes the following steps:
[0089] 1. Both parties (the business district partner and the partner) agree on the rules for handling user IDs: At the start of the project, the business district partner and the partner define a unified user data ID format for ID alignment. Here, the ID should be a unique identifier. The aligned ID type can vary depending on the marketing scenario. Furthermore, a unified ID encryption method must be defined. IDs cannot be transmitted in plaintext; they must be encrypted, and the encryption method must be standardized. Multiple encryption methods can be used, such as MD5 or hash encryption, or a combination of multiple encryption methods, to prevent the encrypted ID from being decrypted.
[0090] 2. Both parties (the business district and the partner) prepare ID sets and feature sets: The business district provides user IDs from the past year based on its own business (e.g., catering, accommodation, tourism, travel, etc.) consumption data. Similarly, the partner provides user IDs from the past year based on its own business data.
[0091] 3. ID Alignment Preparation: The operator of the business district can initiate an ID alignment task at this time and send it to the partner so that the partner can choose to accept or reject the ID alignment task.
[0092] 4. ID Alignment Processing: Both parties (the business circle and the partner) encrypt the IDs that need to be aligned and send them separately to a third-party arbitration institution (an independent, neutral, and secure third-party node). The arbitration institution performs ID alignment (e.g., taking the intersection) and then transmits the aligned IDs to the partner. In practice, besides aggregating the IDs to the arbitration institution, if the business circle and the partner trust each other, the partner can also send the IDs directly to the business circle, and after the business circle completes the ID alignment, it can send them to the partner, or vice versa. This depends on which method meets business and security compliance requirements. The ID alignment method is not fixed, and multiple methods can be used for security considerations, but the core is taking the intersection. If the actual business requires, other methods such as taking the union, left outer join, or right outer join can also be used. The partner can be a mobile app, online shopping website, offline shopping platform, online or offline tourism product sales platform, etc., mainly depending on the business needs of the business circle and the actual information that the partner can provide.
[0093] 5. Obtain the alignment ID's trajectory information: After obtaining the encrypted alignment ID (i.e., the data sample shared by both parties), the partner queries the sample's mobile trajectory information for the past year based on its own business data (such as mobile signaling). The signaling data is data collected by the mobile operator at a fixed frequency, showing the mobile terminal's current time and location.
[0094] 6. Calculate the dense source areas of visiting users: First, after obtaining the aligned sample trajectory information, the partner determines the residential and workplace locations of the samples based on the trajectory information. Specifically, the high-frequency dwell points of the samples between 9:00 AM and 5:00 PM on weekdays are considered workplaces, and the high-frequency dwell points between 10:00 PM and 8:00 AM are considered residential locations (dwelling points: locations where the sample trajectory information lingers for an extended period). Then, the residential and workplace locations of all samples are calculated, and a spatial clustering algorithm is used to cluster the residential and workplace locations, resulting in clustered regions. Finally, the 10 clustered regions with the highest spatial distribution of visiting users are selected from the clustering results and displayed on a map, i.e., the top ten dense source areas of visiting users in the business district.
[0095] 7. Depicting Visitor Profiles: After identifying the top ten most frequent sources of visitors to the business district, calculate the profiles of all sample groups from these sources. Profile dimensions include: visitor volume over a fixed time period, the temporal distribution of visitors (e.g., counting the number of samples visiting at different times), component analysis, and social graph analysis. Component analysis can include structured features such as gender, age, presence of children, spending power, and shopping preferences. Social graph analysis primarily utilizes data from social media platforms, extracting representations of first- and second-order similarity based on relationships with friends, family, and colleagues of previously visited users using the LINE algorithm in Graph Embedding. This data serves as the basis for later matching the visitor profiles with the profiles of target new customers. Meanwhile, in the processing of structured features, different approaches are taken for the time attributes of user tags: For static features, their representations are stored in the content library for easy later use; for dynamic features, such as user preferences and interests that change frequently over time, long-term representation vectors can be calculated on a daily or weekly basis, and short-term representation vectors can be calculated on an hourly or minutely basis. The short-term and long-term representation vectors are then weighted and used to calculate the matching degree with target new customers, thereby achieving real-time user recommendations. Here, the group profiling does not determine the specific profile or tags of an individual.
[0096] 8. Identifying Target New Customers: The partner will use a similar user mining algorithm to obtain user IDs among unvisited users who have a high similarity to visited users, based on the user profiles of users who visited the business district and the historical user profiles of the business district. These are the target new customer IDs. The similar user mining algorithm is as follows: First, for structured profile data (e.g., age, gender, spending power), categorical encoding is performed (e.g., male is labeled as 0001, female as 1000), and dimensionality reduction is used to obtain a more concise profile representation vector. Then, for unstructured relationships (e.g., social graphs), the LINE algorithm of Graph Embedding is used to extract social network representations, obtaining network representation vectors, which are then concatenated with the profile representation vectors. Finally, the similarity between samples is calculated based on the representation vectors of the samples; samples exceeding a certain threshold are considered target new customers.
[0097] 9. Identify the densest sources of the target new customer group: Once the target new customer group is identified, use the method described in step 6 to calculate the top ten densest sources of the target new customer group;
[0098] 10. Results Return: The results are returned to a third-party arbitration institution for aggregation, and then the arbitration institution sends them to the business district operator for subsequent marketing planning. The partners here can be one or more, depending on the content of the marketing campaign.
[0099] 11. The actual marketing activities are carried out by the business district.
[0100] Figure 3 This is a schematic diagram illustrating the customer profile construction and new customer acquisition process according to an embodiment of the present invention. Figure 3 As shown, after the business district and its partners extract and reduce user features, they obtain group profiles. User i and user j in the figure both represent group profiles. Then, intermediate results are obtained by comparing the group profiles for similarity. These intermediate results are then homomorphically encrypted and sent to a third-party arbitration institution. Throughout this process, neither the third-party arbitration institution, the business district, nor the partners directly obtain the original data from the other parties; they only obtain encrypted or privacy-protected data, thus ensuring data security and privacy.
[0101] Figure 4 This is a schematic diagram of the main modules of an information push device according to an embodiment of the present invention. Figure 4 As shown, the information push device 400 of this embodiment mainly includes a user identifier alignment module 401, a user origin determination module 402, a customer profile calculation module 403, and a target user determination module 404.
[0102] User identifier alignment module 401 is used to obtain a set of user identifiers to be aligned and perform identifier alignment processing between the set of user identifiers to be aligned and the set of user identifiers in the business district to obtain aligned user identifiers.
[0103] The user origin determination module 402 is used to obtain the movement trajectory data of the aligned user identifier based on the business data, and determine the dense origin of the visiting user based on the movement trajectory data.
[0104] The customer profile calculation module 403 is used to calculate the customer profile of each dense source area of visitors based on the characteristic data of the visitors.
[0105] The target user determination module 404 is used to compare the similarity between the customer profile and the historical user profile of the business district to determine the target user, and to push information to the target user.
[0106] According to one embodiment of the present invention, the identifier alignment process includes: intersection processing, union processing, left outer join processing, or right outer join processing.
[0107] According to another embodiment of the present invention, the user identifier alignment module 401 can also be used to: send the user identifier set to be aligned to a third-party arbitration institution, so that the third-party arbitration institution can perform identifier alignment processing based on the user identifier set to be aligned and the business district user identifier set.
[0108] According to another embodiment of the present invention, the information push device 400 further includes an identifier set preprocessing module (not shown in the figure), used for:
[0109] Before performing the alignment process between the user identifier set to be aligned and the business district user identifier set, the processing rules for user identifiers agreed upon with the business district in advance are obtained. The processing rules include the data format and encryption method of the user identifiers.
[0110] The processing rules are used to process the set of user identifiers to be aligned, and the processed set of user identifiers to be aligned is used as the set of user identifiers to be aligned.
[0111] According to another embodiment of the present invention, the user origin determination module 402 can also be used for:
[0112] The residence and workplace of each visiting user are determined based on the aforementioned movement trajectory data;
[0113] A spatial clustering algorithm was used to cluster residential and workplace locations, resulting in multiple clustered regions.
[0114] From the multiple cluster regions, select a specified number of cluster regions where the visiting users are most densely distributed in space as the dense source of visiting users.
[0115] According to yet another embodiment of the present invention, the target user determination module 604 can also be used for:
[0116] The structured profile data in the customer profile and the historical user profile of the business district are categorized and encoded, and the profile representation vector is obtained by dimensionality reduction.
[0117] Social network representations are extracted from the unstructured profile data in the customer profile and the historical user profile of the business district, respectively, to obtain network representation vectors;
[0118] By concatenating the profile representation vector and the network representation vector together, feature vectors for the customer profile and the historical user profile of the business district are obtained respectively.
[0119] Based on the similarity between feature vectors, the similarity between the customer profile and the historical user profile of the business district is compared, and the visiting users corresponding to the customer profile with a similarity that meets a preset threshold are identified as target users.
[0120] According to another embodiment of the present invention, the target user determination module 604 can also be used for:
[0121] The first dense source location of the target user is obtained and returned to the business district so that the business district can push information based on the first dense source location.
[0122] According to the technical solution of this invention, by acquiring a set of user identifiers to be aligned and performing identifier alignment processing with a set of user identifiers in the business district to obtain aligned user identifiers; by acquiring the movement trajectory data of the aligned user identifiers based on business data and determining the dense source areas of visiting users based on the movement trajectory data; by calculating the customer profile of visiting users in each dense source area based on the feature data of visiting users; by comparing the similarity between the customer profile and the historical user profile of the business district to determine the target users and pushing information to the target users, it is possible to combine data from multiple sources to characterize user profiles and push information, making up for the deficiency of the single dimension of user data features in the business district operation scenario, realizing the full mining of the source areas of users who have not visited the business district, thereby better pushing information such as advertisements in the scenario of attracting new customers in the business district; based on the spatiotemporal attributes of user trajectory information, the precise dense source areas of users are obtained, thereby enabling better, more real-time and more accurate characterization of user customer profiles, more accurate determination of potential customers in the business district, and facilitating the improvement of the success rate of information push.
[0123] This technical solution first utilizes the temporal attributes of trajectories to filter precise locations of residence as user origins. Then, leveraging spatial attributes, it employs the OPTICS density clustering algorithm and similar user mining algorithm to explore and mine the origins of target new customers through existing customer flow analysis. This locks in potential target customer information, solving the problems of "who to target, where to advertise, how to advertise, and when to advertise," significantly improving the efficiency of lead generation planning and marketing. When constructing customer profiles, user characteristics come from the business district's own tag features, e-commerce tag features, and social graph features from social platforms. The feature vector data from each party cannot be aggregated to calculate cosine similarity on a single platform. Therefore, homomorphic encryption technology is used. The results calculated from the different feature vectors of the three platforms are encrypted using a public key and sent to a third-party arbitration institution. The third-party arbitration institution decrypts the data using a private key and then sends the cosine similarity result back to the business district.
[0124] Figure 5 An exemplary system architecture 500 is shown, in which the information push method or information push apparatus of the present invention can be applied.
[0125] like Figure 5 As shown, system architecture 500 may include terminal devices 501, 502, and 503, a network 504, and a server 505. Network 504 serves as the medium for providing communication links between terminal devices 501, 502, and 503 and server 505. Network 504 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0126] Users can use terminal devices 501, 502, and 503 to interact with server 505 via network 504 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 501, 502, and 503, such as shopping applications, web browser applications, search applications, advertising push tools, social media platform software, etc. (for example only).
[0127] Terminal devices 501, 502, and 503 can be various electronic devices with displays that support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0128] Server 505 can be a server providing various services, such as a backend management server supporting shopping websites browsed by users using terminal devices 501, 502, and 503 (for example only). The backend management server can obtain a set of user identifiers to be aligned from received information push requests and other data, and perform identifier alignment processing between the set of user identifiers to be aligned and the set of user identifiers in the business district to obtain aligned user identifiers; obtain the movement trajectory data of the aligned user identifiers based on business data, and determine the dense source areas of visiting users based on the movement trajectory data; calculate the customer profile of visiting users in each dense source area based on the characteristic data of visiting users; compare the similarity of the customer profile with the historical user profiles of the business district to determine target users, and perform information push and other processing on the target users, and feed back the processing results (e.g., information push results, product information push results—for example only) to the terminal devices.
[0129] It should be noted that the information push method provided in this embodiment of the invention is generally executed by server 505, and correspondingly, the information push device is generally set in server 505.
[0130] It should be understood that Figure 5 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0131] The following is for reference. Figure 6 It shows a schematic diagram of the structure of a computer system 600 suitable for implementing terminal devices or servers of the present invention. Figure 6 The terminal device or server shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.
[0132] like Figure 6As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 602 or programs loaded from storage section 608 into random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of the system 600. The CPU 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0133] The following components are connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 610 as needed so that computer programs read from it can be installed into storage section 608 as needed.
[0134] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by central processing unit (CPU) 601, it performs the functions defined above in the system of this invention.
[0135] It should be noted that the computer-readable medium shown in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0137] The units or modules described in the embodiments of the present invention can be implemented in software or hardware. The described units or modules can also be housed in a processor. For example, a processor can be described as including a user identifier alignment module, a user origin determination module, a customer profile calculation module, and a target user determination module. The names of these units or modules do not necessarily limit the specific unit or module itself. For example, the user identifier alignment module can also be described as "a module for obtaining a set of user identifiers to be aligned and performing identifier alignment processing between the set of user identifiers to be aligned and a set of business district user identifiers to obtain aligned user identifiers."
[0138] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, which, when executed by the device, cause the device to include: acquiring a set of user identifiers to be aligned, and performing identifier alignment processing between the set of user identifiers to be aligned and a set of user identifiers in a business district to obtain aligned user identifiers; acquiring movement trajectory data of the aligned user identifiers based on business data, and determining the dense source areas of visiting users based on the movement trajectory data; calculating customer profiles for each dense source area of visiting users based on the characteristic data of visiting users; comparing the similarity of the customer profiles with the historical user profiles of the business district to determine target users, and pushing information to the target users.
[0139] According to the technical solution of this invention, by acquiring a set of user identifiers to be aligned and performing identifier alignment processing with a set of user identifiers in the business district to obtain aligned user identifiers; by acquiring the movement trajectory data of the aligned user identifiers based on business data and determining the dense source areas of visiting users based on the movement trajectory data; by calculating the customer profile of visiting users in each dense source area based on the feature data of visiting users; by comparing the similarity between the customer profile and the historical user profile of the business district to determine the target users and pushing information to the target users, it is possible to combine data from multiple sources to characterize user profiles and push information, making up for the deficiency of the single dimension of user data features in the business district operation scenario, realizing the full mining of the source areas of users who have not visited the business district, thereby better pushing information such as advertisements in the scenario of attracting new customers in the business district; based on the spatiotemporal attributes of user trajectory information, the precise dense source areas of users are obtained, thereby enabling better, more real-time and more accurate characterization of user customer profiles, more accurate determination of potential customers in the business district, and facilitating the improvement of the success rate of information push.
[0140] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for information push, characterized in that, include: Obtain the set of user identifiers to be aligned, and perform identifier alignment processing between the set of user identifiers to be aligned and the set of user identifiers in the business district to obtain aligned user identifiers; The movement trajectory data of the aligned user identifiers is obtained based on the business data, and the dense source areas of users visiting the business district are determined based on the movement trajectory data. Based on the characteristic data of visitors from the densely populated source areas, a customer profile is calculated for each densely populated source area. The customer profile is compared with the historical user profile of the business district. Unvisited users corresponding to the customer profile with a similarity that meets a preset threshold are identified as target users, and information is pushed to the target users. The historical user profile of the business district is obtained by extracting features from historical users who have visited the business district.
2. The method according to claim 1, characterized in that, The identifier alignment process includes: Processing can be done by taking the intersection, taking the union, or performing a left outer join or a right outer join.
3. The method according to claim 1, characterized in that, The process of aligning the user identifier set to be aligned with the business district user identifier set includes: The set of user identifiers to be aligned is sent to a third-party arbitration institution so that the third-party arbitration institution can perform identifier alignment processing based on the set of user identifiers to be aligned and the set of user identifiers in the business district.
4. The method according to claim 1, characterized in that, Before performing identifier alignment processing between the user identifier set to be aligned and the business district user identifier set, the process also includes: The processing rules for obtaining user identifiers pre-agreed with the business district include the data format and encryption method of the user identifiers; The processing rules are used to process the set of user identifiers to be aligned, and the processed set of user identifiers to be aligned is used as the set of user identifiers to be aligned.
5. The method according to claim 1, characterized in that, Based on the aforementioned movement trajectory data, the dense source areas of users visiting the business district include: Based on the movement trajectory data, the residential and workplace locations of each user visiting the business district are determined; A spatial clustering algorithm was used to cluster residential and workplace locations, resulting in multiple clustered regions. Select a specified number of clusters from the multiple clusters to identify the densest distribution of visitors to the business district in the space, as the dense source of visitors to the business district.
6. The method according to claim 1, characterized in that, The target users are identified by comparing the aforementioned customer profile with the historical user profiles of the business district. The structured profile data in the customer profile and the historical user profile of the business district are categorized and encoded, and the profile representation vector is obtained by dimensionality reduction. Social network representations are extracted from the unstructured profile data in the customer profile and the historical user profile of the business district, respectively, to obtain network representation vectors; By concatenating the profile representation vector and the network representation vector together, feature vectors for the customer profile and the historical user profile of the business district are obtained respectively. Based on the similarity between feature vectors, the similarity between the customer profile and the historical user profile of the business district is compared, and the unvisited users corresponding to the customer profile with a similarity that meets a preset threshold are identified as target users.
7. The method according to claim 1, characterized in that, Pushing information to the target user includes: The first dense source location of the target user is obtained and returned to the business district so that the business district can push information based on the first dense source location.
8. An information push device, characterized in that, include: The user identifier alignment module is used to obtain a set of user identifiers to be aligned and perform identifier alignment processing between the set of user identifiers to be aligned and the set of user identifiers in the business district to obtain aligned user identifiers. The user origin determination module is used to obtain the movement trajectory data of the aligned user identifiers based on business data, and determine the dense origin of users visiting the business district based on the movement trajectory data. The customer profile calculation module is used to calculate the customer profile for each dense source area based on the characteristic data of the visitors from the dense source areas. The target user determination module is used to compare the similarity between the customer profile and the historical user profile of the business district, determine the unvisited users corresponding to the customer profile with a similarity that meets a preset threshold as target users, and push information to the target users; wherein, the historical user profile of the business district is obtained by extracting features from historical users who have visited the business district.
9. An electronic device for information push, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.