Marketing data generation method and device based on portrait data, equipment and medium
By constructing a three-tiered data gateway and a marketing data graph, combined with a dynamic weighting mechanism and federated learning, the limitations of data dimensions and insufficient scenario adaptation in the marketing system were solved, enabling the generation of precise marketing strategies and improving operational efficiency.
Patent Information
- Application Number
- CN202511512609.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-11-21
AI Technical Summary
Existing marketing systems suffer from limitations in data dimensions, insufficient granularity in user profiles, lack of contextualization, inadequate adaptation to gift scenarios, and low operational efficiency, resulting in insufficient recommendation accuracy and an inability to meet the needs of refined operations.
A three-tier data gateway is constructed to collect marketing-related data. Feature fusion is performed through purification and star-shaped collaborative networks. Marketing data graphs are used for scenario identification. Targeted marketing strategies are generated by combining dynamic weighting mechanisms and federated learning.
It achieves multi-dimensional data integration, improves the accuracy of marketing strategies and the ability to adapt to different scenarios, solves the problem of data silos, enhances operational efficiency, and meets the needs of refined operations.
Smart Images

Figure CN120996866A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for generating marketing data based on profile data. Background Technology
[0002] The existing marketing systems mainly suffer from the following defects: 1. Data Dimension Limitations: Relying solely on data from a single merchant leads to a narrow perspective in assessing user needs, failing to capture potential demands arising from cross-domain consumption trends and multi-platform behavioral correlations, thus limiting recommendation accuracy. For example, a clothing brand that only makes recommendations based on its own store's customer data cannot capture the potential needs of its customers in other fashion categories; 2. Coarse audience profiles: The generated audience profiles are coarse-grained and cannot accurately pinpoint the personalized needs of segmented groups. This is especially true in areas with fragmented needs and strong scene dependence (such as the women's clothing industry), making it difficult to meet the needs of refined operations. For example, it is impossible to accurately pinpoint the needs of new employees for simple, fashionable, and cost-effective clothing. 3. Lack of contextualization: There is a lack of in-depth exploration and construction of consumption scenarios. Recommendations remain only at the product level, failing to connect products with the user's situation and emotional needs, making it difficult to evoke emotional resonance and impacting conversion rates. For example, recommending sun-protective clothing and accessories to users who frequently engage in outdoor activities in the context of "high-temperature outdoor scenarios" at the end of summer. 4. Insufficient Adaptation to Gift Scenarios: The lack of a "relationship-scenario-emotion" association system means the system fails to understand the relationship between the giver and receiver, the specificity of the gift's context, and the emotions intended to be conveyed. This results in untargeted gift recommendations and an inability to address users' "gift-giving dilemma." For example, during Valentine's Day, the system cannot recommend suitable gifts based on the couple's relationship, the festive atmosphere, and romantic emotional needs. 5. Low operational efficiency: Relying on manual content creation and strategy formulation, small and medium-sized businesses lack professional teams and find it difficult to implement refined operations. Furthermore, the consistency and accuracy of manually produced content and strategies are insufficient. 6. Inefficient Feature Utilization: The system has a weak ability to process unstructured user features (such as fragmented reviews and vague descriptions of needs), failing to transform them into structured information that can guide operations. This makes it difficult to support AI-driven, refined operational decisions, resulting in a large number of potential needs being overlooked. For example, a user might mention in a product review that they "want a lighter design," but current technology cannot effectively extract this need and use it for product recommendations and improvement suggestions. Summary of the Invention
[0003] In view of the above, it is necessary to provide a marketing data generation method, device, equipment and medium based on profile data, which aims to solve the problems of insufficient granularity of the audience profile, lack of scenario-based engine, imbalance of gift scenario adaptation, low operational efficiency and limited data dimensions in the existing technology.
[0004] A marketing data generation method based on profile data, the marketing data generation method based on profile data includes: A three-tiered data gateway system is constructed, comprising a private domain gateway, an industry domain gateway, and a public domain gateway. Marketing-related data is collected using the three-tiered data gateway system to obtain initial data. The initial data is purified to obtain intermediate data; A star-shaped collaborative network based on a dynamic weighting mechanism and a federated learning mechanism is obtained, and the star-shaped collaborative network is used to perform feature fusion on the intermediate data to obtain fused features; A pre-constructed marketing data map is obtained, and the target scene is obtained by scene recognition of the fused features based on the marketing data map; wherein, the marketing data map is constructed based on a secondary scene classification tree including gift scenes; The target engine is obtained by matching the target scenario in a pre-built strategy generation engine. The marketing data graph is processed using the target engine to obtain a target marketing strategy; Target marketing data is generated based on the target marketing strategy and the marketing data map.
[0005] A marketing data generation device based on profile data, the marketing data generation device based on profile data includes: The data collection unit is used to construct a three-level data gateway, including a private domain gateway, an industry domain gateway, and a public domain gateway, and to use the three-level data gateway to collect marketing-related data to obtain initial data. A purification unit is used to purify the initial data to obtain intermediate data; The fusion unit is used to acquire a star-shaped collaborative network constructed based on a dynamic weighting mechanism and a federated learning mechanism, and to use the star-shaped collaborative network to perform feature fusion on the intermediate data to obtain fused features; The identification unit is used to acquire a pre-constructed marketing data map and perform scene identification on the fused features based on the marketing data map to obtain the target scene; wherein, the marketing data map is constructed based on a secondary scene classification tree including gift scenes; The matching unit is used to perform matching in a pre-built strategy generation engine according to the target scenario to obtain the target engine; The processing unit is used to process the marketing data graph using the target engine to obtain a target marketing strategy; The generation unit is used to generate target marketing data based on the target marketing strategy and the marketing data map.
[0006] A computer device, the computer device comprising: Memory, storing at least one instruction; and The processor executes instructions stored in the memory to implement the marketing data generation method based on profile data.
[0007] A computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to implement the marketing data generation method based on profile data.
[0008] As can be seen from the above technical solutions, this invention can collect and purify marketing-related data based on a three-level data gateway, and use a star-shaped collaborative network built based on dynamic weighting and federated learning mechanisms for feature fusion, thus solving the problems of limited data dimensions and data silos; it can identify scenarios based on a marketing data graph constructed from a two-level scenario classification tree including gift scenarios, thus solving the problems of inefficient use of unstructured data and insufficient granularity of user profiles; it can generate marketing strategies using a target engine that matches scenarios, thus solving the problems of missing scenario-based engines and imbalance in gift scenario adaptation; and it can generate target marketing data based on target marketing strategies and marketing data graphs, thus solving the problems of low operational efficiency and insufficient content accuracy. Attached Figure Description
[0009] Figure 1 This is a flowchart of a preferred embodiment of the marketing data generation method based on profile data of the present invention; Figure 2 This is a functional block diagram of a preferred embodiment of the marketing data generation device based on profile data of the present invention; Figure 3 This is a schematic diagram of the structure of a computer device that implements a preferred embodiment of the marketing data generation method based on profile data according to the present invention. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0011] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the marketing data generation method based on user profile data according to the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different needs.
[0012] The marketing data generation method based on profile data is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0013] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.
[0014] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0015] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0016] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0017] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0018] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).
[0019] S10, construct a three-level data gateway including a private domain gateway, an industry domain gateway, and a public domain gateway, and use the three-level data gateway to collect marketing-related data to obtain initial data.
[0020] In this embodiment, the initial data obtained by using the three-level data gateway to collect marketing-related data includes: The private domain gateway is used to synchronize merchant order data, customer data, and private domain interaction data as private domain data at a preset synchronization frequency; wherein, the preset synchronization frequency is dynamically configured according to the data update volume of each party, and an incremental synchronization mechanism is adopted during peak periods; The industry domain gateway collects vertical domain data and cross-merchant consumption association rules as industry domain data according to a first preset cycle. The public domain gateway collects social media trending topics, e-commerce comment sentiment data, and environmental dynamic data as public domain data according to a second preset cycle. The initial data is obtained by combining the private domain data, the industry domain data, and the public domain data.
[0021] For example, the private domain gateway can synchronize in real time order data from the merchant's ERP (Enterprise Resource Planning) system (including SKU (Stock Keeping Unit) details, payment methods, etc.), CRM (Customer Relationship Management) customer profiles (such as spending levels, tag preferences, etc.), and private domain interaction records (such as community posts, customer service inquiries, etc.). The preset synchronization frequency can be dynamically configured according to the data update volume of each merchant, and an incremental synchronization mechanism is used during peak periods to reduce data transmission volume and resource consumption.
[0022] For example, the industry domain gateway can update the vertical domain data platform (such as style trends and price range distribution in the women's clothing industry) and cross-merchant consumption association rules (such as "70% of users who buy dresses will pair them with high heels").
[0023] The cross-merchant consumption association rules can be generated using federated learning technology, which can avoid cross-entity leakage of raw data in a distributed training mode.
[0024] For example, the public domain gateway can incrementally update social media trending topics, e-commerce comment sentiment, and environmental dynamic data (such as weather, holiday calendar, etc.) every 2 hours.
[0025] Furthermore, the trending topics on the social media platform can be subject to a heat decay coefficient, and the recent trending topics can be given higher weight in the analysis through the design of a time decay function.
[0026] The above embodiments can overcome the limitations of a single data source, integrate multi-dimensional data from merchant private domains, industry domains, and public domains, improve feature dimensions, solve the problems of limited data dimensions and data silos in existing technologies, and provide a comprehensive data foundation for subsequent precision marketing.
[0027] S11, The initial data is purified to obtain intermediate data.
[0028] In this embodiment, the purification process of the initial data to obtain intermediate data includes: The interquartile range (IQR) algorithm is used to identify outliers in the initial data, and the outliers are removed from the initial data to obtain the first data. The purchase timestamps in the first data are corrected using a Long Short-Term Memory (LSTM) network trained based on an error penalty coefficient to obtain the second data; wherein, during the training of the LSTM network using the error penalty coefficient, when the model output exceeds the multi-dimensional time constraints, the current model is penalized using the error penalty coefficient. Obtain order data generated by the same user within a preset time period as baseline data, and calculate the difference between the second data and the baseline data; When the difference between the second data and the baseline data is greater than a preset threshold, the second data is sent to the verification platform for secondary correction to obtain the third data; The third data is converted into a configuration format to obtain the fourth data; The fourth data is anonymized to obtain the intermediate data; wherein, sensitive personal information is encrypted and de-identified; the user's mobile phone number is encrypted using the MD5 (Message-Digest Algorithm 5) algorithm to obtain the ciphertext of the mobile phone number, and random noise is added to the ciphertext of the mobile phone number; the city-level information in the address information is retained; the SHA-256 (Secure Hash Algorithm 256-bit) hash value of the user's email address is calculated, and the SHA-256 hash value is masked.
[0029] For example, when the amount of a single order is more than 3 times the average (this value can be set based on industry characteristics and can be dynamically adjusted), the amount of the single order can be determined as an outlier based on the interquartile range algorithm.
[0030] For example, an LSTM time series model can be used to correct missing purchase timestamps caused by network latency, equipment failure, etc., and a triple mechanism can be used to ensure that the correction error is controlled within ±5 minutes. Specifically, a multi-dimensional time series constraint is constructed, integrating related data such as order payment time, logistics pickup time, and inventory change time to form a reasonable time stamp interval boundary (e.g., generating an order timestamp within 10 minutes after payment). An error penalty function is introduced during the model training phase, imposing a high penalty weight on prediction results that exceed the reasonable range, guiding the model to prioritize learning prediction logic with smaller errors. Before the corrected result (i.e., the second data) is output, it can be compared with baseline data such as the recent order time patterns of the same user and the order acceptance time range of the merchant. If it exceeds the threshold, a second correction is triggered. For example, a "automated pre-judgment + hierarchical manual verification" mechanism can be used on the verification platform. Basic anomalies are automatically handled by the verification platform, while high-risk anomalies can trigger manual verification. At the same time, logs are recorded, and the proportion of manual verification is gradually reduced through machine learning optimization of the model, and the verification results are incorporated into the model training closed loop.
[0031] For example, dates can be standardized to the format "yyyy-MM-ddHH:mm:ss", and product categories can adopt a three-level category system, such as "Apparel > Women's Clothing > Dresses", thereby ensuring consistency and compliance across data sources. The three-level category system can also be mapped to merchant-defined categories, and naming differences can be resolved through fuzzy matching algorithms (such as edit distance calculation).
[0032] For example, when performing masking, the first three characters can be masked with "*".
[0033] Through the above embodiments, data quality and security can be ensured based on three-dimensional purification operations, meeting relevant requirements, while protecting user privacy, enhancing data credibility, and providing high-quality data support for subsequent model training and strategy generation.
[0034] S12, obtain a star-shaped collaborative network constructed based on dynamic weighting mechanism and federated learning mechanism, and use the star-shaped collaborative network to perform feature fusion on the intermediate data to obtain fused features.
[0035] In this embodiment, before obtaining the star-shaped collaborative network constructed based on the dynamic weight mechanism and the federated learning mechanism, the method further includes: The star-shaped collaborative network, comprising edge nodes and a central node, is constructed based on the dynamic weighting mechanism and the federated learning mechanism. Each merchant is identified as an edge node. The central node includes a model parameter aggregation center, a homomorphic encryption module, and a dynamic weight adjustment engine. The weight dynamic adjustment engine is used to configure basic weights based on data source relevance and data quality. The weight dynamic adjustment engine adjusts the dynamic correction coefficient according to data freshness and domain relevance at preset time intervals, calculates the adjustment ratio by summing the adjusted dynamic correction coefficient with 1, calculates the product of the adjustment ratio and the basic weight to obtain the dynamic weight, and distributes the dynamic weight to each edge node. The adjustment ratio is less than or equal to the configured ratio. During the training of the star-shaped collaborative network, each edge node receives dynamic weights from the dynamic weight adjustment engine and fuses local data according to the dynamic weights to train a local model. The integrity of the model parameters of each local model is then verified, and the verified model parameters are uploaded to the model parameter aggregation center. The local data has the same data dimension as the initial data, and the local data also includes structured feature vectors of spatiotemporal data. These structured feature vectors are used to fuse user behavior features from the local data with an attention mechanism to form a basic model for scene prediction. The model parameter aggregation center is used to aggregate the uploaded model parameters to obtain aggregated parameters. The homomorphic encryption module uses the Commercial Cryptography Algorithm 4 (SM4) block cipher algorithm to encrypt and distribute the aggregation parameters to each edge node.
[0036] For example, the weight dynamic adjustment engine can fine-tune the weight of the data source weekly, with the adjustment range not exceeding 10% of the base weight, and then distribute the adjusted dynamic weight to each edge node.
[0037] For example, each merchant node can retain the original data in its local storage unit to prevent data leakage. Each merchant node can also prevent model parameters from being tampered with during transmission through methods such as checksum verification.
[0038] For example: Dynamic weight = Base weight × (1 + Dynamic adjustment coefficient). The base weight is configured based on data source relevance and data quality; for example, the weight for private domain data can be configured as 0.6, the weight for industry domain data as 0.3, and the weight for public domain data as 0.1. The dynamic adjustment coefficient can be adjusted based on data freshness and domain relevance; for example, the adjustment coefficient for private domain data from the past 7 days can be configured as 0.2 (weight increase of 20%), the adjustment coefficient for unrelated children's clothing industry data from women's clothing merchants can be configured as -0.3 (weight decrease of 30%), and the adjustment coefficient for highly timely social media trending data from the public domain can be configured as 0.15 (weight increase of 15%).
[0039] For example, spatiotemporal data such as weather and holidays can be transformed into structured feature vectors. For example, the "rainstorm day" vector can be defined as [weather type - rainstorm (1), holiday - no (0), weekday - yes (1), precipitation intensity - 0.8], with the last digit being the standardized value of precipitation intensity. Then, it is fused with user behavior features through an attention mechanism to form a basic model for scene prediction.
[0040] In the above embodiments, dynamic weighted federated learning based on a star-shaped collaborative network architecture, while ensuring data privacy, allocates higher weights to high-value data (such as recent private domain consumption data) during fusion through differentiated weight distribution. This improves the model's adaptability to multi-source data, making the generated user feature vectors more reflective of users' real needs and behavioral characteristics. Simultaneously, environmental feature embedding technology transforms spatiotemporal data into structured features, providing a foundation for scenario prediction and enabling subsequent marketing strategies to better integrate spatiotemporal factors, thereby improving scenario adaptability.
[0041] In this embodiment, a secure environment where "data is available but not visible" can also be constructed, supporting cross-domain feature interaction (such as the calculation of overlapping preferences between brand users and competitor users), and establishing a data usage traceability mechanism in conjunction with blockchain notarization to ensure that data usage is auditable and traceable.
[0042] Privacy computing sandboxes and blockchain-based evidence storage mechanisms further enhance data security, achieving "data usable but not visible," resolving privacy conflicts in cross-domain data fusion, and balancing the needs of data dimension expansion and privacy protection.
[0043] S13, obtain a pre-constructed marketing data map, and perform scene recognition on the fused features based on the marketing data map to obtain the target scene; wherein, the marketing data map is constructed based on a secondary scene classification tree including gift scenes.
[0044] In this embodiment, before obtaining the pre-constructed marketing data map, the method further includes: The three-level data gateway is used to collect data according to the data dimensions of the initial data to obtain the data to be processed. The data to be processed is segmented into words and tagged with parts of speech, and stop words are removed from the data to obtain the text preprocessing result; Deep semantic analysis of the social unstructured text in the text preprocessing results is performed based on the BERT-wwm (Bidirectional Encoder Representations from Transformers-Whole Word Masking) pre-trained model to obtain the core feature vector. Identify high-frequency scene words in the core feature vector and construct a scene feature library using the high-frequency scene words; The core feature vector is classified into sentiment tendencies to obtain the sentiment feature labels of the core feature vector; wherein, a reason sub-label is added to the data whose sentiment feature labels are negative sentiment. Named entity recognition technology is used to identify gift-related context words in the core feature vector, and a gift feature subset is constructed using these gift-related context words; wherein, the gift-related context words include relation words and event words; Calculate the mutual information entropy between the core feature vectors, and filter candidate features from the core feature vectors based on the mutual information entropy; A time decay factor is introduced into the density-based noisy spatial clustering method, and the clustering optimization of the selected candidate features is combined with the scene demand heat value to obtain clustered features. In the clustering process, the effect of different numbers of clusters is evaluated by the silhouette coefficient to select the optimal number of clusters, so that the similarity of samples within the same category is greater than or equal to a first value, and the similarity of samples between different categories is less than or equal to a second value, wherein the first value is greater than the second value. The time decay factor is configured according to the performance of each candidate feature within a specified time period. Based on the core feature vector, determine the urgency, relevance, and consumption potential of each user, obtain the urgency weight, relevance weight, and consumption potential weight, and perform a weighted calculation based on each user's urgency, relevance weight, and consumption potential weight to obtain the scenario demand heat value of each user. A dual mechanism of rule matching and probability prediction is adopted to construct the secondary scene classification tree based on the scene demand popularity value of each user; wherein, rule matching is used first to clarify the scene description, and then probability prediction is used to cover the ambiguous text; wherein, for the gift scene, the emotional intent is quantitatively scored. A knowledge graph is constructed based on the scene feature library, the emotional feature tags, the gift feature subset, the clustering features, and the secondary scene classification tree to obtain the marketing data graph. The marketing data graph includes an entity layer, a relationship layer, and an attribute layer. Each entity in the entity layer is uniquely identified using hash encoding. The relationship weights in the relationship layer are dynamically adjusted based on behavior frequency and timeliness. The attribute layer includes multi-dimensional attribute features for each entity.
[0045] For example, the Jieba segmentation tool can be used to complete word segmentation. It also supports custom dictionaries to expand industry terminology, remove stop words, and perform part-of-speech tagging.
[0046] For example, the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm can be used to identify high-frequency scene keywords (e.g., term frequency ≥ 5% of the total sample size) in order to construct the scene feature library.
[0047] For example, the TextCNN (Text Convolutional Neural Network) model can be used for sentiment classification, dividing text sentiment into three categories: positive, negative, and neutral, to generate corresponding sentiment feature labels. Negative sentiment can be further labeled with reason sub-labels (such as "poor quality" or "high price" as negative reasons).
[0048] For example, in the context of gift-giving, named entity recognition technology can be used to capture relational terms such as "best friend" and "boss" and event terms such as "birthday" and "promotion" separately, forming a subset of gift features that are associated and bound together.
[0049] For example, core features (such as purchase frequency in the last 30 days, average order value, and top 3 category preferences) can be selected from over 300 candidate features identified through multiple dimensions such as user behavior and product attributes. The importance of these features is calculated using mutual information entropy, and the top 50 core features are selected for subsequent clustering. Furthermore, the effect of different clustering numbers can be evaluated using the silhouette coefficient, and the optimal clustering number (default k=5~15) can be selected, requiring similarity within the same cluster to be ≥0.7 and similarity between different clusters to be ≤0.3. Based on the density-based spatial clustering of applications with noise (DBSCAN) method, a time decay factor is introduced, setting higher weight coefficients for consumption behavior in the last 3 months. Clustering is optimized by combining scenario demand heat values (adjusted according to urgency, relevance, and consumption potential), such as increasing the urgency weight of rain gear demand from 5% to 35% during heavy rain.
[0050] For example, the secondary scene classification tree may include: Level 1: Daily routines, holidays, and special events; Level 2: Detailed scenarios such as commuting and dating.
[0051] A dual mechanism of rule matching and probability prediction is adopted (rule matching takes priority). Rule matching handles explicit scene descriptions, while probability prediction covers ambiguous text.
[0052] External data (including weather changes, social hot events, etc.) is captured by the scene dynamic perception network and transformed into the scene demand heat value. The scene demand heat value = α × urgency + β × relevance + γ × consumption potential, where α + β + γ = 1, and α, β, and γ are dynamic adjustment coefficients (such as urgency weight can be configured to 30%-50%, relevance weight can be configured to 20%-40%, and consumption potential weight can be configured to 20%-40%). Urgency, relevance, and consumption potential are all mapped to the [0, 1] interval through standardization processing, and the population clustering weight is dynamically adjusted.
[0053] For the gift scenario, an emotion score of 1-5 can be applied (e.g., "thank you" can be scored 4 points) to achieve a quantitative expression of emotional intent. The scoring criteria can be generated based on expert-annotated samples. Alternatively, an emotion tendency probability distribution (e.g., 60% positive, 30% neutral, and 10% negative) can be used to replace the scoring. The probability values can be output through a multimodal emotion perception model (which integrates text and behavioral data).
[0054] For example, in the marketing data graph, the entity layer can include user IDs (Identity Documents), product IDs, scene tags, sentiment tags, etc. Unique entity identifiers can be hashed to ensure the security of sensitive information. The relationship layer can cover behavioral associations such as purchasing, browsing, collecting, and gifting. Relationship weights (1-10 points) are dynamically adjusted based on behavior frequency and timeliness to reflect changes in user interests in real time. The attribute layer can record user characteristics (age, purchasing power), product characteristics (price, style), and scene characteristics (frequency, duration). Attribute values support multi-dimensional expansion (such as adding product "material," "suitable audience," etc.) to comprehensively characterize object features.
[0055] Among them, Neo4j graph database can be used to store triples (such as user A, who frequently appears in commuting scenarios), adapting to the needs of efficient storage and querying of graph structure data.
[0056] Among them, newly added triples can be periodically checked for integrity (including format, entity relationship, etc.) before being stored in the database, and failed data is put into an exception queue for modification and processing.
[0057] Among these features, node weights can be optimized periodically using the PageRank algorithm, and low-association entities with association frequencies less than a preset frequency (such as 10 times) can be removed (such as retaining a 30-day backup to support accidental deletion recovery) in order to simplify the graph structure and ensure data security.
[0058] Among them, the scene knowledge graph can be self-evolved periodically based on graph neural networks to update the scene classification tree and automatically discover new scene association rules (such as making a strong association between "camping" and "sunscreen").
[0059] Through the above embodiments, unstructured data that was originally difficult to utilize (such as fragmented comments and social content) can be transformed into structured features, solving the problem of inefficient utilization of unstructured data in existing technologies and providing more dimensional data support for the construction of accurate user profiles. Dynamic clustering iteration, combined with time decay factors and scene demand heat values, optimizes the clustering effect, solves the problem of insufficient granularity of user profiles, and can more accurately segment user groups and capture the personalized needs of different groups. Especially in fields with fragmented needs and strong scene dependence (such as the women's clothing industry), it can meet the needs of refined operation. Scene tag mapping and emotional intent quantification realize the structured expression of scenes and emotions, lays the foundation for subsequent scene-based marketing and emotional recommendation, improves scene adaptability and emotional resonance, and solves the problem of missing scenes in existing technologies. By establishing close connections between users, products, scenes and emotions through knowledge graphs, it supports efficient relationship query and analysis. At the same time, the self-evolution mechanism of knowledge graphs can incorporate new scenes and new association rules in a timely manner, maintaining the timeliness and integrity of knowledge graphs and providing strong relational data support for intelligent strategy generation.
[0060] In this embodiment, a "consumer gene sequence" can also be constructed for each user, converting user basic information, transaction records, browsing preferences, and other data into feature fragments through a "genetic encoding algorithm." Taking beauty users as an example, fragments such as "shade preference gene" and "skincare efficacy emphasis gene" constitute their consumer gene sequence. A "genetic evolution prediction model" is then used to infer the evolution of user needs based on historical user behavior and current market trends, providing a reference for merchants' product development and marketing strategies.
[0061] Through the above embodiments, the constructed consumer gene sequence and gene evolution prediction model can help businesses plan ahead for demand trends, provide forward-looking guidance for product development and marketing strategy formulation, and enhance the market competitiveness of businesses.
[0062] S14, Match the target scenario in the pre-built strategy generation engine to obtain the target engine.
[0063] In this embodiment, the strategy generation engine may include, but is not limited to, a crowd-scene matching engine, a relationship-emotion dynamic matching engine, etc.
[0064] The audience-scenario matching engine is used to generate basic strategies based on the relationship between users and scenarios in the marketing data graph.
[0065] The relationship-emotion dynamic matching engine is used in the gift scenario, and additional relationship recognition and emotion analysis logic is introduced to improve the targeting of the strategy.
[0066] S15, The marketing data graph is processed using the target engine to obtain a target marketing strategy.
[0067] In this embodiment, the step of processing the marketing data graph using the target engine to obtain the target marketing strategy includes: When the target scenario is the gift scenario, the target engine is activated, and the fusion features are used to traverse the marketing data graph to obtain the target relationship type; The semantic cosine similarity between the product label features and the scene label features in the fused features is calculated by using a BERT bidirectional encoder to obtain the scene fit. The relationship intimacy is calculated based on the historical interaction frequency and relationship duration in the fusion features. Obtain the scene adaptability weight and relationship intimacy weight; The emotional matching degree is obtained by calculating the weighted sum of the scene adaptability and the relationship intimacy based on the scene adaptability weight and the relationship intimacy weight. Obtain a dynamic mapping table of real-time adjusted relationship types and price ranges, and use the target relationship type to match in the dynamic mapping table to obtain the target price range; The target marketing strategy is generated based on the target relationship type, the emotional matching degree, and the target price range.
[0068] For example, when the target scenario is the gift scenario, a three-dimensional recognition system of "text input - behavior history - knowledge graph" is constructed. Named entity recognition technology is used to parse the relational words input by the user (such as "colleague" or "elder"), combined with recipient tags from historical gift-giving records (such as "gifted to A twice a year"), and matched with "user-relationship-interaction frequency" triples in the knowledge graph (such as: user A, colleague relationship, user B, association frequency > 5 times) to comprehensively determine the relationship type.
[0069] This embodiment can also construct a dynamic evolution model of the social relationship lifecycle, and make refined stage divisions for social relationships such as "relatives, colleagues, and friends". For example, colleague relationships can be subdivided into "stranger acquaintance stage, project cooperation stage, daily familiarity stage, and workplace close friend stage", with different recommendation strategy tendencies corresponding to different stages.
[0070] For example, a weighted fusion model can be used to calculate the emotion matching degree, such as Emotion Matching Degree = 0.6 × Scene Adaptability + 0.4 × Relationship Intimacy. Where: Scene fit: The semantic cosine similarity between product tags (such as "high-end gift box") and scene tags (such as "promotion celebration") is calculated using the BERT model. The value range is [0, 1]. The closer the semantics are, the higher the scene fit. Relationship Intimacy: Based on the frequency of historical interactions (average number of gifts given per year × 30% + chat interaction time × 70%), the value range is [0, 1]. The intimacy of close friends is set to >0.7, and that of business partners is set to <0.5. The intimacy parameter is automatically updated according to the duration of the relationship (e.g., the intimacy coefficient of friends who have known each other for more than 3 years is increased by 20%).
[0071] This embodiment can also construct a multimodal emotion perception model, breaking through the limitations of traditional text sentiment analysis. In addition to analyzing textual information such as user reviews and search keywords, it also incorporates user behavior data during the purchase process, such as browsing time (e.g., prolonged browsing is associated with "cautious decision-making"), page jump frequency (e.g., frequent jumps are associated with "hesitation"), and purchase time (e.g., late-night snack purchases are associated with "loneliness"), further improving the accuracy of emotion judgment. For example, in a maternal and infant scenario, if a user browses the "pacifier" product page for a long time in the early morning and frequently compares different brands, combined with their recent search for keywords such as "frequent night awakenings of baby," the model determines that the user is in an "anxious" state, and can then push suitable combination solutions.
[0072] The dynamic mapping table can be adjusted in real time based on the consumer spending characteristics of the industry data platform (such as the average order value of users over the past 3 months). For example: Business Partners: The base price range is 300-500 yuan. If the user's average order value in the past 3 months is greater than 2,000 yuan, the price range will automatically increase by 20% (adjusted to 360-600 yuan). High-end business gift boxes will be recommended first. Close friends: The base price range is 100-300 yuan. If it is a birthday, anniversary or other holiday occasion, the price range will be temporarily extended to a maximum of 500 yuan, and the proportion of personalized customized products recommended will be increased. Relatives (elders): The benchmark price range is 200-400 yuan. Considering the user's spending power, if the average order value is more than 1.5 times the industry average, then we recommend high-end health and wellness gifts, such as quality bird's nest and massage devices.
[0073] Furthermore, a strategy generator can be used to generate the target marketing strategy based on the target relationship type, the sentiment matching degree, and the target price range.
[0074] For example, product combination schemes: In general scenarios, based on user clustering results and scenario tags, the combination logic of "core demand products + related complementary products" is adopted. For example, in the commuting scenario, "anti-wrinkle suit (core) + wrinkle-free shirt (related)" is recommended. In the gift scenario, the combination of "main gift + auxiliary gift + packaging service" is generated by combining relationship type, emotional matching degree and price range. For example, in the business partner promotion scenario, "high-end fountain pen gift box (main gift) + customized thank-you card (auxiliary gift) + business gift box packaging (service)" is recommended. Moreover, the association path of products in the combination must be traceable in the knowledge graph (e.g., user → business gift-giving scenario → high-end demand → fountain pen gift box).
[0075] Customer care initiatives: Dynamically designed based on scenario attributes and user emotional tags. Basic care such as "double membership points and worry-free after-sales returns and exchanges" is provided for general scenarios (such as daily consumption); gift scenarios are designed differently, such as "hiding invoice details and timed delivery" for business gifts, and "handwritten blessings and personalized gift packaging" for gifts for close relationships. In addition, care initiatives need to be linked to emotional matching degree. When the emotional matching degree is >0.8, an extra small gift (such as a customized keychain) is given.
[0076] Channel adaptation recommendations: Based on user reach habits and the urgency of the scenario, channels are recommended. For urgent scenarios (such as temporary gift delivery), "same-city express delivery platforms and offline store self-pickup" are preferred; for non-urgent scenarios (such as daily consumption), "brand APP and e-commerce platform flagship stores" are recommended; for business scenarios, "dedicated customer service on WeChat and customized channels on the official website" are preferred. Channel recommendations should also refer to historical conversion data. For example, if a user's last 3 purchases were all completed through the APP, then APP push notifications should be given priority.
[0077] In the above embodiments, the lack of a scenario-based engine in existing technologies is solved by using scene recognition and a dual-engine adaptation mechanism. This differentiates general and specific scenarios, especially in the gift scenario, where deep analysis of "relationship-emotion" enables precise matching of user needs, addressing the pain point of imbalanced gift scenario adaptation. For example, it can improve the accuracy of business gift recommendations by more than 30%. Through a social relationship lifecycle model and a multimodal emotion perception model, refined relationship segmentation and accurate emotion capture are achieved, making strategy generation more personalized. For instance, user satisfaction with gifts recommended to colleagues during the "workplace friendship period" is 25% higher than with traditional strategies. Dynamic price ranges and differentiated strategy content balance user spending power and scenario needs, improving the acceptance of product recommendations and customer care satisfaction. For example, after dynamically adjusting the price range, the product add-to-cart rate can increase by 18%.
[0078] In this embodiment, the following dynamic strategy adjustments can also be used to adjust marketing strategies, etc.
[0079] (1) Real-time feature feedback: Deploy a streaming computing engine (such as Flink) to incrementally update user behavior feature vectors every hour and dynamically adjust strategies based on feature changes: Short-term feature response: If a user browses a product for more than 5 minutes, the product's priority in the bundled solution will be increased (weight +0.3), and a "detailed explanation of the product" service will be added to customer care; After a user adds a product to their cart, related products will be automatically added to the recommendations (e.g., cufflinks of the same brand will be recommended after adding a tie), and the frequency of channel push will be adjusted (an additional APP reminder will be added within 1 hour after adding the product to the cart). Long-term feature response: If a user views the same scenario product (such as commuter clothing) more than 3 times in the past 7 days, the weight of that scenario tag will be automatically increased by 50%. When generating the strategy, products and care measures for that scenario will be prioritized, such as increasing the recommendation ratio of commuter scenario products and providing exclusive care such as "commuter outfit matching suggestions".
[0080] (2) A / B testing configuration: Stratified sampling is used for multivariate experiments to ensure that the strategy effect is verifiable and optimizable. Experimental groups: Test group 1 consisted of "contextualized copywriting + combined recommendations," with a sample size of ≥2000 people. Evaluation metrics included click-through rate, conversion rate, average order value, and copywriting resonance (the proportion of positive emotional words in comments); Test group 2 consisted of "regular copywriting + single product recommendations," with a sample size of ≥2000 people. Evaluation metrics were the same as for test group 1; The control group consisted of "no recommendations (organic traffic)," with a sample size of ≥2000 people. Evaluation metrics were also click-through rate, conversion rate, and average order value. Testing period and judgment: The testing period is 14 days (including 2 weekends, covering different consumption periods). The significance of the strategy is judged by t test (e.g., p < 0.05, where p value represents the probability that the difference is caused by random factors, p < 0.05 means that the difference is statistically significant). If the conversion rate of test group 1 is more than 15% higher than that of the control group and p < 0.05, then "scenario-based copywriting + combination recommendation" will be determined as the optimal strategy and promoted to all users. Variable iteration: Test variables can be designed for different dimensions, such as copywriting style (pain point type: "Struggling with what to wear to work?" vs. benefit type: "Time-saving and stylish commuting outfits"), recommendation format (combination recommendation vs. single item recommendation), and reach channel (APP push vs. SMS). Optimize strategy details through multiple rounds of testing.
[0081] Through the above embodiments, based on real-time feature feedback and A / B testing mechanisms, it is possible to ensure that the strategy responds to changes in user behavior in a timely manner. At the same time, the effectiveness of the strategy is verified through scientific experiments, avoiding blind adjustments and providing data support for strategy optimization. For example, the click-through rate of scenario-based copy optimized through A / B testing increased by 22% compared to the initial version.
[0082] S16, Generate target marketing data based on the target marketing strategy and the marketing data map.
[0083] In this embodiment, after the target marketing strategy is generated, specific content can be matched in the marketing data graph based on the target marketing strategy to implement the target marketing strategy and generate specific recommended products and related multimodal content.
[0084] For example, the multimodal content may include, but is not limited to: (1) Audience Insight Generation: Based on the clustering results and consumer gene sequences mentioned above, output a three-dimensional profile of "features-needs-scenarios" to clarify the explicit and implicit needs of users. For example, "22-26-year-old startup employees: monthly consumption of 1500-3000 yuan (feature) → need for quick matching of professional attire (explicit need), clothing to enhance workplace confidence (implicit need) → 8 am commute scenario (scenario)", and mark the correlation of implicit needs (such as "workplace confidence" correlation of 0.7) to provide direction for content creation. Audience insights need to be updated weekly, incorporating the latest clustering and gene evolution prediction results to ensure the timeliness of insights; (2) Scene description generation: The scene is described using the four elements of "time-space-action-emotion" to restore the user's real consumption scene and enhance the sense of immersion in the content. Example: "7:30 in the rented room (space), you look through the wardrobe (action), there are only 20 minutes left before work (time), and the anxiety of being reminded by the boss because your clothes are not up to standard surges into your heart (emotion)". The scene description needs to be combined with environmental characteristics (such as weather, holidays). For example, the commuting scene on a rainy day needs to include the detail of "the rain wetting the clothes", and the description content needs to be verified by user research to ensure that it is consistent with the real scene. (3) Marketing copy generation: Combining the AIDA (Attention-Interest-Decision-Action) model to generate copy, the content of each stage needs to match the audience insights and scenario description: Note the following stage: Attract attention by addressing pain points or trending topics, such as "5-Minute Guide to Workplace Attire | A Beginner's Guide to Avoiding Pitfalls"; Interest stage: Highlight the core selling points of the product and match user needs, such as "wrinkle-resistant fabric + versatile color scheme, no wrinkles even after sitting for a long time / no repeats a week"; Desire stage: Connect with the emotional context to stimulate the desire to buy, such as "I was reminded of the wrong outfit yesterday? Today, this outfit will get the boss's approval"; Action phase: Clearly guide conversion, such as "Click to receive a 30 yuan coupon exclusively for new users"; After the copy is generated, it needs to undergo semantic review to ensure that there are no violations and that it is adapted to different channels (e.g., SMS copy should be concise to within 70 characters, while APP copy can be more detailed, etc.). (4) Image Generation: Generate a "product-scenario-user" fusion graph using GAN (Generative Adversarial Networks) to intuitively present the product usage scenario, such as: "A new employee wearing a recommended suit walks in the office corridor, with a clock showing 7:50 in the background, highlighting the 'quick outfit' selling point." The image must conform to the characteristics of the target audience (such as the image of a new employee aged 22-26) and the details of the scene (such as the office environment and commuting time), and must undergo aesthetic review to ensure that the visual effect is attractive to users, while also supporting user customization (such as changing the background color).
[0085] In this embodiment, product recommendations can be made based on a product recommendation engine. For example, the product recommendation engine can make product recommendations in the following ways: (1) Collaborative Filtering Optimization: A "scene similarity" correction factor is introduced to address the limitation of traditional collaborative filtering relying solely on user behavior. Scene similarity is calculated based on the correlation between scene nodes in the knowledge graph. For example, "commuting scene" and "business trip scene" have a similarity of 0.6 and share 30% of the recommendation results; "daily consumption scene" and "holiday gift scene" have a similarity of 0.2 and share only 5% of the recommendation results. At the same time, collaborative filtering needs to incorporate a time decay factor, with the weight of user behavior in the past 3 months being twice that of behavior in the past 3-6 months, to ensure the timeliness of recommendations. (2) Content matching upgrade: The BERT model is used to calculate the semantic similarity between product descriptions, scene tags, and user needs to accurately match products with users. For example, the semantic similarity between the product description "anti-wrinkle, versatile" and the commuting scene tag "convenient, workplace" is 0.82, and the similarity with the user need "quick match" is 0.78, so the product is recommended first; and the content matching needs to cover the product's multi-dimensional attributes (such as material, style, price) to avoid the deviation caused by single-dimensional matching; (3) Real-time adjustment mechanism: A two-stage adjustment logic of "cold start - hot feedback" is set up to solve the problem of recommending new users / new products: Cold start phase: The weight of the first 3 clicks of new users is increased by 30%. Based on the click characteristics, a preliminary user profile is quickly built, and products with "general preferences and common scenarios" are recommended. In the early stage of the launch of new products, a basic recommendation weight (0.8 of the weight of similar products) is given based on the scenario adaptation data of similar products, and gradually adjusted based on user feedback. Heat feedback phase: Recommendations are dynamically adjusted based on real-time behavioral characteristics. For example, if a user clicks on an anti-wrinkle suit, the recommendation ratio of anti-wrinkle products of the same style is immediately increased, while the weight of non-anti-wrinkle products is reduced. (4) Scenario-based bundle recommendation: Generate a three-tiered bundle structure of "core products + related products + emergency products" to meet users' one-stop needs, such as: Commuting scenario combination pack: "Suit (core item, to meet basic dressing needs) + scarf (related item, to enhance the sophistication of the outfit) + portable iron (emergency item, to solve the problem of wrinkles in clothes)"; Gift-giving scenario package: "High-end gift box (core item) + customized greeting card (related item) + gift repair tool (emergency item, such as repairing minor damage to the gift box)"; The bundled package must include a knowledge graph-based connection path (e.g., "user → commuting scenario → anti-wrinkle needs → suit → scarf → portable ironing machine"), and the total price of the bundled products must conform to the price range of the corresponding scenario to avoid prices exceeding user expectations.
[0086] This embodiment can also execute a strategy pre-simulation optimization mechanism, specifically including: (1) Building a virtual market simulation platform: Stress testing of recommendation strategies and content using Monte Carlo simulation algorithms to simulate the effects of strategies under different market environments (such as surges in user traffic and competitor promotions): Risk scenario simulation: Assuming "the user does not adopt the recommended commuting package", the system automatically extrapolates the subsequent consumption path: "unsatisfied with commuting outfit → poor work image → reduced brand repurchase → reduced annual spending", quantifying the potential losses from the strategy's failure (e.g., a single user's annual spending loss of 500 yuan). Optimize the output direction: Identify the weaknesses of the strategy based on the simulation results, such as the low click-through rate of emergency items in the bundle, and then adjust the category of emergency items (e.g., replace the portable iron with anti-wrinkle spray). (2) Construct a dynamic response model of "loss-intervention": Based on the simulation results and actual operational data, set a loss threshold and trigger precise intervention for users with high loss risk: Risk identification: If a user's schedule includes important meetings (extracted from scenario features) and they have not clicked on the commuting package recommendation, they are identified as a high-risk user (potential loss > 300 yuan). Intervention measures: Send a reminder 45 minutes in advance with AR (Augmented Reality) makeup / fitting function, along with a "response kit" recommendation (such as a mini powder compact and a portable scarf), and provide a "30-minute express delivery" service to reduce losses from strategy failure.
[0087] In the above embodiments, multimodal content generation realizes the contextualized presentation of "text + image", solving the problems of single marketing content and weak sense of immersion in the existing technology. For example, the resonance (positive comment ratio) of contextualized copywriting is 40% higher than that of conventional copywriting, and the click-through rate of GAN-generated images is 28% higher than that of traditional product images. The product recommendation engine improves the accuracy of recommendations and user acceptance through multi-dimensional optimization. After collaborative filtering and scenario similarity, the recommendation hit rate can be improved by 35%. Contextualized bundled package recommendations increase the average order value by 50% compared to single-item recommendations, solving the pain point of "users' one-stop needs not being met". The strategy pre-deduction optimization mechanism avoids risks in advance and reduces the loss of strategy failure. For example, the emergency product combination optimized through virtual simulation improves the overall conversion rate of the bundled package by 22%. The "loss-intervention" model has a 60% success rate in intervening with high-risk users, effectively reducing user churn.
[0088] In this embodiment, after generating target marketing data based on the target marketing strategy and the marketing data map, the method further includes: Construct a full-link data monitoring indicator system; wherein, the full-link data monitoring indicator system includes click-through rate, add-to-cart rate, payment conversion rate, user response rate, average order value, and repurchase rate; The target marketing data is evaluated using the aforementioned end-to-end data monitoring indicator system to obtain evaluation results; The target marketing strategy will be adjusted based on the evaluation results.
[0089] The click-through rate (CTR) statistic is the proportion of users clicking on recommended content / products. It needs to be broken down by channel (e.g., APP CTR, SMS CTR) to identify high-conversion channels.
[0090] The payment conversion rate statistics, which show the percentage of users who complete payment after clicking, need to be broken down by scenario (such as commuting scenario conversion rate, gift scenario conversion rate) to identify the shortcomings of the scenario strategy.
[0091] The calculation of the average order value (AOP) for each payment order needs to be linked to recommended combinations (such as the AOP of a bundled package and the AOP of a single item) to evaluate the effectiveness of the bundled recommendations.
[0092] The user response rate can be used to deeply evaluate user experience and strategy adaptability, including average browsing time, number of page jumps, copywriting resonance (the proportion of positive emotional words in comments), and scenario matching (the consistency between user's posting scenario and the recommended scenario).
[0093] The average browsing time is used to reflect the attractiveness of the content; a browsing time of less than 30 seconds indicates that the content needs to be optimized.
[0094] The scenario matching score is used to calculate the consistency ratio between the user's post-show scenario and the recommended scenario. If the matching score is less than 60%, the scenario tag mapping logic needs to be adjusted.
[0095] Among these, the monitoring of each indicator needs to be visualized in real time (e.g., a dashboard can be used for display). If an abnormal indicator (e.g., a sudden drop in conversion rate of 10%) triggers an alarm, the operations staff should be notified in a timely manner to investigate.
[0096] In the above embodiments, the dual evaluation of "conversion effect + user experience" is achieved through the full-link indicator monitoring system, which solves the problem that traditional monitoring only focuses on conversion and ignores user emotions. For example, by discovering and optimizing the scene tag mapping logic through the scene matching degree indicator, the scene matching degree can be increased from 60% to 85%.
[0097] In this embodiment, policy attribution and A / B testing iterations can also be performed. Specifically, this includes: (1) Calculation of strategy attribution coefficient: The strategy attribution coefficient is calculated by distinguishing between "natural conversion" and "strategy-driven conversion" through a causal inference model (such as dual machine learning). Attribution logic: Select a user group similar to the target users but who have not received strategy intervention as a control group, compare the conversion differences between the two groups, and the difference is the strategy-driven conversion; Coefficient application: If the attribution coefficient is >0.5, it indicates that the conversion is mainly driven by the strategy and the strategy needs to be promoted more widely; if the attribution coefficient is <0.2, it indicates that the strategy is weak and iterative optimization is needed. (2) A / B testing iterative optimization: Based on the results of previous tests, continuously optimize test variables and strategies: Variable expansion: In addition to "copywriting style, recommendation format, and channel", new variables such as "price anchor (e.g., 'original price 399, now 299' vs. 'direct discount of 100') and recommendation timing (e.g., 1 hour before work vs. 2 hours after get off work)" have been added. Application of Results: If the test finds that the conversion rate of "pain point-based copywriting" is 20% higher than that of "benefit-based copywriting," then all copywriting across all scenarios will be uniformly adjusted to a pain point-based format. If the conversion rate of the "APP push notification" channel is 35% higher than that of "SMS," then traffic will be prioritized for the APP push notification channel, while reducing the frequency of SMS push notifications (from twice a day to once a day). Test results must form a closed loop of "test report - strategy adjustment - implementation tracking." For example, after adjusting the copywriting style, the conversion rate change must be continuously monitored for 14 days to ensure the stability of the optimization effect. In addition, test parameters will be adjusted according to the characteristics of different industries. For example, in the women's clothing industry, where seasonal demand changes rapidly, the test period can be shortened to 7 days; in the maternal and infant industry, where demand is stable, the test period can be extended to 21 days, thereby ensuring that the test results conform to the actual business scenarios of the industry.
[0098] In the above embodiments, the marketing value is accurately quantified by the strategy attribution coefficient, avoiding overestimation or underestimation of the strategy effect. For example, if the original statistical conversion rate of a certain strategy increased by 25%, after attribution, it was found that the natural conversion rate accounted for 10%, and the actual conversion driven by the strategy was 15%, providing an accurate basis for resource allocation and avoiding ineffective investment.
[0099] This embodiment can also employ the following mechanisms to optimize the models and knowledge graphs involved: (1) Daily fine-tuning (recommendation weight): Based on real-time monitoring data from the previous 24 hours, the recommendation weight is dynamically adjusted to ensure that the strategy responds to immediate changes in demand. Scene weight adjustment: For example, after the "rainstorm" environmental data is triggered, the recommendation weight of related products such as rain gear, waterproof shoe covers, and car interior cleaning products will be automatically increased (from the basic weight of 1.0 to 1.8), while the weight of outdoor sun protection products will be reduced (from 1.0 to 0.5). Product weight adjustment: If the click-through rate of a certain commuter suit suddenly increases by 30% in the first 24 hours and the conversion rate is 20% higher than the category average, its recommendation weight will be increased from 1.2 to 1.5, and it will be displayed in commuter scenarios first. (2) Weekly optimization (clustering and matching model): Based on weekly A / B test results and user behavior trends, update the core algorithm model: Clustering algorithm optimization: If weekly data shows behavioral differences between "high-frequency commuting (more than 5 times per week)" and "low-frequency commuting (1-2 times per week)" in the "commuter group" (high-frequency users pay more attention to anti-wrinkle and portability, while low-frequency users pay more attention to cost-effectiveness), then the original "commuter group" cluster is split, two new sub-clusters are added, and the recommendation strategy is adjusted accordingly. Emotion matching model optimization: If the weekly sentiment index shows that the matching accuracy of the "thank you" sentiment has decreased (from 85% to 75%), then the multimodal sentiment perception model is retrained and supplemented with the latest "thank you" text and behavior samples (such as keywords such as "thank you" and "thank you for your trouble", and behaviors such as "checking thank you cards multiple times") to improve the model accuracy; (3) Monthly upgrades (knowledge graph node relationships): Reconstruct knowledge graph entities, relationships, and attributes, incorporate new scenarios and new association rules, and expand the graph's coverage capabilities: New scenarios included: If monthly monitoring finds that the search volume for products related to "camping scenarios" has increased by 200%, then add the "camping scenario" entity to the knowledge graph, associate it with product entities such as "sunscreen equipment, portable kitchenware, and outdoor tents", and establish triples such as "user-camping scenario-purchase of sunscreen equipment"; Relationship rule update: Based on monthly cross-merchant consumption data, the association rules are updated. For example, if it is found that "80% of users who buy camping tents also buy moisture-proof mats", the relationship weight of "camping tent - associated purchase - moisture-proof mat" will be strengthened in the knowledge graph (from 5 points to 8 points) to guide subsequent combination recommendations. Attribute expansion: Add new attributes to products, such as "sun protection index" and "breathability" for clothing products, and "suitable number of people" and "duration" for scenarios (such as "suitable for 2-3 people" and "duration of 1-2 days" for camping scenarios) to improve matching accuracy.
[0100] Through the above embodiments, the dual-cycle iteration mechanism can ensure continuous adaptation to business changes. Daily fine-tuning allows the strategy to respond to immediate needs (such as a 40% increase in the conversion rate of rain gear recommendations during heavy rain), weekly optimization allows the model to align with user behavior trends (such as a 25% increase in recommendation accuracy after segmenting commuter groups), and monthly upgrades allow the knowledge graph to cover emerging scenarios (such as a 180% increase in sales of related products after the launch of camping scenario recommendations).
[0101] In this embodiment, a strategy evolutionary tree can also be constructed, and gene transfer can be performed.
[0102] Specifically, by constructing a strategy evolution tree, historically effective strategies (such as holiday marketing rhetoric and scenario combination recommendation logic) are encoded into transferable "strategy genes". Each gene contains three parts: "scenario adaptation conditions, strategy content template, and effect parameters". For example, the "Valentine's Day gift box strategy gene" contains "adapted scenario: couples giving gifts / Valentine's Day, content template: romantic copywriting + gift box combination, and effect parameters: conversion rate increased by 30%".
[0103] When a new scenario strategy is generated, it can be quickly adapted through gene recombination. For example, if the "Valentine's Day gift box strategy gene" is migrated to the "520 scenario", the system will automatically adjust: the scenario adaptation condition will be changed from "Valentine's Day" to "520", the emotional intensity of the copywriting will be increased from "romantic" to "sweet", and the price range will be expanded from "200-500 yuan" to "200-800 yuan". At the same time, the core logic of "gift box combination + customized card" will be retained, reducing the development cycle of the new scenario strategy (from 7 days to 1 day).
[0104] In the above embodiments, the development efficiency of new scenario strategies can be greatly improved and the labor cost can be reduced by using strategy evolution trees and gene transfer. For example, the 520 scenario strategy can be quickly implemented through gene transfer, improving development efficiency by 85%, and the effect is on par with historical mature strategies (conversion rate difference < 5%), solving the pain point of "insufficient new scenario operation capabilities" for small and medium-sized businesses.
[0105] This embodiment optimizes the entire closed loop to form a positive cycle of "data feedback - strategy adjustment - model upgrade - effect improvement", which continuously improves recommendation accuracy, operational efficiency and user satisfaction, and ultimately drives substantial growth in business indicators such as repurchase rate and marketing ROI (Return on Investment). For example, after the entire process is optimized, the average repurchase rate of merchants can be increased by 35% and the marketing ROI can be increased by 50%.
[0106] As can be seen from the above technical solutions, this invention can collect and purify marketing-related data based on a three-level data gateway, and use a star-shaped collaborative network built based on dynamic weighting and federated learning mechanisms for feature fusion, thus solving the problems of limited data dimensions and data silos; it can identify scenarios based on a marketing data graph constructed from a two-level scenario classification tree including gift scenarios, thus solving the problems of inefficient use of unstructured data and insufficient granularity of user profiles; it can generate marketing strategies using a target engine that matches scenarios, thus solving the problems of missing scenario-based engines and imbalance in gift scenario adaptation; and it can generate target marketing data based on target marketing strategies and marketing data graphs, thus solving the problems of low operational efficiency and insufficient content accuracy.
[0107] like Figure 2The diagram shown is a functional block diagram of a preferred embodiment of the marketing data generation device based on user profile data of the present invention. The marketing data generation device 11 based on user profile data includes a data acquisition unit 110, a purification unit 111, a fusion unit 112, a recognition unit 113, a matching unit 114, a processing unit 115, and a generation unit 116. The module / unit referred to in this invention refers to a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0108] The data collection unit 110 is used to construct a three-level data gateway including a private domain gateway, an industry domain gateway, and a public domain gateway, and to use the three-level data gateway to collect marketing-related data to obtain initial data. The purification unit 111 is used to purify the initial data to obtain intermediate data; The fusion unit 112 is used to acquire a star-shaped collaborative network constructed based on a dynamic weighting mechanism and a federated learning mechanism, and to use the star-shaped collaborative network to perform feature fusion on the intermediate data to obtain fused features. The identification unit 113 is used to acquire a pre-constructed marketing data map and perform scene recognition on the fused features based on the marketing data map to obtain the target scene; wherein, the marketing data map is constructed based on a secondary scene classification tree including gift scenes; The matching unit 114 is used to perform matching in a pre-built strategy generation engine according to the target scenario to obtain the target engine; The processing unit 115 is used to process the marketing data graph using the target engine to obtain a target marketing strategy; The generation unit 116 is used to generate target marketing data based on the target marketing strategy and the marketing data map.
[0109] As can be seen from the above technical solutions, this invention can collect and purify marketing-related data based on a three-level data gateway, and use a star-shaped collaborative network built based on dynamic weighting and federated learning mechanisms for feature fusion, thus solving the problems of limited data dimensions and data silos; it can identify scenarios based on a marketing data graph constructed from a two-level scenario classification tree including gift scenarios, thus solving the problems of inefficient use of unstructured data and insufficient granularity of user profiles; it can generate marketing strategies using a target engine that matches scenarios, thus solving the problems of missing scenario-based engines and imbalance in gift scenario adaptation; and it can generate target marketing data based on target marketing strategies and marketing data graphs, thus solving the problems of low operational efficiency and insufficient content accuracy.
[0110] like Figure 3The diagram shown is a schematic representation of the computer device used in a preferred embodiment of the marketing data generation method based on profile data according to the present invention.
[0111] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a marketing data generation program based on profile data.
[0112] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0113] It should be noted that all data involved in this case was legally obtained. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A marketing data generation method based on profile data, characterized in that, The marketing data generation method based on profile data includes: A three-tiered data gateway system is constructed, comprising a private domain gateway, an industry domain gateway, and a public domain gateway. Marketing-related data is collected using the three-tiered data gateway system to obtain initial data. The initial data is purified to obtain intermediate data; A star-shaped collaborative network based on a dynamic weighting mechanism and a federated learning mechanism is obtained, and the star-shaped collaborative network is used to perform feature fusion on the intermediate data to obtain fused features; A pre-constructed marketing data map is obtained, and the target scene is obtained by scene recognition of the fused features based on the marketing data map; wherein, the marketing data map is constructed based on a secondary scene classification tree including gift scenes; The target engine is obtained by matching the target scenario in a pre-built strategy generation engine. The marketing data graph is processed using the target engine to obtain a target marketing strategy; Target marketing data is generated based on the target marketing strategy and the marketing data map.
2. The marketing data generation method based on profile data as described in claim 1, characterized in that, The initial data obtained by using the three-level data gateway to collect marketing-related data includes: The private domain gateway is used to synchronize merchant order data, customer data, and private domain interaction data as private domain data at a preset synchronization frequency; wherein, the preset synchronization frequency is dynamically configured according to the data update volume of each party, and an incremental synchronization mechanism is adopted during peak periods; The industry domain gateway collects vertical domain data and cross-merchant consumption association rules as industry domain data according to a first preset cycle. The public domain gateway collects social media trending topics, e-commerce comment sentiment data, and environmental dynamic data as public domain data according to a second preset cycle. The initial data is obtained by combining the private domain data, the industry domain data, and the public domain data.
3. The marketing data generation method based on profile data as described in claim 1, characterized in that, The purification process for the initial data to obtain intermediate data includes: The interquartile range algorithm is used to identify outliers in the initial data, and the outliers are removed from the initial data to obtain the first data. The purchase timestamp in the first data is corrected using a long short-term memory network trained based on an error penalty coefficient to obtain the second data; wherein, during the process of training the long short-term memory network using the error penalty coefficient, when the model output exceeds the multi-dimensional time constraints, the current model is penalized using the error penalty coefficient; Obtain order data generated by the same user within a preset time period as baseline data, and calculate the difference between the second data and the baseline data; When the difference between the second data and the baseline data is greater than a preset threshold, the second data is sent to the verification platform for secondary correction to obtain the third data; The third data is converted into a configuration format to obtain the fourth data; The fourth data is anonymized to obtain the intermediate data; wherein, sensitive personal information is encrypted and de-identified; the user's mobile phone number is encrypted using the MD5 message digest algorithm to obtain the ciphertext of the mobile phone number, and random noise is added to the ciphertext of the mobile phone number; the city-level information in the address information is retained; the SHA-256 hash value of the user's email address is calculated, and the SHA-256 hash value is masked.
4. The marketing data generation method based on profile data as described in claim 1, characterized in that, Before obtaining the star-shaped collaborative network constructed based on the dynamic weight mechanism and the federated learning mechanism, the method further includes: The star-shaped collaborative network, comprising edge nodes and a central node, is constructed based on the dynamic weighting mechanism and the federated learning mechanism. Each merchant is identified as an edge node. The central node includes a model parameter aggregation center, a homomorphic encryption module, and a dynamic weight adjustment engine. The weight dynamic adjustment engine is used to configure basic weights based on data source relevance and data quality. The weight dynamic adjustment engine adjusts the dynamic correction coefficient according to data freshness and domain relevance at preset time intervals, calculates the adjustment ratio by summing the adjusted dynamic correction coefficient with 1, calculates the product of the adjustment ratio and the basic weight to obtain the dynamic weight, and distributes the dynamic weight to each edge node. The adjustment ratio is less than or equal to the configured ratio. During the training of the star-shaped collaborative network, each edge node receives dynamic weights from the dynamic weight adjustment engine and fuses local data according to the dynamic weights to train a local model. The integrity of the model parameters of each local model is then verified, and the verified model parameters are uploaded to the model parameter aggregation center. The local data has the same data dimension as the initial data, and the local data also includes structured feature vectors of spatiotemporal data. These structured feature vectors are used to fuse user behavior features from the local data with an attention mechanism to form a basic model for scene prediction. The model parameter aggregation center is used to aggregate the uploaded model parameters to obtain aggregated parameters. The homomorphic encryption module uses the national standard SM4 block cipher algorithm to encrypt and distribute the aggregation parameters to each edge node.
5. The marketing data generation method based on profile data as described in claim 1, characterized in that, Before acquiring the pre-built marketing data map, the method further includes: The three-level data gateway is used to collect data according to the data dimensions of the initial data to obtain the data to be processed. The data to be processed is segmented into words and tagged with parts of speech, and stop words are removed from the data to obtain the text preprocessing result; Based on the BERT-wwm pre-trained model, deep semantic analysis is performed on the social unstructured text in the text preprocessing results to obtain the core feature vector; Identify high-frequency scene words in the core feature vector and construct a scene feature library using the high-frequency scene words; The core feature vector is classified into sentiment tendencies to obtain the sentiment feature labels of the core feature vector; wherein, a reason sub-label is added to the data whose sentiment feature labels are negative sentiment. Named entity recognition technology is used to identify gift-related context words in the core feature vector, and a gift feature subset is constructed using these gift-related context words; wherein, the gift-related context words include relation words and event words; Calculate the mutual information entropy between the core feature vectors, and filter candidate features from the core feature vectors based on the mutual information entropy; A time decay factor is introduced into the density-based noisy spatial clustering method, and the clustering optimization of the selected candidate features is combined with the scene demand heat value to obtain clustered features. In the clustering process, the effect of different numbers of clusters is evaluated by the silhouette coefficient to select the optimal number of clusters, so that the similarity of samples within the same category is greater than or equal to a first value, and the similarity of samples between different categories is less than or equal to a second value, wherein the first value is greater than the second value. The time decay factor is configured according to the performance of each candidate feature within a specified time period. Based on the core feature vector, determine the urgency, relevance, and consumption potential of each user, obtain the urgency weight, relevance weight, and consumption potential weight, and perform a weighted calculation based on each user's urgency, relevance weight, and consumption potential weight to obtain the scenario demand heat value of each user. A dual mechanism of rule matching and probability prediction is adopted to construct the secondary scene classification tree based on the scene demand popularity value of each user; wherein, rule matching is used first to clarify the scene description, and then probability prediction is used to cover the ambiguous text; wherein, for the gift scene, the emotional intent is quantitatively scored. A knowledge graph is constructed based on the scene feature library, the emotional feature tags, the gift feature subset, the clustering features, and the secondary scene classification tree to obtain the marketing data graph. The marketing data graph includes an entity layer, a relationship layer, and an attribute layer. Each entity in the entity layer is uniquely identified using hash encoding. The relationship weights in the relationship layer are dynamically adjusted based on behavior frequency and timeliness. The attribute layer includes multi-dimensional attribute features for each entity.
6. The marketing data generation method based on profile data as described in claim 1, characterized in that, The process of using the target engine to process the marketing data graph to obtain the target marketing strategy includes: When the target scenario is the gift scenario, the target engine is activated, and the fusion features are used to traverse the marketing data graph to obtain the target relationship type; The semantic cosine similarity between the product label features and the scene label features in the fused features is calculated by using a BERT bidirectional encoder to obtain the scene fit. The relationship intimacy is calculated based on the historical interaction frequency and relationship duration in the fusion features. Obtain the scene adaptability weight and relationship intimacy weight; The emotional matching degree is obtained by calculating the weighted sum of the scene adaptability and the relationship intimacy based on the scene adaptability weight and the relationship intimacy weight. Obtain a dynamic mapping table of real-time adjusted relationship types and price ranges, and use the target relationship type to match in the dynamic mapping table to obtain the target price range; The target marketing strategy is generated based on the target relationship type, the emotional matching degree, and the target price range.
7. The marketing data generation method based on profile data as described in claim 1, characterized in that, After generating the target marketing data based on the target marketing strategy and the marketing data map, the method further includes: Construct a full-link data monitoring indicator system; wherein, the full-link data monitoring indicator system includes click-through rate, add-to-cart rate, payment conversion rate, user response rate, average order value, and repurchase rate; The target marketing data is evaluated using the aforementioned end-to-end data monitoring indicator system to obtain evaluation results; The target marketing strategy will be adjusted based on the evaluation results.
8. A marketing data generation device based on profile data, characterized in that, The marketing data generation device based on profile data includes: The data collection unit is used to construct a three-level data gateway, including a private domain gateway, an industry domain gateway, and a public domain gateway, and to use the three-level data gateway to collect marketing-related data to obtain initial data. A purification unit is used to purify the initial data to obtain intermediate data; The fusion unit is used to acquire a star-shaped collaborative network constructed based on a dynamic weighting mechanism and a federated learning mechanism, and to use the star-shaped collaborative network to perform feature fusion on the intermediate data to obtain fused features; The identification unit is used to acquire a pre-constructed marketing data map and perform scene identification on the fused features based on the marketing data map to obtain the target scene; wherein, the marketing data map is constructed based on a secondary scene classification tree including gift scenes; The matching unit is used to perform matching in a pre-built strategy generation engine according to the target scenario to obtain the target engine; The processing unit is used to process the marketing data graph using the target engine to obtain a target marketing strategy; The generation unit is used to generate target marketing data based on the target marketing strategy and the marketing data map.
9. A computer device, characterized in that, The computer device includes: Memory, storing at least one instruction; and The processor executes instructions stored in the memory to implement the marketing data generation method based on profile data as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the marketing data generation method based on profile data as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Data analysis method for precision marketing based on data management platform
CN114579544A
Large-model intelligent community marketing system
CN120031615A
Industrial graph construction method and system based on scene type marketing
CN120687621A
Marketing decision analysis method based on artificial intelligence
CN120746635A
Marketing strategy optimization management system based on six elements of order transaction
CN120765352A
Cited By
AIGC cross-medium-based meta-universe scene dynamic generation method
CN121614035A
Multi-agent collaborative product closed-loop optimization system and method based on user feedback data
CN121615880A
Product selection decision-making method based on dynamic consumption trend prediction and related products
CN121810332A
A product selection decision method based on dynamic consumption trend prediction and related products
CN121810332B