Marketing suggestion generation method and device for target community, equipment and storage medium

By extracting and compressing user feature data from the user profiling system, generating a refined user profile summary, and then embedding it into the LLM model, the token length limitation of massive user data is solved, enabling efficient and accurate marketing strategy generation and improving the quality and efficiency of marketing suggestions.

CN121032533APending Publication Date: 2025-11-28XIAMEN CHANGYUEYUAN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511128401.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

In existing technologies, when massive amounts of user data cannot be directly input into an LLM model to generate marketing strategies, there are issues such as token length limitations, logical breaks, and distraction, leading to unstable quality of marketing recommendations.

Method used

By selecting user feature data streams that conform to preset rules from the user profiling system of the target community, multi-dimensional analysis is performed using data compression operators to generate compressed feature description text fragments, forming user profile summaries. Based on the few-shot prompt word construction method, these summaries are embedded into the LLM large language model to generate structured marketing strategy suggestions.

Benefits of technology

It effectively solves the token length limitation problem, ensuring that LLM can generate accurate marketing strategies, avoiding data truncation and attention loss, and improving the quality and efficiency of marketing suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121032533A_ABST
    Figure CN121032533A_ABST
Patent Text Reader

Abstract

The invention provides a marketing suggestion generation method, device and equipment for a target community and a storage medium, and the method comprises the steps: selecting a full-field user feature data stream of a target user group, which accords with a preset rule, from a user portrait system of the target community, reading a data compression operator combination corresponding to a marketing report type from an operator configuration library, each operator defines a specific data extraction and compression rule. And sequentially executing each data compression operator, performing multi-dimensional analysis and information extraction on the user feature data, generating a compressed feature description text fragment, and compressing the original mass data of millions of tokens to the scale of hundreds of tokens. According to a predefined template structure, all feature description text segments are combined in order to form a user portrait abstract, the user portrait abstract is embedded into a complete cue word template based on a feed-shot cue word construction method, it is ensured that LLM can generate a precise marketing strategy based on complete and refined user information, and the problems of data truncation and attention loss are effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to a marketing suggestion generation method and device for a target community, an equipment and a storage medium. BACKGROUND

[0002] With the rapid development of big data technology and artificial intelligence, enterprise marketing decisions are transforming from traditional experience-driven to data-intelligent-driven. In the prior art, the formulation of marketing strategies uses a large language model (LLM) to assist decision-making. Users manually write prompts and ask LLM for marketing suggestions. Although this method can utilize the knowledge reserve of AI, it has the following significant problems, such as the difference in the ability of different operators to write prompts, which directly affects the output quality. There is a lack of standardized prompt construction method. The context window of LLM is limited, such as DeepSeek R1 supports 64K Tokens (about 50,000 Chinese characters) by default, but in actual use, more than 4000 words may cause logical breakage. The user data of an enterprise is often hundreds of thousands, far exceeding the processing capacity of LLM. Even if the user data is input into LLM in batches, it is easy to cause the model to pay attention to the dispersion, and the output content is low in relevance, and it is difficult to form a systematic marketing suggestion.

[0003] In view of this, the present application is proposed. SUMMARY

[0004] The present application discloses a marketing suggestion generation method and device for a target community, an equipment and a storage medium, which solves the token length limitation problem that massive user data cannot be directly input into an LLM model for marketing strategy generation.

[0005] The first embodiment of the present application provides a marketing suggestion generation method for a target community, comprising: selecting full-field user feature data streams of a target user group meeting a preset rule from a user portrait system of the target community, wherein the user feature data stream includes a user unique identifier, demographic characteristics, behavior characteristics and value characteristics; obtaining a marketing report type, and reading a data compression operator combination corresponding to the marketing report type from a preset operator configuration library, wherein each operator defines a specific data extraction and compression rule; sequentially executing each data compression operator, performing multi-dimensional analysis and information extraction on the user feature data, and generating a compressed feature description text segment; According to a predefined template structure, all feature description text segments are sequentially combined to form a user portrait summary, based on a few-shot prompt word construction method, the user portrait summary is embedded into a complete prompt word template containing role setting, task description and examples, the constructed prompt word is submitted to an LLM large language model, and a structured marketing strategy suggestion for the target user group is obtained. Preferably, the data compression operator includes a statistical operator and a business operator. The statistical operator is used to extract mathematical relationships between user features, including a correlation operator, a distribution operator, and a clustering operator. The business operator is used to extract user group features with business implications, including a value assessment operator, a behavior pattern operator, and a risk assessment operator.

[0006] Preferably, the execution process of the correlation operator includes: Read the specified two continuous feature fields, calculate the Pearson correlation coefficient, and its expression is: wherein, is the user age, is the consumption amount, is the mean of user age, is the mean of consumption amount; According to a predetermined correlation strength mapping table, the correlation coefficient is converted into a natural language description, and a formatted feature description text segment is output.

[0007] Preferably, the execution process of the value assessment operator includes: Based on the RFM model, define a first value user judgment function:

[0008] wherein , , represent the user's i's recent consumption time, consumption frequency and consumption amount, , , are the quantile thresholds of the corresponding dimensions; Calculate the proportion of first value users to generate a feature description text segment.

[0009] Preferably, before the marketing report type is obtained, and a combination of data compression operators corresponding to the marketing report type is read from a preset operator configuration library, the method further includes: Quality assessment is performed on the user feature data stream, and the operator execution strategy is dynamically adjusted based on a data quality scoring mechanism. When the data quality score of a feature field is lower than a preset threshold, the summary operator that depends on the field is automatically disabled.

[0010] Preferably, the step of orderly combining all feature description text fragments according to a predefined template structure to form a user profile summary specifically involves: Semantic deduplication is performed on all feature-descriptive text fragments by calculating text similarity, identifying and merging semantically repetitive segments; The deduplicated feature description text fragments are sorted based on preset importance weights, wherein the weights are dynamically adjusted according to historical marketing performance. Calculate the total number of tokens in the current text fragment set. When the number exceeds the preset limit, remove text fragments in order of importance from low to high until the token limit is met. Insert semantic connectors between the remaining text fragments to form a semantically coherent user profile summary text.

[0011] Preferably, it further includes: Record the complete prompts generated each time and the corresponding LLM output results, and evaluate the quality of the marketing suggestions output by the LLM through automated metrics; Based on the evaluation results, reinforcement learning algorithms are used to adjust the weight parameters of each operator and the text template, and the operator configuration library and prompt word template library are updated regularly.

[0012] A second embodiment of the present invention provides a marketing suggestion generation device for a target community, comprising: The data stream extraction unit is used to select a full-field user feature data stream of the target user group that conforms to preset rules from the user profile system of the target community. The user feature data stream includes user unique identifiers, demographic features, behavioral features and value features. The operator reading unit is used to obtain the marketing report type and read the data compression operator combination corresponding to the marketing report type from the preset operator configuration library, wherein each operator defines specific data extraction and compression rules; The text fragment generation unit is used to execute each data compression operator in sequence, perform multi-dimensional analysis and information extraction on user feature data, and generate compressed feature description text fragments. The suggestion generation unit is used to orderly combine all feature description text fragments according to a predefined template structure to form a user profile summary. Based on the few-shot prompt word construction method, the user profile summary is embedded into a complete prompt word template containing role settings, task descriptions and examples. The constructed prompt words are submitted to the LLM large language model to obtain structured marketing strategy suggestions for the target user group.

[0013] The third embodiment of the present invention provides a marketing suggestion generation device for a target community, including a memory and a processor. The memory stores a computer program, which can be executed by the processor to implement a marketing suggestion generation method for a target community as described in any of the above embodiments.

[0014] The fourth embodiment of the present invention provides a computer-readable storage medium storing a computer program, which can be executed by a processor of the device in which the computer-readable storage medium is located, to implement a marketing suggestion generation method for a target community as described in any of the above embodiments.

[0015] Based on the marketing suggestion generation method, apparatus, device, and storage medium for target communities provided by this invention, a complete dataset containing user unique identifiers, demographic characteristics, behavioral characteristics, and value characteristics is obtained by selecting a full-field user characteristic data stream of the target user group that conforms to preset rules from the user profile system of the target community. Next, a marketing report type is obtained, and a combination of data compression operators corresponding to the marketing report type is read from a preset operator configuration library. Each operator defines specific data extraction and compression rules. Then, by sequentially executing each data compression operator, multi-dimensional analysis and information extraction are performed on the user characteristic data to generate compressed feature description text fragments, compressing the original massive data of millions of tokens to a scale of hundreds of tokens. Finally, according to a predefined template structure, all feature description text fragments are combined in an orderly manner to form a user profile summary. Based on the few-shot prompt word construction method, the user profile summary is embedded into a complete prompt word template containing role settings, task descriptions, and examples, ensuring that LLM can generate accurate marketing strategies based on complete and concise user information, effectively avoiding data truncation and attention loss problems. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating a method for generating marketing suggestions for a target community, provided in the first embodiment of the present invention. Figure 2 This is a schematic diagram of a marketing suggestion generation device for a target community provided in the second embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] To better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0019] It should be understood that the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0020] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0021] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0022] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0023] The terms "first" and "second" used in the embodiments are merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" can be interchanged in a specific order or sequence where permissible. It should be understood that the objects distinguished by "first" and "second" can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein.

[0024] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0025] This invention discloses a method, apparatus, device, and storage medium for generating marketing suggestions for a target community, which solves the problem of token length limitation that prevents the direct input of massive user data into an LLM model for generating marketing strategies.

[0026] Please see Figure 1The first embodiment of the present invention provides a method for generating marketing suggestions for a target community, which can be executed by a marketing suggestion generation device for the target community (hereinafter referred to as the generation device or system), and in particular, by one or more processors within the generation device, to at least implement the following steps: S101, Select a full-field user feature data stream of the target user group that conforms to preset rules from the user profile system of the target community, wherein the user feature data stream includes user unique identifier, demographic features, behavioral features and value features; In this embodiment, the generating device can be a terminal with data processing capabilities, such as a desktop computer, laptop computer, server, or workstation. The generating device can be equipped with a corresponding operating system and application software, and the functions required in this embodiment can be realized through the combination of the operating system and application software. Specifically, in this embodiment, a data connection needs to be established with the enterprise's user profiling system. As the core data asset management platform for an enterprise, the user profiling system typically employs a distributed storage architecture, storing millions or even tens of millions of user data points. When operations personnel need to formulate marketing strategies, they first define the target user group through the selection function in the user profiling system's visual interface. For example, operations personnel can set the filtering rules to "users aged 18-24 who have spent between 100 and 500 yuan on the platform," and the system will filter out a set of users who meet the criteria from the database based on these preset rules.

[0027] After filtering, the system extracts a batch of full-field user characteristic data streams for these target users through a data interface. This data stream is organized in a structured format, with each user corresponding to a complete data record. A unique user identifier serves as the primary key to ensure data uniqueness and traceability, typically using an encrypted user ID or UUID format. Demographic characteristics include basic user attribute information such as age, gender, city, occupation, and education level. This data primarily originates from information provided during user registration and subsequent completed profiles.

[0028] Behavioral characteristics record various user interactions on the platform, including but not limited to login frequency, distribution of active time periods, feature usage preferences, content browsing history, and number of social interactions. Taking dating platforms as an example, behavioral characteristics record detailed multi-dimensional behavioral data such as daily online time, number of messages sent, number of times users viewed other users' profiles, and frequency of participation in topic discussions. This behavioral data is collected in real time through the platform's event tracking system, and after cleaning and aggregation, it is stored in the user profile system.

[0029] Value characteristics primarily reflect a user's commercial value, including metrics such as total historical recharge amount, average monthly spending, payment conversion time, renewal frequency, and average order value. The system also calculates some derived value metrics, such as lifetime value (LTV), willingness to pay score, and price sensitivity level.

[0030] Data extraction dynamically adjusts batch size based on the number of users, typically processing 1000-5000 user records per batch. For large groups containing hundreds of thousands of users, the system automatically performs sharding to ensure the stability and efficiency of data extraction. The extracted raw data is temporarily stored in a cache to provide high-speed data access for subsequent operator processing.

[0031] It's worth noting that the full-field data stream may contain some missing or outlier values, which is common in real-world business scenarios. For example, some users may not have filled in complete personal information, or newly registered users may have limited behavioral data. The system preserves these raw values ​​during data extraction, and subsequent data compression operators perform appropriate processing and labeling to ensure the accuracy of data analysis.

[0032] S102, Obtain the marketing report type, and read the data compression operator combination corresponding to the marketing report type from the preset operator configuration library, wherein each operator defines specific data extraction and compression rules; It should be noted that when the operations staff selects the type of marketing report to be generated in the system interface, the system automatically triggers the operator configuration loading process. In this embodiment, the operations staff selected the "Online Chat Room Activity Marketing Planning" report type. After receiving this selection, the system accesses the pre-built operator configuration library, which stores operator combination schemes for different marketing scenarios in JSON format. The operator configuration library is designed to take into account the diversity of marketing scenarios; each report type corresponds to a set of operator combinations, which are summarized based on historical marketing experience and best practices in data analysis.

[0033] The system reads the configuration file corresponding to "Online Chat Room Activity Marketing Plan" from the configuration library. This configuration file defines in detail the list of data compression operators to be executed and their parameter settings. In this scenario, the configuration includes several predefined summary templates, such as age summary template, consumption summary template, and recharge summary template. Each template is actually a set of related operators, which are arranged in a specific logical order.

[0034] Data compression operators, as core components of the system, are divided into two main categories: statistical operators and business operators. Statistical operators are primarily responsible for mathematically uncovering the intrinsic relationships between user characteristics. Among them, correlation operators are specifically used to analyze the strength of associations between different feature dimensions. When the system executes the age-spending correlation operator, it first extracts the age field and historical spending amount field of all target users from the user feature data stream. Both of these fields are continuous variables, making them suitable for analysis using the Pearson correlation coefficient.

[0035] The internal calculation logic of the operator strictly follows statistical principles. The system first calculates the average age and average spending of all users. Then, for each user, it calculates the difference between their age and the average age, and the difference between their spending and the average spending. By summing the products of these differences and dividing by the product of the standard deviations, the final correlation coefficient r is obtained. In a practical case, when r = 0.62, the system queries a pre-defined correlation strength mapping table. This mapping table defines the interpretation rules for the correlation coefficient: a correlation coefficient with an absolute value less than 0.3 is considered weak, between 0.3 and 0.6 is considered moderate, and greater than or equal to 0.6 is considered strong.

[0036] Based on the lookup results, the correlation operator automatically generates a natural language description: "The Pearson correlation coefficient between user age and consumption is 0.62, showing a strong correlation trend within the reference range." This transformation converts abstract mathematical indicators into textual descriptions that are easy for operations personnel to understand, greatly improving the readability of data analysis results. Similarly, the distribution operator analyzes the distribution characteristics of feature values, such as skewness and kurtosis, while the clustering operator identifies natural clustering patterns in user groups using algorithms such as K-means.

[0037] The business operators are designed to better align with actual marketing needs. The value assessment operator uses the classic RFM model to identify high-value users. The system first calculates the distribution of all users across three dimensions: Recency, Frequency, and Monetary Amount, and determines the 90th quantile for each dimension as a threshold. This threshold selection is based on the Pareto principle, which states that 20% of users contribute 80% of the value. For each user, the operator checks whether all three RFM metrics exceed the corresponding thresholds; only when all three conditions are met simultaneously is the user marked as a high-value user.

[0038] When processing data from tens of thousands of users, the value assessment operator iterates through each user record, executing a decision function. Ultimately, it counts the number of high-value users and calculates their percentage of the total user base. For example, if the calculated percentage of high-value users is 6.8%, the operator generates the descriptive text: "High-value users account for 6.8% of the total user base." This information is of significant reference value for operations personnel in developing differentiated strategies.

[0039] Behavioral pattern operators focus on analyzing user behavior patterns, such as identifying the distribution of users' active time periods and the sequence of function usage preferences. Risk assessment operators assess user churn probability and the risk of declining willingness to pay by building predictive models. The outputs of all these operators follow a unified text format to ensure smooth subsequent summary combination. Through this operator design, the system can flexibly respond to the needs of different marketing scenarios.

[0040] S103, execute each data compression operator in sequence, perform multi-dimensional analysis and information extraction on user feature data, and generate compressed feature description text fragments; It should be noted that after loading the operator configuration, the system begins executing the core data compression process. The operator execution engine sequentially calls each data compression operator to process user feature data according to the order defined in the configuration file. Each operator acts as a specialized information extractor, extracting the most valuable insights from massive amounts of raw data. Taking the user group in this embodiment as an example, the system needs to process data from tens of thousands of users aged 18-24 with spending amounts between 100-500 yuan, and the original data volume may reach millions of tokens.

[0041] When the first operator is executed, it reads the required feature fields from the data stream and performs corresponding calculations and analyses. For example, after the age-consumption correlation operator completes its calculation, it will output a text snippet such as "The Pearson correlation coefficient between a user's age and consumption is 0.62, showing a strong correlation trend within the reference range." Next, the consumption distribution operator might output "User consumption amounts follow a normal distribution, with a mean of 287 yuan and a standard deviation of 95 yuan." The high-value user percentage operator will output "High-value users account for 6.8% of the total number of users." As operators are executed successively, the system will accumulate more and more feature description text snippets.

[0042] However, semantic repetition often exists in these text fragments. For example, different operators may describe similar user characteristics from different perspectives. The system employs advanced text similarity calculation methods to identify these repetitions. Specifically, the system uses a word vector-based cosine similarity algorithm to convert each text fragment into a high-dimensional vector representation, and then calculates the similarity score between the fragments. When the similarity between two text fragments exceeds a threshold of 0.85, the system determines that they contain semantic repetition.

[0043] When semantic redundancy is detected, the system does not simply delete one of them, but intelligently merges them. For example, if one fragment describes "average user spending is 287 yuan" and another fragment describes "user spending is mainly concentrated in the 200-400 yuan range," the system will merge them into "average user spending is 287 yuan, mainly concentrated in the 200-400 yuan range." This merging strategy avoids information redundancy while retaining the key information points of each fragment.

[0044] After deduplication, the system needs to rank all text fragments by importance. Each operator has a base weight value in the system, which is not fixed but dynamically adjusted based on the performance feedback of historical marketing campaigns. The system maintains a weight learning module that tracks the performance metrics of each marketing campaign and analyzes which feature descriptions contribute most to the final marketing results. For example, if historical data shows that the feature "percentage of high-value users" is highly correlated with marketing ROI, then the weight of the operator that generates this description will be increased accordingly.

[0045] In this embodiment, it is assumed that after sorting, "users with high spending value account for 6.8% of the total number of users" receives the highest weight of 0.95, "user age and spending show a strong correlation" has a weight of 0.82, and "user login time is mainly concentrated between 8-11 pm" has a weight of 0.45. The system will reorganize these text fragments in descending order of weight.

[0046] The next crucial step is token count control. The system uses the same tokenizer as the target LLM model to accurately calculate the total number of tokens for all current text segments. In this case, the initially generated text segments may contain 1500 tokens, but to ensure the LLM can effectively handle them and avoid attention mechanism degradation, the system sets an upper limit of 800 tokens. This necessitates intelligent pruning.

[0047] The system removes text fragments one by one, starting with the least weighted fragment, and recalculates the total token count after each fragment is removed. This process continues until the total token count drops below 800. In practice, relatively minor information such as "user registration channels are relatively scattered" or "user avatar completion rate is 45%" may be removed. It is worth noting that the system considers the completeness of information during the pruning process, ensuring that the retained fragments comprehensively reflect the core characteristics of the user group.

[0048] Finally, the system needs to organize these independent text fragments into a coherent summary text. This is not a simple splicing, but requires the addition of appropriate semantic connectives to ensure the fluency of the text. The system predefines a set of connectives, including "at the same time," "in addition," "notably," and "from a consumption perspective," etc. The system selects appropriate connectives based on the semantic relationships between adjacent text fragments.

[0049] For example, the final user profile summary might look like this: "High-spending users account for 6.8% of the total users in this user group. Meanwhile, user age and spending show a strong correlation, with younger users having relatively higher spending power. Behavioral characteristics show that these users are active an average of 18 days per month, indicating high platform stickiness. Furthermore, user payment conversions mainly occur within the first week after registration, suggesting a clear new user payment window." This summary text contains only 702 tokens, but it condenses the most crucial information from the original millions of tokens of user data.

[0050] S104. According to the predefined template structure, all feature description text fragments are combined in an orderly manner to form a user profile summary. Based on the few-shot prompt word construction method, the user profile summary is embedded into a complete prompt word template containing role settings, task descriptions and examples. The constructed prompt words are submitted to the LLM large language model to obtain structured marketing strategy suggestions for the target user group. It's important to note that after generating and optimizing the feature description text fragments, the system enters the crucial stage of prompt word construction. The system first loads a predefined template structure, which is an optimal structure derived from extensive practical experience, maximizing the capabilities of LLM. The template is designed using a few-shot approach and includes four core modules: role setting, user profile summary, task description, and examples. This ensures that the generated prompt words contain sufficient background information, clear task direction, and provide referable examples.

[0051] When the system begins constructing prompts, it first processes the role setting module. Based on the previously selected "Online Chat Room Activity Marketing Planning" report type, the system extracts the corresponding role setting text from the role template library. In this embodiment, the system sets the role as "Senior Dating Project Operation Expert." This role setting is not arbitrary but carefully designed. By giving the LLM a specific professional identity, the knowledge modules related to this field in the model can be effectively activated, making the generated content more professional and closer to actual business scenarios. The role setting text typically includes elements such as identity description, professional background, and work experience, for example: "You are a senior expert with 8 years of experience in operating dating platforms, proficient in user growth, event planning, and ROI optimization." Next, the system embeds the previously generated user profile summary into the second module of the template. This summary has been compressed and optimized, containing refined information from 702 tokens. During embedding, the system adds appropriate guiding text, such as "Based on the following user group analysis:", and then inserts the complete summary content. The summary includes core characteristics of the user group, such as "High-spending users account for 6.8% of the total users in this user group, and user age shows a strong correlation with spending," etc. This information provides the necessary background knowledge for LLM to understand the target user group.

[0052] The task description module is more complex and flexible in its construction. The system allows operations personnel to select or customize marketing goals and constraints through the interface. In this case, the operations personnel selected two preset options: "Cost gradient distribution (low / medium / high cost strategy ratio = 3:5:2)" and "Covering at least 80% of target users". The system converts these selections into standardized task description text: "Based on the above user profile, design a tiered marketing campaign plan. Requirements: 1) Design three levels of strategies according to marketing costs: low, medium, and high, with a quantity ratio of 3:5:2; 2) All strategies must cover at least 80% of target users in total; 3) Each strategy must clearly specify the target subgroup, activity format, expected results, and cost budget." This clear task description helps LLM accurately understand the requirements and avoids generating content that deviates from the objectives.

[0053] The few-shot example module is key to improving the quality of LLM output. The system maintains a library of successful case studies, storing historically effective marketing strategies. When constructing new prompts, the system uses a vector similarity matching algorithm to retrieve 1-2 of the most relevant cases from the library, based on the current user demographics and marketing objectives. For example, the system might find a case study targeting a "recharge promotion for 20-25 year old users," where the input is a similar user profile and the output is a detailed tiered marketing plan. The system will then incorporate the complete input-output pair of this case study into the prompt, serving as a learning example for LLM.

[0054] Once all modules are prepared, the system combines them into a complete prompt in a predetermined order. During the combination process, the system adds appropriate separators and transitional statements between modules to ensure overall coherence. The final generated prompt might have the following structure: "You are a senior expert with 8 years of experience operating a dating platform... [Role Setting]. Based on the following user group analysis: High-spending users account for 6.8% of the total users in this user group... [User Profile Summary]. Based on the above user profile, please design a tiered marketing campaign plan... [Task Description]. Below is a successful case for reference: Input: [Historical Case Input], Output: [Historical Case Output]." Once the suggested words are constructed, they are submitted to the LLM (Large Language Model) via API. The system supports several mainstream LLMs, such as Kimi and DeepSeek, allowing users to choose according to their needs. During submission, the system sets appropriate parameters, such as a temperature coefficient of 0.7 to balance creativity and stability, and a maximum output length of 2000 tokens to ensure policy integrity.

[0055] After receiving prompts, LLM generates detailed marketing strategy suggestions based on its trained knowledge and few-shot examples. Typical outputs include multi-layered strategy proposals, each clearly outlining the target audience, activity mechanisms, expected results, and other elements. For example, a low-cost strategy might suggest "pushing limited-time recharge coupons to users who have registered within the last 7 days, expected to cover 30% of users with an ROI of 3.5"; a medium-cost strategy might include "organizing online social events with a tiered reward system, expected to increase user activity by 50%"; and a high-cost strategy might be "providing exclusive privileges and one-on-one service to high-value users, although only covering 6.8% of users, it could contribute 40% to revenue growth."

[0056] After receiving the LLM output, the system performs structured parsing and formatting, converting the text content into a standard marketing strategy report. Each strategy in the report is extracted as an independent module, containing clear implementation steps, resource requirements, risk warnings, and other practical information. Through this automated process, strategies that previously required marketing experts to spend days analyzing and developing can now be completed in minutes, significantly improving the efficiency of marketing decision-making. According to actual application statistics, after adopting this system, the average decision-making time for operations personnel has been reduced by 82%, and the generated strategies have achieved excellent results in actual implementation. In one possible implementation of the present invention: Before the system executes the data compression operator, a comprehensive quality assessment of the acquired user feature data stream is required. The data quality assessment module is started immediately after the system receives the user feature data stream and uses a multi-dimensional detection mechanism to quantitatively assess the data quality of each feature field.

[0057] The system first performs an integrity check on each feature field. Taking the user data in this embodiment as an example, when the system processes feature data from tens of thousands of users, it counts the number of null values, missing values, and outliers for each field. For the age field, the system checks for null values ​​and outliers exceeding reasonable ranges (e.g., less than 0 or greater than 150). For the spending amount field, in addition to checking for null values, it also identifies negative numbers, extremely large values, and other anomalies. The system calculates the missing rate for each field, which is the percentage of missing values ​​out of the total number of records.

[0058] Building upon integrity checks, the system further performs consistency checks. This includes verifying the logical consistency between related fields; for example, a user's registration time should not be later than their first purchase time, and the total monthly spending amount should equal the total historical spending amount. The system also checks the reasonableness of data distribution, identifying any significant anomalies by calculating statistical indicators such as skewness and kurtosis. For categorical fields, the system checks whether the values ​​are within predefined ranges and whether there are any spelling errors or encoding inconsistencies.

[0059] Based on these multi-dimensional detection results, the system employs a weighted scoring mechanism to calculate a comprehensive data quality score for each feature field. The scoring formula comprehensively considers multiple dimensions, including completeness, consistency, and accuracy, with each dimension having a corresponding weight coefficient. For example, the quality score for a certain field might be calculated as: Quality Score = Completeness Score × 0.4 + Consistency Score × 0.3 + Accuracy Score × 0.3. The completeness score is calculated based on the missing rate; a perfect score of 100 is awarded when the missing rate is 0%, and 20 points are deducted for every 10% increase in the missing rate. The consistency score is calculated based on the proportion of logical errors, and the accuracy score is calculated based on the outlier detection rate.

[0060] In this embodiment, it is assumed that after evaluation, basic user information fields such as age and gender received high quality scores (average above 85 points) because this information must be filled in during user registration and is relatively complete. However, some behavioral characteristic fields, such as "user interest tags," may have more missing data because they are generated through algorithmic inference, resulting in a quality score of only 45 points. Consumption-related fields, due to their involvement in real transactions, typically have higher data quality, scoring around 90 points.

[0061] The system has a preset data quality threshold, typically set to 60 points. This threshold is determined based on extensive historical experience; data fields scoring below this threshold are considered to be of poor quality, and using them for analysis may lead to erroneous conclusions. When the system detects that the quality score of a feature field is below 60 points, it triggers a dynamic adjustment mechanism for the operator execution strategy.

[0062] The core logic of dynamic adjustment is to establish a field dependency graph. During system initialization, the system analyzes the dependency relationship of each operator on feature fields. For example, the age-spending correlation operator depends on the fields of age and spending amount; the user activity operator depends on behavioral fields such as login frequency and online duration. When a field is marked as low quality, the system automatically identifies all operators that depend on that field and marks these operators as "unexecutable".

[0063] Taking the user interest tag field as an example, when its quality score is only 45 points, below the threshold of 60 points, the system will automatically disable operators that rely on this field, such as the "interest preference analysis operator" and the "interest-consumption association operator." However, this disabling is not a simple skipping; the system will generate corresponding alternative descriptive text, such as "Insufficient user interest tag data quality, unable to perform interest preference analysis." This design ensures that even if some operators cannot be executed, the final user profile summary remains complete, and the data limitations are clearly communicated.

[0064] Furthermore, the system implements an intelligent degradation mechanism. When a core operator cannot be executed due to data quality issues, the system will attempt to find alternative solutions. For example, if the direct age field is of poor quality, but the user's registration time and ID number fields are of good quality, the system can estimate the age using the ID number, or use "account age" (registration duration) as a substitute indicator. This flexible approach maximizes the use of available data and avoids the overall analysis results being affected by the quality issues of individual fields.

[0065] In one possible implementation of the present invention, it further includes: Each time the system performs a marketing suggestion generation task, it automatically initiates a full-process recording mechanism. Once the suggestion is constructed and sent to the LLM model, the system stores complete information about this interaction in a dedicated historical database. This database uses structured storage, with each record containing a timestamp, user group identifier, selected report type, execution results of each operator, the generated complete suggestion text, the LLM model version used, and the raw output returned by the model. This detailed recording provides a valuable data foundation for subsequent quality assessment and system optimization.

[0066] In this embodiment, after the system generates marketing suggestions for users aged 18-24 with a spending range of 100-500 yuan, the complete prompts (including role settings, user profile summaries, task descriptions, and examples) and the marketing strategy plan output by the LLM are fully saved. The LLM output typically contains multi-level marketing strategies, each with detailed descriptions of the target audience, activity mechanisms, expected performance metrics, etc. The system not only saves the text content but also records metadata generated during the process, such as API call time, the number of tokens used, and model parameter settings.

[0067] The quality assessment module begins working immediately after the LLM output is saved. The system employs a multi-dimensional automated evaluation index system to quantify the quality of marketing recommendations. The first step is structural integrity assessment. The system parses the LLM output text to check if it contains all the elements required in the task description. For example, if the task requires generating low, medium, and high-level strategies, the system will check if the output indeed contains three types of strategies, and whether each strategy has a clear cost positioning and target audience description.

[0068] Secondly, there's the logical consistency assessment. The system analyzes whether there are logical conflicts between strategies, whether there's unreasonable overlap in target audiences, and whether budget allocation meets preset ratio requirements. For example, if a low-cost strategy claims to cover 60% of users, and a mid-cost strategy covers 50%, resulting in a total coverage exceeding 100%, the system will identify this logical error and deduct points. The system also checks whether the various metrics mentioned in the strategies are within reasonable ranges, such as whether ROI is exaggerated and whether user conversion rates conform to industry norms.

[0069] A deeper quality assessment involves content relevance and originality. The system uses text similarity analysis to determine whether the generated strategy corresponds to the characteristics in the user profile summary. If the user profile shows "6.8% high-value users," but the marketing strategy makes no specific arrangements for this user group, the relevance score will decrease. Originality is assessed by comparing the generated strategy with a historical strategy library; if the generated strategy is highly similar to past cases and lacks originality, the originality score will decrease accordingly.

[0070] The system combines the scores from these dimensions to calculate a total quality score for each generated result. In actual operation, the quality score may present the following distribution: structural integrity 90 points (out of 100), logical consistency 85 points, content relevance 88 points, and originality 75 points, with a weighted average total score of 85.5 points. The system will compare this score with a preset quality benchmark (e.g., 80 points) to determine whether the generated result meets the standard.

[0071] Based on accumulated evaluation data, the system begins executing a reinforcement learning optimization process. The reinforcement learning algorithm treats each marketing suggestion generation as a decision-making process, where operator selection and weight setting are actions, and the final quality score is the reward signal. The system employs an improved Q-learning algorithm, continuously learning and experimenting to find the optimal operator configuration strategy.

[0072] In practice, the system maintains a Q-value table, recording the expected returns of different operator combinations under different states. State definitions include user group characteristics (such as age group, spending level), data quality, etc.; action definitions include which operators to use and the weight settings for each operator. After each task is generated, the system updates the Q-value table based on the quality score. If a high-weighted "age-spending correlation operator" receives a high-quality score, the Q-value of that operator will increase in similar scenarios.

[0073] As experience accumulates, the system gradually learns which operator combinations work best in specific scenarios. For example, for younger users, the combination of "behavioral pattern operators" and "social network operators" may be more effective than traditional RFM analysis; while for high-spending users, the weights of "value assessment operators" and "loyalty analysis operators" should be increased. The system will periodically (e.g., weekly) adjust the default weight settings in the operator configuration library based on updates to the Q-value table.

[0074] Optimizing text templates is equally important. The system analyzes the characteristics of prompt word templates corresponding to high-quality output, identifying which expressions are more likely to guide LLM to generate high-quality content. For example, the system may find that adding a description of "focusing on data-driven decision-making" to the role setting can improve the completeness of the quantitative indicators of the output strategy; explicitly stating "the differentiated needs of different user groups need to be considered" in the task description can improve the granularity of the strategy. These findings will be compiled into template optimization suggestions and updated in the prompt word template library.

[0075] The updates to the operator configuration library and prompt word template library are not simple replacements, but rather employ a version management mechanism. Each update creates a new version, while retaining historical versions for retrospection and comparison. The system also implements an A / B testing mechanism; new configurations are first tested on a small set of tasks, and only after proving that they effectively improve output quality are they fully rolled out. Through this continuous learning and optimization mechanism, the system's performance is constantly improving, and the quality of generated marketing suggestions is steadily increasing, truly achieving a self-evolving intelligent system.

[0076] Please see Figure 2 The second embodiment of the present invention provides a marketing suggestion generation device for a target community, comprising: Data stream extraction form 201 is used to select a full-field user feature data stream of a target user group that conforms to preset rules from the user profile system of the target community. The user feature data stream includes a unique user identifier, demographic features, behavioral features, and value features. The operator reading unit 202 is used to obtain the marketing report type and read the data compression operator combination corresponding to the marketing report type from the preset operator configuration library, wherein each operator defines specific data extraction and compression rules; The text fragment generation unit 203 is used to execute each data compression operator in sequence, perform multi-dimensional analysis and information extraction on user feature data, and generate compressed feature description text fragments. It is suggested that generation unit 204 be used to combine all feature description text fragments in an orderly manner according to a predefined template structure to form a user profile summary. Based on the few-shot prompt word construction method, the user profile summary is embedded into a complete prompt word template containing role settings, task descriptions and examples. The constructed prompt words are then submitted to the LLM large language model to obtain structured marketing strategy suggestions for the target user group.

[0077] The third embodiment of the present invention provides a marketing suggestion generation device for a target community, including a memory and a processor. The memory stores a computer program, which can be executed by the processor to implement a marketing suggestion generation method for a target community as described in any of the above embodiments.

[0078] The fourth embodiment of the present invention provides a computer-readable storage medium storing a computer program, which can be executed by a processor of the device in which the computer-readable storage medium is located, to implement a marketing suggestion generation method for a target community as described in any of the above embodiments.

[0079] Based on the marketing suggestion generation method, apparatus, device, and storage medium for target communities provided by this invention, a complete dataset containing user unique identifiers, demographic characteristics, behavioral characteristics, and value characteristics is obtained by selecting a full-field user characteristic data stream of the target user group that conforms to preset rules from the user profile system of the target community. Next, a marketing report type is obtained, and a combination of data compression operators corresponding to the marketing report type is read from a preset operator configuration library. Each operator defines specific data extraction and compression rules. Then, by sequentially executing each data compression operator, multi-dimensional analysis and information extraction are performed on the user characteristic data to generate compressed feature description text fragments, compressing the original massive data of millions of tokens to a scale of hundreds of tokens. Finally, according to a predefined template structure, all feature description text fragments are combined in an orderly manner to form a user profile summary. Based on the few-shot prompt word construction method, the user profile summary is embedded into a complete prompt word template containing role settings, task descriptions, and examples, ensuring that LLM can generate accurate marketing strategies based on complete and concise user information, effectively avoiding data truncation and attention loss problems.

[0080] Exemplary examples show that the computer program described in the third and fourth embodiments of the present invention can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the device for generating marketing suggestions for a target community. For example, the apparatus described in the second embodiment of the present invention.

[0081] The processor referred to can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the marketing suggestion generation method for a target community, connecting the various parts of the method using various interfaces and lines.

[0082] The memory can be used to store the computer program and / or modules. The processor, by running or executing the computer program and / or modules stored in the memory and calling the data stored in the memory, implements various functions of a marketing suggestion generation method for a target community. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, text conversion function, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, text message data, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0083] If the implemented module is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0084] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0085] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating marketing suggestions for a target community, characterized in that, include: Select a full-field user feature data stream of the target user group that meets the preset rules from the user profiling system of the target community. The user feature data stream includes user unique identifier, demographic features, behavioral features and value features. Obtain the marketing report type, and read the data compression operator combination corresponding to the marketing report type from the preset operator configuration library, wherein each operator defines specific data extraction and compression rules; Each data compression operator is executed sequentially to perform multi-dimensional analysis and information extraction on user feature data, generating compressed feature description text fragments. Following a predefined template structure, all feature description text fragments are sequentially combined to form a user profile summary. Based on the few-shot prompt word construction method, the user profile summary is embedded into a complete prompt word template containing role settings, task descriptions, and examples. The constructed prompt words are then submitted to the LLM large language model to obtain structured marketing strategy suggestions for the target user group.

2. The method for generating marketing suggestions for a target community according to claim 1, characterized in that, The data compression operators include statistical operators and business operators; The statistical operators are used to extract mathematical relationships between user features, including correlation operators, distribution operators, and clustering operators; Business operators are used to extract user group characteristics with business implications, including value assessment operators, behavior pattern operators, and risk assessment operators.

3. The method for generating marketing suggestions for a target community according to claim 2, characterized in that, The execution process of the correlation operator includes: Read two specified continuous feature fields and calculate the Pearson correlation coefficient, the expression of which is: ,in, For user age, For the amount of consumption, The average age of users This represents the average amount spent. Based on the preset correlation strength mapping table, the correlation coefficients are converted into natural language descriptions, and formatted feature description text fragments are output.

4. The method for generating marketing suggestions for a target community according to claim 2, characterized in that, The execution process of the value evaluation operator includes: Define the first-value user determination function based on the RFM model: in , , Let represent the recent purchase time, purchase frequency, and purchase amount of user i, respectively. , , This represents the quantile threshold for the corresponding dimension; Calculate the percentage of users with the highest value and generate a text fragment describing their characteristics.

5. The method for generating marketing suggestions for a target community according to claim 2, characterized in that, Before obtaining the marketing report type and reading the data compression operator combination corresponding to the marketing report type from the preset operator configuration library, the method further includes: The user feature data stream is evaluated for quality, and the operator execution strategy is dynamically adjusted based on the data quality scoring mechanism. When the data quality score of a feature field is lower than a preset threshold, the summary operator that depends on that field is automatically disabled.

6. The method for generating marketing suggestions for a target community according to claim 1, characterized in that, The step involves sequentially combining all feature description text fragments according to a predefined template structure to form a user profile summary, specifically as follows: Semantic deduplication is performed on all feature-descriptive text fragments by calculating text similarity, identifying and merging semantically repetitive segments; The deduplicated feature description text fragments are sorted based on preset importance weights, wherein the weights are dynamically adjusted according to historical marketing performance. Calculate the total number of tokens in the current text fragment set. When the number exceeds the preset limit, remove text fragments in order of importance from low to high until the token limit is met. Insert semantic connectors between the remaining text fragments to form a semantically coherent user profile summary text.

7. The method for generating marketing suggestions for a target community according to claim 1, characterized in that, Also includes: Record the complete prompts generated each time and the corresponding LLM output results, and evaluate the quality of the marketing suggestions output by the LLM through automated metrics; Based on the evaluation results, reinforcement learning algorithms are used to adjust the weight parameters of each operator and the text template, and the operator configuration library and prompt word template library are updated regularly.

8. A marketing suggestion generation device for a target community, characterized in that, include: The data stream extraction unit is used to select a full-field user feature data stream of the target user group that conforms to preset rules from the user profile system of the target community. The user feature data stream includes user unique identifiers, demographic features, behavioral features and value features. The operator reading unit is used to obtain the marketing report type and read the data compression operator combination corresponding to the marketing report type from the preset operator configuration library, wherein each operator defines specific data extraction and compression rules; The text fragment generation unit is used to execute each data compression operator in sequence, perform multi-dimensional analysis and information extraction on user feature data, and generate compressed feature description text fragments. The suggestion generation unit is used to orderly combine all feature description text fragments according to a predefined template structure to form a user profile summary. Based on the few-shot prompt word construction method, the user profile summary is embedded into a complete prompt word template containing role settings, task descriptions and examples. The constructed prompt words are submitted to the LLM large language model to obtain structured marketing strategy suggestions for the target user group.

9. A marketing suggestion generation device for a target community, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program that can be executed by the processor to implement a marketing suggestion generation method for a target community as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The device contains a computer program that can be executed by a processor of the device on which the computer-readable storage medium is located, to implement a marketing suggestion generation method for a target community as described in any one of claims 1 to 7.