Activity strategy generation method and device, equipment, storage medium and program product
By combining structural equation modeling and machine learning models in the activity strategy generation process, optimizing the path coefficient set, and automatically filtering target observation variables, the problems of inaccurate strategies and low efficiency in traditional methods are solved, and more efficient activity strategy generation is achieved.
Patent Information
- Application Number
- CN202511604428.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2026-02-10
AI Technical Summary
Traditional campaign strategy development relies on business experience and single metrics, resulting in the underutilization of high-value data. Structural equation modeling assumptions do not align with the distribution of marketing strategies, leading to inaccurate and inefficient strategies.
By identifying candidate observed variables and latent variable sample data, structural equation modeling is used to fit path coefficients, and machine learning modeling is combined to iteratively correct residuals, optimize the path coefficient set, and automatically select the target observed variable generation strategy.
It improves the accuracy and generation efficiency of activity strategies, can accurately estimate path coefficients under non-normal distribution conditions, and automatically optimizes the model without manual adjustment.
Smart Images

Figure CN121504526A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, and in particular relates to an activity strategy generation method, apparatus, device, storage medium and program product. Background Technology
[0002] Traditional campaign strategy development relies heavily on business experience and domain knowledge, or is determined by simple business metrics such as average order value. This approach leads to the underutilization of some high-value data, resulting in less precise campaign strategies.
[0003] In related technologies, to improve the accuracy of activity strategies, structural equation modeling (SEM) is used to analyze the relationships between multiple observed variables used in braking activity strategies, and the activity strategy is formulated based on the analysis results. However, the core assumptions of structural equation modeling include the linear relationship assumption and the normality assumption of the observed variables. In the field of activity strategy design, different observed variables usually do not conform to a multivariate normal distribution, which makes the analysis results of structural equation modeling contain large errors, resulting in the final activity strategy being less accurate. Moreover, traditional structural equation modeling requires manual and repeated adjustments to settings such as paths, which consumes a lot of time when used to generate activity strategies, resulting in low service efficiency. Summary of the Invention
[0004] This application provides an activity strategy generation method, apparatus, device, storage medium, and program product, which can improve the accuracy and generation efficiency of activity strategies.
[0005] In a first aspect, embodiments of this application provide an activity strategy generation method, including: Determine the candidate observation variable sample data and the corresponding latent variable sample data. The candidate observation variable sample data includes the parameter values of at least two candidate observation variables. The candidate observation variables are the factors that affect the activity effect of the activity strategy to be generated. The latent variable sample data includes the parameter values of at least one latent variable. Each latent variable has a business relationship with at least one candidate observation variable. The structural equation model is fitted based on the sample data of candidate observed variables and the sample data of latent variables to obtain the path coefficient set of the structural equation model. The path coefficient set includes the path coefficient of each candidate observed variable in at least two candidate observed variables. The path coefficient characterizes the strength of the relationship between the candidate observed variable and the corresponding latent variable. The residuals of the structural equation model are fitted using the first machine learning model, and the path coefficient set of the structural equation model is iteratively corrected based on the residual fitting results to obtain the target path coefficient set. Based on the target path coefficient set, identify at least one target observation variable that affects the activity effect from at least two candidate observation variables; Generate an activity strategy based on at least one target observation variable.
[0006] Secondly, embodiments of this application provide an activity strategy generation apparatus, comprising: The sample data determination module is used to determine the candidate observation variable sample data and the latent variable sample data corresponding to the candidate observation variable sample data. The candidate observation variable sample data includes the parameter values of at least two candidate observation variables. The candidate observation variables are factors that affect the activity effect of the activity strategy to be generated. The latent variable sample data includes the parameter values of at least one latent variable. Each latent variable has a business relationship with at least one candidate observation variable. The fitting module is used to fit the structural equation model based on the sample data of candidate observed variables and the sample data of latent variables to obtain the path coefficient set of the structural equation model. The path coefficient set includes the path coefficient of each candidate observed variable in at least two candidate observed variables. The path coefficients characterize the strength of the relationship between the candidate observed variable and the corresponding latent variable. The iterative optimization module is used to fit the residuals of the structural equation model through the first machine learning model, and iteratively correct the path coefficient set of the structural equation model based on the residual fitting results to obtain the target path coefficient set. The variable selection module is used to determine at least one target observation variable that affects the activity effect from at least two candidate observation variables based on the target path coefficient set. The strategy generation module is used to generate activity strategies based on at least one target observed variable.
[0007] Thirdly, embodiments of this application provide an electronic device, the electronic device comprising: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the activity policy generation method as described in the first aspect.
[0008] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the activity strategy generation method as described in the first aspect.
[0009] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform the activity strategy generation method as described in the first aspect.
[0010] In this embodiment, during the generation of activity strategies based on structural equation modeling (SEM), a first machine learning model optimizes the SEM fitted to the candidate observed variable sample data and latent variable sample data corresponding to the activity strategy. This ensures that even when the candidate observed variables do not conform to the assumptions of normality and independent and identically distributed distribution of the SEM, the path coefficients of the candidate observed variables can still be accurately estimated. Based on these path coefficients, the target observed variables affecting the activity effect are accurately selected, and the activity strategy is generated based on these target observed variables, thereby improving the accuracy of the final generated activity strategy. Furthermore, the SEM is automatically optimized throughout the entire strategy generation process, eliminating the need for manual adjustments and improving the efficiency of activity strategy generation. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating an activity strategy generation method provided in some embodiments of this application; Figure 2 This is a flowchart illustrating an activity strategy generation method provided in some embodiments of this application; Figure 3 This is a schematic diagram of the structure of an activity strategy generation device provided in some embodiments of this application; Figure 4 These are schematic diagrams of the structure of electronic devices provided in some embodiments of this application. Detailed Implementation
[0013] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0014] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0015] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0016] XGBoost is a high-efficiency gradient boosting decision tree algorithm that improves performance through second-order Taylor expansion and regularization strategies. It supports regression and classification tasks and has stronger optimization capabilities and computational efficiency compared to random forests.
[0017] AIDA Model: A marketing communication process. From exposure to external marketing information to the completion of a purchase, consumers go through four consecutive stages based on their level of response: attention, interest, desire, and action. AIDA is an acronym for the first letters of these four stages.
[0018] Before providing a more detailed description of the embodiments of this application, the activity strategy design methods in related technologies are introduced. Currently, activity strategy design and optimization methods mainly fall into two categories: qualitative and quantitative. The first category is methods for designing campaign strategies based on qualitative business models such as the 4Ps marketing mix model. Examples include the 4Ps marketing mix model, the 7Ps marketing mix model, and the AIDA model. The 4Ps marketing mix model includes four basic elements: Product, Price, Place, and Promotion. The 7Ps marketing mix model adds three more basic elements—People, Processes, and Physical Evidence—to the 4Ps model, used for designing overall marketing strategies. The AIDA model describes the consumer decision-making process, including Attention, Interest, Desire, and Action, providing a clear framework and checkpoints for planning marketing campaigns (especially advertising and content marketing).
[0019] The second category relies on quantitative methods such as average order value (AOV) and structural equation modeling (SEM). Event organizers typically refer to the design and format of similar past events, or base their decisions on the target scenario and AOV level, to identify marketing partners and event intensity. SEM is primarily used to test and estimate complex relationships between variables, especially when dealing with latent variables (such as attitudes, abilities, and other concepts that cannot be directly measured). It is widely used in social sciences, psychology, and economics. In marketing, SEM is applied to explore complex relationships between consumer behavior and advertising effectiveness, effectively identifying causal paths between different variables.
[0020] The second type of quantitative method mentioned above mainly has the following three problems: (1) In terms of data, some high-value data has not been utilized, resulting in inaccurate strategies. At present, the strategy mainly relies on business experience and domain knowledge, or only uses a few data such as the average order value of merchants or industries, while key information such as the intensity, validity period and effect of historical marketing activities are not included in the scope of consideration.
[0021] (2) In terms of models, structural equation modeling assumes linear relationships and the normality of observed variables. This model can better reveal the causal relationships between different variables, but it cannot analyze nonlinear relationships between variables. In the field of marketing strategy design, when the observed variables do not conform to a multivariate normal distribution, the estimation of path coefficients will have a large error. At the same time, traditional structural equation modeling requires manual and repeated adjustment of path, covariance and other settings. If a large number of observed variables are involved, it will consume a lot of time and the service efficiency is low.
[0022] (3) In terms of functionality, current methods or technologies mainly focus on optimizing individual activity configuration items, such as the optimization threshold, or their main purpose is to identify the relationship between activity effects and variables, including activity configuration items. For example, structural equation modeling cannot identify the quantitative relationship between different factors on activity effects and cannot calculate the optimal activity configuration. Overall, there is a lack of an end-to-end technical solution in the field of marketing strategy design.
[0023] In view of this, in order to solve the problems of the above-mentioned quantitative methods, this application provides an activity strategy generation method, apparatus, device, storage medium and program product, which provides better support for scientifically setting activity strategies and improving activity effectiveness.
[0024] The activity strategy generation method provided in this application can be applied to activity strategy design scenarios, such as formulating an electronic coupon distribution strategy. The activity strategy generation method provided in this application will be described in detail below with reference to the accompanying drawings. It should be noted that the execution entity of the activity strategy generation method provided in this application can be an activity strategy generation device. This application uses an activity strategy generation device executing the activity strategy generation method as an example to illustrate the activity strategy generation method provided in this application.
[0025] Figure 1 A flowchart illustrating an embodiment of the activity policy generation method provided in this application is shown. Figure 1 As shown, the method includes the following steps 110-140, which will be explained in detail below.
[0026] Step 110. Determine the candidate observed variable sample data and the corresponding latent variable sample data.
[0027] The candidate observation variable sample data includes parameter values of at least two candidate observation variables, which are factors that affect the effectiveness of the activity strategy to be generated.
[0028] The latent variable sample data includes parameter values for at least one latent variable. Each latent variable has a business relationship with at least one candidate observed variable and affects the effectiveness of the campaign strategy to be generated. For example, in the scenario of developing an e-coupon campaign strategy, the latent variable includes "coupon value," which can be measured through two candidate observed variables, "discount rate" and "discount threshold," and affects the campaign effectiveness indicator "redemption rate." In some embodiments of this application, candidate observed variables and candidate observed variable sample data can be determined based on historical campaign data of historical campaign strategies similar to the current campaign strategy to be generated. For example, if the current campaign strategy to be developed is an e-coupon campaign strategy, then candidate observed variables and their parameter values, i.e., candidate observed variable sample data, can be determined based on historical campaign data of historical e-coupon campaigns.
[0029] In some embodiments of this application, after determining the sample data of candidate observed variables, business modeling is performed. Based on the candidate observed variables and prior knowledge, the business relationship between latent variables and candidate related variables is constructed. Specifically, the AIDA model can be used to determine at least one latent variable that has a business relationship with at least two candidate observed variables. For example, taking the activity strategy to be formulated as an electronic coupon distribution activity as an example, based on the classic AIDA model in the fields of marketing communication and advertising, combined with practical experience, the variables that mainly affect the interests, desires, and actions of potential consumers in the design of the electronic coupon distribution activity can be determined. These variables are represented by four latent variables: acquisition cost, coupon value, willingness to use, and cost of use, and their corresponding candidate observed variables. The parameter value of each latent variable can be measured by the parameter value of the candidate observed variable that has a business relationship with it. Based on this, after determining the latent variable, the parameter value of the latent variable can be determined based on the parameter value of the candidate observed variable that has a business relationship with it in the sample data of the candidate observed variable, thereby obtaining the sample data of the latent variable.
[0030] Step 120. Fit the structural equation model based on the candidate observed variable sample data and the latent variable sample data to obtain the path coefficient set of the structural equation model.
[0031] The path coefficient set of the structural equation model includes the path coefficient of each of the at least two candidate observed variables. The path coefficient represents the strength of the relationship between the candidate observed variable and the corresponding latent variable. The latent variable corresponding to the candidate observed variable refers to the latent variable that has a business relationship with it.
[0032] After obtaining the candidate observed variable sample data and the latent variable sample data, substitute the candidate observed variable sample data and the latent variable sample data into the following structural equation model expressions (2) and (3) to obtain the path coefficient set of the structural equation model.
[0033] Formulas (2) and (3) above are the measurement models for structural equation modeling, where, This represents the candidate observed variables that reflect each latent variable. Indicators representing the effectiveness of the activity and Represents the path coefficient. Representing latent variables, and This represents the regression error of the model.
[0034] Step 130. Using the first machine learning model, fit the residuals of the structural equation model, and iteratively correct the path coefficient set of the structural equation model based on the residual fitting results to obtain the target path coefficient set.
[0035] To improve the ability of structural equation model to fit non-normally distributed variables and enhance the stability of path coefficient estimation, a first machine learning model is used to further fit the residuals of the structural equation model. Based on the fitting results, the path coefficient set of the original structural equation model is iteratively corrected to achieve accurate estimation of path coefficients when the candidate observed variables do not conform to the assumptions of normality and independent and identically distributed distribution of the structural equation model.
[0036] In this embodiment of the application, the iterative correction of the path coefficient set follows the following formula (4): In the above formula (4), This represents the updated set of path coefficients. This represents the set of path coefficients before the update. The learning rate controls the update step size, preventing excessively large updates in a single step. This represents a matrix of candidate observed variables, where each column represents a candidate observed variable and each row represents a sample. for The transpose of the matrix, This represents the residual of the current iteration of the structural equation model. This represents the residual fitting result, (ε'-ε) represents the change in residuals, and (ε'-ε) represents the residuals that can be fitted in the current model. The explanation section. This information will be conveyed through... transpose matrix Feedback to the path coefficient set can correct the path coefficient set, thereby enabling the corrected structural equation model to better fit the data.
[0037] Based on the above formula (4), step 130 may specifically include the following steps 1301-1303.
[0038] Step 1301. Determine the residuals of the structural equation model.
[0039] Based on the candidate observed variable sample data and latent variable sample data used in this iteration, the residuals of the structural equation model in this iteration are calculated. and path coefficient set .
[0040] Among them, the residual of this iteration It can be calculated using the following formula (5).
[0041] In the above formula (5), This represents the residual of the current iteration. This represents the predicted value of the activity performance index obtained through this iteration of the structural equation model. This represents the actual value of the activity performance metrics used in this iteration.
[0042] Step 1302. Fit the residuals using the first machine learning model to obtain the residual fitting results.
[0043] Using the first machine learning model, with the candidate observed variable sample data used in this iteration as input, and the residual of this iteration... Perform residual fitting on the target and obtain the residual fitting result. .
[0044] In some embodiments of this application, the first machine learning model may be the xgboost model.
[0045] Step 1303. Based on the residual fitting results, update the candidate observed variable sample data and return to the step of performing model fitting based on the candidate observed variable sample data and latent variable sample data to obtain the structural equation model, i.e., step 120 above, until the preset first iteration stopping condition is met to obtain the target path coefficient set.
[0046] After obtaining the residual fitting results, the true value of the activity effect index corresponding to the candidate observation variable sample data used in the next iteration can be updated according to the update formula shown in formula (6). Then, the candidate observation variable sample data is updated based on the true value of the activity effect index used in the next iteration, so as to obtain the candidate observation variable sample data used in the next iteration.
[0047] In the above formula (6), This represents the true value of the activity performance metrics set to be used in the next iteration. This represents the predicted value of the activity performance index set obtained from this iteration. This represents the residual fitting result. In this embodiment, the activity performance index is affected by the candidate observed variables; based on this, after obtaining... Afterwards, it can be based on The relationship between activity performance indicators and candidate observed variables is analyzed, and the sample data of candidate observed variables are updated.
[0048] After updating the candidate observed variable sample data, the structural equation model is refitted based on the updated candidate observed variable sample data, and then the path coefficient set is updated until the first iteration stopping condition is reached. The structural equation model obtained by the final iteration update is determined as the enhanced structural equation model, and the path coefficient set of the enhanced structural equation model is determined as the target path coefficient set.
[0049] In some embodiments of this application, the first iteration stopping condition may include at least one of the following: the number of iterations equals a threshold, and the residual of the structural equation model is less than a threshold. The threshold for the number of iterations and the threshold for the residual can be set according to actual needs and are not specifically limited thereto.
[0050] In some embodiments of this application, when the first machine learning model is the xgboost model, overfitting can be avoided by restricting the xgboost tree structure and setting an early stopping strategy.
[0051] The structural equation model is enhanced and optimized based on the first machine learning model. On the one hand, it retains the good interpretability of the traditional structural equation model and clearly describes how different factors affect the activity effect. On the other hand, it can improve the model's fitting ability, fitting accuracy and stability.
[0052] Step 140. Based on the target path coefficient set, determine at least one target observation variable in the candidate observation variable sample data that affects the activity effect.
[0053] The target path coefficient set includes the path coefficients of each of at least two candidate observed variables. Based on this, after obtaining the target path coefficient set, at least two candidate observations can be sorted in descending order of path coefficients to obtain an observed variable sequence. The first N candidate observed variables in the observed variable sequence are taken as the target observed variables. Here, N is a positive integer.
[0054] As mentioned earlier, the path coefficient of a candidate observed variable represents the strength of the relationship between the candidate observed variable and its corresponding latent variable. The stronger the relationship, the greater the influence of the candidate observed variable on its corresponding latent variable, and thus the greater the influence on the activity effect. In this way, the target observed variables that have the greatest impact on the activity effect can be screened out.
[0055] Step 150. Generate an activity strategy based on at least one target observation variable.
[0056] After obtaining at least the target observed variables, the parameter values of the target observed variables are set according to actual needs, and then the activity strategy is generated based on the target observed variables and their parameter values.
[0057] In this embodiment, during the generation of activity strategies based on structural equation modeling (SEM), a first machine learning model optimizes the SEM fitted to the candidate observed variable sample data and latent variable sample data corresponding to the activity strategy. This ensures that even when the candidate observed variables do not conform to the assumptions of normality and independent and identically distributed distribution of the SEM, the path coefficients of the candidate observed variables can still be accurately estimated. Based on these path coefficients, the target observed variables affecting the activity effect are accurately selected, and the activity strategy is generated based on these target observed variables, thereby improving the accuracy of the final generated activity strategy. Furthermore, the SEM is automatically optimized throughout the entire strategy generation process, eliminating the need for manual adjustments and improving the efficiency of activity strategy generation.
[0058] In some embodiments of this application, in step 110 above, the candidate observed variable sample data can be determined through the following steps 1101-1104.
[0059] Step 1101. Obtain historical activity data for historical activity strategies similar to the one to be generated. The historical activity data includes activity configuration variable data and activity performance indicator data.
[0060] Here, the activity configuration variable data includes at least two activity configuration variables and their parameter values, while the activity performance metric data includes at least one activity performance metric and its parameter values. Activity configuration variables refer to configuration items used to generate activity strategies, such as discount strength, discount threshold, and coupon validity period—items that can be configured with parameters. Activity performance metrics refer to metrics used to indicate activity performance, such as redemption rate, number of participating users, number of participating merchants, and leverage ratio.
[0061] Step 1102. Preprocess the historical activity data to remove outliers and missing values, and to convert the historical activity data into a form that the structural equation model can understand.
[0062] In some embodiments of this application, historical activity data may include different types of data. For example, historical activity data may include numerical features such as discount rates, discount thresholds, and ticket validity periods. Historical activity data may also include categorical variables, such as unordered variables like e-coupon type, activity channel, merchant type, and activity time period type, and ordered variables such as discount intensity level, activity scale level, user participation enthusiasm, and merchant cooperation level. Different preprocessing methods can be used for different types of data in historical activity data. Specifically, for numerical features in historical activity data, outliers and missing values can be detected and removed based on the 3sigma principle to improve data accuracy. One-hot encoding is performed on unordered variables in historical activity data to convert them into a form that is easier for the model to understand. Format conversion is performed on ordered variables in historical activity data to convert them into numerical features, thus converting them into a form that is easier for the model to understand.
[0063] Step 1103. Based on the preprocessed historical activity data, determine the candidate observation variable set, which includes at least two candidate observation variables.
[0064] After obtaining the preprocessed historical activity data, the activity configuration variables to be used in the structural equation model can be determined based on the preprocessed historical activity data, and the determined activity configuration variables can be used as candidate observation variables.
[0065] In some embodiments of this application, if the number of activity configuration variables in the preprocessed historical activity data is less than or equal to a quantity threshold, the activity configuration variables in the activity configuration variable data can be directly identified as candidate observation variables. The quantity threshold can be set according to actual needs; for example, the quantity threshold can be a value greater than 5.
[0066] In some embodiments of this application, when the number of activity configuration variables in the preprocessed historical activity data exceeds a threshold, candidate observation variables can be screened through the following steps 11031-11033 to improve subsequent modeling efficiency.
[0067] Step 11031. Using the second machine learning model, fit the fitting relationship between each activity effect index in the activity effect index data and all activity configuration variables in the activity configuration variable data.
[0068] The activity configuration variables in the preprocessed activity configuration variable data are used as independent variables, and the activity effect indicators in the preprocessed activity effect indicator data are used as dependent variables for fitting.
[0069] When the campaign performance metrics data includes multiple metrics, such as redemption rate and leverage ratio, a fitting process is performed for each metric. For each metric, a second machine learning model is used to fit the metric based on the parameter values of all campaign configuration variables in the campaign configuration variable data. This involves training a machine learning model corresponding to that metric, which describes the fitting relationship between the metric and all campaign configuration variables in the campaign configuration variable data.
[0070] In some embodiments of this application, the second machine learning model may be the XGBoost model. Based on this, for each activity performance metric, a machine learning model as shown in equation (1) can be fitted respectively: In the above formula (1), Indicators representing the effectiveness of the activity This indicates the XGBoost model fit. This refers to the activity configuration variables in the activity configuration variable data.
[0071] The XGBoost model can handle both continuous and discrete values and has a good ability to capture nonlinear relationships, thus providing a reliable assessment of feature importance.
[0072] In some embodiments of this application, in addition to the xgboost model, the second machine learning model may also employ a model that can be used for feature importance analysis, such as a Lightweight Gradient Boosting Machine (LightGBM) or a Random Forest.
[0073] Step 11302. Based on the fitted relationship, determine the importance score of each activity configuration variable. The importance score is positively correlated with the degree of influence of the activity configuration variable on the activity effect.
[0074] By fitting the model, a machine learning model can be obtained for each activity performance indicator. For each activity performance indicator, the importance score of each activity configuration variable relative to that indicator can be extracted from the corresponding machine learning model. The importance score represents the degree to which the activity configuration variable is important in predicting the activity performance indicator. Thus, when the activity performance indicator data includes P activity performance indicators, each activity configuration variable can correspond to P importance scores. For each activity configuration variable, a comprehensive score can be determined based on its corresponding P importance scores, and this comprehensive score is determined as the importance score of that activity configuration variable.
[0075] In some embodiments of this application, the weight of each activity effect indicator in the activity effect indicator set can be preset based on actual needs. For each activity configuration variable, the P importance scores corresponding to the activity configuration variable can be weighted and summed based on the weight of each activity effect indicator to obtain the comprehensive score of the activity configuration variable. The comprehensive score is determined as the importance score of the activity configuration variable.
[0076] Step 11033. Select activity configuration variables from the activity configuration variable data whose importance scores meet the score selection criteria to form a candidate observation variable set.
[0077] After obtaining the importance score of each activity configuration variable in the activity configuration variable data, the activity configuration variables in the activity configuration variable data are filtered based on the score filtering conditions, so as to select the activity configuration variables whose importance scores meet the score filtering conditions as candidate observation variables, and the set of candidate observation variables is used as the candidate observation variable sample data.
[0078] In some embodiments of this application, the score filtering criteria may include: the importance score of the activity configuration variable is greater than a score threshold, and the score threshold can be set according to the actual situation. Based on this, in step 11033 above, the importance score of each activity configuration variable can be compared with the score threshold, and the activity configuration variables with an importance score greater than the score threshold can be filtered as candidate observation variables. All the filtered candidate observation variables are then combined into a candidate observation variable set.
[0079] In some other embodiments of this application, the score filtering criteria may include: the activity configuration variables are the top M activity configuration variables in terms of importance. Based on this, in step 11033 above, the activity configuration variables in the activity configuration variable data can be sorted in descending order of importance score to obtain an activity configuration variable sequence. The top M activity configuration variables in the activity configuration variable sequence are selected as candidate observation variables, and all the selected candidate observation variables form a candidate observation variable set. Here, M is a positive integer less than or equal to the quantity threshold.
[0080] Importance scores based on the second machine model can quickly filter observed variables, thereby significantly improving the speed of variable selection and the efficiency of structural equation model construction.
[0081] Step 1104. Construct candidate observation variable sample data based on the parameter values of the candidate observation variable set in the historical activity data.
[0082] By using the above methods to determine candidate observation variable sample data based on historical activity data, we can make full use of activity configuration variable data such as historical activity intensity, ticket validity period, and ticket type, as well as activity effect indicator data such as participating merchants, users, and transactions. This will result in comprehensive and high-quality candidate observation variable sample data, which will help improve the accuracy of structural equation modeling.
[0083] In traditional methods, configuration variable parameter values are typically set manually based on business experience. This approach suffers from low efficiency and accuracy. Therefore, to further improve the accuracy of the final generated strategy and enhance strategy generation efficiency, some embodiments of this application refer to... Figure 2 The above step 150 can be implemented through the following steps 1501-1505.
[0084] Step 1501. Based on all activity performance metrics corresponding to the activity strategy to be generated, determine the comprehensive activity performance metrics.
[0085] The effectiveness of an event strategy can be measured by multiple performance metrics, such as redemption rate, revenue generated, and user participation. The combined parameter values of these metrics reflect the overall effectiveness of the event strategy. Based on this, a comprehensive event performance metric can be determined from these multiple metrics to measure the overall effectiveness of the event strategy. By constructing a comprehensive event performance metric, multiple business objectives can be integrated into a single quantifiable indicator, thereby achieving multi-objective optimization.
[0086] In some embodiments of this application, the weight value of each activity effect indicator can be determined based on prior knowledge, and the comprehensive activity effect indicator can be determined based on multiple activity effect indicators and the weight value of each activity effect indicator according to the following formula (7): In the above formula (7), Indicators representing the overall effectiveness of the activity. Indicates the effectiveness metrics of a single activity. This indicates the weight value of the activity's effectiveness metrics.
[0087] Step 1502. Construct a generalized additive model with at least one target observation variable as input and a comprehensive activity effect index as output.
[0088] The relationship between the target observed variables and the comprehensive activity effect indicators may not be linear. However, the generalized additive model can capture the quantitative relationship between the target observed variables and the comprehensive activity effect indicators, making subsequent optimization more in line with actual business rules.
[0089] Specifically, the expression for the constructed generalized additive model is shown in equation (8): In the above formula (8), g Indicates the link function, This indicates the expected outcome of the overall activity. Represents the intercept term. Indicates the effect on the target observed variable The smoothing function, This indicates the error term.
[0090] Step 1503. Using the generalized additive model as the objective function, construct an optimization problem with the goal of optimizing the value of the comprehensive activity effect index and the range of values of each objective observation variable as the constraint.
[0091] By setting constraints, activity risks can be controlled, extreme situations can be avoided, and the final configuration information can be ensured to be feasible in business.
[0092] In some embodiments of this application, the constraint St is shown in equation (9) below: In the above formula (9), For boundary constraints, it indicates that the range of values for the target observed variable is... ,in This represents the lower limit of the value range. This represents the upper limit of the value range; different target observed variables can correspond to different values. And β, the corresponding target observed variable β can be set according to prior knowledge and actual needs. For example, the discount rate can be set to 10% to 50%, the discount threshold can be set to 10 yuan to 500 yuan, and the validity period of the voucher can be set to 1 day to 90 days, etc. This is an inequality constraint, indicating that the value of each target observation variable is greater than 0.
[0093] Step 1504. Solve the optimization problem to obtain the activity configuration information, which includes the parameter values of each target observation variable.
[0094] In some embodiments of this application, the above optimization problem can be solved by the following steps 210-240.
[0095] Step 210. Generate a value sequence for each target observation variable based on the search range and step size of each target observation variable.
[0096] The range of values for the target observed variable can be directly obtained from the constraints.
[0097] In some embodiments of this application, for continuous target observation variables, the step size of the target observation variable can be determined according to the type of the target observation variable and the business accuracy requirements. The step size refers to the size of the interval in which each parameter is divided in the continuous parameter space. The smaller the step size, the higher the accuracy, but the greater the computational load. The value sequence of the target observation variable is generated based on the step size and value range of the target observation variable. For example, if the target observation variable is a discount rate with a value range of [10%, 50%], and its corresponding step size is set to 5%, the resulting value sequence may include values such as 10%, 15%, 20%, 25%, ..., 50%, and the interval between two adjacent values is the step size of 5%.
[0098] In some embodiments of this application, for discrete target observation variables, the corresponding value sequence can be set by enumeration based on their corresponding value range.
[0099] Step 220. Based on the value sequence of all target observation variables in the target observation variable set, generate a value grid, where a point in the value grid represents a combination of values of all target observation variables.
[0100] The values in the sequence of values of all target observed variables are arranged and combined to obtain all possible value combinations. A value grid is constructed based on all value combinations, and a point in the grid represents a value combination.
[0101] Step 230. For each point in the value grid that satisfies the constraints, calculate the comprehensive activity effect index value corresponding to the point based on the objective function.
[0102] For each point in the value grid, first check whether it meets the constraints. If it does, calculate the objective function value, i.e., the comprehensive activity effect index value; otherwise, skip the point.
[0103] Step 240. The combination of values represented by the point with the largest comprehensive activity effect index value among all points in the value grid that meet the constraints is determined as the activity configuration information.
[0104] Compare the objective function values of all points that satisfy the constraints, and select the point that maximizes the comprehensive activity effect index value as the optimal solution, which is used as the activity configuration information.
[0105] By employing grid search in the above manner to solve the optimization problem, all possible configuration combinations can be traversed, avoiding getting trapped in local optima.
[0106] In some embodiments of this application, the optimization problem can be solved by the following steps 310-330.
[0107] Step 310. Construct the Lagrangian function based on the objective function and constraints.
[0108] The purpose of constructing the Lagrange function is to transform a constrained optimization problem into an unconstrained optimization problem, thereby enabling the solution of the aforementioned optimization problem using unconstrained optimization methods. The Lagrange function integrates the constraints into the objective function by introducing Lagrange multipliers. This allows us to find possible extreme points of the original optimization problem by solving for the stationary points of the Lagrange function (e.g., points where the gradient is zero).
[0109] Step 320. Determine a point that satisfies the constraints. The point represents a combination of values for all target observation variables.
[0110] Step 330. Based on the Lagrange function, iteratively update the points until the second iteration stopping condition is met. The combination of values represented by the points obtained from the final iteration is determined as the activity configuration information.
[0111] The stopping condition for the second iteration may include at least one of the following: The gradient norm of the Lagrange function at the updated point is less than the first threshold. The constraint violation is less than the second threshold, where constraint violation refers to the degree to which a candidate solution to the optimization problem does not satisfy the pre-defined constraint conditions; The change in the objective function value between two consecutive iterations is less than the third threshold; The parameter change in two consecutive iterations is less than the fourth threshold; The number of iterations is greater than or equal to the threshold number; The iteration time is greater than or equal to the preset duration.
[0112] During the iterative update process, steps 3301-3309 are executed at each iteration.
[0113] Step 3301. Calculate the function value, first gradient, and Hessian distance of the objective function at the point in this iteration.
[0114] Step 3302. Calculate the second gradient of the constraint at the point in this iteration.
[0115] Step 3303. Based on the first and second gradients, calculate the third gradient of the Lagrangian function at the point in this iteration.
[0116] Step 3304. Based on the function value of the objective function at the point of this iteration, the first gradient, and the Hessian distance, construct a quadratic approximation of the objective function.
[0117] Step 3305. Perform a first-order Taylor expansion of the constraint conditions at the points in this iteration to obtain a linear approximation of the constraint conditions.
[0118] Step 3306. Based on the quadratic approximation of the objective function and the linear approximation of the constraints, construct the subproblem.
[0119] Step 3307. Solve the subproblems to obtain the search direction.
[0120] Step 3308. Determine the search step size in the search direction using a line search algorithm.
[0121] Step 3309. Update the points for this iteration based on the search direction and search step size.
[0122] When updating the points in the current iteration, the search step size can be increased in the search direction of the current point to obtain the updated point.
[0123] Solving optimization problems using the above method can quickly generate activity configuration information that accurately satisfies the constraints.
[0124] Step 1505. Generate activity policies based on activity configuration information.
[0125] By using the above method, the parameter configuration values are transformed into a constrained optimization problem. Compared with the qualitative value selection method that mainly relies on business experience in traditional event planning, the above data-driven quantitative calculation method can more scientifically determine the optimal configuration, reduce subjective bias, and improve the accuracy and efficiency of event strategies.
[0126] This application provides an end-to-end architecture for generating optimal activity strategies. It identifies target observed variables that affect the effectiveness of activities based on an enhanced structural equation model, and obtains the optimal result by fitting the quantitative relationship between the target observed variables and the activity effectiveness indicators based on a generalized additive model. This automatically generates the optimal activity strategy, significantly improving the accuracy and efficiency of activity strategy generation as a whole.
[0127] Based on the activity strategy generation method provided in the above embodiments, this application also provides specific implementations of the activity strategy generation apparatus. Please refer to the following embodiments.
[0128] See Figure 3 The activity strategy generation device 300 provided in this application embodiment includes the following modules: The sample data determination module 301 is used to determine the candidate observation variable sample data and the latent variable sample data corresponding to the candidate observation variable sample data. The candidate observation variable sample data includes the parameter values of at least two candidate observation variables. The candidate observation variables are factors that affect the activity effect of the activity strategy to be generated. The latent variable sample data includes the parameter values of at least one latent variable. Each latent variable has a business relationship with at least one candidate observation variable. The fitting module 302 is used to fit the structural equation model based on the sample data of the candidate observed variables and the sample data of the latent variables to obtain the path coefficient set of the structural equation model. The path coefficient set includes the path coefficient of each candidate observed variable in at least two candidate observed variables. The path coefficient characterizes the strength of the relationship between the candidate observed variable and the corresponding latent variable. The iterative optimization module 303 is used to fit the residuals of the structural equation model through the first machine learning model, and iteratively correct the path coefficient set of the structural equation model based on the residual fitting results to obtain the target path coefficient set. The variable selection module 304 is used to determine, based on the target path coefficient set, at least one target observation variable that affects the effect of the activity from at least two candidate observation variables. The strategy generation module 305 is used to generate an activity strategy based on at least one target observation variable.
[0129] In this embodiment, during the generation of activity strategies based on structural equation modeling (SEM), a first machine learning model optimizes the SEM fitted to the candidate observed variable sample data and latent variable sample data corresponding to the activity strategy. This ensures that even when the candidate observed variables do not conform to the assumptions of normality and independent and identically distributed distribution of the SEM, the path coefficients of the candidate observed variables can still be accurately estimated. Based on these path coefficients, the target observed variables affecting the activity effect are accurately selected, and the activity strategy is generated based on these target observed variables, thereby improving the accuracy of the final generated activity strategy. Furthermore, the SEM is automatically optimized throughout the entire strategy generation process, eliminating the need for manual adjustments and improving the efficiency of activity strategy generation.
[0130] In some embodiments, the sample data determination module 301 includes: The data acquisition unit is used to acquire historical activity data of historical activity strategies similar to the activity strategy to be generated. The historical activity data includes activity configuration variable data and activity performance indicator data. The activity configuration variable data includes at least two activity configuration variables and their parameter values. The activity performance indicator data includes at least one activity performance indicator and its parameter values. The preprocessing unit is used to preprocess the historical activity data to remove outliers and missing values from the historical activity data, and to convert the historical activity data into a form that the structural equation model can understand. The candidate observation variable determination unit is used to determine a set of candidate observation variables based on preprocessed historical activity data, wherein the set of candidate observation variables includes at least two candidate observation variables. The observation sample data construction unit is used to construct the candidate observation variable sample data based on the parameter values of the candidate observation variable set in the historical activity data.
[0131] In some embodiments, the candidate observation variable determination unit is specifically used for: The second machine learning model is used to fit the relationship between each activity performance index in the activity performance index data and all activity configuration variables in the activity configuration variable data. Based on the fitting relationship, the importance score of each activity configuration variable is determined, and the importance score is positively correlated with the degree of influence of the activity configuration variable on the activity effect; At least two activity configuration variables whose importance scores meet the score filtering criteria are selected from the activity configuration variable data to form the candidate observation variable set.
[0132] In some embodiments, the sample data determination module 301 includes: The latent variable determination unit is used to determine, through the IDA model, at least one latent variable that has a business relationship with the at least two candidate observed variables; A latent sample data determination unit is used to determine the parameter value of each latent variable based on the parameter values of candidate observed variables that have a business relationship with the latent variable.
[0133] In some embodiments, the iterative optimization module 303 includes: A residual determination unit is used to determine the residuals of the structural equation model; The residual fitting unit is used to fit the residuals using the first machine learning model to obtain the residual fitting result. The iterative update unit is used to update the candidate observed variable sample data based on the residual fitting result, and return to the step of fitting the structural equation model based on the candidate observed variable sample data and the latent variable sample data to obtain the path coefficient set of the structural equation model, until the first iteration stopping condition is met to obtain the target path coefficient set.
[0134] In some embodiments, the iterative update unit is specifically used for: Based on the residual fitting results, the true values of the activity effect indicators corresponding to the candidate observation variable sample data used in the next iteration are updated according to the update formula. The candidate observation variable sample data is updated based on the true value of the activity effect index to be used in the next iteration, so as to obtain the candidate observation variable sample data to be used in the next iteration. The update formula includes: in, This indicates the actual value of the activity performance metric to be used in the next iteration. This represents the predicted value of the activity performance indicator obtained from this iteration. This represents the residual fitting result.
[0135] In some embodiments, the variable filtering module 304 is specifically used for: The at least two candidate observed variables are sorted in descending order of their path coefficients to obtain a sequence of observed variables. The candidate observed variables ranked in the top N positions of the observed variable sequence are determined as the target observed variables, where N is a positive integer.
[0136] In some embodiments, the policy generation module 305 includes: The comprehensive indicator determination unit is used to determine the comprehensive activity effect indicator based on all activity effect indicators corresponding to the activity strategy to be generated. The model building unit is used to build a generalized additive model that takes the at least one target observation variable as input and the comprehensive activity effect index as output. The optimization problem construction unit is used to construct an optimization problem with the generalized additive model as the objective function, with the optimal value of the comprehensive activity effect index as the objective and the range of values of each objective observation variable as the constraint. The problem-solving unit is used to solve the optimization problem and obtain activity configuration information, which includes the parameter values of each target observation variable. The strategy generation unit is used to generate activity strategies based on the activity configuration information.
[0137] In some embodiments, the problem-solving unit is specifically used for: Based on the range of values for each target observation variable, determine the value sequence for each target observation variable; Based on the value sequence of all target observed variables, a value grid is generated, where a point in the value grid represents a combination of values of all target observed variables; For each point in the value grid that satisfies the constraints, the comprehensive activity effect index value corresponding to the point is calculated based on the objective function; The combination of values represented by the points in the value grid that satisfy the constraints and have the largest corresponding comprehensive activity effect index value is determined as the activity configuration information.
[0138] In some embodiments, the problem-solving unit is specifically used for: Based on the objective function and the constraints, construct the Lagrange function; Determine a point that satisfies the constraints, where the point represents a combination of values for all target observed variables; Based on the Lagrange function, the points are iteratively updated until the second iteration stopping condition is met. The combination of values represented by the points obtained in the final iteration is determined as the activity configuration information. In each iteration, the following steps are performed: Calculate the function value, first gradient, and Hessian distance of the objective function at the point in this iteration; Calculate the second gradient of the constraint at the point in this iteration; Based on the first gradient and the second gradient, calculate the third gradient of the Lagrange function at the point of this iteration; Based on the function value, first gradient, and Hessian distance of the objective function at the point of this iteration, a second approximation of the objective function is constructed; A first-order Taylor expansion of the constraint at the point in this iteration yields a linear approximation of the constraint. Based on the quadratic approximation of the objective function and the linear approximation of the constraints, a subproblem is constructed; Solve the subproblem to obtain the search direction; The search step size in the search direction is determined using a line search algorithm; Update the points for this iteration based on the search direction and the search step size.
[0139] The activity strategy generation device provided in this application embodiment can achieve... Figures 1 to 2 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0140] Figure 4 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0141] Electronic device 400 may include processor 401 and memory 402 storing computer program instructions.
[0142] Specifically, the processor 401 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0143] Memory 402 may include a large-capacity memory for data or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 402 is non-volatile solid-state memory. Memory 402 may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Thus, typically, memory 402 includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it can perform the operations described in any of the activity policy generation methods in the above embodiments.
[0144] The processor 401 implements any of the activity strategy generation methods in the above embodiments by reading and executing computer program instructions stored in the memory 402.
[0145] In one example, the electronic device 400 may also include a communication interface 403 and a bus 410. For example, Figure 4 As shown, the processor 401, memory 402, and communication interface 403 are connected through bus 410 and complete communication with each other.
[0146] The communication interface 403 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0147] Bus 410 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 410 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0148] Furthermore, in conjunction with the activity policy generation method in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the activity policy generation methods in the above embodiments.
[0149] This application also provides a computer program product, including a computer program, which, when executed, implements any of the activity strategy generation methods described in the above embodiments.
[0150] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0151] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0152] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0153] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0154] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method for generating activity strategies, characterized in that, include: Determine candidate observation variable sample data and corresponding latent variable sample data. The candidate observation variable sample data includes parameter values of at least two candidate observation variables. The candidate observation variables are factors that affect the activity effect of the activity strategy to be generated. The latent variable sample data includes parameter values of at least one latent variable. Each latent variable has a business relationship with at least one candidate observation variable. The structural equation model is fitted based on the candidate observed variable sample data and the latent variable sample data to obtain the path coefficient set of the structural equation model. The path coefficient set includes the path coefficient of each candidate observed variable among the at least two candidate observed variables. The path coefficient characterizes the strength of the relationship between the candidate observed variable and the corresponding latent variable. The residuals of the structural equation model are fitted using a first machine learning model, and the path coefficient set of the structural equation model is iteratively corrected based on the residual fitting results to obtain the target path coefficient set. Based on the target path coefficient set, determine at least one target observation variable that affects the activity effect from the at least two candidate observation variables; An activity strategy is generated based on the at least one target observation variable.
2. The method according to claim 1, characterized in that, Determine the sample data for candidate observed variables, including: Obtain historical activity data of historical activity strategies similar to the activity strategy to be generated. The historical activity data includes activity configuration variable data and activity performance indicator data. The activity configuration variable data includes at least two activity configuration variables and their parameter values. The activity performance indicator data includes at least one activity performance indicator and its parameter value. The historical activity data is preprocessed to remove outliers and missing values, and to convert the historical activity data into a form that the structural equation model can understand. Based on the preprocessed historical activity data, a set of candidate observation variables is determined, which includes at least two candidate observation variables. Based on the parameter values of the candidate observation variable set in the historical activity data, the candidate observation variable sample data is constructed.
3. The method according to claim 2, characterized in that, The process of determining the candidate set of observed variables based on preprocessed historical activity data includes: The second machine learning model is used to fit the relationship between each activity performance index in the activity performance index data and all activity configuration variables in the activity configuration variable data. Based on the fitting relationship, the importance score of each activity configuration variable is determined, and the importance score is positively correlated with the degree of influence of the activity configuration variable on the activity effect; At least two activity configuration variables whose importance scores meet the score filtering criteria are selected from the activity configuration variable data to form the candidate observation variable set.
4. The method according to claim 1, characterized in that, Determining the latent variable sample data corresponding to the candidate observed variable sample data includes: Using the IDA model, at least one latent variable that has a business relationship with the at least two candidate observed variables is identified; For each latent variable, the parameter value of the latent variable is determined based on the parameter values of candidate observed variables that have a business relationship with the latent variable.
5. The method according to claim 1, characterized in that, The step involves fitting the residuals of the structural equation model using a first machine learning model, and iteratively refining the path coefficient set of the structural equation model based on the residual fitting results to obtain the target path coefficient set, including: Determine the residuals of the structural equation model; The residuals are fitted using the first machine learning model to obtain residual fitting results; Based on the residual fitting results, the candidate observed variable sample data is updated, and the process of fitting the structural equation model based on the candidate observed variable sample data and the latent variable sample data to obtain the path coefficient set of the structural equation model is returned until the first iteration stopping condition is met, and the target path coefficient set is obtained.
6. The method according to claim 5, characterized in that, The step of updating the candidate observed variable sample data based on the residual fitting result includes: Based on the residual fitting results, the true values of the activity effect indicators corresponding to the candidate observation variable sample data used in the next iteration are updated according to the update formula. The candidate observation variable sample data is updated based on the true value of the activity effect index to be used in the next iteration, so as to obtain the candidate observation variable sample data to be used in the next iteration. The update formula includes: in, This indicates the actual value of the activity performance metric to be used in the next iteration. This represents the predicted value of the activity performance indicator obtained from this iteration. This represents the residual fitting result.
7. The method according to claim 1, characterized in that, The step of determining at least one target observed variable that affects the activity effect from the at least two candidate observed variables based on the target path coefficient set includes: Based on the target path coefficient set, the at least two candidate observation variables are sorted in descending order of path coefficients to obtain an observation variable sequence; The candidate observed variables ranked in the top N positions of the observed variable sequence are determined as the target observed variables, where N is a positive integer.
8. The method according to any one of claims 1-7, characterized in that, The generation of the activity strategy based on the at least one target observed variable includes: Based on all activity performance metrics corresponding to the activity strategy to be generated, determine the comprehensive activity performance metric. Construct a generalized additive model that takes the at least one target observation variable as input and the comprehensive activity effect index as output; Using the generalized additive model as the objective function, an optimization problem is constructed with the goal of optimizing the index value of the comprehensive activity effect index and the range of values of each objective observation variable as the constraint. Solve the optimization problem to obtain activity configuration information, which includes the parameter values of each target observation variable; Based on the activity configuration information, an activity strategy is generated.
9. The method according to claim 8, characterized in that, Solving the optimization problem yields activity configuration information, including: Based on the range of values for each target observation variable, determine the value sequence for each target observation variable; Based on the value sequence of all target observed variables, a value grid is generated, where a point in the value grid represents a combination of values of all target observed variables; For each point in the value grid that satisfies the constraints, the comprehensive activity effect index value corresponding to the point is calculated based on the objective function; The combination of values represented by the points in the value grid that satisfy the constraints and have the largest corresponding comprehensive activity effect index value is determined as the activity configuration information.
10. The method according to claim 8, characterized in that, Solving the optimization problem yields activity configuration information, including: Based on the objective function and the constraints, construct the Lagrange function; Determine a point that satisfies the constraints, where the point represents a combination of values for all target observed variables; Based on the Lagrange function, the points are iteratively updated until the second iteration stopping condition is met. The combination of values represented by the points obtained in the final iteration is determined as the activity configuration information. In each iteration, the following steps are performed: Calculate the function value, first gradient, and Hessian distance of the objective function at the point in this iteration; Calculate the second gradient of the constraint at the point in this iteration; Based on the first gradient and the second gradient, calculate the third gradient of the Lagrange function at the point of this iteration; Based on the function value, first gradient, and Hessian distance of the objective function at the point of this iteration, a second approximation of the objective function is constructed; A first-order Taylor expansion of the constraint at the point in this iteration yields a linear approximation of the constraint. Based on the quadratic approximation of the objective function and the linear approximation of the constraints, a subproblem is constructed; Solve the subproblem to obtain the search direction; The search step size in the search direction is determined using a line search algorithm; Update the points for this iteration based on the search direction and the search step size.
11. An activity strategy generation device, characterized in that, include: The sample data determination module is used to determine the candidate observed variable sample data and the latent variable sample data corresponding to the candidate observed variable sample data. The candidate observed variable sample data includes the parameter values of at least two candidate observed variables. The candidate observed variables are factors that affect the activity effect of the activity strategy to be generated. The latent variable sample data includes the parameter values of at least one latent variable. Each latent variable has a business relationship with at least one candidate observed variable. The fitting module is used to fit the structural equation model based on the candidate observed variable sample data and the latent variable sample data to obtain the path coefficient set of the structural equation model. The path coefficient set includes the path coefficient of each candidate observed variable among the at least two candidate observed variables, and the path coefficient characterizes the strength of the relationship between the candidate observed variable and the corresponding latent variable. The iterative optimization module is used to fit the residuals of the structural equation model using a first machine learning model, and iteratively correct the path coefficient set of the structural equation model based on the residual fitting results to obtain the target path coefficient set. The variable screening module is used to determine, based on the target path coefficient set, at least one target observation variable that affects the activity effect from among the at least two candidate observation variables; The strategy generation module is used to generate an activity strategy based on the at least one target observation variable.
12. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the activity strategy generation method as described in any one of claims 1-10.
13. A computer storage medium, characterized in that, The computer storage medium stores computer program instructions, which, when executed by a processor, implement the activity strategy generation method as described in any one of claims 1-10.
14. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the activity strategy generation method as described in any one of claims 1-10.