SAAS cloud platform-based store member management system
By constructing customer state vectors and multi-objective reinforcement learning modules in the store membership management system of the SaaS cloud platform, personalized promotional strategies are generated, solving the problems of homogenized membership promotions and wasted marketing resources in the membership management system. This enables personalized and dynamic adjustment of promotional strategies, thereby improving customer loyalty and operational efficiency.
Patent Information
- Application Number
- CN202511771440.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-03-20
AI Technical Summary
In existing SaaS cloud platform store membership management systems, membership promotion strategies lack personalization, leading to fatigue among high-value customers and limited activation of low-value customers. It is difficult to effectively balance multiple objectives, resulting in wasted marketing resources and decreased customer loyalty.
The store membership management system, based on a SaaS cloud platform, constructs customer state vectors through data acquisition, state modeling, action generation, multi-objective reinforcement learning, and strategy execution modules. This generates personalized promotional actions, and the strategy is optimized through the multi-objective reinforcement learning module, enabling personalized and dynamic adjustment of promotional strategies.
It enables real-time adjustment of promotional strategies based on customer behavior and environmental changes, reducing customer fatigue, improving the relevance and timeliness of promotional strategies, reducing marketing resource waste, and enhancing customer loyalty and operational efficiency.
Smart Images

Figure CN121707642A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a store membership management system based on a SaaS cloud platform. Background Technology
[0002] Store membership management systems are widely used in the retail industry. Current technologies typically employ SaaS-based cloud platforms for management. These systems, deployed in the cloud, provide unified membership data storage and business processing capabilities for multiple stores. Through multi-tenant databases and unified interface services, they achieve centralized management of basic member information, points accounts, transaction records, and promotional activity data. Enterprises can leverage such platforms to configure membership levels, points rules, and promotional activity templates at headquarters. Each store can then access the cloud system through POS terminals or online channels to perform functions such as member identification, points accumulation and deduction, and coupon issuance and redemption. This reduces the deployment and maintenance costs of localized systems to some extent and supports collaborative operations between online and offline channels.
[0003] Under existing application models, store membership promotion strategies mainly rely on pre-configured uniform templates and static customer segmentation methods. For example, based on historical transaction data, indicators such as RFM are constructed to group customers into several fixed levels, and the same or similar incentive rules such as discounts, spending thresholds, and points rewards are configured for each group. Due to the lack of fine-grained differentiation of real-time customer behavior and individual preferences, customers in the same group often receive the same type and frequency of promotional content. This leads to high-value customers gradually becoming fatigued with discounts or points incentives after frequent exposure to similar promotional information, resulting in a decrease in response rate. Low-value or dormant customers may have limited activation effects due to mismatched discount strength or form. For the system, this promotion management model based on static rules is difficult to automatically adjust discount strategies according to changes in member behavior, which easily leads to a waste of marketing resources and is not conducive to stabilizing the customer base and improving long-term loyalty.
[0004] To address these issues, some solutions have incorporated machine learning algorithms into cloud-based membership management systems to segment customers more finely and attempt to push differentiated promotional content based on preference characteristics. For example, cluster analysis can be used to identify different consumer groups, or collaborative filtering can be used to recommend products or discount combinations. Furthermore, the timing and intensity of push notifications can be adjusted to some extent by incorporating behavioral data such as recent browsing, adding to cart, and store visit frequency. However, these solutions still face limitations in balancing multi-objective optimization (such as short-term conversion rate versus long-term customer value, promotional cost versus discount intensity) and in continuously tracking and responding to individual customer feedback. Consequently, they struggle to effectively and sustainably alleviate the problems of promotional homogenization and customer fatigue in real-world large-scale retail scenarios. Summary of the Invention
[0005] In view of the aforementioned existing problems, the present invention is proposed.
[0006] This invention provides a store membership management system based on a SaaS cloud platform to solve the problems of severe homogenization of promotions and difficulty in balancing short-term conversion and long-term customer value in existing SaaS membership systems, and static grouping.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a store membership management system based on a SaaS cloud platform, deployed on a cloud server and connected to multiple store POS systems via a network. The system includes: The data acquisition module is used to collect multi-source data, including at least customer historical transaction data, customer real-time behavior data, and external environment data. A state modeling module, connected to the data acquisition module, is used to construct a customer state vector representing the current state of the customer based on the multi-source data. An action generation module, connected to the state modeling module, is used to generate at least one personalized promotional action based on the customer state vector. A multi-objective reinforcement learning module, connected to the action generation module, is used to perform reinforcement learning training and strategy optimization on the personalized promotional action based on a multi-objective reward function to obtain an optimized promotional strategy. And a strategy execution module, connected to the multi-objective reinforcement learning module, is used to convert the optimized promotion strategy into a store-executable promotion configuration and distribute it to the multiple store POS systems for execution.
[0008] As a preferred embodiment of the SAAS cloud platform-based store membership management system described in this invention, the customer state vector constructed by the state modeling module includes customer RFM score, real-time behavior sequence and external environment indicators, wherein the real-time behavior sequence includes customer scanning behavior, dwelling behavior and clicking behavior in the store, and the external environment indicators include competitor promotion intensity and weather data.
[0009] As a preferred embodiment of the SaaS cloud platform-based store membership management system described in this invention, the construction of the customer state vector includes: Feature extraction is performed on the real-time behavior sequence to obtain behavioral features that characterize behavior frequency, behavior sequence, and behavior duration; and the external environment indicators are normalized.
[0010] As a preferred embodiment of the SAAS cloud platform-based store membership management system described in this invention, the personalized promotional actions generated by the action generation module include a combination of at least one of the following: discount level, points multiplier, gift options, limited-time purchase qualification, and cross-store benefit redemption options. The personalized promotional actions are used to configure corresponding promotional incentive schemes for individual customers or customer subgroups.
[0011] As a preferred embodiment of the SAAS cloud platform-based store membership management system of the present invention, the multi-objective reinforcement learning module evaluates and updates the personalized promotional actions based on the multi-objective deep Q-network algorithm, and uses a non-dominated ranking mechanism to handle the multi-objective optimization problem.
[0012] As a preferred embodiment of the SAAS cloud platform-based store membership management system described in this invention, the multi-objective reward function includes a weighted combination of short-term sales rewards, customer satisfaction rewards, and long-term customer value rewards. The short-term sales rewards are used to measure the increase in sales revenue generated within a preset time window after a promotion. The customer satisfaction rewards are used to measure customer interaction behavior or evaluation results after a promotion. The long-term customer value rewards are used to measure the expected contribution of customers over a long period.
[0013] As a preferred embodiment of the SAAS cloud platform-based store membership management system described in this invention, it further includes a model update module connected to the multi-objective reinforcement learning module, used to dynamically adjust the weights of short-term sales rewards, customer satisfaction rewards, and long-term customer value rewards in the multi-objective reward function based on customer feedback data.
[0014] As a preferred embodiment of the SAAS cloud platform-based store membership management system of the present invention, the data acquisition module supports a streaming data processing framework, which is used to capture customer behavior events in real time in the form of event streams, and to perform deduplication, anomaly detection and noise filtering on the customer behavior events.
[0015] As a preferred embodiment of the SAAS cloud platform-based store membership management system described in this invention, the system adopts a multi-tenant architecture, with retail enterprises as tenants managing their respective membership data and promotional rules, and allowing multiple stores within the same tenant to share the optimized promotional strategies, enabling customers to enjoy cross-store benefit redemption options and a unified membership discount strategy when making purchases at different stores.
[0016] As a preferred solution of the store membership management system based on the SAAS cloud platform described in the present invention, wherein: the policy execution module is further configured to map the optimized promotion policy into a distribution plan for different channels and time windows, including differential promotion parameter configurations for offline store POS systems, online shopping malls, and mobile applications.
[0017] The beneficial effects of the present invention are as follows: The store membership management system based on the SAAS cloud platform provided by the present invention breaks the homogeneous operation mode of traditional membership systems that rely on static grouping and unified promotion templates by constructing a data collection module, a state modeling module, an action generation module, a multi-objective reinforcement learning module, and a policy execution module in the cloud. The system centrally collects customer historical transactions, in-store real-time behaviors, and external environment data such as competitor promotion intensity and weather under a multi-tenant architecture, and uniformly maps them into customer state vectors that can finely depict individual customer states, achieving a comprehensive representation of customer value, current interests, and environmental factors. On this basis, the present invention abstracts incentive means such as discount strength, integral multiples, gift options, flash sale qualifications, and cross-store privilege redemption into configurable personalized promotion actions, introduces a multi-objective deep Q-network and a non-dominated sorting mechanism, and jointly optimizes three business objectives: short-term sales increase, customer satisfaction, and long-term customer value, avoiding problems such as excessive promotion disturbance or resource waste caused by a single conversion rate orientation. By setting an adjustable multi-objective reward function and its weights, and配合 the model update module to dynamically adjust the weights of each objective according to customer feedback data, the system can adaptively balance between boosting short-term sales and accumulating long-term value at different business stages, making the promotion policy more in line with the current strategic demands of the enterprise. With the help of a streaming data processing framework, customer behavior events can be cleaned in real time in the form of an event stream and incorporated into the reinforcement learning process, enabling the promotion policy to be iteratively updated in response to customer behaviors and environmental changes, thereby reducing the promotion fatigue of high-value customers and improving the relevance and timeliness of preferential offers. At the same time, after the optimized promotion policy is uniformly generated in the cloud, it can be mapped by the policy execution module into differential parameter configurations for multi-store POS, online shopping malls, and mobile applications, supporting cross-store privilege sharing and multi-channel linkage execution while ensuring unified rules, which helps to reduce the local configuration workload of each store, improve the system maintenance efficiency, and form a coherent and consistent membership operation experience across all channels. Brief Description of the Drawings
[0018] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope of the present application.
[0019] Figure 1This is a schematic diagram of the framework of the store membership management system based on the SaaS cloud platform in the embodiment.
[0020] Figure 2 This is a flowchart illustrating the design of a weighted combination of multi-objective rewards in an embodiment.
[0021] Figure 3 This is a flowchart illustrating the process of using a non-dominated ranking mechanism to handle multi-objective optimization problems in the multi-objective reinforcement learning module of the embodiment. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0023] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0024] For example, the terms “first” and “second” used in this application are only used to distinguish and describe similar objects, to differentiate the first object from another object, and are not used to describe a specific order or sequence, nor should they be interpreted as indicating or implying relative importance.
[0025] This application proposes a store membership management system based on a SaaS cloud platform, deployed on a cloud server and connected to multiple store POS systems via a network, combining... Figure 1 As shown, the system includes: The data acquisition module is used to collect multi-source data, including at least customer historical transaction data, customer real-time behavior data, and external environment data. The state modeling module, connected to the data acquisition module, is used to construct a customer state vector that represents the current state of a customer based on multi-source data. The action generation module, connected to the state modeling module, is used to generate at least one personalized promotional action based on the customer's state vector. The multi-objective reinforcement learning module, connected to the action generation module, is used to perform reinforcement learning training and strategy optimization on personalized promotional actions based on the multi-objective reward function, so as to obtain the optimized promotional strategy. And a strategy execution module, which is connected to a multi-objective reinforcement learning module, to convert the optimized promotion strategy into a promotion configuration that can be executed by the store and distribute it to multiple store POS systems for execution; In one embodiment, the customer state vector constructed by the state modeling module includes the customer RFM score, real-time behavior sequence, and external environment indicators. The real-time behavior sequence includes the customer's scanning behavior, dwelling behavior, and clicking behavior in the store, and the external environment indicators include the intensity of competitor promotions and weather data. In this embodiment, the customer's RFM score can be obtained by categorizing and weighting the number of days between the customer's most recent purchase and the current time, the number of purchases within a preset statistical period, and the cumulative purchase amount within that period. Historical transaction data is collected from the store's POS system and online store order records via the network to the cloud data acquisition module. Customer scanning, dwelling, and clicking behaviors in the store can be obtained from cashier scanning records, in-store equipment perception logs of shelf areas, and page browsing and clicking logs in mobile applications or mini-programs, respectively. The intensity of competitor promotions in the external environment indicators can be manually entered by operations personnel in the management interface or synchronized with promotional information from third-party retail data platforms through an interface. Weather data can be automatically obtained at fixed time intervals by establishing an interface with the meteorological service provider. Furthermore, the statistical period for the RFM score can be set monthly or quarterly, typically ranging from 30 to 90 days. The upper limit for the interval between the most recent purchases can be 180 days. The upper limit for the number of purchases can be set based on the high quantile of the company's historical data. The upper limit for the purchase amount can be taken as the high quantile reference value of historical purchase amounts. Values exceeding the upper limit are treated as the upper limit to reduce the impact of extreme values. Similarly, competitor promotion intensity can be calculated as a percentage based on factors such as the average discount level, frequency of activities, or percentage of promotional shelf space over the past 7 to 30 days. Weather data can be categorized into several levels based on temperature ranges, precipitation, and holiday markers, and mapped to preset numerical labels. Optionally, when an external data source is unavailable or the data loss rate exceeds a preset threshold during the current period, the state modeling module can revert to using only historical transaction data and real-time behavioral data to construct a customer state vector. The corresponding external environment indicators can temporarily use the most recently successfully acquired values or the system's global default values to ensure the customer state vector remains available. In the event of a prolonged external data source outage, the system can prompt maintenance personnel to investigate through the management interface and automatically reactivate the data source upon recovery.
[0026] In one embodiment, the construction of a customer state vector includes: Feature extraction is performed on real-time behavior sequences to obtain behavioral features that characterize behavior frequency, behavior sequence, and behavior duration; and external environment indicators are normalized to map external environment data of different dimensions to a unified numerical range. Furthermore, the normalization of external environmental indicators can employ a linear mapping method based on historical statistical intervals. For each external environmental indicator, the minimum and maximum values of its observed values are first calculated within a preset statistical period. The current observed value is then proportionally compressed to between zero and one within this interval. When the current observed value is lower than the historical minimum, it is treated as zero; when it is higher than the historical maximum, it is treated as one, thereby reducing the interference of outliers on the customer state vector. Specifically, the preset statistical period can be set to 30 to 90 days based on business rhythm. When retail enterprises have more stable operational data over a longer period, the interval range can also be recalculated on an annual or semi-annual basis. The system default value can be set to 90 days, and operators can adjust this in the configuration interface. For external environmental indicators with obvious periodicity, such as indicators related to holidays or weekend effects, upper and lower limits can be calculated separately for weekdays and non-working days before normalization. Then, the observed values within each interval are mapped to avoid the normalization result being biased due to distribution differences between different types of dates. Optionally, for external environment indicators with highly skewed distributions, the state modeling module can also preprocess the original data before normalization by truncating or performing logarithmic transformations to improve the numerical stability of the state vector during reinforcement learning training. When the number of valid samples for a certain type of external environment indicator is less than a preset lower limit in the current statistical period, the indicator can be temporarily removed from the state vector, and the customer state vector can be constructed using only the remaining external environment indicators and historical transaction and behavioral features. The indicator will be automatically reintroduced after the sample size is restored.
[0027] In one embodiment, the personalized promotional actions generated by the action generation module include a combination of at least one of discount level, points multiplier, gift options, limited-time purchase qualification and cross-store benefit redemption options. The personalized promotional actions are used to configure corresponding promotional incentive programs for individual customers or customer subgroups. Specifically, discount levels can be pre-divided into several tiers as percentages, for example, multiple selectable tiers can be set in increments of 5% between 5% and 30%. Points multipliers can be divided into integers between 1x and 5x. Gift options can be selected from one or more items available in the current store's gift list. Limited-time purchase eligibility can be defined by setting the promotion start and end times, as well as the maximum number of times each customer or store can participate. Cross-store benefit redemption options can be described by specifying the range of participating stores and the types of benefits available. Furthermore, to reduce the configuration burden on store operations staff, the system can pre-configure several typical promotional action combination templates in the cloud. For example, default discount levels, points multipliers, and gift combinations can be provided for newly registered members, dormant members, and high-value members respectively. Operations staff can fine-tune the values of each tier based on actual operating conditions. In the default configuration, the discount level is no higher than 30%, and the points multiplier is no higher than three times, to prevent excessive impact on the normal pricing system. Optionally, when some stores have not yet enabled cross-store benefits or their inventory does not include items from the gift list, the action generation module can automatically remove relevant options and generate personalized promotional actions only in the dimensions of discounts and points, keeping the action space aligned with the actual business capabilities of the stores. During strategy execution, if feedback is received from stores regarding insufficient inventory or full activity slots, the strategy execution module can automatically disable gift options and limited-time purchase qualifications related to that product or slot, while retaining promotional configurations in other dimensions, ensuring that the promotional strategy can still be executed smoothly in resource-constrained scenarios.
[0028] In one embodiment, the multi-objective reinforcement learning module evaluates and updates personalized promotional actions based on a multi-objective deep Q-network algorithm, and uses a non-dominated ranking mechanism to handle multi-objective optimization problems, so as to obtain a set of compromise candidate promotional strategies among multiple conflicting optimization objectives. Reference Figure 3 The steps for handling multi-objective optimization problems using a non-dominated sorting mechanism include: Step a: During store operations, the multi-objective reinforcement learning module obtains multi-source data from the data acquisition module, including historical customer transaction data, real-time customer behavior data, and external environment data. It combines the current customer state vector and personalized promotional actions to form a sample sequence of state-action-multidimensional reward-next state, which is used to drive the training of the multi-objective deep Q-network. This multidimensional reward vector is subsequently decomposed into three components: short-term sales reward, customer satisfaction reward, and long-term customer value reward. Step b: During training, the multi-objective reinforcement learning module uses a multi-objective deep Q-network to simultaneously update the value of the same state-action pair on three optimization objectives. The update rule is written as follows: , in, Indicates the moment of decision-making Customer status is Execute promotional activities At that time, the three-dimensional Q-value vector output by the current multi-object depth Q-network, Indicates at time The constructed customer state vector, Indicates at time Personalized promotional actions issued to this customer. This represents the parameter vector of the current multi-objective deep Q-network. This represents the learning rate coefficient used to control the degree of fusion between the current estimate and the target estimate. Indicates at time A three-dimensional reward vector consisting of short-term sales incentives, customer satisfaction incentives, and long-term customer value incentives. This represents the discount factor used to discount future earnings. Indicates the next customer status Below, for all candidate promotional actions The component-wise maximum value vector is obtained by comparing the component-wise Q-value vectors of the multi-objective Q-value vectors. Indicates the moment of decision-making Execute action The next customer state vector obtained after the transition, This represents the target network parameter vector used to construct the target Q-value. This represents the time index of the current decision-making moment. This represents the time index of the next decision point obtained after the current decision is completed; Step c: After updating the Q-value, the multi-objective reinforcement learning module uses the updated multi-dimensional Q-value vector as the comprehensive value evaluation result of the action across the three business objective dimensions. Given a customer state, it compares the merits of different promotional actions, and the evaluation relationship is written as: , in, Indicates the customer status Below, regarding promotional activities The resulting three-dimensional action value assessment vector, Indicates at time Candidate personalized promotional actions, Indicates at time The customer state vector, Consistent with the aforementioned meaning, it reflects the action value of the short-term sales target dimension, the customer satisfaction target dimension, and the long-term customer value target dimension in three components respectively. This represents the parameter vector of the current multi-object deep Q-network, with subscripts... This indicates the corresponding component at the decision time. efficient; Step d: Under the same customer state, for a set of candidate promotional actions, the multi-objective reinforcement learning module constructs a multi-objective solution set based on the aforementioned action value evaluation vector; for any two candidate actions... and The multi-objective value vectors were calculated separately. and The dominance relationship is written as: , in, Indicates the first in the set of candidate promotional actions A promotional activity, Indicates the first in the set of candidate promotional actions A promotional activity, This indicates a partial order relation of "dominance" in a multi-objective context. This indicates that the logical conditions on both sides are equivalent. Indicates promotional actions Value assessment vector across three business objective dimensions Indicates promotional actions Value assessment vector across three business objective dimensions express In the Component values in each business objective dimension express In the Component values in each business objective dimension Indexes representing business objective dimensions. This represents the target index set, where index 1 corresponds to the short-term sales target dimension, index 2 corresponds to the customer satisfaction target dimension, and index 3 corresponds to the long-term customer value target dimension. Represents a set All indexes All conditions within parentheses are met. Indicates in set There exists at least one index in it. The conditions within the parentheses must be met. This indicates that both logical conditions must be true simultaneously. According to this judgment condition, if If established, then promotional activities will be effective across all business objectives. No less effective than promotional activities And it outperforms promotional activities in at least one business objective dimension. By traversing all promotional action pairs and applying the above dominance determination, multiple non-dominance frontier levels can be constructed. Promotional actions that are not dominated by any other actions are classified into the first level of non-dominance frontier, and promotional actions that are still not dominated by the remaining actions after removing the first level are classified into the second level of non-dominance frontier, and so on. Step e: After obtaining the set of candidate promotional actions divided by non-dominance level, the multi-objective reinforcement learning module can combine the retail enterprise's tendency towards the three types of business objectives at different operating stages, select the candidate promotional actions with strong compromise from the leading non-dominance frontier, form a set of optimized promotional strategies, and the strategy execution module converts the promotional strategy into promotional configuration parameters for the POS systems of each store, so as to realize the unified strategy distribution and linkage execution within the store network. Specifically, this section describes how a software system can simultaneously consider multiple business objectives and transform them into executable store promotion strategies. The system first uses multi-source data collected from the cloud to form state, action, and multi-dimensional reward samples. Through a deep network, it provides a value assessment of each action across multiple objective dimensions, thus unifying previously scattered business metrics into a computable representation framework. Subsequently, it introduces the concept of non-dominated ranking, treating each candidate promotion action as a solution under multiple objectives given a customer state. By comparing the performance of different solutions across various business objectives, actions that are inferior in all objectives are eliminated, resulting in a set of compromise solutions composed of several layers. Each layer's non-dominated front represents a different degree of balance, allowing operations personnel to select suitable candidate solutions from the higher layers based on phased objectives, such as focusing on boosting short-term sales or accumulating long-term value. In this way, the system can adaptively balance short-term conversion and long-term relationship maintenance without changing the underlying learning structure, providing the store network with a clearly structured and scalable intelligent promotion decision-making method. For example, the multi-objective reinforcement learning module can extract recently accumulated interaction samples from cloud storage at fixed time intervals for batch parameter updates. This time interval can be configured from five minutes to one hour, depending on the business's requirements for policy timeliness. When there are many stores and customer interaction events are frequent, a shorter interval is preferred to increase the policy update frequency. In the small-scale trial phase, a longer interval can be used to reduce computing resource consumption. In each round of updates, the module can set a lower and upper limit for the number of samples used at one time. Typical settings are between several thousand and tens of thousands of samples. When the number of samples collected in the current time interval is less than the lower limit, it can be combined with samples from the previous time interval to ensure the stability of the network parameter update process. Furthermore, sample generation can be automatically recorded by the cloud service after the policy execution module completes the promotion configuration distribution. The recorded content includes the anonymized customer identifier, the corresponding customer state vector, the actual personalized promotion actions distributed, and the three-dimensional reward vector obtained by summarizing within the preset observation window, thus eliminating the need to add complex processing logic locally at the store. Optionally, for retail enterprises that have just joined the system and have limited historical data, the multi-objective reinforcement learning module can be pre-trained using existing offline historical data. Before the cumulative number of online interaction samples reaches a preset threshold, the system can use an initial promotional strategy generated by empirical rules or expert configuration as a fallback. After the model converges and stabilizes, the proportion of the model's output strategy in actual deployment can be gradually increased. If an abnormally large loss value or parameter update failure is detected during training, the system can automatically freeze the model parameters from the most recent convergence, revert to the previously validated and stable promotional strategy, and issue an alert to the operations and maintenance personnel to ensure the continuity of store operations.
[0029] In one embodiment, refer to Figure 2 The multi-objective reward function includes a weighted combination of short-term sales rewards, customer satisfaction rewards, and long-term customer value rewards. The short-term sales rewards are used to measure the increase in sales generated within a preset time window after the promotion, the customer satisfaction rewards are used to measure the customer's interaction behavior or evaluation results after the promotion, and the long-term customer value rewards are used to measure the customer's expected contribution in the long-term period. The weighted combination of short-term sales incentives, customer satisfaction incentives, and long-term customer value incentives is defined as follows: Step f: During the store interaction process, the multi-objective reinforcement learning module obtains a multi-dimensional reward vector based on customer status and personalized promotional actions. , written as: , in, Indicates the moment of decision-making The obtained three-dimensional multi-objective reward vector, Indicates at time The short-term sales incentive component generated by promotions is used to reflect the sales increase effect within a preset time window. Indicates at time The customer satisfaction reward component generated by promotions is used to reflect changes in customer interaction behavior or evaluation results. Indicates at time The long-term customer value reward component generated by promotions is used to reflect changes in expected customer contribution over a longer period. This indicates the corresponding reward at the decision-making time. Effective; Step g, in order to combine the three types of rewards into a single adjustable comprehensive reward signal in the same decision-making process, the multi-objective reinforcement learning module constructs a weighted summation form of the reward function: , in, Indicates the moment of decision-making The comprehensive reward scalar used to drive updates in multi-objective deep Q-networks. , , These correspond to the aforementioned meanings, specifically the three components: short-term sales incentives, customer satisfaction incentives, and long-term customer value incentives. This indicates the weighting coefficient for short-term sales incentives. The weighting coefficients representing the customer satisfaction reward components. These represent the weighting coefficients for long-term customer value rewards; to ensure the overall reward is interpretable and adjustable, the three types of weighting coefficients satisfy the following constraints: , in, This indicates the weighting coefficients for the three categories in the index. The summation operation within the range equals 1. Indicates that the index is The weighting coefficients corresponding to the business objectives, when When corresponding to short-term sales incentives, When corresponding customer satisfaction rewards are given, The time corresponds to the long-term customer value reward, symbol This indicates a summation operation on each item within a specified index range, with constraints. This indicates that each weight coefficient is a non-negative value, and the set... This represents the set of indices representing the three types of business objectives; through the above weighted combination, the system can simultaneously reflect the impact of the three types of business objectives in a single scalar reward. Step h: In order to dynamically adjust the focus of the three types of business objectives according to the store's strategy preferences at different operating stages, the multi-objective reinforcement learning module can introduce an operating stage index. The weight coefficients are updated in stages; assuming stages The weights below are The following formula can be used for updating: , in, Indicating the next stage of operation Time The weighting coefficients corresponding to each business objective Indicating the current stage of operation Time The weighting coefficients corresponding to each business objective This represents the weight adjustment step size coefficient, used to control the speed at which the new weights converge towards the target weights. Indicating during the operation phase The target weight reference value is calculated based on actual operational feedback (such as sales target achievement, changes in satisfaction indicators, and long-term value growth performance), and the subscript is... Index indicating the current operational stage, subscript Indicates the index for the next stage of operation; After this update, the constraints in step g can be applied. Normalization is performed to ensure that the three weights sum to 1 and are non-negative in each operational stage. In this way, the system can shift towards short-term transformation or long-term value at different stages while maintaining the learning structure of the multi-objective reinforcement learning module. Step i: After completing the construction of the comprehensive reward and dynamic adjustment of the weights, the multi-objective reinforcement learning module can select the comprehensive reward. It can directly drive the update of a single-objective deep Q-network, or the comprehensive reward can be used as a reference signal for a certain dimension of the multi-objective Q-value vector to calibrate the existing three-dimensional reward structure. Combining the aforementioned multi-objective deep Q-network and non-dominated ranking mechanism, the system continuously uses updated weight coefficients and comprehensive reward signals to iteratively optimize the promotion strategy during the training phase, and distributes the optimized promotion configuration to multiple store POS systems through the strategy execution module; Specifically, this approach transforms the previously fragmented three categories of business metrics into an operational reward structure. By treating short-term sales, customer satisfaction, and long-term value as independent components, the system can generate structured reward information for each customer interaction. A weighted combination mechanism assigns different weights to the three categories of metrics, allowing operators to drive the learning process with a comprehensive reward while retaining fine-tuning space for the importance of each metric. The constraint that the sum of the weights is one and non-negative ensures that the comprehensive reward has a stable numerical scale and clear business meaning. Furthermore, by introducing weight update rules at the time or operational stage level, the system can gradually adjust the weights based on sales target achievement, changes in customer feedback, and long-term value accumulation, favoring short-term sales or strengthening customer relationship maintenance at different stages without altering the underlying algorithm structure. In this way, the multi-objective reward function not only aligns with the dynamic needs of store operations but also works collaboratively with the aforementioned multi-objective Q-learning and non-dominated ranking mechanism to form a sustainable, iterative, and stage-adjustable intelligent promotional strategy generation method for store members in the cloud. Similarly, regarding the weight configuration of multi-objective reward functions, retail enterprises can preset multiple weight schemes for different operational stages through a cloud-based management interface. For example, during new store openings or large-scale promotional events, the weight of short-term sales rewards can be increased; during membership operations and repeat purchase cultivation stages, the weight of customer satisfaction and long-term customer value rewards can be increased. The three weights are summed to a total of one at any stage, and the typical value range for each weight is between 0.1 and 0.7, with specific values set by the enterprise according to its own operational strategy. Furthermore, the model update module can statistically analyze business indicators such as sales growth rate, customer evaluation metrics, and customer retention rate for each store under the current weight configuration on a weekly or monthly basis. Based on the business rules preset by operations personnel, it automatically calculates the target weight reference value for the next stage and provides adjustment suggestions on the management interface. Operations personnel can then confirm and implement the new weight scheme. Optionally, when the system detects that a certain business indicator is significantly lower than a preset threshold for multiple consecutive statistical periods, it can suggest appropriately increasing the reward weight corresponding to that indicator, while limiting the single adjustment to no more than a certain percentage of the original weight, such as no more than 10%, to avoid drastic changes in the reward structure that could lead to strategy instability. When a corrupted weight configuration file or missing parameters are detected, the model update module can automatically revert to the factory default weights or the most recently successfully saved weight snapshot, and record the abnormal information for subsequent auditing and investigation.
[0030] In one embodiment, a model update module is also included, connected to the multi-objective reinforcement learning module, for dynamically adjusting the weights of short-term sales rewards, customer satisfaction rewards, and long-term customer value rewards in the multi-objective reward function based on customer feedback data, so that the reinforcement learning module can adaptively balance short-term conversion and long-term customer value at different business stages. In one embodiment, the data acquisition module supports a streaming data processing framework for ingesting customer behavior events in real time as an event stream, and performing deduplication, anomaly detection, and noise filtering on the customer behavior events to generate cleaned behavior data for state modeling. Optionally, real-time event stream ingestion can be achieved through a distributed message queue or log subscription mechanism. The end-to-end latency from the generation of a single event at the store to its consumption by the cloud data acquisition module is preferably controlled within the range of hundreds of milliseconds to several seconds to ensure timely updates to the customer state vector. In this embodiment, deduplication can be based on a combination key of event timestamp, customer identifier, and event type, retaining only the first or latest event within a preset duplicate detection window. The duplicate detection window can be configured from one second to ten seconds depending on network fluctuations. Anomaly detection can be achieved by setting a maximum number of events allowed per customer per unit time. When the actual number exceeds this limit, the excess is marked as abnormal and discarded. A typical value for this limit can be set to tens of times per minute to prevent an event storm caused by abnormal clients or system failures from impacting the downstream processing chain. Furthermore, noise filtering can utilize simple threshold-based rules, such as directly ignoring events with a dwell time significantly less than several seconds and truncating clicks that clearly exceed a reasonable range in a single session. Specific thresholds can be set by the operations and data teams based on historical distribution data and configured in the system. When a data source from a store or channel is detected to be completely interrupted or the event latency is significantly increased within a specified time, the data acquisition module can enter a degraded mode. In this mode, only historical transaction data and cached recent behavior data are retained for state modeling, and the data source connection status is retried periodically. Once normal operation is restored, the module will automatically exit the degraded mode.
[0031] In one embodiment, the system adopts a multi-tenant architecture, with retail enterprises as tenants managing their own membership data and promotional rules, and allowing multiple stores within the same tenant to share optimized promotional strategies, enabling customers to enjoy cross-store benefit redemption options and a unified membership discount strategy when making purchases at different stores. Furthermore, a multi-tenant architecture can be implemented by introducing a tenant identifier field into all storage records related to member data, behavioral data, and promotional strategies. Each transaction record, behavioral event record, customer state vector, and corresponding promotional strategy record are simultaneously associated with a unique tenant identifier and store identifier. During read and write operations, the system always uses the tenant identifier as one of the filtering conditions, thereby ensuring logical data isolation between different retail enterprises. When sharing optimized promotional strategies, the strategy execution module first retrieves the strategy version list under the tenant based on the tenant identifier of the logged-in account, and then combines the store identifier and store business configuration to determine whether the current store should adopt the headquarters' unified strategy or the store's own coverage strategy. For stores that have not enabled cross-store benefits, only the strategy entries applicable to the store are issued. Optionally, to prevent cross-tenant data leakage due to configuration errors, the system can add tenant consistency verification during the strategy release stage. When the target store list is detected to contain store identifiers from different tenants, the release is automatically blocked and the operations personnel are prompted to correct the configuration. If, during daily operations, some records are found to be missing tenant identifiers or have tenant identifiers that do not conform to the expected format, the multi-tenant management logic can temporarily store such records in an isolated area, preventing them from participating in business queries and policy calculations, and triggering a data quality audit process for verification and repair by operations or data personnel.
[0032] In one embodiment, the strategy execution module is also used to map the optimized promotion strategy to a distribution plan for different channels and time windows, including differentiated promotion parameter configurations for offline store POS systems, online malls and mobile applications, so as to automatically complete discount calculation, points accumulation and deduction and gift distribution in different channels according to the promotion configuration. In this embodiment, the distribution plans for different channels and time windows can be generated at the granularity of activity cycle, natural day, or hour. By default, the promotion configuration for offline store POS systems is fully synchronized once before the start of each business day. When a significant adjustment to the promotion strategy output by the reinforcement learning model is detected, incremental synchronization can be performed at intervals of five to thirty minutes during business hours to balance the timeliness of the strategy and the load on the store system. For online stores and mobile applications, promotion parameters can be retrieved in real time when customers visit through a unified configuration service, or they can be cached within a short period. The caching time can be set from several minutes to tens of minutes depending on the access volume and update frequency. Furthermore, to avoid putting pressure on the store side due to frequent updates, the strategy execution module can merge multiple strategy change requests received by the same store in a short period of time in the cloud, and push the latest version of the merged version only at the end of the merging window. The typical value of the merging window can be several minutes. When the offline store's network connection is temporarily interrupted, the local POS system can continue to use the most recently successfully synchronized promotional configuration for discount calculation, points accumulation and deduction, and gift distribution. After the connection is restored, the strategy execution module can compare the local configuration with the latest cloud strategy, fill in the missing strategy version based on their respective effective time and validity period, and clean up expired configurations. Optionally, each strategy distribution operation can record the distribution time, target channel, applicable time window, and strategy version identifier in the cloud. When abnormal discounts or gift distribution records occur later, the relevant strategy version and distribution batch can be quickly located based on this metadata, assisting in problem investigation and compliance auditing.
[0033] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0034] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of this application and form different embodiments. For example, all the embodiments above can be used in any combination. The information disclosed in this background section is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.
Claims
1. A store membership management system based on a SaaS cloud platform, deployed on a cloud server and connected to multiple store POS systems via a network, characterized in that: The system includes: The data acquisition module is used to collect multi-source data, including at least customer historical transaction data, customer real-time behavior data, and external environment data. A state modeling module, connected to the data acquisition module, is used to construct a customer state vector representing the current state of the customer based on the multi-source data. An action generation module, connected to the state modeling module, is used to generate at least one personalized promotional action based on the customer state vector. A multi-objective reinforcement learning module, connected to the action generation module, is used to perform reinforcement learning training and strategy optimization on the personalized promotional action based on a multi-objective reward function to obtain an optimized promotional strategy. And a strategy execution module, connected to the multi-objective reinforcement learning module, is used to convert the optimized promotion strategy into a store-executable promotion configuration and distribute it to the multiple store POS systems for execution.
2. The store membership management system based on a SaaS cloud platform as described in claim 1, characterized in that, The customer state vector constructed by the state modeling module includes the customer RFM score, real-time behavior sequence, and external environment indicators. The real-time behavior sequence includes the customer's scanning behavior, dwelling behavior, and clicking behavior in the store, and the external environment indicators include the intensity of competitor promotions and weather data.
3. The store membership management system based on a SaaS cloud platform as described in claim 2, characterized in that, The construction of the customer state vector includes: Feature extraction is performed on the real-time behavior sequence to obtain behavioral features that characterize behavior frequency, behavior sequence, and behavior duration; and the external environment indicators are normalized.
4. The store membership management system based on a SaaS cloud platform as described in claim 1, characterized in that, The personalized promotional actions generated by the action generation module include a combination of at least one of the following: discount level, points multiplier, gift options, limited-time purchase qualification, and cross-store benefit redemption options. The personalized promotional actions are used to configure corresponding promotional incentive schemes for individual customers or customer subgroups.
5. The store membership management system based on a SaaS cloud platform as described in claim 1, characterized in that, The multi-objective reinforcement learning module evaluates and updates the personalized promotional actions based on the multi-objective deep Q-network algorithm, and uses a non-dominated ranking mechanism to handle the multi-objective optimization problem.
6. The store membership management system based on a SaaS cloud platform as described in claim 5, characterized in that, The multi-objective reward function includes a weighted combination of short-term sales rewards, customer satisfaction rewards, and long-term customer value rewards. The short-term sales rewards are used to measure the increase in sales generated within a preset time window after the promotion. The customer satisfaction rewards are used to measure customer interaction behavior or evaluation results after the promotion. The long-term customer value rewards are used to measure the expected contribution of customers in the long term.
7. The store membership management system based on a SaaS cloud platform as described in claim 6, characterized in that, It also includes a model update module, which is connected to the multi-objective reinforcement learning module, and is used to dynamically adjust the weights of short-term sales rewards, customer satisfaction rewards and long-term customer value rewards in the multi-objective reward function based on customer feedback data.
8. The store membership management system based on a SaaS cloud platform as described in claim 1, characterized in that, The data acquisition module supports a streaming data processing framework, which is used to capture customer behavior events in real time in the form of an event stream, and to perform deduplication, anomaly detection and noise filtering on the customer behavior events.
9. The store membership management system based on a SaaS cloud platform as described in claim 1, characterized in that, The system adopts a multi-tenant architecture, with retail enterprises managing their own membership data and promotional rules as tenants. It also allows multiple stores within the same tenant to share the optimized promotional strategies, enabling customers to enjoy cross-store benefit redemption options and a unified membership discount strategy when making purchases at different stores.
10. The store membership management system based on a SaaS cloud platform as described in claim 1, characterized in that, The strategy execution module is also used to map the optimized promotion strategy into a distribution plan for different channels and time windows, including differentiated promotion parameter configurations for offline store POS systems, online malls, and mobile applications.