Multi-mode full-link cooperation method for right supermarket

By integrating multimodal data and employing a full-link collaboration mechanism, the issues of fragmented rules and value differences among user groups in the rights and benefits business have been resolved, enabling precise matching and efficient response of rights and benefits services, and improving service adaptability and stickiness of high-value users.

CN122022909APending Publication Date: 2026-05-12CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-01-26
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively address the issues of fragmented rules, frequent updates across multiple scenarios, and differences in user segment value in rights and benefits businesses. This results in poor module synergy, lagging adaptation of regional rules, and mismatches in matching high-value users, making it difficult to meet the needs for precise and efficient business operations.

Method used

By employing a multimodal data fusion approach, a rights-specific modal system is constructed. Through behavioral, environmental, emotional intent, and business rule modalities, combined with rights-specific dimensions, a full-link collaborative mechanism is built to achieve precise matching between fragmented rules and user needs. This includes data fusion, modeling, prediction, rule optimization, and feedback iteration, thereby strengthening the relevance and adaptability of regional rights rules.

Benefits of technology

It improved the adaptability and response efficiency of rights and benefits services, achieved dynamic adaptation of rules, improved service accuracy and stickiness of high-value users, and optimized operating costs and user matching satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122022909A_ABST
    Figure CN122022909A_ABST
Patent Text Reader

Abstract

The invention provides a right supermarket-oriented multi-modal full-link cooperation method, belongs to the field of right personalized services, and is used for solving the problems of right rule fragmentation, slow multi-scene adaptation and inaccurate user grouping matching in related technologies. The method comprises the following steps: collecting behavior, environment and other multi-modal data, fusing features through a right-business correlation matrix and an attention mechanism, and generating user groups through adaptive density clustering; the method combines Bayesian weight inference and LTV prediction modeling influence degree, adopts a hybrid pre-judgment model to pre-judge a matching result, generates a grouping exclusive rule through multi-objective optimization, and is matched with an intention perception conflict mediation scheme to realize full-link dynamic collaboration and improve service precision and operation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of personalized rights services, and in particular to a multimodal full-link collaborative method for rights supermarkets. Background Technology

[0002] With the widespread adoption of digital services, the "Benefits Supermarket" has rapidly developed as a platform aggregating various user benefits. Here, "benefits" broadly refers to various digital perks or privileges that users can obtain or redeem. Typical forms include, but are not limited to: exclusive data packages for telecommunications or internet platforms, membership vouchers, offline merchant coupons, extended movie / entertainment memberships, points redemption for physical goods, interest-free installment plans, and dedicated customer service channels. The platform's core need is to provide users with precise, personalized benefit matching services to enhance user engagement and business conversion.

[0003] Currently, our benefits services cover multiple regions and scenarios, and the user base and the complexity of their needs continue to grow, placing higher demands on the accuracy and response speed of our services. Achieving personalized matching typically involves several key business parameters, such as:

[0004] Points: Virtual points accumulated by users through consumption or active behavior, used for redeeming benefits or level assessment;

[0005] Membership levels: Tiers are defined based on factors such as user history and spending power, with different levels corresponding to different eligibility and priority for obtaining benefits;

[0006] Region: The province or city where the user is located or is currently located, which directly affects the compliance and applicable rules of the rights and benefits that can be offered.

[0007] In existing technologies, general AI personalization methods are mostly applied to e-commerce, content recommendation, and other fields. While these methods can achieve basic user preference matching, they are not customized for the specific characteristics of the aforementioned benefits-based services. Most solutions use fixed-dimensional data modeling and rely on general behavioral characteristics, failing to fully integrate core elements specific to the benefits-based services, such as points consumption patterns, tiered membership benefits, and regionally differentiated compliance strategies.

[0008] The aforementioned existing technologies have significant drawbacks: due to the lack of an identification and matching model deeply integrated with rights and benefits parameters, the system cannot adapt to the characteristics of rights and benefits services, such as "fragmented provincial rules, frequent updates across multiple scenarios, and significant differences in the value of user groups." Specifically, this manifests as poor module synergy, lagging adaptation of regional rules, and mismatches in matching high-value users, making it difficult to meet the business needs for precision and efficiency in rights and benefits services and hindering the large-scale development of rights and benefits services. Summary of the Invention

[0009] This application provides a multimodal, end-to-end collaborative method for rights and benefits supermarkets, which can accurately adapt to the characteristics of rights and benefits business, solve the problem of fragmented rules and personalized matching, and improve service accuracy and business efficiency.

[0010] Firstly, this application provides a multimodal end-to-end collaborative method for rights and benefits supermarkets. This method addresses the fragmented nature of rights and benefits business rules, the high-frequency updates across multiple scenarios, and the differences in value among user groups. It collects multimodal data containing rights-specific dimensions. These rights-specific dimensions are feature dimensions designed specifically for rights and benefits supermarket business and distinct from general service dimensions. These include dimensions directly related to rights and benefits business, such as rights redemption sequences and points consumption characteristics. This data includes behavioral modalities, environmental modalities, emotional intent modalities, and business rule modalities. Behavioral modalities include rights redemption sequences and points consumption characteristics; environmental modalities include provincial regional labels and terminal types; emotional intent modalities include rights consultation semantic features; and business rule modalities include provincial compliance labels and points thresholds. Based on this data, an end-to-end link of "data fusion - modeling - prediction - rule optimization - conflict resolution - feedback iteration" is constructed to achieve dynamic adaptation of fragmented rules.

[0011] By adopting the above technical solutions and building a full-link collaborative mechanism based on rights-specific multimodal data, the limitations of dimensionality and module isolation of general methods are broken, enabling fragmented rules to be accurately matched with user needs, and improving the adaptability and response efficiency of rights services.

[0012] Furthermore, the multimodal data fusion process includes constructing a rights-specific modality system, which includes a rights combination selection sequence for the behavioral modality, provincial regional codes for the environmental modality, emotional scores for consultation texts for the emotional intent modality, and rights validity period labels for the business rule modality. Based on this system, a rights-business relevance matrix is ​​constructed, and an attention mechanism with provincial correlation adjustment is used to fuse features, outputting fusion features that conform to provincial rules.

[0013] By adopting the above technical solutions, the correlation between multimodal data and regional rights rules is strengthened, the business adaptability of fused features is improved, high-quality input is provided for subsequent accurate modeling, and regional rule adaptation errors are reduced.

[0014] Furthermore, the process of modeling the impact of rights and interests includes constructing a set of rights and interests-specific dimensions, which includes basic dimensions, scenario-specific dimensions, and group-specific dimensions. The basic dimensions include membership level and rights and interests combination preferences. The scenario-specific dimensions are feature dimensions preset for different rights and interests scenarios. The group-specific dimensions are feature dimensions customized for different user groups. The dimension weights are determined by using rights-oriented Bayesian weight inference. Dynamic thresholds are determined based on user group value, scenario importance, and the number of rights and interests. The impact score is adjusted in combination with the tolerance for points gap.

[0015] By adopting the above technical solutions, the rights and interests of the dimensions and weights can be dynamically adjusted in a scenario-based manner, highlighting the impact of core elements such as membership level and points, improving the accuracy of impact assessment, and providing a reliable basis for matching high-value users.

[0016] Furthermore, the prediction module adopts a hybrid prediction model that integrates gradient boosting, meta-learning and adversarial learning. The generator generates simulated matching data based on provincial regional rule labels and integral constraints. The discriminator distinguishes between real and generated data. In a small number of sample scenarios, a meta-gradient accumulation mechanism is used to accelerate adaptation. The cause of failure is located and a calibration strategy is triggered by a rights-specific error attribution network. The small number of sample scenarios are rights promotion scenarios where the number of existing samples does not exceed a preset sample number threshold.

[0017] By adopting the above technical solutions, the model's adaptation speed to new rights and a small number of sample scenarios is improved, the specific reasons for the failure of predicted rights are accurately located, the stability and reliability of the prediction results are ensured, and the adaptation cycle after rule updates is shortened.

[0018] Furthermore, the rule optimization process includes constructing a rights compliance attribution network, solving for Pareto optimal rules through a multi-objective optimization module, optimizing objectives including conversion efficiency, compliance rate, and rights cost, generating exclusive rules for different user groups, relaxing the points gap constraint for high-value member rules, simplifying the constraints for new customer rules, and assigning high-value member rules to high-value groups, which are user groups whose long-term value prediction value is not lower than the preset user value. After the candidate rules pass compliance verification and matching verification, they are synchronized to the corresponding provincial rule subset.

[0019] By adopting the above technical solutions, multi-objective optimization and localization of rules were achieved, balancing conversion and cost while ensuring compliance. Group-specific rules improved user matching satisfaction and reduced operating costs.

[0020] Furthermore, the conflict resolution process includes calculating conflict priority based on the core priority of rights, which resets the effective rights options to the highest level, followed by membership level, provincial region, and points. A customized solution library for different groups is constructed. The high-value membership solution includes points coupons and rights extension, while the new customer solution includes points acceleration packages and simplified rules. The high-value membership solution corresponds to a high-value group, which is a group of users whose long-term value prediction is not lower than the preset user value. A constrained reinforcement learning optimization strategy is adopted to establish a mediation-guidance-upgrade closed loop. After the user's task is completed, the group is automatically upgraded and the rule threshold is lowered.

[0021] By adopting the above technical solutions, conflicts are resolved with a focus on core rights and interests. Customized solutions for different groups increase the acceptance of mediation, and a closed-loop system promotes user activity, achieving the dual goals of conflict resolution and user value enhancement.

[0022] Furthermore, the user segmentation process includes extracting a set of exclusive behavioral features for benefits, which includes the frequency of benefit redemption, the proportion of points consumption, the click rate of provincial benefits, and the long-term value prediction value. A benefit-adaptive density clustering algorithm is used to generate core segments. Temporary segments are triggered by activity scenarios, and they are merged according to long-term value after the scenario ends. Resource allocation is implemented based on the segment value. High-value segments are matched with high-quality benefit rules. High-value segments are user segments whose long-term value prediction value is not lower than the preset user value.

[0023] By adopting the above technical solutions, we have achieved precise segmentation based on the value of rights and benefits, dynamically adapted to the needs of activity scenarios, improved the stickiness of high-value users through resource allocation strategies, and optimized the efficiency of rights and benefits resource allocation.

[0024] Furthermore, it also includes the steps of customizing the equity mathematical model, imposing mathematical constraints on the weights of equity-specific dimensions, strengthening the weight ratio of core dimensions through prior distribution, constructing an equity segmented reward mechanism, setting reward values ​​based on the achievement of rigid constraints, setting gradient rewards based on the points gap of flexible constraints, constructing a long-term value model based on the equity consumption time series, and transforming the upper limit of the group cost into a constraint embedded in the objective function.

[0025] By adopting the above technical solutions, the mathematical model can be deeply aligned with the constraints and objectives of equity business, strengthen the orientation of core elements, balance compliance and value, and avoid the business disconnect caused by pure mathematical modeling.

[0026] Furthermore, it also includes scenario-based feedback iteration steps. When a user triggers a core behavior related to their rights, the influence weight and multimodal feature weight are adjusted in real time. During peak periods, the fine-tuning cycle of the hybrid prediction model is shortened and the learning rate is improved. During off-peak periods, data and driving rule optimization are updated in stages. The modality matrix and cluster model are updated weekly. Emergency feedback is triggered when the prediction error exceeds the standard.

[0027] By adopting the above technical solutions, we have achieved dynamic iteration across the entire chain, quickly responding to changes in user needs and peak demand, ensuring the continuous adaptability of models and rules, and maintaining stable service performance.

[0028] Furthermore, it also includes: provincial rule subset synchronization and verification steps: after the provincial rule subset is updated, it is synchronized to the corresponding regional rights data cache node, and the adaptability of the updated rules and user groups is verified through a feature comparison mechanism. When the adaptation deviation exceeds the set range, the dimension weight is adjusted.

[0029] By adopting the above technical solutions, the adaptation verification and closed-loop optimization after rule updates are strengthened, the deviation of dimension weights is corrected in a timely manner, the matching consistency between rules and groups is ensured, and the stability of regional services is improved.

[0030] In summary, this application has at least the following beneficial effects:

[0031] It provides a full-link collaboration method specific to rights and interests, solving the problem of fragmented rule adaptation and improving service accuracy;

[0032] By segmenting users and optimizing rules, we can achieve efficient allocation of rights and resources and enhance the stickiness of high-value users.

[0033] Contextualized feedback and iteration ensure continuous service adaptation and maintain a stable level of business conversion.

[0034] It should be understood that the description in the Summary Section is not intended to limit the key or essential features of the embodiments of this application, nor is it intended to restrict the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0035] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0036] Figure 1 A schematic diagram of an exemplary operating environment in which embodiments of this application can be implemented is shown.

[0037] Figure 2 A flowchart of a multimodal end-to-end collaborative method for a rights and interests supermarket is shown in an embodiment of this application. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0039] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0040] This application provides a multimodal, end-to-end collaborative method for rights and benefits supermarkets. It accurately adapts to the characteristics of rights and benefits business, solves the problem of fragmented rules, achieves efficient matching of rights and benefits with user needs, and improves service accuracy and stickiness of high-value users.

[0041] Figure 1 A schematic diagram of an exemplary operating environment in which embodiments of this application can be implemented is shown.

[0042] Reference Figure 1 The operating environment includes a multi-dimensional hardware collaboration system that supports the realization of the "multimodal full-link collaboration method for equity supermarket". This system takes "data collection-storage-computation-transmission-security" as the core link and meets the real-time, accuracy and compliance requirements of the entire equity business chain through the orderly connection and functional cooperation of various hardware modules.

[0043] The core hardware components of this operating environment include a multi-terminal data acquisition cluster, a distributed storage architecture, a tiered computing resource cluster, a low-latency network transmission system, and security and compliance support equipment. The multi-terminal data acquisition cluster covers various service terminals such as apps, TVs, and smart home devices. Each type of terminal is equipped with a front-end behavior tracking module, environmental status sensors, and a semantic recognition unit. The front-end behavior tracking module collects behavioral data such as benefit redemption and points consumption; the environmental status sensors acquire environmental data such as provincial region and terminal type; and the semantic recognition unit extracts emotional features from the consultation text. All collected data is transmitted to the storage architecture via the network.

[0044] The distributed storage architecture adopts a two-tier deployment mode of "edge-center". Each of the 31 provincial edge nodes is configured with a local storage server to cache a subset of the rights and interests rules of the province and real-time data of local users. The central end deploys a Redis cache cluster, a Neo4j graph database cluster and an Elasticsearch cluster. The Redis cache cluster stores multimodal real-time features and core parameters of the model, the Neo4j graph database cluster constructs a "user-rights-provincial rule" association network, and the Elasticsearch cluster indexes fragmented rights and interests rules. The edge storage and the central storage achieve data synchronization and interaction through a dedicated link.

[0045] Tiered computing resource clusters and storage architectures are deployed accordingly. Provincial edge nodes are configured with edge computing servers to undertake real-time lightweight computing tasks such as multimodal feature extraction and influence weight iteration. The central end is configured with GPU-accelerated clusters to carry heavy computing tasks such as adversarial training and multi-objective rule optimization. The edge computing servers and GPU-accelerated clusters are connected through a computing power scheduling system, which allocates computing resources according to task priority.

[0046] The low-latency network transmission system takes the provincial backbone network as its core, builds a dedicated transmission channel between edge nodes and the central end, and configures the nearest access node for each collection terminal to ensure that the transmission latency of the collected data from the terminal to the edge storage is ≤10ms, and the synchronization latency of rules and parameters between the edge and the center is ≤5ms, providing network guarantee for real-time collaboration across the entire link.

[0047] The security and compliance support equipment includes a data anonymization server, a provincial compliance verification server, and an access control gateway. The data anonymization server is connected to the data collection cluster and performs real-time anonymization processing on the collected data. The provincial compliance verification server is connected to the rights and interests rule database of 31 provinces and works in conjunction with the distributed storage architecture and computing resource cluster to complete the compliance verification of rule updates and model output. The access control gateway is deployed at the network entrance to restrict unauthorized data access and operations.

[0048] Each hardware component forms a closed-loop collaboration through a preset network topology. Data from the collection cluster is anonymized and then transmitted to the storage architecture. The computing resource cluster calls data from the storage architecture to complete the computation. The computation results are fed back to the terminal or used to update the rule data in the storage through the network. Security and compliance devices are involved in all aspects of the data flow, ultimately forming a stable hardware environment that supports the end-to-end operation of the aforementioned method.

[0049] This application discloses a multimodal full-link collaborative method for rights and interests supermarkets.

[0050] Figure 2 A flowchart of a multimodal end-to-end collaborative method for a rights and interests supermarket is shown in an embodiment of this application.

[0051] Reference Figure 2 The method specifically includes the following steps:

[0052] S1: To address the fragmented nature of rights and benefits business rules, the high frequency of updates across multiple scenarios, and the differences in value among user groups, collect multimodal data including rights and benefits-specific dimensions.

[0053] The data includes behavioral modalities, environmental modalities, emotional intent modalities, and business rule modalities. Behavioral modalities include rights-claiming sequences and points-based consumption features; environmental modalities include provincial regional tags and terminal types; emotional intent modalities include semantic features of rights-based consultations; and business rule modalities include provincial compliance tags and points thresholds. During the data collection process, targeted preprocessing was performed on each modality: Rights-claiming sequences and points-based consumption features of the behavioral modality were collected in real-time through front-end tracking, and Z-score outlier removal and KNN missing value filling were used to ensure data validity ≥98%; provincial regional tags of the environmental modality were obtained through IP address mapping, and terminal types were determined through device UA resolution, both standardized to a unified encoding format; semantic features of consultation texts in the emotional intent modality were extracted using a BERT model with fine-tuning, outputting a three-dimensional emotion vector of "interested - neutral - impatient," which was then normalized and converted into a single-dimensional emotion score; provincial compliance tags and points thresholds of the business rule modality were synchronously obtained from the rights-based rule database interface of 31 provinces and stored according to "province + rights type."

[0054] The specific methods in this step include: the multimodal data fusion process includes constructing a rights-specific modality system, which includes rights combination selection sequences for behavioral modality, provincial regional codes for environmental modality, emotional scores of consultation texts for emotional intent modality, and rights validity period labels for business rule modality. Based on this system, a rights-business relevance matrix is ​​constructed, and an attention mechanism with provincial correlation adjustment is used to fuse features and output fusion features that conform to provincial rules.

[0055] The formula for constructing the equity-business relevance matrix is: ,in Indicates the first Modal pairs Contribution to individual equity business targets For the first Modal features Corresponding behavioral modalities Corresponding environmental modes, Corresponding emotional intention modality (corresponding business rule modality) For equity business objectives ( Corresponding matching accuracy Corresponding compliance rate Corresponding points conversion Corresponding membership upgrade rate, (corresponding consultation satisfaction) Modal features With business objectives Mutual information between them For business objectives The preset weight value, (in the embodiments of this application) ,satisfy .

[0056] In the process of integrating the attention mechanism for provincial-level relevance adjustment, the query vector is first constructed. AND key matrix ,in The learnable parameter matrix has dimensions of "feature dimension × 64" (64 being the attention head dimension); then the basic attention weights are calculated. ,in, The i-th modal feature is processed by a learnable parameter matrix The resulting query vector, where K is the learnable parameter matrix of all modal features. The key matrix obtained by mapping, For attention head dimension; finally, provincial relevance is introduced. (This province's modality) out-of-province modality This value is set manually to enhance the influence of the province's rules in feature fusion. In practical applications, it can be adjusted within the range of 1.0 to 1.5. (This is related to the maximum business contribution of the modality.) The final attention weights are obtained. fusion features The output dimension is 512.

[0057] The user segmentation process includes extracting a set of exclusive behavioral features for benefits, which includes the frequency of benefit redemption, the proportion of points consumption, the click rate of provincial benefits, and the long-term value prediction value. A benefit-adaptive density clustering algorithm is used to generate core segments. Temporary segments are triggered by activity scenarios and merged according to long-term value after the scenario ends. Resource allocation is implemented based on the segment value. High-value segments are matched with high-quality benefit rules. High-value segments are user segments whose long-term value prediction value is not lower than the preset user value.

[0058] Thus, through data collection, multimodal fusion, and user segmentation in S1, unified fusion features and precise user segments conforming to provincial rules have been generated. The output of this "data fusion" stage will serve as direct input to support the collaborative operation of a series of logical stages such as modeling, prediction, and optimization in the subsequent S2 full-link process.

[0059] The equity-specific behavioral characteristics are concentrated, and the long-term value prediction is calculated using a temporal attention LSTM model, with the following formula: ,in Features of the time series of rights and benefits redemption / consumption over the past 90 days (dimensions: Attention is a temporal attention mechanism. Attention weight matrix ( ), To output the projection matrix For bias terms, The Sigmoid activation function has an output range of... .

[0060] In the equity adaptive density clustering algorithm, the local density of each sample is first calculated. ,in For the sample of Neighborhood The median distance between all samples is used; then the cluster center is determined as the point of local density maximum, satisfying... For all Established; finally, cluster labels are assigned based on the distance between the sample and the cluster center, and automatically generated. The core customer segments (high-value members, points-sensitive new customers, region-specific users, dormant users, etc.) are evaluated for purity using a silhouette coefficient and must meet the Silhouette criteria. .

[0061] The clustering algorithm reflects the adaptability of the benefits business in the following ways: First, it incorporates the weighted Euclidean distance of benefits-specific features (such as the frequency of benefits redemption and the proportion of points consumption) into the distance calculation; second, it uses the long-term value prediction value (LTV) as the sample weight when calculating the density, so that high-value user samples have a higher influence in the clustering; and third, it dynamically adjusts the neighborhood parameter $\epsilon$ according to the intensity of recent benefits activities, and appropriately expands the neighborhood during the activity period to attract potential active users.

[0062] When temporary clustering is triggered by event scenarios (such as Member's Day or Marathon), temporary clusters are quickly generated using the K-means algorithm based on temporary features such as participation frequency and task completion rate during the event. After the event ends, the long-term value similarity between the temporary clusters and the core clusters is calculated. If a cluster is merged into the corresponding core cluster, it will be retained as an independent cluster. When resource allocation is implemented based on cluster value, the cluster value score (ValueScore) will be used. LTV Frequency of claiming benefits Points spending percentage, high-value segment (ValueScore) The proportion of high-quality rights and interests rules matched is not less than .

[0063] S2: Based on the collected multimodal data, construct an end-to-end link of "data fusion - modeling - prediction - rule optimization - conflict resolution - feedback iteration" to achieve dynamic adaptation of fragmented rules.

[0064] This step builds upon the multimodal fusion features and user segmentation output by S1, and then constructs an end-to-end link of "modeling-prediction-rule optimization-conflict mediation-feedback iteration" to achieve dynamic adaptation of fragmented rules.

[0065] The specific methods in this step include: the rights and interests impact modeling process includes constructing a set of rights and interests-specific dimensions, which includes basic dimensions, scenario-specific dimensions, and group-specific dimensions. The basic dimensions include membership level and rights and interests combination preferences. The scenario-specific dimensions are feature dimensions preset for different rights and interests scenarios. The group-specific dimensions are feature dimensions customized for different user groups. Rights and interests-oriented Bayesian weight inference is used to determine the dimension weights. Dynamic thresholds are determined based on user group value, scenario importance, and the number of rights and interests. The impact score is adjusted in combination with the points gap tolerance.

[0066] When determining dimensional weights using a rights-oriented Bayesian weight inference method and establishing dynamic thresholds based on user segment value, scenario importance, and the number of rights, a reinforcement learning reward function integrating multi-dimensional indicators is introduced to achieve the dual goals of short-term matching accuracy and long-term user value optimization. The formula is as follows: ,in Let be the matching rate between users and benefits at time t. For the benefit of click-through rate, To improve the efficiency of points conversion, Emotion The sentiment score of the consultation text extracted from S1, This is the predicted long-term user value calculated in S1. An adaptive learning rate is used to adjust the weight update efficiency, ensuring timely adjustments even with larger errors. The learning rate formula is... In the formula Let be the learning rate at time t. The number of model iterations. The actual matching result label at time t.

[0067] In the construction of the dimension set, the basic dimension set Membership level, provincial / regional location, preferred combination of benefits, points spending habits Scene-specific dimensions The Member's Day event now includes a "Points Balance" dimension, and the Marathon event includes a "Event Participation History" dimension, creating separate dimensions for different user groups. For high-value members, a "history of exclusive benefits redemption" was added; for new customers, a "registration duration" was added. The final set of dimensions was optimized by removing redundancies. ,in For mutual information Redundant dimensions. The equity-oriented Bayesian weight inference treats weights as random variables, with a prior distribution using a Dirichlet distribution. Membership level and provincial region General dimensions Likelihood function ,in To match observation data with rights and interests, Let r be the matching index between user u and equity r at time t. For feature similarity, These are the dimension weights; the final output weights are... ,satisfy And the weighting of core dimensions The formula for calculating the dynamic threshold is as follows: group, scene GroupValueScore Scene Importance RightsQuantity, where GroupValueScore is the group value score (high-value members). SceneImportance is the importance of a scene (marathon) RightsQuantity represents the total number of rights within the scenario; the score reflects the impact of the points adjustment. GapTolerance Let d be the dimension feature of user u. For the dimension d of the rights and benefits r, GapTolerance is the tolerance for points gap (gap for high-value members). hour ,otherwise ).

[0068] The steps for customizing the equity mathematical model include applying mathematical constraints to the weights of equity-specific dimensions, strengthening the weight ratio of core dimensions through prior distribution, constructing a segmented reward mechanism for equity, setting reward values ​​for rigid constraints based on achievement status, setting gradient rewards for flexible constraints based on points gaps, constructing a long-term value model based on equity consumption time series, and transforming the upper limit of segmented cost into constraints embedded in the objective function.

[0069] The weight mathematical constraints follow the Dirichlet prior distribution mentioned above to ensure the weight proportions of the core dimensions. In the segmented reward function, rigid constraint (compliance / validity period) reward (Meets the standard) or 0 (Does not meet the standard), flexible constraints (points / tasks) rewards (Meets standards), 0.8 (Gap) )or gap Total Rewards The long-term value model employs a temporal attention LSTM, and its output... , For the time series of rights redemption / consumption, Attention is directed towards high-value rights (average order value). The tilted attention mechanism of (yuan) and For model parameters, The sigmoid function is used. The cluster cost constraint is transformed into the inequality RightsCost. group The embedding rule optimizes the objective function, where group Cost cap for segmentation (high-value members) Yuan, new customer Yuan).

[0070] The prediction module adopts a hybrid prediction model that integrates gradient boosting, meta-learning and adversarial learning. The generator generates simulated matching data based on provincial regional rule labels and integral constraints. The discriminator distinguishes between real and generated data. In a small number of sample scenarios, a meta-gradient accumulation mechanism is used to accelerate adaptation. The failure cause is located and a calibration strategy is triggered by a rights-specific error attribution network. The small number of sample scenarios are rights promotion scenarios where the number of existing samples does not exceed a preset sample number threshold.

[0071] In scenarios with a small number of samples, a meta-gradient accumulation mechanism is used to accelerate adaptation, and a rights-specific error attribution network is used to locate the cause of failure and trigger a calibration strategy. The final output of the prediction model is obtained by a weighted fusion of the results of the base model and the meta-learning model to balance stability and adaptability. The fusion formula is as follows: ,in To increase the base matching probability of the model output by gradient boosting, This represents the fit matching probability output by the meta-learning model. The equity-specific error attribution network quantifies the probability of each failure cause using Bayes' theorem to pinpoint the core problem. The formula is: ErrorCause Error In the formula ErrorCause Prior probability of the cause of failure (rule mutation) Points gap Membership level drift ), ErrorCause The conditional probability of the error caused by this reason (based on historical data statistics, this value is 0.85 when the rule changes abruptly). This represents the total probability of the error occurring.

[0072] In the hybrid prediction model, the generator is an MLP network, which takes "provincial regional rule labels + score thresholds + membership constraints" as input and generates simulated data covering "score attainment / deficiency" and "membership fit / mismatch"; the adversarial training objective function is... ,in, This represents the distribution followed by real 'user-benefit' matching data, with the data source being actual matching records from the benefits supermarket within the past 90 days. The noise follows a distribution with the same dimension as the generator input, and is used to drive the generator to output diverse simulated data. To predict the matching rate of the model output, The actual matching result label corresponding to this matching data. The discriminator's judgment result on the real data x. The generator output is shown, with the last term representing the consistency loss. The meta-gradient accumulation mechanism is designed for... The temporary equity samples are updated with parameters through the meta-task gradient, as shown in the formula. ,in, This is the initial parameter set for the meta-learning model, including the model's weight matrix and bias terms. The gradient of the loss function corresponding to the support set data of the k-th meta-task. The gradient of the loss function corresponding to the query set data of the k-th meta-task is used together to update the parameters of the meta-learning model, achieving rapid adaptation with a small number of samples and an adaptation period of ≤30 seconds. The rights-specific error attribution network is a Bayesian network, with nodes containing "provincial rule update, points gap, and member level drift". After locating the cause of failure, repair is triggered: rule update → synchronize dimension weights, points gap → adjust. .

[0073] The last term in the objective function is the consistency loss. This is used to ensure that the model's prediction logic for real data and generated data is consistent, where To predict the model's output on real "user-rights" matching data x, To predict the simulated matching data output by the generator G from the model The output result is 0.05, which is the weighting coefficient of the consistency loss, balancing the impact of prediction loss and adversarial loss.

[0074] The rule optimization process includes constructing a rights compliance attribution network, solving for Pareto optimal rules through a multi-objective optimization module, and optimizing objectives including conversion efficiency, compliance rate, and rights cost. Dedicated rules are generated for different user segments. High-value member rules have relaxed points gap constraints, and new customer rules have simplified constraints. High-value member rules correspond to high-value segments, which are user segments whose long-term predicted value is no less than a preset user value. Candidate rules, after passing compliance verification and matching validation, are synchronized to the corresponding provincial rule subset. The rights compliance attribution network is a Bayesian network that learns parameters through the EM algorithm to quantify the impact weights of reasons such as "Shandong points rules not synchronized," achieving high attribution accuracy. The multi-objective optimization objective function is: (Conversion efficiency) (Compliance rate) (Cost of equity) The constraints include the number of new customer segmentation rules. The gap in flexible constraints for high-value members Shandong user points threshold ...and so on, using the NSGA-III algorithm for output. There are 10 candidate rules. Candidate rules must be validated by compliance interfaces in 31 provinces using UpdateValid. Once approved, the rules will be automatically synchronized to the provincial rule subset, with an update delay. minute.

[0075] When the rule optimization process includes "constructing a rights compliance attribution network and solving for Pareto optimal rules through a multi-objective optimization module," the rights compliance attribution network adopts a Bayesian network structure. It locates the causes of rule failure by quantifying the relationships between multiple factors, as shown in the formula: ,in As a unique dimension of rights and interests, group Dimensional features of a certain group The probability of its occurrence, The network parameters are learned through the EM algorithm to determine the probability that feature combinations lead to rule failure. The constraint condition for multi-objective optimization is explicitly defined as "the number of rule constraints for new customer segmentation". 3. "The flexible constraint points gap for high-value members" "300", candidate rules must pass quantitative verification before they can be updated. The verification formula is: In the formula For the mathematical expectation of the segmented users, Compliance ( ) represents the compliance metric for user u under the rule (1 indicates compliance, 0 indicates non-compliance). When UpdateValid... At that time, the rules can be synchronized to the provincial rule subset.

[0076] After the provincial rule subset is updated, it is synchronized to the corresponding rights and interests data cache node. The adaptability of the updated rules and user groups is verified through the feature comparison mechanism. When the adaptation deviation exceeds the set range, the dimension weights in the rights and interests impact modeling process are adjusted a second time.

[0077] The provincial rule subset is synchronized to the local cache of the corresponding provincial edge node, and the feature comparison mechanism calculates the cosine similarity between the updated rule and the cluster feature. , group , ,when When the adaptation deviation exceeds the standard, the dimensional weights of the rights and interests impact modeling are recalculated, with a focus on adjusting the dimensional weights that are highly correlated with provincial rules.

[0078] After the secondary adjustment of dimensional weights is completed, the rights-business relevance matrix constructed in S1 needs to be updated synchronously to ensure multimodal fusion and adaptation to the latest rules, forming end-to-end data collaboration. The matrix update formula is as follows: Compliance ,in For the updated matrix elements, For the elements of the original matrix, Compliance represents the change in the provincial compliance rate after the rule update (positive when the compliance rate increases and negative when it decreases), and 0.1 is the impact coefficient of the compliance rate on the matrix, ensuring that the adjustment range of the matrix is ​​strongly correlated with the business effect.

[0079] The conflict resolution process includes calculating conflict priority based on the core priority of rights and interests. This priority resets the effective options of rights and interests to the highest level, followed by membership level, provincial region, and points. A customized solution library for different groups is constructed. The high-value membership solution includes points vouchers and rights extension, while the new customer solution includes points acceleration packages and simplified rules. The high-value membership solution corresponds to a high-value group, which is a group of users whose long-term value prediction value is not lower than the preset user value. A constrained reinforcement learning optimization strategy is adopted to establish a closed loop of mediation-guidance-upgrade. After the user completes the task, the group is automatically upgraded and the rule threshold is lowered.

[0080] Conflict Priority Calculation Formula

[0081] , Priority for equity business (validity period) Membership Level ,area ,integral For scenario priority, IntentionStrength represents the user's intention strength (calculated based on the similarity between the consultation text and historical interests). In constrained reinforcement learning, the action space... For the combination of clustered solutions, compliance constraints The reward function is only applied if the scheme meets the provincial rigid rules. AcceptRate(a) represents the acceptance rate of the proposal, and Conversion(a) represents the conversion efficiency. In the onboarding loop, the completion rate of onboarding tasks (such as "viewing membership benefits 3 times") is... At that time, new customers will be automatically upgraded to potential users, and the threshold for corresponding benefits rules will be lowered. .

[0082] When the conflict resolution process includes calculating conflict priority based on core equity priority, which resets the effective equity option to the highest level, the conflict priority calculation needs to incorporate user intent strength to improve the user adaptability of the resolution plan. Intent strength is calculated through the similarity between the current query and historical interests, using the formula IntentionStrength. ,in The semantic features of the user's current rights consultation Behavioral characteristics of users' historical high-value benefits, Let be the set of all conflicts in the current scenario, and sim be the cosine similarity calculation function. The conflict priority formula, after incorporating intent strength, is updated as follows: In the formula Prioritizing equity-related business, Prioritize the scenario. In the guided loop, a reward mechanism is triggered after the user completes the mediation guidance task, with the formula being: LTV Boost (Completion) is the long-term value gain after the task is completed. The value is 0.2 when a high-value member completes the task, and 0.15 when a new customer completes the task.

[0083] The scenario-based feedback iteration steps include real-time fine-tuning of the impact weight and multimodal feature weight when users trigger core rights behaviors, shortening the fine-tuning cycle of the hybrid prediction model and improving the learning rate during peak scenarios, updating data and driving rule optimization at different levels during off-peak periods, updating the modality matrix and cluster model weekly, and triggering emergency feedback when the prediction error exceeds the standard.

[0084] Real-time feedback is provided at the millisecond level. The impact weights of core user behaviors such as claiming benefits and consuming points are finely adjusted, with an adjustment step size of 0.1 times the base step size. During peak scenarios (Member Day / Marathon), the prediction model is incrementally fine-tuned every 10 minutes, with the learning rate... Improved from 0.05 to 0.08; during off-peak hours, provincial edge node caches are updated hourly, and the rule optimization engine is driven daily based on "provincial rule matching rate and points conversion efficiency". The rights-business relevance matrix in S1 is updated weekly based on full data. User segmentation model; emergency feedback trigger condition is prediction error. After triggering, the impact threshold is adjusted synchronously. The adjustment range for the predicted model parameters is twice that of a regular fine-tuning.

[0085] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0086] The multimodal, end-to-end collaborative approach for the rights and benefits supermarket employs a progressive technical architecture of "S1 data collection and fusion - S2 end-to-end modeling and optimization." This architecture leverages the complementary functions and data flow of each module to form a collaborative closed loop, ultimately achieving precise adaptation of fragmented rules and enhancing the value of rights and benefits services. The specific derivation logic is as follows:

[0087] As the foundational support for the entire chain, S1 first ensures the integrity and validity of the original data by collecting multimodal data covering behavior, environment, emotional intent, and business rules, combined with targeted preprocessing methods such as Z-score outlier removal, KNN missing value imputation, and BERT semantic extraction. The sequence features of the behavior modality capture user rights and preferences, the regional coding of the environment modality anchors regional rule constraints, the emotional scores of the emotional intent modality reflect user demand tendencies, and the compliance labels of the business rule modality solidify the compliance bottom line. The comprehensive collection of these four types of data avoids decision-making biases caused by single-dimensional data from the source. Building upon this foundation, the equity-business relevance matrix quantifies modal contribution through mutual information and business objective weights. The attention mechanism of provincial correlation adjustment strengthens the binding between regional rules and multimodal features, ensuring that the fused features not only meet quality standards (efficiency ≥ 98%) but also possess strong business orientation, providing high-value input for subsequent modeling. In the user segmentation process, equity adaptive density clustering is used to achieve automatic cluster identification through local density calculation and time-series LTV prediction. This not only adapts to the "core dense + edge sparse" distribution characteristics of equity users but also ensures that the segmentation results can respond to dynamic changes in the scenario and anchor the long-term value of users through the activity-based temporary segmentation and long-term value merging mechanism, providing accurate user profile support for subsequent segmented exclusive services.

[0088] Based on the high-quality fusion features and accurate clustering results output by S1, S2 completes the full-link value transformation through multi-module collaboration: In the rights and interests impact modeling stage, the construction of a set of rights and interests-specific dimensions achieves full coverage of basic features, scenarios, and clustering features. Bayesian weight inference treats weights as random variables and introduces Dirichlet priors, effectively reducing the uncertainty of weight estimation. Furthermore, the reinforcement learning reward function and adaptive learning rate that integrate LTV further solve the short-term benefit orientation problem of traditional modeling, so that the impact score can accurately reflect the matching degree between rights and interests and users, while also taking into account the long-term value improvement of users, providing a reliable basis for matching high-value users. The customized rights and interests mathematical model transforms the rigid compliance requirements and flexible integral constraints of rights and interests business into calculable mathematical objectives by applying mathematical constraints to the weights of core dimensions and constructing a segmented reward mechanism, avoiding the disconnect between pure mathematical modeling and business scenarios, and ensuring that the model output always fits the actual rights and interests operation.

[0089] The prediction module employs a hybrid prediction model combining gradient boosting, meta-learning, and adversarial learning. It uses a generator to simulate multi-scenario matching data and a discriminator to enhance feature discrimination, improving the model's generalization ability to complex benefit scenarios. For a small number of samples with newly added temporary benefits, the meta-gradient accumulation mechanism achieves rapid parameter updates by accumulating gradients across multiple tasks, overcoming the adaptation bottleneck of traditional models that rely on large amounts of data. Combined with the Bayesian probabilistic localization of the benefit-specific error attribution network, it can accurately trigger corrective strategies such as weight adjustments or rule synchronization, ensuring the stability and reliability of the prediction results. In the rule optimization stage, the Bayesian compliance attribution network quantifies the impact of multi-factor failures. The multi-objective optimization module uses the NSGA-III algorithm to find the Pareto optimal solution, achieving a balance between conversion efficiency, compliance rate, and cost. The segmented rule design (relaxed point constraints for high-value members and simplified rules for new customers) further enhances the user adaptability of the rules. Rules that have passed compliance verification are synchronized to the provincial cache node, ensuring regional adaptability while reducing response latency.

[0090] During conflict resolution, a priority calculation model integrating user intent strength is used to quantify user demand tendencies by measuring the similarity between current queries and historical interests, making conflict ranking more aligned with core user needs. The combination of a segmented customized solution library and constrained reinforcement learning enhances acceptance through differentiated solutions such as high-value membership points coupons and new customer points acceleration packages, while mitigating business risks through compliance and cost constraints. The closed-loop guidance, achieved through segment upgrades and rule reductions after task completion, transforms the resolution effect into user value. The scenario-based feedback iteration module triggers real-time weight fine-tuning based on core behaviors and differentiated update strategies for peak and off-peak scenarios, enabling the end-to-end model and rules to dynamically respond to changes in user behavior and scenario pressure. Weekly updates to the modality matrix and segmentation model ensure continuous adaptation between technical solutions and business needs.

[0091] The collaborative closed loop of each module further amplifies the technical effect: the high-quality data of S1 provides reliable input for S2 modeling; the prediction error of S2 and the rule update results drive the modality matrix update of S1 in reverse; and the secondary adjustment of the dimension weights after the provincial rule subset update achieves the adaptation and calibration of "rule-grouping-feature". This closed-loop mechanism enables the method to continuously adapt to the core characteristics of the rights and interests business: "fragmented rules, high-frequency changes in scenarios, and different group values", ultimately deriving three core effects: first, the dynamic adaptation capability of fragmented rules, solving the problem of the dispersion of regional and scenario rules; second, the precise matching of rights and interests with users, improving service targeting through grouping and influence modeling; and third, the two-way improvement of business value and user value, improving operational efficiency through cost constraints and conversion optimization, and enhancing user stickiness through differentiated services and guidance closed loop.

[0092] In summary, the end-to-end operation flow of the multimodal full-link collaborative method for rights and benefits supermarkets provided in this embodiment can be summarized as follows:

[0093] Step 1 (Data Foundation): Collect four types of multimodal data: behavior, environment, emotional intent, and business rules. After preprocessing, the data is integrated into a unified feature vector through the rights-business correlation matrix and the provincial attention mechanism. At the same time, core user groups are generated based on rights-adaptive density clustering.

[0094] Step 2 (Core Modeling): Construct a set of rights-specific dimensions including basic, scenario, and group dimensions; determine dynamic weights using Bayesian weight inference; and complete the rights influence modeling by combining LTV prediction and reinforcement learning reward function; output matching prediction results through a hybrid prediction model (gradient boosting + meta-learning + adversarial learning).

[0095] Step 3 (Optimization and Mediation): Generate group-specific rules based on the rights compliance attribution network and multi-objective optimization (NSGA-III), and synchronize them to the provincial rule subset after verification; when user rights conflict, call the group-customized solution library according to the conflict priority, and complete mediation and guidance through constrained reinforcement learning.

[0096] Step 4 (Continuous Iteration): Fine-tune the weights in real time based on core user behaviors, update the model and rules according to the rhythm of the scenario, and form a closed loop of "data-model-rules-feedback" through the provincial rule synchronization verification mechanism, so as to achieve dynamic and accurate adaptation of fragmented rights and interests rules and continuous improvement of user value.

[0097] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A multimodal end-to-end collaborative method for rights and benefits supermarkets, characterized in that, In response to the fragmented nature of rights and benefits business rules, the high frequency of updates across multiple scenarios, and the differences in value among user groups, The data collected includes multimodal data with rights-specific dimensions. These rights-specific dimensions refer to feature dimensions designed specifically for the rights supermarket business and distinct from general service dimensions. These include dimensions directly related to the rights business, such as rights redemption sequences and points consumption characteristics. This data includes behavioral modalities, environmental modalities, emotional intent modalities, and business rule modalities. The behavioral modality includes the sequence of rights and benefits redemption and the characteristics of points consumption; the environmental modality includes provincial regional tags and terminal types; the emotional intent modality includes the semantic characteristics of rights and benefits consultation; and the business rule modality includes provincial compliance tags and points thresholds. Based on this data, an end-to-end link of "data fusion - modeling - prediction - rule optimization - conflict resolution - feedback iteration" is constructed to achieve dynamic adaptation of fragmented rules.

2. The multimodal end-to-end collaborative method for rights and interests supermarkets according to claim 1, characterized in that, The multimodal data fusion process includes constructing a rights-specific modality system. The system includes a rights combination selection sequence for the behavioral modality, provincial regional codes for the environmental modality, emotional scores for consultation texts in the emotional intention modality, and rights validity period tags for the business rules modality. Based on this system, an equity-business relevance matrix is ​​constructed, and a feature fusion mechanism with provincial relevance adjustment is adopted to output fusion features that conform to provincial rules.

3. The multimodal end-to-end collaborative method for rights and interests supermarkets according to claim 2, characterized in that, The process of modeling the impact of rights and interests includes constructing a set of rights and interests-specific dimensions. This set includes basic dimensions, scenario-specific dimensions, and segment-specific dimensions. The basic dimensions include membership level and preferred combination of benefits. The scenario-specific dimensions are feature dimensions preset for different benefit scenarios. The segment-specific dimensions are feature dimensions customized for different user segments. The dimension weights are determined by using rights-oriented Bayesian weight inference, and dynamic thresholds are determined based on user group value, scenario importance and the number of rights. The impact score is adjusted by combining the tolerance for points gap.

4. The multimodal end-to-end collaborative method for rights and interests supermarkets according to claim 3, characterized in that, The prediction module employs a hybrid prediction model that integrates gradient boosting, meta-learning, and adversarial learning. The generator produces simulated matching data based on provincial regional rule labels and integral constraints, while the discriminator distinguishes between real and generated data. In scenarios with a small number of samples, a meta-gradient accumulation mechanism is used to accelerate adaptation. The cause of failure is located and a calibration strategy is triggered by a rights-specific error attribution network. The scenarios with a small number of samples are rights promotion scenarios where the number of existing samples does not exceed a preset sample number threshold.

5. The multimodal end-to-end collaborative method for rights and interests supermarkets according to claim 3, characterized in that, The rule optimization process includes building a rights compliance attribution network. The Pareto optimality rule is solved through a multi-objective optimization module, with optimization objectives including conversion efficiency, compliance rate, and equity cost. Customized rules are generated for different user segments. The points gap constraint is relaxed for high-value members, and the constraints are simplified for new customer rules. High-value members are defined as users whose long-term predicted value is no less than the preset user value. After the candidate rules pass compliance verification and matching verification, they are synchronized to the corresponding provincial rule subset.

6. The multimodal end-to-end collaborative method for rights and interests supermarkets according to claim 3, characterized in that, The conflict resolution process includes calculating conflict priority based on core rights priorities. This priority resets the effective equity options to the highest level, followed by membership level, provincial region, and points. We have built a customized solution library for different user groups. High-value membership programs include points vouchers and extended benefits, while new customer programs include points acceleration packages and simplified rules. High-value membership programs correspond to high-value user groups, which are user groups whose long-term predicted value is no less than the preset user value. A constrained reinforcement learning optimization strategy is adopted to establish a mediation-guidance-upgrade closed loop. After the user's task is completed, the group is automatically upgraded and the rule threshold is lowered.

7. The multimodal end-to-end collaborative method for rights and interests supermarkets according to claim 2, characterized in that, The user segmentation process includes extracting a set of behavioral characteristics specific to user rights. This collection includes the frequency of claiming benefits, the proportion of points spent, the click-through rate of provincial benefits, and the predicted long-term value. Core clusters are generated using an equity-adaptive density clustering algorithm. Temporary clusters are triggered by activity scenarios, and clusters are merged based on long-term value after the scenario ends. Resource allocation is implemented based on the value of user groups, and high-value groups are matched with high-quality rights and benefits rules. High-value groups are user groups whose long-term value prediction is no less than the preset user value.

8. The multimodal end-to-end collaborative method for rights and interests supermarkets according to claim 3, characterized in that, It also includes the step of customizing the mathematical model of equity. Mathematical constraints are imposed on the weights of rights-specific dimensions, and the weight proportions of core dimensions are strengthened through prior distribution. Establish a tiered reward mechanism, with rigid constraints setting reward values ​​based on achievement of targets, and flexible constraints setting tiered rewards based on points gaps. A long-term value model is constructed based on the time series of equity consumption, and the upper limit of group cost is transformed into a constraint embedded in the objective function.

9. The multimodal end-to-end collaborative method for rights and interests supermarkets according to claim 4, characterized in that, It also includes scenario-based feedback iteration steps. When a user triggers a core behavior that infringes their rights, the impact weight and multimodal feature weight are adjusted in real time. In peak scenarios, the fine-tuning cycle of the hybrid prediction model is shortened and the learning rate is improved; during off-peak periods, data and driving rules are updated in stages for optimization. The modality matrix and cluster model are updated weekly, and emergency feedback is triggered when the prediction error exceeds the standard.

10. The multimodal end-to-end collaborative method for rights and interests supermarkets according to claim 5, characterized in that, Also includes: Provincial rule subset synchronization and verification steps: After the provincial rule subset is updated, it is synchronized to the corresponding regional rights data cache node. The adaptability of the updated rules and user groups is verified through the feature comparison mechanism. When the adaptation deviation exceeds the set range, the dimension weight is adjusted.