Member order processing method, device, equipment and storage medium
By constructing a member graph structure and a deep learning model, and combining data from multiple platforms for member order processing, the problems of multi-platform resource integration and member value assessment were solved, achieving efficient order allocation and personalized benefits distribution, and improving the accuracy of member management and order processing.
Patent Information
- Application Number
- CN202411805332.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-12-10
AI Technical Summary
Traditional member order processing methods cannot fully utilize multi-platform data resources, resulting in inaccurate member value assessment, inefficient order allocation, and a lack of deep integration of member and order characteristics, making it difficult to achieve accurate order classification and prioritization, as well as optimal resource allocation and personalized adjustments to member benefits.
By collecting member information from multiple platforms to construct a member graph structure, extracting node features, and combining graph neural networks and soft attention mechanisms, the complex relationships between member attributes are captured. Adversarial bidirectional encoders and bidirectional LSTM networks are used to fuse member-order features for deep integration. Order classification and priority allocation are performed based on multilayer perceptron networks and support vector regression models. A multi-platform collaborative integer programming model is used for order allocation, combined with entropy weighting and analytic hierarchy process (AHP) for member value assessment. A dynamic rights allocation model based on deep Q-networks is used for personalized rights combinations.
It improved the accuracy and comprehensiveness of member characteristic representation, achieved efficiency and accuracy in order processing, optimized order resource allocation, enhanced the reliability of member value assessment and the real-time nature of personalized benefits allocation, and improved member satisfaction and loyalty.
Smart Images

Figure CN119295188B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of order processing technology, and in particular to a method, apparatus, equipment and storage medium for processing member orders. Background Technology
[0002] Traditional membership order processing methods are often limited to a single platform, failing to fully utilize multi-platform data resources, leading to problems such as inaccurate membership value assessment and inefficient order allocation. Furthermore, existing order processing systems generally lack deep integration of membership and order characteristics, making it difficult to achieve accurate order classification and prioritization.
[0003] Furthermore, in a multi-platform collaborative environment, effectively allocating order resources and balancing the processing capacity and expertise of each platform has become a pressing challenge. Traditional order allocation methods often employ simple rule matching or random allocation, failing to achieve optimal resource allocation and easily leading to situations where some platforms are overloaded while others are idle. On the other hand, membership benefit allocation strategies typically rely on static rules, lacking flexibility and personalization, and unable to adjust benefit combinations in a timely manner based on members' dynamic value and preferences. This not only affects member satisfaction but also limits the potential for businesses to enhance member value through targeted marketing. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and storage medium for processing member orders. This invention can comprehensively utilize data from multiple platforms to achieve intelligent order processing and dynamic rights allocation.
[0005] In a first aspect, the present invention provides a member order processing method, the member order processing method comprising:
[0006] Collect member information from multiple platforms and construct a member graph structure to extract node features, thereby obtaining multi-dimensional member feature vectors;
[0007] The multi-dimensional member feature vector and the order data to be processed are fused to obtain the member-order fused feature vector;
[0008] Orders are classified based on the member-order fusion feature vector to obtain the order classification results and processing priorities.
[0009] Based on the order classification results and the processing priorities, a multi-platform collaborative integer programming and solution is performed to obtain a cross-platform order allocation scheme.
[0010] Based on the cross-platform order allocation scheme and order processing results, member value is evaluated to obtain dynamic member value scores and hierarchical classification information.
[0011] The dynamic member value score and the hierarchical classification information are input into a dynamic rights allocation model based on a deep Q-network to optimize the rights allocation strategy and obtain a personalized combination of member rights.
[0012] Secondly, the present invention provides a member order processing device, the member order processing device comprising:
[0013] The data acquisition module is used to collect member information from multiple platforms and construct a member graph structure to extract node features and obtain multi-dimensional member feature vectors.
[0014] The fusion module is used to perform feature fusion on the multi-dimensional member feature vector and the order data to be processed to obtain the member-order fusion feature vector;
[0015] The classification module is used to classify orders based on the member-order fusion feature vector, and obtain the order classification result and processing priority;
[0016] The planning module is used to perform multi-platform collaborative integer programming and solving based on the order classification results and the processing priorities to obtain a cross-platform order allocation scheme;
[0017] The evaluation module is used to evaluate member value based on the cross-platform order allocation scheme and order processing results, and obtain dynamic member value scores and hierarchical classification information.
[0018] The strategy optimization module is used to input the dynamic member value score and the hierarchical division information into a dynamic rights allocation model based on a deep Q network to optimize the rights allocation strategy and obtain a personalized member rights combination.
[0019] A third aspect of the present invention provides a member order processing device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the member order processing device to perform the above-described member order processing method.
[0020] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the above-described member order processing method.
[0021] The technical solution provided by this invention employs graph neural networks and soft attention mechanisms for member feature extraction, effectively capturing the complex relationships between member attributes and improving the accuracy and comprehensiveness of member feature representation. By utilizing adversarial bidirectional encoders and bidirectional LSTM networks for member-order feature fusion, deep integration of member information and order data is achieved. The order classification method based on multilayer perceptron networks and support vector regression models accurately identifies order types and rationally allocates processing priorities, improving the efficiency and accuracy of order processing. A multi-platform collaborative integer programming model is used for order allocation, fully considering the processing capabilities and expertise of each platform, achieving optimal allocation of order resources and improving overall order processing efficiency. Combining entropy weighting and analytic hierarchy process (AHP) for member value assessment considers both the distribution characteristics of objective data and incorporates expert experience, making the member value assessment results more comprehensive and reliable. A dynamic rights allocation model based on deep Q-networks enables personalized allocation and real-time optimization of member rights, allowing timely adjustment of rights combinations based on members' dynamic value and preferences, improving member satisfaction and loyalty. This invention improves the efficiency and accuracy of member management and order processing, providing enterprises with a comprehensive member order processing solution.
[0022] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of one embodiment of the member order processing method in this invention;
[0025] Figure 2 This is a schematic diagram of one embodiment of the member order processing device in this invention;
[0026] Figure 3 This is a schematic diagram of one embodiment of the member order processing device in this invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] The terms "comprising" and "having," and any variations thereof, used in the embodiments of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0029] To facilitate understanding of this embodiment, a member order processing method disclosed in this invention will first be described in detail. For example... Figure 1 As shown, this method includes the following steps:
[0030] 101. Collect member information from multiple platforms and construct a member graph structure to extract node features and obtain multi-dimensional member feature vectors;
[0031] It is understood that the executing entity of this invention can be a member order processing device, a terminal, or a server; no specific limitation is made here. This embodiment of the invention will be described using a server as an example.
[0032] Specifically, data cleaning and format standardization are performed on member information from multiple platforms to remove duplicate data, handle missing values, and eliminate invalid data items. Simultaneously, the data formats of each platform are standardized to conform to a unified data standard, resulting in a standardized member dataset. Based on this standardized dataset, a member graph structure is constructed. The graph structure uses members as nodes and the relationships between member attributes as edges. Through the setting of nodes and edges, a graph representing members and their relationships is formed. Each node represents a member, and the edges represent the associations between members, such as interaction frequency, similar purchasing behaviors, etc., thus forming a member network graph structure reflecting multi-dimensional relationships. The member graph structure is represented by a sparse matrix, effectively compressing the information in the graph and improving storage and computation efficiency. The member graph structure is encoded using an adjacency matrix and a feature matrix, resulting in a graph structure encoding matrix. The adjacency matrix represents the connections between nodes, while the feature matrix represents the attribute features of each node. These two matrices are combined to form the overall encoded representation of the graph structure, thus completely inputting the structured information of the graph into the subsequent neural network. The graph structure encoding matrix is then input into a self-attention gated graph neural network for processing. Graph neural networks (GNNs) can effectively extract and aggregate features from nodes in a graph. Through multi-layer graph convolution operations, node features are updated layer by layer, gradually aggregating information from neighboring nodes to obtain an updated node feature matrix. In this process, a self-attention mechanism is used to enhance the recognition of node importance during graph convolution, allowing the network to assign different weights to different neighboring nodes during information aggregation, thus focusing on important features. Based on the updated node feature matrix, attention weights between nodes are calculated to capture the dynamic relationships between them. The attention weights are calculated based on the updated node feature matrix, and the attention scores are normalized using the softmax function to obtain a node attention weight matrix. The node attention weight matrix is multiplied by the updated node feature matrix to obtain a weighted node feature matrix. The weighted node feature matrix is then summed in column vectors and normalized using the L2 norm to obtain a graph-level representation vector. Column vector summation aggregates node-level information to the overall graph level, resulting in a feature vector describing the entire graph. L2 norm normalization is used to standardize this vector, ensuring that its magnitude is not affected by the feature dimension. The graph-level representation vector is concatenated with the member's basic attribute vector. A fully connected layer is then used for feature fusion and dimensionality reduction to obtain the final multi-dimensional member feature vector. Feature fusion through the fully connected layer combines graph structure information and member's basic attributes to generate a comprehensive feature vector containing both graph structure and attribute features. Simultaneously, dimensionality reduction effectively reduces the feature dimension, lowers computational complexity, and avoids overfitting in high-dimensional data.
[0033] 102. Perform feature fusion on the multi-dimensional member feature vector and the order data to be processed to obtain the member-order fused feature vector;
[0034] Specifically, the multi-dimensional member feature vector is expanded. Through copying and zero-padding operations, the multi-dimensional member feature vector is expanded to the same length as the order data sequence to be processed, resulting in the expanded member feature vector. The expanded member feature vector is then concatenated with the order data to be processed along the feature dimension, yielding the initial fused feature sequence. In this fused feature sequence, each time step contains the member's features and their corresponding order features. The initial fused feature sequence is then input into a bidirectional Long Short-Term Memory (LSTM) network. The bidirectional LSTM consists of a forward LSTM and a backward LSTM, each consisting of two layers, each containing 128 hidden units, and using the tanh activation function to introduce non-linearity. Through bidirectional processing, the temporal dependencies between the two directions in the fused feature sequence are captured. The forward LSTM feature sequence and the backward LSTM feature sequence extract the potential temporal correlation between member features and order features from the two time directions, respectively. The outputs of the forward and backward LSTM feature sequences at the last hidden layer are concatenated to obtain a bidirectional LSTM feature sequence. This sequence not only contains global temporal information of the input sequence but also effectively combines the contextual information from both the forward and backward directions. The bidirectional LSTM feature sequence is then fed into a multi-head self-attention layer. This layer extracts important information from the feature sequence, using eight attention heads, each with a dimension of 64. A query, key, and value matrix is generated through linear transformation. Then, the self-attention feature sequence is obtained by calculating scaled dot product attention and performing softmax normalization. The multi-head attention mechanism captures hidden feature relationships in the sequence from different angles and scales, allowing the model to focus more precisely on the correlations between different time steps. Parallel computation of multiple attention heads simultaneously captures various information patterns in the input sequence, enhancing the model's feature representation capability. The self-attention feature sequence is then fed into a two-layer feedforward neural network for feature processing. The first layer uses the ReLU activation function to introduce non-linearity to enhance the model's expressive power; the second layer uses a linear activation function to linearly transform the features. The hidden layer dimension is set to 512 to ensure the network can process sufficiently rich information and perform complex feature transformations to obtain the encoder output feature sequence. The encoder's feedforward network fuses and transforms the features extracted by LSTM and self-attention, improving the abstractness and effectiveness of the features. To enhance the discriminative ability of the encoder's extracted features, the encoder output feature sequence is input into an adversarial discriminator for optimization. The adversarial discriminator consists of three one-dimensional convolutional layers with kernel sizes of 3, 4, and 5, designed to capture local features at different time steps. Each convolutional layer has 100 output channels to ensure sufficient capacity for extracting rich feature patterns. The discriminator is followed by a fully connected layer to output the binary classification result.Through adversarial training, the encoder and discriminator engage in a game of strategy, enabling the encoder to learn more discriminative and generalizable features, resulting in an adversarially optimized feature sequence. Max pooling is then performed on the adversarially optimized feature sequence along the time dimension, and the maximum value in each feature dimension is selected to obtain the member-order fusion feature vector.
[0035] 103. Classify orders based on the member-order fusion feature vector to obtain order classification results and processing priorities;
[0036] Specifically, the member-order fused feature vector is input into a multilayer perceptron network. This multilayer perceptron network consists of three hidden layers: the first layer contains 512 neurons, the second layer contains 256 neurons, and the third layer contains 128 neurons. A ReLU activation function is used between each layer to introduce non-linearity, enabling the model to learn more complex feature patterns. The final layer uses a linear activation function to maintain feature continuity, resulting in a representation of order features. The multilayer structure of the multilayer perceptron effectively abstracts the input fused feature vector layer by layer, extracting deep features representing orders and reflecting the correlation between orders and members, as well as the various attributes of the orders. L2 regularization is applied to the order feature representation. The L2 norm of the order feature vector is calculated, and the original feature vector is divided by this L2 norm to obtain a normalized order feature vector. L2 regularization ensures consistent scale of the feature vectors, preventing adverse effects on subsequent classification due to scale differences in different feature values. Simultaneously, normalization can, to some extent, avoid overfitting and improve the model's generalization ability. The normalized order feature vectors are input into a softmax classifier. The softmax classifier calculates the probability of each category using the softmax function, obtaining the order category probability distribution. The softmax function maps the feature vectors to a probability space, so that the output is the probability value for each possible category, with the sum of probabilities being 1. Based on the order category probability distribution, the argmax function is used to find the index of the category with the highest probability value, and the category corresponding to this index is taken as the order classification result. Entropy is calculated on the order category probability distribution to assess the uncertainty of the classification result. Entropy is an important concept in information theory, used to measure the degree of uncertainty of a probability distribution. For the order category probability distribution, its entropy value is calculated to obtain a scalar value representing the classification uncertainty. The scalar entropy value is compared with a predefined threshold. If the entropy value is less than the threshold, it indicates that the classification is highly certain, and the order is considered to belong to a definite category; conversely, if the entropy value is greater than the threshold, it indicates that the model has high uncertainty in classifying the order, and it is classified as an uncertain category. For orders with a definite category, a predefined priority mapping table is queried based on the order classification result. This priority mapping table is a rule that maps each order category to an integer priority level from 1 to 5, predefining the processing priority for each category. For example, orders in certain important categories are assigned higher priorities to ensure they are processed first. The initial processing priority for an order is obtained by querying the priority mapping table. For orders with uncertain classifications, a more refined priority determination method is used. The normalized order feature vectors are input into a pre-trained support vector regression model. This support vector regression model uses the RBF kernel function to capture the complex nonlinear relationships of order features in the feature space.A continuous priority score is obtained through a support vector regression model, reflecting the importance of the order. This continuous score is then transformed into a discrete priority level of 1 to 5 by setting interval boundaries, providing a more flexible and refined priority determination strategy for orders with uncertain classifications.
[0037] 104. Based on the order classification results and processing priorities, perform multi-platform collaborative integer programming and solve the problem to obtain a cross-platform order allocation scheme;
[0038] Specifically, each order is assigned a weight coefficient based on its classification and processing priority. The formula for calculating the weight coefficient is: Weight = α * Classification Score + β * Priority, where α and β are preset balancing parameters used to balance the influence of the order's classification score and processing priority on the final weight. By calculating the weight coefficient, the relative importance of each order is quantified, resulting in an order weight list. A platform-order fit matrix is constructed based on the processing capabilities and areas of expertise of each platform. The element aij in the matrix represents the fit of platform i with order j, ranging from 0 to 1, used to quantify the degree of matching between the platform and the order. For example, a platform that excels at handling a certain type of order will have a higher fit, while its fit for other types of orders it is not good at will be lower. Based on the order weight list and the platform-order fit matrix, an objective function for order allocation is constructed, with three constraints set. The objective function aims to maximize the overall order processing efficiency, typically defined as the sum of the weighted scores of all orders across all platforms. The weighted score of each order is obtained by multiplying its weight by its platform suitability. By maximizing the weighted score, it is ensured that important orders are assigned to the most suitable platform for processing. While constructing the objective function, three constraints are set to ensure the rationality of order allocation. The first constraint is that each order can only be assigned to one platform, ensuring the independence of order allocation; the second constraint is that the order processing volume of each platform cannot exceed the platform's maximum processing capacity, thus avoiding overload; the third constraint is that the cross-platform allocation ratio of preset order types cannot exceed a preset threshold, which is used to ensure that the allocation of certain specific order types meets business needs and strategy requirements. Combining the objective function and the three constraints forms a multi-platform collaborative integer programming model. The core of the multi-platform collaborative integer programming model is the decision variable, which represents the weighted score of each order on each platform. The model's goal is to maximize the overall efficiency of order processing through appropriate order allocation. The design of the multi-platform collaborative integer programming model can comprehensively consider the order weight, platform suitability, and the processing capacity and order type allocation requirements of each platform to achieve a globally optimal solution for cross-platform order allocation. The multi-platform collaborative integer programming model is input into the interactive optimization solver for solution. Before inputting the model into the solver, several key parameters are set, including the maximum number of iterations, the convergence threshold, and the time limit. The maximum number of iterations controls the number of iterations in the optimization algorithm to prevent the solution process from entering an infinite loop; the convergence threshold determines whether the solution has reached the optimal solution or is close enough to it; and the time limit limits the solution time to prevent excessively long solution times from affecting the overall system performance. After setting the parameters, the interactive optimization solver is started to begin the solution process. During the solution process, the interactive optimization solver continuously adjusts the allocation of orders across the various platforms to maximize the value of the objective function while ensuring that all constraints are met.After the solver completes the solution, the optimal values of the decision variables are extracted. Order-platform pairs with a decision variable of 1 are extracted and combined into a cross-platform order allocation scheme. The resulting cross-platform order allocation scheme achieves reasonable and efficient order distribution, fully utilizing the processing capacity of each platform while maximizing the overall efficiency of order processing.
[0039] 105. Based on the cross-platform order allocation scheme and order processing results, evaluate member value to obtain dynamic member value scores and hierarchical classification information;
[0040] Specifically, based on the cross-platform order allocation scheme and the order processing results of pending orders, various indicators for member value assessment are extracted, including consumption frequency, consumption amount, order completion rate, platform activity, and customer service demand. These indicators reflect members' consumption behavior, activity level, and interaction with the platform, thus constructing a raw indicator matrix for assessing member value. The raw indicator matrix is normalized to eliminate dimensional differences between indicators, ensuring comparison and calculation on the same scale. The range transformation method is used to convert indicators of different dimensions into dimensionless values, resulting in a standardized indicator matrix. Based on the standardized indicator matrix, the information entropy of each indicator is calculated. Information entropy measures the uncertainty and information content of each indicator throughout the assessment process. The entropy weight method is used to calculate the objective weight of each indicator based on its information entropy, resulting in an objective weight vector. The entropy weight method measures the importance of each indicator using objective data, reflecting its contribution to the overall value in actual member behavior. A hierarchical analysis model is constructed, dividing the member value assessment indicators into three levels: the target layer, the criterion layer, and the indicator layer. In the target layer, the ultimate goal is to evaluate member value. In the criteria layer, the main aspects influencing member value are identified, such as consumption behavior, platform activity, and customer service needs. The indicator layer contains specific evaluation indicators, such as consumption frequency, consumption amount, and order completion rate. A judgment matrix is established to compare the relative importance of each indicator. The judgment matrix is constructed using expert scoring to represent pairwise comparisons between indicators. A consistency check is performed on the judgment matrix to ensure its logical consistency. The maximum eigenvalue λmax of the judgment matrix is calculated, and the consistency index CI is calculated using the formula CI=(λmax-n) / (n-1), where n is the order of the judgment matrix. The consistency ratio CR=CI / RI is calculated, where RI is the random consistency index. If the consistency ratio CR is less than 0.1, the judgment matrix is considered to have passed the consistency check, indicating good logical consistency in subjective evaluation. If CR ≥ 0.1, the values in the judgment matrix need to be adjusted until the consistency check is passed. Based on the judgment matrix that has passed the consistency check, the subjective weights of each indicator are calculated using the eigenvalue method, resulting in a subjective weight vector. Subjective weights reflect experts' subjective judgments on the relative importance of each indicator, incorporating human experience and domain knowledge. To comprehensively consider objective data and expert judgment, the objective and subjective weight vectors are weighted and averaged to obtain a comprehensive weight vector. Based on the comprehensive weight vector and the standardized indicator matrix, each member's dynamic value score is calculated. The comprehensive weight vector is then weighted and summed with the indicator data of each member in the standardized indicator matrix to obtain each member's total value score, reflecting the member's comprehensive performance in multiple aspects such as consumption and activity. Member value scores are clustered to obtain hierarchical classification information.Members are grouped according to their value scores using clustering algorithms (such as K-means clustering) to divide them into different tiers.
[0041] 106. Input the dynamic member value score and hierarchical classification information into the dynamic rights and benefits allocation model based on deep Q network to optimize the rights and benefits allocation strategy and obtain a personalized member rights and benefits combination.
[0042] Specifically, the dynamic member value score and tier classification information are combined into a state vector S. State vector S contains the member's value score, member tier, and historical usage rate of benefits. State vector S is input into a deep Q-network containing three hidden layers with 256, 128, and 64 neurons respectively, using the ReLU activation function between layers. The ReLU activation function introduces non-linearity, enabling the model to better learn the complex relationships between states and actions. Through this deep structure, the deep Q-network extracts high-dimensional features from the state vector, outputting a Q-value matrix. Each element in the Q-value matrix represents the expected reward obtained by taking a benefit allocation action in the current state. Based on the Q-value matrix, an ε-greedy strategy is used to select the benefit allocation action. The ε-greedy strategy is an exploration and exploitation strategy that randomly selects an action with probability ε and chooses the action with the highest current Q-value with probability (1-ε). To better adapt to the needs of different membership levels, the ε value is dynamically adjusted based on the member's level. For higher-level members, the ε value is smaller to allow for more frequent selection of actions with the highest expected returns, thus improving the accuracy of benefits allocation. For lower-level members, the ε value is relatively larger to explore more combinations of benefits and uncover potential optimal solutions. This process yields the initial benefits allocation plan. A feasibility check is performed on the initial plan to ensure that the allocated benefits meet the preset benefits allocation rules and resource constraints. These rules and constraints include the maximum number of benefits, platform resource limitations, and the types of benefits each member can obtain within a certain period. By iteratively adjusting the benefits items in the plan, the initial plan is optimized while satisfying all rules and constraints to obtain an effective benefits allocation plan. This effective plan is then applied to specific members, and member feedback on the use of allocated benefits is collected within a fixed time window. The actual usage of benefits by members is a crucial basis for evaluating the effectiveness of the benefits allocation strategy. An instant reward value is calculated based on a preset reward function. The reward function is designed based on factors such as the frequency of member benefit usage, satisfaction, and further consumption behavior on the platform. The immediate reward value reflects the effectiveness of the current benefits allocation scheme and member satisfaction. Accumulated rewards are calculated based on the immediate reward value and a preset discount factor. The discount factor is used to balance the relationship between current rewards and future rewards, thereby optimizing the member benefits allocation strategy in the long term. Each experience sample (i.e., a quadruple of state, action, reward, and next state) is stored in an experience replay pool with a capacity of 10,000. The experience replay pool stores historical experience samples. Through random sampling, the model training process is independent of the order of samples, reducing the correlation between samples and improving the model's stability and generalization ability. Mini-batch experience data of size 64 is randomly sampled from the experience replay pool to update the parameters of the deep Q-network. The parameters of the target Q-network are updated using a double-Q learning algorithm.Dual-Q learning effectively reduces overestimation in Q-value estimation and improves model stability and convergence speed by introducing two independent Q-networks to estimate action value and make action selection, respectively. After each parameter update, the updated target Q-network parameters are copied to the evaluation Q-network using a soft update. The soft update, with a small learning rate, allows the target network to gradually approach the evaluation network, ensuring model training stability and gradual convergence. Upon reaching a preset convergence condition, a personalized membership benefit package is output. This personalized benefit package is based on members' dynamic value scores and hierarchical classification information, combined with historical benefit usage feedback, and continuously optimized through deep reinforcement learning to maximize member satisfaction with benefit usage and improve the platform's overall operational efficiency.
[0043] In this invention, by employing graph neural networks and soft attention mechanisms for member feature extraction, the complex relationships between member attributes can be effectively captured, improving the accuracy and comprehensiveness of member feature representation. Member-order feature fusion is achieved using adversarial bidirectional encoders and bidirectional LSTM networks, realizing deep integration of member information and order data. The order classification method based on multilayer perceptron networks and support vector regression models can accurately identify order types and rationally allocate processing priorities, improving the efficiency and accuracy of order processing. Order allocation is performed using a multi-platform collaborative integer programming model, fully considering the processing capabilities and expertise of each platform, achieving optimal allocation of order resources and improving overall order processing efficiency. Member value assessment is performed by combining entropy weighting and analytic hierarchy process, considering both the distribution characteristics of objective data and incorporating expert experience, making the member value assessment results more comprehensive and reliable. A dynamic rights allocation model based on deep Q-networks enables personalized allocation and real-time optimization of member rights, allowing for timely adjustment of rights combinations based on members' dynamic value and preferences, improving member satisfaction and loyalty. This invention improves the efficiency and accuracy of member management and order processing, providing enterprises with a comprehensive member order processing solution.
[0044] In one specific embodiment, the process of performing step 101 may specifically include the following steps:
[0045] Data cleaning and format unification processing of member information from multiple platforms are performed to obtain a standardized member dataset. Based on the standardized member dataset, member attributes are set as nodes and the relationships between attributes are set as edges to construct a member graph structure.
[0046] The member graph structure is represented by a sparse matrix. The member graph structure information is encoded by the adjacency matrix and the feature matrix to obtain the graph structure encoding matrix. The graph structure encoding matrix is then input into the self-attention gated graph neural network. The node features are updated through multi-layer graph convolution operations to obtain the updated node feature matrix.
[0047] The attention weights between nodes are calculated based on the updated node feature matrix. The attention scores are normalized using the softmax function to obtain the node attention weight matrix. The node attention weight matrix is then multiplied with the updated node feature matrix to obtain the weighted node feature matrix.
[0048] The column vectors of the weighted node feature matrix are summed and then normalized using the L2 norm to obtain a graph-level representation vector. This graph-level representation vector is then connected to the member information attribute vector, and feature fusion and dimensionality reduction are performed through a fully connected layer to obtain a multi-dimensional member feature vector.
[0049] Specifically, member data from different platforms undergoes data cleaning and format standardization. Data cleaning removes noise, fills in missing values, and merges duplicate data, ensuring semantic and format consistency across all platforms. Format standardization converts data from different platforms into a unified format, such as converting data from different units to a consistent unit, and mapping all member information to a common structured data table, forming a standardized member dataset. Based on this standardized dataset, a member graph structure is constructed. In this graph, member attributes are set as nodes, and relationships between attributes are set as edges. For example, if member A and member B share common purchase records or similar browsing behavior, this is represented as an edge in the graph; member A's attributes (such as purchase frequency and browsing habits) are represented as nodes. In the member graph structure, each node represents a member, and the existence of edges indicates a relationship between these members, such as shared interests or similar consumption behaviors. A sparse matrix representation of the member graph structure effectively compresses the information content, improving storage and computational efficiency. Adjacency matrices are used. Describes the connectivity of a graph, where Represents a node and nodes There are edges connecting them, and This indicates that there is no direct relationship between the two. Simultaneously, through the feature matrix... This represents the attribute characteristics of each node, such as purchase frequency and amount, resulting in a matrix containing node feature information. The graph structure encoding matrix combines the adjacency matrix... and characteristic matrix The result, expressed by the formula, is:
[0050] ;
[0051] in, This represents the initial graph structure encoding matrix. It is an adjacency matrix, and It is the feature matrix. Encoding the graph structure matrix. The input is fed into a self-attention gated graph neural network, where node features are updated through multiple layers of graph convolution operations. Graph convolution aggregates the features of neighboring nodes and updates the node's own features. The application of the self-attention mechanism in graph convolution assigns different neighbor weights to each node, focusing on the neighboring nodes most important for updating the node's features. Through multiple layers of graph convolution operations, information from distant nodes is gradually aggregated onto each node, resulting in an updated node feature matrix, denoted as . ,in This indicates the number of layers in the graph convolution. Based on the updated node feature matrix. The attention weights between nodes are calculated to quantify the importance of each node to other nodes during feature update. An attention score is calculated for each pair of nodes using a self-attention mechanism, and the attention scores are normalized using a softmax function to obtain the attention weight matrix of each node. The soft attention mechanism ensures that the sum of the attention scores of all neighboring nodes is 1 through normalization, generating a probability distribution. The node attention weight matrix... With the updated node feature matrix Perform matrix multiplication to obtain the weighted node feature matrix:
[0052] ;
[0053] in, This represents the weighted node feature matrix, which aggregates the features of neighboring nodes using attention weights to reflect the importance and dependencies between nodes. Performing column vector summation aggregates all node features into a single graph-level representation vector. Column vector summation consolidates node-level information to the overall graph level, resulting in a feature vector describing the entire graph, denoted as . To eliminate the bias caused by inconsistent eigenvalue magnitudes, the following measures are taken: After performing L2 norm normalization, the standardized graph-level representation vector is obtained:
[0054] ;
[0055] in, Representing vectors L2 norm, This is the normalized graph-level representation vector. The graph-level representation vector... By concatenating the basic attribute vectors of members, a comprehensive feature vector containing graph structure information and basic member information is obtained. This comprehensive feature vector is then processed through a fully connected layer for feature fusion and dimensionality reduction. The fully connected layer can perform linear transformations on the input high-dimensional features and introduce non-linearity through an activation function to extract more meaningful features, mapping the high-dimensional feature vector to a lower-dimensional space, ultimately yielding a multi-dimensional member feature vector.
[0056] In one specific embodiment, the process of performing step 102 may specifically include the following steps:
[0057] The multi-dimensional member feature vector is expanded by copying and zero-padding operations to extend the multi-dimensional member feature vector to the same length as the order data sequence to be processed, thus obtaining the expanded member feature vector.
[0058] The extended member feature vector is concatenated with the order data to be processed along the feature dimension to obtain the initial fused feature sequence, where each time step contains member features and corresponding order features;
[0059] The initial fused feature sequence is input into a bidirectional LSTM network, where the forward LSTM and the reverse LSTM each contain two layers, each containing 128 hidden units. The tanh activation function is used to obtain the forward LSTM feature sequence and the reverse LSTM feature sequence.
[0060] The outputs of the forward LSTM feature sequence and the backward LSTM feature sequence at the last hidden layer are concatenated to obtain the bidirectional LSTM feature sequence, which contains temporal dependency information captured from both directions.
[0061] The bidirectional LSTM feature sequence is input into a multi-head self-attention layer with 8 attention heads, each with a dimension of 64. The query, key, and value matrices are generated through linear transformation. The scaled dot product attention is calculated and softmax normalization is performed to obtain the self-attention feature sequence.
[0062] The self-attention feature sequence is input into a two-layer feedforward neural network. The first layer uses the ReLU activation function, the second layer uses the linear activation function, and the hidden layer dimension is 512. A non-linear transformation is performed to obtain the encoder output feature sequence.
[0063] The encoder output feature sequence is input into the adversarial discriminator, which contains three one-dimensional convolutional layers with kernel sizes of 3, 4, and 5, and each layer has 100 output channels. Finally, a fully connected layer is connected to output the binary classification result. The encoder parameters are optimized through adversarial training to obtain the adversarially optimized feature sequence.
[0064] Max pooling is performed on the time dimension of the optimized feature sequence to select the maximum value in each feature dimension, thus obtaining the member-order fusion feature vector.
[0065] Specifically, the multi-dimensional member feature vector undergoes dimensionality expansion processing. Through copying and zero-padding operations, the multi-dimensional member feature vector is expanded to the same length as the order data sequence to be processed, ensuring it shares the same time dimension with the order data and aligning member and order features. The expanded member feature vector is then concatenated with the order data to be processed along the feature dimensions to obtain the initial fused feature sequence. Assume the expanded member feature vector is... Order data is ,in The length of the order sequence. For the dimensions of member characteristics, This represents the dimension of the order features. Concatenating the two along this feature dimension yields the initial fused feature sequence. Each time step contains corresponding member features and order features. The initial fused feature sequence... The input is fed into a bidirectional LSTM network. The bidirectional LSTM network consists of a forward LSTM and a backward LSTM, with two layers in each direction, each containing 128 hidden units. LSTM is a recurrent neural network used to process temporal data, capable of capturing temporal dependencies in the data. The forward LSTM starts from time steps... arrive The input sequence is processed, while the inverse LSTM processes the input sequence from time step. arrive The input sequence is processed in reverse to capture temporal features from both directions. Each LSTM layer uses the tanh activation function to introduce non-linear features. After bidirectional LSTM processing, the forward LSTM feature sequence is obtained. and inverse LSTM feature sequences The forward and backward LSTM feature sequences are concatenated at the output of the last hidden layer to obtain the bidirectional LSTM feature sequence. Bidirectional LSTM feature sequences contain temporal dependency information captured from two directions, reflecting the relationship between order and member features at each time step. The bidirectional LSTM feature sequences... The input is fed into a multi-head self-attention layer. This layer computes dependencies between different time steps, using eight attention heads, each with a dimension of 64. A query matrix is generated through a linear transformation. Key matrix Sum matrix ,in:
[0066] ;
[0067] in, The linear transformation matrix is used to generate the query, key, and value matrices, respectively. By calculating the scaled dot product attention and performing softmax normalization, the attention weight matrix is obtained, which is then multiplied by the value matrix to obtain the self-attention feature sequence.
[0068] ;
[0069] in, The dimension of the key matrix is used, and the softmax function is used to normalize the attention scores, ensuring an effective measurement of the relationships between time steps. Through a multi-head attention mechanism, information patterns in the sequence are captured simultaneously from different perspectives, resulting in a self-attention feature sequence. Self-attention feature sequences The input is fed into a two-layer feedforward neural network. The first layer uses the ReLU activation function to introduce non-linearity and enhance the expressive power of the features; the hidden layer has a dimension of 512. The second layer uses a linear activation function to maintain the continuity of the output. Through the non-linear transformation of the feedforward neural network, deep features in the sequence are extracted, resulting in the encoder output feature sequence. To enhance the discriminative power of the feature sequence, the encoder outputs the feature sequence. The input is fed into the adversarial discriminator. The adversarial discriminator consists of three one-dimensional convolutional layers with kernel sizes of 3, 4, and 5, and each layer has 100 output channels. Convolutional operations extract local patterns from the feature sequence, and finally, a fully connected layer outputs a binary classification result to determine the authenticity or quality of the feature sequence. Through adversarial training, the encoder and discriminator compete against each other, forcing the encoder to generate more discriminative features, thereby optimizing its parameters and obtaining the adversarially optimized feature sequence. Max pooling is then performed on the adversarially optimized feature sequence in the time dimension, selecting the maximum value in each feature dimension to retain the most salient information in the feature sequence. This effectively reduces the dimensionality of the features while extracting the most representative features, ultimately yielding the member-order fusion feature vector.
[0070] In one specific embodiment, the process of performing step 103 may specifically include the following steps:
[0071] The member-order fusion feature vector is input into a multilayer perceptron network. The multilayer perceptron network contains three hidden layers: the first layer has 512 neurons, the second layer has 256 neurons, and the third layer has 128 neurons. The ReLU activation function is used between each layer, and the last layer uses a linear activation function to obtain the order feature representation.
[0072] The order feature representation is subjected to L2 regularization, the L2 norm of the feature vector is calculated, and the feature vector is divided by the L2 norm to obtain the normalized order feature vector.
[0073] The normalized order feature vector is input into the softmax classifier. The softmax function is used to calculate the probability of each category to obtain the order category probability distribution. Based on the order category probability distribution, the argmax function is used to find the index with the highest probability, and the category corresponding to the index is taken as the order classification result.
[0074] The entropy of the order category probability distribution is calculated to obtain a scalar entropy value representing the classification uncertainty. The scalar entropy value representing the classification uncertainty is compared with a predefined threshold. If the scalar entropy value representing the classification uncertainty is less than the threshold, the order is classified into a definite category; otherwise, it is classified into an uncertain category.
[0075] For orders with a defined category, a pre-defined priority mapping table is queried based on the order category result. This priority mapping table maps each order category to an integer priority of 1-5 to obtain an initial processing priority. For orders with an uncertain category, the normalized order feature vector is input into a pre-trained support vector regression model. The pre-trained support vector regression model uses the RBF kernel function to obtain a continuous priority score, and the continuous score is converted into a discrete processing priority of 1-5 by setting interval boundaries.
[0076] Specifically, the member-order fusion feature vector is input into a multilayer perceptron network. This multilayer perceptron network contains three hidden layers: the first layer contains 512 neurons, the second layer contains 256 neurons, and the third layer contains 128 neurons. The ReLU activation function is used between each layer. By introducing the non-linear activation function ReLU, the model can capture complex feature relationships and increase its representational power. After each layer of neurons and ReLU activation, a linear activation function is used in the last layer to maintain the continuity of the output, resulting in the feature representation of the order. The input feature vector is , No. The output of the layer is represented as:
[0077] ;
[0078] in, It is the first The weight matrix of the layer, It is the first Layer bias vector, This indicates the output of the previous layer. This is the output after ReLU activation. After passing through all hidden layers, the order feature representation is obtained. To ensure scale consistency and numerical stability of order feature representations, L2 regularization is applied. The L2 norm of the feature vector is calculated, and then the feature vector is divided by the L2 norm to obtain the normalized order feature vector. The formula for calculating the L2 norm is:
[0079] ;
[0080] in, The L2 norm of the eigenvectors. For the first in the order feature representation One portion, The dimension of the feature vector. The normalized order feature vector. Represented as:
[0081] ;
[0082] Normalization ensures that the feature vector has a length of 1, eliminating the impact of inconsistent feature values on model training and prediction. The normalized order feature vectors are then input into a softmax classifier to classify orders. The softmax classifier uses the softmax function to calculate the probability of each class; the expression for the softmax function is:
[0083] ;
[0084] in, This indicates that the order is categorized. The probability, and They are categories The weight vector and bias, This represents the total number of categories. By calculating the probability of each category, the probability distribution of order categories is obtained. The `argmax` function is used to find the index with the highest probability, and the category corresponding to that index is taken as the order's classification result. In this way, each order is classified, and its category is determined. To evaluate the certainty of the classification results, the entropy of the order category probability distribution is calculated to measure the uncertainty of the classification results. The formula for calculating entropy is:
[0085] ;
[0086] in, The entropy value of the probability distribution of order categories. It is a category The probability. Entropy value. The larger the entropy value, the higher the uncertainty of the classification; the smaller the entropy value, the more certain the classification result. (The entropy value is then considered.) The entropy value is compared with a predefined threshold. If the entropy value is less than the threshold, the classification result is considered highly certain, and the order is classified into a definite category. If the entropy value is greater than the threshold, the classification result is considered uncertain, and the order is classified into an uncertain category. For orders with a definite category, a pre-defined priority mapping table is queried based on the order classification result. This priority mapping table maps each order category to an integer priority of 1-5, obtaining the initial processing priority of the order. For example, if an order is classified as a high-value category, it is mapped to a priority of 1, indicating that it needs to be processed first. For orders with uncertain classification, a more refined priority determination method is used. The normalized order feature vector is input into a pre-trained support vector regression model to obtain a continuous priority score for the order. The support vector regression model uses the RBF kernel function to capture the complex nonlinear relationship between order features and priorities. The output of the support vector regression is a continuous score representing the order priority. To transform the continuous priority score into a discrete priority, a method of setting interval boundaries is used, such as dividing the score into an integer priority interval of 1-5, making the priority of an order comparable to other orders, and finally obtaining the processing priority of each order.
[0087] In one specific embodiment, the process of performing step 104 may specifically include the following steps:
[0088] Based on the order classification results and processing priority, a weight coefficient is assigned to each order. The weight coefficient is calculated as follows: weight = α * classification score + β * priority, where α and β are preset balancing parameters, resulting in a list of order weights.
[0089] Based on the processing capabilities and areas of expertise of each platform, a platform order compatibility matrix is constructed. The matrix element aij represents the compatibility of platform i with order j, and the compatibility range is 0-1.
[0090] Based on the order weight list and the platform order compatibility matrix, an objective function is constructed, and three constraints are set: each order can only be assigned to one platform, the order processing volume of each platform does not exceed the platform's maximum processing capacity, and the cross-platform allocation ratio of preset type orders does not exceed a preset threshold.
[0091] By combining the objective function and three constraints, a multi-platform collaborative integer programming model is constructed. The decision variable of the multi-platform collaborative integer programming model is the weighted score of each order on each platform, and the objective is to maximize the overall order processing efficiency.
[0092] Input the multi-platform collaborative integer programming model into the interactive optimization solver, set the solution parameters including the maximum number of iterations, convergence threshold and time limit, start the solution process, obtain the solution results from the interactive optimization solver, extract the optimal values of the decision variables, and combine the order-platform pairs with decision variables of 1 into a cross-platform order allocation scheme.
[0093] Specifically, each order is assigned a weight coefficient based on its classification and processing priority. The formula for calculating the weight coefficient is: Weight Category score Priority, where and Preset balancing parameters are used to adjust the relative importance of classification scores and priorities in the final weighting. The weight of each order is calculated using a formula, forming an order weight list. A platform-order fit matrix is constructed based on the processing capabilities and areas of expertise of each platform. The elements of the platform-order fit matrix represent the platform... Processing orders The degree of compatibility, using The fit score ranges from 0 to 1, representing the degree of matching between the platform and the order. A higher fit score indicates that the platform is better at handling that type of order. Based on the order weight list and the platform-order fit score matrix, an objective function for order allocation is constructed, and relevant constraints are set. The core of the objective function is to maximize the overall order processing efficiency, defined as the sum of the weighted scores of all orders across various platforms. The weighted score is expressed by the following formula:
[0094] ;
[0095] in, Indicates the overall order processing efficiency. Let be the decision variable, representing the order. Whether to assign to the platform If allocated Otherwise, it is 0; Indicates order The weight, Indicates platform For orders Adaptability; and These represent the number of platforms and the number of orders, respectively. Objective function The aim is to maximize the weighted score, ensuring that important orders are allocated to the most suitable platforms and fully utilizing their processing capacity. While constructing the objective function, three constraints are set to ensure the rationality and operability of order allocation. The first constraint is that each order can only be allocated to one platform, expressed by the following constraint formula:
[0096] ;
[0097] This formula represents each order Orders can only be assigned to one platform, ensuring the independence and clarity of order allocation. The second constraint is that the order processing volume of each platform cannot exceed its maximum processing capacity, expressed by the following formula:
[0098] ;
[0099] in, Indicates platform The first constraint is the platform's maximum processing capacity, ensuring it won't overload due to excessive order volume and maintaining processing stability. The second constraint is that the cross-platform allocation ratio of preset order types cannot exceed a preset threshold, ensuring that the cross-platform allocation of certain specific order types conforms to business needs and strategies; for example, certain high-value orders can only be allocated to specific platforms. Combining the above objective function and three constraints, a multi-platform collaborative integer programming model is constructed. The model's decision variable is the weighted score of each order on each platform, with the goal of maximizing overall order processing efficiency. The integer programming form allows the system to find the optimal solution in discrete space, determining how orders are allocated to maximize efficiency while satisfying constraints. The multi-platform collaborative integer programming model is input into an interactive optimization solver for solving. Before solving, some key parameters are set, including the maximum number of iterations, convergence threshold, and time limit. The maximum number of iterations controls the computational load of the optimization, ensuring the solution process doesn't continue indefinitely; the convergence threshold determines whether the solution has reached the optimal solution or is close enough to it; and the time limit limits the solution time to avoid affecting the system's real-time performance if the solution time is too long. By setting these parameters, the solver can find a better order allocation scheme within a reasonable time. After the solver starts the solution process, the optimization algorithm continuously adjusts the allocation of orders across various platforms to maximize the value of the objective function and ensure that all constraints are met. The interactive optimization solver gradually approaches the optimal solution by iteratively updating the decision variables. When a preset stopping condition is reached, the final solution result is obtained from the solver, and the optimal values of the decision variables are extracted. For each order... and platform If decision variables This indicates an order. Assigned to the platform The order-platform pairs with a decision variable of 1 are extracted and combined to form the final cross-platform order allocation scheme.
[0100] In one specific embodiment, the process of performing step 105 may specifically include the following steps:
[0101] Based on the order processing results of the cross-platform order allocation scheme and pending order data, member value assessment indicators are extracted, including consumption frequency, consumption amount, order completion rate, platform activity and customer service demand, and an original indicator matrix is constructed.
[0102] The original index matrix is normalized, and the range transformation method is used to convert the indices of different dimensions into dimensionless values to obtain a standardized index matrix. Based on the standardized index matrix, the information entropy of each index is calculated, and the objective weight of each index is calculated using the entropy weight method to obtain the objective weight vector.
[0103] A hierarchical analysis structure model is constructed, dividing the member value assessment indicators into three levels: the target level, the criteria level, and the indicator level, and a judgment matrix is established.
[0104] Perform a consistency test on the judgment matrix, calculate the maximum eigenvalue λmax and the consistency index CI=(λmax-n) / (n-1), where n is the matrix order, and calculate the consistency ratio CR=CI / RI, where RI is the random consistency index. If the consistency ratio CR<0.1, the judgment matrix passes the consistency test; otherwise, adjust the values in the judgment matrix until it passes the consistency test.
[0105] Based on the judgment matrix that has passed the consistency test, the subjective weight of each indicator is calculated using the eigenvalue method to obtain the subjective weight vector. The objective weight vector and the subjective weight vector are then weighted and averaged to obtain the comprehensive weight vector.
[0106] Based on the comprehensive weight vector and standardized index matrix, the value score of each member is calculated, and the member value scores are clustered to obtain hierarchical classification information.
[0107] Specifically, data is collected across multiple key dimensions related to each member, including purchase frequency, purchase amount, order completion rate, platform activity, and customer service demand. Each indicator represents a member's performance on the platform in a specific aspect. Purchase frequency measures the number of times a member makes purchases within a given timeframe; purchase amount is the member's total purchase amount; order completion rate indicates the percentage of orders completed by the member; platform activity reflects the member's interaction with the platform, such as the frequency of browsing and leaving messages; and customer service demand measures the frequency of a member's need for customer service assistance. This information is extracted from cross-platform order allocation schemes and order processing results to form an initial indicator matrix, where each row represents a member and each column represents an indicator. The initial indicator matrix is then normalized. Using the range transformation method, indicators with different dimensions are converted into dimensionless values, resulting in a standardized indicator matrix. The normalization formula for the range transformation method is:
[0108] ;
[0109] in, These are the original index values. and These are the minimum and maximum values of the indicator, respectively. These are the standardized indicator values. Through normalization, all indicators are transformed to the range of [0,1], thereby eliminating dimensional differences between different indicators and allowing the values of each indicator to be weighted and compared under the same standard. Based on the standardized indicator matrix, the information entropy of each indicator is calculated to assess its importance and information content. Information Entropy The calculation formula is:
[0110] ;
[0111] in, For the first Information entropy of each indicator It is a positive constant, usually taken as 1. This is used to ensure that the entropy value is within the range of [0,1]. For the first The first of the indicators The proportion of a member's standardized value to the total of this indicator. This refers to the number of members. By calculating the information entropy of each metric, the amount of information contained in the metric within the entire dataset is measured. The lower the entropy value, the stronger the discriminative power and the greater the importance of the metric. The objective weight of each metric is calculated using the entropy weight method. The formula for calculating the objective weight is:
[0112] ;
[0113] in, For the first The objective weight of each indicator, For the first Information entropy of each indicator This represents the total number of indicators. The objective weight of each indicator is calculated using this formula, reflecting the relative importance of each indicator to member value. A hierarchical analysis model is constructed, dividing the member value assessment indicators into three levels: the target level, the criteria level, and the indicator level. The target level assesses member value; the criteria level includes the main aspects affecting member value, such as consumption behavior, activity level, and service needs; and the indicator level includes specific quantitative indicators, such as consumption frequency and consumption amount. Then, a judgment matrix is established to compare the relative importance of each indicator in pairs. The elements of the judgment matrix are... Indicates the first The first indicator is relative to the first The importance of each indicator is scored by experts based on experience. A consistency check is performed on the judgment matrix to ensure its logical consistency. The largest eigenvalue of the judgment matrix is calculated. Then calculate the consistency index.
[0114] ;
[0115] in, To determine the largest eigenvalue of a matrix, To determine the order of the matrix, calculate the consistency ratio.
[0116] ;
[0117] in, This is a random consistency index, dependent on the order of the judgment matrix. When If the judgment matrix passes the consistency test, it indicates good logical consistency; otherwise, the values in the judgment matrix need to be adjusted until the consistency test is passed. After passing the consistency test, the subjective weights of each indicator are calculated using the eigenvalue method. The eigenvalue method calculates the eigenvector corresponding to the largest eigenvalue of the judgment matrix and performs normalization to obtain the subjective weight vector of each indicator. To consider both objective and subjective weights, the objective weight vector and the subjective weight vector are weighted and averaged to obtain the comprehensive weight vector.
[0118] ;
[0119] in, This is a weighting factor, ranging from [0,1], used to balance the influence of subjective and objective weights. Based on the comprehensive weight vector and standardized indicator matrix, the comprehensive value score for each member is calculated. (Member's comprehensive value score) The calculation formula is:
[0120] ;
[0121] in, For the first The overall value score of each member. For the first The combined weight of each indicator For the first The member in the The standardized values of each indicator are calculated. A weighted sum is then used to obtain the final value score for each member, which is used to assess the member's overall value and contribution. Cluster analysis is performed on the members' comprehensive value scores to categorize them into different tiers. For example, the K-means clustering algorithm can be used to divide members into high-value, medium-value, and low-value categories. This clustering tier information helps the platform develop targeted marketing and service strategies, thereby better managing member resources and improving member satisfaction and loyalty.
[0122] In one specific embodiment, the process of performing step 106 may specifically include the following steps:
[0123] The dynamic member value score and hierarchical classification information are combined into a state vector S, which contains the member value score, member level, and historical benefit usage rate. The state vector S is input into a deep Q network, which contains 3 hidden layers, each with 256, 128, and 64 neurons respectively. The ReLU activation function is used to obtain the Q-value matrix.
[0124] Based on the Q-value matrix, an ε-greedy strategy is used to select the rights and benefits allocation action, where the ε value is dynamically adjusted according to the membership level to obtain the initial rights and benefits allocation scheme;
[0125] A feasibility check is performed on the initial equity allocation plan to ensure that it meets the preset equity allocation rules and resource constraints. An effective equity allocation plan is obtained by iteratively adjusting the equity items in the plan.
[0126] Apply effective benefits allocation schemes to members, collect member feedback on the use of benefits within a fixed time window, and calculate instant reward values based on a preset reward function;
[0127] Based on the instant reward value and the preset discount factor, the cumulative reward is calculated and the experience tuple is stored in the experience replay pool with a capacity of 10,000.
[0128] Randomly sample small batches of experience data of size 64 from the experience replay pool, update the parameters of the target Q network using the double Q learning algorithm, and then copy the updated target Q network parameters to the evaluation Q network in a soft update manner until the preset convergence condition is reached, and output a personalized membership benefit combination.
[0129] Specifically, the dynamic value score of members, hierarchical classification information, and historical usage rate of benefits are combined to form a state vector. State vector As a description of a member, it includes three parts: the member's value score. This is used to reflect a member's overall contribution to the platform; member hierarchy information. This is used to indicate a member's level on the platform and the usage rate of historical benefits. This describes the frequency with which a member has historically used the rights allocated to them; this information helps determine a member's interest and need for different types of rights. The state vector... As input, it is fed into a deep Q-network. The deep Q-network contains three hidden layers, each with 256, 128, and 64 neurons respectively. The ReLU activation function is used between the hidden layers. The ReLU activation function is a non-linear function defined as... This allows for the introduction of nonlinearity, thereby enhancing the model's learning ability. After passing through each layer of the neural network, the deep Q-network learns complex patterns and features in the state vector, outputting a Q-value matrix. Each element in the Q-value matrix... This represents the expected return that can be obtained by taking a certain equity allocation action in a given state, where Indicates state, This represents the possible actions for allocating benefits. Based on the Q-value matrix, an ϵ-greedy strategy is used to select the action for benefit allocation. The ϵ-greedy strategy is a strategy that strikes a balance between exploration and utilization, where ϵ represents the probability of exploration. In each decision step, there is an ϵ probability of choosing a random action to explore more possible benefit allocation methods, and a 1−ϵ probability of choosing the action with the highest current Q-value to utilize the learned best strategy. To better adapt to the needs of different membership levels, the ϵ value is dynamically adjusted according to the membership level. For example, for higher-level members, the ϵ value is smaller to select more high-return actions, thus ensuring the refinement and personalization of benefit allocation; while for lower-level members, the ϵ value is larger to explore more different combinations of benefit allocation, thus uncovering possible optimal strategies. In this way, an initial benefit allocation scheme is obtained. The feasibility of the initial benefit allocation scheme is checked to ensure that the allocated benefits meet the platform's preset benefit allocation rules and resource constraints. The preset rights and benefits allocation rules include limits on the number of rights members can obtain and total platform resource limits. The feasibility of the initial plan is verified, and the allocation plan is improved by iteratively adjusting the rights and benefits as needed. Assuming that resources for a certain rights and benefits are limited and cannot meet the needs of all members, the initial plan is adjusted, for example, by reducing the quantity of certain rights and benefits or replacing them with other alternatives, resulting in an effective rights and benefits allocation plan that satisfies all constraints. The effective rights and benefits allocation plan is applied to members, and member feedback on the use of allocated rights and benefits is collected within a fixed time window. By tracking members' actual use of rights and benefits during this period, member satisfaction and usage frequency are obtained. Based on the usage feedback, an instant reward value is calculated using a preset reward function. The reward function is designed based on factors such as the frequency of member usage of benefits, satisfaction, and further consumption behavior on the platform. It typically takes the following form:
[0130] ;
[0131] in, For the first The instant reward value at each time step. and The reward coefficients represent the importance of the use of rights and consumption behavior, respectively. For the utilization rate of rights, This represents the amount of money a member spends on the platform after the allocation of benefits. Based on instant reward values. and preset discount factor Calculate cumulative rewards
[0132] ;
[0133] in, The discount factor, typically ranging from [0,1], is used to balance the relationship between immediate and future rewards, ensuring the model considers both the current equity allocation effect and the long-term benefits. The calculated cumulative reward measures the overall effect of an equity allocation scheme over the entire time series. Experience tuples (i.e., a quadruple of state, action, immediate reward, and next state) are stored in an experience replay pool of 10,000. The experience replay pool stores past experience samples and randomly samples these samples during training to update model parameters, effectively reducing the correlation between samples and improving the model's generalization ability and training stability. Small batches of experience data of size 64 are randomly sampled from the experience replay pool each time, and these data are used to update the parameters of the deep Q-network. During parameter updates, a double-Q learning algorithm is used to avoid overestimating the Q-value. Double-Q learning introduces two independent Q-networks: an evaluation Q-network and a target Q-network. The evaluation Q-network is used to select actions, while the target Q-network is used to evaluate the value of actions. During training, the parameters of the target Q-network are updated using the following formula:
[0134] ;
[0135] in, Indicates the target Q-network in state Select action Q value, For instant rewards, For the next state, This represents the Q-value for selecting the optimal action in the next state. The model iteratively approximates the optimal Q-value function, improving its decision-making ability. After each update, the updated target Q-network parameters are copied to the evaluation Q-network using a soft update method, implemented through the following formula:
[0136] ;
[0137] in, and The parameters are used to evaluate the Q-network and the target Q-network, respectively. This is the soft update coefficient, typically chosen as a small value (e.g., 0.01) to allow the target network to gradually approach the parameters of the evaluation network. This ensures the model gradually stabilizes during training and eventually reaches the preset convergence condition. Once the deep Q-network converges, it outputs a personalized membership benefit package.
[0138] The above describes the member order processing method in the embodiments of the present invention. The following describes the member order processing apparatus in the embodiments of the present invention. Please refer to [link / reference]. Figure 2 One embodiment of the member order processing device in this invention includes:
[0139] The acquisition module 201 is used to collect member information from multiple platforms and construct a member graph structure to extract node features and obtain multi-dimensional member feature vectors.
[0140] The fusion module 202 is used to perform feature fusion on the multi-dimensional member feature vector and the order data to be processed to obtain the member-order fusion feature vector;
[0141] Classification module 203 is used to classify orders based on member-order fusion feature vectors, and obtain order classification results and processing priorities;
[0142] Planning module 204 is used to perform multi-platform collaborative integer programming and solution based on order classification results and processing priorities to obtain cross-platform order allocation scheme;
[0143] Evaluation module 205 is used to evaluate member value based on cross-platform order allocation scheme and order processing results, and obtain dynamic member value score and hierarchical classification information;
[0144] The strategy optimization module 206 is used to input dynamic member value scores and hierarchical division information into a dynamic rights allocation model based on a deep Q network to optimize the rights allocation strategy and obtain a personalized combination of member rights.
[0145] Through the collaborative efforts of the aforementioned components, and by employing graph neural networks and soft attention mechanisms for member feature extraction, the complex relationships between member attributes can be effectively captured, improving the accuracy and comprehensiveness of member feature representation. The use of adversarial bidirectional encoders and bidirectional LSTM networks for member-order feature fusion achieves deep integration of member information and order data. The order classification method based on multilayer perceptron networks and support vector regression models accurately identifies order types and rationally allocates processing priorities, improving the efficiency and accuracy of order processing. The use of a multi-platform collaborative integer programming model for order allocation fully considers the processing capabilities and expertise of each platform, achieving optimal allocation of order resources and improving overall order processing efficiency. The combination of entropy weighting and analytic hierarchy process (AHP) for member value assessment considers both the distribution characteristics of objective data and incorporates expert experience, making the member value assessment results more comprehensive and reliable. The use of a dynamic rights allocation model based on deep Q-networks enables personalized allocation and real-time optimization of member rights, allowing for timely adjustments to rights combinations based on members' dynamic value and preferences, improving member satisfaction and loyalty. This invention improves the efficiency and accuracy of member management and order processing, providing enterprises with a comprehensive member order processing solution.
[0146] above Figure 2 The member order processing device in this embodiment of the invention will be described in detail from the perspective of modular functional entities. The member order processing device in this embodiment of the invention will be described in detail from the perspective of hardware processing.
[0147] Figure 3 This is a schematic diagram of a member order processing device 300 provided in an embodiment of the present invention. The member order processing device 300 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 310 (e.g., one or more processors) and a memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) for storing application programs 333 or data 332. The memory 320 and storage media 330 can be temporary or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the member order processing device 300. Furthermore, the processor 310 may be configured to communicate with the storage media 330 and execute the series of instruction operations in the storage media 330 on the member order processing device 300 to implement the steps of the aforementioned member order processing method.
[0148] The member order processing device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 3 The illustrated member order processing device structure does not constitute a limitation on the member order processing device provided by the present invention. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0149] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the member order processing method.
[0150] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0151] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0152] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing member orders, characterized in that, The method includes: Collect member information from multiple platforms and construct a member graph structure to extract node features, thereby obtaining multi-dimensional member feature vectors; The multi-dimensional member feature vector and the order data to be processed are fused to obtain the member-order fused feature vector; Orders are classified based on the member-order fusion feature vector to obtain the order classification results and processing priorities. Based on the order classification results and processing priorities, a multi-platform collaborative integer programming solution is performed to obtain a cross-platform order allocation scheme. Specifically, this includes: assigning a weight coefficient to each order based on the order classification results and processing priorities. The weight coefficient is calculated using the formula: Weight = α * Classification Score + β * Priority, where α and β are preset balancing parameters, resulting in an order weight list; and constructing a platform order compatibility matrix based on the processing capabilities and expertise of each platform, where matrix element a... ij Let represent the suitability of platform i for processing order j, with a suitability range of 0-1. Based on the order weight list and the platform order suitability matrix, an objective function is constructed, and three constraints are set: each order can only be assigned to one platform, the order processing volume of each platform does not exceed the platform's maximum processing capacity, and the cross-platform allocation ratio of preset type orders does not exceed a preset threshold. The objective function and the three constraints are combined to construct a multi-platform collaborative integer programming model. The decision variable of the multi-platform collaborative integer programming model is the weighted score of each order on each platform, and the objective is to maximize the overall order processing efficiency. The multi-platform collaborative integer programming model is input into an interactive optimization solver, and the solution parameters include the maximum number of iterations, the convergence threshold, and the time limit. The solution process is started, and the solution results are obtained from the interactive optimization solver. The optimal values of the decision variables are extracted, and the order-platform pairs with a decision variable of 1 are combined into a cross-platform order allocation scheme. Based on the cross-platform order allocation scheme and order processing results, member value is evaluated to obtain dynamic member value scores and hierarchical classification information. The dynamic member value score and the hierarchical classification information are input into a dynamic rights allocation model based on a deep Q-network for rights allocation strategy optimization to obtain a personalized member rights combination. Specifically, this includes: combining the dynamic member value score and the hierarchical classification information into a state vector S, where state vector S contains member value score, member level, and historical rights usage rate; inputting state vector S into a deep Q-network containing three hidden layers with 256, 128, and 64 neurons respectively; using the ReLU activation function to obtain a Q-value matrix; based on the Q-value matrix, using an ε-greedy strategy to select rights allocation actions, where the ε value is dynamically adjusted according to the member level to obtain an initial rights allocation scheme; and further refining the initial rights allocation scheme. A feasibility check is performed to ensure that the preset rights allocation rules and resource constraints are met. An effective rights allocation scheme is obtained by iteratively adjusting the rights items in the scheme. This effective scheme is then applied to members, and member feedback on rights usage is collected within a fixed time window. An instant reward value is calculated based on a preset reward function. Based on the instant reward value and a preset discount factor, cumulative rewards are calculated, and experience tuples are stored in an experience replay pool with a capacity of 10,000. Small batches of experience data of size 64 are randomly sampled from the experience replay pool, and the parameters of the target Q-network are updated using a double-Q learning algorithm. The updated target Q-network parameters are then copied to the evaluation Q-network using a soft update method until a preset convergence condition is met, outputting a personalized member rights combination.
2. The member order processing method according to claim 1, characterized in that, The process involves collecting member information from multiple platforms and constructing a member graph structure to extract node features, resulting in a multi-dimensional member feature vector, including: Data cleaning and format unification processing are performed on member information from multiple platforms to obtain a standardized member dataset. Based on the standardized member dataset, member attributes are set as nodes, and the relationships between attributes are set as edges to construct a member graph structure. The member graph structure is represented by a sparse matrix. The member graph structure information is encoded by the adjacency matrix and the feature matrix to obtain the graph structure encoding matrix. The graph structure encoding matrix is then input into the self-attention gated graph neural network. The node features are updated by multi-layer graph convolution operations to obtain the updated node feature matrix. Based on the updated node feature matrix, the attention weights between nodes are calculated, and the attention scores are normalized using the softmax function to obtain the node attention weight matrix. The node attention weight matrix is then multiplied by the updated node feature matrix to obtain the weighted node feature matrix. The column vectors of the weighted node feature matrix are summed and then normalized using the L2 norm to obtain a graph-level representation vector. This graph-level representation vector is then connected to the member information attribute vector, and feature fusion and dimensionality reduction are performed through a fully connected layer to obtain a multi-dimensional member feature vector.
3. The member order processing method according to claim 2, characterized in that, The step of fusing the multi-dimensional member feature vector and the order data to be processed to obtain a member-order fused feature vector includes: The multi-dimensional member feature vector is expanded by copying and zero-padding operations to extend it to the same length as the order data sequence to be processed, thus obtaining the expanded member feature vector. The extended member feature vector is concatenated with the order data to be processed along the feature dimension to obtain an initial fused feature sequence, wherein each time step contains member features and corresponding order features; The initial fused feature sequence is input into a bidirectional LSTM network, where the forward LSTM and the reverse LSTM each contain two layers, each containing 128 hidden units. The tanh activation function is used to obtain the forward LSTM feature sequence and the reverse LSTM feature sequence. The outputs of the forward LSTM feature sequence and the reverse LSTM feature sequence at the last hidden layer are concatenated to obtain a bidirectional LSTM feature sequence, which contains temporal dependency information captured from both directions. The bidirectional LSTM feature sequence is input into a multi-head self-attention layer with 8 attention heads, each with a dimension of 64. The query, key, and value matrices are generated through linear transformation. The scaled dot product attention is calculated and softmax normalization is performed to obtain the self-attention feature sequence. The self-attention feature sequence is input into a two-layer feedforward neural network. The first layer uses the ReLU activation function, the second layer uses the linear activation function, and the hidden layer dimension is 512. A non-linear transformation is performed to obtain the encoder output feature sequence. The encoder output feature sequence is input into an adversarial discriminator, which contains three one-dimensional convolutional layers with kernel sizes of 3, 4, and 5, and each layer has 100 output channels. Finally, a fully connected layer is connected to output a binary classification result. The encoder parameters are optimized through adversarial training to obtain the adversarially optimized feature sequence. Max pooling is performed on the adversarial optimized feature sequence in the time dimension, and the maximum value in each feature dimension is selected to obtain the member-order fusion feature vector.
4. The member order processing method according to claim 3, characterized in that, The order classification based on the member-order fusion feature vector, to obtain the order classification result and processing priority, includes: The member-order fusion feature vector is input into a multilayer perceptron network, which contains three hidden layers: the first layer has 512 neurons, the second layer has 256 neurons, and the third layer has 128 neurons. The ReLU activation function is used between each layer, and the last layer uses a linear activation function to obtain the order feature representation. The order feature representation is subjected to L2 regularization, the L2 norm of the feature vector is calculated, and the feature vector is divided by the L2 norm to obtain the normalized order feature vector. The normalized order feature vector is input into the softmax classifier. The softmax function is used to calculate the probability of each category to obtain the order category probability distribution. Based on the order category probability distribution, the argmax function is used to find the index with the highest probability, and the category corresponding to the index is taken as the order classification result. Entropy calculation is performed on the probability distribution of the order categories to obtain a scalar entropy value representing the classification uncertainty. The scalar entropy value representing the classification uncertainty is compared with a predefined threshold. If the scalar entropy value representing the classification uncertainty is less than the threshold, the order is classified into a definite category; otherwise, it is classified into an uncertain category. For orders with a defined category, a preset priority mapping table is queried based on the order category result. This priority mapping table maps each order category to an integer priority of 1-5 to obtain a preliminary processing priority. For orders with an uncertain category, the normalized order feature vector is input into a pre-trained support vector regression model. The pre-trained support vector regression model uses the RBF kernel function to obtain a continuous priority score, and the continuous score is converted into a discrete processing priority of 1-5 by setting interval boundaries.
5. The member order processing method according to claim 1, characterized in that, The process of evaluating member value based on the cross-platform order allocation scheme and order processing results to obtain dynamic member value scores and hierarchical classification information includes: Based on the cross-platform order allocation scheme and the order processing results of the pending order data, member value assessment indicators are extracted, including consumption frequency, consumption amount, order completion rate, platform activity and customer service demand, and an original indicator matrix is constructed. The original index matrix is normalized, and the range transformation method is used to convert the indices of different dimensions into dimensionless values to obtain a standardized index matrix. Based on the standardized index matrix, the information entropy of each index is calculated, and the objective weight of each index is calculated using the entropy weight method to obtain the objective weight vector. A hierarchical analysis structure model is constructed, dividing the member value assessment indicators into three levels: the target level, the criteria level, and the indicator level, and a judgment matrix is established. Perform a consistency check on the judgment matrix and calculate the largest eigenvalue λ. max Consistency index CI=(λ) max -n) / (n-1), where n is the matrix order, calculate the consistency ratio CR=CI / RI, where RI is the random consistency index. If the consistency ratio CR<0.1, the judgment matrix passes the consistency test; otherwise, adjust the values in the judgment matrix until it passes the consistency test. Based on the judgment matrix that has passed the consistency test, the subjective weight of each indicator is calculated using the eigenvalue method to obtain the subjective weight vector. The objective weight vector and the subjective weight vector are then weighted and averaged to obtain the comprehensive weight vector. Based on the comprehensive weight vector and the standardized index matrix, the value score of each member is calculated, and the member value scores are clustered to obtain hierarchical classification information.
6. A member order processing device, characterized in that, For performing the member order processing method as described in any one of claims 1-5, the member order processing apparatus comprises: The data acquisition module is used to collect member information from multiple platforms and construct a member graph structure to extract node features and obtain multi-dimensional member feature vectors. The fusion module is used to perform feature fusion on the multi-dimensional member feature vector and the order data to be processed to obtain the member-order fusion feature vector; The classification module is used to classify orders based on the member-order fusion feature vector, and obtain the order classification result and processing priority; The planning module is used to perform multi-platform collaborative integer programming and solving based on the order classification results and the processing priorities to obtain a cross-platform order allocation scheme; The evaluation module is used to evaluate member value based on the cross-platform order allocation scheme and order processing results, and obtain dynamic member value scores and hierarchical classification information. The strategy optimization module is used to input the dynamic member value score and the hierarchical division information into a dynamic rights allocation model based on a deep Q network to optimize the rights allocation strategy and obtain a personalized member rights combination.
7. A member order processing device, characterized in that, The member order processing device includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the member order processing device to perform the member order processing method as described in any one of claims 1-5.
8. A computer-readable storage medium storing instructions thereon, characterized in that, When the instruction is executed by the processor, it implements the member order processing method as described in any one of claims 1-5.
Citation Information
Patent Citations
Double-layer vehicle path planning method based on Lagrange relaxation algorithm
CN118247111A
Order maintenance management platform and method based on machine learning, and electronic equipment
CN119067789A