Online housekeeping service order allocation method considering comprehensive benefits of platform
Through reinforcement learning and Transformer network optimization of housekeeping service order allocation, the local optimal problem of existing algorithms under large-scale data is solved, the balance between customer satisfaction and service personnel income is achieved, and the intelligence and rationality of order allocation is improved.
Patent Information
- Application Number
- CN202510789197.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing housekeeping service order allocation algorithm converges slowly under large-scale data and is prone to fall into local optimal solutions. It fails to comprehensively consider factors such as customer satisfaction, service personnel income and order interval time, resulting in poor real-time performance and low matching efficiency.
A reinforcement learning-based method is adopted, combined with Transformer network and K-means clustering, objective functions and constraints are constructed, and order allocation strategies are optimized by destroying and repairing order sequences, customer satisfaction, service staff ratings and order intervals are considered, and order intelligence and rationality of order allocation are improved.
It improves the intelligence and rationality of order allocation, shortens the passage distance and order interval time of service personnel, improves customer satisfaction and service personnel income, and enhances the comprehensive benefits of the platform.
Smart Images

Figure CN120297709A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to an online housekeeping service order allocation method considering the comprehensive benefits of the platform. Background Art
[0002] In recent years, with the continuous progress of technology and the booming development of the economy, the housekeeping service industry has gradually entered the golden age. For the mechanism of the online housekeeping service platform, fully considering factors such as customer satisfaction and the overall income of service personnel, and designing a more intelligent and reasonable order matching algorithm play an important role in improving the mechanism of the online housekeeping service platform, enhancing the comprehensive benefits of the online housekeeping service platform, and expanding the application market of the housekeeping service industry.
[0003] For the order allocation problem, a relatively common solution method is the Large Neighborhood Search (LNS) algorithm. However, with the increase in the problem scale and the number of optimization objective categories, certain limitations of the LNS algorithm have been exposed during the solution process, mainly reflected in aspects such as slow convergence speed, limited global search ability, and being prone to falling into local optimal solutions.
[0004] Most existing studies aim to minimize the time consumption during the process of service personnel arriving at the service destination (for example, in the paper "Vehicle Routing and Scheduling for Regular Mobile Healthcare Services" published in ICTAI, the optimization objectives of home healthcare services are divided into two parts according to importance. Firstly, it is necessary to minimize the number of medical vehicles used within the specified time window. Secondly, on the premise that the vehicle can return to the origin station normally after the service ends, minimize the total driving time, the amount of fuel consumed, and the number of service routes. And in the paper "A Simultaneous Facility Location and Vehicle Routing Problem Arising in Health Care Logistics in the Netherlands" published in the European Journal of Operational Research, the optimization objective is to minimize the logistics cost of transporting drugs in lockers and the travel cost of patients picking up drugs from lockers, screen from a group of potential locker locations, and generate routes to visit these lockers and routes to visit patients within the coverage area of the lockers) or maximize customer satisfaction (for example, in the paper "Online Doctor-Patient Dynamic Stable Matching Model Based on Regret Theory Under Incomplete Information" published in Socio-Economic Planning Sciences, an online doctor-patient matching mechanism is proposed. Doctors are scored according to patients' disease types and doctors' skill information. Only when the doctor and the patient are successfully matched bilaterally does the doctor have the opportunity to participate in the treatment of the patient. This method takes into account both patients' health and doctors' medical needs; in the paper "Optimal quality regulation on the online health platform" published in Electronic Markets, by excluding low-level doctors and adding patient satisfaction to the specific benefits of the platform, the quality of online doctor-patient services is better guaranteed; in the paper "Medical Service Matching Decision Method Considering Patients' Individualized Needs" published in Operations Research and Management Science, by constructing a stable matching scheme and a satisfactory matching scheme for online doctor-patient services, the matching degree between doctors and patients is calculated, and then the matching decision of online doctor-patient services is realized), while ignoring the comprehensive consideration of the two and the impact of factors such as the time interval between adjacent service orders on the matching decision.In addition, the real-time performance of the existing domestic service order allocation algorithm is poor, especially obvious when the number of orders is at the peak. There is often a phenomenon that customers cannot match with service personnel for a long time after placing an order. Moreover, with the increase of the problem scale and the increase of the optimization target categories, the LNS algorithm exposes certain limitations in the solution process, mainly reflected in aspects such as slow convergence speed, limited global search ability, and easy to fall into local optimal solutions. Summary of the Invention
[0005] In view of this, the present invention aims to provide an online domestic service order allocation method considering the comprehensive benefits of the platform. On the basis of the traditional order allocation problem, factors such as customer satisfaction, the travel distance of service personnel for each order, and the time interval between adjacent orders are comprehensively considered, with the goal of improving the comprehensive benefits of the online domestic service platform, providing effective decision-making support for improving the intelligence and rationality of order allocation in the online domestic service platform.
[0006] To achieve the above object, the technical solution of the present invention is realized as follows: An online domestic service order allocation method considering the comprehensive benefits of the platform, including: S1: Determine the constraint conditions and objective function of the prediction model for predicting the order allocation result according to the influence parameters affecting the comprehensive benefits of the domestic service platform; S2: Sort the orders to be analyzed according to the importance degree, and divide the sorted order sequence to obtain the initial optimal solution of the predicted order allocation result, and the initial optimal solution satisfies the constraint conditions; S3: Based on the constraint conditions and objective function obtained in step S1, perform local search based on reinforcement learning on the initial optimal solution obtained in step S2 to predict the optimal order allocation result that satisfies the constraint conditions and objective function.
[0007] Further, the objective function in step S1 is: ; where obj represents the optimization target value, α, β, and γ are weight parameters in the objective function, α + β + γ = 1, I represents the total number of service personnel, J represents the total number of orders, f i represents the service score of the i-th service personnel, 0.5 < f i < 1, R ij takes a value of 0 or 1, R ij = 0 indicates that the j-th order is not served by the i-th service personnel, R ij = 1 indicates that the j-th order is served by the i-th service personnel, |I serve | represents the number of service personnel actually participating in order service, D jj’represents the Euclidean distance between the j-th order and the j'-th order, X ijj’ takes values of 0 or 1, X ijj’ when = 0, it means that the next order of the j-th order of the i-th service staff is not the j'-th order, X ijj’ when = 1, it means that the next order of the j-th order of the i-th service staff is the j'-th order, represents the time when the i-th service staff executes the j'-th order, represents the time when the i-th service staff executes the j-th order, represents the service duration of the j-th order.
[0008] Furthermore, the constraint conditions in step S1 include the service staff number constraint for each order, the service time constraint for each order, the time connectivity and space connectivity of the order route in each order, and the position constraint that all service staff start from the origin point at the first order, and the i-th service staff starts from the origin position and goes to the position where the j-th order is located; among them: The service staff number constraint is: ; The service time constraint is the half-hour integer multiple moment between the earliest service start time and the latest service start time of each order, and all service staff must arrive at the position where the order is located before the latest service start time of each order, that is:
[0009] Among them, represents the latest start service time of the j-th order, represents the earliest start service time of the j-th order, a ij’ represents the time when the i-th service staff arrives at the position of the j'-th order, M represents a positive integer; The time connectivity satisfies: ; Among them, v represents the passing speed of the service staff; The space connectivity satisfies:
[0010] The position constraint satisfies:
[0011] Among them, p i represents the initial position of the i-th service staff, represents that the i-th service staff starts from the initial position p i whether it directly reaches the position of the j'-th order, takes values of 0 or 1, When the value is 0, it means that the $i$-th service staff starts from the initial position $p$ i and does not directly reach the $j'$-th order position, When the value is 1, it means that the $i$-th service staff starts from the initial position $p$ i and directly reaches the $j'$-th order position.
[0012] Furthermore, step S2 includes: S21: Construct an order node set from service orders and service staff, and construct a complete undirected graph from the order node set; Using the adjacency matrix method, construct a graph Laplacian matrix based on the complete undirected graph, and calculate the eigenvalues of the graph Laplacian matrix, where the eigenvalues characterize the importance of the corresponding orders; S22: Perform K-means clustering on the positions and order times of the orders to obtain several clusters; Sort the orders within the clusters and the orders between the clusters respectively according to the eigenvalues of the orders within the clusters and the number of connections between the orders between the clusters; S23: Use the K-means clustering algorithm to segment the order sequence obtained in step S22 to obtain multiple subsequences, and calculate the comprehensive score of each order in the current subsequence; Sort all the orders according to the comprehensive scores to obtain the initial optimal solution.
[0013] Furthermore, the process of calculating the comprehensive score of each order in the current subsequence in step S23 includes: Calculate the Euclidean distance between each order and the first order in the current subsequence, and perform weighted calculation with the service score of each order to obtain the comprehensive score corresponding to each order.
[0014] Furthermore, step S3 includes: S31: Input the initial optimal solution into multiple feature encoding modules based on the Transformer network to extract order features, obtaining the initial state and initial features, where the orders in the initial state satisfy the constraint conditions; S32: Input the initial state and initial features obtained in step S31 into the destruction module to obtain and execute the current destruction action of deleting some order nodes, obtaining the current destruction state, where the orders in the current destruction state satisfy the constraint conditions, and determining the current destruction reward corresponding to the current destruction action; S33: Input the current destruction state obtained in step S32 into the repair module to obtain and execute the current repair action, obtaining the current repair state and the corresponding order sequence, determining the current repair reward corresponding to the current repair action, where the orders in the current repair state satisfy the constraint conditions; S34: Input the order sequence obtained in step S33 into the feature encoding module for feature extraction. Replace the initial features and the initial state in step S32 with the extracted features and the current repair status, and repeat steps S32 - S34, and output the optimal order allocation result from the repair module.
[0015] Further, in each feature encoding module, input the initial optimal solution into the linear transformation layer for linear transformation; the output of the linear transformation layer is then successively subjected to two feature extractions and order rearrangement by the first attention sub-module and the second attention sub-module to obtain the initial state; In the first attention sub-module, use the multi-head attention mechanism to perform feature encoding processing on the input features, and then perform batch normalization processing on the encoded features to obtain the output; In the second attention sub-module, the input features are first processed by a feed-forward neural network, and then the processed features are subjected to batch normalization processing to obtain the output.
[0016] Further, in the disruption module: Decode the initial state or the repair state at the previous moment through the disruption policy network to obtain the current disruption action; After executing the current disruption action to obtain the current disruption state, use the disruption reward function to evaluate the current disruption state to obtain the current disruption reward.
[0017] Further, the disruption reward function is: ; where represents the disruption reward function, represents the current disruption state, represents the current disruption action, m and n respectively represent the m-th order node and the n-th order node, represents the correlation degree between the m-th order node and the n-th order node, which is obtained by the following formula: ; where ε and η are respectively weight values, d mn represents the Euclidean distance between the m-th order node and the n-th order node, that is: ; where and respectively represent the position coordinates of the m-th order node and the n-th order node.
[0018] Furthermore, the destruction policy network includes a multi-head self-attention module and an LSTM network; among them, the multi-head self-attention module performs feature interaction on the initial state and the repair state at the previous moment, and then uses the LSTM network to decode the execution sequence of the interaction features to obtain the removal probability score of the removed order node; according to the removal probability distribution, the removed order node is removed from the order node set at the previous moment to complete the current destruction action.
[0019] Furthermore, in the repair module: The current repair action is obtained by decoding the destruction state at the previous moment through the repair policy network; After the current repair action is executed to obtain the current repair state, the repair reward function takes the change in the objective function of the current repair state at the previous and current moments as the current repair reward.
[0020] Furthermore, in the repair policy network, the multi-head self-attention is used to decode the current destruction state and the removed order nodes to obtain the repair probability distribution of the order node insertion positions; the first several positions with the same number as the number of removed order nodes are selected as the insertion areas; according to the repair probability distribution, the removed order nodes are sequentially inserted into the order node set at the previous moment to complete the current repair action.
[0021] Furthermore, the repair reward function is: ; Among them, represents the repair reward function, represents the current repair state, represents the current repair action, m and n respectively represent the m-th order node and the n-th order node, represents the correlation degree between the m-th order node and the n-th order node, which is obtained by the following formula: ; Among them, ε and η are weight values respectively, and d mn represents the Euclidean distance between the m-th order node and the n-th order node, that is: ; Among them, and respectively represent the position coordinates of the m-th order node and the n-th order node.
[0022] Compared with the prior art, the present invention can achieve the following beneficial effects: The online housekeeping service order allocation method considering the comprehensive benefits of the platform in the present invention for creations solves well the problems such as slow search speed and easy getting stuck in local optimality of traditional service scheduling algorithms in large-scale data, and has achieved certain improvements in terms of algorithm intelligence and reasoning speed compared with the commonly used solution methods (LNS algorithm) for problems such as order allocation. At the same time, the objective function constructed in the present invention is more reasonable and meets the actual needs. The present invention can provide effective decision-making support for improving the intelligence and rationality of order allocation in the online housekeeping service platform, and can also have a positive impact on improving the comprehensive benefits of the platform and expanding the market scale of the online housekeeping service industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings constituting a part of the present invention for creations are used to provide a further understanding of the present invention for creations. The schematic embodiments and descriptions thereof of the present invention for creations are used to explain the present invention for creations and do not constitute an improper limitation on the present invention for creations. In the drawings: Figure 1 is a schematic flowchart of the online housekeeping service order allocation method considering the comprehensive benefits of the platform according to an embodiment of the present invention for creations; Figure 2 is a flowchart block diagram of the online housekeeping service order allocation method considering the comprehensive benefits of the platform according to an embodiment of the present invention for creations; Figure 3 is a schematic diagram of the reinforcement learning algorithm according to an embodiment of the present invention for creations; Figure 4 is a schematic diagram of the feature encoding module according to an embodiment of the present invention for creations. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In order to make the objectives, technical solutions and advantages of the present invention for creations clearer, the present invention for creations will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention for creations and do not constitute a limitation on the present invention for creations.
[0025] It should be noted that, without conflict, the embodiments in the present invention for creations and the features in the embodiments can be combined with each other.
[0026] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more.
[0027] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0028] The present invention will be described in detail below with reference to the drawings and in conjunction with embodiments.
[0029] As Figure 1 and Figure 2 shown, the online housekeeping order allocation method considering the comprehensive benefits of the platform described in the embodiments of the present invention includes: S1: Determine the constraint conditions and objective function of the prediction model for predicting the order allocation result according to the influencing parameters affecting the comprehensive benefits of the housekeeping service platform.
[0030] In some online housekeeping service scenarios, the order allocation problem can be described as follows: There are no more than I service personnel participating in J orders. It is known that the order location coordinates of the j-th order among the J orders, the earliest start time of service for the j-th order , the latest start time of service for the j-th order , the service duration of the j-th order , and the service score f of the i-th service personnel i, and the passing speed v of each service staff. No more than M service staff depart from their origin at the same time, return to the origin after completing all order services, and aim to maximize the service scores of all service staff matching the orders, while minimizing the average passing distance of service staff for each order and the average interval time between adjacent orders. A reasonable matching strategy between orders and service staff is designed. During this process, each service staff departs from the origin, serves only one order at each moment, and must arrive at the location of the order within the specified time window of the order. At the same time, one order can only be served by one service staff. In addition, to simplify the calculation, it is stipulated that the start time of order service is an integer multiple of half an hour within the time window. In the objective function, since the passing speed of the service staff is constant, the time interval between serving adjacent orders depends on the Euclidean distance between adjacent orders. In the embodiments of the present invention, the passing speed v of each service staff is the same and is a constant value, that is, v = 15 km / h.
[0031] In some embodiments, the objective function is: ; where obj represents the optimized objective value, α, β, and γ are weight parameters in the objective function, and α + β + γ = 1. Since the variables corresponding to the β term and γ term have a negative impact on the final objective function value, a negative sign is used to balance; 0.5 < f i <1, R ij takes a value of 0 or 1, R ij = 0 indicates that the j-th order is not served by the i-th service staff, R ij = 1 indicates that the j-th order is served by the i-th service staff, |I serve | represents the number of service staff actually participating in order service, D jj’ represents the Euclidean distance between the j-th order and the j'-th order, X ijj’ takes a value of 0 or 1, X ijj’ = 0 indicates that the next order of the j-th order of the i-th service staff is not the j'-th order, X ijj’ = 1 indicates that the next order of the j-th order of the i-th service staff is the j'-th order, represents the moment when the i-th service staff executes the j'-th order, represents the moment when the i-th service staff executes the j-th order, represents the service duration of the j-th order. In the embodiments of the present invention, the weight parameters α, β, and γ take values of 0.78, 0.025, and 0.195 respectively.
[0032] Analyzing each component in the objective function, it can be seen that to maximize the value of the objective function, it is required that the average score of the service personnel participating in the order service is as high as possible, the average Euclidean distance between the current location of the service personnel and the order location in each order service is as short as possible, and the average time interval between the service personnel serving adjacent orders is as short as possible.
[0033] In some embodiments, the constraint conditions in step S1 include the service personnel number constraint for each order, the service time constraint for each order, the time connectivity and space connectivity of the order route in each order, and the location constraint that all service personnel start from the origin at the first order, and the i-th service personnel starts from the origin location and travels to the location where the j-th order is located; where: The service personnel number constraint requires that each order can only be served by one and only one service personnel. Otherwise, it is not conducive to the design of the model, and it is specifically expressed as: ; The service time constraint is the half-hour integer multiple moment between the earliest service start time and the latest service start time of each order. All service personnel must arrive at the order location before the latest service start time of each order, that is:
[0034] where, a ij’ represents the moment when the i-th service personnel arrives at the location of the j'-th order, and M represents a positive integer; The time connectivity satisfies: ; The space connectivity satisfies:
[0035] The location constraint satisfies: ; where, p i represents the initial location of the i-th service personnel, represents whether the i-th service personnel directly reaches the location of the j'-th order from the initial location p i , takes a value of 0 or 1, when taking a value of 0, it means that the i-th service personnel does not directly reach the location of the j'-th order from the initial location p i , when taking a value of 1, it means that the i-th service personnel directly reaches the location of the j'-th order from the initial location p i .
[0036] S2: Sort the orders to be analyzed according to their importance levels, and partition the sorted orders to obtain an initial optimal solution for the predicted order allocation result, and the initial optimal solution satisfies the constraint conditions. In the present invention, the order is constructed in the form of graph data, and an order-service staff matching mechanism is formulated based on the properties of the graph Laplacian matrix to generate an initial solution. In some embodiments, step S2 includes: S21: Construct an order node set from service orders and service staff, and construct a complete undirected graph from the order node set; using the adjacency matrix method, construct a graph Laplacian matrix based on the complete undirected graph, and calculate the eigenvalues of the graph Laplacian matrix, where the eigenvalues characterize the importance levels of the corresponding orders.
[0037] In the embodiments of the present invention, a complete undirected graph G = (V, E) is defined, where V is the set of vertices in the undirected graph, and here V = I ∪ J represents the order node set in the undirected graph, and E represents the set of edges in the undirected graph. If a service staff can continuously serve two orders within adjacent time windows, then it is said that there is a connectivity relationship between these two orders. In the embodiments of the present invention, the connectivity relationship between all orders is represented in the form of an adjacency matrix, and the graph Laplacian matrix is constructed based on the complete undirected graph G using the adjacency matrix. By calculating the eigenvalues of the graph Laplacian matrix, the importance level of each order in the graph can be obtained. The larger the eigenvalue, the higher the importance level of the order in the graph. By analyzing the importance levels of each order and preferentially matching the orders with higher importance levels, the efficiency and timeliness of the service can be effectively guaranteed.
[0038] S22: Perform K-means clustering on the order locations and order times to obtain several clusters; perform in-cluster order sorting and inter-cluster order sorting respectively based on the eigenvalues of the orders within the clusters and the number of connections between the orders in different clusters.
[0039] The eigenvector of an order consists of the importance of the order and the two-dimensional coordinates of the order location. Since it is easy to cause problems such as slow algorithm search speed and getting stuck in local optima under large-scale data, the strategy of divide and conquer is adopted in the present invention. Specifically, in the embodiments of the present invention, based on the K-means algorithm, clustering is performed according to the abscissa of the order location, the ordinate of the order location, and the order time to obtain several clusters. After clustering all the orders, orders with similar characteristics within the same cluster can make full use of the surrounding service staff resources in the local order sequence, thereby improving the algorithm convergence speed and minimizing the travel distance of the service staff to serve the orders as much as possible. Then, in-cluster order sorting and inter-cluster order sorting are performed respectively based on the importance of the orders within the clusters and the number of connections between the orders in different clusters.
[0040] S23: Use the K-means clustering algorithm to segment the order sequence obtained in step S22 to obtain multiple subsequences, and calculate the comprehensive score of each order in the current subsequence; sort all orders according to the comprehensive score to obtain an initial optimal solution that meets the constraint conditions. In some embodiments, the process of calculating the comprehensive score of each order in the current subsequence includes calculating the Euclidean distance between each order and the first order in the current subsequence, and performing weighted calculation with the service score of each order, which can reflect the fitness of the service personnel to serve this order. The obtained comprehensive score corresponding to each order indicates that the higher the comprehensive score, the more suitable this order is for allocation to this service personnel. In addition, in the present invention, the passing speed of the service personnel and the time window of order service are comprehensively considered, and the order sequence is reasonably segmented into multiple subsequences, so as to accurately evaluate the performance of each service personnel in each subsequence and allocate the best service personnel to each subsequence. This strategy can ensure that each sub-path is served by the most suitable service personnel, thereby improving service quality and customer satisfaction.
[0041] S3: Based on the constraint conditions and objective function obtained in step S1, perform local search based on reinforcement learning on the initial optimal solution obtained in step S2, and predict an optimal order allocation result that meets the constraint conditions and objective function. In the present invention, a reinforcement learning algorithm as shown in Figure 3 is used to guide the agent to select the optimal action in a specific state, perform neighborhood search by destroying and repairing the order sequence, use the constraint conditions and objective function as the basis, and comprehensively consider the current state and possible future states to solve the problem, so as to reduce the risk of falling into a local optimal solution.
[0042] In some embodiments, step S3 includes: S31: Input the initial optimal solution into the feature encoding module based on the Transformer network to extract order features, and obtain the initial state and initial features.
[0043] In the embodiments of the present invention, using N feature encoding modules to embed the initial optimal solution X0 into a high-dimensional feature space to obtain the initial state s0 can be expressed as: ; where represents the encoded feature output by the v-th feature encoding module among the N feature encoding modules in the 0th iteration. When v = 0, that is, represents the initial feature after initial feature encoding of the initial optimal solution X0. Transformer represents the feature extraction operation of the feature encoding module. After the p-th iteration, the feature obtained after the operation of N feature encoding modules is 。The feature encoding aims to utilize N feature encoding modules based on the Transformer network to enrich the feature representation, thereby providing context information in the initial order sequence for subsequent disruption actions during the initial iteration of the algorithm.
[0044] In some embodiments, the feature encoding module uses the feature encoding module based on the Transformer network in the paper "Attention is all you need" published in "Advances in neural information processing systems". Its structure is as Figure 4 shown. Among them, the initial optimal solution is input into the linear transformation layer for linear transformation. The output of the linear transformation layer is then successively subjected to two feature extractions and order rearrangement by the first attention sub-module and the second attention sub-module to obtain the initial state, and the orders in the initial state satisfy the constraint conditions. Among them, in the first attention sub-module, the multi-head attention mechanism is used to perform feature encoding processing on the input features, and the encoded features are then subjected to batch normalization processing to obtain the output. In the second attention sub-module, the input features are first processed by a feed-forward neural network, and the processed features are then subjected to batch normalization processing to obtain the output. In the embodiments of the present invention, the feed-forward neural network adopted is the long short-term memory network in the recurrent feed-forward neural network.
[0045] S32: Input the initial state and the initial features obtained in step S31 into the disruption module, obtain and execute the current disruption action of deleting some order nodes to obtain the current disruption state, and the orders in the current disruption state satisfy the constraint conditions, and determine the current disruption reward corresponding to the current disruption action. An effective disruption mechanism can introduce a new search space and expand the range of service personnel that orders can match, thereby increasing the probability of high-scoring service personnel participating in order services.
[0046] In some embodiments, in the disruption module, the disruption policy network decodes the initial state or the repair state at the previous moment to obtain the current disruption action; after executing the current disruption action to obtain the current disruption state, the disruption reward function is used to evaluate the current disruption state to obtain the current disruption reward. The disruption policy network includes a multi-head self-attention module and a long short-term memory network (LSTM network). Among them, the multi-head self-attention module performs feature interaction on the initial state and the repair state at the previous moment, and then the LSTM network decodes the execution sequence of the interaction features to obtain the removal probability score of the order nodes; the order nodes in the order node set at the previous moment are removed according to the removal probability distribution to complete the current disruption action.
[0047] The disruption module in the present invention adopts a disruption mechanism, such as Figure 3 shown. The disruption mechanism includes 4 parts, namely the current disruption state , the current destruction action , the current destruction reward and the destruction policy network for formulating the current optimal destruction action .
[0048] For the current destruction state , the destruction module executes the current destruction action on the initial state s0 or the repair state at the previous moment , so as to obtain the currently damaged order sequence, that is, the current destruction state , where represents the set of all possible sequences of all order features after removing a certain number of orders from the initial order sequence or the order sequence at the previous moment.
[0049] For the current destruction action , through the destruction policy network decodes the repair state at the previous moment to obtain the current destruction action for the current order sequence . The content of the destruction action includes: continuously selecting a certain number of orders to be removed from the current order sequence, and the current destruction action can be expressed as: ; where P t represents the current order sequence, represents the set of features of the removed orders at the current moment, , A represents the set of actions of selecting a certain number of orders from the current order sequence P t .
[0050] For the current destruction reward , after obtaining the current destruction action , the current destruction state is evaluated through the destruction reward function to obtain the current destruction reward of the current destruction action . Since deleting orders with too high similarity of features can, to a certain extent, expand the scope of search solutions, thereby increasing the probability of matching high-scoring service personnel and achieving the purpose of improving customer satisfaction. Therefore, in order to guide the destruction policy network to achieve a balance between exploration and exploitation, the similarity of the orders deleted at the moment is used as the reward index of the destruction action. The destruction policy network is encouraged by the reward function to explore orders with high similarity more actively, so as to achieve better results in the deletion operation. In some embodiments, the destruction reward function is: ; Among them, m and n respectively represent two different orders, and x m and x n respectively represent the m-th order node and the n-th order node, represents the degree of association between the m-th order node and the n-th order node, and is obtained by the following formula: ; Among them, ε and η are respectively weight values, which are adaptively defined and assigned according to the actual situation, and d mn represents the Euclidean distance between the m-th order node and the n-th order node, that is: .
[0051] Among them, and respectively represent the position coordinates of the m-th order node and the n-th order node.
[0052] For the destruction policy network , by using the initial state s0 or the repair state at the previous moment to perceive the context information before the execution of the action, the current destruction action for destroying the current order sequence can be obtained. In the present invention, multi-head self-attention is used to perform feature interaction among order sequences, and then an LSTM network is used to decode the execution sequence of the interaction features, and then the probability distribution of the removed orders is output in sequence. In the embodiment of the present invention, specifically, in the destruction policy network, first perform feature extraction of multi-head self-attention on the feature encoding module, and at the same time perform decoding operations of the LSTM network on the input initial state or the repair state at the previous moment. A function that performs masking and softmax operations on the decoded features in sequence to obtain the removed order; remove the order at the specified position from the initial state or the repair state at the previous moment, complete the current destruction action, and obtain the current destruction state. The processing process of the destruction policy network can be represented by the following formula: ; ; ; ; Among them, represents the features extracted by multi-head self-attention in the destruction policy network in the p-th iteration, Attention represents multi-head self-attention, , and represent the weights in multi-head self-attention, and respectively represent the feature vector and output vector of the output of the s-th layer network in the LSTM network, Ω represents the removed orders, masked_softmax represents the function that performs masking and softmax operations on the feature vector in sequence, Repeat represents the repetition operation, and Delete represents the repair state from the previous moment The operation of removing the order elements at the specified positions from , k represents the number of times the Delete operation is executed, and the number of times k is adaptively adjusted according to the actual situation and updated to the optimal value through parameter tuning.
[0053] After the destruction policy network outputs the current destruction action and finishes executing the current destruction action, the current destruction state can be obtained, which in turn provides an object for the repair module to execute the repair mechanism.
[0054] S33: Input the current destruction state obtained in step S32 into the repair module, obtain and execute the current repair action, obtain the current repair state and the corresponding order sequence, determine the current repair reward corresponding to the current repair action, and the orders in the current repair state satisfy the constraint conditions.
[0055] Specifically, in the repair module: the repair policy network decodes the destruction state of the previous moment to obtain the current repair action; after executing the current repair action to obtain the current repair state, the repair reward function takes the change in the objective function of the current repair state at the previous and current moments as the current repair reward. In the repair policy network, multi-head self-attention is used to decode the current destruction state and the removed order nodes to obtain the repair probability distribution of the insertion positions of the order nodes; select the first several positions with the same number as the removed order nodes as the insertion area; insert the removed order nodes into the order node set of the previous moment in sequence according to the repair probability distribution to complete the current repair action.
[0056] The repair module in the present invention adopts a repair mechanism, as Figure 3 shown, the repair mechanism includes 4 parts, namely the current repair state , the current repair action , the current repair reward and the repair policy network for formulating the current optimal repair action .
[0057] For the current repair state , the repair module executes the current repair action on the current destruction state , thereby obtaining the currently repaired order sequence, that is, the current repair state , where represents the set of all possible order permutations when inserting a certain number of orders into the damaged large path.
[0058] For the current repair action , through the repair policy network decode the current damaged state to obtain the current repair action for the current order sequence . The content of the repair action includes: inserting the removed order into the damaged current order sequence. The current repair action can be expressed as: ; wherein, represents the damaged current order sequence , and A’ represents the set of actions of inserting the removed order into the damaged current order sequence
[0059] For the current repair reward , after obtaining the current repair state , evaluate the current repair state through the repair reward function to obtain the current repair reward of the current repair action , so as to guide the balance between exploration and exploitation of the repair policy network. In some embodiments, the repair reward function is: .
[0060] For the repair policy network , using the current damaged state , and the feature set of the removed order at the current moment to perceive the context information before the action execution, the current repair action for damaging the current order sequence can be obtained . In the embodiments of the present invention, the multi-head self-attention is used to decode the current damaged state to obtain the probability distribution of the order insertion position, and finally select the top multiple positions consistent with the number of removed order nodes as the insertion area, and insert the removed order nodes into the selected insertion area in turn. Specifically: using the multi-head self-attention to decode the current damaged state to obtain the probability distribution of the order insertion position, and finally select the top positions as the insertion area, and insert the order feature set into the selected insertion area in turn. The specific process can be expressed by the following formula: ; ; wherein, represents the current damaged state and the current order feature The mutual encoding feature vector between represents the set of order feature insertion positions, W 0 and b 0 Represents the weight value of the softmax function.
[0061] S34: Input the order sequence obtained in step S33 into the feature encoding module for feature extraction, replace the initial features and initial state in step S32 with the extracted features and the current repair state, repeat steps S32 to S34, and output the optimal order allocation result from the repair module. In the embodiment of the present invention, the number of iterations is 20 times, that is, after repeating steps S32 to S34 20 times, the optimal order allocation result can be output from the repair module.
[0062] Through the above steps, the final order allocation result can be obtained.
[0063] In the embodiment of the present invention, the good performance of the online housekeeping order allocation method considering the comprehensive benefits of the platform provided by the present invention is verified through the following experiments.
[0064] Obtain a verified data set. The embodiment of the present invention obtains a data set consisting of 2304 groups of order data and 2975 groups of service personnel data from the "58 Home" housekeeping service platform. Each group of order data includes 7 fields, namely, order number, order placement time, earliest time for the order service to start, latest time for the order service to start, order service duration, order service location horizontal coordinate x, and order service location vertical coordinate y. Each group of service personnel data includes the service personnel number, service personnel service score, and the horizontal and vertical coordinates of the service personnel's initial departure position. The order data time is from August 25, 2022 to September 10, 2022, with a time span of 17 days.
[0065] Since the overall scale of the experimental data set is large and the LNS algorithm itself is random, the results obtained from each experiment may be different, and even abnormal experimental data may appear. Therefore, in order to avoid abnormal experimental data interfering with the analysis of experimental conclusions and to prove that the order allocation scheme involved in the present invention is more reliable, the running time Time(s) and the optimization target value obj before and after the algorithm is improved under a large-scale data set are calculated by repeated experiments, so as to more accurately analyze the experimental results, as shown in Table 1. Among them, the lower the Time(s) value and the higher the obj value, the better the performance of the algorithm.
[0066] Table 1:
[0067] According to the data in Table 1, compared with the existing benchmark LNS algorithm, the order matching algorithm involved in the present invention has better comprehensive performance, can have a faster inference speed while ensuring a relatively high objective function value, is more in line with the actual needs, and can effectively improve the intelligence and rationality of order allocation in the online domestic service platform.
[0068] To further verify the rationality of the objective function model constructed in the present invention, in the experiment, the participation service ratio Rate of high-scoring service personnel (setting those with a score higher than 0.85 as high-scoring service personnel, with a full score of 1.00), the average travel distance Distance of service personnel for each order, and the average score Score of all service personnel participating in order services are used as evaluation indicators. Experiments are respectively conducted with only considering customer satisfaction, only considering service personnel income, comprehensively considering the average score of service personnel, the travel distance of service personnel in the process of serving each order, and the time interval factor between adjacent orders served by service personnel as the objective function. All experiments are based on the online domestic order allocation method considering the comprehensive benefits of the platform involved in the present invention. To avoid interference from abnormal experimental data that may be generated due to algorithm randomness, the strategy of repeated experiments is still adopted, and the average value of ten groups of non-abnormal experimental data is taken as the final experimental result, so as to ensure the reliability of the experimental results and conclusions. The calculation methods of the service ratio Rate, average travel distance Distance, and average score Score are as follows: ; ; ; Among them, Num grate_work represents the number of service personnel with a score higher than 0.85 and participating in order services, Num grate represents the number of service personnel with a score higher than 0.85 in the dataset, Num work represents the number of all service personnel participating in services, Distance i represents the average travel distance of the i-th service personnel participating in services for each order, Score i represents the score of the i-th service personnel participating in services. For customer satisfaction, the average score of service personnel participating in all order services is used to measure. The higher this average score value, the higher the customer satisfaction. And the income of service personnel is measured by the average number of orders received per person of all service personnel participating in order services. The higher this value, the higher the average income of practitioners in this industry. The comparison experimental results are shown in Table 2.
[0069] Table 2:
[0070] Analysis of the experimental data in Table 2 shows that: (1) When only considering customer satisfaction, although the proportion of highly rated service personnel participating in services is slightly higher than that under the objective function constructed based on the present invention, there is only an increase of 1.44 percentage points. In addition, the average rating of the service personnel actually participating in order services has only increased by 0.012. These increases have little impact on the overall efficiency of order matching. However, when using the objective function constructed according to the present invention, the average travel distance of service personnel for each order is shortened by nearly 2 kilometers compared to only considering customer satisfaction. This means that service personnel can participate in the services of more other orders per unit time, which to a certain extent promotes the increase of their own income. Therefore, the objective function constructed by the present invention is more reasonable than only considering customer satisfaction.
[0071] (2) When only considering the income of service personnel, although the travel distance for each order on average is shortened by about 0.5 km, when using the objective function constructed according to the present invention, the proportion of highly rated service personnel participating in services leads by 22.41 percentage points, and the average rating of the service personnel participating in order services also increases by 0.035. This means that when taking the objective function constructed by the present invention as the optimization target, it can take into account customer satisfaction more while not overly increasing the travel distance of service personnel during the process of serving orders, and is more in line with the expectations for the order allocation strategy in the online domestic service platform.
[0072] The above experimental results can effectively prove the effectiveness of an online domestic order allocation method considering the comprehensive benefits of the platform involved in the present invention, and at the same time can also verify the rationality of the objective function constructed in the present invention.
[0073] It should be understood that various forms of the processes shown above can be used, reordering, adding or deleting steps. For example, the steps recorded in the disclosure of the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in the present invention can be achieved, and no limitations are imposed herein.
[0074] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An online housekeeping service order allocation method considering the comprehensive benefits of the platform, characterized in that, Including: S1: Determine the constraint conditions and objective function of the prediction model for predicting the order allocation result according to the influence parameters affecting the comprehensive benefit of the domestic service platform; S2: Sort the orders to be analyzed according to the importance level, and divide the sorted order sequence to obtain the initial optimal solution for predicting the order allocation result, and the initial optimal solution satisfies the constraint conditions; S3: Based on the constraint conditions and objective function obtained in step S1, perform local search based on reinforcement learning on the initial optimal solution obtained in step S2, and predict the optimal order allocation result that satisfies the constraint conditions and the objective function.
2. The online housekeeping service order allocation method considering the comprehensive benefits of the platform according to claim 1, characterized in that, The objective function in step S1 is: ; Among them, obj represents the optimized target value, α, β, and γ are the weight parameters in the objective function, α + β + γ = 1, I represents the total number of service personnel, J represents the total number of orders, and f i represents the service score of the i-th service personnel, 0.5 < f i < 1, R ij takes values of 0 or 1, and when R ij = 0, it means that the j-th order is not served by the i-th service personnel. When R ij = 1, it means that the j-th order is served by the i-th service personnel. |I serve | represents the number of service personnel actually participating in order service. D jj’ represents the Euclidean distance between the j-th order and the j'-th order. X ijj’ takes values of 0 or 1, and when X ijj’ = 0, it means that the next order of the j-th order of the i-th service personnel is not the j'-th order. When X ijj’ = 1, it means that the next order of the j-th order of the i-th service personnel is the j'-th order, represents the time when the i-th service personnel executes the j'-th order, represents the time when the i-th service personnel executes the j-th order, represents the service duration of the j-th order.
3. The online housekeeping service order allocation method considering the comprehensive benefits of the platform according to claim 2, characterized in that The constraint conditions in step S1 include the service personnel number constraint for each order, the service time constraint for each order, the time connectivity and spatial connectivity of the order route in each order, and the position constraint that all service personnel start from the origin point at the first order, and the position of the i-th service personnel starting from the origin position to the position where the j-th order is located; where: The service personnel number constraint is: ; The service time constraint is at half-hour integer multiples between the earliest service start time and the latest service start time of each order, and all service personnel must arrive at the position where the order is located before the latest service start time of each order, that is: Among them, represents the latest start service time of the j-th order, represents the earliest start service time of the j-th order, a ij’ represents the moment when the i-th service staff arrives at the location of the j'-th order, and M represents a positive integer; The time connectivity satisfies: ; where v represents the passing speed of the service personnel; The spatial connectivity satisfies: The position constraint satisfies: Among them, p i represents the initial position of the i-th service staff, represents whether the i-th service staff directly reaches the j'-th order location from the initial position p i and its value is 0 or 1. When the value is 0, it means that the i-th service staff does not directly reach the j'-th order location from the initial position p When the value is 1, it means that the i-th service staff directly reaches the j'-th order location from the initial position p i without directly reaching the j'-th order location, and when the value is 1, it means that the i-th service staff directly reaches the j'-th order location from the initial position p i directly reaches the j'-th order location.
4. The online housekeeping service order allocation method considering the comprehensive benefits of the platform according to claim 2, characterized in that Step S2 includes: S21: Construct an order node set from service orders and service personnel, and construct a complete undirected graph from the order node set; use the adjacency matrix method to construct a graph Laplacian matrix based on the complete undirected graph, and calculate the eigenvalues of the graph Laplacian matrix, and the eigenvalues characterize the importance level of the corresponding order; S22: Perform K-means clustering on the order positions and order times to obtain several clusters; sort the orders within the clusters and the orders between the clusters respectively according to the eigenvalues of the orders within the clusters and the number of connections between the orders between the clusters; S23: Use the K-means clustering algorithm to divide the order sequence obtained in step S22 to obtain multiple subsequences, and calculate the comprehensive score of each order in the current subsequence; sort all orders according to the comprehensive score to obtain the initial optimal solution.
5. The online domestic service order allocation method considering the comprehensive benefits of the platform according to claim 4, characterized in that, The process of calculating the comprehensive score of each order in the current subsequence in step S23 includes: Calculate the Euclidean distance of each order from the first order in the current subsequence, and perform weighted calculation with the service score of each order to obtain the corresponding comprehensive score of each order.
6. The online housekeeping service order allocation method considering the comprehensive benefits of the platform according to claim 1, wherein Step S3 includes: S31: Input the initial optimal solution into multiple feature encoding modules based on the Transformer network to extract order features, obtain the initial state and initial features, and the orders in the initial state satisfy the constraint conditions; S32: Input the initial state and initial features obtained in step S31 into the disruption module, obtain and execute the current disruption action of deleting some order nodes, obtain the current disrupted state, where the orders in the current disrupted state satisfy the constraint condition, and determine the current disruption reward corresponding to the current disruption action; S33: Input the current disrupted state obtained in step S32 into the repair module, obtain and execute the current repair action, obtain the current repaired state and the corresponding order sequence, determine the current repair reward corresponding to the current repair action, where the orders in the current repaired state satisfy the constraint condition; S34: Input the order sequence obtained in step S33 into the feature encoding module for feature extraction, replace the initial features and initial state in step S32 with the extracted features and the current repaired state, and repeat steps S32 - S34, and output the optimal order allocation result from the repair module.
7. The online housekeeping service order allocation method considering the comprehensive benefits of the platform according to claim 6, characterized in that In each feature encoding module, input the initial optimal solution into the linear transformation layer for linear transformation; the output of the linear transformation layer is then successively subjected to two feature extractions and order rearrangement by the first attention sub-module and the second attention sub-module to obtain the initial state; In the first attention sub-module, use the multi-head attention mechanism to perform feature encoding processing on the input features, and the encoded features are then subjected to batch normalization processing to obtain the output; In the second attention sub-module, the input features are first processed by a feed-forward neural network, and the processed features are then subjected to batch normalization processing to obtain the output.
8. The online housekeeping service order allocation method considering the comprehensive benefits of the platform according to claim 6, characterized in that In the disruption module: Decode the initial state or the repaired state at the previous moment through the disruption policy network to obtain the current disruption action; After executing the current disruption action to obtain the current disrupted state, use the disruption reward function to evaluate the current disrupted state to obtain the current disruption reward.
9. The online housekeeping service order allocation method considering the comprehensive benefits of the platform according to claim 8, characterized in that, The disruption reward function is: ; Among them, represents the damage reward function, represents the current damage state, represents the current damage action, where m and n respectively represent the m-th order node and the n-th order node, represents the correlation degree between the m-th order node and the n-th order node, which is obtained by the following formula: ; where ε and η are weight values respectively, and d mn represents the Euclidean distance between the m-th order node and the n-th order node, that is: ; Among them, and respectively represent the position coordinates of the m-th order node and the n-th order node.
10. The online housekeeping service order allocation method considering the comprehensive benefits of the platform according to claim 8, characterized in that, The disruption policy network includes a multi-head self-attention module and an LSTM network; among them, the multi-head self-attention module performs feature interaction on the initial state and the repaired state at the previous moment, and then uses the LSTM network to decode the execution sequence of the interaction features to obtain the removal probability distribution of removing order nodes; remove order nodes from the order node set at the previous moment according to the removal probability distribution to complete the current disruption action.
11. The online domestic service order allocation method considering the comprehensive benefits of the platform according to claim 6, wherein, In the repair module: Decode the disrupted state at the previous moment through the repair policy network to obtain the current repair action; After executing the current repair action to obtain the current repaired state, the repair reward function takes the change in the objective function of the current repaired state at the previous and current moments as the current repair reward.
12. The online housekeeping service order allocation method considering the comprehensive benefits of the platform according to claim 11, characterized in that, In the repair policy network, use multi-head self-attention to decode the current disrupted state and the removed order nodes to obtain the repair probability distribution of the order node insertion positions; Select the first several positions with the same number as the number of removed order nodes as the insertion area; Insert the removed order nodes into the order node set at the previous moment in sequence according to the repair probability distribution to complete the current repair action.
13. The online domestic service order allocation method considering the comprehensive benefits of the platform according to claim 11, characterized in that, The repair reward function is: ; Among them, represents the repair reward function, represents the current repair status, represents the current repair action, where m and n respectively represent the m-th order node and the n-th order node, represents the degree of association between the m-th order node and the n-th order node, which is obtained by the following formula: ; where ε and η are weight values respectively, and d mn represents the Euclidean distance between the m-th order node and the n-th order node, that is: ; wherein, and respectively represent the position coordinates of the m-th order node and the n-th order node.
Citation Information
Patent Citations
Cloud household service platform service distribution system
CN107766996A
Carpooling order allocation method and system, terminal and storage medium
CN114331026A
Order selection scheduling optimization method based on self-learning
CN115829218A
Vehicle path planning method considering three-dimensional boxing problem
CN116136990A
Order delivery method and device, storage medium and electronic equipment
CN116205696A
Cited By
Order matching method and system based on machine learning
CN120579787A
A machine learning based order matching method and system
CN120579787B