B2b order intelligent management method and system based on reinforcement learning
By using a reinforcement learning-based B2B order management system, order data is collected in real time, and a quantitative evaluation model of the impact of order insertion and a multi-objective optimization reward function are employed to solve the problems of inflexible order processing and inaccurate decision-making, thus achieving efficient and intelligent order management.
Patent Information
- Application Number
- CN202510378028.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-03-26
AI Technical Summary
Existing B2B order management systems lack flexibility, rely on human experience leading to inaccurate decision-making, impacting customer satisfaction and increasing operating costs.
By employing a reinforcement learning-based approach, order data is collected in real time. Through a quantitative evaluation model of the impact of order insertion and a multi-objective optimization reward function, target orders are accurately identified and order insertion strategies are optimized, thereby improving the intelligence and accuracy of order processing.
It improves the flexibility and efficiency of order processing, ensures timely processing of high-priority orders, reduces operational risks and costs, and enhances the intelligence and decision-making accuracy of B2B order management.
Smart Images

Figure CN120494920B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electronic commerce, and particularly relates to a B2B order intelligent management method and system based on reinforcement learning. BACKGROUND
[0002] In the B2B (Business-to-Business) business environment, order management is one of the core links of enterprise operation. With the increasing fierce market competition and the diversification of customer demand, enterprises are facing more and more complex order processing challenges.
[0003] At present, most of the B2B order management systems of enterprises still adopt the traditional processing mode. In this mode, orders are processed in a fixed order, which lacks flexibility. Once a new order is inserted into the demand, the entire order processing flow often needs to be adjusted again, which not only consumes a lot of time and manpower, but also easily leads to order processing delay, affecting customer satisfaction.
[0004] Among them, in the order insertion decision, the existing method mainly relies on manual experience. The staff decides whether to insert an order and how to insert an order according to their subjective judgment of the order situation. However, this decision-making method has great limitations. On the one hand, manual judgment is easily affected by factors such as personal emotions and fatigue levels, leading to inaccurate decision-making; on the other hand, due to the lack of scientific quantitative evaluation tools, staff are difficult to comprehensively and accurately evaluate the impact of order insertion on the entire order processing flow, which is prone to decision-making errors, causing losses to the enterprise. In addition, unreasonable order insertion will also have a negative impact on the utilization of enterprise resources. For example, in the production link, inappropriate order insertion may lead to frequent adjustment of production equipment, increasing production preparation time and cost; in the logistics link, order insertion may lead to re-allocation of goods and changes in transportation routes, reducing logistics efficiency and increasing logistics cost. SUMMARY
[0005] The present application provides a B2B order intelligent management method and system based on reinforcement learning, which can improve the processing flexibility and intelligent degree of B2B order management, and improve the accuracy of order insertion decision-making and the utilization rate of order resources.
[0006] In order to solve the above technical problems, the first aspect of the present application discloses a B2B order intelligent management method based on reinforcement learning, which comprises:
[0007] Real-time collection of order data corresponding to the first order queue to be processed, the order data at least including order priority and delivery time;
[0008] determining, according to the order data, whether there is a target order in the first order queue that meets preset order insertion conditions, when the determination result is yes, extracting order feature information of the target order from the order data, and associating each order feature information with a second order queue currently processed;
[0009] performing quantitative evaluation on the order feature information according to a preset order insertion influence quantitative evaluation model, to obtain a quantitative evaluation result for the target order, the quantitative evaluation result including influence information of the target order inserted into the second order queue; and the quantitative evaluation result being used to indicate a specific order insertion matter of inserting the target order into the second order queue.
[0010] As an optional implementation, in the first aspect of the present application, the method further comprises:
[0011] determining, according to the quantitative evaluation result, a plurality of order insertion strategies for the target order;
[0012] calculating, according to a preset multi-objective optimization reward function, a reward value corresponding to each order insertion strategy in combination with the quantitative evaluation result; the multi-objective optimization reward function including a plurality of optimization parameters; each optimization parameter corresponding to an order requirement of a customer; and different customers having different attention proportions for all the order requirements;
[0013] determining, according to the reward value corresponding to each order insertion strategy, an optimal order insertion strategy from all the order insertion strategies.
[0014] As an optional implementation, in the first aspect of the present application, the calculating, according to a preset multi-objective optimization reward function, a reward value corresponding to each order insertion strategy in combination with the quantitative evaluation result comprises:
[0015] determining a plurality of optimization parameters corresponding to the preset multi-objective optimization reward function; the plurality of optimization parameters including an order overall delay parameter, a core customer order delay parameter, and an order resource utilization rate;
[0016] obtaining a target order requirement corresponding to an order customer of the target order, and determining a target parameter weight corresponding to each optimization parameter according to the target order requirement, and updating the multi-objective optimization reward function according to the target parameter weight corresponding to each optimization parameter;
[0017] for each order insertion strategy, determining a predicted parameter value associated with each optimization parameter in the order insertion strategy, and inputting the predicted parameter value associated with each optimization parameter in the order insertion strategy into the multi-objective optimization reward function to calculate a reward value corresponding to the order insertion strategy.
[0018] As an optional implementation, in the first aspect of the present application, the function formula corresponding to the multi-objective optimization reward function is:
[0019] R = ω1xf1(D overall ) + ω2xf2(D key-customer ) + ω3xf3(U resource )
[0020] Wherein, R is the reward value of each said single order strategy corresponding to the multi-objective optimization reward function; D overall is the order overall delay parameter; f1(D overall ) is the conversion function corresponding to the order overall delay parameter, and ω1 is the parameter weight corresponding to the order overall delay parameter; D key-customer is the core customer order delay parameter; f2(D key-customer ) is the conversion function corresponding to the core customer order delay parameter, and ω2 is the parameter weight corresponding to the core customer order delay parameter; U resource is the order resource utilization rate; f3(U resource ) is the conversion function corresponding to the order resource utilization rate; and ω3 is the parameter weight corresponding to the order resource utilization rate.
[0021] As an optional implementation, in the first aspect of the present application, the order characteristic information of the target order includes order basic information, time characteristic information, customer characteristic information and production and logistics characteristic information; the order basic information includes order amount, product type and quantity; the time characteristic information includes delivery time and order placement time; the customer characteristic information includes customer priority and customer historical order situation; the production and logistics characteristic information includes production process complexity and transportation requirements;
[0022] According to the preset single order influence quantitative evaluation model, the quantitative evaluation of the order characteristic information is performed to obtain the quantitative evaluation result for the target order, which includes:
[0023] The order processing information of the second order queue is obtained; the order processing information at least includes the processing progress of each second order in the second order queue;
[0024] The order characteristic information is input into the preset single order influence quantitative evaluation model, and the quantitative evaluation of the order characteristic information is performed by the single order influence quantitative evaluation model combined with the order processing information based on the preset single order influence factor, so as to obtain the quantitative evaluation result of the order characteristic information as the quantitative evaluation result for the target order;
[0025] The quantitative evaluation result comprises a calculation result corresponding to the order insertion influence factor, the order insertion influence factor is used to indicate influence information on processing progress of all the second orders after the target order is inserted into the second order queue, and the order insertion influence factor comprises delay days, production resource occupation rate and logistics resource occupation rate.
[0026] As an optional implementation, in the first aspect of the application, the order data further comprises goods type, production resource demand, logistics resource demand and customer level, and the first order queue comprises a plurality of to-be-processed first orders.
[0027] The method further comprises the following steps of:
[0028] The order priority and the delivery time are determined as first reference parameters, and the goods type, the production resource demand, the logistics resource demand and the customer level are determined as second reference parameters.
[0029] For each first order, a numerical quantification operation is performed on the first reference parameters and the second reference parameters corresponding to the first order according to a preset numerical quantification rule, so as to obtain a numerical quantification result corresponding to the first order.
[0030] According to the numerical quantification result corresponding to the first order, order classification is performed on the first order in combination with a preset order classification regulation, so as to obtain an order classification result corresponding to the first order; the order classification regulation comprises at least emergency order classification and regular order classification; the processing priority corresponding to the emergency order classification is higher than that of the regular order classification; and the order classification result corresponding to the first order is used to indicate that the first order belongs to the emergency order classification or the regular order classification.
[0031] According to the order classification result corresponding to each first order, it is determined whether there is a target order meeting a preset order insertion condition in all the first orders.
[0032] The target order meeting the preset order insertion condition is specifically that the order classification result corresponding to a first order indicates that the first order belongs to the emergency order classification.
[0033] As an optional implementation, in the first aspect of the application, the numerical quantification result corresponding to the first order comprises a first quantification value corresponding to the first reference parameter corresponding to the first order and a second quantification value corresponding to the second reference parameter corresponding to the first order.
[0034] The numerical quantification result corresponding to the first order is combined with a preset order classification regulation to perform order classification on the first order to obtain an order classification result corresponding to the first order, including:
[0035] According to the first quantification value corresponding to the first order, an upper order classification corresponding to the first order is determined in combination with a preset order classification regulation, and the upper order classification corresponding to the first order is used to indicate the order classification of the first order, and the order classification includes an emergency order or a regular order.
[0036] According to the second quantification value corresponding to the first order, a lower order classification corresponding to the first order is determined, and the lower order classification corresponding to the first order is used to indicate the ranking information of the first order in the order classification corresponding to the first order, and the ranking information is a ranking value or a ranking level; the higher the ranking value corresponding to the first order or the higher the ranking level corresponding to the first order, the higher the processing priority of the first order.
[0037] According to the upper order classification and the lower order classification corresponding to the first order, a target order classification of the first order is determined in combination with the order classification regulation as the order classification result corresponding to the first order.
[0038] The second aspect of the present application discloses a B2B order intelligent management system based on reinforcement learning, and the system comprises:
[0039] The acquisition module is used for acquiring order data corresponding to a first order queue to be processed in real time, and the order data at least includes order priority and delivery time.
[0040] The judgment module is used for judging whether a target order meeting a preset order insertion condition exists in the first order queue according to the order data.
[0041] The information extraction module is used for extracting order feature information of the target order from the order data when the judgment result of the judgment module is yes, and each order feature information is associated with a second order queue currently processed.
[0042] The quantitative evaluation module is used for performing quantitative evaluation on the order feature information according to a preset order insertion influence quantitative evaluation model to obtain a quantitative evaluation result for the target order, and the quantitative evaluation result includes influence information of the target order inserted into the second order queue; and the quantitative evaluation result is used to indicate specific order insertion matters of the target order inserted into the second order queue.
[0043] As an optional implementation, in the second aspect of the present application, the system further comprises:
[0044] determining module configured to determine a plurality of insertion order strategies for the target order according to the quantitative evaluation result;
[0045] a calculating module configured to calculate a reward value corresponding to each of the insertion order strategies according to a preset multi-objective optimization reward function in combination with the quantitative evaluation result; the multi-objective optimization reward function comprises a plurality of optimization parameters; each of the optimization parameters corresponds to an order requirement of a customer; and different customers have different attention proportions for all the order requirements;
[0046] The determining module is further configured to determine an optimal insertion order strategy from all the insertion order strategies according to the reward value corresponding to each of the insertion order strategies.
[0047] As an optional implementation form, in the second aspect of the present application, the manner in which the calculating module calculates the reward value corresponding to each of the insertion order strategies according to the preset multi-objective optimization reward function in combination with the quantitative evaluation result specifically comprises:
[0048] determining a plurality of optimization parameters corresponding to the preset multi-objective optimization reward function; the plurality of optimization parameters comprises an order overall delay parameter, a core customer order delay parameter and an order resource utilization rate;
[0049] obtaining a target order requirement corresponding to an order customer of the target order, and determining a target parameter weight corresponding to each of the optimization parameters according to the target order requirement, and updating the multi-objective optimization reward function according to the target parameter weight corresponding to each of the optimization parameters;
[0050] for each of the insertion order strategies, determining a predicted parameter value associated with each of the optimization parameters from the insertion order strategy, and inputting the predicted parameter value associated with each of the optimization parameters in the insertion order strategy into the multi-objective optimization reward function to calculate a reward value corresponding to the insertion order strategy.
[0051] As an optional implementation form, in the second aspect of the present application, a function formula corresponding to the multi-objective optimization reward function is:
[0052] R = ω1 × f1(D overall ) + ω2 × f2(D key-customer ) + ω3 × f3(U resource )
[0053] wherein, R is the reward value of each of the insertion order strategies corresponding to the multi-objective optimization reward function; D overall is the order overall delay parameter; f1(D overall ) is a conversion function corresponding to the order overall delay parameter, and ω1 is a parameter weight corresponding to the order overall delay parameter; D key-customera delay parameter of the core customer order; f2(D key-customer a conversion function corresponding to the delay parameter of the core customer order, and ω2 is a parameter weight corresponding to the delay parameter of the core customer order; resource an order resource utilization rate; f3(U resource a conversion function corresponding to the order resource utilization rate; and ω3 is a parameter weight corresponding to the order resource utilization rate.
[0054] As an optional implementation, in the second aspect of the present application, the order feature information of the target order includes order basic information, time feature information, customer feature information, and production and logistics feature information; the order basic information includes order amount, product type and quantity; the time feature information includes delivery time and order placement time; the customer feature information includes customer priority and customer historical order situation; the production and logistics feature information includes production process complexity and transportation requirements.
[0055] The manner in which the quantitative evaluation module performs quantitative evaluation on the order feature information according to the preset order insertion influence quantitative evaluation model to obtain the quantitative evaluation result for the target order specifically includes:
[0056] Obtaining order processing information of the second order queue; the order processing information at least includes processing progress of each second order in the second order queue;
[0057] Inputting the order feature information into a preset order insertion influence quantitative evaluation model, and taking a preset order insertion influence factor as a benchmark, the order insertion influence quantitative evaluation model performs quantitative evaluation on the order feature information in combination with the order processing information to obtain a quantitative evaluation result for the order feature information as the quantitative evaluation result for the target order;
[0058] The quantitative evaluation result includes a calculation result corresponding to the order insertion influence factor; the order insertion influence factor is used to indicate influence information on the processing progress of all the second orders after the target order is inserted into the second order queue; the order insertion influence factor includes delay days, production resource occupation rate and logistics resource occupation rate.
[0059] As an optional implementation, in the second aspect of the present application, the order data further includes cargo type, production resource demand, logistics resource demand and customer level; the first order queue includes a plurality of to-be-processed first orders.
[0060] The manner in which the judgment module judges whether there is a target order meeting a preset order insertion condition in the first order queue according to the order data specifically includes:
[0061] The order priority and the delivery time are determined as first reference parameters, and the cargo type, production resource demand, logistics resource demand, and customer level are determined as second reference parameters;
[0062] For each first order, a numerical quantification operation is performed on the first reference parameters and the second reference parameters corresponding to the first order according to a preset numerical quantification rule, to obtain a numerical quantification result corresponding to the first order;
[0063] According to the numerical quantification result corresponding to the first order, order classification is performed on the first order in combination with a preset order classification regulation, to obtain an order classification result corresponding to the first order; the order classification regulation at least includes emergency order classification and regular order classification; a processing priority corresponding to the emergency order classification is higher than a processing priority of the regular order classification; the order classification result corresponding to the first order is used to indicate that the first order belongs to the emergency order classification or the regular order classification;
[0064] According to the order classification result corresponding to each first order, it is determined whether a target order meeting a preset order insertion condition exists in all the first orders;
[0065] The target order meeting the preset order insertion condition is specifically a first order whose order classification result indicates that the first order belongs to the emergency order classification.
[0066] As an optional implementation, in the second aspect of the present application, the numerical quantification result corresponding to the first order includes a first quantification value corresponding to the first reference parameter corresponding to the first order and a second quantification value corresponding to the second reference parameter corresponding to the first order;
[0067] The manner in which the determining module performs order classification on the first order in combination with a preset order classification regulation according to the numerical quantification result corresponding to the first order to obtain an order classification result corresponding to the first order specifically includes:
[0068] According to the first quantification value corresponding to the first order, an upper order classification corresponding to the first order is determined in combination with a preset order classification regulation; the upper order classification corresponding to the first order is used to indicate an order category of the first order, and the order category includes an emergency order or a regular order;
[0069] determine a lower order level corresponding to the first order according to the second quantization value corresponding to the first order; the lower order level corresponding to the first order is used to indicate ranking information of the first order in the order category corresponding to the first order, and the ranking information is a ranking value or a ranking level; the higher the ranking value corresponding to the first order or the higher the ranking level corresponding to the first order, the higher the processing priority corresponding to the first order;
[0070] determine a target order level of the first order as an order level result corresponding to the first order according to the upper order level corresponding to the first order, the lower order level, and the order level regulation.
[0071] The third aspect of the present application discloses another B2B order intelligent management device based on reinforcement learning, and the device comprises:
[0072] a memory storing executable program codes;
[0073] a processor coupled with the memory;
[0074] The processor invokes the executable program codes stored in the memory to execute the B2B order intelligent management method based on reinforcement learning disclosed in the first aspect of the present application.
[0075] The fourth aspect of the present application discloses a computer storage medium storing computer instructions, and the computer instructions are used to execute the B2B order intelligent management method based on reinforcement learning disclosed in the first aspect of the present application when being invoked.
[0076] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0077] In the embodiment of the present application, a B2B order intelligent management method based on reinforcement learning is provided, which comprises: collecting order data corresponding to a first order queue to be processed in real time, the order data at least including order priority and delivery time; judging whether there is a target order meeting preset order insertion conditions in the first order queue according to the order data, when the result of the judgment is yes, extracting order feature information of the target order from the order data, and associating each order feature information with a second order queue currently processed; performing quantitative evaluation on the order feature information according to a preset order insertion influence quantitative evaluation model, obtaining a quantitative evaluation result for the target order, the quantitative evaluation result including influence information of the target order inserted into the second order queue; and the quantitative evaluation result is used to indicate specific order insertion matters of the target order inserted into the second order queue. It can be seen that, by collecting order data (including order priority and delivery time) in real time, the target order meeting the preset order insertion conditions can be accurately identified, and its order feature information can be extracted, that is, the customer order with temporary / special needs can be responded in time, which is beneficial to improving the flexibility and efficiency of order processing. Further, the order feature information is deeply quantitatively evaluated by using the preset order insertion influence quantitative evaluation model, and detailed influence information of the target order inserted into the second order queue is obtained. By scientifically evaluating the order insertion influence, the order queue management is optimized, the timely processing of high-priority orders is ensured, the operation risk and cost caused by improper order insertion are reduced, and the intelligent degree of B2B order management and the accuracy of decision-making are greatly improved. BRIEF DESCRIPTION OF DRAWINGS
[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0079] Figure 1 is a flow diagram of a B2B order intelligent management method based on reinforcement learning disclosed by the embodiment of the present application;
[0080] Figure 2 is a flow diagram of another B2B order intelligent management method based on reinforcement learning disclosed by the embodiment of the present application;
[0081] Figure 3 is a structure diagram of a B2B order intelligent management system based on reinforcement learning disclosed by the embodiment of the present application;
[0082] Figure 4 is a structure diagram of another B2B order intelligent management system based on reinforcement learning disclosed by the embodiment of the present application;
[0083] Figure 5 is another structure schematic diagram of the B2B order intelligent management device based on reinforcement learning disclosed by the embodiment of the present application. DETAILED DESCRIPTION
[0084] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0085] The terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish different objects, not to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or end including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product, or end.
[0086] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily mutually exclusive or alternative to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0087] The present application discloses a B2B order intelligent management method and system based on reinforcement learning. By collecting order data in real time (including order priority and delivery time), the target order meeting the preset order insertion condition can be accurately identified, and the order feature information thereof can be extracted, that is, the customer order with temporary / special needs can be responded in a timely manner, which is beneficial to improving the flexibility and efficiency of order processing. Further, the order feature information is quantitatively evaluated in depth by using a preset order insertion influence quantitative evaluation model, and detailed influence information of the target order inserted into the second order queue is obtained. By scientifically evaluating the order insertion influence, the order queue management is optimized, the timely processing of high-priority orders is ensured, the operation risk and cost caused by improper order insertion are reduced, and the intelligent degree of B2B order management and the accuracy of decision-making are greatly improved. The following will be described in detail.
[0088] Embodiment one
[0089] Referring to Figure 1 , Figure 1 is a flowchart of a B2B order intelligent management method based on reinforcement learning disclosed by an embodiment of the present application. Wherein, Figure 1 The B2B order intelligent management method based on reinforcement learning described can be applied to a B2B order intelligent management system based on reinforcement learning, and the present application is not limited. As Figure 1 shown, the B2B order intelligent management method based on reinforcement learning can include the following operations:
[0090] 101, real-time collection of order data corresponding to the first order queue to be processed, the order data at least including order priority and delivery time.
[0091] In the present application, the first order queue includes one or more first orders to be processed, and the order data includes sub-order data corresponding to each first order. Further, the sub-order data includes sub-order priority and sub-delivery time corresponding to each first order.
[0092] 102, according to the order data, determine whether there is a target order in the first order queue that meets the preset order insertion condition.
[0093] In the present application, the target order of the preset order insertion condition can be an urgent order that needs to be processed first. It should be noted that the number of target orders can be one or more, and the present application is not limited.
[0094] 103, when the result of step 102 is yes, extract the order feature information of the target order from the order data, and associate each order feature information with the second order queue currently processed.
[0095] In the present application, when the number of target orders is more than one, the order feature information of the target order includes sub-order feature information corresponding to each target order.
[0096] 104, according to the preset order insertion influence quantitative evaluation model, performing quantitative evaluation on the order feature information to obtain the quantitative evaluation result for the target order, and the quantitative evaluation result includes the influence information of the target order inserted into the second order queue.
[0097] In the present application, the quantitative evaluation result is used to indicate the specific order insertion matter of inserting the target order into the second order queue.
[0098] As can be seen, the implementation Figure 1The described B2B order intelligent management method based on reinforcement learning can accurately identify target orders that meet the preset order insertion conditions by collecting order data (including order priority and delivery time) in real time, and extract order feature information, that is, timely respond to customer orders with temporary / special needs, which helps to improve the flexibility and efficiency of order processing. Further, the order feature information is quantitatively evaluated in depth using a preset order insertion influence quantitative evaluation model to obtain detailed influence information of the target order inserted into the second order queue. By scientifically evaluating the influence of order insertion, the order queue management is optimized to ensure timely processing of high-priority orders, while reducing operational risks and costs caused by improper order insertion, greatly improving the intelligent level of B2B order management and the accuracy of decision-making.
[0099] In an optional embodiment, the order feature information of the target order includes order basic information, time feature information, customer feature information, and production and logistics feature information; the order basic information includes order amount, product type and quantity; the time feature information includes delivery time and order placement time; the customer feature information includes customer priority and customer historical order situation; the production and logistics feature information includes production process complexity and transportation requirements;
[0100] The manner of performing quantitative evaluation on the order feature information according to the preset order insertion influence quantitative evaluation model to obtain the quantitative evaluation result for the target order in the above step 104 specifically includes:
[0101] Obtain order processing information of the second order queue; the order processing information at least includes processing progress of each second order in the second order queue;
[0102] Input the order feature information into the preset order insertion influence quantitative evaluation model, and take the preset order insertion influence factor as the benchmark, and perform quantitative evaluation on the order feature information by the order insertion influence quantitative evaluation model combined with the order processing information to obtain the quantitative evaluation result for the order feature information as the quantitative evaluation result for the target order;
[0103] Wherein, the quantitative evaluation result includes a calculation result corresponding to the order insertion influence factor; the order insertion influence factor is used to indicate the influence information of the processing progress of all second orders after the target order is inserted into the second order queue; the order insertion influence factor includes delay days, production resource occupation rate and logistics resource occupation rate.
[0104] It can be seen that in the optional embodiment, by comprehensively integrating the order characteristic information of the target order, covering multi-dimensional data such as order basis, time, customer, production and logistics, a rich and detailed data basis is provided for accurate evaluation of the impact of inserting an order. Specifically, by introducing a preset quantitative evaluation model of the impact of inserting an order, combined with real-time processing information of the second order queue, the order characteristic information is analyzed in depth based on the impact factor of inserting an order. The processing operation of the deep quantitative analysis not only considers the attributes of the order itself, but also dynamically integrates the queue processing state, ensuring the comprehensiveness and accuracy of the evaluation results. In addition, the quantitative evaluation results are directly related to key indicators such as delay days, production resource occupation rate and logistics resource occupation rate, providing intuitive and quantitative information on the impact of inserting an order for decision makers. Through the detailed quantitative evaluation operation of the order characteristic information, the intelligent level of order insertion decision-making is improved, while effectively balancing order priority, processing efficiency and resource utilization, further improving the precision and efficiency of B2B order management.
[0105] In another optional embodiment, the order data further includes goods type, production resource demand, logistics resource demand and customer level; the first order queue includes a plurality of first orders to be processed;
[0106] The manner of determining whether there is a target order satisfying the preset insertion order condition in the first order queue according to the order data in the above step 102 specifically includes:
[0107] The order priority and the delivery time are determined as the first reference parameter, and the goods type, the production resource demand, the logistics resource demand and the customer level are determined as the second reference parameter;
[0108] For each first order, according to a preset numerical quantification rule, a numerical quantification operation is performed on the first reference parameter and the second reference parameter corresponding to the first order, to obtain a numerical quantification result corresponding to the first order;
[0109] According to the numerical quantification result corresponding to the first order, combined with a preset order classification regulation, an order classification is performed on the first order to obtain an order classification result corresponding to the first order; the order classification regulation at least includes an urgent order classification and a regular order classification; the processing priority corresponding to the urgent order classification is higher than the processing priority of the regular order classification; the order classification result corresponding to the first order is used to indicate that the first order belongs to the urgent order classification or the regular order classification;
[0110] According to the order classification result corresponding to each first order, it is determined whether there is a target order satisfying the preset insertion order condition in all first orders;
[0111] The target order meeting the preset single insertion condition is specifically an order classification result of a first order, indicating that the first order belongs to an urgent order classification.
[0112] It can be seen that in the optional embodiment, by introducing multi-order data such as goods type, production resource demand, logistics resource demand, and customer level, the evaluation dimension for each first order is enriched. In addition, by establishing a double-standard parameter system, the order priority and delivery time are set as the first standard, and the remaining key elements are classified as the second standard, and accurate quantification is performed according to the preset numerical quantification rule, ensuring that the order characteristics are comprehensive and objective. Further, the order emergency degree can be intelligently divided in combination with the order classification regulations, improving the identification accuracy of urgent orders that need to be processed first. Through this method, not only the flexibility and response speed of order processing are greatly improved, but also accurate classification and intelligent single insertion judgment are ensured to ensure that urgent orders are processed in a timely and efficient manner, realize the optimal allocation of resources, and improve customer satisfaction.
[0113] In yet another optional embodiment, the numerical quantification result corresponding to the first order includes a first quantification value corresponding to the first standard parameter corresponding to the first order, and a second quantification value corresponding to the second standard parameter corresponding to the first order.
[0114] The above method of performing order classification on the first order according to the numerical quantification result corresponding to the first order in combination with the preset order classification regulations to obtain an order classification result corresponding to the first order specifically includes:
[0115] According to the first quantification value corresponding to the first order, in combination with the preset order classification regulations, the upper order classification corresponding to the first order is determined, and the upper order classification corresponding to the first order is used to indicate the order classification of the first order. The order classification includes an urgent order or a regular order.
[0116] According to the second quantification value corresponding to the first order, the lower order classification corresponding to the first order is determined. The lower order classification corresponding to the first order is used to indicate the ranking information of the first order in its corresponding order classification. The ranking information is a ranking numerical value or a ranking level. The higher the ranking numerical value corresponding to the first order, or the higher the ranking level corresponding to the first order, the higher the processing priority of the first order.
[0117] According to the upper order classification and the lower order classification corresponding to the first order, in combination with the order classification regulations, the target order classification of the first order is determined as the order classification result corresponding to the first order.
[0118] In the optional embodiment, the specific calculation method of the first quantification value is as follows:
[0119] According to a preset classification standard, the order priority is subjected to priority assessment and quantification processing, and an order priority-quantitative value corresponding to the order priority is obtained.
[0120] A difference between the delivery time and the current time is calculated, and the difference is subjected to quantification processing according to a preset time threshold, and a delivery time-quantitative value corresponding to the delivery time is obtained.
[0121] The order priority-quantitative value and the delivery time-quantitative value are subjected to weighted calculation, and a first weighted calculation result is obtained as a first quantitative value.
[0122] In the optional embodiment, for the order priority, the highest priority can be set to 5, and the priority decreases in turn, and the lowest priority is set to 1, so as to intuitively reflect the difference in the urgency of the order in the form of quantitative data. For the interaction time, the difference between the delivery time and the current time can be calculated according to the urgency of the delivery time, and the difference is quantified in combination with the preset time threshold; for example, if the delivery time is within 24 hours after the current time, a higher quantitative value is given; if the delivery time is more relaxed, the quantitative value is correspondingly reduced. In this way, the influence of the delivery time on the order processing can be accurately measured.
[0123] In the optional embodiment, the calculation method of the second quantitative value is similar to that of the first quantitative value. The specific calculation method of the second quantitative value is as follows:
[0124] Determine the goods type and characteristic information corresponding to the goods type;
[0125] Determine the demand quantity and scarcity corresponding to the production resource demand;
[0126] Determine the logistics attributes corresponding to the goods transportation mode, transportation distance and transportation difficulty corresponding to the logistics resource demand;
[0127] Determine the customer information corresponding to the historical order situation, cooperation time and credit level corresponding to the customer level;
[0128] According to the preset quantification rule corresponding to each sub-reference parameter in the second reference parameter, the sub-reference parameter is subjected to quantification calculation and weighted calculation, and a second weighted calculation result corresponding to all sub-reference parameters is obtained as a second quantitative value; wherein all sub-reference parameters of the second reference parameter include the goods type, the production resource demand, the logistics resource demand and the customer level.
[0129] In the optional embodiment, the cargo category and characteristic information can be a perishable, easily damaged or high-value cargo label, and the presence of the cargo label can be assigned a higher quantitative value to reflect its special needs in the processing process. The production resource demand corresponding demand quantity can include equipment demand, personnel demand and raw material demand; further, the ratio of resource demand quantity to available resource quantity can be used as a quantitative index, and the higher the ratio, the greater the quantitative value, indicating that the production resource demand has a greater impact on order processing.
[0130] In the optional embodiment, for logistics resource demand, if there is cargo that needs special transportation equipment or long-distance transportation, a higher quantitative value can be assigned to reflect the impact of logistics resource demand on order processing. For customer level, the customer level can be assigned a corresponding quantitative value according to the above-mentioned various customer information. For example, high-level customers can correspond to a higher quantitative value to reflect their important position in enterprise customer service.
[0131] As can be seen, in the optional embodiment, by refining the numerical quantitative result into a first quantitative value and a second quantitative value, a double-precision classification of orders is achieved. First, according to the first quantitative value and the preset order classification regulations, the upper classification of the order (emergency or routine) is quickly determined, and then the second quantitative value is used to further analyze the specific ranking of the order in the lower classification, ensuring the meticulousness of the priority evaluation. Through the double-classification mechanism of this setting, not only the flexibility and accuracy of order processing are improved, but also the quantitative basis for the priority processing of emergency orders is provided through the clear ranking information, effectively optimizing resource allocation and response speed. Further enhancing the intelligent level of order management, and being conducive to improving the accuracy of grasping customer demand, and also being conducive to improving operational efficiency and customer satisfaction.
[0132] Embodiment Two
[0133] Please refer to Figure 2 , Figure 2 is another flowchart of a B2B order intelligent management method based on reinforcement learning disclosed in the embodiments of the present application. Among them, Figure 2 The B2B order intelligent management method based on reinforcement learning described can be applied to a B2B order intelligent management system based on reinforcement learning, and the embodiments of the present application are not limited. As Figure 2 shown, the B2B order intelligent management method based on reinforcement learning can include the following operations:
[0134] 201, real-time collection of order data corresponding to the first order queue to be processed, the order data at least including order priority and delivery time.
[0135] 202. Determine, according to the order data, whether there is a target order in the first order queue that meets the preset order insertion condition.
[0136] 203. When the determination result of step 202 is yes, extract order feature information of the target order from the order data, and associate each order feature information with the second order queue currently processed.
[0137] 204. According to the preset order insertion influence quantitative evaluation model, perform quantitative evaluation on the order feature information to obtain a quantitative evaluation result for the target order, and the quantitative evaluation result includes influence information of the target order inserted into the second order queue.
[0138] In the embodiments of the present application, for other descriptions of steps 201-204, please refer to other specific descriptions of steps 101-104 in Embodiment I, and the embodiments of the present application will not be repeated.
[0139] 205. According to the quantitative evaluation result, determine a plurality of order insertion strategies for the target order.
[0140] 206. According to the preset multi-target optimization reward function, combine the quantitative evaluation result to calculate a reward value corresponding to each order insertion strategy.
[0141] In the embodiments of the present application, the multi-target optimization reward function includes a plurality of optimization parameters; each optimization parameter corresponds to an order requirement of a customer; and different customers have different attention ratios for all order requirements.
[0142] 207. According to the reward value corresponding to each order insertion strategy, determine an optimal order insertion strategy from all order insertion strategies.
[0143] It can be seen that the implementation Figure 2 The described B2B order intelligent management method based on reinforcement learning can generate a plurality of order insertion strategies according to the quantitative evaluation result, and can introduce a preset multi-target optimization reward function for accurate evaluation. Among them, the multi-target optimization reward function integrates a plurality of optimization parameters, each parameter accurately corresponds to a customer order requirement, and flexibly adapts to the differentiated attention ratio of different customers to the order requirement. Then, the reward value of each order insertion strategy can be calculated accordingly, so as to intelligently select the optimal order insertion strategy. Through the selection mechanism of the optimal order insertion strategy, on the basis of significantly enhancing the flexibility and adaptability of order processing, through the multi-target optimization mechanism, it ensures that the order processing scheme can accurately match the diversified needs of customers, effectively improves the satisfaction of customers using the B2B platform for transaction, and also optimizes the resource allocation, reduces the operating cost, and improves the intelligent degree and the degree of refinement of B2B order management.
[0144] In an optional embodiment, the manner of calculating the reward value corresponding to each insertion strategy according to the preset multi-objective optimization reward function and the quantitative evaluation result in step 206 specifically comprises:
[0145] determining a plurality of optimization parameters corresponding to the preset multi-objective optimization reward function; the plurality of optimization parameters comprise an order overall delay parameter, a core customer order delay parameter, and an order resource utilization rate;
[0146] obtaining a target order requirement corresponding to an order customer of a target order, and determining a target parameter weight corresponding to each optimization parameter according to the target order requirement, and updating the multi-objective optimization reward function according to the target parameter weight corresponding to each optimization parameter;
[0147] for each insertion strategy, determining a predicted parameter value associated with each optimization parameter from the insertion strategy, and inputting the predicted parameter value associated with each optimization parameter in the insertion strategy into the multi-objective optimization reward function to calculate a reward value corresponding to the insertion strategy.
[0148] In the optional embodiment, the function formula corresponding to the multi-objective optimization reward function is:
[0149] R = ω1 × f1(D overall ) + ω2 × f2(D key-customer ) + ω3 × f3(U resource )
[0150] wherein, R is the reward value of each insertion strategy corresponding to the multi-objective optimization reward function; D overall is the order overall delay parameter; f1(D overall ) is a conversion function corresponding to the order overall delay parameter, and ω1 is a parameter weight corresponding to the order overall delay parameter; D key-customer is the core customer order delay parameter; f2(D key-customer ) is a conversion function corresponding to the core customer order delay parameter, and ω2 is a parameter weight corresponding to the core customer order delay parameter; U resource is the order resource utilization rate; f3(U resource ) is a conversion function corresponding to the order resource utilization rate; and ω3 is a parameter weight corresponding to the order resource utilization rate.
[0151] Further, the calculation formula of the conversion function f1 corresponding to the order overall delay parameter is:
[0152]
[0153] wherein, T1 is a preset threshold value, when the order overall delay exceeds T1, it is considered that the insertion strategy has too great an impact on the overall delivery of the order, and the reward contribution is 0, in the range 0 ≤ D overallThe smaller the delay is within T1, the greater the reward contribution is.
[0154] The calculation formula of the conversion function f2 corresponding to the core customer order delay parameter is:
[0155]
[0156] T2 is the threshold value of the core customer order delay parameter, which reflects the strictness of the enterprise's requirement for the delivery of core customer orders.
[0157] The calculation formula of the conversion function f3 corresponding to the resource utilization rate is:
[0158] f3(U resource )=U resource
[0159] In this optional embodiment, specifically, it is assumed that there are three insertion order strategies i1, i2 and i3, and the above-mentioned parameter weights are set as ω1=0.5, ω2=0.3 and ω3=0.2, and the threshold values T1=10 days and T2=5 days. The order overall delay parameters of the three insertion order strategies are 8 days, 12 days and 5 days respectively; the core customer order delay parameters of the three insertion order strategies are 3 days, 2 days and 6 days respectively; the resource utilization rates of the three insertion order strategies are 0.8, 0.9 and 0.7 respectively, and further, the reward values of each insertion order strategy are calculated as follows:
[0160] Taking the insertion order strategy i1 as an example:
[0161]
[0162] f3(U resource )=0.8
[0163] R1=0.5×0.2+0.3×0.4+0.2×0.8=0.38
[0164] Similarly, the above data are substituted in turn, and R2=0.36 corresponding to the insertion order strategy i2 and R3=0.39 corresponding to the insertion order strategy i3 can be calculated. By comparing the reward values of each insertion order strategy, it can be obtained that the value corresponding to R3 is the largest, and then i3 is determined as the optimal insertion order strategy.
[0165] It can be seen that in the optional embodiment, the calculation manner of the multi-objective optimization reward function is refined, by explicitly including key optimization parameters such as order overall delay parameter, core customer order delay parameter and order resource utilization, the enterprise operation core concerns are accurately connected. The specific requirements of the target order customer can be dynamically obtained, and the target weight of each optimization parameter is flexibly adjusted accordingly, so as to realize the individual customization of the reward function. For each order insertion strategy, the associated parameter value of each optimization parameter can be analyzed and predicted in depth, and the updated reward function is substituted to calculate the quantitative reward value. Therefore, the evaluation accuracy and adaptability of the order insertion strategy are significantly improved, ensuring that the order insertion strategy not only meets the diversified customer demand, but also optimizes the order processing efficiency, resource utilization and customer satisfaction on a global level.
[0166] Embodiment three
[0167] Please refer to Figure 3 , Figure 3 is a structural schematic diagram of a B2B order intelligent management system based on reinforcement learning disclosed by the embodiment of the application. The B2B order intelligent management system based on reinforcement learning can be used in a B2B transaction scenario, which is not limited by the embodiment of the application. As shown in Figure 3 , the B2B order intelligent management system based on reinforcement learning can include a collection module 301, a judgment module 302, an information extraction module 303 and a quantitative evaluation module 304, wherein:
[0168] The collection module 301 is used to collect order data corresponding to the first order queue to be processed in real time, and the order data at least includes order priority and delivery time.
[0169] The judgment module 302 is used to determine whether there is a target order meeting the preset order insertion condition in the first order queue according to the order data.
[0170] The information extraction module 303 is used to extract order feature information of the target order from the order data when the judgment result of the judgment module 302 is yes, and each order feature information is associated with the second order queue currently processed.
[0171] The quantitative evaluation module 304 is used to perform quantitative evaluation on the order feature information according to the preset order insertion influence quantitative evaluation model, to obtain a quantitative evaluation result for the target order, and the quantitative evaluation result includes influence information of the target order inserted into the second order queue; the quantitative evaluation result is used to indicate specific order insertion matters of inserting the target order into the second order queue.
[0172] It can be seen that the embodiment Figure 3The described reinforcement learning-based B2B order intelligent management system can accurately identify target orders that meet preset order insertion conditions by collecting order data (including order priority and delivery time) in real time, and extract order feature information, that is, timely respond to customer orders with temporary / special needs, which helps to improve the flexibility and efficiency of order processing. Further, the order feature information is quantitatively evaluated using a preset order insertion impact quantitative evaluation model to obtain detailed impact information of the target order inserted into the second order queue. By scientifically evaluating the impact of order insertion, the order queue management is optimized to ensure timely processing of high-priority orders, while reducing operational risks and costs caused by improper order insertion, greatly improving the intelligent level of B2B order management and the accuracy of decision-making.
[0173] In an optional embodiment, please refer to Figure 4 , Figure 4 is another structure diagram of the reinforcement learning-based B2B order intelligent management system disclosed in the embodiment of the application. As Figure 4 shown, the system further includes a determination module 305 and a calculation module 306, wherein:
[0174] The determination module 305 is configured to determine a plurality of order insertion strategies for the target order according to the quantitative evaluation result.
[0175] The calculation module 306 is configured to calculate a reward value corresponding to each order insertion strategy according to a preset multi-objective optimization reward function in combination with the quantitative evaluation result; the multi-objective optimization reward function includes a plurality of optimization parameters; each optimization parameter corresponds to an order requirement of a customer; and different customers have different attention proportions for all order requirements.
[0176] The determination module 305 is further configured to determine an optimal order insertion strategy from all order insertion strategies according to the reward value corresponding to each order insertion strategy.
[0177] As can be seen, in this optional embodiment, a plurality of order insertion strategies are generated according to the quantitative evaluation result, and a preset multi-objective optimization reward function is introduced for accurate evaluation. The multi-objective optimization reward function integrates a plurality of optimization parameters, each parameter accurately corresponds to a customer order requirement, and flexibly adapts to the differentiated attention proportions of different customers to order requirements. Then, the reward values of each order insertion strategy are calculated to intelligently select the optimal order insertion strategy. Through the selection mechanism of the optimal order insertion strategy, on the basis of significantly enhancing the flexibility and adaptability of order processing, through the multi-objective optimization mechanism, the order processing scheme can accurately match the diversified needs of customers, effectively improving the satisfaction of customers using the B2B platform for transactions, while optimizing resource allocation, reducing operating costs, and improving the intelligent level and refinement level of B2B order management.
[0178] In another optional embodiment, the manner in which the computing module 306 calculates the reward value corresponding to each insertion order strategy according to the preset multi-objective optimization reward function in combination with the quantitative evaluation result specifically includes:
[0179] determining a plurality of optimization parameters corresponding to the preset multi-objective optimization reward function; the plurality of optimization parameters include an order overall delay parameter, a core customer order delay parameter, and an order resource utilization rate;
[0180] obtaining a target order requirement corresponding to an order customer of a target order, and determining a target parameter weight corresponding to each optimization parameter according to the target order requirement, and updating the multi-objective optimization reward function according to the target parameter weight corresponding to each optimization parameter;
[0181] for each insertion order strategy, determining a predicted parameter value associated with each optimization parameter from the insertion order strategy, and inputting the predicted parameter value associated with each optimization parameter in the insertion order strategy into the multi-objective optimization reward function to calculate a reward value corresponding to the insertion order strategy.
[0182] In this optional embodiment, the function formula corresponding to the multi-objective optimization reward function is:
[0183] R = ω1 x f1(D overall ) + ω2 x f2(D key-customer ) + ω3 x f3(U resource )
[0184] wherein R is the reward value of each insertion order strategy corresponding to the multi-objective optimization reward function; D overall is the order overall delay parameter; f1(D overall ) is the conversion function corresponding to the order overall delay parameter, and ω1 is the parameter weight corresponding to the order overall delay parameter; D key-customer is the core customer order delay parameter; f2(D key-customer ) is the conversion function corresponding to the core customer order delay parameter, and ω2 is the parameter weight corresponding to the core customer order delay parameter; U resource is the order resource utilization rate; f3(U resource ) is the conversion function corresponding to the order resource utilization rate; and ω3 is the parameter weight corresponding to the order resource utilization rate.
[0185] It can be seen that in the optional embodiment, the calculation method of the multi-objective optimization reward function is refined. By explicitly including key optimization parameters such as order overall delay parameters, core customer order delay parameters, and order resource utilization, the enterprise operation core concerns are accurately connected. The specific requirements of the target order customer can be dynamically obtained, and the target weight of each optimization parameter can be flexibly adjusted accordingly to realize personalized customization of the reward function. For each order insertion strategy, the associated parameter values of each optimization parameter can be analyzed and predicted in depth, and the updated reward function can be used for accurate calculation to obtain the quantitative reward value. Therefore, the evaluation accuracy and adaptability of the order insertion strategy are significantly improved, ensuring that the order insertion strategy not only meets the diversified customer demand, but also optimizes the order processing efficiency, resource utilization, and customer satisfaction on a global level.
[0186] In yet another optional embodiment, the order characteristic information of the target order includes order basic information, time characteristic information, customer characteristic information, and production and logistics characteristic information; the order basic information includes order amount, product type and quantity; the time characteristic information includes delivery time and order placement time; the customer characteristic information includes customer priority and customer historical order situation; the production and logistics characteristic information includes production process complexity and transportation requirements;
[0187] The manner in which the quantitative evaluation module 304 performs quantitative evaluation on the order characteristic information according to the preset order insertion influence quantitative evaluation model to obtain the quantitative evaluation result for the target order includes:
[0188] Obtain order processing information of the second order queue; the order processing information at least includes the processing progress of each second order in the second order queue;
[0189] Input the order characteristic information into the preset order insertion influence quantitative evaluation model, and take the preset order insertion influence factor as the benchmark. The order insertion influence quantitative evaluation model combines the order processing information to perform quantitative evaluation on the order characteristic information, and obtains the quantitative evaluation result for the order characteristic information as the quantitative evaluation result for the target order;
[0190] The quantitative evaluation result includes a calculation result corresponding to the order insertion influence factor; the order insertion influence factor is used to indicate the influence information of the processing progress of all second orders after the target order is inserted into the second order queue; the order insertion influence factor includes delay days, production resource occupation rate and logistics resource occupation rate.
[0191] It can be seen that in the optional embodiment, by comprehensively integrating the order characteristic information of the target order, covering multi-dimensional data such as order basis, time, customer, production and logistics, a rich and detailed data basis is provided for accurate evaluation of the impact of inserting an order. Specifically, by introducing a preset quantitative evaluation model of the impact of inserting an order, combined with real-time processing information of the second order queue, the order characteristic information is analyzed in depth based on the insertion impact factor. The processing operation of the deep quantitative analysis not only considers the attributes of the order itself, but also dynamically integrates the queue processing state, ensuring the comprehensiveness and accuracy of the evaluation results. In addition, the quantitative evaluation results are directly related to key indicators such as delay days, production resource occupation rate and logistics resource occupation rate, providing intuitive and quantitative insertion impact information for decision makers. Through the detailed quantitative evaluation operation of the order characteristic information, the intelligent level of order insertion decision-making is improved, and the balance between order priority, processing efficiency and resource utilization is effectively achieved, further improving the precision and efficiency of B2B order management.
[0192] In another optional embodiment, the order data further includes goods type, production resource demand, logistics resource demand, and customer level; the first order queue includes a plurality of first orders to be processed;
[0193] The manner in which the determination module 302 determines whether there is a target order in the first order queue that meets the preset insertion condition according to the order data specifically includes:
[0194] The order priority and the delivery time are determined as the first reference parameter, and the goods type, the production resource demand, the logistics resource demand, and the customer level are determined as the second reference parameter;
[0195] For each first order, according to a preset numerical quantification rule, a numerical quantification operation is performed on the first reference parameter and the second reference parameter corresponding to the first order to obtain a numerical quantification result corresponding to the first order;
[0196] According to the numerical quantification result corresponding to the first order, combined with a preset order classification regulation, an order classification is performed on the first order to obtain an order classification result corresponding to the first order; the order classification regulation at least includes an urgent order classification and a regular order classification; the processing priority corresponding to the urgent order classification is higher than the processing priority of the regular order classification; the order classification result corresponding to the first order is used to indicate that the first order belongs to the urgent order classification or the regular order classification;
[0197] According to the order classification result corresponding to each first order, it is determined whether there is a target order in all first orders that meets the preset insertion condition;
[0198] The target order meeting the preset single insertion condition is specifically an order classification result of a first order, which indicates that the first order belongs to an urgent order classification.
[0199] It can be seen that in the optional embodiment, by introducing multi-order data such as goods type, production resource demand, logistics resource demand, and customer level, the evaluation dimension for each first order is enriched. In addition, by establishing a double-standard parameter system, the order priority and delivery time are set as the first standard, and the remaining key elements are classified as the second standard, and accurate quantification is performed according to the preset numerical quantification rule, ensuring that the order characteristics are comprehensive and objective. Further, the order emergency degree can be intelligently divided in combination with the order classification regulations, improving the identification accuracy of urgent orders that need to be processed first. Through this method, not only the flexibility and response speed of order processing are greatly improved, but also accurate classification and intelligent single insertion judgment are ensured to ensure that urgent orders are processed in a timely and efficient manner, realize the optimal allocation of resources, and improve customer satisfaction.
[0200] In yet another optional embodiment, the numerical quantification result corresponding to the first order includes a first quantification value corresponding to the first standard parameter corresponding to the first order, and a second quantification value corresponding to the second standard parameter corresponding to the first order.
[0201] The judgment module 302 executes order classification on the first order according to the numerical quantification result corresponding to the first order in combination with the preset order classification regulations to obtain an order classification result corresponding to the first order.
[0202] According to the first quantification value corresponding to the first order, in combination with the preset order classification regulations, the upper order classification corresponding to the first order is determined, and the upper order classification corresponding to the first order is used to indicate the order classification of the first order. The order classification includes an urgent order or a regular order.
[0203] According to the second quantification value corresponding to the first order, the lower order classification corresponding to the first order is determined. The lower order classification corresponding to the first order is used to indicate the ranking information of the first order in its corresponding order classification. The ranking information is a ranking numerical value or a ranking level. The higher the ranking numerical value corresponding to the first order, or the higher the ranking level corresponding to the first order, the higher the processing priority of the first order.
[0204] According to the upper order classification and the lower order classification corresponding to the first order, in combination with the order classification regulations, the target order classification of the first order is determined as the order classification result corresponding to the first order.
[0205] As can be seen, in this optional embodiment, by refining the numerical quantification result into a first quantification value and a second quantification value, a double-layer accurate division of order classification is realized. First, according to the first quantification value and the preset order classification regulations, the upper classification (emergency or regular) of the order is quickly determined, and then the second quantification value is used to further analyze the specific ranking of the order in the lower classification, ensuring the meticulousness of the priority evaluation. Through the double-layer classification mechanism, not only the flexibility and accuracy of order processing are improved, but also the ranking information is clear, which provides a quantitative basis for the priority processing of emergency orders, effectively optimizing resource allocation and response speed. Further enhance the intelligent level of order management, and help to improve the accuracy of grasping customer demand, and also help to improve the operation efficiency and customer satisfaction.
[0206] Embodiment four
[0207] Please refer to Figure 5 , Figure 5 is another structure diagram of the B2B order intelligent management device based on reinforcement learning disclosed by the embodiment of the application. As shown in Figure 5 , the B2B order intelligent management device based on reinforcement learning can include:
[0208] a memory 401 storing executable program codes;
[0209] a processor 402 coupled with the memory 401;
[0210] The processor 402 invokes the executable program codes stored in the memory 401 to execute the steps of the B2B order intelligent management method based on reinforcement learning described in the embodiment one or the embodiment two of the application.
[0211] Embodiment five
[0212] The embodiment of the application discloses a computer storage medium, which stores computer instructions. When the computer instructions are invoked, the steps of the B2B order intelligent management method based on reinforcement learning described in the embodiment one or the embodiment two of the application are executed.
[0213] Embodiment six
[0214] The embodiment of the application discloses a computer program product, which includes a non-transitory computer storage medium storing a computer program, and the computer program is operable to make a computer execute the steps of the B2B order intelligent management method based on reinforcement learning described in the embodiment one or the embodiment two.
[0215] The apparatus embodiments described above are only illustrative, wherein the modules illustrated as separate components can or can not be physically separated, and the components illustrated as modules can or can not be physical modules, i.e., can be located in one place or distributed to multiple network modules. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0216] Through the specific description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage, a magnetic tape storage, or any other computer readable medium that can be used to carry or store data.
[0217] Finally, it should be noted that: the B2B order intelligent management method and system based on reinforcement learning disclosed by the embodiments of the application are only the preferred embodiments of the application, and are used to illustrate the technical solutions of the application, but not to limit them; although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that; it can still modify the technical solutions recorded in the foregoing embodiments, or replace some technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the application.
Claims
1. A B2B order intelligent management method based on reinforcement learning, characterized in that, The method comprises: Real-time collection of order data corresponding to a first order queue to be processed, the order data at least including order priority and delivery time; According to the order data, it is judged whether there is a target order in the first order queue that meets the preset order insertion condition, and when the judgment result is yes, the order characteristic information of the target order is extracted from the order data, and each order characteristic information is associated with a second order queue currently processed; According to a preset order insertion influence quantitative evaluation model, the order characteristic information is subjected to quantitative evaluation to obtain a quantitative evaluation result for the target order, the quantitative evaluation result including influence information of the target order inserted into the second order queue; the quantitative evaluation result is used to indicate a specific order insertion matter of inserting the target order into the second order queue; the quantitative evaluation result includes a calculation result corresponding to a preset order insertion influence factor; the order insertion influence factor includes delay days, production resource occupation rate and logistics resource occupation rate; The method further comprises: According to the quantitative evaluation result, a plurality of order insertion strategies for the target order are determined; According to a preset multi-objective optimization reward function, in combination with the quantitative evaluation result, a reward value corresponding to each order insertion strategy is calculated; the multi-objective optimization reward function includes a plurality of optimization parameters; each optimization parameter corresponds to an order requirement of a customer; and different customers have different attention proportions for all order requirements; According to the reward value corresponding to each order insertion strategy, an optimal order insertion strategy is determined from all order insertion strategies; The function formula corresponding to the multi-objective optimization reward function is: wherein, R is a reward value of the multi-objective optimization reward function corresponding to each of the insertion strategies; is an order overall delay parameter; is a conversion function corresponding to the order overall delay parameter, is a parameter weight corresponding to the order overall delay parameter; is a core customer order delay parameter; is a conversion function corresponding to the core customer order delay parameter, is a parameter weight corresponding to the core customer order delay parameter; is an order resource utilization rate; is a conversion function corresponding to the order resource utilization rate; is a parameter weight corresponding to the order resource utilization rate; The order overall delay parameter corresponds to a conversion function The calculation formula is: wherein, is a preset threshold, when the overall delay of the order exceeds , it indicates that the order insertion strategy has too much impact on the overall delivery of the order, and the corresponding reward contribution is 0. Within the range , the smaller the overall delay of the order, the greater the corresponding reward contribution. The core customer order delay parameter corresponds to a conversion function The calculation formula is: wherein, is a threshold value of the core customer order delay parameter, indicating the degree of strictness of the enterprise's requirement for the delivery of core customer orders; The resource utilization corresponds to a conversion function The calculation formula is: 。 2.The B2B order intelligent management method based on reinforcement learning according to claim 1, wherein, According to the preset multi-objective optimization reward function, in combination with the quantitative evaluation result, a reward value corresponding to each order insertion strategy is calculated, comprising: Determine a plurality of optimization parameters corresponding to the preset multi-objective optimization reward function; the plurality of optimization parameters include the order overall delay parameter, the core customer order delay parameter and the order resource utilization rate; Obtain the target order requirement corresponding to the order customer of the target order, and determine the target parameter weight corresponding to each optimization parameter according to the target order requirement, and update the multi-objective optimization reward function according to the target parameter weight corresponding to each optimization parameter; For each order insertion strategy, determine the predicted parameter value associated with each optimization parameter in the order insertion strategy, and input the predicted parameter value associated with each optimization parameter in the order insertion strategy into the multi-objective optimization reward function to calculate the reward value corresponding to the order insertion strategy. 3.The B2B order intelligent management method based on reinforcement learning according to claim 1 or 2, characterized in that, The order characteristic information of the target order includes order basic information, time characteristic information, customer characteristic information and production and logistics characteristic information; the order basic information includes order amount, product type and quantity; the time characteristic information includes delivery time and order placement time; the customer characteristic information includes customer priority and customer historical order situation; the production and logistics characteristic information includes production process complexity and transportation requirements; The order feature information is quantitatively evaluated according to the preset order insertion influence quantitative evaluation model to obtain a quantitative evaluation result for the target order, which comprises: Obtaining order processing information of the second order queue; the order processing information at least comprises processing progress of each second order in the second order queue; The order feature information is input into a preset order insertion influence quantitative evaluation model, and the order insertion influence factor is taken as a benchmark, and the order insertion influence quantitative evaluation model is combined with the order processing information to quantitatively evaluate the order feature information, and a quantitative evaluation result for the order feature information is obtained as the quantitative evaluation result for the target order; The order insertion influence factor is used to indicate the influence information of the processing progress of all the second orders after the target order is inserted into the second order queue. 4.The B2B order intelligent management method based on reinforcement learning according to claim 1 or 2, characterized in that, The order data further comprises goods type, production resource demand, logistics resource demand and customer level; the first order queue comprises a plurality of to-be-processed first orders; The order data is used to determine whether there is a target order meeting preset order insertion conditions in the first order queue, which comprises: The order priority and the delivery time are determined as first benchmark parameters, and the goods type, production resource demand, logistics resource demand and customer level are determined as second benchmark parameters; For each first order, a numerical value quantitative operation is performed on the first benchmark parameters and the second benchmark parameters corresponding to the first order according to a preset numerical value quantitative rule to obtain a numerical value quantitative result corresponding to the first order; According to the numerical value quantitative result corresponding to the first order, an order classification is performed on the first order in combination with a preset order classification regulation to obtain an order classification result corresponding to the first order; the order classification regulation at least comprises an emergency order classification and a regular order classification; a processing priority corresponding to the emergency order classification is higher than a processing priority corresponding to the regular order classification; the order classification result corresponding to the first order is used to indicate that the first order belongs to the emergency order classification or the regular order classification; According to the order classification result corresponding to each first order, it is determined whether there is a target order meeting preset order insertion conditions in all the first orders; The target order meeting the preset order insertion conditions is specifically that the order classification result corresponding to a certain first order indicates that the first order belongs to the emergency order classification. 5.The B2B order intelligent management method based on reinforcement learning according to claim 4, characterized in that, The numerical value quantitative result corresponding to the first order comprises a first quantitative value corresponding to the first benchmark parameter corresponding to the first order and a second quantitative value corresponding to the second benchmark parameter corresponding to the first order; The order classification result corresponding to the first order is obtained by performing an order classification on the first order in combination with a preset order classification regulation according to the numerical value quantitative result corresponding to the first order, which comprises: determining, according to the first quantitative value corresponding to the first order and in combination with a preset order classification regulation, an upper order classification corresponding to the first order, the upper order classification corresponding to the first order being used to indicate an order category of the first order, the order category including an emergency order or a regular order; determining, according to the second quantitative value corresponding to the first order, a lower order classification corresponding to the first order, the lower order classification corresponding to the first order being used to indicate ranking information of the first order in the order category corresponding to the first order, the ranking information being a ranking value or a ranking level; the first order corresponding to the higher ranking value or the first order corresponding to the higher ranking level, the higher processing priority of the first order; determining, according to the upper order classification corresponding to the first order and the lower order classification corresponding to the first order and in combination with the order classification regulation, a target order classification of the first order as an order classification result corresponding to the first order.
6. A B2B order intelligent management system based on reinforcement learning, characterized in that, The system is used to execute the B2B order intelligent management method based on reinforcement learning according to any one of claims 1-5, and the system comprises: an acquisition module configured to acquire order data corresponding to a first order queue to be processed in real time, the order data at least including order priority and delivery time; a judgment module configured to determine whether a target order meeting a preset order insertion condition exists in the first order queue according to the order data; an information extraction module configured to extract order feature information of the target order from the order data when the determination result of the judgment module is yes, and associate each order feature information with a second order queue currently processed; a quantitative evaluation module configured to perform quantitative evaluation on the order feature information according to a preset order insertion influence quantitative evaluation model, to obtain a quantitative evaluation result for the target order, the quantitative evaluation result including influence information of the target order inserted into the second order queue; and the quantitative evaluation result being used to indicate a specific order insertion matter of the target order inserted into the second order queue.
7. A B2B order intelligent management device based on reinforcement learning, characterized in that, The device comprises: a memory storing executable program codes; a processor coupled with the memory; the processor invokes the executable program codes stored in the memory to execute the B2B order intelligent management method based on reinforcement learning according to any one of claims 1-5.
8. A computer storage medium, characterized in that, The computer storage medium stores computer instructions, the computer instructions being invoked to execute the B2B order intelligent management method based on reinforcement learning according to any one of claims 1-5.
Citation Information
Patent Citations
Photovoltaic module internal and external packaging production scheduling optimization method and system
CN118886628A
Reinforcement learning environment model construction strategy and scheduling algorithm for optimizing workshop scheduling
CN119398360A