B2B order intelligent management method and system based on reinforcement learning
Through the order management system based on reinforcement learning, the order data is collected and evaluated in real time and the order plugging strategy is optimized, the flexibility and intelligence of B2B order processing are solved, and the order processing efficiency and customer satisfaction are improved.
Patent Information
- Application Number
- CN202510378028.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The existing B2B order management system lacks flexibility and intelligence, resulting in delays in order processing and improper resource utilization, affecting customer satisfaction and enterprise efficiency.
Using reinforcement learning-based methods, order data is collected in real time, target order feature information is judged and extracted, and in-depth evaluation is conducted through the quantitative evaluation model for order insertion impact, and the order insertion strategy is optimized to improve the flexibility and accuracy of order processing.
It realizes efficient identification and response to temporary/special order needs, optimizes order queue management, ensures timely processing of high-priority orders, reduces operational risks and costs, and improves the intelligence of order management and decision-making accuracy.
Smart Images

Figure CN120494920A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of e-commerce technology, and in particular to a B2B order intelligent management method and system based on reinforcement learning. Background Art
[0002] In the B2B (business-to-business) business environment, order management is one of the core aspects of business operations. With the increasing fierceness of market competition and the diversification of customer needs, companies are facing increasingly complex order processing challenges.
[0003] Currently, most companies' B2B order management systems still utilize traditional processing methods. Under this model, orders are processed sequentially in a fixed order, lacking flexibility. Whenever a new order is added, the entire order processing process often needs to be readjusted. This not only consumes significant time and manpower, but also easily leads to order processing delays, impacting customer satisfaction.
[0004] Existing methods for order insertion decisions primarily rely on manual experience. Workers decide whether and how to insert an order based on their subjective judgment of the order situation. However, this decision-making approach has significant limitations. For one thing, manual judgment is easily influenced by factors such as personal emotions and fatigue, leading to inaccurate decisions. Furthermore, due to the lack of scientific quantitative assessment tools, workers struggle to comprehensively and accurately assess the impact of order insertion on the entire order processing process, making it prone to decision-making errors and resulting in losses for the company. Furthermore, irrational order insertion can negatively impact a company's resource utilization. For example, in the production process, inappropriate order insertion can lead to frequent adjustments to production equipment, increasing production preparation time and costs. In the logistics process, order insertion can lead to the reallocation of goods and changes in transportation routes, reducing logistics efficiency and increasing logistics costs. Summary of the Invention
[0005] The present invention provides a B2B order intelligent management method and system based on reinforcement learning, which can improve the processing flexibility and management intelligence of B2B orders, while improving the accuracy of order insertion decisions and the utilization rate of order resources.
[0006] In order to solve the above technical problems, the first aspect of the present invention discloses a B2B order intelligent management method based on reinforcement learning, the method comprising:
[0007] Collecting order data corresponding to the first order queue to be processed in real time, the order data including at least order priority and delivery time;
[0008] determining, based on the order data, whether there is a target order in the first order queue that meets a preset order insertion condition; and if so, extracting order feature information of the target order from the order data, and associating each order feature information with the second order queue currently being processed;
[0009] According to a preset quantitative evaluation model for the impact of order insertion, a quantitative evaluation is performed on the order feature information to obtain a quantitative evaluation result for the target order, wherein the quantitative evaluation result includes the impact information of inserting the target order into the second order queue; the quantitative evaluation result is used to indicate the specific order insertion matters for inserting the target order into the second order queue.
[0010] As an optional embodiment, in the first aspect of the present invention, the method further comprises:
[0011] Determining multiple order insertion strategies for the target order based on the quantitative evaluation results;
[0012] Calculating a reward value corresponding to each order insertion strategy based on a preset multi-objective optimization reward function and the quantitative evaluation results; the multi-objective optimization reward function includes multiple optimization parameters; each optimization parameter corresponds to a customer's order requirement; and different customers have different attention ratios for all the order requirements;
[0013] According to the reward value corresponding to each order insertion strategy, an optimal order insertion strategy is determined from all the order insertion strategies.
[0014] As an optional embodiment, in the first aspect of the present invention, calculating the reward value corresponding to each of the insertion strategies based on a preset multi-objective optimization reward function and in combination with the quantitative evaluation results includes:
[0015] Determining multiple optimization parameters corresponding to a preset multi-objective optimization reward function; the multiple optimization parameters include an overall order delay parameter, a core customer order delay parameter, and an order resource utilization rate;
[0016] Obtaining a target order requirement corresponding to the order customer of the target order, determining a target parameter weight corresponding to each of the optimization parameters according to the target order requirement, and updating the multi-objective optimization reward function according to the target parameter weight corresponding to each of the optimization parameters;
[0017] For each of the insertion strategies, the predicted parameter values associated with each of the optimization parameters are determined from the insertion strategy, and the predicted parameter values associated with each of the optimization parameters in the insertion strategy are input into the multi-objective optimization reward function to calculate the reward value corresponding to the insertion strategy.
[0018] As an optional implementation, in the first aspect of the present invention, the function formula corresponding to the multi-objective optimization reward function is:
[0019] R=ω1×f1(D overall )+ω2×f2(D key-customer )+ω3×f3(U resource )
[0020] Wherein, R is the reward value of each insertion strategy corresponding to the multi-objective optimization reward function; D overall is the overall delay parameter of the order; f1(D overall ) is the conversion function corresponding to the overall order delay parameter, ω1 is the parameter weight corresponding to the overall order delay parameter; D key-customer is the core customer order delay parameter; f2(D key-customer ) is the conversion function corresponding to the core customer order delay parameter, ω2 is the parameter weight corresponding to the core customer order delay parameter; U resource The order resource utilization rate; f3(U resource ) is the conversion function corresponding to the order resource utilization rate; ω3 is the parameter weight corresponding to the order resource utilization rate.
[0021] As an optional embodiment, in the first aspect of the present invention, the order characteristic information of the target order includes basic order information, time characteristic information, customer characteristic information, and production and logistics characteristic information; the basic order information includes the order amount, product type and quantity; the time characteristic information includes the delivery time and order placement time; the customer characteristic information includes customer priority and customer historical order status; and the production and logistics characteristic information includes the complexity of the production process and transportation requirements;
[0022] The step of performing a quantitative evaluation on the order feature information according to a preset insertion order impact quantitative evaluation model to obtain a quantitative evaluation result for the target order includes:
[0023] Obtaining order processing information of the second order queue; the order processing information at least includes a processing progress of each second order in the second order queue;
[0024] Inputting the order feature information into a preset order insertion impact quantitative evaluation model, and using a preset order insertion impact factor as a benchmark, the order insertion impact quantitative evaluation model combines the order processing information to perform a quantitative evaluation on the order feature information, thereby obtaining a quantitative evaluation result for the order feature information as the quantitative evaluation result for the target order;
[0025] Among them, the quantitative evaluation result includes the calculation result corresponding to the order insertion impact factor; the order insertion impact factor is used to indicate the impact information on the processing progress of all second orders after the target order is inserted into the second order queue; the order insertion impact factor includes the number of delay days, production resource occupancy rate and logistics resource occupancy rate.
[0026] As an optional implementation, in the first aspect of the present invention, the order data further includes goods type, production resource requirements, logistics resource requirements, and customer level; the first order queue includes a plurality of first orders to be processed;
[0027] The determining, based on the order data, whether there is a target order in the first order queue that meets a preset order insertion condition includes:
[0028] Determining the order priority and the delivery time as first benchmark parameters, and determining the cargo type, production resource requirements, logistics resource requirements, and customer level as second benchmark parameters;
[0029] For each first order, performing a numerical quantization operation on the first reference parameter and the second reference parameter corresponding to the first order according to a preset numerical quantization rule to obtain a numerical quantization result corresponding to the first order;
[0030] Based on the numerical quantification result corresponding to the first order and in combination with a preset order grading rule, order grading is performed on the first order to obtain an order grading result corresponding to the first order; the order grading rule includes at least an urgent order grading and a regular order grading; the processing priority corresponding to the urgent order grading is higher than the processing priority of the regular order grading; the order grading result corresponding to the first order is used to indicate whether the first order belongs to the urgent order grading or the regular order grading.
[0031] determining, based on the order grading result corresponding to each of the first orders, whether there is a target order that meets a preset order insertion condition among all the first orders;
[0032] Among them, the target order that meets the preset order insertion condition is specifically an order classification result corresponding to a certain first order, indicating that the first order belongs to the urgent order classification.
[0033] As an optional implementation manner, in the first aspect of the present invention, the numerical quantization result corresponding to the first order includes a first quantization value corresponding to the first reference parameter corresponding to the first order, and a second quantization value corresponding to the second reference parameter corresponding to the first order;
[0034] The step of performing order grading on the first order based on the numerical quantification result corresponding to the first order and in combination with a preset order grading rule to obtain an order grading result corresponding to the first order includes:
[0035] Determining, based on the first quantified value corresponding to the first order and in combination with a preset order grading rule, an upper-level order grade corresponding to the first order, where the upper-level order grade corresponding to the first order is used to indicate an order classification of the first order, where the order classification includes an urgent order or a regular order;
[0036] determining, based on the second quantized value corresponding to the first order, a lower-level order grade corresponding to the first order; the lower-level order grade corresponding to the first order is used to indicate ranking information of the first order in the corresponding order category, the ranking information being a ranking value or a ranking level; the higher the ranking value corresponding to the first order, or the higher the ranking level corresponding to the first order, the higher the processing priority corresponding to the first order;
[0037] According to the upper-level order classification and the lower-level order classification corresponding to the first order, combined with the order classification regulations, the target order classification of the first order is determined as the order classification result corresponding to the first order.
[0038] A second aspect of the present invention discloses a B2B order intelligent management system based on reinforcement learning, the system comprising:
[0039] a collection module, configured to collect order data corresponding to the first order queue to be processed in real time, wherein the order data includes at least order priority and delivery time;
[0040] A judgment module, configured to judge whether there is a target order meeting a preset order insertion condition in the first order queue according to the order data;
[0041] an information extraction module, configured to extract order feature information of the target order from the order data when the judgment result of the judgment module is yes, and associate each piece of order feature information with the second order queue currently being processed;
[0042] A quantitative evaluation module is used to perform a quantitative evaluation on the order feature information according to a preset quantitative evaluation model for the impact of order insertion, and obtain a quantitative evaluation result for the target order, wherein the quantitative evaluation result includes the impact information of inserting the target order into the second order queue; the quantitative evaluation result is used to indicate the specific order insertion matters of inserting the target order into the second order queue.
[0043] As an optional embodiment, in the second aspect of the present invention, the system further includes:
[0044] a determination module, configured to determine a plurality of order insertion strategies for the target order based on the quantitative evaluation result;
[0045] a calculation module, configured to calculate a reward value corresponding to each of the order insertion strategies based on a preset multi-objective optimization reward function and the quantitative evaluation results; the multi-objective optimization reward function includes a plurality of optimization parameters; each optimization parameter corresponds to an order requirement of a customer; and different customers have different attention ratios for all the order requirements;
[0046] The determination module is further configured to determine an optimal order insertion strategy from all the order insertion strategies based on a reward value corresponding to each order insertion strategy.
[0047] As an optional embodiment, in the second aspect of the present invention, the calculation module calculates the reward value corresponding to each of the insertion strategies according to a preset multi-objective optimization reward function and in combination with the quantitative evaluation result, specifically including:
[0048] Determining multiple optimization parameters corresponding to a preset multi-objective optimization reward function; the multiple optimization parameters include an overall order delay parameter, a core customer order delay parameter, and an order resource utilization rate;
[0049] Obtaining a target order requirement corresponding to the order customer of the target order, determining a target parameter weight corresponding to each of the optimization parameters according to the target order requirement, and updating the multi-objective optimization reward function according to the target parameter weight corresponding to each of the optimization parameters;
[0050] For each of the insertion strategies, the predicted parameter values associated with each of the optimization parameters are determined from the insertion strategy, and the predicted parameter values associated with each of the optimization parameters in the insertion strategy are input into the multi-objective optimization reward function to calculate the reward value corresponding to the insertion strategy.
[0051] As an optional implementation, in the second aspect of the present invention, the function formula corresponding to the multi-objective optimization reward function is:
[0052] R=ω1×f1(D overall )+ω2×f2(D key-customer )+ω3×f3(U resource )
[0053] Wherein, R is the reward value of each insertion strategy corresponding to the multi-objective optimization reward function; D overall is the overall delay parameter of the order; f1(D overall ) is the conversion function corresponding to the overall order delay parameter, ω1 is the parameter weight corresponding to the overall order delay parameter; D key-customeris the core customer order delay parameter; f2(D key-customer ) is the conversion function corresponding to the core customer order delay parameter, ω2 is the parameter weight corresponding to the core customer order delay parameter; U resource The order resource utilization rate; f3(U resource ) is the conversion function corresponding to the order resource utilization rate; ω3 is the parameter weight corresponding to the order resource utilization rate.
[0054] As an optional embodiment, in the second aspect of the present invention, the order characteristic information of the target order includes basic order information, time characteristic information, customer characteristic information, and production and logistics characteristic information; the basic order information includes the order amount, product type and quantity; the time characteristic information includes the delivery time and order placement time; the customer characteristic information includes customer priority and customer historical order status; and the production and logistics characteristic information includes the complexity of the production process and transportation requirements;
[0055] The quantitative evaluation module performs a quantitative evaluation on the order feature information according to a preset insertion order impact quantitative evaluation model, and obtains a quantitative evaluation result for the target order in a manner that specifically includes:
[0056] Obtaining order processing information of the second order queue; the order processing information at least includes a processing progress of each second order in the second order queue;
[0057] Inputting the order feature information into a preset order insertion impact quantitative evaluation model, and using a preset order insertion impact factor as a benchmark, the order insertion impact quantitative evaluation model combines the order processing information to perform a quantitative evaluation on the order feature information, thereby obtaining a quantitative evaluation result for the order feature information as the quantitative evaluation result for the target order;
[0058] Among them, the quantitative evaluation result includes the calculation result corresponding to the order insertion impact factor; the order insertion impact factor is used to indicate the impact information on the processing progress of all second orders after the target order is inserted into the second order queue; the order insertion impact factor includes the number of delay days, production resource occupancy rate and logistics resource occupancy rate.
[0059] As an optional implementation, in the second aspect of the present invention, the order data further includes goods type, production resource requirements, logistics resource requirements, and customer level; the first order queue includes a plurality of first orders to be processed;
[0060] The method in which the judgment module judges whether there is a target order meeting the preset order insertion condition in the first order queue according to the order data specifically includes:
[0061] Determining the order priority and the delivery time as first benchmark parameters, and determining the cargo type, production resource requirements, logistics resource requirements, and customer level as second benchmark parameters;
[0062] For each first order, performing a numerical quantization operation on the first reference parameter and the second reference parameter corresponding to the first order according to a preset numerical quantization rule to obtain a numerical quantization result corresponding to the first order;
[0063] Based on the numerical quantification result corresponding to the first order and in combination with a preset order grading rule, order grading is performed on the first order to obtain an order grading result corresponding to the first order; the order grading rule includes at least an urgent order grading and a regular order grading; the processing priority corresponding to the urgent order grading is higher than the processing priority of the regular order grading; the order grading result corresponding to the first order is used to indicate whether the first order belongs to the urgent order grading or the regular order grading.
[0064] determining, based on the order grading result corresponding to each of the first orders, whether there is a target order that meets a preset order insertion condition among all the first orders;
[0065] Among them, the target order that meets the preset order insertion condition is specifically an order classification result corresponding to a certain first order, indicating that the first order belongs to the urgent order classification.
[0066] As an optional implementation manner, in the second aspect of the present invention, the numerical quantization result corresponding to the first order includes a first quantization value corresponding to the first reference parameter corresponding to the first order, and a second quantization value corresponding to the second reference parameter corresponding to the first order;
[0067] The judgment module performs order grading on the first order based on the numerical quantification result corresponding to the first order and in combination with preset order grading rules. The method of obtaining the order grading result corresponding to the first order specifically includes:
[0068] Determining, based on the first quantified value corresponding to the first order and in combination with a preset order grading rule, an upper-level order grade corresponding to the first order, where the upper-level order grade corresponding to the first order is used to indicate an order classification of the first order, where the order classification includes an urgent order or a regular order;
[0069] determining, based on the second quantized value corresponding to the first order, a lower-level order grade corresponding to the first order; the lower-level order grade corresponding to the first order is used to indicate ranking information of the first order in the corresponding order category, the ranking information being a ranking value or a ranking level; the higher the ranking value corresponding to the first order, or the higher the ranking level corresponding to the first order, the higher the processing priority corresponding to the first order;
[0070] According to the upper-level order classification and the lower-level order classification corresponding to the first order, combined with the order classification regulations, the target order classification of the first order is determined as the order classification result corresponding to the first order.
[0071] A third aspect of the present invention discloses another B2B order intelligent management device based on reinforcement learning, the device comprising:
[0072] a memory storing executable program code;
[0073] a processor coupled to the memory;
[0074] The processor calls the executable program code stored in the memory to execute the B2B order intelligent management method based on reinforcement learning disclosed in the first aspect of the present invention.
[0075] The fourth aspect of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute the B2B order intelligent management method based on reinforcement learning disclosed in the first aspect of the present invention.
[0076] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0077] In an embodiment of the present invention, a B2B order intelligent management method based on reinforcement learning is provided, which includes: real-time collection of order data corresponding to a first order queue to be processed, the order data including at least order priority and delivery time; judging whether there is a target order in the first order queue that meets the preset order insertion conditions based on the order data, and when the judgment result is yes, extracting the order feature information of the target order from the order data, and associating each order feature information with the second order queue currently being processed; performing a quantitative evaluation on the order feature information according to a preset order insertion impact quantitative evaluation model to obtain a quantitative evaluation result for the target order, the quantitative evaluation result including the impact information of the target order inserted into the second order queue; the quantitative evaluation result is used to indicate the specific order insertion matters of inserting the target order into the second order queue. It can be seen that by implementing the present invention, by real-time collection of order data (including order priority and delivery time), it is possible to accurately identify the target order that meets the preset order insertion conditions and extract its order feature information, that is, it is possible to respond to customer orders with temporary / special needs in a timely manner, which is conducive to improving the flexibility and efficiency of order processing. Further, using the preset order insertion impact quantitative evaluation model, the order feature information is deeply quantitatively evaluated to obtain detailed impact information of the target order inserted into the second order queue. By scientifically evaluating the impact of order insertion, we optimized order queue management, ensuring the timely processing of high-priority orders. This also reduced operational risks and costs caused by improper order insertion, greatly improving the intelligence of B2B order management and the accuracy of decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0079] Figure 1 This is a flowchart of a B2B order intelligent management method based on reinforcement learning disclosed in an embodiment of the present invention;
[0080] Figure 2 This is a flow chart of another B2B order intelligent management method based on reinforcement learning disclosed in an embodiment of the present invention;
[0081] Figure 3 This is a schematic diagram of the structure of a B2B order intelligent management system based on reinforcement learning disclosed in an embodiment of the present invention;
[0082] Figure 4 This is a schematic diagram of the structure of another B2B order intelligent management system based on reinforcement learning disclosed in an embodiment of the present invention;
[0083] Figure 5 This is a structural diagram of another B2B order intelligent management device based on reinforcement learning disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0084] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0085] The terms "first," "second," and so on, in the description and claims of the present invention and the accompanying drawings are used to distinguish between different items, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or end comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed therein, or may optionally include other steps or elements inherent to such process, method, product, or end.
[0086] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0087] The present invention discloses a B2B order intelligent management method and system based on reinforcement learning. By collecting order data (including order priority and delivery time) in real time, it can accurately identify target orders that meet preset order insertion conditions and extract their order feature information, that is, it can respond to customer orders with temporary / special needs in a timely manner, which is conducive to improving the flexibility and efficiency of order processing. Furthermore, using the preset quantitative evaluation model for the impact of order insertion, an in-depth quantitative evaluation of order feature information is performed to obtain detailed impact information of the target order being inserted into the second order queue. By scientifically evaluating the impact of order insertion, order queue management is optimized, ensuring the timely processing of high-priority orders, while reducing the operational risks and costs caused by improper order insertion, and greatly improving the intelligence level of B2B order management and the accuracy of decision-making. The following are detailed descriptions.
[0088] Example 1
[0089] See also Figure 1 , Figure 1 This is a flow chart of a B2B order intelligent management method based on reinforcement learning disclosed in an embodiment of the present invention. Figure 1 The B2B order intelligent management method based on reinforcement learning described above can be applied to a B2B order intelligent management system based on reinforcement learning, and the embodiments of the present invention do not limit this. Figure 1 As shown, the B2B order intelligent management method based on reinforcement learning may include the following operations:
[0090] 101. Collect order data corresponding to a first order queue to be processed in real time. The order data includes at least order priority and delivery time.
[0091] In an embodiment of the present invention, the first order queue includes one or more first orders to be processed, and the order data includes sub-order data corresponding to each first order. Further, the sub-order data includes a sub-order priority and a sub-delivery time corresponding to each first order.
[0092] 102. Determine, based on the order data, whether there is a target order in the first order queue that meets a preset order insertion condition.
[0093] In the embodiment of the present invention, the target order of the preset order insertion condition may refer to an urgent order that needs to be processed first. It should be noted that the number of the target orders may be one or more, which is not limited in the embodiment of the present invention.
[0094] 103. When the judgment result of step 102 is yes, extract the order feature information of the target order from the order data, and associate each order feature information with the second order queue currently being processed.
[0095] In an embodiment of the present invention, when there are multiple target orders, the order feature information of the target order includes sub-order feature information corresponding to each target order.
[0096] 104. Perform a quantitative evaluation on the order feature information according to a preset order insertion impact quantitative evaluation model to obtain a quantitative evaluation result for the target order, where the quantitative evaluation result includes impact information of the target order being inserted into the second order queue.
[0097] In the embodiment of the present invention, the quantitative evaluation result is used to indicate a specific order insertion matter of inserting the target order into the second order queue.
[0098] It can be seen that implementation Figure 1The described B2B order intelligent management method based on reinforcement learning can accurately identify target orders that meet the preset order insertion conditions and extract their order feature information by collecting order data (including order priority and delivery time) in real time. In other words, it can respond promptly to customer orders with temporary / special needs, which is conducive to improving the flexibility and efficiency of order processing. Furthermore, using the preset quantitative evaluation model for the impact of order insertion, an in-depth quantitative evaluation of order feature information is performed to obtain detailed impact information of the target order being inserted into the second order queue. By scientifically evaluating the impact of order insertion, order queue management is optimized, ensuring the timely processing of high-priority orders, while reducing the operational risks and costs caused by improper order insertion, greatly improving the intelligence level of B2B order management and the accuracy of decision-making.
[0099] In an optional embodiment, the order characteristic information of the target order includes basic order information, time characteristic information, customer characteristic information, and production and logistics characteristic information; the basic order information includes the order amount, product type and quantity; the time characteristic information includes the delivery time and order placement time; the customer characteristic information includes customer priority and customer order history; the production and logistics characteristic information includes the complexity of the production process and transportation requirements;
[0100] The above step 104 performs a quantitative evaluation on the order feature information according to the preset insertion order impact quantitative evaluation model, and obtains the quantitative evaluation result for the target order in the following manner:
[0101] Obtaining order processing information of the second order queue; the order processing information at least includes a processing progress of each second order in the second order queue;
[0102] Input the order feature information into a preset order insertion impact quantitative evaluation model, and based on the preset order insertion impact factor, the order insertion impact quantitative evaluation model combines the order processing information to perform a quantitative evaluation on the order feature information, thereby obtaining a quantitative evaluation result for the order feature information, which is used as the quantitative evaluation result for the target order;
[0103] Among them, the quantitative evaluation results include the calculation results corresponding to the order insertion impact factor; the order insertion impact factor is used to indicate the impact information on the processing progress of all second orders after the target order is inserted into the second order queue; the order insertion impact factor includes the number of delay days, production resource occupancy rate and logistics resource occupancy rate.
[0104] It can be seen that in this optional embodiment, by comprehensively integrating the order feature information of the target order, covering multi-dimensional data such as order basis, time, customer, production and logistics, a rich and detailed data foundation is provided for accurately evaluating the impact of order insertion. Specifically, by introducing a preset quantitative evaluation model for the impact of order insertion, combined with the real-time processing information of the second order queue, and taking the order insertion impact factor as the benchmark, an in-depth quantitative analysis of the order feature information is performed. The processing operation of this in-depth quantitative analysis not only takes into account the attributes of the order itself, but also dynamically integrates the queue processing status to ensure the comprehensiveness and accuracy of the evaluation results. In addition, the quantitative evaluation results are directly related to key indicators such as the number of delay days, production resource occupancy rate and logistics resource occupancy rate, providing decision makers with intuitive and quantitative information on the impact of order insertion. Through this detailed quantitative evaluation operation on order feature information, while improving the intelligence level of order insertion decision-making, it can also effectively balance order priority, processing efficiency and resource utilization, further improving the precision and efficiency of B2B order management.
[0105] In another optional embodiment, the order data further includes goods type, production resource requirements, logistics resource requirements, and customer level; the first order queue includes a plurality of first orders to be processed;
[0106] The method of determining whether there is a target order that meets the preset order insertion condition in the first order queue according to the order data in step 102 specifically includes:
[0107] Determine order priority and delivery time as the first benchmark parameters, and determine product type, production resource requirements, logistics resource requirements, and customer level as the second benchmark parameters;
[0108] For each first order, performing a numerical quantization operation on the first reference parameter and the second reference parameter corresponding to the first order according to a preset numerical quantization rule to obtain a numerical quantization result corresponding to the first order;
[0109] Based on the numerical quantification result corresponding to the first order and in combination with a preset order grading rule, order grading is performed on the first order to obtain an order grading result corresponding to the first order; the order grading rule includes at least an urgent order grading and a regular order grading; the processing priority corresponding to the urgent order grading is higher than the processing priority of the regular order grading; the order grading result corresponding to the first order is used to indicate whether the first order belongs to the urgent order grading or the regular order grading;
[0110] According to the order classification result corresponding to each first order, it is determined whether there is a target order that meets the preset order insertion conditions among all the first orders;
[0111] The target order that meets the preset order insertion conditions is specifically an order classification result corresponding to a first order, indicating that the first order belongs to the urgent order classification.
[0112] It can be seen that in this optional embodiment, by introducing multiple order data such as cargo type, production resource demand, logistics resource demand and customer level, the evaluation dimension for each first order is enriched. In addition, by establishing a dual benchmark parameter system, order priority and delivery time are set as the first benchmark, and the remaining key factors are classified as the second benchmark, and are accurately quantified according to preset numerical quantification rules to ensure that the order characteristics are comprehensive and objectively presented. Furthermore, it is possible to combine the order grading regulations to intelligently divide the urgency of orders, thereby improving the recognition accuracy of urgent orders that need to be handled first. Through this method, not only the flexibility and response speed of order processing are greatly improved, but also through accurate grading and intelligent order insertion judgment, it is ensured that urgent orders are handled in a timely and efficient manner, and optimal allocation of resources is achieved, which is conducive to improving customer satisfaction.
[0113] In yet another optional embodiment, the numerical quantization result corresponding to the first order includes a first quantization value corresponding to a first reference parameter corresponding to the first order, and a second quantization value corresponding to a second reference parameter corresponding to the first order;
[0114] The method of performing order grading on the first order based on the numerical quantification result corresponding to the first order and combining the preset order grading rules to obtain the order grading result corresponding to the first order specifically includes:
[0115] Determining, based on the first quantified value corresponding to the first order and in combination with a preset order grading rule, an upper-level order grade corresponding to the first order, wherein the upper-level order grade corresponding to the first order is used to indicate an order classification of the first order, where the order classification includes an urgent order or a regular order;
[0116] Determining, based on the second quantitative value corresponding to the first order, a lower-level order grade corresponding to the first order; the lower-level order grade corresponding to the first order is used to indicate ranking information of the first order in its corresponding order category, where the ranking information is a ranking value or a ranking level; the higher the ranking value corresponding to the first order, or the higher the ranking level corresponding to the first order, the higher the processing priority corresponding to the first order;
[0117] According to the upper-level order classification and the lower-level order classification corresponding to the first order, combined with the order classification regulations, the target order classification of the first order is determined as the order classification result corresponding to the first order.
[0118] In this optional embodiment, the specific calculation method of the first quantization value is as follows:
[0119] According to the preset grading standards, the order priority is evaluated and quantified to obtain the order priority-quantified value corresponding to the order priority;
[0120] Calculate the difference between the delivery time and the current time, and perform quantization processing on the difference according to a preset time threshold to obtain a delivery time-quantized value corresponding to the delivery time;
[0121] A weighted calculation is performed on the order priority-quantitative value and the delivery time-quantitative value to obtain a first weighted calculation result as the first quantitative value.
[0122] In this optional embodiment, the order priority can be set to 5 for the highest priority, and then to 1 for the lowest priority, thereby intuitively reflecting the differences in order urgency through quantitative data. Regarding interaction time, the difference between the delivery time and the current time can be calculated based on the urgency of the delivery time, and quantified based on a preset time threshold. For example, if the delivery time is within 24 hours of the current time, a higher quantitative value is assigned; if the delivery time is more relaxed, the quantitative value is correspondingly lowered. In this way, the impact of delivery time on order processing can be accurately measured.
[0123] In this optional embodiment, the calculation method for the second quantization value is similar to the calculation method for the first quantization value. The specific calculation method for the second quantization value is as follows:
[0124] Determine the type of goods and characteristic information corresponding to the type of goods;
[0125] Determine the quantity and scarcity corresponding to the need for production resources;
[0126] Determine the logistics attributes corresponding to the cargo transportation mode, transportation distance and transportation difficulty corresponding to the logistics resource demand;
[0127] Determine the historical order status, cooperation years and credit rating corresponding to the customer level;
[0128] According to the preset quantization rules corresponding to each sub-benchmark parameter in the second benchmark parameter, quantitative calculation and weighted calculation are performed on the sub-benchmark parameter to obtain a second weighted calculation result corresponding to all sub-benchmark parameters as the second quantized value; wherein all sub-benchmark parameters of the second benchmark parameter include the type of goods, production resource requirements, logistics resource requirements and customer level.
[0129] In this optional embodiment, the cargo type and characteristics information can be a label corresponding to perishable, fragile, or high-value goods; and the presence of such a label can be assigned a higher quantitative value to reflect its special processing requirements. The production resource requirements corresponding to the demand quantity can include equipment requirements, personnel requirements, and raw material requirements. Furthermore, the ratio of resource requirements to available resources can be used as a quantitative indicator. The higher the ratio, the greater the quantitative value, indicating a greater impact of production resource requirements on order processing.
[0130] In this optional embodiment, for logistics resource requirements, if there are goods that require specialized transportation equipment or long-distance transport, a higher quantitative value may be assigned to reflect the impact of logistics resource requirements on order processing. Regarding customer tiers, quantitative values can be assigned to customer tiers based on the various customer information mentioned above. For example, high-tier customers may be assigned higher quantitative values to reflect their importance in the company's customer service.
[0131] It can be seen that in this optional embodiment, by refining the numerical quantification results into a first quantitative value and a second quantitative value, a two-layer precise division of order grading is achieved. And, first, based on the first quantitative value and the preset order grading regulations, the upper classification of the order (urgent or regular) is quickly determined, and then the second quantitative value is used to further analyze the specific ranking of the order in the lower classification to ensure that the priority assessment is meticulous. Through this set two-layer grading mechanism, not only the flexibility and accuracy of order processing are improved, but also through clear ranking information, a quantitative basis is provided for the priority processing of urgent orders, effectively optimizing resource allocation and response speed. It further enhances the intelligence level of order management, and is conducive to improving the accuracy of grasping customer needs, as well as improving operational efficiency and customer satisfaction.
[0132] Example 2
[0133] See also Figure 2 , Figure 2 This is a flow chart of another B2B order intelligent management method based on reinforcement learning disclosed in an embodiment of the present invention. Figure 2 The B2B order intelligent management method based on reinforcement learning described above can be applied to a B2B order intelligent management system based on reinforcement learning, and the embodiments of the present invention do not limit this. Figure 2 As shown, the B2B order intelligent management method based on reinforcement learning may include the following operations:
[0134] 201. Collect order data corresponding to a first order queue to be processed in real time, where the order data at least includes order priority and delivery time.
[0135] 202. Determine, based on the order data, whether there is a target order in the first order queue that meets a preset order insertion condition.
[0136] 203. When the judgment result of step 202 is yes, extract the order feature information of the target order from the order data, and associate each order feature information with the second order queue currently being processed.
[0137] 204. Perform a quantitative evaluation on the order feature information according to a preset order insertion impact quantitative evaluation model to obtain a quantitative evaluation result for the target order, where the quantitative evaluation result includes impact information of the target order being inserted into the second order queue.
[0138] In the embodiment of the present invention, for other descriptions of steps 201 to 204, please refer to other specific descriptions of steps 101 to 104 in the first embodiment, which will not be repeated in the embodiment of the present invention.
[0139] 205. Based on the quantitative evaluation results, determine multiple order insertion strategies for the target order.
[0140] 206. Based on the preset multi-objective optimization reward function and the quantitative evaluation results, calculate the reward value corresponding to each insertion strategy.
[0141] In an embodiment of the present invention, the multi-objective optimization reward function includes multiple optimization parameters; each optimization parameter corresponds to an order requirement of a customer; and different customers pay different attention to the proportion of all order requirements.
[0142] 207. Determine the optimal order insertion strategy from all order insertion strategies based on the reward value corresponding to each order insertion strategy.
[0143] It can be seen that implementation Figure 2 The described B2B order intelligent management method based on reinforcement learning generates multiple order insertion strategies based on the quantitative evaluation results, and can introduce a preset multi-objective optimization reward function for accurate evaluation. Among them, the multi-objective optimization reward function integrates multiple optimization parameters, each parameter accurately corresponds to the customer's order requirements, and flexibly adapts to the differentiated attention ratios of different customers to order requirements. Subsequently, the reward value of each order insertion strategy can be calculated based on this, so as to intelligently screen out the optimal order insertion strategy. Through the screening mechanism of the optimal order insertion strategy, on the basis of significantly enhancing the flexibility and adaptability of order processing, through the multi-objective optimization mechanism, it is ensured that the order processing plan can accurately match the diverse needs of customers, effectively improving customers' satisfaction with using the B2B platform for transactions, while also optimizing resource allocation, reducing operating costs, and improving the intelligence and refinement of B2B order management.
[0144] In an optional embodiment, the method of calculating the reward value corresponding to each insertion strategy in step 206 according to the preset multi-objective optimization reward function and the quantitative evaluation results specifically includes:
[0145] Determining multiple optimization parameters corresponding to a preset multi-objective optimization reward function; the multiple optimization parameters include an overall order delay parameter, a core customer order delay parameter, and an order resource utilization rate;
[0146] Obtaining a target order requirement corresponding to an order customer of a target order, determining a target parameter weight corresponding to each optimization parameter according to the target order requirement, and updating a multi-objective optimization reward function according to the target parameter weight corresponding to each optimization parameter;
[0147] For each insertion strategy, the predicted parameter value associated with each optimization parameter in the insertion strategy is determined, and the predicted parameter value associated with each optimization parameter in the insertion strategy is input into the multi-objective optimization reward function to calculate the reward value corresponding to the insertion strategy.
[0148] In this optional embodiment, the function formula corresponding to the multi-objective optimization reward function is:
[0149] R=ω1×f1(D overall )+ω2×f2(D key-customer )+ω3×f3(U resource )
[0150] Among them, R is the reward value of each insertion strategy corresponding to the multi-objective optimization reward function; D overall is the overall delay parameter of the order; f1(D overall ) is the conversion function corresponding to the overall order delay parameter, ω1 is the parameter weight corresponding to the overall order delay parameter; D key-customer is the core customer order delay parameter; f2(D key-customer ) is the conversion function corresponding to the core customer order delay parameter, ω2 is the parameter weight corresponding to the core customer order delay parameter; U resource Order resource utilization; f3(U resource ) is the conversion function corresponding to the order resource utilization rate; ω3 is the parameter weight corresponding to the order resource utilization rate.
[0151] Furthermore, the calculation formula of the conversion function f1 corresponding to the overall delay parameter of the order is:
[0152]
[0153] Among them, T1 is a preset threshold. When the overall order delay exceeds T1, it is considered that the order insertion strategy has too great an impact on the overall order delivery, and the reward contribution is 0. In the range 0≤D overallWithin ≤T1, the smaller the delay, the greater the reward contribution.
[0154] The calculation formula of the conversion function f2 corresponding to the core customer order delay parameter is:
[0155]
[0156] Among them, T2 is the threshold of the core customer order delay parameter, which reflects the company's strict requirements on the delivery of core customer orders.
[0157] The calculation formula of the conversion function f3 corresponding to the resource utilization is:
[0158] f3(U resource )=U resource
[0159] In this optional embodiment, specifically, assume that there are three order insertion strategies i1, i2, and i3, and set the above parameter weights to ω1 = 0.5, ω2 = 0.3, and ω3 = 0.2. At the same time, the thresholds T1 = 10 days and T2 = 5 days. The overall order delay parameters for the three order insertion strategies are 8 days, 12 days, and 5 days, respectively; the core customer order delay parameters for the three order insertion strategies are 3 days, 2 days, and 6 days, respectively; the resource utilization rates for the three order insertion strategies are 0.8, 0.9, and 0.7, respectively. Furthermore, the reward value for each order insertion strategy is calculated as follows:
[0160] Take the insertion order strategy i1 as an example:
[0161]
[0162] f3(U resource )=0.8
[0163] R1=0.5×0.2+0.3×0.4+0.2×0.8=0.38
[0164] Similarly, substituting the above data sequentially, we can calculate R2 = 0.36 for insertion strategy i2 and R3 = 0.39 for insertion strategy i3. By comparing the reward values of each insertion strategy, we can find that R3 has the largest value, and then determine that i3 is the optimal insertion strategy.
[0165] It can be seen that in this optional embodiment, the calculation method of the multi-objective optimization reward function is refined, and by clearly including key optimization parameters such as the overall order delay parameter, the core customer order delay parameter and the order resource utilization rate, the core concerns of the enterprise operation are accurately connected. This makes it possible to dynamically obtain the specific requirements of the target order customers, and flexibly adjust the target weights of each optimization parameter accordingly to achieve personalized customization of the reward function. For each order insertion strategy, it is possible to deeply analyze and predict the parameter values associated with each optimization parameter, substitute them into the updated reward function for accurate calculation, and obtain a quantitative reward value. Therefore, the evaluation accuracy and adaptability of the order insertion strategy are significantly improved, ensuring that the order insertion strategy not only meets the diverse needs of customers, but also optimizes order processing efficiency, resource utilization and customer satisfaction at the global level.
[0166] Example 3
[0167] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of a B2B order intelligent management system based on reinforcement learning disclosed in an embodiment of the present invention. The B2B order intelligent management system based on reinforcement learning can be used in B2B transaction scenarios, which is not limited in the embodiment of the present invention. Figure 3 As shown, the B2B order intelligent management system based on reinforcement learning may include a collection module 301, a judgment module 302, an information extraction module 303, and a quantitative evaluation module 304, wherein:
[0168] The collection module 301 is used to collect order data corresponding to the first order queue to be processed in real time, and the order data at least includes order priority and delivery time.
[0169] The judgment module 302 is used to judge whether there is a target order that meets the preset order insertion conditions in the first order queue according to the order data.
[0170] The information extraction module 303 is configured to extract order feature information of the target order from the order data when the judgment result of the judgment module 302 is yes, and associate each order feature information with the second order queue currently being processed.
[0171] The quantitative evaluation module 304 is used to perform a quantitative evaluation on the order feature information according to a preset quantitative evaluation model of the impact of inserting orders, and obtain a quantitative evaluation result for the target order. The quantitative evaluation result includes the impact information of inserting the target order into the second order queue; the quantitative evaluation result is used to indicate the specific insertion matters of inserting the target order into the second order queue.
[0172] It can be seen that implementation Figure 3The described B2B order intelligent management system based on reinforcement learning can accurately identify target orders that meet the preset order insertion conditions and extract their order feature information by collecting order data (including order priority and delivery time) in real time. In other words, it can respond promptly to customer orders with temporary / special needs, which is conducive to improving the flexibility and efficiency of order processing. Furthermore, using the preset quantitative evaluation model for the impact of order insertion, an in-depth quantitative evaluation of order feature information is performed to obtain detailed impact information of the target order being inserted into the second order queue. By scientifically evaluating the impact of order insertion, order queue management is optimized, ensuring the timely processing of high-priority orders, while reducing the operational risks and costs caused by improper order insertion, greatly improving the intelligence level of B2B order management and the accuracy of decision-making.
[0173] In an alternative embodiment, see Figure 4 , Figure 4 This is a schematic diagram of another B2B order intelligent management system based on reinforcement learning disclosed in an embodiment of the present invention. Figure 4 As shown, the system further includes a determination module 305 and a calculation module 306, wherein:
[0174] A determination module 305 is configured to determine multiple order insertion strategies for the target order based on the quantitative evaluation results;
[0175] Calculation module 306 is configured to calculate the reward value corresponding to each order insertion strategy based on a preset multi-objective optimization reward function and the quantitative evaluation results. The multi-objective optimization reward function includes multiple optimization parameters. Each optimization parameter corresponds to a customer's order requirement. Different customers may pay different attention to all order requirements.
[0176] The determination module 305 is further configured to determine the optimal order insertion strategy from all order insertion strategies according to the reward value corresponding to each order insertion strategy.
[0177] It can be seen that in this optional embodiment, multiple order insertion strategies are generated based on the quantitative evaluation results, and a preset multi-objective optimization reward function can be introduced for accurate evaluation. Among them, the multi-objective optimization reward function integrates multiple optimization parameters, each parameter accurately corresponds to the customer order requirements, and flexibly adapts to the differentiated attention ratios of different customers to order requirements. Subsequently, the reward value of each order insertion strategy can be calculated based on this, so as to intelligently screen out the optimal order insertion strategy. Through the screening mechanism of the optimal order insertion strategy, on the basis of significantly enhancing the flexibility and adaptability of order processing, through the multi-objective optimization mechanism, it is ensured that the order processing plan can accurately match the diverse needs of customers, effectively improving the customer satisfaction with using the B2B platform for transactions, while also optimizing resource allocation, reducing operating costs, and improving the intelligence and refinement of B2B order management.
[0178] In another optional embodiment, the calculation module 306 calculates the reward value corresponding to each insertion strategy according to the preset multi-objective optimization reward function and the quantitative evaluation results, specifically including:
[0179] Determining multiple optimization parameters corresponding to a preset multi-objective optimization reward function; the multiple optimization parameters include an overall order delay parameter, a core customer order delay parameter, and an order resource utilization rate;
[0180] Obtaining a target order requirement corresponding to an order customer of a target order, determining a target parameter weight corresponding to each optimization parameter according to the target order requirement, and updating a multi-objective optimization reward function according to the target parameter weight corresponding to each optimization parameter;
[0181] For each insertion strategy, the predicted parameter value associated with each optimization parameter in the insertion strategy is determined, and the predicted parameter value associated with each optimization parameter in the insertion strategy is input into the multi-objective optimization reward function to calculate the reward value corresponding to the insertion strategy.
[0182] In this optional embodiment, the function formula corresponding to the multi-objective optimization reward function is:
[0183] R=ω1×f1(D overall )+ω2×f2(D key-customer )+ω3×f3(U resource )
[0184] Among them, R is the reward value of each insertion strategy corresponding to the multi-objective optimization reward function; D overall is the overall delay parameter of the order; f1(D overall ) is the conversion function corresponding to the overall order delay parameter, ω1 is the parameter weight corresponding to the overall order delay parameter; D key-customer is the core customer order delay parameter; f2(D key-customer ) is the conversion function corresponding to the core customer order delay parameter, ω2 is the parameter weight corresponding to the core customer order delay parameter; U resource Order resource utilization; f3(U resource ) is the conversion function corresponding to the order resource utilization rate; ω3 is the parameter weight corresponding to the order resource utilization rate.
[0185] It can be seen that in this optional embodiment, the calculation method of the multi-objective optimization reward function is refined, and by clearly including key optimization parameters such as the overall order delay parameter, the core customer order delay parameter and the order resource utilization rate, the core concerns of the enterprise operation are accurately connected. This makes it possible to dynamically obtain the specific requirements of the target order customers, and flexibly adjust the target weights of each optimization parameter accordingly to achieve personalized customization of the reward function. For each order insertion strategy, it is possible to deeply analyze and predict the parameter values associated with each optimization parameter, substitute them into the updated reward function for accurate calculation, and obtain a quantitative reward value. Therefore, the evaluation accuracy and adaptability of the order insertion strategy are significantly improved, ensuring that the order insertion strategy not only meets the diverse needs of customers, but also optimizes order processing efficiency, resource utilization and customer satisfaction at the global level.
[0186] In another optional embodiment, the order characteristic information of the target order includes basic order information, time characteristic information, customer characteristic information, and production and logistics characteristic information; the basic order information includes the order amount, product type and quantity; the time characteristic information includes the delivery time and order placement time; the customer characteristic information includes customer priority and customer order history; and the production and logistics characteristic information includes the complexity of the production process and transportation requirements.
[0187] The quantitative evaluation module 304 performs a quantitative evaluation on the order feature information according to a preset insertion order impact quantitative evaluation model, and obtains a quantitative evaluation result for the target order in the following manner:
[0188] Obtaining order processing information of the second order queue; the order processing information at least includes a processing progress of each second order in the second order queue;
[0189] Input the order feature information into a preset order insertion impact quantitative evaluation model, and based on the preset order insertion impact factor, the order insertion impact quantitative evaluation model combines the order processing information to perform a quantitative evaluation on the order feature information, thereby obtaining a quantitative evaluation result for the order feature information, which is used as the quantitative evaluation result for the target order;
[0190] Among them, the quantitative evaluation results include the calculation results corresponding to the order insertion impact factor; the order insertion impact factor is used to indicate the impact information on the processing progress of all second orders after the target order is inserted into the second order queue; the order insertion impact factor includes the number of delay days, production resource occupancy rate and logistics resource occupancy rate.
[0191] It can be seen that in this optional embodiment, by comprehensively integrating the order feature information of the target order, covering multi-dimensional data such as order basis, time, customer, production and logistics, a rich and detailed data foundation is provided for accurately evaluating the impact of order insertion. Specifically, by introducing a preset quantitative evaluation model for the impact of order insertion, combined with the real-time processing information of the second order queue, and taking the order insertion impact factor as the benchmark, an in-depth quantitative analysis of the order feature information is performed. The processing operation of this in-depth quantitative analysis not only takes into account the attributes of the order itself, but also dynamically integrates the queue processing status to ensure the comprehensiveness and accuracy of the evaluation results. In addition, the quantitative evaluation results are directly related to key indicators such as the number of delay days, production resource occupancy rate and logistics resource occupancy rate, providing decision makers with intuitive and quantitative information on the impact of order insertion. Through this detailed quantitative evaluation operation on order feature information, while improving the intelligence level of order insertion decision-making, it can also effectively balance order priority, processing efficiency and resource utilization, further improving the precision and efficiency of B2B order management.
[0192] In another optional embodiment, the order data further includes goods type, production resource requirements, logistics resource requirements, and customer level; the first order queue includes a plurality of first orders to be processed;
[0193] The method in which the judgment module 302 judges whether there is a target order meeting the preset order insertion condition in the first order queue according to the order data specifically includes:
[0194] Determine order priority and delivery time as the first benchmark parameters, and determine product type, production resource requirements, logistics resource requirements, and customer level as the second benchmark parameters;
[0195] For each first order, performing a numerical quantization operation on the first reference parameter and the second reference parameter corresponding to the first order according to a preset numerical quantization rule to obtain a numerical quantization result corresponding to the first order;
[0196] Based on the numerical quantification result corresponding to the first order and in combination with a preset order grading rule, order grading is performed on the first order to obtain an order grading result corresponding to the first order; the order grading rule includes at least an urgent order grading and a regular order grading; the processing priority corresponding to the urgent order grading is higher than the processing priority of the regular order grading; the order grading result corresponding to the first order is used to indicate whether the first order belongs to the urgent order grading or the regular order grading;
[0197] According to the order classification result corresponding to each first order, it is determined whether there is a target order that meets the preset order insertion conditions among all the first orders;
[0198] The target order that meets the preset order insertion conditions is specifically an order classification result corresponding to a first order, indicating that the first order belongs to the urgent order classification.
[0199] It can be seen that in this optional embodiment, by introducing multiple order data such as cargo type, production resource demand, logistics resource demand and customer level, the evaluation dimension for each first order is enriched. In addition, by establishing a dual benchmark parameter system, order priority and delivery time are set as the first benchmark, and the remaining key factors are classified as the second benchmark, and are accurately quantified according to preset numerical quantification rules to ensure that the order characteristics are comprehensive and objectively presented. Furthermore, it is possible to combine the order grading regulations to intelligently divide the urgency of orders, thereby improving the recognition accuracy of urgent orders that need to be handled first. Through this method, not only the flexibility and response speed of order processing are greatly improved, but also through accurate grading and intelligent order insertion judgment, it is ensured that urgent orders are handled in a timely and efficient manner, and optimal allocation of resources is achieved, which is conducive to improving customer satisfaction.
[0200] In yet another optional embodiment, the numerical quantization result corresponding to the first order includes a first quantization value corresponding to a first reference parameter corresponding to the first order, and a second quantization value corresponding to a second reference parameter corresponding to the first order;
[0201] The judgment module 302 performs order grading on the first order based on the numerical quantification result corresponding to the first order and in combination with a preset order grading rule. The method of obtaining the order grading result corresponding to the first order specifically includes:
[0202] Determining, based on the first quantified value corresponding to the first order and in combination with a preset order grading rule, an upper-level order grade corresponding to the first order, wherein the upper-level order grade corresponding to the first order is used to indicate an order classification of the first order, where the order classification includes an urgent order or a regular order;
[0203] Determining, based on the second quantitative value corresponding to the first order, a lower-level order grade corresponding to the first order; the lower-level order grade corresponding to the first order is used to indicate ranking information of the first order in its corresponding order category, where the ranking information is a ranking value or a ranking level; the higher the ranking value corresponding to the first order, or the higher the ranking level corresponding to the first order, the higher the processing priority corresponding to the first order;
[0204] According to the upper-level order classification and the lower-level order classification corresponding to the first order, combined with the order classification regulations, the target order classification of the first order is determined as the order classification result corresponding to the first order.
[0205] It can be seen that in this optional embodiment, by refining the numerical quantification results into a first quantitative value and a second quantitative value, a two-layer precise division of order grading is achieved. And, first, based on the first quantitative value and the preset order grading regulations, the upper classification of the order (urgent or regular) is quickly determined, and then the second quantitative value is used to further analyze the specific ranking of the order in the lower classification to ensure that the priority assessment is meticulous. Through this set two-layer grading mechanism, not only the flexibility and accuracy of order processing are improved, but also through clear ranking information, a quantitative basis is provided for the priority processing of urgent orders, effectively optimizing resource allocation and response speed. It further enhances the intelligence level of order management, and is conducive to improving the accuracy of grasping customer needs, as well as improving operational efficiency and customer satisfaction.
[0206] Example 4
[0207] See also Figure 5 , Figure 5 This is a structural diagram of another B2B order intelligent management device based on reinforcement learning disclosed in an embodiment of the present invention. Figure 5 As shown, the B2B order intelligent management device based on reinforcement learning may include:
[0208] A memory 401 storing executable program code;
[0209] a processor 402 coupled to the memory 401;
[0210] The processor 402 calls the executable program code stored in the memory 401 to execute the steps of the B2B order intelligent management method based on reinforcement learning described in the first embodiment of the present invention or the second embodiment of the present invention.
[0211] Example 5
[0212] An embodiment of the present invention discloses a computer storage medium storing computer instructions. When the computer instructions are called, they are used to execute the steps of the reinforcement learning-based B2B order intelligent management method described in Embodiment 1 or Embodiment 2 of the present invention.
[0213] Example 6
[0214] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps of the reinforcement learning-based B2B order intelligent management method described in Example 1 or Example 2.
[0215] The device embodiments described above are merely illustrative. Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0216] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, or of course, by means of hardware. Based on this understanding, the above technical solution, in essence, or the portion that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0217] Finally, it should be noted that the B2B order intelligent management method and system based on reinforcement learning disclosed in the embodiments of the present invention are only preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features therein may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A B2B order intelligent management method based on reinforcement learning, characterized in that: The method comprises: Collecting order data corresponding to the first order queue to be processed in real time, the order data including at least order priority and delivery time; determining, based on the order data, whether there is a target order in the first order queue that meets a preset order insertion condition; and if so, extracting order feature information of the target order from the order data, and associating each order feature information with the second order queue currently being processed; According to a preset quantitative evaluation model for the impact of order insertion, a quantitative evaluation is performed on the order feature information to obtain a quantitative evaluation result for the target order, wherein the quantitative evaluation result includes the impact information of inserting the target order into the second order queue; the quantitative evaluation result is used to indicate the specific order insertion matters for inserting the target order into the second order queue.
2. The B2B order intelligent management method based on reinforcement learning according to claim 1 is characterized in that: The method further comprises: Determining multiple order insertion strategies for the target order based on the quantitative evaluation results; Calculating a reward value corresponding to each order insertion strategy based on a preset multi-objective optimization reward function and the quantitative evaluation results; the multi-objective optimization reward function includes multiple optimization parameters; each optimization parameter corresponds to a customer's order requirement; and different customers have different attention ratios for all the order requirements; According to the reward value corresponding to each order insertion strategy, an optimal order insertion strategy is determined from all the order insertion strategies.
3. The B2B order intelligent management method based on reinforcement learning according to claim 2 is characterized in that: The step of calculating the reward value corresponding to each of the insertion strategies according to the preset multi-objective optimization reward function and in combination with the quantitative evaluation result includes: Determining multiple optimization parameters corresponding to a preset multi-objective optimization reward function; the multiple optimization parameters include an overall order delay parameter, a core customer order delay parameter, and an order resource utilization rate; Obtaining a target order requirement corresponding to the order customer of the target order, determining a target parameter weight corresponding to each of the optimization parameters according to the target order requirement, and updating the multi-objective optimization reward function according to the target parameter weight corresponding to each of the optimization parameters; For each of the insertion strategies, the predicted parameter values associated with each of the optimization parameters are determined from the insertion strategy, and the predicted parameter values associated with each of the optimization parameters in the insertion strategy are input into the multi-objective optimization reward function to calculate the reward value corresponding to the insertion strategy.
4. The B2B order intelligent management method based on reinforcement learning according to claim 3 is characterized in that: The function formula corresponding to the multi-objective optimization reward function is: R=ω1×f1(D overall )+ω2×f2(D key-customer )+ω3×f3(U resource ) Wherein, R is the reward value of each insertion strategy corresponding to the multi-objective optimization reward function; D overall is the overall delay parameter of the order; f1(D overall ) is the conversion function corresponding to the overall order delay parameter, ω1 is the parameter weight corresponding to the overall order delay parameter; D key-customer is the core customer order delay parameter; f2(D key-customer ) is the conversion function corresponding to the core customer order delay parameter, ω2 is the parameter weight corresponding to the core customer order delay parameter; U resource The order resource utilization rate; f3(U resource ) is the conversion function corresponding to the order resource utilization rate; ω3 is the parameter weight corresponding to the order resource utilization rate.
5. The B2B order intelligent management method based on reinforcement learning according to any one of claims 1 to 4, characterized in that: The target order's order characteristic information includes basic order information, time characteristic information, customer characteristic information, and production and logistics characteristic information; the basic order information includes the order amount, product type, and quantity; the time characteristic information includes the delivery time and order placement time; the customer characteristic information includes customer priority and customer order history; and the production and logistics characteristic information includes the complexity of the production process and transportation requirements; The step of performing a quantitative evaluation on the order feature information according to a preset insertion order impact quantitative evaluation model to obtain a quantitative evaluation result for the target order includes: Obtaining order processing information of the second order queue; the order processing information at least includes a processing progress of each second order in the second order queue; Inputting the order feature information into a preset order insertion impact quantitative evaluation model, and using a preset order insertion impact factor as a benchmark, the order insertion impact quantitative evaluation model combines the order processing information to perform a quantitative evaluation on the order feature information, thereby obtaining a quantitative evaluation result for the order feature information as the quantitative evaluation result for the target order; Among them, the quantitative evaluation result includes the calculation result corresponding to the order insertion impact factor; the order insertion impact factor is used to indicate the impact information on the processing progress of all second orders after the target order is inserted into the second order queue; the order insertion impact factor includes the number of delay days, production resource occupancy rate and logistics resource occupancy rate.
6. The B2B order intelligent management method based on reinforcement learning according to any one of claims 1 to 4, characterized in that: The order data further includes goods type, production resource requirements, logistics resource requirements, and customer level; the first order queue includes a plurality of first orders to be processed; The determining, based on the order data, whether there is a target order in the first order queue that meets a preset order insertion condition includes: Determining the order priority and the delivery time as first benchmark parameters, and determining the cargo type, production resource requirements, logistics resource requirements, and customer level as second benchmark parameters; For each first order, performing a numerical quantization operation on the first reference parameter and the second reference parameter corresponding to the first order according to a preset numerical quantization rule to obtain a numerical quantization result corresponding to the first order; Based on the numerical quantification result corresponding to the first order and in combination with a preset order grading rule, order grading is performed on the first order to obtain an order grading result corresponding to the first order; the order grading rule includes at least an urgent order grading and a regular order grading; the processing priority corresponding to the urgent order grading is higher than the processing priority of the regular order grading; the order grading result corresponding to the first order is used to indicate whether the first order belongs to the urgent order grading or the regular order grading. determining, based on the order grading result corresponding to each of the first orders, whether there is a target order that meets a preset order insertion condition among all the first orders; Among them, the target order that meets the preset order insertion condition is specifically an order classification result corresponding to a certain first order, indicating that the first order belongs to the urgent order classification.
7. The B2B order intelligent management method based on reinforcement learning according to claim 6 is characterized in that: The numerical quantization result corresponding to the first order includes a first quantization value corresponding to the first reference parameter corresponding to the first order, and a second quantization value corresponding to the second reference parameter corresponding to the first order; The step of performing order grading on the first order based on the numerical quantification result corresponding to the first order and in combination with a preset order grading rule to obtain an order grading result corresponding to the first order includes: Determining, based on the first quantified value corresponding to the first order and in combination with a preset order grading rule, an upper-level order grade corresponding to the first order, where the upper-level order grade corresponding to the first order is used to indicate an order classification of the first order, where the order classification includes an urgent order or a regular order; determining, based on the second quantized value corresponding to the first order, a lower-level order grade corresponding to the first order; the lower-level order grade corresponding to the first order is used to indicate ranking information of the first order in the corresponding order category, the ranking information being a ranking value or a ranking level; the higher the ranking value corresponding to the first order, or the higher the ranking level corresponding to the first order, the higher the processing priority corresponding to the first order; According to the upper-level order classification and the lower-level order classification corresponding to the first order, combined with the order classification regulations, the target order classification of the first order is determined as the order classification result corresponding to the first order.
8. A B2B order intelligent management system based on reinforcement learning, characterized in that: The system comprises: a collection module, configured to collect order data corresponding to the first order queue to be processed in real time, wherein the order data includes at least order priority and delivery time; A judgment module, configured to judge whether there is a target order meeting a preset order insertion condition in the first order queue according to the order data; an information extraction module, configured to extract order feature information of the target order from the order data when the judgment result of the judgment module is yes, and associate each piece of order feature information with the second order queue currently being processed; A quantitative evaluation module is used to perform a quantitative evaluation on the order feature information according to a preset quantitative evaluation model for the impact of order insertion, and obtain a quantitative evaluation result for the target order, wherein the quantitative evaluation result includes the impact information of inserting the target order into the second order queue; the quantitative evaluation result is used to indicate the specific order insertion matters of inserting the target order into the second order queue.
9. A B2B order intelligent management device based on reinforcement learning, characterized in that: The device comprises: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the B2B order intelligent management method based on reinforcement learning according to any one of claims 1 to 7.
10. A computer storage medium, characterized in that The computer storage medium stores computer instructions, which, when called, are used to execute the B2B order intelligent management method based on reinforcement learning according to any one of claims 1 to 7.
Citation Information
Patent Citations
Cloud order dynamic receiving and scheduling method based on deep reinforcement learning
CN113935586A
Flow shop new order insertion optimization method based on deep reinforcement learning
CN115793583A
Photovoltaic module internal and external packaging production scheduling optimization method and system
CN118886628A
Supply chain scheduling optimization method and device for emergency orders
CN118966696A
Reinforcement learning environment model construction strategy and scheduling algorithm for optimizing workshop scheduling
CN119398360A