Multi-dimensional reward-fused stock matching method, electronic equipment and storage medium

By integrating a multi-dimensional reward-based driver-cargo matching method and optimizing the driver-cargo matching degree calculation model using a multi-dimensional reward function, the problem of low matching degree between cargo and drivers in existing technologies is solved, achieving higher matching degree and higher cargo delivery efficiency.

CN120806792AInactive Publication Date: 2025-10-17JIANGSU MANYUN LOGISTICS INFORMATION CO LTD

Patent Information

Application Number
CN202511301326.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing driver-cargo matching methods result in low matching rates between cargo and drivers, leading to problems such as high vacancy rates, low satisfaction, and driver turnover. They also lack the ability to model and dynamically respond to the specific business objectives of the logistics industry.

Method used

A driver-cargo matching method integrating multi-dimensional rewards is adopted. By acquiring various feature information of cargo sources and drivers, the matching degree score is calculated using a trained driver-cargo matching degree calculation model. A multi-dimensional reward function is introduced for optimization to generate a matching degree priority sequence, supporting end-to-end gradient backpropagation and personalized rearrangement.

Benefits of technology

This improved the matching rate between cargo and drivers, reduced drivers' empty-running rate, and enhanced the platform's cargo delivery efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806792A_ABST
    Figure CN120806792A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of goods source matching, in particular to a multi-dimensional award fused stock matching method, electronic equipment and a storage medium. The shipping matching method comprises the following steps: firstly, acquiring cargo source feature information and driver feature information corresponding to each cargo source, and then inputting the cargo source feature information and the driver feature information into a trained shipping matching degree calculation model to obtain a matching degree score of the cargo source and a target driver; and sorting the plurality of matching degree scores from high to low to obtain a matching degree priority sequence. According to the shipping matching method provided by the invention, various kinds of driver feature information are fused when cargo source matching is carried out on the driver, so that the matching degree priority sequence with the matching degree scores from high to low is finally obtained, and the driver selects the cargo source with the high priority according to the matching degree priority sequence, so that the driver can select the cargo source with the high priority according to the matching degree priority sequence. Therefore, the matching degree between the goods source and the driver is higher, the empty driving rate of the driver is reduced, and the goods pushing efficiency and the goods distribution efficiency of the platform are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cargo matching, and in particular to a driver-cargo matching method fusing multi-dimensional rewards, an electronic device and a storage medium. BACKGROUND

[0002] In the related art, in the driver-cargo matching system of a highway freight platform, cargo and driver matching is a core link connecting drivers and cargos. The current mainstream matching strategy usually relies on a server sorting architecture, and constructs a supervised learning objective function based on explicit user behavior signals such as click rate and transaction conversion rate, such as using a deep learning model for training and prediction. Such methods have achieved certain results in improving short-term conversion efficiency, but still face many challenges in actual business landing, resulting in the driver-cargo matching method in the related art having a low matching degree of cargos and drivers, which may cause problems such as high empty running, low satisfaction, and driver loss, deviating from the long-term operation goal of the platform. Therefore, the cargo and driver matching method provided by the related art still has certain defects and room for improvement. SUMMARY

[0003] The present application provides a driver-cargo matching method fusing multi-dimensional rewards, an electronic device and a storage medium to solve the technical problem of the low matching degree of cargos and drivers in the driver-cargo matching method in the related art.

[0004] In a first aspect, the present application provides a driver-cargo matching method fusing multi-dimensional rewards, comprising: obtaining a plurality of to-be-accepted cargos, for any one of the cargos, obtaining cargo feature information corresponding to the cargo; obtaining driver feature information of a target driver, the driver feature information including vehicle type information of the driver, historical departure time information of the driver, and usage information of the driver for a client; inputting the cargo feature information and the driver feature information into a trained driver-cargo matching degree calculation model to obtain a matching degree score of the cargo and the target driver; the driver-cargo matching degree calculation model is used to calculate the matching degree score between the cargo and the target driver; obtaining a plurality of matching degree scores corresponding to the plurality of to-be-accepted cargos, sorting the plurality of matching degree scores from high to low to obtain a matching degree priority sequence.

[0005] In a possible design, the cargo feature information includes multiple types of vehicle types required by the cargo, vehicle lengths required by the cargo, ton squares of the cargo, distances required to be transported by the cargo, departure locations of the cargo, destinations of the cargo, whether the cargo requires insurance, and whether the cargo is fragile. In a possible design, the driver's vehicle model information includes a vehicle model of the driver's vehicle, a vehicle length of the driver's vehicle, a common driving route of the driver, and a route frequently clicked by the driver. The historical driving time information of the driver includes a time period when the driver usually looks for goods, a weather condition when the driver usually drives, a driving time length of the driver in a day, and a driving probability of the driver on holidays. The driver's usage information for the client includes a stay time length of the driver on a goods looking page of the client, a sliding speed of the driver on the goods looking page, and a click rate of the driver on the goods looking page.

[0006] In a possible design, the driver-goods matching degree calculation model includes a matching degree calculation unit and a reward function optimization unit. The inputting of the freight source feature information and the driver feature information into the trained driver-goods matching degree calculation model to obtain the matching degree score of the freight source and the target driver includes: The inputting of the freight source feature information and the driver feature information into the trained matching degree calculation unit, where the matching degree calculation unit is configured to perform matching degree calculation on the freight source feature information and the driver feature information to obtain an initial matching degree score of the freight source feature information and the driver feature information. The inputting of the initial matching degree score into the trained reward function optimization unit, where the reward function optimization unit is configured to perform optimization on the initial matching degree score by using a preset reward function to obtain an optimized matching degree score.

[0007] In a possible design, the driver-goods matching method further includes: obtaining freight source feature information samples and driver feature information samples; inputting the freight source feature information samples and the driver feature information samples into an initial matching degree calculation unit to obtain initial matching degree score samples; inputting the initial matching degree score samples into an initial reward function optimization unit to obtain optimized matching degree score samples; The reward function optimization unit includes a differentiable sampling module, which is configured to add noise to the initial matching degree score samples output by the initial matching degree calculation unit and perform relaxed sampling to generate a differentiable ranking matrix, rearrange the differentiable ranking matrix into a first matching degree sequence, perform optimization on the first matching degree sequence by using a preset reward function, and obtain the optimized matching degree score samples.

[0008] In a possible design, the reward function optimization unit includes a multi-dimensional reward function. The multi-dimensional reward function is represented as:

[0009] wherein, is a multi-dimensional reward function value, is a driver behavior feedback reward value, is a matching quality reward value, is a driver retention reward in a preset time, , , respectively represent reward coefficients corresponding to each reward value; wherein,

[0010] wherein, represents whether a find goods page on the client is clicked, represents whether a transaction is made after the driver clicks the find goods page, , respectively represent corresponding reward coefficients; wherein,

[0011] wherein, represents an empty driving distance reward function value, if the distance between the origin of the goods and the current position of the driver exceeds the reward threshold, a negative reward is applied, otherwise, a positive reward is applied; represents a vehicle and goods matching reward value, if the driver's vehicle type / length matches the required vehicle type / length of the current goods, a positive reward is applied, otherwise, a negative reward is applied; represents a driver complaint reward value, if the driver is complained by the customer, a negative reward is applied, otherwise, a positive reward is applied; wherein, , and are reward coefficients.

[0012] In one possible design, the method for matching goods and vehicles further includes: after obtaining the matching degree priority sequence, displaying the matching degree priority sequence; receiving a user selection of a target goods based on the matching degree priority sequence; allocating the target goods to the target driver.

[0013] In a second aspect, the present application also provides an electronic device, including a memory and a processor, the memory stores a computer program capable of running on the processor, and the processor implements the method for matching goods and vehicles by fusing multi-dimensional rewards according to any one of the above aspects when executing the computer program.

[0014] In a third aspect, the present application also provides a computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the multi-dimensional reward fusion-based driver-cargo matching method according to any one of the preceding aspects.

[0015] In a fourth aspect, the present application also provides a computer program product, comprising computer program code which, when executed on a computer, causes the computer to perform the multi-dimensional reward fusion-based driver-cargo matching method according to any one of the preceding aspects.

[0016] The multi-dimensional reward fusion-based driver-cargo matching method provided by the first aspect above first acquires a plurality of to-be-accepted cargos. For any one of the cargos, the cargo feature information corresponding to the cargo is acquired. The driver feature information of a target driver is acquired, and the driver feature information includes the vehicle model information of the driver, the historical departure time information of the driver, and the use information of the driver for the client end. Then, the cargo feature information and the driver feature information are input into the trained driver-cargo matching degree calculation model to obtain the matching degree score of the cargo and the target driver. The driver-cargo matching degree calculation model is used to calculate the matching degree score between the cargo and the target driver. Finally, a plurality of matching degree scores corresponding to the plurality of to-be-accepted cargos are acquired, and the plurality of matching degree scores are sorted from high to low to obtain a matching degree priority sequence. It can be seen that, according to the multi-dimensional reward fusion-based driver-cargo matching method provided by the present application, a plurality of driver feature information is fused when the driver is matched with the cargo, so that the matching degree priority sequence from high to low is finally obtained. In this way, the driver selects the cargo with high priority according to the matching degree priority sequence, so that the matching degree between the cargo and the driver is higher, the empty running rate of the driver is reduced, and the cargo pushing efficiency and the cargo distribution efficiency of the platform are improved.

[0017] The beneficial effects provided by the other aspects and the possible designs of the other aspects described above can be referred to the beneficial effects brought by the first aspect and the possible implementation manners of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 One of the multi-dimensional reward fusion-based driver-cargo matching method flowcharts provided by the embodiments of the present application; Figure 2 The driver-cargo matching degree calculation model structure diagram provided by the embodiments of the present application; Figure 3 The driver-cargo matching degree calculation method flowchart provided by the embodiments of the present application; Figure 4 The training method flowchart of the driver-cargo matching degree calculation model provided by the embodiments of the present application; Figure 5A schematic diagram of a work flow of a service matching degree calculation model provided by an embodiment of the present application is shown in FIG. 1. Figure 6 A schematic diagram of an electronic device structure provided by an embodiment of the present application is shown in FIG. 2. Figure 7 A schematic diagram of deployment of a service matching degree calculation model provided by an embodiment of the present application on a server and a client is shown in FIG. 3. DETAILED DESCRIPTION

[0019] In the present application, “at least one” means one or more, and “multiple” means two or more. “And / or” describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character “ / ” generally represents an “or” relationship between the associated objects before and after it. “At least one of the following” or similar expressions means any combination of these items, including single item or any combination of multiple items. For example, at least one of a, b or c alone can represent: a alone, b alone, c alone, combination of a and b, combination of a and c, combination of b and c, or combination of a, b and c, where a, b and c can be single or multiple. In addition, the terms “first” and “second” are only for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0020] The terms “center”, “longitudinal”, “transverse”, “upper”, “lower”, “left”, “right”, “front”, “back” and the like indicate the orientation or positional relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, and cannot be understood as indicating or implying that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application.

[0021] The terms “connected” and “connected” should be broadly understood, for example, the “connected” or “connected” of the circuit structure can mean not only physical connection, but also electrical connection or signal connection, for example, it can be directly connected, that is, physically connected, or indirectly connected through at least one intermediate element, as long as the circuit is connected, it can also be the internal connection of two elements; signal connection can not only be signal connection through a circuit, but also signal connection through media medium, for example, radio waves. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0022] In the related art, in the truck matching system of the highway freight platform, the matching of the cargo source and the driver is the core link connecting the driver and the cargo source. The current mainstream matching strategy usually relies on a server ranking architecture, constructs a supervised learning objective function based on explicit user behavior signals such as Click-Through-Rate (CTR) and Conversion Rate (CVR), and trains and predicts by using a deep learning model. Such methods have achieved certain results in improving short-term conversion efficiency, but still face many challenges in actual business landing, resulting in the matching degree of the cargo source and the driver being low in the related art truck matching method, which may cause high empty running, low satisfaction, and driver loss, etc., deviating from the long-term operation goal of the platform.

[0023] Firstly, the ranking model in the related art focuses on short-term behavior feedback (such as user clicks and user orders), and lacks the ability to model the business goals specific to the logistics industry. For example, whether a driver is willing to take on a certain cargo source depends not only on whether to click or transact, but also on multiple factors such as empty running distance, vehicle length and model matching degree, and future retention willingness. However, in the related art, these key business indicators are often not effectively included in the optimization target, resulting in the model ranking result improving the matching efficiency, but the matching degree between the cargo source and the driver being low, which easily causes high empty running, low satisfaction, and driver loss, etc.

[0024] Secondly, traditional ranking methods generally use a static supervised learning paradigm, which is difficult to dynamically respond to real-time behavior changes and environmental feedback of drivers. Especially in the re-ranking stage, ranking itself is a discrete and non-differentiable operation, making it difficult to implement end-to-end optimization based on reinforcement learning. Although existing research has attempted to model ranking as a sequence decision problem (such as using Pointer Network or REINFORCE algorithm), due to the non-differentiability of the ranking process, the gradient cannot be directly backpropagated to the model parameters, causing unstable training and convergence difficulties, which limits the application of reinforcement learning in actual systems.

[0025] In addition, most existing solutions complete all ranking logic on the server side, and the client side only serves as a display terminal, lacking the ability to rearrange locally and individually. This not only increases network transmission delay, but also makes it difficult to use behavior data collected in real time by the driver side (such as slide times, page dwell time, and context state) for immediate response, affecting user experience and matching timeliness.

[0026] In order to overcome the defects in the above related technologies, the application provides a driver and cargo matching method fusing multi-dimensional rewards. According to the driver and cargo matching method fusing multi-dimensional rewards provided by the application, a plurality of driver feature information is fused when a driver is matched with a cargo source by using a driver and cargo matching degree calculation model, so that a matching degree priority sequence in descending order of matching degree score is finally obtained. In this way, the driver selects a cargo source with high priority according to the matching degree priority sequence, so that the matching degree between the cargo source and the driver is higher, the empty running rate of the driver is reduced, and the cargo pushing efficiency and the cargo distribution efficiency of the platform are improved.

[0027] Figure 1 For the driver and cargo matching method fusing multi-dimensional rewards provided by the embodiment of the application, see Figure 1 for a schematic diagram of the process. The driver and cargo matching method fusing multi-dimensional rewards provided by the embodiment includes: S101, obtaining a plurality of to-be-accepted cargo sources. For any one of the cargo sources, obtaining cargo source feature information corresponding to the cargo source.

[0028] It can be understood that generally, in a city at the same time period, a plurality of to-be-accepted cargo sources will be simultaneously published on the platform. For each cargo source, the feature extraction method is used to extract the cargo source feature information corresponding to the cargo source.

[0029] In an embodiment of the application, the cargo source feature information includes a plurality of types of vehicle models required by the cargo source, vehicle lengths required by the cargo source, tonnage of the cargo source, distance required to be transported by the cargo source, departure place of the cargo source, destination of the cargo source, whether the cargo source requires insurance, and whether the cargo source is fragile.

[0030] For example, different cargos require different vehicle models. For example, according to the tonnage of the cargo, the required vehicle model can be determined. For example, cargos with large tonnage require large vehicles, and cargos with small tonnage require small trucks. In addition, cargos that require cold chain preservation require vehicles with cold storage preservation. In addition, the length of the vehicle required by different cargo sources is also different. For example, some cargos with a certain length (such as reinforcing steel bars and pipes) require a certain length of the vehicle model. For example, the length of the vehicle for transporting pipes is not less than 6 meters. In addition, the tonnage of the cargo source also determines the length of the vehicle required by the cargo or the large, medium and small vehicle model.

[0031] For example, during the operation of cargo, other factors that drivers are more concerned about are: the distance the cargo needs to be transported, the departure point of the cargo, and the destination of the cargo; for example, some drivers may not want to transport cargo with a long delivery distance in the near future based on their work and time arrangements; or, the driver determines the distance between his current location and the departure point of the cargo, and prefers to choose cargo that is closer to the departure point of the cargo; or, some drivers will choose cargo whose destination is closer to their target place (such as their home) based on the destination of the cargo.

[0032] For example, during the operation of cargo, for some special cargoes, it is also necessary to consider their fragility and whether they need to be stored in a cold chain. For valuable cargoes or easily damaged goods, the driver also needs to consider whether the cargoes are insured.

[0033] S102: Obtain driver characteristic information of the target driver, where the driver characteristic information includes the driver's vehicle type information, the driver's historical vehicle departure time information, and the driver's usage information for the client.

[0034] It is understandable that the method provided in this embodiment is mainly used for the driver's client (such as a mobile phone), etc. For a single target driver, it is only necessary to consider the matching degree between multiple cargo sources and himself. The target driver only needs to pay attention to the matching degree between each cargo source and himself.

[0035] In one embodiment of the present application, the driver's vehicle model information includes the model of the driver's vehicle, the length of the driver's vehicle, the driver's frequently traveled routes, and the routes that the driver frequently clicks on.

[0036] Understandably, the driver's vehicle model (large truck, medium truck, or small truck) and vehicle length are the most influential factors in determining driver-cargo matching, and therefore should be considered first. During model training, these factors have the greatest impact on matching. In some application scenarios, a driver's frequent routes are also a major factor influencing the matching between cargo sources and drivers. For example, a driver's frequent routes may be a preferred route based on certain factors, such as routes close to home or routes familiar to the driver. Furthermore, drivers typically browse their client app (Application) to search for cargo sources. Routes that users frequently click on or have a high click-through rate are their preferred routes, and these preferred routes can also be considered as a factor in matching.

[0037] In one embodiment of the present application, the driver's historical driving time information includes the time period when the driver usually searches for goods, the weather conditions when the driver usually drives, the driver's usual daily driving time, and the probability of the driver driving on holidays.

[0038] In some application scenarios, some drivers are limited by various factors and may have certain requirements for the departure time. For example, some drivers cannot depart at too early time or cannot depart on holidays. Therefore, the historical departure time of the driver is also one of the important factors affecting the matching degree of the driver and the cargo.

[0039] For example, some drivers can only depart at some time of the day due to family factors, which is one of the main factors affecting the matching degree between the driver and the cargo. Similarly, the average departure time of the driver in the historical data also needs to be considered.

[0040] In an embodiment of the present application, the use information of the driver for the client includes the stay duration of the driver on the cargo finding page of the client, the sliding speed of the driver on the cargo finding page, and the click rate of the driver on the cargo finding page.

[0041] In addition, the use information of the driver for the client is also introduced in the present embodiment. The use information of the driver for the client can also reflect the matching degree of the driver for the cargo to some extent.

[0042] In an embodiment of the present application, the use information of the driver for the client includes the stay duration of the driver on the cargo finding page of the client. For example, the user stays on the page of the cargo information A for a long time, which can indicate that the driver has a tendency for the cargo to some extent, and can indicate that the driver thinks that the matching degree between the cargo and the driver is high. If the driver quickly slides on the page of the cargo information A, it indicates that the driver thinks that the matching degree between the cargo and the driver is low. In addition, the high click rate of the driver on the page of the cargo information A can also indicate that the driver has a tendency for the cargo to some extent, and can also indicate that the driver thinks that the matching degree between the cargo and the driver is high.

[0043] S103, input the cargo feature information and the driver feature information into the trained driver-cargo matching degree calculation model to obtain the matching degree score of the cargo and the target driver. The driver-cargo matching degree calculation model is used to calculate the matching degree score between the cargo and the target driver.

[0044] It can be understood that the neural network model is introduced in the present embodiment to calculate the matching degree between the cargo feature information and the driver feature information. First, the initial driver-cargo matching degree calculation model can be trained by using sample data until the trained driver-cargo matching degree calculation model meets the calculation accuracy, and then the trained driver-cargo matching degree calculation model can be deployed for use. The trained driver-cargo matching degree calculation model can quickly calculate the matching degree score between each cargo and driver.

[0045] The matching degree calculation model in the embodiment can adopt a BST (Behavior Sequence Transformer) model.

[0046] Figure 2 A structure diagram of the matching degree calculation model is shown in FIG. 2. Figure 2 The matching degree calculation model 200 in the embodiment includes a matching degree calculation unit 201 and a reward function optimization unit 202. The matching degree calculation unit 201 is configured to calculate the matching degree of the freight information and the driver information to obtain an initial matching degree score of the freight information and the driver information. The reward function optimization unit 202 is configured to optimize the initial matching degree score by using a preset reward function to obtain an optimized matching degree score.

[0047] Figure 3 A flowchart of the matching degree calculation method is shown in FIG. 3. Figure 3 In the embodiment, the freight information and the driver information are input into the trained matching degree calculation model to obtain a matching degree score of the freight and the target driver. Specifically, the method includes the following steps. S301: The freight information and the driver information are input into the trained matching degree calculation unit. The matching degree calculation unit is configured to calculate the matching degree of the freight information and the driver information to obtain an initial matching degree score of the freight information and the driver information.

[0048] S302: The initial matching degree score is input into the trained reward function optimization unit. The reward function optimization unit is configured to optimize the initial matching degree score by using a preset reward function to obtain an optimized matching degree score.

[0049] Figure 4 A flowchart of the training method of the matching degree calculation model is shown in FIG. 4. Figure 4 In the embodiment, the training method of the matching degree calculation model includes the following steps. S401: Freight information samples and driver information samples are obtained.

[0050] Specifically, the specific information types in the driver information samples can include the vehicle model information of the driver, the historical driving time information of the driver, and the use information of the driver for the client. The specific information types in the freight information samples can include the vehicle model required by the freight, the vehicle length required by the freight, the tonnage of the freight, the distance required to be transported by the freight, the departure location of the freight, the destination of the freight, whether the freight requires insurance, and whether the freight is fragile.

[0051] S402, input the cargo source feature information sample and the driver feature information sample into the initial matching degree calculation unit to obtain an initial matching degree score sample.

[0052] The training of the driver-cargo matching degree calculation model in this embodiment can be divided into two stages. The first stage is to train the initial matching degree calculation unit using the cargo source feature information sample and the driver feature information sample. The second stage is to realize the differentiability of the sorting operation by using the Gumbel-Softmax technology, thereby supporting the end-to-end gradient back propagation. At the same time, a multi-dimensional reward is introduced to train the initial reward function optimization unit, thereby optimizing the matching degree value output by the initial matching degree calculation unit to further improve the matching quality.

[0053] S403, input the initial matching degree score sample into the initial reward function optimization unit to obtain an optimized matching degree score sample; wherein the reward function optimization unit includes a differentiable sampling module, and the differentiable sampling module is configured to: add Gumbel noise to the initial matching degree score sample output by the initial matching degree calculation unit and perform Softmax relaxation sampling to generate a differentiable sorting matrix, rearrange the differentiable sorting matrix into a first matching degree sequence, optimize the first matching degree sequence by using a preset reward function, and obtain the optimized matching degree score sample.

[0054] Figure 5 For the driver-cargo matching degree calculation model workflow provided by the embodiments of the present application, please refer to Figure 5 According to the driver-cargo matching degree calculation model workflow provided by the embodiments of the present application, the optimized matching degree score sample is finally obtained.

[0055] Specifically, please refer to Figure 5As shown, first, driver features, cargo source features, and scene information are acquired by a feature acquisition module; the driver features include driver registered vehicle length, vehicle model, resident city, frequently traveled route, and historical clicked route; the cargo source features include required vehicle length, required vehicle model, tonnage, loading time, loading location, and transportation distance; and the scene information includes homepage recommendation, today's cargo source, subscription, and route scene. Then, the driver features, cargo source features, and scene information are input into a feature information extraction layer, which extracts the required feature information and compresses the user features and cargo source features into a 128-dimensional semantic space, facilitating subsequent feature interaction. Further, the feature information is input into a feature gating layer, which automatically learns the importance of user-cargo source interaction through a gating mechanism. Further, the feature information is input into an attention mechanism layer, which performs sequence modeling on the fused features through a multi-head self-attention mechanism. The attention mechanism layer captures complex interaction patterns between cargo source sequences while considering effective sequence masks. Then, the feature information is input into a fusion layer, which maps the feature vectors after feature interaction and attention processing into final prediction scores. Then, it is determined whether to start a fine-tuning mode. If not, the ranking sequence is directly output. If yes, the Gumbel-softmax sampling method is used to fine-tune the model based on the defined multi-objective reward function, and the optimal matching sequence is obtained.

[0056] In the process of determining the need for fine-tuning, the fine-tuning optimization method in this embodiment specifically includes: 1. Joint loss optimization: In the training step, the total loss is the weighted sum of the policy loss (RL loss) and the supervision loss (BCE loss). The BCE loss is the cross-entropy loss based on the click label, and the RL loss is the policy gradient loss based on reinforcement learning. The BCE weight decays over time, so the model relies more on the supervision signal in the early training stage and gradually increases the weight of reinforcement learning in the later stage. This combination can improve the early training effect by using the stability of supervised learning, and optimize the long-term return through reinforcement learning.

[0057] 2. Policy gradient optimization: In the reinforcement learning part, the model uses the Gumbel-Softmax sampling strategy to generate the ranking of items and calculates the policy gradient based on the position-aware reward. Randomness is introduced through Gumbel noise, and the advantage function is used to reduce variance, so that the policy parameters are updated more effectively. The goal of this optimization is to maximize the expected multi-objective reward.

[0058] It should be noted that in the related art, how to break through the technical bottleneck that the sorting operation is not derivable, make the discrete reordering process support gradient backpropagation, and thus realize end-to-end optimization based on the idea of reinforcement learning has always been a difficulty in the field. In the present embodiment, the Gumbel-Softmax is introduced to make the initial matching degree calculation unit output score derivable and relaxed, an approximately derivable permutation matrix is constructed, so that the reward signal can be effectively backpropagated to the bottom layer parameters of the model, and the response ability of the model to dynamic feedback and the training stability are improved.

[0059] In addition, in the training of the driver-goods matching degree calculation model in the present embodiment, a two-stage training strategy (the first stage uses a supervised learning initialization strategy, and the second stage fine-tunes the training based on the reward signal) is used to obtain a model with high precision and strong generalization ability, and support conversion to TFLite format for deployment on the client side. Real-time driver feature information is used for low-latency personalized rearrangement to improve driver-goods matching efficiency and driver experience. Specifically, in the first training stage, the cargo source feature information samples and the driver feature information samples are used for supervised training using cross-entropy loss. This stage makes the model preliminarily learn the driver preferences and matching rules, and provides a stable initial sorting strategy.

[0060] Among them, the reward function optimization unit 202 of the present embodiment includes a multi-dimensional reward function; The multi-dimensional reward function is represented as: (1) Among them, is the multi-dimensional reward function value, is the driver behavior feedback reward value, is the matching quality reward value, is the driver retention reward in the preset time, , , respectively represent the reward coefficients corresponding to each reward value; the reward coefficients , , The weight value of can be set according to the weight of each reward value.

[0061] For example, the driver retention reward in the preset time can be understood as whether the driver reviews the current matching platform within the preset time period. For example, if the driver uses the driver-goods matching system again (specifically, opens the APP of the client side and browses) within seven days, it is determined that the retention reward is a positive reward, and if the driver does not use the driver-goods matching system again within seven days, it is determined that the retention reward is a negative reward.

[0062] Among them, (2) Among them, indicates whether the user clicks on the find goods page on the client, for example, if the user clicks on the find goods page on the client, it is determined that may be 1, if the user does not click on the find goods page on the client, it is determined that may be 0; indicates whether the driver clicks on the find goods page and then completes the transaction, for example, if the driver clicks on the find goods page and then completes the transaction, it is determined that is 1, if the driver clicks on the find goods page and does not complete the transaction, it is determined that is 0, , respectively indicate the corresponding reward coefficients, , determined according to the weight of the corresponding behavior.

[0063] wherein, (3) wherein, indicates the empty running distance reward function value, if the distance between the origin of the goods source and the current position of the driver exceeds the reward threshold, a negative reward is applied, otherwise, a positive reward is applied; indicates the vehicle and goods matching reward value, if the driver's vehicle type / length matches the vehicle type / length required by the current goods source, a positive reward is applied, otherwise, a negative reward is applied; indicates whether the driver is rewarded for being complained by the customer, if the driver is complained by the customer, a negative reward is applied, otherwise, a positive reward is applied; wherein, , and are reward coefficients, the size of each reward coefficient is determined according to the weight of each reward value.

[0064] S104, obtain a plurality of matching degree scores corresponding to a plurality of to-be-accepted goods sources, sort the plurality of matching degree scores from high to low to obtain a matching degree priority sequence.

[0065] It can be understood that after obtaining the plurality of matching degree scores corresponding to the plurality of to-be-accepted goods sources output by the above-mentioned driver-goods matching degree calculation model 200, then the plurality of matching degree scores are sorted from high to low to obtain a matching degree priority sequence, that is, the priority of the goods source corresponding to the matching degree score is higher, the user browses the goods source matching page, and preferentially selects the goods source with high priority (the page is in front) as the target goods source, so that the success rate of driver and goods matching is higher, the empty running rate of the driver is reduced, and the user experience and work efficiency are improved.

[0066] In an embodiment of the present application, the driver-goods matching method further comprises: after obtaining the matching degree priority sequence, displaying the matching degree priority sequence; then receiving a user selection of a target goods source selected based on the matching degree priority sequence; and assigning the target goods source to the target driver, so as to complete the accurate matching of goods and drivers.

[0067] Figure 6 For a schematic diagram of the electronic device structure provided in the embodiment of the present application, please refer to Figure 6 As shown, the electronic device includes a memory 601 and a processor 602, and the memory 601 stores a computer program that can be run on the processor 602. When the processor 602 executes the computer program, it implements the driver-cargo matching method integrating multi-dimensional rewards as described above. The driver-cargo matching method integrating multi-dimensional rewards includes: obtaining multiple sources of cargo to be accepted, and for any current source of cargo, obtaining source feature information corresponding to the source of cargo; obtaining driver feature information of the target driver, the driver feature information including the driver's vehicle model information, the driver's historical departure time information and the driver's usage information for the client; inputting the source feature information and the driver feature information into the trained driver-cargo matching calculation model to obtain the matching score between the source of cargo and the target driver; the driver-cargo matching calculation model is used to calculate the matching score between the source of cargo and the target driver; obtaining multiple matching scores corresponding to multiple sources of cargo to be accepted, sorting the multiple matching scores from high to low, and obtaining a matching priority sequence.

[0068] Among them, the trained driver-cargo matching calculation model provided in this embodiment can be deployed on a client (such as a mobile phone) to push a matching priority sequence of delivery sources to the driver, so that the driver user can easily select a delivery source with a high matching degree with himself.

[0069] It can be seen from the above description that the driver-cargo matching calculation model of this embodiment can be pre-trained in the background of the server, and then pushed to the user's client after training.

[0070] Figure 7 For the deployment diagram of the driver-cargo matching calculation model provided in the embodiment of this application on the server and client, please refer to Figure 7 As shown, the server first trains a driver-cargo matching calculation model and pushes the trained model to the user's client. This way, each time the client obtains cargo source feature information from the server, it uses the driver-cargo matching calculation model to calculate the matching score between the cargo source and the target driver. It then obtains multiple matching scores corresponding to multiple cargo sources for pending orders, sorts them from high to low, and ultimately generates a matching priority sequence, which is then displayed.

[0071] See Figure 7As shown, the server specifically includes a recall module, a coarse ranking module, a fine ranking module, a re-ranking module, a rendering module, a feature platform, a machine learning platform, an end-side backflow sample module, a model training module, and a model quantization / evaluation module; wherein the recall module, the coarse ranking module, the fine ranking module, the re-ranking module, and the rendering module are specifically configured to, at each request of the driver, recall appropriate freight sources for the driver, then perform ranking, and perform interface element rendering on the freight sources, and return to the client for display to the driver; wherein the feature platform and the machine learning platform are mainly used for feature storage, feature query, and model training.

[0072] In the model training process of the embodiment, first, samples returned by the client are used for model training through the machine learning platform, and the model file format for deployment of the client is converted; meanwhile, the model is quantized and compressed to reduce inference time consumption, and the model effect is evaluated offline, and the model file is pushed to the client model management position for subsequent packaging and deployment.

[0073] Please continue to see Figure 7 As shown, the client of the embodiment specifically includes a trigger condition module, an on-end re-ranking service module, a business logic processing module, an end-side inference engine, a feature engineering module, a model resource management module, a cloud-side feature transmission module, a machine learning platform, and a monitoring and alarm module; wherein the trigger condition in the trigger condition module is that the driver clicks a certain ticket freight in the list, triggering the calling of the on-end model service; the on-end re-ranking service in the on-end re-ranking service module is that, for freight sources in the list returned by the server that have not been exposed, based on the features transmitted by the server and the features collected and processed by the client, re-ranking is performed based on the end-side inference engine; the end-side inference engine is used for inference, and if the inference is abnormal, the original server-side ranking is maintained, without affecting the driver's experience perception; the business logic processing module is used for related processing of business logic, for example, some unexposed freight sources need to be inserted into a specified position.

[0074] The embodiment of the application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the driver-freight matching method fusing multi-dimensional rewards.

[0075] The embodiment of the application also provides a computer program product, which comprises computer program code, and when the computer program code is run on a computer, the computer program code causes the computer to execute the driver-freight matching method fusing multi-dimensional rewards.

[0076] The above-described embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available medium sets. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium. The semiconductor medium can be a solid state drive (SSD).

[0077] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of the present application.

[0078] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0079] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the embodiments of the device described above are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0080] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units. They can be located in one position or distributed on a plurality of network units. Some or all of the units can be selected according to the actual needs to achieve the purposes of the embodiments of the present application.

[0081] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.

[0082] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application essentially or the parts that make contributions to the prior art or the parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a processor (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various media that can store program codes.

[0083] Finally, it should be noted that: the above embodiments are merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A driver-cargo matching method integrating multi-dimensional rewards, characterized in that: include: Obtain multiple sources of goods to be accepted, and for any current source of goods, obtain the source characteristic information corresponding to the source of goods; Obtain driver characteristic information of the target driver, including the driver's vehicle type information, the driver's historical vehicle departure time information, and the driver's client usage information; Inputting the cargo source characteristic information and the driver characteristic information into a trained driver-cargo matching calculation model to obtain a matching score between the cargo source and the target driver; the driver-cargo matching calculation model is used to calculate the matching score between the cargo source and the target driver; A plurality of matching scores corresponding to the plurality of sources of goods to be ordered are obtained, and the plurality of matching scores are sorted from high to low to obtain a matching priority sequence.

2. The driver-cargo matching method integrating multi-dimensional rewards according to claim 1 is characterized in that: The cargo source characteristic information includes the type of vehicle required for the cargo source, the length of the vehicle required for the cargo source, the size of the cargo source, the distance the cargo source needs to be transported, the departure point of the cargo source, the destination of the cargo source, whether the cargo source needs insurance, and whether the cargo source is fragile.

3. The driver-cargo matching method integrating multi-dimensional rewards according to claim 1 is characterized in that: The vehicle type information of the driver includes the vehicle type of the driver, the length of the driver's vehicle, the routes the driver frequently travels, and the routes the driver frequently clicks; The driver's historical driving time information includes the time period when the driver usually searches for goods, the weather conditions when the driver usually drives, the driver's typical daily driving time, and the probability of the driver driving on holidays; The driver's usage information for the client includes the driver's stay time on the client's search page, the driver's sliding speed on the search page, and the driver's click rate on the search page.

4. The driver-cargo matching method integrating multi-dimensional rewards according to claim 1 is characterized in that: The driver-cargo matching degree calculation model includes a matching degree calculation unit and a reward function optimization unit; Inputting the cargo source characteristic information and the driver characteristic information into the trained driver-cargo matching calculation model to obtain the matching score between the cargo source and the target driver includes: Inputting the cargo source characteristic information and the driver characteristic information into a trained matching degree calculation unit, wherein the matching degree calculation unit is configured to perform matching degree calculation on the cargo source characteristic information and the driver characteristic information to obtain an initial matching degree score for the cargo source characteristic information and the driver characteristic information; The initial matching score is input into a trained reward function optimization unit, and the reward function optimization unit is used to optimize the initial matching score using a preset reward function to obtain an optimized matching score.

5. The driver-cargo matching method integrating multi-dimensional rewards according to claim 4 is characterized in that: The driver-cargo matching method further includes: Obtain cargo source characteristic information samples and driver characteristic information samples; Inputting the cargo source characteristic information sample and the driver characteristic information sample into an initial matching degree calculation unit to obtain an initial matching degree score sample; Inputting the initial matching score sample into the initial reward function optimization unit to obtain an optimized matching score sample; The reward function optimization unit includes a differentiable sampling module, which is used to: add noise to the initial matching score sample output by the initial matching calculation unit and perform relaxation sampling to generate a differentiable sorting matrix, rearrange the differentiable sorting matrix to obtain a first matching sequence, and optimize the first matching sequence using a preset reward function to obtain an optimized matching score sample.

6. The driver-cargo matching method integrating multi-dimensional rewards according to claim 4 is characterized in that: The reward function optimization unit includes a multi-dimensional reward function; The multi-dimensional reward function is expressed as: in, is the multi-dimensional reward function value, Feedback reward value for driver behavior, To match the quality reward value, Retention rewards for drivers within a preset time period, 、 、 Respectively represent the reward coefficient corresponding to each reward value; in, in, Indicates whether to click on the product search page on the client. Indicates whether the transaction is completed after the driver clicks the cargo search page. 、 Respectively represent the corresponding reward coefficients; in, in, Represents the value of the empty distance reward function. If the distance between the source of the cargo and the driver's current location exceeds the reward threshold, a negative reward is applied; otherwise, a positive reward is applied. Represents the reward value for matching vehicles and cargo sources. If the driver's vehicle type / length matches the vehicle type / length required by the current cargo source, a positive reward is applied; otherwise, a negative reward is applied. Indicates the reward value of whether the driver has been complained by the customer. If the driver has been complained by the customer, a negative reward is applied; otherwise, a positive reward is applied. 、 as well as is the reward coefficient.

7. The driver-cargo matching method integrating multi-dimensional rewards according to any one of claims 1 to 6, characterized in that: The driver-cargo matching method further includes: After obtaining the matching priority sequence, displaying the matching priority sequence; Receive a user's selection of a target source of supply based on the matching priority sequence; Allocate the target cargo source to the target driver.

8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor executes the computer program, the driver-cargo matching method integrating multi-dimensional rewards as described in any one of claims 1 to 7 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the driver-cargo matching method integrating multi-dimensional rewards as described in any one of claims 1 to 7 is implemented.

10. A computer program product, characterized in that The computer program product includes: a computer program code, which, when executed on a computer, enables the computer to execute the driver-cargo matching method integrating multi-dimensional rewards as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Training method, device, storage medium and electronic device for vehicle-cargo matching model

    CN109242044A

  • Transport capacity matching method, system and equipment based on multi-source data and storage medium

    CN117495023A

  • Method for driver to accurately find goods based on goods source recommendation strategy

    CN117635005A

  • Intelligent freight source order grabbing method and system

    CN119180584A

  • Vehicle and goods matching method, electronic equipment and readable storage medium

    CN119477126A

Cited By

  • Recommendation sorting method and system in combination with driver experience

    CN121146464A

  • Goods source intelligent recall method and system based on multi-dimensional condition matrix

    CN121190168A

  • Method and system for supporting local intelligent sorting of client

    CN121235567A

  • Model training method and device

    CN121258357A

  • Training method of prediction model, prediction method, electronic equipment and storage medium

    CN121303784A