Training method for journey planning model and journey planning method
By training the journey planning model based on time budget information, the problems of unreasonable and time-consuming planning results in the existing technology are solved, and efficient and personalized journey planning is achieved.
Patent Information
- Application Number
- CN202210648943.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-09
AI Technical Summary
In the journey planning process, the planning results based on the integer planning method are unreasonable and time-consuming, while the method based on the recurrent neural network does not consider the time budget information, resulting in the planning results not being personalized enough.
By obtaining the training sample set, selecting the journey request sample and the target journey sample, and training the initial journey planning model based on time budget information until the number of times threshold conditions are met, the target journey planning model is obtained.
It improves the efficiency and accuracy of journey planning, makes the planning results more reasonable and more in line with user needs, and achieves end-to-end real-time planning.
Smart Images

Figure CN114969576B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the field of deep learning technology, and can be applied to scenarios such as journey planning. In particular, it relates to a method for training a journey planning model, a journey planning method, a device, a device, a storage medium, and a computer program product. Background Art
[0002] Currently, when performing journey planning, the journey planning problem is usually transformed into an integer programming problem for solution, or the time budget information is not considered, and a journey is generated based on a recurrent neural network. However, the planning result of the planning method based on the integer programming problem may not be reasonable, and the solution process is time-consuming. The method based on the recurrent neural network does not consider the time budget information, and the planning result is not personalized enough. Summary of the Invention
[0003] The present disclosure provides a method for training a journey planning model, a journey planning method, a device, a device, a storage medium, and a computer program product, which improves the efficiency of journey planning.
[0004] According to one aspect of the present disclosure, there is provided a method for training a journey planning model, including: obtaining a training sample set, where the training sample includes a journey request sample and a corresponding target journey sample; performing the following training steps: selecting a pair of journey request sample and target journey sample from the training sample set; training an initial journey planning model based on the time budget information in the selected target journey sample and journey request sample to obtain a trained journey planning model; and in response to the number of training times meeting the first number threshold condition, determining the trained journey planning model as the target journey planning model.
[0005] According to another aspect of the present disclosure, there is provided a journey planning method, including: obtaining a journey request, where the journey request includes time budget information and departure location information; inputting the journey request into the target journey planning model to obtain a target journey.
[0006] According to still another aspect of the present disclosure, there is provided a device for training a journey planning model, including: an obtaining module configured to obtain a training sample set, where the training sample includes a journey request sample and a corresponding target journey sample; a training module configured to perform the following training steps: selecting a pair of journey request sample and target journey sample from the training sample set; training an initial journey planning model based on the time budget information in the selected target journey sample and journey request sample to obtain a trained journey planning model; and in response to the number of training times meeting the first number threshold condition, determining the trained journey planning model as the target journey planning model.
[0007] According to another aspect of the present disclosure, there is provided a journey planning device, including: an acquisition request module configured to acquire a journey request, the journey request including time budget information and departure location information; a planning module configured to input the journey request into a target journey planning model to obtain a target journey
[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the training method and the journey planning method of the above-mentioned journey planning model.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the training method and the journey planning method of the above-mentioned journey planning model.
[0010] According to still another aspect of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, the training method and the journey planning method of the above-mentioned journey planning model are implemented.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0013] Figure 1 is an exemplary system architecture diagram to which the present disclosure can be applied;
[0014] Figure 2 is a flowchart of an embodiment of the training method of the journey planning model according to the present disclosure;
[0015] Figure 3 is a flowchart of another embodiment of the training method of the journey planning model according to the present disclosure;
[0016] Figure 4 is a schematic diagram of the training method of the journey planning model according to the present disclosure;
[0017] Figure 5 is a flowchart of an embodiment of the journey planning method according to the present disclosure;
[0018] Figure 6It is a schematic structural diagram of an embodiment of a training device for a journey planning model according to the present disclosure;
[0019] Figure 7 It is a schematic structural diagram of an embodiment of a journey planning device according to the present disclosure;
[0020] Figure 8 It is a block diagram of an electronic device for implementing a training method or a journey planning method of a journey planning model according to an embodiment of the present disclosure. Detailed implementation manners
[0021] The exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] Figure 1 An exemplary system architecture 100 is shown, which can apply an embodiment of a training method or a journey planning method of a journey planning model or a training device or a journey planning device of a journey planning model according to the present disclosure.
[0023] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0024] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to obtain a target journey planning model or a journey planning result, etc. Various client applications, such as a sample acquisition application, etc., may be installed on the terminal devices 101, 102, 103.
[0025] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices, including but not limited to smartphones, tablets, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-mentioned electronic devices. They may be implemented as multiple software or software modules, or may be implemented as a single software or software module. No specific limitation is made herein.
[0026] Server 105 may provide various services based on a determined target journey planning model or journey planning result. For example, server 105 may analyze and process journey planning requests obtained from terminal devices 101, 102, and 103, and generate a processing result (such as determining time budget information, departure location information, etc.).
[0027] It should be noted that server 105 may be hardware or software. When server 105 is hardware, it may be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it may be implemented as multiple software or software modules (such as for providing distributed services), or as a single software or software module. Specific limitations are not made here.
[0028] It should be noted that the training method of the journey planning model or the journey planning method provided by the embodiments of the present disclosure is generally executed by server 105. Correspondingly, the training device or journey planning device of the journey planning model is generally disposed in server 105.
[0029] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0030] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. Figure 2 Continuing to refer to
[0031] which shows a flow 200 of an embodiment of the training method of the journey planning model according to the present disclosure. The training method of the journey planning model includes the following steps:
[0032] In this embodiment, the execution subject of the training method of the journey planning model (such as Figure 1 the server 105 shown) may obtain a training sample set. Among them, the execution subject may obtain an existing sample set stored in a publicly available database, or may collect samples through terminal devices (such as Figure 1 the terminal devices 101, 102, and 103 shown). In this way, the execution subject may receive the samples collected by the terminal devices and store these samples locally, thereby generating a training sample set.
[0033] The training sample set may include at least one pair of samples. Among them, the sample may include a journey request sample and a corresponding target journey sample. The journey request sample includes at least one piece of request information. Exemplarily, the request information may be the departure place, travel mode, regional range where the destination is located, time budget, etc., and the present disclosure does not limit this. The target journey sample corresponds to the journey request sample one by one, and is a journey pre-selected by the user under the constraint of the corresponding journey request sample. Among them, a target journey sample includes multiple locations and the order of the multiple locations. Exemplarily, a target journey sample includes three locations: the departure place, the first destination, and the second destination, and the order of the three locations is the departure place, the second destination, and the first destination.
[0034] Step 202: Select a pair of journey request samples and target journey samples from the training sample set.
[0035] In this embodiment, after obtaining the training sample set, the above-mentioned execution subject may select a pair of journey request samples and target journey samples from the training sample set. Specifically, a pair of journey request samples and target journey samples may be randomly selected from the training sample set, or a pair of journey request samples and target journey samples may be selected from the training sample set based on a preset sample selection rule, and the present disclosure does not limit this.
[0036] Step 203: Train the initial journey planning model based on the time budget information in the selected target journey sample and journey request sample to obtain a trained journey planning model.
[0037] In this embodiment, the above-mentioned execution subject may train the initial journey planning model based on the time budget information in the selected target journey sample and journey request sample. Among them, the initial journey planning model is a journey planning model that can calculate a complete journey information based on information such as the time budget information in an input journey request sample. Among them, the time budget information is the duration that the user plans to spend in a complete journey. The selected journey request sample may be used as an input data and input into the initial journey planning model. From the output end of the initial journey planning model, a journey information is output. By comparing the calculated journey information with the selected target journey sample, the initial journey planning model is trained, and the initial journey planning model obtained after one training is determined as the trained journey planning model, and the training times are accumulated by one.
[0038] After obtaining the accumulated training times, the accumulated training times may be compared with the first number threshold condition, and according to the comparison result, step 204 or 205 may be continued to be executed.
[0039] Step 204: In response to the training times not meeting the first threshold condition, use the trained journey planning model as the initial journey planning model and execute the training steps again.
[0040] In this embodiment, after obtaining the accumulated training times, the above-mentioned execution entity may, in response to the training times not meeting the first threshold condition, use the trained journey planning model as the initial journey planning model and execute the training steps again. Among them, the first threshold condition may be that the training times are equal to the first threshold. Exemplarily, the first threshold may be 100 times. Specifically, the accumulated training times may be compared with the first threshold. If the accumulated training times are less than the first threshold, the training times do not meet the first threshold condition, and the trained journey planning model may be used as the initial journey planning model to execute the above-mentioned training steps again. Specifically, steps 202 - 203 may be executed again based on the trained journey planning model.
[0041] Step 205: In response to the training times meeting the first threshold condition, determine the trained journey planning model as the target journey planning model.
[0042] In this embodiment, after obtaining the accumulated training times, the above-mentioned execution entity may, in response to the training times meeting the first threshold condition, determine the trained journey planning model as the target journey planning model. Specifically, the accumulated training times may be compared with the first threshold. If the accumulated training times are equal to the first threshold, the training times meet the first threshold condition. At this time, it may be determined that the trained journey planning model is trained, and the trained journey planning model is determined as the target journey planning model.
[0043] The training method of the journey planning model provided by the embodiments of the present disclosure first obtains a training sample set, and then executes the following training steps: select a pair of journey request samples and target journey samples from the training sample set; train the initial journey planning model based on the time budget information in the selected target journey samples and journey request samples to obtain a trained journey planning model; in response to the training times meeting the first threshold condition, determine the trained journey planning model as the target journey planning model. Training based on the time budget information enables the journey planning result obtained by the target journey planning model to consider time planning, making the journey planning result more reasonable and accurate.
[0044] Further continue to refer to Figure 3 , which shows the flow 300 of another embodiment of the training method of the journey planning model according to the present disclosure. The training method of the journey planning model includes the following steps:
[0045] Step 301: Obtain a training sample set, where the training samples include journey request samples and corresponding target journey samples.
[0046] Step 302: Select a pair of a journey request sample and a target journey sample from the training sample set.
[0047] In this embodiment, the specific operations of steps 301 - 302 have been described in detail in steps 201 - 202 of the embodiment shown in Figure 2 and will not be elaborated here.
[0048] Step 303: Obtain the time budget information and departure location information in the selected journey request sample.
[0049] In this embodiment, after selecting the journey request sample, the above-mentioned execution entity can obtain the time budget information and departure location information in the selected journey request sample. Specifically, the information in the selected journey request sample can be directly read. Among them, the time budget information is the duration that the user plans to spend in a complete journey. The time budget information can be several years, several days, or several hours, and the present disclosure does not limit this. The departure location information may include the name, coordinates, category of the departure location, and the average stay time of the departure location. Among them, the average stay time of the departure location can be the average value of the stay times of all users who have visited the departure location within a certain period.
[0050] In some alternative implementation manners of this embodiment, the user information of the user who submits the journey request can also be obtained from the selected journey request sample. Exemplarily, the user information can be the age of the user. Based on the age of the user, a journey route closer to the user's travel needs can be planned.
[0051] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0052] After obtaining the time budget information and departure location information, a joint calculation can be performed on the time budget information, departure location information, and a plurality of pre-determined candidate location information based on the initial journey planning model, and at least one target location information can be selected from the plurality of candidate location information and combined with the departure location information to form initial journey information. Specifically, steps 304 - 307 can be executed.
[0053] In some alternative implementation manners of this embodiment, the initial journey planning model may include a preprocessing layer, an attention layer, a feed-forward network layer, and a planning layer.
[0054] Step 304: Preprocess the plurality of candidate location information through the preprocessing layer to obtain a first vector matrix.
[0055] In this embodiment, the above-mentioned execution entity can preprocess multiple candidate location information through a preprocessing layer to obtain a first vector matrix. Among them, the multiple candidate location information is predetermined. Exemplarily, some or all of the candidate location information can be selected from a known candidate location database as the multiple candidate location information. Each candidate location information can include the name, coordinates, category, and average stay time information of a candidate location. Among them, the category can be, for example, a park, a shopping mall, a cultural attraction, etc. The average stay time information can be the average value of the stay times of all users who have visited the candidate location within a period of time. Specifically, all check-in information of users who have visited the candidate location within a period of time can be obtained. Among them, each check-in information records the arrival time and departure time of a user. The stay time of the user can be calculated based on the arrival time and departure time, and the average stay time of the candidate location can be obtained based on the stay times of all users.
[0056] Each piece of information of each candidate location can be concatenated based on the preprocessing layer in the initial journey planning model to form a row of the first vector matrix. After all the selected locations are concatenated, a complete first vector matrix is obtained.
[0057] In some alternative implementation manners of this embodiment, the following operations can be performed based on the preprocessing layer: converting the multiple candidate location information into corresponding multiple vector groups, where each vector group includes a coordinate embedding vector, a category embedding vector, and a stay time embedding vector; concatenating the coordinate embedding vector, category embedding vector, and stay time embedding vector in the same group into a first representation vector; and determining the obtained multiple first representation vectors as the first vector matrix.
[0058] Specifically, the coordinates, category, and average stay time of each candidate location can be respectively converted into a coordinate embedding vector, a category embedding vector, and a stay time embedding vector based on an embedding method. Among them, the coordinate embedding vector, category embedding vector, and stay time embedding vector of the same candidate location are a vector group, and multiple candidate locations correspond to multiple vector groups. Then, the coordinate embedding vector, category embedding vector, and stay time embedding vector in each group are concatenated, and the concatenated vector is linearly transformed to obtain a first representation vector. Multiple first representation vectors are obtained based on multiple vector groups, and each first representation vector serves as a row of the first vector matrix, thereby obtaining the first vector matrix. Among them, the coefficient of the linear transformation is a variable to be solved, and the first vector matrix is an N*d-dimensional matrix, where N is the number of rows of the first vector matrix and d is the dimension of each first representation vector. It should be noted that the average stay duration is a continuous value. When converting it into an embedding vector, it needs to be discretized first. Specifically, the average stay duration can be divided by 15 first, then rounded, and then converted into a stay time embedding vector by an embedding method.
[0059] In some alternative implementation manners of this embodiment, user information can be converted into a user embedding vector based on an embedding method, the user embedding vector is concatenated with each group of coordinate embedding vectors, category embedding vectors, and residence time embedding vectors, and the concatenated vector is linearly transformed to obtain a first representation vector. Multiple first representation vectors are obtained based on multiple vector groups. The user embedding vectors in each first representation vector are the same, and each first representation vector serves as a row of a first vector matrix, thereby obtaining the first vector matrix.
[0060] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0061] Step 305: Input the first vector matrix into the attention layer for calculation to obtain a second vector matrix.
[0062] In this embodiment, the above-mentioned execution subject can input the first vector matrix into the attention layer for calculation to obtain a second vector matrix. Specifically, the first vector matrix can be used as input data and input into the attention layer, and the second vector matrix is output from the output end of the attention layer.
[0063] In some alternative implementation manners of this embodiment, the input first vector matrix can be calculated through multiple attention heads in the attention layer to obtain multiple attention head matrices; the multiple attention head matrices are concatenated to obtain the second vector matrix.
[0064] Specifically, the attention layer includes multiple attention heads. Exemplarily, it can include 6 attention heads. The first vector matrix can be input into the first attention head, and the output of the first attention head is calculated based on the scaled dot product attention strategy. Then, the output of the first attention head is used as the input of the second attention head, and the calculation is continued based on the scaled dot product attention strategy, and so on, to obtain multiple attention head matrices corresponding to the multiple attention heads. The multiple attention heads can capture attention information in different aspects. After obtaining the multiple attention head matrices, the multiple attention head matrices can be concatenated, and then a linear transformation is performed based on the concatenated matrix to obtain the second vector matrix. Among them, the coefficients of the linear transformation are variables to be solved.
[0065] Step 306: Input the second vector matrix into the feed-forward network layer for calculation to obtain a candidate vector matrix.
[0066] In this embodiment, the above-mentioned execution entity may input the second vector matrix into the feedforward network layer for calculation to obtain a candidate vector matrix. Specifically, the second vector matrix may be used as input data and input into the feedforward network layer, and the candidate vector matrix is output from the output end of the feedforward network layer.
[0067] In some alternative implementation manners of this embodiment, the input second vector matrix may be non-linearly transformed through multiple feedforward network sub-layers in the feedforward network layer to obtain a candidate vector matrix.
[0068] Specifically, through multiple feedforward network sub-layers in the feedforward network layer, based on the ReLU (Rectified Linear Unit) function, the second vector matrix may be non-linearly transformed, and the calculation result is determined as the candidate vector matrix.
[0069] Step 307: Jointly calculate the candidate vector matrix, time budget information, and departure location information through a planning layer, and select at least one target location information from multiple candidate location information based on the calculation result.
[0070] In this embodiment, after obtaining the candidate vector matrix, the above-mentioned execution entity may jointly calculate the candidate vector matrix, time budget information, and departure location information through a planning layer, and select at least one target location information from multiple candidate location information based on the calculation result. Specifically, the candidate vector matrix, time budget information, and departure location information may be used as input data and input into the planning layer, and the probability of selecting each candidate location is output from the output end of the planning layer. The candidate location information corresponding to the candidate location with a selection probability greater than the probability threshold is used as the at least one target location information selected.
[0071] In some alternative implementation manners of this embodiment, the following operations may be performed based on the planning layer: jointly calculate the candidate vector matrix, time budget information, and departure location information to generate a context vector, where the context vector includes available time; select one target location information from multiple candidate location information based on the context vector; and in response to the available time not meeting the time threshold condition, perform the joint calculation of the candidate vector matrix, time budget information, and departure location information again.
[0072] Specifically, candidate vectors for each row can be obtained from the candidate vector matrix, where each candidate vector corresponds to a candidate location and includes the average residence time information and coordinate information of the candidate location. The travel mode can be obtained from the selected trip request samples. Exemplarily, the travel mode obtained is walking. Then, the information of the target locations that have been selected can be obtained. The information of the target locations that have been selected can include the departure location information or the departure location information and the information of the candidate locations that have been selected. Then, the candidate vectors corresponding to the selected target locations are obtained from the candidate vector matrix, and the coordinates and average residence time of each target location are obtained from the obtained candidate vectors. Based on the average walking speed of humans and the coordinates of each target location, the elapsed transfer time is calculated. The time budget is subtracted by the average residence time of each target location and the elapsed transfer time to obtain the available time. The available time can be converted into a vector form to form a context vector. After obtaining the context vector, based on the available time information in the context vector, a comparison is made with the average residence time of each unselected candidate location. Among the candidate locations where the average residence time is greater than the available time, the candidate location with the smallest average residence time is selected as the next target location to be selected. The time threshold condition can be that the available time is less than the average residence time of any unselected candidate location. If the available time does not meet the time threshold condition, the above steps are executed again to select the target location. If the available time meets the time threshold condition, the selection ends, and the names of the selected target locations are arranged in the order of selection to obtain a complete journey.
[0073] In some alternative implementation manners of this embodiment, the candidate vector matrix can be split and calculated to generate a global representation vector; the candidate vector matrix, the time budget information, and the departure location information are jointly calculated to obtain the available time, and the available time is converted into an available time embedding vector; the representation vector of the information of the last selected target location is obtained from the candidate vector matrix; the global representation vector, the time embedding vector, and the representation vector of the information of the last selected target location are concatenated into a context vector; a time masking operation is performed on the context vector to obtain an improved context vector; the probability of selecting each candidate location information is calculated based on the improved context vector; the candidate location information with the highest probability is determined as the target location information; in response to the available time not meeting the time threshold condition, the above steps are executed again.
[0074] Specifically, each row vector in the candidate vector matrix can be obtained, and the average value of each row vector is used as the global representation vector. Then, the candidate vector matrix, the time budget information, and the departure location information are jointly calculated to obtain the available time. The calculation steps of the available time have been introduced in detail and will not be elaborated here. After obtaining the available time, the available time can be discretized, and the available time is converted into an available time embedding vector by an embedding method. Then, the information of the finally selected target location is obtained, and a row vector corresponding to the finally selected target location is obtained from the candidate vector matrix as the representation vector of the finally selected target location information. The obtained global representation vector, the available time embedding vector, and the representation vector of the finally selected target location information are concatenated into a context vector. After obtaining the context vector, a time masking operation can be performed on the context vector. Specifically, the unselected candidate locations can be screened based on a unit step function. First, the time required to reach and stay at each unselected candidate location from the finally selected target location can be calculated, and then the available time is compared with the calculated multiple required times. The unselected candidate locations are screened using the comparison results and the unit step function. If the available time minus the required time is greater than or equal to 0, the corresponding candidate location can be selected. If the available time minus the required time is less than 0, the corresponding candidate location is screened out and no longer selected because selecting this candidate location will exceed the time budget. The screening information is added to the context vector to obtain an improved context vector. Then, based on the improved context vector, the probability of selecting each candidate location among the selectable candidate locations is calculated, and the candidate location information with the highest probability is determined as the target location information. The time threshold condition can be that the available time is less than the average stay time of any unselected candidate location. If the available time does not meet the time threshold condition, the above steps are executed again to select the target location. If the available time meets the time threshold condition, the selection ends, and the names of the selected target locations are arranged in the order of selection to obtain a complete journey as the initial journey information.
[0075] Step 308: Calculate the loss value based on the initial journey information and the selected target journey samples, and adjust the parameters of the initial journey planning model based on the loss value to obtain a trained journey planning model.
[0076] In this embodiment, after obtaining the initial journey information, the above-mentioned execution entity can calculate a loss value based on the initial journey information and the selected target journey sample, and adjust the parameters of the initial journey planning model based on the loss value to obtain a trained journey planning model. Specifically, the probability distributions of the initial journey information and the selected target journey sample can be obtained respectively, the loss value can be calculated based on the two probability distributions, the parameters of the initial journey planning model can be adjusted based on the loss value, and the adjusted initial journey planning model can be used as the trained journey planning model.
[0077] Increment the training count by one, and compare the accumulated training count with the first count threshold condition and the second count threshold condition, and determine to execute step 309, 310 or 311 based on the comparison result.
[0078] Step 309: In response to the training count not meeting the first count threshold condition, use the trained journey planning model as the initial journey planning model and execute the training step again.
[0079] In this embodiment, the specific operation of step 309 has been introduced in detail in Figure 2 step 204 of the illustrated embodiment, and will not be elaborated here.
[0080] Step 310: In response to the training count meeting the first count threshold condition and not meeting the second count threshold condition, adjust the parameters of the trained journey planning model based on the policy gradient algorithm to obtain an optimized journey planning model, and use the optimized journey planning model as the trained journey planning model, and execute the parameter adjustment of the trained journey planning model based on the policy gradient algorithm again.
[0081] In this embodiment, after obtaining the accumulated number of training times, the above-mentioned execution entity may, in response to the training times satisfying the first number threshold condition and not satisfying the second number threshold condition, adjust the parameters of the trained journey planning model based on the policy gradient algorithm to obtain an optimized journey planning model. Among them, the first number threshold condition may be that the training times are equal to the first number threshold. Exemplarily, the first number threshold may be 100 times; the second number threshold condition may be that the training times are equal to the second number threshold. Exemplarily, the second number threshold may be 200 times. It should be noted that the value of the first number threshold condition is less than the value of the second number threshold condition. Specifically, the accumulated number of training times may be first compared with the first number threshold. If the accumulated number of training times is greater than or equal to the first number threshold, the training times satisfy the first number threshold condition. On the premise that the training times satisfy the first number threshold condition, the accumulated number of training times may be further compared with the second number threshold. If the accumulated number of training times is less than the second number threshold, the training times do not satisfy the second number threshold condition. At this time, the context vector generated during the process of generating the initial journey information may be determined as the state information, and the step of selecting the target location during the process of generating the initial journey information may be determined as the action information. The initial journey planning model may further include a discriminator, and the discriminator may generate a reward value based on the generated initial journey information. The parameters of the trained journey planning model may be adjusted based on the reward value, the adjusted journey planning model may be determined as the optimized journey planning model, and the optimized journey planning model may be used as the trained journey planning model to execute the above-mentioned training steps again. Specifically, steps 302-307 may be executed again based on the trained journey planning model.
[0082] Step 311: In response to the training times satisfying the second number threshold condition, determine the optimized journey planning model as the target journey planning model.
[0083] In this embodiment, the above-mentioned execution entity may, in response to the training times satisfying the second number threshold condition, determine the optimized journey planning model as the target journey planning model. Specifically, the accumulated number of training times may be compared with the second number threshold condition. If the accumulated number of training times is equal to the second number threshold, the training times satisfy the second number threshold condition. At this time, it is determined that the training is completed, and the optimized journey planning model is determined as the target journey planning model.
[0084] From Figure 3 it can be seen that compared with Figure 2Compared with the corresponding embodiments, the training method of the journey planning model in this embodiment is trained based on information such as the category of candidate locations, making the planning result of the target journey planning model more reasonable. For example, it will not continuously take two shopping centers as the target locations; based on the context vector and time masking operation, the planning result of the target journey planning model can be made more personalized, more in line with user needs, and the user experience is improved; training based on two frequency threshold conditions improves the training efficiency; the target journey planning model trained based on the above method can achieve end-to-end journey planning and can plan in real time, improving the journey planning efficiency.
[0085] Further referring to Figure 4 , which shows a schematic diagram 400 of the training method of the journey planning model according to the present disclosure. It can be seen from Figure 4 that a journey request sample and a target journey sample can be obtained in advance. Read user information, time budget information, and departure location information from the journey request sample, then fuse the user information with a plurality of pre-determined POIs (candidate location information) to obtain a candidate vector matrix, input the candidate vector matrix into the attention layer and the feed-forward network layer for calculation to obtain the calculation result of the feed-forward network layer, generate a context vector based on the calculation result of the feed-forward network layer, time budget information, and departure location information, improve the context vector based on time encoding operation to obtain an improved context vector, select a candidate location from the unselected candidate locations as the next target location and add it to the journey based on the improved context vector, generate a new context vector again, continue to select target locations until the available time calculated based on the time budget information is not enough to select another target location, at this time the selection ends, arrange the selected target locations in the order of selection to generate a complete journey. The loss value can be calculated based on the generated journey and the target journey sample, and the journey planning model can be optimized based on the loss value. Repeat the above training steps, and when the number of training times meets the preset frequency threshold condition, the training ends to obtain the target journey model. The efficiency and accuracy of journey planning are improved.
[0086] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0087] Further referring to Figure 5 , which shows a flow chart 500 of an embodiment of the journey planning method according to the present disclosure. The journey planning method includes the following steps:
[0088] Step 501, obtain a journey request, where the journey request includes time budget information and departure location information.
[0089] In this embodiment, the above-mentioned execution entity may obtain a journey request. Among them, the journey request includes time budget information and departure location information. The specific operations of the time budget information and the departure location information have been Figure 3 detailedly introduced in step 303 of the embodiment shown, and will not be elaborated here.
[0090] Step 502: Input the journey request into the target journey planning model to obtain a target journey.
[0091] In this embodiment, after obtaining the journey request, the above-mentioned execution entity may input the journey request into the target journey planning model to obtain a target journey. Specifically, the journey request may be used as input data and input into the target journey planning model, and the target journey is output from the output end of the target journey planning model. Among them, the target journey includes multiple destinations and the play order of multiple destinations.
[0092] From Figure 5 it can be seen that the journey planning method in this embodiment can perform journey planning based on a pre-trained target journey planning model, making the target journey more reasonable, with higher planning efficiency and greater accuracy.
[0093] Further referring to Figure 6 , as an implementation of the above-mentioned training method of the journey planning model, the present disclosure provides an embodiment of a training device for a journey planning model. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0094] As Figure 6 shown, the training device 600 for the journey planning model in this embodiment may include an acquisition module 601 and a training module 602. Among them, the acquisition module 601 is configured to acquire a training sample set, where the training samples include journey request samples and corresponding target journey samples; the training module 602 is configured to perform the following training steps: select a pair of journey request samples and target journey samples from the training sample set; based on the time budget information in the selected target journey sample and journey request sample, train the initial journey planning model to obtain a trained journey planning model; in response to the number of training times satisfying the first threshold condition, determine the trained journey planning model as the target journey planning model.
[0095] In this embodiment, for the training device 600 of the journey planning model: the specific processing of the acquisition module 601 and the training module 602 and the technical effects brought by them can be respectively referred to Figure 2 the relevant descriptions of steps 201-205 in the corresponding embodiment, and will not be elaborated here.
[0096] In some alternative implementation manners of this embodiment, the training device 600 of the journey planning model further includes: a re-execution module configured to, in response to the number of training times not meeting the first threshold condition, use the trained journey planning model as the initial journey planning model and re-execute the training step.
[0097] In some alternative implementation manners of this embodiment, the training module 602 includes: an acquisition sub-module configured to acquire the time budget information and the departure location information in the selected journey request samples; a selection sub-module configured to perform joint calculation on the time budget information, the departure location information, and a plurality of pre-determined candidate location information based on the initial journey planning model, select at least one target location information from the plurality of candidate location information, and form initial journey information with the departure location information; an adjustment sub-module configured to calculate a loss value based on the initial journey information and the selected target journey samples, and perform parameter adjustment on the initial journey planning model based on the loss value to obtain the trained journey planning model.
[0098] In some alternative implementation manners of this embodiment, the initial journey planning model includes a preprocessing layer, an attention layer, a feed-forward network layer, and a planning layer; the selection sub-module includes: a preprocessing unit configured to preprocess the plurality of candidate location information through the preprocessing layer to obtain a first vector matrix; a first calculation unit configured to input the first vector matrix into the attention layer for calculation to obtain a second vector matrix; a second calculation unit configured to input the second vector matrix into the feed-forward network layer for calculation to obtain a candidate vector matrix; a planning unit configured to perform joint calculation on the candidate vector matrix, the time budget information, and the departure location information through the planning layer, and select at least one target location information from the plurality of candidate location information based on the calculation result.
[0099] In some alternative implementation manners of this embodiment, the preprocessing unit includes: a conversion sub-unit configured to convert the plurality of candidate location information into corresponding vector groups, and each vector group includes a coordinate embedding vector, a category embedding vector, and a stay time embedding vector; a first splicing sub-unit configured to splice the coordinate embedding vector, the category embedding vector, and the stay time embedding vector in the same group into a first representation vector; a determination sub-unit configured to determine the obtained plurality of first representation vectors as the first vector matrix.
[0100] In some alternative implementation manners of this embodiment, the first calculation unit includes: a first calculation sub-unit configured to calculate the input first vector matrix through a plurality of attention heads in the attention layer to obtain a plurality of attention head matrices; a second splicing sub-unit configured to splice the plurality of attention head matrices to obtain a second vector matrix.
[0101] In some alternative implementation manners of this embodiment, the second calculation unit includes: a second calculation subunit, configured to perform a non-linear transformation on the input second vector matrix through a plurality of feed-forward network sub-layers in the feed-forward network layer to obtain a candidate vector matrix.
[0102] In some alternative implementation manners of this embodiment, the planning unit includes: a third calculation subunit, configured to perform a joint calculation on the candidate vector matrix, the time budget information, and the departure location information to generate a context vector, where the context vector includes available time; a selection subunit, configured to select a target location information from a plurality of candidate location information based on the context vector; and a re-execution subunit, configured to, in response to the available time not satisfying the time threshold condition, perform the joint calculation on the candidate vector matrix, the time budget information, and the departure location information again.
[0103] In some alternative implementation manners of this embodiment, the third calculation subunit includes: a splitting subunit layer, configured to perform a splitting calculation on the candidate vector matrix to generate a global representation vector; a first calculation subunit layer, configured to perform a joint calculation on the candidate vector matrix, the time budget information, and the departure location information to obtain the available time, and convert the available time into an available time embedding vector; an obtaining subunit layer, configured to obtain a representation vector of the last selected target location information from the candidate vector matrix; and a splicing subunit layer, configured to splice the global representation vector, the available time embedding vector, and the representation vector of the last selected target location information into a context vector.
[0104] In some alternative implementation manners of this embodiment, the selection subunit includes: a masking subunit layer, configured to perform a time masking operation on the context vector to obtain an improved context vector; a second calculation subunit layer, configured to calculate the probability of selecting each candidate location information based on the improved context vector; and a determination subunit layer, configured to determine the candidate location information with the maximum probability as the target location information.
[0105] In some alternative implementation manners of this embodiment, the training module 602 further includes: a re-execution sub-module, configured to, in response to the number of training times satisfying the first number threshold condition and not satisfying the second number threshold condition, adjust the parameters of the trained journey planning model based on the policy gradient algorithm to obtain an optimized journey planning model, and use the optimized journey planning model as the trained journey planning model to perform the parameter adjustment on the trained journey planning model based on the policy gradient algorithm again; a determination sub-module, configured to, in response to the number of training times satisfying the second number threshold condition, determine the optimized journey planning model as the target journey planning model; where the value of the first number threshold condition is less than the value of the second number threshold condition.
[0106] Further referenceFigure 7 , as an implementation of the above journey planning method, the present disclosure provides an embodiment of a journey planning device, which corresponds to the method embodiment shown in Figure 5 and can be specifically applied to various electronic devices.
[0107] As shown in Figure 7 , the journey planning device 700 of this embodiment may include an acquisition request module 701 and a planning module 702. Among them, the acquisition request module 701 is configured to acquire a journey request, and the journey request includes time budget information and departure location information; the planning module 702 is configured to input the journey request into a target journey planning model to obtain a target journey.
[0108] In this embodiment, for the journey planning device 700: the specific processing of the acquisition request module 701 and the planning module 702 and the technical effects brought by them can respectively refer to the relevant descriptions of steps 501-502 in the Figure 5 corresponding embodiment, which will not be elaborated here.
[0109] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0110] Figure 8 The schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0111] As shown in Figure 8 , the device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 802 or the computer program loaded from the storage unit 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.
[0112] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as a keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as a disk, optical disc, etc.; and communication unit 809, such as a network card, modem, wireless communication transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0113] Computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 801 executes the various methods and processes described above, such as the training method or journey planning method of the journey planning model. For example, in some embodiments, the training method or journey planning method of the journey planning model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by computing unit 801, one or more steps of the training method or journey planning method of the journey planning model described above can be executed. Alternatively, in other embodiments, computing unit 801 can be configured to execute the training method or journey planning method of the journey planning model in any other suitable way (e.g., by means of firmware).
[0114] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0115] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program codes can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0116] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0117] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0118] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0119] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology. The server can be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.
[0120] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0121] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A training method for a journey planning model, comprising: Obtaining a training sample set, wherein the training samples include journey request samples and corresponding target journey samples; Performing the following training steps: Selecting a pair of a journey request sample and a target journey sample from the training sample set; obtaining the time budget information and departure location information in the selected journey request sample; preprocessing a plurality of predetermined candidate location information through a preprocessing layer to obtain a first vector matrix; inputting the first vector matrix into an attention layer for calculation to obtain a second vector matrix; inputting the second vector matrix into a feed-forward network layer for calculation to obtain a candidate vector matrix; Performing joint calculation on the candidate vector matrix, the time budget information, and the departure location information through a planning layer, and selecting at least one target location information from the plurality of candidate location information based on the calculation result, including: performing the following operations based on the planning layer: performing joint calculation on the candidate vector matrix, the time budget information, and the departure location information to generate a context vector, where the context vector includes available time; selecting a target location information from the plurality of candidate location information based on the context vector; in response to the available time not satisfying the time threshold condition, performing the joint calculation on the candidate vector matrix, the time budget information, and the departure location information again; Combining the selected target location information with the departure location information to form initial journey information; calculating a loss value based on the initial journey information and the selected target journey sample, and adjusting the parameters of the initial journey planning model based on the loss value to obtain a trained journey planning model, where the initial journey planning model includes the preprocessing layer, the attention layer, the feed-forward network layer, and the planning layer; In response to the number of training times satisfying the first number threshold condition, determining the trained journey planning model as the target journey planning model.
2. The method according to claim 1, further comprising: In response to the number of training times not satisfying the first number threshold condition, using the trained journey planning model as the initial journey planning model and performing the training steps again.
3. The method according to claim 1, wherein, The preprocessing the plurality of candidate location information through the preprocessing layer to obtain a first vector matrix includes: Performing the following operations based on the preprocessing layer: Converting the plurality of candidate location information into corresponding plurality of vector groups, where each vector group includes a coordinate embedding vector, a category embedding vector, and a stay time embedding vector; Concatenating the coordinate embedding vector, the category embedding vector, and the stay time embedding vector in the same group into a first representation vector; Determining the obtained plurality of first representation vectors as the first vector matrix.
4. The method according to claim 3, wherein The inputting the first vector matrix into the attention layer for calculation to obtain a second vector matrix includes: Calculating the input first vector matrix through a plurality of attention heads in the attention layer to obtain a plurality of attention head matrices; Concatenating the plurality of attention head matrices to obtain the second vector matrix.
5. The method according to claim 4, wherein Inputting the second vector matrix into a feed-forward network layer for calculation to obtain a candidate vector matrix includes: Performing a non-linear transformation on the input second vector matrix through a plurality of feed-forward network sub-layers in the feed-forward network layer to obtain the candidate vector matrix.
6. The method according to claim 1, wherein Jointly calculating the candidate vector matrix, the time budget information, and the departure location information to generate a context vector includes: Performing a split calculation on the candidate vector matrix to generate a global representation vector; Jointly calculating the candidate vector matrix, the time budget information, and the departure location information to obtain the available time, and converting the available time into an available time embedding vector; Obtaining a representation vector of the target location information that was last selected from the candidate vector matrix; Concatenating the global representation vector, the available time embedding vector, and the representation vector of the target location information that was last selected into the context vector.
7. The method according to claim 6, wherein, Selecting a target location information from the plurality of candidate location information based on the context vector includes: Performing a time masking operation on the context vector to obtain an improved context vector; Calculating the probability of selecting each candidate location information based on the improved context vector; Determining the candidate location information with the highest probability as the target location information.
8. The method according to any one of claims 1-7, wherein, Determining the trained journey planning model as the target journey planning model in response to the number of training times satisfying a first number threshold condition includes: In response to the number of training times satisfying the first number threshold condition and not satisfying the second number threshold condition, adjusting the parameters of the trained journey planning model based on the policy gradient algorithm to obtain an optimized journey planning model, and using the optimized journey planning model as the trained journey planning model, and again performing parameter adjustment on the trained journey planning model based on the policy gradient algorithm; In response to the number of training times satisfying the second number threshold condition, determining the optimized journey planning model as the target journey planning model; Wherein, the value of the first number threshold condition is less than the value of the second number threshold condition.
9. A journey planning method, including: Obtaining a journey request, where the journey request includes time budget information and departure location information; Inputting the journey request into a target journey planning model to obtain a target journey, where the target journey planning model is trained based on the training method according to any one of claims 1-8.
10. A training device for a journey planning model, the device includes: An acquisition module configured to acquire a training sample set, where the training samples include journey request samples and corresponding target journey samples; A training module configured to perform the following training steps: selecting a pair of journey request samples and target journey samples from the training sample set; training an initial journey planning model based on the target journey sample and the time budget information in the selected journey request sample to obtain a trained journey planning model; determining the trained journey planning model as the target journey planning model in response to the number of training times satisfying a first number threshold condition; Among them, the training module includes: an acquisition sub-module configured to acquire the time budget information and departure location information in the selected journey request sample; a selection sub-module configured to perform a joint calculation on the time budget information, the departure location information, and a plurality of pre-determined candidate location information based on the initial journey planning model, select at least one target location information from the plurality of candidate location information, and form initial journey information with the departure location information; an adjustment sub-module configured to calculate a loss value based on the initial journey information and the selected target journey sample, and adjust the parameters of the initial journey planning model based on the loss value to obtain the trained journey planning model; Among them, the initial journey planning model includes a preprocessing layer, an attention layer, a feed-forward network layer, and a planning layer; the selection sub-module includes: a preprocessing unit configured to preprocess the plurality of candidate location information through the preprocessing layer to obtain a first vector matrix; a first calculation unit configured to input the first vector matrix into the attention layer for calculation to obtain a second vector matrix; a second calculation unit configured to input the second vector matrix into the feed-forward network layer for calculation to obtain a candidate vector matrix; a planning unit configured to perform a joint calculation on the candidate vector matrix, the time budget information, and the departure location information through the planning layer, and select at least one target location information from the plurality of candidate location information based on the calculation result; Among them, the planning unit includes: a third calculation sub-unit configured to perform a joint calculation on the candidate vector matrix, the time budget information, and the departure location information to generate a context vector, where the context vector includes available time; a selection sub-unit configured to select one target location information from the plurality of candidate location information based on the context vector; a re-execution sub-unit configured to, in response to the available time not satisfying the time threshold condition, re-execute the joint calculation on the candidate vector matrix, the time budget information, and the departure location information.
11. The apparatus according to claim 10, wherein the apparatus further includes: A re-execution module configured to, in response to the number of training times not satisfying the first number threshold condition, use the trained journey planning model as the initial journey planning model and re-execute the training step.
12. The apparatus according to claim 11, wherein, The preprocessing unit includes: A conversion sub-unit configured to convert the plurality of candidate location information into corresponding vector groups, and each vector group includes a coordinate embedding vector, a category embedding vector, and a stay time embedding vector; A first splicing sub-unit configured to splice the coordinate embedding vector, the category embedding vector, and the stay time embedding vector in the same group into a first representation vector; A determination sub-unit configured to determine the obtained plurality of first representation vectors as the first vector matrix.
13. The device according to claim 12, wherein, The first calculation unit includes: A first computing subunit, configured to calculate the input first vector matrix through a plurality of attention heads in the attention layer to obtain a plurality of attention head matrices; A second splicing subunit, configured to splice the plurality of attention head matrices to obtain the second vector matrix.
14. The apparatus according to claim 13, wherein, The second computing unit includes: A second computing subunit, configured to perform a non-linear transformation on the input second vector matrix through a plurality of feed-forward network sub-layers in the feed-forward network layer to obtain the candidate vector matrix.
15. The device according to claim 10, wherein, The third computing unit includes: A splitting subunit layer, configured to perform splitting calculation on the candidate vector matrix to generate a global representation vector; A first computing subunit layer, configured to jointly calculate the candidate vector matrix, the time budget information, and the departure location information to obtain the available time, and convert the available time into an available time embedding vector; An obtaining subunit layer, configured to obtain a representation vector of the finally selected target location information from the candidate vector matrix; A splicing subunit layer, configured to splice the global representation vector, the available time embedding vector, and the representation vector of the finally selected target location information into the context vector.
16. The apparatus according to claim 15, wherein, The selection subunit includes: A masking subunit layer, configured to perform a time masking operation on the context vector to obtain an improved context vector; A second computing subunit layer, configured to calculate the probability of selecting each candidate location information based on the improved context vector; A determining subunit layer, configured to determine the candidate location information with the highest probability as the target location information.
17. The device according to any one of claims 10-16, wherein, The training module further includes: A re-execution sub-module, configured to, in response to the training times satisfying the first number threshold condition and not satisfying the second number threshold condition, adjust the parameters of the trained journey planning model based on the policy gradient algorithm to obtain an optimized journey planning model, and use the optimized journey planning model as the trained journey planning model to re-execute the parameter adjustment of the trained journey planning model based on the policy gradient algorithm; A determining sub-module, configured to, in response to the training times satisfying the second number threshold condition, determine the optimized journey planning model as the target journey planning model; wherein the value of the first number threshold condition is less than the value of the second number threshold condition.
18. A journey planning device, wherein, The device includes: A request obtaining module, configured to obtain a journey request, where the journey request includes time budget information and departure location information; A planning module, configured to input the journey request into the target journey planning model to obtain a target journey, where the target journey planning model is trained by the training device according to any one of claims 10-17.
19. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-9.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are for causing the computer to execute the method according to any one of claims 1-9.
21. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-9.
Citation Information
Patent Citations
Method and device for establishing route planning model and planning tour route
CN107633317A