Method and system for carrying out travel transfer sorting by applying word vector and cross network

By constructing a neural network model of the city and site embedding layer, using word vectors and cross networks for feature interaction, the problem of lack of accuracy in the recommendation of transit solution is solved, and a more accurate transit solution sorting is achieved.

CN120296241APending Publication Date: 2025-07-11SUZHOU CHUANGLUTIANXIA INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510223906.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing methods of recommendation of transit schemes cannot fully understand the deep correlation between cities and sites, resulting in a lack of accuracy in recommendation results.

Method used

By constructing a neural network model that includes urban embedding layer and site embedding layer, using word vectors to capture the semantic relationship between cities and sites, and perform feature interactions through cross-networks, combining logistic regression functions to calculate the estimated purchase rate, and using batch training and test verification methods to optimize the model.

Benefits of technology

The accuracy and generalization capabilities of the recommendation results of the transit plan are improved, ensuring that the model can understand and utilize the geographical correlation characteristics between cities and sites, and provide reliable sorting basis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296241A_ABST
    Figure CN120296241A_ABST
Patent Text Reader

Abstract

The invention provides a method, system and device for carrying out travel transfer sorting by applying word vectors and a cross network and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: collecting feature data and mark data of a transfer scheme, and judging whether city word vectors and site word vectors in the feature data meet preset conditions or not; if the preset condition is met, constructing a neural network model; performing iterative training on the neural network model through the training data set until a loss function in the neural network model converges to a preset threshold value, and obtaining a target neural network model; verifying the accuracy rate of the target neural network model, and if the accuracy rate is greater than a preset accuracy rate, inputting the plurality of to-be-sorted transfer schemes into the target neural network model to obtain an estimated purchase rate of each to-be-sorted transfer scheme; and sorting the to-be-sorted transfer schemes according to the estimated purchase rate to generate a sorting result. The method has the technical effect that the accuracy of the transfer scheme recommendation result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, system, device and storage medium for sorting travel transfers using word vectors and cross networks. Background Art

[0002] With the continuous development and improvement of the transportation network, travel modes are becoming more and more diversified. Users often face multiple options when choosing transfer routes. How to quickly find the most suitable option for user needs from a large number of transfer options and provide users with personalized recommendation results has become a technical problem that needs to be solved in the current transportation field.

[0003] At present, common transfer plan recommendation methods are mainly based on rule matching and simple statistical models. Such methods usually score and sort transfer plans according to preset rules (such as fare, duration, number of transfers, etc.), or use simple statistical features of historical data to make recommendations. Although the above methods can provide users with basic transfer plan recommendations, they cannot fully understand the deep connection between cities and stations, and it is difficult to capture the complex patterns in user selection behavior, resulting in a lack of accuracy in the transfer plan recommendation results. Summary of the invention

[0004] The present application provides a method, system, device and storage medium for sorting travel transfers using word vectors and cross networks, which are used to improve the accuracy of transfer scheme recommendation results.

[0005] In a first aspect, the present application provides a method for sorting travel transfers using word vectors and a cross network. The method includes: collecting feature data of transfer plans and labeled data corresponding to the feature data, and dividing the feature data and the labeled data into a training data set and a test data set according to a preset ratio; obtaining the city word vectors and station word vectors in the feature data, and determining whether both the city word vectors and the station word vectors meet preset conditions; if both the city word vectors and the station word vectors meet the preset conditions, constructing a neural network model, the neural network model including a city embedding layer and a station embedding layer; loading the city word vectors into the city embedding layer, and loading the station word vectors into the station embedding layer, and setting the city embedding layer and the station embedding layer to be trainable; iteratively training the neural network model through the training data set until the loss function in the neural network model converges to a preset threshold to obtain a target neural network model; verifying the accuracy of the target neural network model through the test data set, if the accuracy is greater than a preset accuracy, inputting multiple transfer plans to be sorted into the target neural network model to obtain the estimated purchase rates of the transfer plans to be sorted; sorting the transfer plans to be sorted according to the estimated purchase rates to generate a sorting result.

[0006] By adopting the above technical solution, by constructing a neural network model including a city embedding layer and a station embedding layer, it is possible to effectively learn the feature representations of cities and stations and capture the semantic relationships between them in the form of word vectors. During the model training process, by pre-verifying whether the word vectors meet the preset conditions, the quality of the input data is ensured; by dividing the data set into a training set and a test set and setting the loss function convergence threshold and accuracy requirements, the generalization ability and prediction accuracy of the model are guaranteed. Finally, the model can accurately estimate the purchase rates of each transfer plan, providing a reliable basis for plan sorting, thereby improving the accuracy of the transfer plan recommendation results.

[0007] Optionally, the determination of whether both the urban word vector and the station word vector meet the preset conditions includes: constructing an urban word vector distance matrix based on the cosine distance between any two of the urban word vectors; constructing a station word vector distance matrix based on the cosine distance between any two of the station word vectors; selecting associated cities with a distance less than a first preset value from the urban word vector distance matrix for the urban word vector of the target city; selecting associated stations with a distance less than a second preset value from the station word vector distance matrix for the station word vector of the target station; determining whether the associated cities are in the same province or adjacent provinces as the target city, and whether the associated stations are in the same city or adjacent cities as the target station; when at least a third preset proportion of the cities in the associated cities are in the same province or adjacent provinces as the target city, and at least a fourth preset proportion of the stations in the associated stations are in the same city or adjacent cities as the target station, it is determined that the urban word vector and the station word vector meet the preset conditions.

[0008] By adopting the above technical solution, by constructing the distance matrices of urban and station word vectors and verifying in combination with geographical location information, it is ensured that the word vectors can accurately reflect the geographical relevance between cities and stations. Specifically, the cosine distance is calculated to measure the similarity of the word vectors, and preset values are set to screen associated cities and stations, and then the screening results are compared and verified with the actual geographical relationships. When the preset proportion requirements are met, it indicates that the word vectors have successfully captured the proximity relationship in the geographical space, thus ensuring that the subsequent neural network model can accurately understand and utilize the geographical association features between cities and stations during feature learning, and improving the prediction accuracy of the model.

[0009] Optionally, the construction of the neural network model includes: constructing a cross network including multiple cross layers, where the output of each cross layer is obtained by performing feature interaction calculation on the input features; connecting the urban embedding layer and the station embedding layer to the input end of the cross network respectively, and setting a logistic regression function layer at the output end of the cross network, where the logistic regression function layer is used to convert the output value of the cross network into the estimated purchase rate; setting the exposure data with purchase behavior as positive sample data, setting the exposure data without purchase behavior as negative sample data, and training the positive sample data and the negative sample data using the logarithmic loss function.

[0010] By adopting the above technical solution, through constructing a multi-layer cross network structure, the deep interaction between the urban embedding features and the station embedding features is realized, and the non-linear combination relationship between the features can be effectively captured. A logistic regression function layer is set at the network output end to convert the cross features into the estimated purchase rate, so that the model output has a clear business meaning. At the same time, by dividing the user behavior data into positive and negative samples and using the logarithmic loss function for training, the model can accurately learn the purchase preferences of users and improve the accuracy of the estimated purchase rate. This structural design not only ensures the sufficiency of feature interaction but also ensures the interpretability of the model prediction results, thus improving the overall effect of the transfer plan ranking.

[0011] Optionally, the iterative training of the neural network model with the training data set until the loss function in the neural network model converges to a preset threshold to obtain the target neural network model includes: dividing the training data set into multiple training batch data according to a preset batch size; sequentially inputting each training batch data into the neural network model and calculating the current loss function value of the neural network model; according to the current loss function value, using the backpropagation algorithm to sequentially update the parameters in the urban embedding layer and the station embedding layer; until the current loss function value is less than the preset threshold to obtain the target neural network model.

[0012] By adopting the above technical solution, the neural network model is iteratively optimized by means of batch training, effectively balancing the training efficiency and memory consumption. By dividing the training data set into multiple batches, inputting them into the model for training batch by batch, and using the backpropagation algorithm to dynamically update the parameters of the urban embedding layer and the station embedding layer until the loss function value converges to the preset threshold. This training strategy not only ensures that the model can fully learn the feature patterns in the data, but also avoids the computational resource pressure caused by loading all the data at once. At the same time, the setting of the loss function threshold ensures the convergence quality of the model training and improves the performance of the final target neural network model.

[0013] Optionally, verifying the accuracy of the target neural network model with the test data set includes: dividing the test data set into multiple verification batch data; sequentially inputting each verification batch data into the target neural network model to obtain the prediction results of each verification batch data; comparing the prediction results of each verification batch data with the corresponding labeled data, and counting the number of correctly predicted samples; obtaining the accuracy of the target neural network model according to the ratio of the number of correctly predicted samples to the total number of samples in the test data set.

[0014] By adopting the above technical solution, through the method of batch partitioning and batch-by-batch verification of the test data set, a comprehensive evaluation of the performance of the target neural network model is achieved. By comparing the prediction results of each verification batch with the true labeled data, counting the number of correctly predicted samples, and calculating the accuracy rate, not only can the prediction ability of the model be accurately measured, but also the computational pressure caused by processing a large amount of data at one time can be effectively avoided. This way of batch verification not only ensures the reliability of the verification process but also improves the verification efficiency, providing an objective and accurate measurement standard for evaluating the model performance.

[0015] Optionally, the step of inputting multiple to-be-sorted transfer plans into the target neural network model to obtain the estimated purchase rates of the to-be-sorted transfer plans includes: obtaining the city features and station features in each of the to-be-sorted transfer plans, inputting the city features into the city embedding layer, and inputting the station features into the station embedding layer; inputting the output features of the city embedding layer and the station embedding layer into the cross network, and performing feature interaction calculations through each of the cross layers to obtain the interacted feature vectors; inputting the interacted feature vectors into the logistic regression function layer to obtain the estimated purchase rates corresponding to the to-be-sorted transfer plans.

[0016] By adopting the above technical solution, by separately inputting the city and station features of the to-be-sorted transfer plans into the trained neural network model, multi-level processing and transformation of the features are achieved. First, vector representations of the features are obtained through the city embedding layer and the station embedding layer, then deep feature interactions are carried out using the cross network to capture the complex correlation relationships between the features, and finally, through the logistic regression function layer, the interacted features are transformed into the estimated purchase rates with actual business meanings. This hierarchical feature processing flow not only ensures the sufficiency of feature interaction but also the interpretability of the prediction results, thus providing a reliable quantitative basis for the sorting of transfer plans.

[0017] Optionally, after sorting the to-be-sorted transfer plans according to the estimated purchase rates to generate a sorting result, the method further includes: calculating the actual purchase rates of the to-be-sorted transfer plans according to the historical display data and historical purchase data of the to-be-sorted transfer plans; calculating the differences between the estimated purchase rates and the actual purchase rates of the to-be-sorted transfer plans; obtaining the supplementary feature data and supplementary labeled data of the transfer plans with the differences greater than the preset deviation value; adding the supplementary feature data and the supplementary labeled data to the training data set to retrain the neural network model.

[0018] By adopting the above technical solution, a feedback mechanism for the prediction effect of the model is established by comparing the estimated purchase rate and the actual purchase rate of the transfer plan. When the prediction deviation exceeds the preset threshold, by collecting supplementary feature data and labeled data of the relevant transfer plan and adding them to the training data set for retraining the model, the dynamic update of the data and the continuous optimization of the model are realized. This feedback optimization mechanism based on the prediction deviation can not only timely detect and correct the prediction deviation of the model in specific scenarios, but also continuously enrich the diversity of the training data, improve the prediction accuracy of the model for various transfer plans, so that the sorting result is more in line with the actual business scenario.

[0019] In a second aspect, the present application provides a system for sorting travel transfers using word vectors and cross networks. The system includes: a collection module, a judgment module, a construction module, a loading module, a training module, and an output module; wherein, the collection module is used to collect the feature data of the transfer plan and the labeled data corresponding to the feature data, and divide the feature data and the labeled data into a training data set and a test data set according to a preset ratio; the judgment module is used to obtain the city word vector and the station word vector in the feature data, and judge whether both the city word vector and the station word vector meet the preset conditions; the construction module is used to construct a neural network model if both the city word vector and the station word vector meet the preset conditions, and the neural network model includes a city embedding layer and a station embedding layer; the loading module is used to load the city word vector into the city embedding layer, and load the station word vector into the station embedding layer, and set the city embedding layer and the station embedding layer to a trainable state; the training module is used to iteratively train the neural network model through the training data set until the loss function in the neural network model converges to a preset threshold to obtain a target neural network model; the output module is used to verify the accuracy of the target neural network model through the test data set. If the accuracy is greater than the preset accuracy, input multiple transfer plans to be sorted into the target neural network model to obtain the estimated purchase rate of each transfer plan to be sorted; sort each transfer plan to be sorted according to the estimated purchase rate to generate a sorting result.

[0020] In a third aspect, the present application provides an electronic device, adopting the following technical solution: including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes a computer program of any one of the above methods for sorting travel transfers using word vectors and cross networks.

[0021] Fourthly, the present application provides a computer-readable storage medium, adopting the following technical solution: storing a computer program that can be loaded and executed by a processor to perform any of the above methods for sorting travel transfer plans using application word vectors and cross networks.

[0022] In summary, the present application includes at least one of the following beneficial technical effects: By constructing a neural network model including a city embedding layer and a station embedding layer, it is possible to effectively learn the feature representations of cities and stations, and capture the semantic relationships between them in the form of word vectors. During the model training process, by pre-verifying whether the word vectors meet the preset conditions, the quality of the input data is ensured; by dividing the data set into a training set and a test set, and setting the convergence threshold of the loss function and the accuracy requirement, the generalization ability and prediction accuracy of the model are guaranteed. Finally, the model can accurately estimate the purchase rates of each transfer plan, providing a reliable basis for plan sorting, thereby improving the accuracy of the transfer plan recommendation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flowchart of a method for sorting travel transfers using application word vectors and cross networks provided by an embodiment of the present application; Figure 2 is a structural diagram of a system for sorting travel transfers using application word vectors and cross networks provided by an embodiment of the present application; Figure 3 is a structural diagram of an electronic device provided by an embodiment of the present application.

[0024] Description of the reference numerals: 1000, electronic device; 1001, processor; 1002, communication bus; 1003, user interface; 1004, network interface; 1005, memory. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0026] In the description of the embodiments of the present application, words such as "exemplary", "for example" or "for illustration" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary", "for example" or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "exemplary", "for example" or "for illustration" is intended to present relevant concepts in a specific manner.

[0027] Figure 1 1 is a flow chart of a method for applying word vectors and cross networks to perform travel transfer sorting according to an embodiment of the present application. Figure 1 As shown, the method includes S101-S106: S101, collecting feature data of a transfer plan and label data corresponding to the feature data, and dividing the feature data and the label data into a training data set and a test data set according to a preset ratio.

[0028] In order to realize the intelligent sorting of transit plans, it is necessary to first obtain the data required for training and validating the model. Specifically, from the historical records of the list of intelligently recommended transit plans, the characteristic data of each transit plan is collected, including: total duration of the plan, total price of the plan, departure station, arrival station, transit station, departure city, arrival city, transit city, user preferred transit station and user preferred transit city. Among them, the user preferred transit station refers to the transit station that appears most frequently in the historical purchase records of the user from the same departure city to the arrival city; the definition of the user preferred transit city is similar to that of the user preferred transit station, which refers to the transit city that appears most frequently in the historical purchase records. These characteristic data reflect the basic attributes of the transit plan and the user's historical preferences, which helps the model learn the user's selection pattern.

[0029] For each feature data, its corresponding label data also needs to be recorded, that is, whether the user purchased the transfer plan after being exposed. Specifically, if the user completes the purchase after viewing the transfer plan, the label data is 1, indicating a positive sample; if the user does not purchase after viewing, the label data is 0, indicating a negative sample. This labeling method enables the model to learn to distinguish the features of the transfer plan that the user is more likely to purchase.

[0030] In order to ensure the generalization ability of the model and evaluate the actual effect of the model, the collected data set needs to be divided into a training data set and a test data set according to a preset ratio (such as 8:2). The training data set is used for parameter learning and optimization of the model, while the test data set is used to evaluate the actual prediction effect of the model. This division method can avoid model overfitting, that is, the model overfits the training data and leads to poor performance in actual applications.

[0031] The structure of a single piece of data can be represented as [Y, X1, X2,..., Xn], where Y is the labeled data and X1 to Xn are the feature data. For example, a specific piece of data can be represented as [0, 120, 12.5, 16666, 18888, 17777, 266, 288, 277, 36666, 388], corresponding to [whether to purchase, total plan duration (minutes), total plan price (yuan), departure station number, arrival station number, transfer station number, departure city number, arrival city number, transfer city number, user-preferred transfer station number, user-preferred transfer city number] respectively. Through this structured data collection and division method, a reliable data foundation is provided for subsequent model training and verification.

[0032] S102, obtain the city word vectors and station word vectors in the feature data, and determine whether both the city word vectors and the station word vectors meet the preset conditions.

[0033] Specifically, use a mature Chinese text embedding model (such as word2vec, text2vec, or Chinese BERT, etc.) to train a large-scale Chinese corpus, and extract the word vectors corresponding to the city names and station names from it. These word vectors can reflect the relationship between cities and stations in the semantic space, and the dimension is usually between 64 and 512. Input all the city names that appear in the training data into the pre-trained model to obtain the corresponding city word vectors, and store the corresponding relationship between the city numbers and the word vectors as a city_emd file, for example, save it in the format of "1_Suzhou City_[0.1261, 0.1475, 0.9130, 0.4347...]"; similarly, obtain the station word vectors for all station names, and store the corresponding relationship between the station numbers and the word vectors as a station_emd file.

[0034] In order to verify whether the obtained word vectors truly reflect the geographical and semantic relationships between cities and stations, it is necessary to make a judgment on the preset conditions. First, calculate the cosine distance between any two city word vectors to construct a city word vector distance matrix; construct a station word vector distance matrix in the same way. For the target city, select the cities with a word vector distance less than the first preset value (such as 0.3) as the associated cities; for the target station, select the stations with a word vector distance less than the second preset value (such as 0.25) as the associated stations.

[0035] To further verify the rationality of these correlation relationships, check whether the associated cities are in the same province or adjacent provinces as the target city, and whether the associated stations are in the same city or adjacent cities as the target station. When at least a third preset ratio (such as 80%) of the cities in the associated cities are in the same province or adjacent provinces as the target city, and at least a fourth preset ratio (such as 75%) of the stations in the associated stations are in the same city or adjacent cities as the target station, it is considered that the obtained word vectors meet the preset conditions.

[0036] For example, through calculation, it is found that the cities with the word vectors closest to the word vector of "Suzhou City" include "Wuxi City", "Nantong City", "Changzhou City", etc., and these cities are indeed all within Jiangsu Province or adjacent regions; the stations with the word vectors closest to the word vector of "Suzhou Station" include "Suzhou North Station", "Kunshan Station", etc., and these stations are indeed all within Suzhou City or surrounding areas. This verification method ensures that the word vectors can reasonably express the geographical and semantic correlations between cities and stations, providing effective feature representations for subsequent model training.

[0037] Based on the above embodiments, as an alternative implementation, in S102, determining whether both the city word vectors and the station word vectors meet the preset conditions specifically includes S21 - S24: S21, construct a city word vector distance matrix according to the cosine distance between any two city word vectors; construct a station word vector distance matrix according to the cosine distance between any two station word vectors.

[0038] First, it is necessary to calculate the similarity relationship between word vectors. The cosine distance is used as the similarity metric standard, and the calculation formula for the cosine distance is: cos_distance = 1 - (A·B) / (||A||·||B||), where A and B respectively represent two word vectors, · represents the vector inner product, and ||A|| represents the vector norm. For N cities, construct an N×N city word vector distance matrix, and each element city_dist[i][j] in the matrix represents the cosine distance between city i and city j; construct a station word vector distance matrix station_dist[m][n] in the same way, where m and n respectively represent the numbers of different stations. This matrix structure facilitates quickly querying the semantic distance between any two cities or stations.

[0039] S22, select the associated cities whose city word vector distances from the target city are less than a first preset value according to the city word vector distance matrix; select the associated stations whose station word vector distances from the target station are less than a second preset value according to the station word vector distance matrix.

[0040] Next, based on the constructed distance matrix, identify cities and stations with similar semantics. For each target city, select cities with a word vector distance less than the first preset value (e.g., 0.3) as its associated cities. For example, when analyzing "Suzhou City", cities such as "Wuxi City", "Nantong City", and "Changzhou City" with a word vector distance less than 0.3 may be found as its associated cities. Similarly, for each target station, select stations with a word vector distance less than the second preset value (e.g., 0.25) as its associated stations. The first preset value and the second preset value here can be adjusted according to the actual application scenario. A smaller preset value means a stricter similarity requirement.

[0041] S23, determine whether the associated cities are in the same province or adjacent provinces as the target city, and whether the associated stations are in the same city or adjacent cities as the target station.

[0042] S24, when at least a third preset proportion of the cities among the associated cities are in the same province or adjacent provinces as the target city, and at least a fourth preset proportion of the stations among the associated stations are in the same city or adjacent cities as the target station, determine that the city word vectors and the station word vectors meet the preset conditions.

[0043] To verify the consistency between the semantic relationship and the geographical relationship, it is necessary to check whether these associated relationships conform to the actual geographical distribution. The system determines whether each associated city is in the same province or adjacent provinces as the target city, and whether each associated station is in the same city or adjacent cities as the target station through the geographical location database of cities and stations. For example, for the associated cities of "Suzhou City", check whether they are all in Jiangsu Province or provinces adjacent to Jiangsu Province; for the associated stations of "Suzhou Station", check whether they are all in Suzhou City or adjacent cities.

[0044] Finally, set a reasonable threshold to evaluate the overall consistency. When at least a third preset proportion (e.g., 80%) of the cities among the associated cities are in the same province or adjacent provinces as the target city, and at least a fourth preset proportion (e.g., 75%) of the stations among the associated stations are in the same city or adjacent cities as the target station, it is considered that the word vectors meet the preset conditions. For example, if there are 10 associated cities for "Suzhou City", and at least 8 of them are in Jiangsu Province or adjacent provinces, then the verification requirement for the city word vectors is met; if there are 8 associated stations for "Suzhou Station", and at least 6 of them are in Suzhou City or adjacent cities, then the verification requirement for the station word vectors is met.

[0045] S103, if both the city word vectors and the station word vectors meet the preset conditions, then construct a neural network model, and the neural network model includes a city embedding layer and a station embedding layer.

[0046] After confirming that the city word vectors and station word vectors meet the preset conditions, it is necessary to construct a neural network model that can effectively utilize the information of these word vectors. The core goal of this model is to combine the pre-trained semantic information with the actual selection behavior of users, so as to more accurately predict the purchase tendency of users for transfer plans.

[0047] The basic architecture of the neural network model includes a city embedding layer and a station embedding layer, which are used to process city features and station features respectively. The city embedding layer is responsible for processing city-related features such as the departure city, arrival city, transfer city, and user-preferred transfer city, while the station embedding layer is responsible for processing station-related features such as the departure station, arrival station, transfer station, and user-preferred transfer station. Each embedding layer is actually a trainable feature transformation layer, and its initial weights are determined by the pre-trained word vectors, but can be adjusted according to the actual user behavior data in subsequent training.

[0048] The model also includes multiple cross layers for capturing the interaction relationships between features. The design of the cross layers can automatically learn the high-order combination relationships between features, such as the impact of the geographical relevance between the departure city and the transfer city on the user's choice. Through the multi-layer cross structure, the model can gradually build complex associations between features, so as to better understand the various factor combinations that affect user decisions.

[0049] To handle the synergistic effects of continuous features (such as the total duration of the plan, the total price of the plan, etc.) and discrete features (such as city numbers, station numbers, etc.), the model adopts a parallel feature processing structure. The continuous features are directly input into the cross network after being standardized, while the discrete features are first converted into dense vector representations through the corresponding embedding layers and then input into the cross network. This design enables the model to consider the impacts of different types of features simultaneously.

[0050] At the output end of the cross network, a logistic regression function layer (sigmoid function layer) is set to map the output of the model to the interval [0, 1], representing the predicted purchase probability. The calculation formula of the logistic regression function is: y = 1 / (1 + e^(-z)), where z is the output value of the cross network and y is the final estimated purchase rate. The reason for choosing the logistic regression function is that its output conforms to the domain of probability, and the function has good differentiability, which is convenient for model training.

[0051] The training of the entire neural network model adopts the mini-batch stochastic gradient descent method, and uses the logarithmic loss function (logloss) as the optimization objective. The calculation formula of the logarithmic loss function is: loss = -[y * log(p) + (1 - y) * log(1 - p)], where y is the actual purchase label (0 or 1), and p is the purchase probability predicted by the model. This loss function can effectively measure the difference between the predicted value and the actual value, and guide the update of the model parameters through the backpropagation algorithm.

[0052] Based on the above embodiments, as an optional implementation manner, in S103, constructing the neural network model specifically includes S31 - S33: S31, construct a cross network including multiple cross layers, where the output of each cross layer is obtained by performing feature interaction calculation on the input features.

[0053] First, construct a cross network as the core component of the model, which includes multiple cross layers for feature interaction. Each cross layer can automatically learn the combination relationship between features, and the calculation formula of each layer is as follows: X l1 = X0X l T W l b l X l , X0: represents the initial input feature vector (original feature) of the model; X l : represents the output feature vector of the l-th layer (the output of the previous layer); W l : represents the weight matrix of the l-th layer, used for weight calculation of feature interaction; b l : represents the bias term of the l-th layer, used to control the offset of the feature interaction output; X l1 : represents the output feature vector of the current layer (the l + 1-th layer), X0X l T : represents the outer product of X0 (original feature vector) and X l (the output feature vector of the previous layer). The outer product result is a matrix, representing the interaction relationship between the original feature and the feature of the previous layer. By explicitly constructing feature interaction, the cross network can efficiently learn the complex relationship between features. Therefore, the cross network is selected to help the learning of city and station vectors. For example, the popularity of the "Beijing - Shanghai" route may vary at different times and different prices. By setting multiple cross layers (such as 8 layers), the model can learn more complex feature combination patterns layer by layer. Each layer generates new feature combinations based on the previous layer, enabling the model to understand deeper patterns, such as complex patterns like "economy travelers are more likely to choose certain specific transfer stations during holidays".

[0054] S32. Connect the city embedding layer and the station embedding layer to the input end of the cross network respectively, and set a logistic regression function layer at the output end of the cross network. The logistic regression function layer is used to convert the output value of the cross network into an estimated purchase rate.

[0055] S33. Set the exposure data with purchase behavior as positive sample data, set the exposure data without purchase behavior as negative sample data, and use the logarithmic loss function to train the positive sample data and the negative sample data.

[0056] The input end of the cross network is connected to the city embedding layer and the station embedding layer. When a city number (such as the number of "Suzhou City") is input, the city embedding layer will convert it into a vector containing 100 values, which contains various characteristic information of the city. Similarly, the station number (such as the number of "Suzhou Station") will be converted into the corresponding vector representation after passing through the station embedding layer. This design enables the model to understand the correlation between the city and the station. For example, the system can recognize the geographical connection between "Suzhou Station" and "Wuxi Station".

[0057] Set a logistic regression function layer at the output end of the cross network, whose function is to convert the value output by the cross network into an estimated purchase rate. For example, a certain transfer plan obtains an estimated purchase rate of 0.8 after being calculated by the model, indicating that the model believes that the user has a relatively high possibility of choosing this plan. This conversion ensures that the estimated purchase rate output by the model is always between 0 and 1, facilitating subsequent plan sorting and comparison.

[0058] When preparing the training data, use the user's actual purchase behavior as the learning target of the model. For example, if the user views three transfer plans (Plan A, B, and C) and finally purchases Plan A, then Plan A is used as the positive sample (indicating the plan liked by the user), and Plans B and C are used as negative samples (indicating the plans not liked by the user). This data marking method enables the model to learn the user's true preferences.

[0059] The logarithmic loss function is used in the training process to evaluate the accuracy of the model prediction. When the model's prediction for a certain plan does not match the user's actual choice, the loss function will give a relatively large penalty value, prompting the model to adjust its parameters to improve the prediction accuracy. For example, if the model wrongly gives a high estimated purchase rate to a plan that the user finally does not purchase, the loss function will generate a relatively large penalty, guiding the model to reduce the estimated purchase rate of similar plans in subsequent training.

[0060] S104. Load the city word vector into the city embedding layer, and load the station word vector into the station embedding layer, and set the city embedding layer and the station embedding layer to the trainable state.

[0061] After constructing the basic architecture of the neural network model, it is necessary to reasonably integrate the word vector information obtained from pre-training into the model. The core purpose of this step is to achieve transfer learning of word vectors, that is, to transfer the semantic knowledge learned from general corpora to the specific transfer plan recommendation scenario and allow these knowledge to be adaptively adjusted during the training process.

[0062] First, read the city word vector information from the previously saved city_emd file. This file contains the mapping relationship between city numbers and corresponding word vectors, such as "1_Suzhou City_[0.1261, 0.1475, 0.9130, 0.4347...]". Organize these word vectors into a weight matrix in the order of city numbers, where the number of rows of the matrix is equal to the total number of cities, and the number of columns is equal to the dimension of the word vectors. Process the station word vectors in the station_emd file in the same way to construct the weight matrix of the station word vectors. This organization method ensures that the corresponding word vector representation can be quickly found by city or station number when the model runs.

[0063] Next, load the weight matrix of the city word vectors into the city embedding layer of the neural network model, and load the weight matrix of the station word vectors into the station embedding layer. This loading process is actually using the pre-trained word vectors as the initial weight values of the embedding layer. For example, when the model inputs a city number, the city embedding layer will find the corresponding word vector from the weight matrix as the feature representation of this city. This mechanism enables the model to utilize the semantic relationship information between cities and stations contained in the pre-trained word vectors.

[0064] Set the trainable parameters of the city embedding layer and the station embedding layer to True, which means that the weights (i.e., word vectors) of these layers can be updated during the subsequent model training process. The importance of this setting lies in that although the pre-trained word vectors contain valuable semantic information, this information is learned from general corpora and may have certain differences from the specific scenario of transfer plan recommendation. By setting the embedding layer to be trainable, the model can fine-tune the word vectors according to the actual user behavior data to better adapt to the needs of specific tasks.

[0065] S105, iteratively train the neural network model through the training dataset until the loss function in the neural network model converges to a preset threshold to obtain the target neural network model.

[0066] The training dataset is divided into multiple training batch data according to a preset batch size (e.g., batch_size = 256). Selecting an appropriate batch size has an important impact on the training effect: too small a batch size will lead to unstable training, while too large a batch size will reduce the training efficiency. Each training batch data contains feature data (such as city number, station number, plan duration, plan price, etc.) and corresponding label data (whether to purchase). Before inputting into the model, continuous features (such as plan duration, price) are normalized so that their numerical ranges are unified to the interval [0, 1], which can avoid the unbalanced impact of features with different dimensions on model training.

[0067] In each training iteration, each training batch data is sequentially input into the neural network model. The data is first converted into a dense vector representation through the city embedding layer and the station embedding layer, then undergoes feature interaction through the cross network, and finally the predicted purchase rate is obtained through the logistic regression function layer. The predicted purchase rate is compared with the actual purchase label to calculate the loss function value of the current batch data. The loss function uses the logarithmic loss function (logloss), and its calculation formula is: loss = -[y * log(p) + (1 - y) * log(1 - p)], where y is the actual purchase label and p is the predicted purchase rate.

[0068] According to the calculated loss function value, the backpropagation algorithm is used to calculate the gradients of the parameters of each layer. Backpropagation starts from the output layer and calculates the partial derivatives of the loss function with respect to the parameters of each layer layer by layer. To improve the training efficiency and stability, the Adam optimizer is used for parameter update. The Adam optimizer combines the advantages of the momentum method and the adaptive learning rate, and its parameter update formula is: where θt represents the current parameter, η is the learning rate (e.g., 0.001), mt and vt are the first-order momentum and the second-order momentum respectively, and ε is a small constant to prevent division by zero.

[0069] When updating the parameters, focus on the parameter adjustment of the city embedding layer and the station embedding layer. The parameter updates of these layers need to gradually adapt to the actual user selection behavior while maintaining the original semantic information of the pre-trained word vectors. For this reason, a relatively small learning rate (e.g., 0.1 times the original learning rate) can be used for these layers to make their parameter changes relatively gentle.

[0070] To monitor the training process, calculate the loss value on the validation set every certain number of steps (e.g., 100 steps). When the change amplitude of the validation set loss value for multiple consecutive rounds (e.g., 10 rounds) is less than a preset threshold (e.g., 0.0001), or the validation set loss value starts to rise, it is considered that the model training has reached the convergence state. The neural network model at this time is the target neural network model. This early stopping strategy can effectively prevent model overfitting.

[0071] The entire training process usually requires multiple epochs (e.g., 50 epochs), and each epoch will iterate through the entire training dataset. During the training process, the loss function value of the model usually shows a trend of first rapidly decreasing and then gradually stabilizing. For example, the initial loss value may be around 0.8, and after training, it stabilizes near 0.3. This changing trend of the loss value indicates that the model gradually masters the patterns in the data and can better predict users' purchase behaviors.

[0072] Based on the above embodiments, as an optional implementation manner, in S105, the neural network model is iteratively trained with the training dataset until the loss function in the neural network model converges to a preset threshold to obtain the target neural network model, which specifically includes S51 - S54: S51, divide the training dataset into multiple training batch data according to a preset batch size.

[0073] First, the training data needs to be reasonably organized. The entire training dataset is divided into multiple training batches according to a preset batch size (e.g., batch_size = 256). Selecting an appropriate batch size has an important impact on the training effect: if the batch size is too small (e.g., 32), although the memory usage is less, the training process is prone to instability; if the batch size is too large (e.g., 1024), although the training is more stable, it will reduce the training efficiency and increase the memory requirement. Each training batch contains the complete records of users' queries and selected transfer plans, including features such as the departure location, destination, transfer station, duration, price, etc., as well as the marking information of whether the user makes a purchase.

[0074] S52, sequentially input each training batch data into the neural network model and calculate the current loss function value of the neural network model.

[0075] During the training process, the system sequentially inputs the data of each training batch into the neural network model. The data first converts the city number into a vector representation through the city embedding layer and converts the station number into a vector representation through the station embedding layer, then undergoes feature interaction through the cross network, and finally outputs the estimated purchase rate through the logistic regression function layer. By comparing the estimated purchase rate with the actual purchase marking, the loss function value of the current batch is calculated. This loss function value reflects the accuracy of the model prediction. The larger the loss value, the more inaccurate the prediction. For example, if the model gives a very low estimated purchase rate for a plan that the user actually purchases, a large loss value will be generated.

[0076] S53, according to the current loss function value, use the backpropagation algorithm to sequentially update the parameters in the city embedding layer and the station embedding layer.

[0077] Based on the calculated loss function value, the system updates the model parameters using the backpropagation algorithm. This update process pays particular attention to the parameters of the city embedding layer and the station embedding layer because these parameters determine the quality of the feature representations of cities and stations. The direction and magnitude of parameter updates are determined by the gradient of the loss function, which indicates how to adjust the parameters to reduce the loss value. For example, if the model finds that the feature representation of a certain city leads to a large prediction deviation, it will correspondingly adjust the parameters of that city in the embedding layer. To maintain the stability of training, a relatively small learning rate (such as 0.001) is usually used for parameter updates.

[0078] S54, until the current loss function value is less than the preset threshold, to obtain the target neural network model.

[0079] The training process continuously repeats the loop of "forward calculation - loss calculation - parameter update" until the loss function value drops below the preset threshold (such as 0.1). The selection of this threshold needs to balance the fitting degree and generalization ability of the model: too low a threshold may lead to overfitting, while too high a threshold may cause the model to be underfitted. During the training process, the loss function value usually shows a trend of first rapidly decreasing and then leveling off. For example, initially the loss value may be around 0.8, and after multiple rounds of training, it gradually decreases below 0.1.

[0080] S106, verify the accuracy of the target neural network model through the test data set. If the accuracy is greater than the preset accuracy, input multiple to-be-sorted transfer plans into the target neural network model to obtain the estimated purchase rates of each to-be-sorted transfer plan; sort each to-be-sorted transfer plan according to the estimated purchase rates to generate a sorting result.

[0081] Use the pre-divided test data set to verify the target neural network model. Input the feature data (including total plan duration, total plan price, departure station, arrival station, transfer station, etc.) in the test data set into the model to obtain the estimated purchase rate of each test data. Convert the estimated purchase rate into a binary prediction result: when the estimated purchase rate is greater than 0.5, it is determined as predicted purchase (1), otherwise it is determined as predicted non-purchase (0). By comparing the prediction result with the actual purchase label in the test data set, calculate the accuracy of the model. The calculation formula is: accuracy = (number of correctly predicted samples) / (total number of samples).

[0082] For example, in a scenario with 1000 test data, if the model correctly predicts the purchase situations of 850 data, the accuracy is 85%. Compare the calculated accuracy with the preset accuracy (such as 80%). Only when the actual accuracy exceeds the preset accuracy can the model be considered to have the ability for practical application. This strict verification mechanism ensures that the model put into use has reliable prediction performance.

[0083] When the model passes the accuracy verification, it can be applied to the actual transfer plan sorting task. For the departure and destination of the user's query, the system will generate multiple feasible transfer plans. Each of these transfer plans to be sorted contains complete feature information, such as the total duration of the plan, the total price of the plan, the specific departure station, arrival station, transfer station, etc. Input these feature information into the verified target neural network model in sequence, and the model will output an estimated purchase rate for each transfer plan, with the value range between [0, 1].

[0084] The estimated purchase rate reflects the model's prediction of the possibility of the user choosing this plan. For example, a certain query generates three transfer plans: the estimated purchase rate of Plan A is 0.85, Plan B is 0.62, and Plan C is 0.43. These values comprehensively consider the impact of the various features of the plan and their combinations on the user's decision-making. The larger the value, the higher the possibility that the user will choose this plan.

[0085] According to the obtained estimated purchase rates, the system sorts all the transfer plans to be sorted in descending order to generate the final sorting result. In the above example, the sorting result will be: Plan A > Plan B > Plan C. This sorting method ensures that the plans with higher estimated purchase rates will be displayed to the user first, thus improving the efficiency of the user finding a satisfactory plan.

[0086] Based on the above embodiments, as an alternative implementation, in S106, verifying the accuracy of the target neural network model through the test dataset specifically includes S61 - S64: S61, divide the test dataset into multiple verification batch data.

[0087] To efficiently verify the model, it is necessary to reasonably organize the test data. Divide the test dataset into multiple verification batch data according to an appropriate batch size (such as 512 records in one batch). This batch division method can not only ensure the computational efficiency of the verification process but also ensure reasonable memory usage. Each verification batch contains complete transfer plan information, including features such as the departure place, destination, transfer stations, price, duration, etc., as well as the marking information of whether the user actually purchases. For example, a dataset containing 10,000 test data can be divided into 20 verification batches, with each batch containing 500 records.

[0088] S62, input each verification batch data into the target neural network model in sequence to obtain the prediction results of each verification batch data.

[0089] The verification process adopts a batch-by-batch processing method, and inputs the data of each verification batch into the target neural network model in turn. The model processes each piece of data: first, it obtains the feature representation through the city embedding layer and the station embedding layer, then conducts feature interaction through the cross network, and finally outputs the estimated purchase rate through the logistic regression function layer. Samples with an estimated purchase rate greater than 0.5 are predicted as "will purchase", and samples less than or equal to 0.5 are predicted as "will not purchase". For example, for a certain transfer plan, if the estimated purchase rate output by the model is 0.7, the prediction result is "will purchase"; if the estimated purchase rate is 0.3, the prediction result is "will not purchase".

[0090] S63. Compare the prediction results of each verification batch of data with the labeled data corresponding to each verification batch of data, and count the number of samples with correct predictions.

[0091] For each verification batch, compare the prediction results of the model with the actual purchase labels, and count the number of samples with correct predictions. Correct predictions include two cases: one is that the model predicts "will purchase" and the user actually does purchase; the other is that the model predicts "will not purchase" and the user actually does not purchase. For example, in a verification batch containing 500 records, if the prediction results of 420 records match the actual situation, the number of correct predictions for this batch is 420.

[0092] S64. Obtain the accuracy rate of the target neural network model according to the ratio of the number of samples with correct predictions to the total number of samples in the test data set.

[0093] After completing the predictions for all verification batches, add up the number of samples with correct predictions in all batches to obtain the total number of samples with correct predictions. Divide this number by the total number of samples in the test data set to obtain the accuracy rate of the model. For example, if 8500 out of 10000 test data are predicted correctly, the accuracy rate of the model is 85%. This accuracy rate reflects the prediction ability of the model in the actual application scenario.

[0094] Based on the above embodiments, as an alternative implementation, in S106, inputting multiple transfer plans to be sorted into the target neural network model to obtain the estimated purchase rates of each transfer plan to be sorted specifically includes S71 - S73: S71. Obtain the city features and station features in each transfer plan to be sorted, input the city features into the city embedding layer, and input the station features into the station embedding layer.

[0095] First, it is necessary to extract the feature information of the transfer plans to be sorted. Each transfer plan contains city features (such as departure city, arrival city, transfer city) and station features (such as departure station, arrival station, transfer station). These features are input into the model in the form of numbers. For example, a transfer plan may contain city features of "Beijing (number 101) - Shijiazhuang (number 102) - Shanghai (number 103)", and station features of "Beijing Station (number 1001) - Shijiazhuang Station (number 1002) - Shanghai Station (number 1003)". These features are respectively input into the city embedding layer and the station embedding layer and converted into dense vector representations. This conversion enables the model to understand the semantic relationships between cities and stations. For example, the model can recognize that cities or stations with similar geographical locations may have similar feature representations.

[0096] S72, Input the output features of the city embedding layer and the station embedding layer into the cross network, and perform feature interaction calculations through each cross layer to obtain the feature vectors after interaction.

[0097] Next, input the feature vectors output by the city embedding layer and the station embedding layer into the cross network for deep feature interaction. The cross network constructs the combined relationship of features layer by layer through multiple cross layers (such as 8 layers). Each layer can discover new feature interaction patterns. For example, the first layer may learn simple "city - station" combined relationships, and deeper layers may discover complex combined relationships such as "early morning train - transfer station - price". This layer-by-layer deep feature interaction enables the model to fully understand various factor combinations that affect user choices.

[0098] S73, Input the feature vectors after interaction into the logistic regression function layer to obtain the estimated purchase rates corresponding to each transfer plan to be sorted.

[0099] Finally, input the feature vectors after interaction into the logistic regression function layer to calculate the estimated purchase rate of each transfer plan. The estimated purchase rate is a value between 0 and 1, which reflects the possibility of a user choosing this plan. For example, if there are three transfer plans A, B, and C to be sorted, the estimated purchase rates obtained after model calculation may be 0.8, 0.6, and 0.3 respectively. This indicates that plan A is most likely to be chosen by the user, while the possibility of choosing plan C is lower.

[0100] Sort each transfer plan to be sorted according to the estimated purchase rate. After generating the sorting result, it also includes: Calculate the actual purchase rate of the transfer options to be sorted based on the historical display data and historical purchase data of each transfer option to be sorted; calculate the difference between the estimated purchase rate and the actual purchase rate of each transfer option to be sorted; obtain the supplementary feature data and supplementary label data of the transfer options whose differences are greater than the preset deviation value; add the supplementary feature data and supplementary label data to the training dataset and retrain the neural network model.

[0101] In one example, the system first needs to collect and count the historical data of each transfer option. For each transfer option, the number of times it is shown to users (historical display data) and the number of times users actually purchase it (historical purchase data) are recorded. For example, if a certain transfer option is shown 1000 times in the past week and 80 of them are purchased by users, then the actual purchase rate of this option is 8%. This kind of statistics based on historical data provides an objective reflection of users' real choice behavior.

[0102] Next, the system compares the purchase rate estimated by the model with the actual purchase rate. For example, if the estimated purchase rate of a certain transfer option by the model is 15% while the actual purchase rate is only 5%, then the difference is 10%. This difference reflects the accuracy of the model prediction. Usually, a reasonable preset deviation value (such as 5%) is set as the judgment criterion. When the difference between the estimated value and the actual value exceeds this threshold, it indicates that there is an obvious deviation in the model's prediction of this type of option and optimization and adjustment are needed.

[0103] For those transfer options with large prediction deviations, the system will collect more relevant feature data and label data as supplementary training data. The supplementary feature data includes the complete information of these options, such as the departure place, destination, transfer station, price, time, etc.; the supplementary label data is the actual purchase behavior record of users for these options. For example, if it is found that the model has a large deviation in predicting the transfer options of some popular routes during holidays, relevant data of this type of option will be focused on collecting.

[0104] Add these supplementary data to the original training dataset, and then retrain the neural network model. The retraining process uses the same method as the initial training, and feature learning and interaction are carried out through the city embedding layer, the station embedding layer, and the cross network. This incremental learning method enables the model to continuously adapt to new data patterns and changes in user preferences. For example, if the transfer conditions of a certain transfer station have been recently improved, resulting in a significant increase in its actual purchase rate, the model can adjust the estimated purchase rate of the options for this station in a timely manner through retraining.

[0105] Based on the above method, the present application also discloses a system for travel transfer sorting using word vectors and cross networks, as Figure 2 shown Figure 2It is a schematic structural diagram of a system for sorting travel transfers using word vectors and cross networks provided by an embodiment of the present application. The system includes: a collection module, a judgment module, a construction module, a loading module, a training module, and an output module. Among them, the collection module is used to collect feature data of transfer plans and labeled data corresponding to the feature data, and divide the feature data and labeled data into a training data set and a test data set according to a preset ratio. The judgment module is used to obtain the city word vector and the station word vector in the feature data, and judge whether both the city word vector and the station word vector meet the preset conditions. The construction module is used to construct a neural network model if both the city word vector and the station word vector meet the preset conditions. The neural network model includes a city embedding layer and a station embedding layer. The loading module is used to load the city word vector into the city embedding layer, and load the station word vector into the station embedding layer, and set the city embedding layer and the station embedding layer to a trainable state. The training module is used to iteratively train the neural network model through the training data set until the loss function in the neural network model converges to a preset threshold to obtain a target neural network model. The output module is used to verify the accuracy rate of the target neural network model through the test data set. If the accuracy rate is greater than the preset accuracy rate, input multiple transfer plans to be sorted into the target neural network model to obtain the estimated purchase rate of each transfer plan to be sorted; sort each transfer plan to be sorted according to the estimated purchase rate to generate a sorting result.

[0106] It should be noted that when the system provided in the above embodiment realizes its functions, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0107] Please refer to Figure 3 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 3 shown, the electronic device 1000 may include: at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.

[0108] Among them, the communication bus 1002 is used to realize the connection and communication between these components.

[0109] Among them, the user interface 1003 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface.

[0110] Among them, the network interface 1004 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface).

[0111] Among them, the processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the entire server using various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling data stored in the memory 1005, it performs various functions of the server and processes data. Optionally, the processor 1001 may be implemented in at least one of the following hardware forms: Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 1001 may integrate a combination of one or several of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 1001 and can be implemented separately through a single chip.

[0112] Among them, the memory 1005 may include Random Access Memory (RAM), and may also include Read-Only Memory. Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store the data involved in the above-mentioned various method embodiments. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As Figure 3 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a method of applying word vectors and a cross network to perform travel transfer sorting.

[0113] In Figure 3 In the electronic device 1000 shown, the user interface 1003 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 1001 can be used to call an application program stored in the memory 1005 for a method of using an application word vector and a cross network to perform travel transfer sorting. When executed by one or more processors, the electronic device is caused to execute one or more of the methods as described in the above embodiments.

[0114] An electronic device-readable storage medium stores instructions. When executed by one or more processors, the electronic device is caused to execute one or more of the methods as described in the above embodiments.

[0115] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be adopted in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0116] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0117] In the several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some service interfaces. The indirect couplings or communication connections of the devices or units can be in an electrical or other form.

[0118] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0119] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0120] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.

[0121] The foregoing are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and practicing the present disclosure herein, those skilled in the art will readily think of other implementation manners of the present disclosure. The present application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A method for travel transfer sorting using word vectors and cross networks, characterized in that The method includes: Collecting feature data of the transfer plan and the corresponding labeled data, and dividing the feature data and the labeled data into a training data set and a test data set according to a preset ratio; Obtaining the city word vector and the station word vector in the feature data, and determining whether both the city word vector and the station word vector meet the preset conditions; If both the city word vector and the station word vector meet the preset conditions, constructing a neural network model, where the neural network model includes a city embedding layer and a station embedding layer; Loading the city word vector into the city embedding layer, and loading the station word vector into the station embedding layer, and setting the city embedding layer and the station embedding layer to be in a trainable state; Iteratively training the neural network model with the training data set until the loss function in the neural network model converges to a preset threshold to obtain a target neural network model; Verifying the accuracy rate of the target neural network model through the test data set. If the accuracy rate is greater than the preset accuracy rate, inputting multiple transfer plans to be sorted into the target neural network model to obtain the estimated purchase rates of the transfer plans to be sorted; sorting the transfer plans to be sorted according to the estimated purchase rates to generate a sorting result.

2. The method for sorting travel transfers using word vectors and a cross network according to claim 1, characterized in that The determination of whether both the city word vector and the station word vector meet the preset conditions includes: Constructing a city word vector distance matrix according to the cosine distance between any two city word vectors; constructing a station word vector distance matrix according to the cosine distance between any two station word vectors; Selecting associated cities with a city word vector distance less than a first preset value from the target city according to the city word vector distance matrix; selecting associated stations with a station word vector distance less than a second preset value from the target station according to the station word vector distance matrix; Determining whether the associated cities are in the same province or adjacent provinces as the target city, and whether the associated stations are in the same city or adjacent cities as the target station; When at least a third preset ratio of the associated cities are in the same province or adjacent provinces as the target city, and at least a fourth preset ratio of the associated stations are in the same city or adjacent cities as the target station, it is determined that the city word vector and the station word vector meet the preset conditions.

3. The method for sorting travel transfers using word vectors and a cross network according to claim 1, wherein The construction of the neural network model includes: Constructing a cross network including multiple cross layers, where the output of each cross layer is obtained by performing feature interaction calculation on the input features; Connecting the city embedding layer and the station embedding layer to the input end of the cross network respectively, and setting a logistic regression function layer at the output end of the cross network, where the logistic regression function layer is used to convert the output value of the cross network into the estimated purchase rate; Setting the exposure data with purchase behavior as positive sample data, setting the exposure data without purchase behavior as negative sample data, and training the positive sample data and the negative sample data using a logarithmic loss function.

4. The method for sorting travel transfers using word vectors and a cross network according to claim 1, characterized in that, Iteratively training the neural network model with the training dataset until the loss function in the neural network model converges to a preset threshold to obtain a target neural network model, including: Dividing the training dataset into multiple training batch data according to a preset batch size; Sequentially inputting each of the training batch data into the neural network model and calculating the current loss function value of the neural network model; According to the current loss function value, using the backpropagation algorithm to sequentially update the parameters in the city embedding layer and the station embedding layer; Until the current loss function value is less than the preset threshold to obtain a target neural network model.

5. The method for sorting travel transfers using word vectors and a cross network according to claim 1, wherein Verifying the accuracy of the target neural network model with the test dataset, including: Dividing the test dataset into multiple validation batch data; Sequentially inputting each of the validation batch data into the target neural network model to obtain the prediction results of each of the validation batch data; Comparing the prediction results of each of the validation batch data with the labeled data corresponding to each of the validation batch data and counting the number of correctly predicted samples; Obtaining the accuracy of the target neural network model according to the ratio of the number of correctly predicted samples to the total number of samples in the test dataset.

6. The method for sorting travel transfers using word vectors and a cross network according to claim 3, wherein Inputting multiple transfer schemes to be sorted into the target neural network model to obtain the estimated purchase rates of each of the transfer schemes to be sorted, including: Obtaining the city features and station features in each of the transfer schemes to be sorted, inputting the city features into the city embedding layer, and inputting the station features into the station embedding layer; Inputting the output features of the city embedding layer and the station embedding layer into the cross network, performing feature interaction calculations through each of the cross layers to obtain an interaction feature vector; Inputting the interaction feature vector into the logistic regression function layer to obtain the estimated purchase rates corresponding to each of the transfer schemes to be sorted.

7. The method for sorting travel transfers using word vectors and a cross network according to claim 1, characterized in that After sorting each of the transfer schemes to be sorted according to the estimated purchase rate to generate a sorting result, further including: Calculating the actual purchase rate of each of the transfer schemes to be sorted according to the historical display data and historical purchase data of each of the transfer schemes to be sorted; Calculating the difference between the estimated purchase rate and the actual purchase rate of each of the transfer schemes to be sorted; Obtaining the supplementary feature data and supplementary labeled data of the transfer schemes whose difference is greater than a preset deviation value; Adding the supplementary feature data and the supplementary labeled data to the training dataset to retrain the neural network model.

8. A system for sorting travel transfers using word vectors and cross networks, characterized in that, The system includes: a collection module, a judgment module, a construction module, a loading module, a training module, and an output module; wherein, The collection module is used to collect the feature data of the transfer scheme and the labeled data corresponding to the feature data, and divide the feature data and the labeled data into a training dataset and a test dataset according to a preset ratio; The judgment module is used to obtain the city word vector and the station word vector in the feature data and judge whether both the city word vector and the station word vector meet the preset conditions; The building module is configured to build a neural network model if both the city word vector and the station word vector meet the preset conditions. The neural network model includes a city embedding layer and a station embedding layer; The loading module is configured to load the city word vector into the city embedding layer, load the station word vector into the station embedding layer, and set the city embedding layer and the station embedding layer to a trainable state; The training module is configured to iteratively train the neural network model through the training data set until the loss function in the neural network model converges to a preset threshold to obtain a target neural network model; The output module is configured to verify the accuracy rate of the target neural network model through the test data set. If the accuracy rate is greater than the preset accuracy rate, input multiple to-be-sorted transfer plans into the target neural network model to obtain the estimated purchase rates of the to-be-sorted transfer plans; sort the to-be-sorted transfer plans according to the estimated purchase rates to generate a sorting result.

9. An electronic device, characterized in that, It includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program is stored that can be loaded and executed by a processor to execute the method according to any one of claims 1-7.