Travel Location Selection Method and System Based on Multimodal Fusion
By adopting a multimodal fusion deep neural network model in the travel site selection system, integrating site-scale data, popularity POI data and traffic timing data, the problem of low accuracy in travel site selection decisions in the existing technology is solved, and higher decision accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202411930369.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The existing travel site selection system relies on a single data source or a simple statistical model, resulting in low decision accuracy and reliability, and cannot effectively solve the problem of low travel site selection decision accuracy.
The travel site selection method based on multimodal fusion is adopted, and the site scale data, the heat POI data and traffic timing data are obtained, and the pre-constructed multimodal fusion deep neural network model is used for feature extraction and feature analysis, and the features of different data types are fused to generate the recommended values of candidate points.
By comprehensively considering the location scale attributes, popularity POI data and traffic timing data, one-sided decision-making basis caused by a single data type is avoided, and the accuracy and reliability of site selection decisions are improved.
Smart Images

Figure CN119358987B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data analysis, and particularly to a travel location selection method and system based on multimodal fusion. Background Art
[0002] With the acceleration of urbanization and the diversification of travel demands, reasonably planning the location selection of travel service points has become an important link in improving user experience and optimizing resource allocation.
[0003] Existing location selection systems usually rely on a single data source or adopt simple statistical models. A single data source usually has difficulty providing sufficiently diverse information and cannot comprehensively capture various potential variables, resulting in one-sided decision-making basis. Simple statistical models usually cannot capture the correlations between data, resulting in low accuracy and reliability of location selection decisions.
[0004] Currently, no effective solution has been proposed for the problem of low accuracy of travel location selection decisions in related technologies. Summary of the Invention
[0005] Embodiments of this application provide a travel location selection method, system, electronic device, and storage medium based on multimodal fusion to at least solve the problem of low accuracy of travel location selection decisions in related technologies.
[0006] In a first aspect, embodiments of this application provide a travel location selection method based on multimodal fusion. The method includes:
[0007] Obtain the venue scale data, popularity POI data, and traffic time series data of candidate points;
[0008] Through a pre-constructed multimodal fusion deep neural network model, perform feature extraction and feature analysis on the venue scale data, the popularity POI data, and the traffic time series data respectively to obtain scale features, POI features, and time series features, as well as cross-modal feature interaction relationships;
[0009] Through the multimodal fusion deep neural network model, based on the cross-modal feature interaction relationships, fuse the scale features, the POI features, and the time series features, and generate a recommendation value for the candidate points based on the fusion result.
[0010] In some embodiments, the multimodal fusion deep neural network model includes a data preprocessing layer, a feature fusion layer, and a decision layer;
[0011] The data preprocessing layer performs feature extraction on the venue scale data, the popularity POI data, and the traffic time series data to obtain scale features, POI features, and time series features, and respectively obtains the feature interaction relationships of the scale features, the POI features, and the time series features;
[0012] The feature fusion layer obtains a cross-modal feature interaction relationship according to the feature interaction relationships of the scale feature, the POI feature, and the temporal feature, and fuses the scale feature, the POI feature, and the temporal feature based on the cross-modal feature interaction relationship to obtain a fused feature;
[0013] The decision-making layer generates a recommendation value for the candidate point based on the fused feature.
[0014] In some embodiments, obtaining the feature interaction relationship of the scale feature includes:
[0015] Screening the scale feature through a gradient boosting tree to obtain first feature data;
[0016] Inputting the first feature data into a multi-layer perceptron to obtain the feature interaction relationship of the scale feature.
[0017] In some embodiments, obtaining the feature interaction relationship of the POI feature includes:
[0018] Modeling the spatial relationship of the POI feature through a graph neural network to obtain a spatial relationship model;
[0019] Embedding the node structure information of the spatial relationship model into the vector representation of the POI feature to obtain second feature data;
[0020] Inputting the second feature data into a multi-layer perceptron to obtain the feature interaction relationship of the POI feature.
[0021] In some embodiments, obtaining the feature interaction relationship of the temporal feature includes:
[0022] Constructing the global time dependence relationship of the temporal feature through a Transformer network;
[0023] Capturing the long-term trend according to the global time dependence relationship through a long short-term memory network;
[0024] Obtaining the feature interaction relationship of the temporal feature based on the long-term trend.
[0025] In some embodiments, the method further includes:
[0026] Constructing the multi-modal fusion deep neural network model based on the binary cross-entropy loss function;
[0027] Training the multi-modal fusion deep neural network model through a data augmentation strategy, where the multi-modal fusion deep neural network model is a binary classification model;
[0028] The performance of the trained multimodal fusion deep neural network model was evaluated by accuracy, recall and F1 score.
[0029] In some embodiments, the method further comprises:
[0030] Performing weighted calculation on the scale feature, the POI feature and the time series feature to obtain a comprehensive score for each of the candidate points;
[0031] Based on the preset grading rules and the comprehensive score, determine the recommended star rating corresponding to each of the candidate points;
[0032] A point recommendation list is generated according to the recommendation star ratings and the recommendation values.
[0033] In some embodiments, before extracting features from the venue scale data, the popular POI data, and the traffic time series data, the method further includes:
[0034] By forward filling and backward filling, missing values of the traffic time series data are supplemented;
[0035] Supplement the missing values of the venue scale data and the popular POI data by filling in the mean or the category mode;
[0036] Based on the quantile detection method, extreme values in the venue scale data, the popular POI data and the traffic time series data are eliminated.
[0037] In a second aspect, an embodiment of the present application provides a travel site selection system based on multimodal fusion, the system comprising:
[0038] The data acquisition module is used to obtain the venue scale data, popular POI data and traffic time series data of the candidate points;
[0039] A feature extraction module is used to extract and analyze the features of the venue scale data, the popular POI data and the traffic time series data respectively through a pre-built multi-modal fusion deep neural network model to obtain scale features, POI features and time series features, as well as cross-modal feature interaction relationships;
[0040] The recommendation module is used to fuse the scale feature, the POI feature and the time series feature through the multimodal fusion deep neural network model according to the cross-modal feature interaction relationship, and generate a recommendation value for the candidate point based on the fusion result.
[0041] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the travel location selection method based on multimodal fusion as described in the first aspect above.
[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the travel location selection method based on multimodal fusion as described in the first aspect above.
[0043] Compared with the related art, the travel location selection method based on multimodal fusion provided by the embodiment of the present application obtains the venue scale data, popularity POI data, and traffic time series data of candidate locations, and respectively performs feature extraction and feature analysis on the venue scale data, popularity POI data, and traffic time series data through a pre-constructed multimodal fusion deep neural network model to obtain scale features, POI features, and time series features, as well as cross-modal feature interaction relationships. According to the cross-modal feature interaction relationships, the scale features, POI features, and time series features are fused, and a recommendation value for the candidate location is generated based on the fusion result, solving the problem of low accuracy in travel location selection decisions. By comprehensively considering the venue scale attributes, popularity POI data, and traffic time series data, it avoids one-sided decision-making basis caused by a single data type. At the same time, a deep neural network is used to fuse multimodal data to analyze the deep-level associations between data, improving the accuracy of location selection decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0045] Figure 1 is a flowchart of a travel location selection method based on multimodal fusion according to an embodiment of the present application;
[0046] Figure 2 is a schematic structural diagram of a multimodal fusion deep neural network model according to an embodiment of the present application;
[0047] Figure 3 is a flowchart of a method for training a multimodal fusion deep neural network model according to an embodiment of the present application;
[0048] Figure 4 is a structural block diagram of a travel location selection system based on multimodal fusion according to an embodiment of the present application;
[0049] Figure 5 is a schematic internal structure diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0050] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be described and explained below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without creative efforts fall within the scope of protection of the present application.
[0051] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar scenarios based on these drawings without creative efforts. In addition, it can also be understood that although the efforts made in such a development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as the content disclosed in the present application being insufficient.
[0052] Referring to "embodiments" in the present application means that the specific features, structures or characteristics described in combination with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0053] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "an", "one", "the" and the like involved in this application do not indicate a quantity limitation and may represent a singular or plural number. The terms "comprising", "including", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may further include steps or units not listed, or may further include other steps or units inherent to these processes, methods, products or devices. The similar words such as "connected", "coupled" and "linked" involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The term "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0054] This embodiment provides a travel location selection method based on multimodal fusion. Figure 2 It is a flowchart of the travel location selection method based on multimodal fusion according to the embodiment of the present application. As Figure 2 shown, the process includes the following steps:
[0055] Step S101, obtain the venue scale data, popularity POI data and traffic time series data of the candidate points.
[0056] Optionally, the venue scale data is collected through government open data, including but not limited to the venue area, function type (such as residential area, shopping mall, office area), and floor height. The popularity POI data is collected from the data development platform, including but not limited to the number of people, access frequency, stay time and distribution of public facilities, which can reflect the regional economic vitality and potential user needs. The traffic time series data is obtained by processing the charging pile data and cabinet data of the network platform, including but not limited to the average daily battery replacement order number and the number of passing riders within a preset number of days (such as 3 days).
[0057] The venue scale data reflects the type of the candidate points; the popularity POI data focuses on the popularity of the candidate points and the activity level in a specific time period, reflecting the popularity distribution of different interest points; the traffic time series data focuses on analyzing the change of traffic over time, such as the traffic fluctuation of a certain road, business district or website, to identify trends or make predictions.
[0058] By analyzing three types of data, namely venue scale data, popularity POI data, and traffic time-series data, it avoids one-sided decision-making based on a single data type and improves the accuracy of the analysis results.
[0059] Step S102: Through a pre-constructed multi-modal fusion deep neural network model, perform feature extraction and feature analysis on the venue scale data, popularity POI data, and traffic time-series data respectively to obtain scale features, POI features, and time-series features, as well as cross-modal feature interaction relationships.
[0060] Process and extract features from the collected venue scale data, popularity POI data, and traffic time-series data. In terms of feature processing, standardize the input data and compress the numerical range to the interval [0,1] to reduce the scale difference between features and avoid affecting the learning effect of the model.
[0061] Step S103: Through the multi-modal fusion deep neural network model, based on the cross-modal feature interaction relationships, fuse the scale features, POI features, and time-series features, and generate a recommended value for the candidate location based on the fusion result.
[0062] In some embodiments, the multi-modal fusion deep neural network model includes a data preprocessing layer, a feature fusion layer, and a decision-making layer.
[0063] The data preprocessing layer performs feature extraction on the venue scale data, popularity POI data, and traffic time-series data to obtain scale features, POI features, and time-series features, and respectively obtains the feature interaction relationships of the scale features, POI features, and time-series features.
[0064] The feature fusion layer, based on the feature interaction relationships of the scale features, POI features, and time-series features, obtains cross-modal feature interaction relationships, and fuses the scale features, POI features, and time-series features based on the cross-modal feature interaction relationships to obtain fused features.
[0065] The decision-making layer generates a recommended value for the candidate location based on the fused features.
[0066] Figure 2 is a schematic diagram of the architecture of a multi-modal fusion deep neural network model according to an embodiment of the present application. As Figure 2 shown, the multi-modal fusion deep neural network model includes a data preprocessing layer, a feature fusion layer, a decision-making layer, and a prediction output layer.
[0067] This embodiment improves the deep neural network, extracts and fuses multi-modal data, fully explores the deep-level associations between data, and improves the accuracy of location selection prediction.
[0068] In some of these embodiments, obtaining the feature interaction relationship of scale features includes:
[0069] Step S201: Screen the scale features through a gradient boosting tree to obtain first feature data.
[0070] Step S202: Input the first feature data into a multi-layer perceptron to obtain the feature interaction relationship of scale features.
[0071] Preferably, key features are screened out by XGBoost to reduce redundancy, and then the refined feature vector is input into the MLP network to model the non-linear feature interaction relationship.
[0072] In some of these embodiments, obtaining the feature interaction relationship of POI features includes:
[0073] Step S301: Model the spatial relationship of POI features through a graph neural network to obtain a spatial relationship model.
[0074] Step S302: Embed the node structure information of the spatial relationship model into the vector representation of POI features to obtain second feature data.
[0075] Step S303: Input the second feature data into a multi-layer perceptron to obtain the feature interaction relationship of POI features.
[0076] Preferably, model the spatial relationship of POI features through a graph neural network (GNN), embed the node structure information into the vector representation, and further learn complex feature patterns through a multi-layer perceptron (MLP) to obtain the feature interaction relationship of POI features.
[0077] In some of these embodiments, obtaining the feature interaction relationship of temporal features includes:
[0078] Step S401: Build the global time dependence of temporal features through a Transformer network.
[0079] Step S402: Capture the long-term trend according to the global time dependence through a long short-term memory network.
[0080] Step S403: Obtain the feature interaction relationship of temporal features based on the long-term trend.
[0081] In this embodiment, the global time dependence is modeled through a Transformer network, and then the long short-term memory network (LSTM) is used to capture the long-term trend to obtain a complete temporal representation.
[0082] After obtaining the cross-modal feature interaction relationships of the scale feature, POI feature, and temporal feature, the feature fusion layer uses a Transformer network to model the interaction relationships between cross-modal features, and through the self-attention mechanism, selectively focuses on the relevant features between different modalities while suppressing irrelevant noise features. The fused feature vector is passed to the decision layer, and the binary classification task is completed through an MLP, and the recommended value of the candidate location is output. Preferably, the recommended value includes 0 and 1, where 0 indicates non-recommendation and 1 indicates recommendation.
[0083] The design of the model fully considers the heterogeneity and correlation of the data to ensure the accuracy and reliability of the final output result.
[0084] Through the above steps S101 to S103, the venue scale data, popularity POI data, and traffic temporal data of the candidate location are obtained. Through the pre-constructed multi-modal fusion deep neural network model, the feature extraction and feature analysis are respectively performed on the venue scale data, popularity POI data, and traffic temporal data to obtain the scale feature, POI feature, and temporal feature, as well as the cross-modal feature interaction relationship. According to the cross-modal feature interaction relationship, the scale feature, POI feature, and temporal feature are fused, and the recommended value of the candidate location is generated based on the fusion result, solving the problem of low accuracy of travel location decision-making. Considering the venue scale attribute, popularity POI data, and traffic temporal data comprehensively, it avoids the one-sidedness of the decision-making basis caused by a single data type. At the same time, a deep neural network is used to fuse multi-modal data to analyze the deep-level association between data, improving the accuracy of the location decision-making.
[0085] In some of these embodiments, the method further includes:
[0086] Step S501, constructing a multi-modal fusion deep neural network model based on the binary cross-entropy loss function.
[0087] Step S502, training the multi-modal fusion deep neural network model through a data augmentation strategy. The multi-modal fusion deep neural network model is a binary classification model;
[0088] Step S503, evaluating the performance of the trained multi-modal fusion deep neural network model through accuracy, recall rate, and F1 score.
[0089] The model uses the binary cross-entropy loss function (Binary Cross Entropy, BCE) to measure the difference between the predicted value and the true label, and its expression is:
[0090]
[0091] where y i represents the sample category, and its value is 0 or 1; p iIt represents the probability of predicting the category of the sample. The ultimate goal of the model is to minimize this loss function.
[0092] Optionally, the model optimizer adopts the Adam algorithm, and the learning rate is set to 0.0002 to ensure the stability of gradient update and the training convergence speed. The training dataset is divided into a training set and a validation set in the ratio of 8:2. During the training process, the accuracy of the validation set is monitored, and the model parameters are dynamically adjusted. To improve the generalization ability of the model, data augmentation strategies are used, including but not limited to randomly discarding some features and adding Gaussian noise to time series features. The number of model training epochs is set to 1000 times, and the model with the best performance on the validation set during the training process is saved as the final output. The performance of the model is evaluated by metrics such as accuracy, recall rate, and F1 score to ensure its good prediction ability and generalization. Figure 3 It is a flowchart of a method for training a multi-modal fusion deep neural network model according to an embodiment of the present application.
[0093] When the model is deployed for service, the trained model weights are encapsulated through the microservice framework Flask. The service receives the input request address. For each point, the network predicts the probability value of its suitability for setting up a target service (such as a venue car rental point), and determines the final result (0 or 1) through threshold judgment.
[0094] After the model is deployed, operational data can also be collected through a continuous dynamic feedback mechanism, and the actual effect is used as the input of a new round of training data to gradually improve the prediction accuracy and adaptability of the model.
[0095] In some of these embodiments, the method further includes:
[0096] Step S601, perform weighted calculation on the scale feature, POI feature, and time series feature to obtain the comprehensive score of each candidate point.
[0097] Step S602, based on a preset grading rule, determine the recommended star rating corresponding to each candidate point according to the comprehensive score.
[0098] Step S603, generate a point recommendation list according to each recommended star rating and recommendation value.
[0099] To further improve the decision-making quality, the comprehensive score is calculated in combination with the prediction result. The scoring formula is the weighted sum of the scores of each modal feature. Optionally, the comprehensive score range is [0, 100]. The points are graded according to the score. Points with a score of 90 - 100 are 3-star priority recommendation points, points with a score of 75 - 90 are 2-star secondary priority recommendation points, points with a score of 60 - 75 are 1-star ordinary recommendation points, and points with a score below 60 are not recommended. Finally, combined with the recommendation list output by the system, it assists in optimizing resource allocation for operation and maintenance to achieve precise site selection.
[0100] In some embodiments, before extracting features from the venue scale data, the popular POI data, and the traffic time series data, the method further includes:
[0101] Step S701, supplementing the missing values of the traffic time series data by forward filling and backward filling;
[0102] Step S702, supplementing the missing values of the venue scale data and the popular POI data by filling in the mean value or the category mode;
[0103] Step S703: based on the quantile detection method, remove extreme values in the venue scale data, popular POI data and traffic time series data.
[0104] After obtaining the site scale data, popular POI data and traffic time series data of the candidate points, these data need to be processed to ensure the quality and consistency of the input data.
[0105] Processing of missing values: For time series data, a combination of forward filling and backward filling is used to ensure the integrity of key time series features; for non-time series data, mean filling or category mode filling is used.
[0106] Detect and correct outliers: Use quantile-based detection methods to remove extreme values to improve the robustness of the model.
[0107] After cleaning, all data are formatted uniformly.
[0108] The above method uses multimodal data fusion and deep neural network models to integrate venue scale attributes, popular POI data and traffic time series data to achieve accurate and efficient travel service point layout. The multimodal fusion deep neural network model can not only improve the scientificity and accuracy of site selection prediction and optimize resource allocation, but also reduce the operating costs of enterprises while taking into account the user experience.
[0109] It should be noted that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0110] This embodiment also provides a travel location selection system based on multimodal fusion. This system is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0111] Figure 4 is a structural block diagram of a travel location selection system based on multimodal fusion according to an embodiment of the present application. As Figure 4 shown, the system includes:
[0112] A data acquisition module 41, configured to acquire venue scale data, popularity POI data, and traffic time series data of candidate locations.
[0113] A feature extraction module 42, configured to perform feature extraction and feature analysis on the venue scale data, popularity POI data, and traffic time series data respectively through a pre-constructed multimodal fusion deep neural network model, to obtain scale features, POI features, and time series features, as well as cross-modal feature interaction relationships.
[0114] A recommendation module 43, configured to fuse the scale features, POI features, and time series features according to the cross-modal feature interaction relationships through a multimodal fusion deep neural network model, and generate a recommendation value for the candidate location based on the fusion result.
[0115] In some embodiments, the multimodal fusion deep neural network model includes a data preprocessing layer, a feature fusion layer, and a decision layer.
[0116] The data preprocessing layer performs feature extraction on the venue scale data, popularity POI data, and traffic time series data to obtain scale features, POI features, and time series features, and respectively obtains the feature interaction relationships of the scale features, POI features, and time series features.
[0117] The feature fusion layer obtains cross-modal feature interaction relationships according to the feature interaction relationships of the scale features, POI features, and time series features, and fuses the scale features, POI features, and time series features based on the cross-modal feature interaction relationships to obtain fusion features.
[0118] The decision layer generates a recommendation value for the candidate location based on the fusion features.
[0119] In some embodiments, the data preprocessing layer includes: a first preprocessing module, configured to screen the scale features through a gradient boosting tree to obtain first feature data, and input the first feature data into a multi-layer perceptron to obtain the feature interaction relationship of the scale features.
[0120] In some of these embodiments, the data preprocessing layer includes: a second preprocessing module for modeling the spatial relationships of POI features through a graph neural network to obtain a spatial relationship model; embedding the node structure information of the spatial relationship model into the vector representation of the POI features to obtain second feature data; and inputting the second feature data into a multi-layer perceptron to obtain the feature interaction relationships of the POI features.
[0121] In some of these embodiments, the data preprocessing layer includes: a third preprocessing module for constructing the global temporal dependencies of the temporal features through a Transformer network; capturing the long-term trends according to the global temporal dependencies through a long short-term memory network; and obtaining the feature interaction relationships of the temporal features based on the long-term trends.
[0122] In some of these embodiments, the system further includes: a model construction module, a model training module, and a model evaluation module.
[0123] The model construction module is used to construct a multi-modal fusion deep neural network model based on the binary cross-entropy loss function.
[0124] The model training module is used to train the multi-modal fusion deep neural network model through a data augmentation strategy, and the multi-modal fusion deep neural network model is a binary classification model.
[0125] The model evaluation module is used to evaluate the performance of the trained multi-modal fusion deep neural network model through accuracy, recall rate, and F1 score.
[0126] In some of these embodiments, the system further includes:
[0127] A scoring module for performing weighted calculations on the scale features, POI features, and temporal features to obtain the comprehensive scores of each candidate location.
[0128] A recommended list generation module for determining the recommended star ratings corresponding to each candidate location based on the comprehensive scores according to a preset grading rule, and generating a location recommendation list based on the recommended star ratings and recommended values.
[0129] In some of these embodiments, the data acquisition module further includes:
[0130] A first data correction module for supplementing the missing values of the traffic time series data through forward filling and backward filling.
[0131] A second data correction module for supplementing the missing values of the venue scale data and the popularity POI data through mean filling or category mode filling.
[0132] The third data correction module is used to eliminate extreme values in the venue scale data, popularity POI data, and traffic time series data based on the quantile detection method.
[0133] Through the above system, the problem of low accuracy in travel location selection decision-making is solved. By comprehensively considering the venue scale attribute, popularity POI data, and traffic time series data, it avoids one-sided decision-making based on a single data type. At the same time, a deep neural network is used to fuse multi-modal data to analyze the deep-level correlations between data, improving the accuracy of location selection decision-making.
[0134] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combined form.
[0135] This embodiment also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0136] Optionally, the above-mentioned electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.
[0137] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through a computer program:
[0138] S1. Obtain the venue scale data, popularity POI data, and traffic time series data of the candidate location.
[0139] S2. Respectively perform feature extraction and feature analysis on the venue scale data, popularity POI data, and traffic time series data through a pre-constructed multi-modal fusion deep neural network model to obtain scale features, POI features, and time series features, as well as cross-modal feature interaction relationships.
[0140] S3. Through the multi-modal fusion deep neural network model, according to the cross-modal feature interaction relationship, fuse the scale features, POI features, and time series features, and generate a recommended value for the candidate location based on the fusion result.
[0141] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiment and the optional implementation manners, and will not be repeated here.
[0142] In one embodiment, Figure 5 is a schematic internal structure diagram of an electronic device according to an embodiment of the present application, as Figure 5As shown, an electronic device is provided. The electronic device may be a server, and its internal structure diagram may be as shown in Figure 5 . The electronic device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store data. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a travel location selection method based on multimodal fusion.
[0143] Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0144] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to the memory, storage, database, or other media used in the embodiments provided by the present application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0145] Those skilled in the art should understand that the technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0146] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A travel location selection method based on multimodal fusion, characterized in that: The method comprises: Obtain site size data, popular POI data, and traffic time series data of candidate locations; Through a pre-built multimodal fusion deep neural network model, feature extraction and feature analysis are performed on the venue scale data, the popular POI data and the traffic time series data respectively to obtain scale features, POI features and time series features, as well as cross-modal feature interaction relationships. The multimodal fusion deep neural network model includes a data preprocessing layer, a feature fusion layer and a decision layer; The data preprocessing layer performs feature extraction on the venue scale data, the popular POI data and the traffic time series data to obtain scale features, POI features and time series features, and respectively obtains feature interaction relationships among the scale features, the POI features and the time series features; The feature fusion layer obtains a cross-modal feature interaction relationship according to the feature interaction relationship among the scale feature, the POI feature and the time series feature, and fuses the scale feature, the POI feature and the time series feature based on the cross-modal feature interaction relationship to obtain a fused feature; The decision layer generates a recommendation value for the candidate point based on the fusion feature; Through the multimodal fusion deep neural network model, according to the cross-modal feature interaction relationship, the scale feature, the POI feature and the time series feature are fused, and the recommendation value of the candidate point is generated based on the fusion result.
2. The method according to claim 1, characterized in that The feature interaction relationship of obtaining the scale feature includes: The scale feature is screened by a gradient boosting tree to obtain first feature data; The first feature data is input into a multi-layer perceptron to obtain the feature interaction relationship of the scale feature.
3. The method according to claim 1, characterized in that The feature interaction relationship of acquiring the POI feature includes: Performing spatial relationship modeling on the POI features through a graph neural network to obtain a spatial relationship model; Embed the node structure information of the spatial relationship model into the vector representation of the POI feature to obtain second feature data; The second feature data is input into a multi-layer perceptron to obtain the feature interaction relationship of the POI features.
4. The method according to claim 1, characterized in that The acquiring of the feature interaction relationship of the time series feature comprises: Through the Transformer network, a global time dependency relationship of the time series features is constructed; By using a long short-term memory network, long-term trends are captured according to the global time dependency; Based on the long-term trend, the feature interaction relationship of the time series features is obtained.
5. The method according to claim 1, characterized in that The method further comprises: Constructing the multimodal fusion deep neural network model based on the binary cross entropy loss function; The multimodal fusion deep neural network model is trained through a data enhancement strategy, and the multimodal fusion deep neural network model is a binary classification model; The performance of the trained multimodal fusion deep neural network model was evaluated by accuracy, recall and F1 score.
6. The method according to claim 1, characterized in that The method further comprises: Performing weighted calculation on the scale feature, the POI feature and the time series feature to obtain a comprehensive score for each of the candidate points; Based on the preset grading rules and the comprehensive score, determine the recommended star rating corresponding to each of the candidate points; A point recommendation list is generated according to the recommendation star ratings and the recommendation values.
7. The method according to claim 1, characterized in that Before extracting features from the venue scale data, the popular POI data, and the traffic time series data respectively, the method further includes: By forward filling and backward filling, missing values of the traffic time series data are supplemented; Supplement the missing values of the venue scale data and the popular POI data by filling in the mean or the category mode; Based on the quantile detection method, extreme values in the venue scale data, the popular POI data and the traffic time series data are eliminated.
8. A travel location selection system based on multimodal fusion, characterized in that: The system comprises: The data acquisition module is used to obtain the venue scale data, popular POI data and traffic time series data of the candidate points; A feature extraction module is used to extract and analyze the site scale data, the popular POI data and the traffic time series data through a pre-built multimodal fusion deep neural network model, so as to obtain scale features, POI features and time series features, as well as cross-modal feature interaction relationships. The multimodal fusion deep neural network model includes a data preprocessing layer, a feature fusion layer and a decision layer. The data preprocessing layer performs feature extraction on the venue scale data, the popular POI data and the traffic time series data to obtain scale features, POI features and time series features, and respectively obtains feature interaction relationships among the scale features, the POI features and the time series features; The feature fusion layer obtains a cross-modal feature interaction relationship according to the feature interaction relationship among the scale feature, the POI feature and the time series feature, and fuses the scale feature, the POI feature and the time series feature based on the cross-modal feature interaction relationship to obtain a fused feature; The decision layer generates a recommendation value for the candidate point based on the fusion feature; The recommendation module is used to fuse the scale feature, the POI feature and the time series feature through the multimodal fusion deep neural network model according to the cross-modal feature interaction relationship, and generate a recommendation value for the candidate point based on the fusion result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the travel site selection method based on multimodal fusion as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Site selection method and device, equipment and storage medium
CN110837930A
Depth feature fusion and optimization method and system for multi-modal data
CN117909922A