Multi-source multi-scale intelligent fusion scenic spot passenger flow prediction method, device and equipment and storage medium

By using multi-source data feature fusion and multi-scale time modeling, the problems of insufficient fusion of multi-source heterogeneous data and coarse spatial granularity in passenger flow prediction at scenic spots have been solved, enabling fine-grained passenger flow prediction for each site within the scenic area.

CN120875154APending Publication Date: 2025-10-31GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510997280.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies lack the ability to fuse multi-source heterogeneous data, have coarse spatial prediction granularity, and cannot effectively capture the dynamic changes and transmission relationships of passenger flow at scenic spots.

Method used

By acquiring multi-source data from sites within the scenic area, including historical ticketing data, weather data, holiday data, search index data, and thermal image data, the data is converted into multi-source data feature vectors, fused, and input into a multi-scale time modeling module, and then used to make predictions using a regression decoder.

Benefits of technology

It enables fine-grained prediction of passenger flow data at various stations within the scenic area, improves the fusion capability and prediction accuracy of multi-source heterogeneous data, and solves the problem of coarse spatial prediction granularity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875154A_ABST
    Figure CN120875154A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source multi-scale intelligent fusion scenic spot passenger flow prediction method, device and equipment and a storage medium, and relates to the technical field of data processing. According to the scheme, the method comprises the steps that feature fusion and multi-scale time modeling module processing are conducted on multi-source data of stations in a scenic spot, prediction of passenger flow data of all the stations is achieved through a regression decoder, multi-source deep fusion is conducted on the multi-source heterogeneous data with the stations in the scenic spot as the minimum prediction unit, and a space fine-grained result is output. The problems of insufficient multi-source heterogeneous data fusion capability and rough spatial prediction granularity in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, device, equipment and storage medium for predicting tourist flow in scenic areas using multi-source and multi-scale intelligent fusion. Background Technology

[0002] In recent years, the tourism industry has been booming, with scenic spots facing a constant stream of tourists during holidays. This surge in tourism has also brought about numerous problems. On the one hand, in terms of reception capacity, some scenic spots have exceeded their own capacity, leading to insufficient local infrastructure and reception capacity, resulting in problems such as tourist congestion. On the other hand, whether scenic spot management needs to plan ahead for the next holiday requires relevant visitor flow forecasts to provide strong support for their decision-making. Therefore, accurately predicting the number of tourists to scenic spots is of paramount importance for making relevant preparations.

[0003] In related technologies, prediction models based on historical passenger flow time series are commonly used. Although auxiliary features such as weather and holidays are also introduced, most of them adopt simple splicing strategies for multimodal data and focus on predicting the overall passenger flow of the entire scenic area. This results in problems such as insufficient ability to integrate multi-source heterogeneous data and coarse spatial prediction granularity. Summary of the Invention

[0004] This specification provides a method for predicting tourist flow in scenic areas by intelligent fusion of multiple sources and scales, in order to solve the problems of insufficient fusion capability of multi-source heterogeneous data and coarse spatial prediction granularity in the prior art.

[0005] To solve the above-mentioned technical problems, the embodiments in this specification are implemented as follows:

[0006] Firstly, the embodiments of this specification provide a multi-source, multi-scale intelligent fusion method for predicting visitor flow in scenic areas, applied to stations within the scenic area, including:

[0007] Obtain multi-source data from the site; the multi-source data includes historical ticketing data, weather data, holiday data, search index data, heat map data, and text data;

[0008] The multi-source data is converted into a multi-source data feature vector;

[0009] The feature vectors of the multi-source data are fused to obtain a fused feature vector;

[0010] The fused feature vector is input into the multi-scale time modeling module to obtain the multi-scale feature vector;

[0011] Based on the multi-scale feature vector, the predicted passenger flow data of the station is predicted through a regression decoder.

[0012] Secondly, the multi-source, multi-scale intelligent fusion scenic area visitor flow prediction device provided in the embodiments of this specification, applied to stations within a scenic area, includes:

[0013] The acquisition module is used to acquire multi-source data from the site; the multi-source data includes historical ticketing data, weather data, holiday data, search index data, heat map data, and text data.

[0014] The conversion module is used to convert the multi-source data into multi-source data feature vectors;

[0015] The fusion module is used to fuse the feature vectors of the multi-source data to obtain a fused feature vector;

[0016] A multi-scale time modeling module is used to input the fused feature vector into the multi-scale time modeling module to obtain a multi-scale feature vector;

[0017] The prediction module is used to predict the predicted passenger flow data of the station based on the multi-scale feature vector and through a regression decoder.

[0018] Thirdly, the embodiments of this specification provide a multi-source, multi-scale intelligent fusion scenic area visitor flow prediction device, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the multi-source, multi-scale intelligent fusion scenic area visitor flow prediction method in Scheme 1.

[0019] Fourthly, the embodiments of this specification provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the multi-source, multi-scale intelligent fusion method for predicting tourist flow in scenic areas as described in Scheme 1.

[0020] One embodiment of this specification achieves the following beneficial effects: multi-source data from stations within a scenic area are processed through feature fusion and multi-scale time modeling modules, and then the passenger flow data of each station is predicted through a regression decoder. By using stations within the scenic area as the smallest prediction unit, multi-source heterogeneous data is deeply fused, and spatial fine-grained results are output, thus solving the problems of insufficient multi-source heterogeneous data fusion capability and coarse spatial prediction granularity in the prior art. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart illustrating a multi-source, multi-scale intelligent fusion method for predicting visitor flow in scenic areas, provided as an embodiment of this specification.

[0023] Figure 2 This is a schematic diagram of the process for generating multi-source data feature vectors provided in the embodiments of this specification;

[0024] Figure 3 A schematic diagram of the framework for generating fused feature vectors provided in the embodiments of this specification;

[0025] Figure 4 A schematic diagram of the framework for generating multi-scale feature vectors provided in the embodiments of this specification;

[0026] Figure 5 This is a schematic diagram of the station passenger flow prediction result output process provided in the embodiments of this specification;

[0027] Figure 6 A schematic diagram of the structure of a multi-source, multi-scale intelligent fusion scenic area visitor flow prediction device provided in the embodiments of this specification;

[0028] Figure 7 This is a schematic diagram of a multi-source, multi-scale intelligent fusion scenic area visitor flow prediction device provided in the embodiments of this specification. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of one or more embodiments of this specification clearer, the technical solutions of one or more embodiments of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of one or more embodiments of this specification.

[0030] A typical scenic area includes multiple stations, such as docks, ticket gates, and zone entrances. Taking the Yulong River Scenic Area in Guilin as an example, it contains multiple docks and activity nodes, with significant differences in visitor flow at different stations at different times. Traditional forecasting methods cannot effectively capture this dynamic change at the spatial level, nor can they reflect the flow of visitors between stations.

[0031] In order to overcome the deficiencies in the prior art, the technical solutions provided by the various embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0032] The following is a detailed description of a multi-source, multi-scale intelligent fusion method for predicting visitor flow in scenic areas, based on the embodiments provided in the specification, in conjunction with the accompanying drawings.

[0033] Figure 1 This is a flowchart illustrating a multi-source, multi-scale intelligent fusion method for predicting tourist flow in scenic areas, as provided in an embodiment of this specification. From a programming perspective, the entity executing the process can be a program mounted on an application server or an application client. From a hardware perspective, the entity executing the process can be a terminal device; this embodiment does not impose any particular limitation on this.

[0034] The multi-source, multi-scale intelligent fusion method for predicting visitor flow in scenic areas, as provided in the embodiments of this manual, is applied to stations within scenic areas, such as... Figure 1 As shown, the process may include the following steps:

[0035] Step 110: Obtain multi-source data from the site; the multi-source data includes historical ticketing data, weather data, holiday data, search index data, heat map data, and text data.

[0036] In the embodiments described in this specification, multi-source data is obtained for each station within the scenic area. The multi-source data may include historical ticketing data, weather data, holiday data, search index data, thermal image data, and text data.

[0037] To facilitate understanding, examples of data, data format, and modality for each type of data are provided in Table 1.

[0038] Table 1

[0039] Data types Data Example Data format Modal Historical ticketing data Ticket checking data for each station by hour Time series numerical values Weather data Temperature, wind speed, weather conditions, etc. Array structure Numerical value + category Holiday data Weekday / Holiday / Event Decoration Marking Category scalar category Search index data User search volume, total number of posts across the entire network Time series numerical values Thermal image data Digital twin heat map of mobile signals false color image image Text data Tourists are numerous in the morning and sparse in the afternoon. Sentence-level text Natural Language

[0040] Traversing all stations within the scenic area i We can obtain the set S of all stations within the entire scenic area, S = {s1, s2, ..., s...} n}, where n is the number of stations.

[0041] Step 120: Convert the multi-source data into a multi-source data feature vector.

[0042] In the embodiments described in this specification, data features are extracted and multi-source data are converted into feature vectors to facilitate subsequent fusion and processing, enabling different types of data to be calculated and analyzed in a unified vector space.

[0043] Step 130: Fuse the feature vectors of the multi-source data to obtain a fused feature vector.

[0044] In the embodiments described in this specification, by fusing feature vectors from multiple sources, information from various data sources can be comprehensively utilized to improve the generalization ability and accuracy of the prediction model.

[0045] Step 140: Input the fused feature vector into the multi-scale time modeling module to obtain the multi-scale feature vector.

[0046] In the embodiments of this specification, the changes in tourist flow in scenic areas have multi-scale temporal patterns. Considering the multi-scale characteristics of time, by inputting the multi-scale time modeling module, the patterns of tourist flow changes at different time granularities can be captured, and future tourist flow can be predicted more accurately.

[0047] Step 150: Based on the multi-scale feature vector, predict the predicted passenger flow data of the station using a regression decoder.

[0048] In the embodiments of this specification, the predicted passenger flow data can be the passenger flow at future times. Based on multi-scale feature vectors, the predicted passenger flow data is obtained through a regression decoder, which can provide continuous prediction results.

[0049] It should be understood that the order of some steps in the methods described in one or more embodiments of this specification may be interchanged according to actual needs, or some steps may be omitted or deleted.

[0050] In the embodiments of this specification, multi-source data from stations within the scenic area are processed through feature fusion and multi-scale time modeling modules. Through a regression decoder, the passenger flow data of each station is predicted. By using stations within the scenic area as the smallest prediction unit, multi-source heterogeneous data is deeply fused to output fine-grained spatial results, thus solving the problems of insufficient multi-source heterogeneous data fusion capability and coarse spatial prediction granularity in the prior art.

[0051] based on Figure 1 In addition to the method described in the embodiments of this specification, some specific implementation schemes of the method are also provided, which will be described below.

[0052] Optionally, the conversion of the multi-source data into a multi-source data feature vector as described in the embodiments of this specification may specifically include:

[0053] The multi-source data is transformed by an encoder corresponding to the multi-source data to obtain the feature vector corresponding to the multi-source data;

[0054] The feature vectors are concatenated to obtain multi-source data feature vectors.

[0055] In the embodiments of this specification, data features of each type of data are extracted to obtain feature data corresponding to historical ticketing data, weather data, holiday data, search index data, thermal image data, and text data.

[0056] For different types of data, a dedicated encoder is used for transformation, resulting in feature vectors. Here, Encoder is the encoder.

[0057] In practical applications, data sources include: (historical ticketing data, weather data, holiday data, search index data, heat map data, and text data).

[0058] X_input = [Time Series Features] (Historical passenger flow at ticket checkpoints, daily / hourly temperature records) + [Category Features] + [Image Density Features] + [Image and Text Summary Features]

[0059] For example:

[0060] X_input = timeSequense[ticketNum,tempDegree] + [isHoliday = 1,weatherCode = 0,weekCode = 3,wind_degree = 2,hotWordFreq = 12] + [imgCrowdDensity = 0.65] + ['Crowds are dense in the morning and sparse in the afternoon']

[0061] Feature extraction methods:

[0062] [Time Series Features]: Encoded using Transformer

[0063] [Category Feature]: one-hot

[0064] [Image Density Features]: First, the YOLOv11 object detection model is used to count the number of heads in the image to roughly estimate the crowd size. For high-density occluded scenes, a density map regression network (CSRNet) is further used to generate a pixel-level crowd distribution map, and the integral sum is normalized and used as the image density index feature v. img .

[0065] [Image and Text Summary Features]: For surveillance images and heatmap snapshot data, we first extract temporal information from multiple frames collected within a continuous time period. Then, we use an image semantic generation model (such as BLIP-2) combined with specific prompts to generate a natural language summary of passenger flow status, reflecting the changing trend and spatial distribution characteristics of crowd density during the current time period. Subsequently, we input this summary text into a language encoder (such as BERT) to extract its semantic vector as the image and text summary feature vector v. txt It is used to integrate into a unified multimodal prediction input.

[0066] The multi-source dataset S is split by site Si. For each site's data, it is categorized and joined according to the ts / cat / img / txt type, i.e., (j = Concat(ts, cat, img, txt)) rule, integrating data of the same type to form a single-site multi-type subset. The processed data from all sites are aggregated to form a global multi-source dataset X, realizing the transformation from raw, scattered multi-source data to a structured organization by site and type. This facilitates subsequent encoding, feature extraction, and fusion (normalized and concatenated into h). (Si) Prepare the data.

[0067] Extract all site-related features. After classifying and connecting the data types ts (historical ticketing, search index data), cat (holiday data, weather data), img (thermal image data), and txt (text data), the subsets of each data type represent the original multi-source data organization form before feature encoding.

[0068] X={x ts ,x cat ,x img ,x txt}, where X is the global set of multi-source data from all stations within the scenic area after type-based classification and connection. Specific category data is assigned to corresponding encoders for encoding, resulting in a feature vector set V = {v ts ,v cat ,v img ,v txt Historical ticketing data and search index data are transformed by an encoder to obtain feature vector V. ts Holiday data and weather data are transformed by the encoder to obtain feature vector V. cts The thermal image data is converted by an encoder to obtain the feature vector V. img The text data is transformed by the encoder to obtain the feature vector V. txt .

[0069] To facilitate understanding, examples are provided to illustrate the encoder used for each type of data and the feature vector after transformation for each type of data, as shown in Table 2.

[0070] Table 2

[0071] Feature type encoder Feature vector <![CDATA[Time series type (historical ticketing, search index) x ts > TransformerEncoder <![CDATA[Timing vector v ts > <![CDATA[Categorical (Holiday, Weather) x cat > One-Hot <![CDATA[Embedded vector v cat > <![CDATA[Image type (thermal image) x img > CNN (ResNet-50 Encoder) <![CDATA[v space feature vector img > <![CDATA[Textual (Text) x txt > BERT <![CDATA[Semantic vector v txt >

[0072] By concatenating the feature vectors of the same site, we can obtain the multi-source data feature vector of that site. Here, ConcatProj represents the concatenation and projection transformation operations. First, the input feature vector... The features are concatenated dimensionally, and then linearly projected through a fully connected layer to map the concatenated high-dimensional features to the target dimension. Simultaneously, sub-operations such as LayerNorm normalization and non-linear activation are integrated to achieve multi-source data feature fusion and dimensional adaptation, outputting a regularized multi-source data feature vector h for subsequent tasks. Si) .

[0073] Figure 2This is a schematic diagram of the process for generating multi-source data feature vectors provided in the embodiments of this specification.

[0074] like Figure 2 As shown, the target visit site set (the set of all sites within the entire scenic area) S = {s1, s2, ..., s} n}, Select the time prediction window T f Collect the original input feature set X = {x ts ,x cat ,x img ,x txt Dataset loading and encoder allocation, including historical ticketing data and search index data for each site. Use Transformer Encoder to perform the transformation and obtain the feature vector. Holiday data, weather data Use One-Hot transformation to obtain the feature vector. Thermal image data The feature vectors are obtained by using ResNet-50Encoder for transformation. Text data Use BERT to transform and obtain the feature vector. LayerNorm normalizes and concatenates the data to form the complete fused input representation vector h of Si. Si) .

[0075] To further adapt to the spatial modeling structure and enhance the contextual association between structures, the multi-source data feature vectors of each station are projected into 2K-dimensional channel vectors using an MLP (Multilayer Perceptron). All stations are then rearranged into an H×W two-dimensional structure, constructing a set of multi-source data feature vectors for all stations.

[0076] Where R is the real number field. H×W×2K H represents a set of real tensors of dimension H×W×2K, used to limit the elements h in the set H. (Si) The data type and dimensional characteristics, i.e., h (Si) It is a tensor with real values ​​and dimensions satisfying the H×W×2K specification, which defines the attributes of feature data in the numerical domain and dimensional space.

[0077] Optionally, the fusion of the multi-source data feature vectors to obtain a fused feature vector, as described in the embodiments of this specification, may specifically include:

[0078] The feature vectors of the multi-source data are split by convolutional splitting channels to obtain query vectors and key-value vectors;

[0079] The query vector and the key-value vector are fused using the Cross-Attention mechanism to obtain a fused feature vector.

[0080] In the embodiments described in this specification, the input multi-source feature vector is split into a query vector and a key-value vector through convolution operations. Attention weights are generated by calculating the similarity between the query vector (Q) and the key vector (K), and then weighted and summed on the value vector (V) to achieve cross-modal feature fusion. This efficiently integrates multi-source heterogeneous data and improves the accuracy of scenic area visitor flow prediction.

[0081] Figure 3 This is a schematic diagram of the framework for generating fused feature vectors provided in the embodiments of this specification.

[0082] like Figure 3 As shown, h (si) ∈R d The feature vectors from the multi-source data serve as the input vectors (H×W×2048) for the semantically guided Cross-Attention fusion module.

[0083] Query Vector (Q) Construction: Multi-source data feature vectors are split into main feature vectors through convolutional channel splitting. Then multiply by the weight parameter matrix W Q Perform a linear transformation to obtain the query vector Q(H×W×1024), then perform a linear transformation using the weight matrix to obtain...

[0084]

[0085] Key-value vector (K, V) construction: Multi-source data feature vectors are split into text guidance vectors through convolutional channel splitting. As guiding features, each is dot-multiplied by the weight parameter matrix W. K W V as a key vector Where K is the "key" for attention matching and V is the "value" for attention weighting.

[0086] The association between aligned features is calculated through attention calculation. First, the dot product Q·K of the transposes of Q and K is calculated. T , divided by (d is the feature dimension, used for scaling to avoid excessively large dot products), then use Softmax to calculate the attention weights, multiply the resulting weight matrix by V, and output the fused feature vector. By "finding associations" to align features, Q extracts important features from V based on their matching degree with K.

[0087] In practice, to enhance the fusion result, the attention output can be fused with the feature vector z. (si)The feature vectors from multiple sources are segmented by convolution to split the channels. Vector addition (Add) is: The multi-source fusion feature enhancement vector h after incorporating semantic information is obtained. '(si) .

[0088] Optionally, the multi-scale time modeling module described in the embodiments of this specification includes an hourly prediction sub-network, a daily prediction sub-network, and a weekly prediction sub-network;

[0089] The step of inputting the fused feature vector into the multi-scale time modeling module to obtain the multi-scale feature vector may specifically include:

[0090] The fused feature vector is input into the hourly prediction sub-network to obtain the first-scale feature vector;

[0091] The fused feature vector is input into the daily prediction sub-network to obtain the second-scale feature vector;

[0092] The fused feature vector is input into the weekly prediction sub-network to obtain the third-scale feature vector;

[0093] The first-scale feature vector, the second-scale feature vector, and the third-scale feature vector are concatenated to obtain a multi-scale feature vector.

[0094] In the embodiments of this specification, passenger flow characteristics at the hourly (short-term fluctuations), daily (intraday patterns), and weekly (periodic trends) levels are captured respectively. By splicing and integrating the characteristics of different time scales, the independence of each scale is preserved while enhancing the expressive power.

[0095] Multi-scale time modeling modules can be Transformer, LSTM, or Temporal Convolution, etc.

[0096] The changes in visitor flow in scenic areas exhibit significant multi-scale temporal patterns: the hourly scale reflects immediate fluctuations in visitor flow (such as morning peak and increases after lunch break); the daily scale reflects structural trends within a day (such as the proportion of visitors in the morning / afternoon); and the weekly scale reflects cross-day patterns (such as weekends being hotter and the beginning of the week being flatter).

[0097] Figure 4 This is a schematic diagram of the framework for generating multi-scale feature vectors provided in the embodiments of this specification.

[0098] like Figure 4 As shown, based on the multi-source fusion feature enhancement vector h '(si) Construct three sets of input windows to obtain the hourly time window historical sequence. Historical time window sequence (representative moments of each day over the past 7 days) Historical sequence of weekday time windows (taking the same time period from the previous week)

[0099] The multi-scale time modeling module can include hourly, daily, and weekly prediction subnetworks. The hourly prediction subnetwork... Daily prediction subnetwork Weekly prediction subnetwork

[0100] The multi-source fusion feature vectors are input into the hourly, daily, and weekly prediction sub-networks, respectively. The feature vectors at the three scales are then concatenated dimensionally to obtain the multi-scale feature vector.

[0101] Optionally, the method described in the embodiments of this specification may include:

[0102] The multi-scale feature vector is input into a multilayer perceptron to obtain a time feature vector.

[0103] In the embodiments described in this specification, multi-scale feature vectors are input into a multilayer perceptron (MLP) to complete the structured integration of scale information and nonlinear interactive modeling.

[0104] To model the nonlinear interaction between scales, a shallow multilayer perceptron structure containing two layers of linear transformations and activation functions is designed, as follows:

[0105]

[0106] r (si) =W2·h1+b2∈R d”

[0107] Where W1∈R d'×3d ,b1∈R d' Here are the weights and biases of the first layer; σ1(·) is the ReLU activation function; W2∈R d ”×d' b2∈R d” For output layer transformation; d” represents the unified fusion output dimension, which is usually equal to the prediction target dimension or the decoding input dimension, and the final output is a temporal feature vector.

[0108] Based on time feature vector r (si) The high-dimensional feature vector is mapped to a future passenger flow sequence within the target time window through a set of fully connected transformations. Passenger flow prediction is a multivariate regression task, decoding the high-dimensional feature vector into a multi-step time series prediction result. A lightweight regression decoder is used to transform the station time feature vector r... (si)The data is mapped to target values ​​within the prediction time window to generate numerical passenger flow forecasts. Let the target prediction time window be Tp (e.g., the next hour or the entire next day), then the regression decoder structure is as follows: in, For the predicted passenger flow data of site si in the next Tp time steps, W out ∈R Tp×d b out ∈R Tp These are the weights and bias parameters of the fully connected layer.

[0109] Optionally, the method described in the embodiments of this specification may include:

[0110] Based on the Gaussian distribution, calculate the confidence interval of the predicted passenger flow data;

[0111] or,

[0112] Based on the auxiliary prediction annotation mechanism, the semantic information of the predicted passenger flow data is obtained.

[0113] In the embodiments described in this specification, passenger flow prediction in real-world scenarios involves a certain degree of uncertainty, especially during holidays and unforeseen events where prediction deviations may increase. To enhance the reliability of the prediction, a confidence interval output mechanism can be introduced. This confidence interval output mechanism assumes that each predicted value follows a Gaussian distribution and calculates its standard deviation σ based on a regression bias or residual learning network. t Therefore, the prediction confidence interval is estimated as follows:

[0114]

[0115] in, The primary predicted value, i.e., the point prediction output; To predict variance, estimates are derived from the main model or auxiliary error network; The upper and lower boundaries of the confidence interval represent the 95% confidence range of the prediction. The introduction of this confidence interval output mechanism can add "prediction confidence level" information to the predicted value, helping managers to estimate the risk range and express uncertainty in the form of "prediction bands" in the visualization interface.

[0116] To further enhance the interpretability and practical usability of the prediction results, an auxiliary prediction annotation mechanism based on historical averages, temporal anomaly patterns, and semantic discrimination can be introduced. This auxiliary prediction annotation mechanism can automatically annotate possible outliers in the prediction sequence, such as predicting passenger flow at time t. Compare with historical averages (i.e., the average passenger flow statistics for the same period in the past) In comparison, the γ coefficient can be set empirically or obtained through training to adjust the sensitivity of anomaly detection (for example, a value of 2 corresponds to "2 times the standard deviation"). σ represents the standard deviation of historical data, which measures the dispersion of data from the same period in the past and outputs semantic information for outlier labeling.

[0117] Identifying sudden surges in passenger flow: Judgment criteria are as follows The semantic information is "higher than the historical average plus 2 standard deviations, possibly due to the pre-holiday effect".

[0118] Sudden Drop in Passenger Flow Identification: Judgment Criteria are as follows The semantic information is "far below the historical value for the same period, suspected to be due to extreme weather interference".

[0119] Highly sensitive holiday tags: based on the embedded channel activation value 'a' during holidays. t =Sigmoid(GlobalAvgpool(e t Quantify the impact of holidays and the degree of prediction deviation. To perform rule matching and extract corresponding semantic information; where the activation value a t The feature embedding vector e of holidays t Global average pooling is performed to compress the value into a 1-dimensional scalar, and then the Sigmoid (·) activation function is used to map the value to [0,1]. The closer the value is to 1, the more significant the impact of holidays. ε is a local minimum value to avoid a denominator of 0. Dev t ∈[0,+∞), the larger the value, the more serious the prediction deviation.

[0120] Based on activation value a t Deviation t The threshold rules are mapped to business semantic hints, and the logic is as follows:

[0121] Set threshold (parameters can be adjusted according to business scenarios) v t ≥T a Holidays have a significant impact (e.g., T) a =0.7), indicating that the embedding layer determines that the current situation is a "strong holiday scenario"; Dev t ≥T d : The prediction deviates significantly (e.g., (T) d =0.2), indicating a deviation exceeding 20%.

[0122] The rules and hints mapping is shown in Table 3.

[0123] Table 3

[0124]

[0125] Figure 5 This is a schematic diagram of the station passenger flow prediction result output process provided in the embodiments of this specification.

[0126] like Figure 5 As shown, the temporal semantic vector generates the temporal feature vector r. (si) Linear regression modeling The predicted value calculation allows you to choose between confidence interval calculation and outlier labeling, and outputs the prediction results for each station.

[0127] The prediction results for each station within the scenic area are shown in Table 4.

[0128] Table 4

[0129]

[0130] The multi-source, multi-scale intelligent fusion method for predicting tourist flow in scenic areas provided in the embodiments of this specification has the following advantages:

[0131] 1. Supports fine-grained prediction at the site level, solving the problem of insufficient spatial granularity.

[0132] Traditional methods typically focus on modeling visitor flow trends for the entire scenic area, neglecting the differences in traffic flow and structural fluctuations between stations within the scenic area. This invention constructs an independent spatiotemporal modeling process using stations such as visitor docks and entrances / exits as basic prediction units, enabling parallel visitor flow prediction tasks across multiple stations, time windows, and outputs, thus meeting the needs of refined scenic area operations.

[0133] 2. Introduce a text semantic guidance mechanism to enhance the ability to express multimodal context fusion.

[0134] By introducing the semantic description text embedding vector of tourist density from monitoring images into the fusion module, a cross-modal Cross-Attention structure is constructed, which endows the prediction model with the ability to perceive and respond to the context of the operating state, improves the modeling ability of "implicit time period signals" such as peak congestion and switching between hot and cold spots, and improves the problem of insufficient adaptability of traditional methods to abrupt changes.

[0135] 3. The multi-source feature fusion structure is reasonably designed and has strong scalability.

[0136] Structured encoders are designed for different data types (historical ticketing data, weather data, holiday data, search index data, thermal image data, and text data, etc.), and a fusion representation space is constructed through a unified alignment mapping, exhibiting good modal compatibility and system scalability. It can support the introduction of more data sources in the future, such as GPS mobility trajectories and social platform data, and has the potential for continuous evolution.

[0137] 4. Construct multi-scale time forecasting models to improve the overall performance of short, medium and long-term forecasts.

[0138] By modeling and integrating hourly, daily, and weekly time windows separately, it is possible to simultaneously capture short-term fluctuation trends and long-term cycle patterns, solving the problem of oscillation or trend drift in existing model predictions, and improving the overall stability and robustness of predictions.

[0139] 5. Supports uncertainty modeling and anomaly prediction labeling, enhancing the interpretability and practicality of results.

[0140] Optionally, the confidence interval for each predicted value, as well as anomaly labeling information such as sudden increases / decreases, can be output to facilitate decision support operations such as risk warning, service scheduling, and dynamic window adjustment in the operation platform, thereby further expanding the service value of the prediction model.

[0141] 6. It has good deployment adaptability and industry versatility.

[0142] The algorithm structure of this invention is lightweight and has a clear modular division. It is suitable for various tourist attractions centered around sites (such as mountain and water scenic areas, theme parks, street attractions, etc.) and can be integrated into existing ticketing systems, video surveillance systems and big data platforms, with a clear path for industrial implementation.

[0143] Figure 6 This is a schematic diagram of a multi-source, multi-scale intelligent fusion scenic area visitor flow prediction device provided in the embodiments of this specification.

[0144] Corresponding to the method embodiment, this embodiment also provides a multi-source, multi-scale intelligent fusion scenic area visitor flow prediction device, applied to stations within the scenic area, and may include:

[0145] The acquisition module 602 is used to acquire multi-source data of the site; the multi-source data includes historical ticketing data, weather data, holiday data, search index data, heat map data, and text data;

[0146] The conversion module 604 is used to convert the multi-source data into a multi-source data feature vector;

[0147] The fusion module 606 is used to fuse the feature vectors of the multi-source data to obtain a fused feature vector;

[0148] The multi-scale time modeling module 608 is used to input the fused feature vector into the multi-scale time modeling module to obtain the multi-scale feature vector.

[0149] The prediction module 610 is used to predict the predicted passenger flow data of the station based on the multi-scale feature vector and through a regression decoder.

[0150] Optionally, the fusion of the multi-source data feature vectors to obtain a fused feature vector, as described in the embodiments of this specification, may specifically include:

[0151] The feature vectors of the multi-source data are split by convolutional splitting channels to obtain query vectors and key-value vectors;

[0152] The query vector and the key-value vector are fused using the Cross-Attention mechanism to obtain a fused feature vector.

[0153] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.

[0154] Figure 7 This is a schematic diagram of a multi-source, multi-scale intelligent fusion scenic area visitor flow prediction device provided as an embodiment of this specification. Figure 7 As shown in the embodiments of this specification, a multi-source, multi-scale intelligent fusion scenic area visitor flow prediction device 700 includes a memory 730, a processor 710, and a computer program 720 stored in the memory. The processor 710 executes the computer program 720 to implement the multi-source, multi-scale intelligent fusion scenic area visitor flow prediction method described in any of the above embodiments.

[0155] The embodiments of this specification provide a multi-source, multi-scale intelligent fusion scenic area visitor flow prediction device, which may include a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the multi-source, multi-scale intelligent fusion scenic area visitor flow prediction method described in any of the above embodiments.

[0156] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the multi-source, multi-scale intelligent fusion method for predicting tourist flow in scenic areas as described in any of the above embodiments.

[0157] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, for... Figure 7 As the device shown is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0158] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0159] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0160] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0161] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0162] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0163] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0164] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0165] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0166] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0167] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0168] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0169] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0170] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0171] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0172] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A multi-source, multi-scale intelligent fusion method for predicting visitor flow in scenic areas, characterized in that, Sites used within the scenic area include: Obtain multi-source data from the site; the multi-source data includes historical ticketing data, weather data, holiday data, search index data, heat map data, and text data; The multi-source data is converted into a multi-source data feature vector; The feature vectors of the multi-source data are fused to obtain a fused feature vector; The fused feature vector is input into the multi-scale time modeling module to obtain the multi-scale feature vector; Based on the multi-scale feature vector, the predicted passenger flow data of the station is predicted using a regression decoder.

2. The method according to claim 1, characterized in that, The step of converting the multi-source data into a multi-source data feature vector specifically includes: The multi-source data is transformed by an encoder corresponding to the multi-source data to obtain the feature vector corresponding to the multi-source data; The feature vectors are concatenated to obtain multi-source data feature vectors.

3. The method according to claim 1, characterized in that, The process of fusing the feature vectors from the multi-source data to obtain a fused feature vector specifically includes: The feature vectors of the multi-source data are split by convolutional splitting channels to obtain query vectors and key-value vectors; The query vector and the key-value vector are fused using the Cross-Attention mechanism to obtain a fused feature vector.

4. The method according to claim 1, characterized in that, The multi-scale time modeling module includes an hourly prediction subnetwork, a daily prediction subnetwork, and a weekly prediction subnetwork. The step of inputting the fused feature vector into the multi-scale time modeling module to obtain the multi-scale feature vector specifically includes: The fused feature vector is input into the hourly prediction sub-network to obtain the first-scale feature vector; The fused feature vector is input into the daily prediction sub-network to obtain the second-scale feature vector; The fused feature vector is input into the weekly prediction sub-network to obtain the third-scale feature vector; The first-scale feature vector, the second-scale feature vector, and the third-scale feature vector are concatenated to obtain a multi-scale feature vector.

5. The method according to claim 4, characterized in that, The method includes: The multi-scale feature vector is input into a multilayer perceptron to obtain a time feature vector.

6. The method according to claim 1, characterized in that, The method includes: Based on the Gaussian distribution, calculate the confidence interval of the predicted passenger flow data; or, Based on the auxiliary prediction annotation mechanism, the semantic information of the predicted passenger flow data is obtained.

7. A multi-source, multi-scale intelligent fusion device for predicting visitor flow in scenic areas, characterized in that, Sites used within the scenic area include: The acquisition module is used to acquire multi-source data from the site; the multi-source data includes historical ticketing data, weather data, holiday data, search index data, heat map data, and text data. The conversion module is used to convert the multi-source data into multi-source data feature vectors; The fusion module is used to fuse the feature vectors of the multi-source data to obtain a fused feature vector; A multi-scale time modeling module is used to input the fused feature vector into the multi-scale time modeling module to obtain a multi-scale feature vector; The prediction module is used to predict the predicted passenger flow data of the station based on the multi-scale feature vector and through a regression decoder.

8. The apparatus according to claim 7, characterized in that, The process of fusing the feature vectors from the multi-source data to obtain a fused feature vector specifically includes: The feature vectors of the multi-source data are split by convolutional splitting channels to obtain query vectors and key-value vectors; The query vector and the key-value vector are fused using the Cross-Attention mechanism to obtain a fused feature vector.

9. A multi-source, multi-scale intelligent fusion scenic area visitor flow prediction device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 6.