Amusement park real-time crowd density monitoring method and system based on visual recognition

By deploying video capture equipment arrays in amusement parks to collect data from multiple perspectives and extract stereoscopic visual features, the problems of large errors in manual counting and inaccurate sensor monitoring in existing technologies have been solved. This has enabled real-time and accurate monitoring of crowd density and optimization of crowd flow, thereby improving the operational efficiency of amusement parks and the visitor experience.

CN122157154APending Publication Date: 2026-06-05BEIJING MIAOXIANG SCIENCE & TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING MIAOXIANG SCIENCE & TECHNOLOGY CO LTD
Filing Date
2026-03-03
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing technologies for monitoring crowd flow in amusement parks suffer from high costs and large errors due to manual counting, and sensor-based monitoring cannot accurately obtain the spatial distribution and movement status of visitors, making it difficult to achieve comprehensive real-time crowd density monitoring.

Method used

By deploying an array of video capture devices to acquire multi-view video stream data, performing stereoscopic visual feature extraction, combining it with spatial semantic map data of the amusement park area for coordinate mapping, generating a real-time density distribution grid data matrix, and analyzing visitor flow trends to generate guidance signals to optimize pedestrian flow distribution.

Benefits of technology

It enables real-time and accurate monitoring of tourist density and analysis of flow trends, effectively guiding tourists to flow rationally, avoiding localized congestion, and improving operational efficiency and tourist experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122157154A_ABST
    Figure CN122157154A_ABST
Patent Text Reader

Abstract

The application provides a kind of amusement park real-time crowd density monitoring method and system based on visual identification, it is related to computer vision technical field, first acquisition amusement park each area video acquisition equipment array synchronous collection multi-angle video stream data set, obtains tourist spatial position distribution and individual motion vector field feature set by stereoscopic vision feature extraction processing, calls spatial semantic map data and carries out coordinate mapping alignment, generates tourist real-time density distribution grid data matrix, again, the tourist real-time density distribution grid data matrix is analyzed to the tourist flow trend, obtains density change rate characteristic field and flow main direction characteristic vector set, finally, based on the above information generates the coding sequence of area-by-area flow guide instruction signal, drives prompt light array difference dynamic indication.The application realizes the real-time, accurate, comprehensive crowd density monitoring and effective flow guide of amusement park.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and more specifically, to a method and system for real-time crowd density monitoring in amusement parks based on visual recognition. Background Technology

[0002] In the daily operation of amusement parks, real-time monitoring of crowd density in various areas is crucial. Accurate crowd density monitoring helps amusement park managers to allocate resources rationally, optimize visitor routes, and prevent safety accidents such as stampedes, thereby improving the visitor experience and the overall operational safety of the amusement park.

[0003] Currently, the main methods for monitoring crowd flow in amusement parks are manual counting and monitoring based on simple sensors. Manual counting requires a large workforce, is not only costly but also prone to errors, making real-time and comprehensive monitoring difficult. Monitoring based on simple sensors, such as infrared sensors and pressure sensors, can reflect crowd flow to some extent, but these sensors typically only acquire localized, discrete information. They cannot accurately obtain detailed information such as the spatial distribution and movement patterns of visitors, making it difficult to create a comprehensive and accurate crowd density distribution map or effectively analyze visitor flow trends. Summary of the Invention

[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for real-time crowd density monitoring in amusement parks based on visual recognition, the method comprising: Acquire a collection of multi-view video stream data of the amusement park area during the current monitoring period, which is synchronously collected by the video acquisition equipment array deployed in various areas of the amusement park; Based on the video stream data set, stereoscopic visual features of the amusement park scene are extracted to obtain a set of spatial location distribution features of tourists based on pixel coordinates and a set of individual tourist motion vector field features. The pre-configured spatial semantic map data of the amusement park area is called, and the spatial coordinate mapping and alignment processing of the tourist spatial location distribution feature set and the tourist individual motion vector field feature set is performed to generate a real-time tourist density distribution grid data matrix of each amusement park area based on the geographic coordinate system during the current monitoring period. The real-time density distribution grid data matrix of tourists is processed by tourist flow trend analysis to obtain the characteristic field of tourist density change rate and the set of characteristic vectors of tourist flow main direction in each amusement park area. Based on the real-time visitor density distribution grid data matrix, the visitor density change rate feature field, and the visitor flow main direction feature vector set, a region-by-region flow guidance signal encoding sequence is generated to drive the indicator light arrays distributed in various areas of the amusement park to provide differentiated dynamic indications. Furthermore, this embodiment of the invention also provides a visual recognition-based real-time visitor density monitoring system for amusement parks, including: A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the above-described method for real-time crowd density monitoring in amusement parks based on visual recognition by executing the machine-executable instructions.

[0005] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, the processor of the vision-based amusement park real-time crowd density monitoring system reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the vision-based amusement park real-time crowd density monitoring system to perform the aforementioned vision-based amusement park real-time crowd density monitoring method.

[0006] Based on the above, by acquiring multi-view video stream data sets synchronously collected by video capture equipment arrays in various areas of the amusement park, scene information of each area can be captured. Based on the above video stream data, stereoscopic visual feature extraction processing is performed to obtain a set of spatial location distribution features and a set of individual visitor motion vector field features. Using pre-configured spatial semantic map data of the amusement park area, the above feature sets are aligned by regional spatial coordinate mapping, generating a real-time visitor density distribution grid data matrix based on a geographic coordinate system. This achieves a spatial representation of visitor density information, allowing managers to intuitively understand the visitor flow distribution in each area. Visitor flow trend analysis is performed on the real-time visitor density distribution grid data matrix to obtain a visitor density change rate feature field and a set of visitor flow main direction feature vectors, further revealing the dynamic changes in visitor flow. Finally, based on the above information, a region-by-region guidance indicator signal encoding sequence is generated to drive the indicator light array for differentiated dynamic indication. This can guide visitors to flow rationally in real time and effectively, avoiding overcrowding in local areas and improving the operational efficiency of the amusement park and the visitor experience. Attached Figure Description

[0007] Figure 1 This is a schematic diagram of the execution flow of the real-time crowd density monitoring method for amusement parks based on visual recognition provided in an embodiment of the present invention.

[0008] Figure 2This is a schematic diagram of exemplary hardware and software components of a vision recognition-based real-time crowd density monitoring system for amusement parks provided in an embodiment of the present invention. Detailed Implementation

[0009] Figure 1 This is a flowchart illustrating a method for real-time crowd density monitoring in amusement parks based on visual recognition, provided in one embodiment of the present invention. A detailed description follows.

[0010] Step S110: Obtain the set of multi-view video stream data of the amusement park area during the current monitoring period, which is synchronously collected by the video acquisition equipment array deployed in various areas of the amusement park.

[0011] In this embodiment, taking a "Fantasy Valley" amusement park as an example, to achieve real-time monitoring of crowd density in various areas of the "Fantasy Valley" amusement park, raw visual data is first acquired. The amusement park deploys video acquisition equipment arrays in areas such as "Adventure Island," "Fairy Tale Castle," "Future World," and "Happy Avenue." This video acquisition equipment array consists of multiple high-definition network cameras, covering different perspectives of each area. During the current monitoring period from 2:00 PM to 2:05 PM, all cameras synchronize their clocks and acquire data via the Network Time Protocol. The acquired data is multi-perspective, meaning that the same area may be simultaneously captured by multiple cameras at different angles. The above data together constitute the video stream data set for the current monitoring period. This video stream data set is a composite data structure containing multiple data streams. For example, the video stream generated by a camera in the "Adventure Island" area is a series of video frame image units arranged in chronological order. Each unit is a pixel matrix with a resolution of 1920*1080, and each element in the matrix represents the red-green-blue color space value of a pixel. The entire video stream data set is logically organized as a five-dimensional tensor, with the dimensions being camera number, time index, image height, image width, and color channels.

[0012] Step S120: Perform stereoscopic visual feature extraction processing on the amusement park scene based on the video stream data set to obtain a set of spatial location distribution features of tourists based on pixel coordinate system and a set of individual tourist motion vector field features.

[0013] Next, feature extraction is performed on the video stream dataset to transform the raw pixel information into structured spatial location and motion state data.

[0014] Step S121: Perform video frame demultiplexing processing on each video stream data in the video stream data set to obtain multiple amusement park area video frame image units with synchronization timestamp markers.

[0015] Each compressed video stream is demultiplexed to separate consecutive image frames. For the video stream from camera A in the "future world" region, each frame is read and a capture timestamp is appended to it, generating a video frame image unit with a timestamp. The data structure of this video frame image unit includes an image header storing metadata such as the timestamp, camera identifier, and image size, followed by the image pixel data body. The timestamp format is a Unix timestamp plus a millisecond-level offset, ensuring that frames captured by different cameras at the same time have highly consistent time identifiers. The image size metadata records the width W_frame and height H_frame of the frame. All camera video streams undergo this processing, forming an image unit pool organized by time and camera index. This image unit pool is organized in memory as a dictionary structure, with the key being a combination of camera ID and timestamp, and the value being the corresponding image pixel data matrix.

[0016] Step S122: Perform initial segmentation processing on the video frame image units of the amusement park area to separate the foreground and background, and obtain a mask image of the foreground pixel area containing the outline of a suspected tourist.

[0017] To separate dynamic visitors from a static background, perform foreground segmentation.

[0018] Step S1221: Perform color space model conversion processing on the video frame image unit of the amusement park area, converting the original red-green-blue color space pixel values ​​to the hue-saturation-brightness color space to obtain the hue-saturation-brightness color space image unit.

[0019] Each video frame image unit is converted from the red-green-blue color space to the hue-saturation-brightness color space. For a video frame in the "Fairytale Castle" region, the original pixel value consists of three components: red (R), green (G), and blue (B), each ranging from 0 to 255. The conversion process first calculates intermediate variables: maximum value V_max = max(R, G, B), minimum value V_min = min(R, G, B), and the difference Δ = V_max - V_min. The brightness component V is directly taken as V_max / 255. The saturation component S is calculated as follows: if V_max == 0, then S = 0; otherwise, S = Δ / V_max. The hue component H is calculated based on the maximum value component: if V_max == R, then H = 60 * ((GB) / Δ) % 360; if V_max == G, then H = 60 * ((BR) / Δ) + 120; if V_max == B, then H = 60 * ((RG) / Δ) + 240. Ultimately, each pixel is represented by three components: hue (H), saturation (S), and lightness (V), with values ​​ranging from 0 to 360, 0 to 1, and 0 to 1, respectively. This color space decouples color information from lightness information, making subsequent color-based foreground detection more robust to changes in illumination.

[0020] Step S1222: Based on the background modeling reference frame image corresponding to the amusement park area video frame image unit, calculate the difference in hue component, saturation component, and brightness component between the amusement park area video frame image unit and the background modeling reference frame image at the corresponding pixel positions in the hue-saturation-brightness color space.

[0021] A background modeling reference frame image is pre-maintained for each camera viewpoint. This reference frame image is calculated using the median pixel values ​​of multiple consecutive frames of a scene without tourists, representing the static baseline of the scene. For example, the background reference frame for the "Happy Street" camera is an image of an empty street, also stored in a hue-saturation-brightness color space. In this hue-saturation-brightness color space, the component differences between the current frame and the background reference frame at each corresponding pixel position are calculated. For a pixel at coordinates (x, y), let the hue, saturation, and brightness components of the current frame be H_cur, S_cur, and V_cur, and the corresponding components of the background frame be H_bg, S_bg, and V_bg. The calculation of the hue component difference ΔH needs to consider the cyclical nature of hue, taking the minimum value (|H_cur - H_bg|, 360 - |H_cur - H_bg|). The saturation component difference ΔS = |S_cur - S_bg|. The brightness component difference ΔV = |V_cur - V_bg|. These three differences describe the degree of change of a pixel relative to the background model from different dimensions.

[0022] Step S1223: Perform weighted fusion processing on the hue component difference, saturation component difference and lightness component difference to generate a pixel difference metric map that represents the degree to which a pixel deviates from the background model.

[0023] The three component differences are weighted and fused. Since the brightness component is sensitive to changes in illumination, while the hue and saturation components are sensitive to the object's own color, higher weights (w_H, w_S) are assigned to the hue and saturation differences, respectively, while a lower weight (w_V) is assigned to the brightness difference. These weights are determined through offline optimization to minimize the fusion value of the background region and maximize the fusion value of the foreground object under varying illumination. For a pixel at coordinates (x, y), its pixel difference metric D is calculated using D = w_H*ΔH + w_S*ΔS + w_V*ΔV. After performing the same calculation on all pixels, a pixel difference metric map D_map with the same size as the original image is obtained. The floating-point value of each pixel in the map represents the degree of difference between that point and the background model; a larger value indicates a higher probability that the pixel belongs to the foreground.

[0024] Step S1224: Perform preliminary binarization processing on the pixel difference measurement map based on global adaptive threshold segmentation to obtain an initial foreground pixel label map.

[0025] The global adaptive threshold algorithm is used to calculate the optimal segmentation threshold \(T_{otsu}\). The core idea is to traverse all possible thresholds, calculate the within-class variance and between-class variance of foreground and background pixels under each threshold, and maximize the ratio of between-class variance to within-class variance. Specifically, let the total number of pixels in the pixel difference metric map be \(N_{total}\), and the number of gray levels be \(L\). For each potential threshold \(t\), calculate the foreground pixel ratio \(w_0(t)\) and background pixel ratio \(w_1(t)\), as well as the foreground mean \(\mu_0(t)\) and background mean \(\mu_1(t)\). The global mean \(\mu = w_0(t) * \mu_0(t)+w_1(t) * \mu_1(t)\). The formula for the between-class variance \(\sigma_b^2(t)\) is \(w_0(t) * (\mu_0(t)-\mu)^2+w_1(t) * (\mu_1(t)-\mu)^2\). Traverse all \(t\) values and find the \(t\) that maximizes \(\sigma_b^2(t)\) as \(T_{otsu}\). Compare each pixel point on the pixel difference metric map \(D_{map}\) with this threshold \(T_{otsu}\): if \(D_{map}(x,y)\geq T_{otsu}\), mark this pixel point as 1, representing the suspected foreground area; if \(D_{map}(x,y)<T_{otsu}\), mark it as 0, representing the background area. Thus, a binary initial foreground pixel label map \(B_{init}\) is generated.

[0026] Step S1225: Perform morphological opening operation on the initial foreground pixel label map to eliminate isolated noisy pixel blocks in the suspected foreground area and connect the broken parts caused by discontinuous segmentation in the suspected foreground area, and obtain a morphologically optimized foreground label map.

[0027] There may be isolated noise points in the initial foreground pixel label map \(B_{init}\) due to leaf shaking and slight light changes, as well as holes or breaks inside the foreground object due to similar colors to the background. Use morphological opening operation for processing. This operation first performs erosion operation and then dilation operation. The structuring element for the erosion operation is a 3*3 square kernel. For each pixel point, only when all pixels in its neighborhood are 1, the central pixel is retained as 1, otherwise it is set to 0. This operation can eliminate isolated white noise points smaller than the size of the structuring element. The dilation operation also uses a 3*3 square kernel. For each pixel point, as long as there is at least one pixel in its neighborhood that is 1, the central pixel is set to 1. This operation can restore the foreground area shrunk due to erosion and connect adjacent broken parts. After the opening operation of erosion first and then dilation, a morphologically optimized foreground label map \(B_{morph}\) with less noise and more complete regions is obtained.

[0028] Step S1226: Perform connected component analysis on the foreground marker map after morphological optimization, extract all connected regions composed of 1 pixel, and calculate the characteristic parameters of the area of the bounding rectangle and the contour perimeter of each connected region.

[0029] Perform connected component analysis on the foreground marker map B_morph after morphological optimization to find all connected blocks of pixels with a value of 1. Use a two-pass scan algorithm. In the first pass, assign a temporary label to each pixel and record the equivalence relationship. In the second pass, merge the equivalent labels. Finally, each independent pixel block is assigned a unique connected region label L_id. Calculate the area A_rect of the minimum bounding rectangle for each connected region. The area is obtained by multiplying the length L of the rectangle by the width W, i.e., A_rect = L * W. At the same time, calculate the perimeter P_contour of its contour. By traversing the boundary pixels of the region, calculate the Euclidean distance between adjacent boundary pixels and accumulate it. For an 8-connected region, the step size in the horizontal and vertical directions is counted as 1, and the step size in the diagonal direction is counted as √2.

[0030] Step S1227: According to the preset minimum bounding rectangle area threshold and contour perimeter range threshold for tourist targets, perform preliminary screening and filtering on the connected regions, and剔除 connected regions with an area smaller than the minimum bounding rectangle area threshold or a contour perimeter exceeding the contour perimeter range threshold to obtain a candidate tourist connected region set.

[0031] According to prior knowledge, the projected area and contour perimeter of a normal tourist in the image are within a reasonable range. Set the minimum bounding rectangle area threshold Th_A_min, which is pre-calibrated according to the installation height and viewing angle of the camera to ensure that small noise regions can be filtered out. At the same time, set the contour perimeter range threshold [Th_P_min, Th_P_max], which is obtained based on ergonomic statistics and can exclude non-tourist targets that are too large or too small. Traverse all connected regions. If the area A_rect of a certain region < Th_A_min, or its contour perimeter P_contour < Th_P_min or P_contour > Th_P_max, then determine that this region is noise or a non-tourist target, and set all pixels of this region to 0 in the marker map. The remaining connected regions form the candidate tourist connected region set C_candidate.

[0032] Step S1228: Perform hole filling on the inside of the contour for each connected region in the candidate tourist connected region set to ensure that each connected region is composed of a continuous pixel area and generate a complete connected region mask without internal holes.

[0033] Some candidate tourist connected regions may contain holes (pixels marked as 0) due to the similarity between the tourist's clothing color and the background. Each connected region is processed using either a flood fill algorithm or a scanline fill algorithm. For a connected region R_i, starting from any foreground pixel within the region, the search expands outwards, filling all background pixels connected to the starting point and surrounded by the region's outline with 1. Specifically, this can be achieved using stack-based recursive filling or scanline-based non-recursive filling. After hole filling, a complete connected region mask M_filled_i with all internal pixel values ​​of 1 is obtained. This process is repeated for all candidate regions to obtain the set of filled connected regions.

[0034] Step S1229: Perform edge smoothing on each connected region mask after hole filling. Use a polygon approximation algorithm to simplify the jagged edges of the connected region contour to obtain a smooth connected region contour line.

[0035] The segmented contours often have jagged edges due to image pixelation. The Douglas-Puk algorithm is used to approximate the contour of each connected region as a polygon. Its input is an ordered sequence of pixels P_seq that constitutes the contour. It recursively searches for points whose perpendicular distance from the line connecting the start and end points of the contour is greater than a preset threshold ε. First, a straight line segment is constructed between the start and end points of the contour, and the perpendicular distance d_perp from all intermediate points to this line segment is calculated. The maximum distance d_max and its corresponding point P_max are found. If d_max > ε, the contour is divided into two segments using P_max as the dividing point, and the same processing is recursively performed on each segment. If d_max ≤ ε, all intermediate points are discarded, and only the start and end points are retained. Through recursive calls, a smooth polygonal contour P_poly composed of a small number of vertices is finally obtained. The value of ε is set according to the image resolution and accuracy requirements, for example, 1.5 pixels.

[0036] Step S12210: Generate a foreground pixel region mask image containing the outline of a suspected tourist based on the smooth edge connected region contour line, wherein each connected region in the foreground pixel region mask image represents a two-dimensional pixel projection of a tourist candidate object.

[0037] The smooth contour line P_poly, after polygon approximation, is used as the boundary of the mask. A polygon filling algorithm is used to assign a value of 1 to all pixels inside the contour and a value of 0 to those outside, generating the final foreground pixel region mask image M_final. This foreground pixel region mask image is a binary image, where each connected region with a value of 1 logically represents the projection of a tourist candidate object onto a two-dimensional image plane. M_final has the same size as the original video frame.

[0038] Step S123: Perform contour edge refinement fitting on each connected region in the foreground pixel region mask image, extract the head vertex pixel coordinates and foot contact point pixel coordinates of the tourist candidate object corresponding to the connected region, and generate a pair of two-dimensional spatial positioning points of the tourist individual in the pixel coordinate system.

[0039] After obtaining the contour mask of the tourists, the key points of each tourist are precisely located, namely the top of the head and the contact point of the feet.

[0040] Step S1231: Perform minimum bounding rectangle orientation calibration on the target connected region in the foreground pixel region mask image, calculate the principal axis direction angle of the target connected region, and rotate the target connected region according to the principal axis direction angle so that the principal axis of the target connected region is parallel to the vertical axis of the image coordinate system, thereby obtaining the pose-normalized connected region mask.

[0041] For a selected connected region R_target, its minimum bounding rectangle is first calculated. The eigenvalues ​​and eigenvectors are obtained by calculating the covariance matrix of the contour points in this region. The covariance matrix C = [[Σ((x_i-x_mean)^2) / n,Σ((x_i-x_mean)(y_i-y_mean)) / n];[Σ((x_i-x_mean)(y_i-y_mean)) / n,Σ((y_i-y_mean)^2) / n]], where x_mean and y_mean are the average coordinates of the contour points. The direction of the eigenvector corresponding to the largest eigenvalue of the covariance matrix is ​​the principal axis direction. The principal axis angle θ_main is the angle between this eigenvector and the horizontal axis. Using the center point (cx, cy) of the region as the rotation center, the entire connected region mask is rotated by -θ_main angles to make the principal axis parallel to the vertical axis of the image. Rotation is achieved through an affine transformation of the image, with the transformation matrix M_rot = [[cosθ_main, sinθ_main, (1-cosθ_main)*cx-sinθ_main*cy]; [-sinθ_main, cosθ_main, sinθ_main*cx+(1-cosθ_main)*cy]]. After applying this transformation, a pose-normalized mask M_norm is obtained.

[0042] Step S1232: Perform vertical pixel projection distribution statistical processing on the pose-normalized connected region mask, calculate the number of foreground pixels contained in each row of pixels on the vertical axis of the connected region mask, and obtain the vertical projection histogram.

[0043] For the pose-normalized mask M_norm, iterate through each row, i.e., each y-coordinate value. For a given y, count the number of pixels in that row with a value of 1, denoted as N_y. Specifically, create a one-dimensional array hist_y with a length of the image height H, initialized to 0. Iterate through all pixels in M_norm; if a pixel has a value of 1, then hist_y[y] += 1. Finally, obtain the vertical projection histogram, where the x-coordinate is the y-coordinate and the y-coordinate is the number of foreground pixels N_y.

[0044] Step S1233: Perform peak detection processing on the vertical projection histogram to identify the inflection point where the number of foreground pixels in the vertical projection histogram changes. The inflection point includes the head starting point, the shoulder peak point, and the foot starting point.

[0045] Analyze the vertical projection histogram `hist_y`. Start scanning downwards from the top of the image (y=0) to `y=H`. The head start point (y_head_start) is defined as the first y-value satisfying `hist_y[y]>0` and `hist_y[y-1]==0`. Continue scanning downwards, searching for local peaks in the histogram, i.e., y-values ​​satisfying `hist_y[y]>hist_y[y-1]` and `hist_y[y]>hist_y[y+1]`. The largest peak is denoted as the shoulder peak (y_shoulder_peak). Continue scanning downwards to the bottom, where the foot start point (y_feet_start) is defined as the last y-value satisfying `hist_y[y]>0` and `hist_y[y+1]==0`. These three turning points characterize the locations of key parts of the human body in the vertical dimension.

[0046] Step S1234: Based on the vertical coordinates corresponding to the head starting point, locate the horizontal coordinates of the geometric center of the foreground pixel in the vertical coordinate row on the connected region mask, and determine the combination of the horizontal coordinates of the geometric center and the vertical coordinates of the head starting point as the first coordinates of the head vertex of the tourist candidate in the pose normalized image.

[0047] In the pose-normalized image M_norm, locate the row with y-coordinate y_head_start. Traverse all pixel columns x from 0 to W in this row, collecting the column coordinates of all pixels with a value of 1, forming a set X_head_set. Calculate the arithmetic mean of these column coordinates to obtain the horizontal coordinate of the geometric center x_head_center = (Σx_i) / |X_head_set|, where x_i belongs to X_head_set. Therefore, the head vertex coordinates in the pose-normalized image are (x_head_center, y_head_start). This point represents the horizontal center of the top of the head.

[0048] Step S1235: Based on the vertical coordinates corresponding to the foot starting point, locate the geometric center horizontal coordinates of the foreground pixel in the vertical coordinate row on the connected region mask, and determine the combination of the geometric center horizontal coordinates and the vertical coordinates of the foot starting point as the second coordinates of the foot contact point of the tourist candidate in the pose normalized image.

[0049] Locate the row with y-coordinate y_feet_start. Traverse all pixel columns x from 0 to W in this row, collecting the column coordinates of all pixels with a value of 1, forming a set X_feet_set. Calculate the arithmetic mean of these column coordinates to obtain the horizontal coordinate of the geometric center x_feet_center = (Σx_i) / |X_feet_set|. The coordinates of the foot contact point in the pose-normalized image are then (x_feet_center, y_feet_start). This point represents the horizontal center of the foot's contact with the ground.

[0050] Step S1236: Based on the rotation angle and rotation center used in the pose normalization process, perform reverse rotation transformation on the first coordinate and the second coordinate to map the first coordinate and the second coordinate back to the pixel coordinate system of the original amusement park area video frame image unit.

[0051] The head vertex coordinates P_head_norm=(x_head_center, y_head_start) and foot contact point coordinates P_feet_norm=(x_feet_center, y_feet_start) in the pose-normalized image are subjected to a reverse rotation transformation. The reverse rotation matrix M_inv=[[cosθ_main, -sinθ_main, (1-cosθ_main)*cx+sinθ_main*cy]; [sinθ_main, cosθ_main, -sinθ_main*cx+(1-cosθ_main)*cy]]. The point coordinates are converted to homogeneous form (x, y, 1) and multiplied by M_inv to obtain the coordinates P_head_orig=(x_head_orig, y_head_orig) and P_feet_orig=(x_feet_orig, y_feet_orig) in the original image. This mapping accurately restores the normalized keypoints to the pose of the original image.

[0052] Step S1237: Perform sub-pixel precision edge search processing within a local neighborhood window on the head vertex coordinates after mapping back to the original pixel coordinate system. Within a preset window around the head vertex coordinates, adjust the head vertex coordinates to the position of the local maximum curvature point of the head contour using image grayscale gradient information.

[0053] In the original image I_orig, a local neighborhood window of size W_wind*W_wind, centered at P_head_orig, is defined, with W_wind set to 11. Within this window, the Sobel operator is first used to calculate the gray-level gradient of each pixel, obtaining the horizontal gradient G_x and the vertical gradient G_y. The gradient magnitude G_mag = √(G_x^2 + G_y^2). A search is performed along the normal direction of the head contour, which is determined by the gradient direction. The location with the largest gradient magnitude is found, corresponding to the point with the strongest edge in the image, i.e., the point of maximum local curvature of the head contour. Since the gradient magnitude is discrete, a quadratic surface is fitted in the maximum point and its neighborhood using bilinear interpolation to solve for the sub-pixel precision offsets Δx_sub and Δy_sub. The final optimized sub-pixel coordinates of the head vertices are x_head_fine = x_head_orig + Δx_sub, y_head_fine = y_head_orig + Δy_sub.

[0054] Step S1238: Optimize the foot contact point coordinates after mapping back to the original pixel coordinate system based on the geometric constraints of the ground contact point. Combine the pixel distribution of the bottom edge of the connected region of the tourist candidate object, adjust the foot contact point coordinates to the midpoint of the bottom edge of the connected region and the coincidence point of the bottom edge and the ground.

[0055] In the original image, obtain the lowest edge of the connected region mask M_final. Traverse all pixels in M_final with a pixel value of 1, and find the set S_lowest of all points with the minimum y-coordinate value y_min. Calculate the average horizontal coordinate of these lowest edge pixels to obtain the horizontal coordinate of the midpoint of the lowest edge: x_feet_lowest = (Σx_i) / |S_lowest|. Simultaneously, using the ground plane equation obtained through camera calibration in the scene, calculate the geometric position of the ground contact line in the image. Determine the point where the lowest edge coincides with the ground contact line; typically, the midpoint is taken as the optimal contact point. Optimize and adjust the coordinates of the foot contact point to (x_feet_lowest, y_min) to ensure it accurately falls on the physical contact point between the tourist and the ground.

[0056] Step S1239: The optimized head vertex coordinates and optimized foot contact point coordinates in the original pixel coordinate system are output as the head vertex pixel coordinates and foot contact point pixel coordinates of the tourist candidate object corresponding to the connected region.

[0057] The final output consists of the head vertex coordinates (x_head_fine, y_head_fine) optimized by subpixel edge search and the foot contact point coordinates (x_feet_lowest, y_min) optimized by ground contact point geometric constraints. Both coordinates are floating-point numbers with subpixel precision, enabling a more accurate description of the tourist's position in the image.

[0058] Step S12310: Associate and store the extracted head vertex pixel coordinates and foot contact point pixel coordinates with the corresponding tourist candidate object identifier to generate a two-dimensional spatial location point pair data structure of the tourist individual in the pixel coordinate system.

[0059] Each detected tourist individual is assigned a unique tourist identifier ID_k, encoded as a 64-bit integer to ensure uniqueness throughout the monitoring period. The head vertex coordinates and foot contact point coordinates output in step S1239 are associated with this ID_k and stored to form a two-dimensional spatial location point pair data structure. This data structure is represented in JSON format as {"id":ID_k, "head":[x_head_fine_k, y_head_fine_k], "feet":[x_feet_lowest_k, y_min_k]}. All detected individuals are compiled into a list and stored in an in-memory database.

[0060] Step S124: Based on the pixel coordinates of the foot contact points in the two-dimensional spatial location point pairs of the tourist individuals in the pixel coordinate system, generate a tourist spatial location distribution feature set labeled with pixel coordinate values. The tourist spatial location distribution feature set contains the sequence of foot contact point pixel coordinates of all detected tourist individuals during the current monitoring period.

[0061] Iterate through all video frames within the current monitoring period. For each individual tourist ID_k detected in each frame t_m, extract the pixel coordinates (x_feet_k_m, y_feet_k_m) of their foot contact point from the data structure in step S12310. Organize all extracted foot contact point coordinates according to time order and tourist identifiers to form a tourist spatial location distribution feature set S_pos. This tourist spatial location distribution feature set is organized in memory as a nested dictionary: the outer key is the timestamp t_m, the inner key is the tourist identifier ID_k, and the value is the corresponding foot contact point coordinates (x_feet_k_m, y_feet_k_m). This tourist spatial location distribution feature set completely records the two-dimensional spatial location information of each tourist at each moment.

[0062] Step S125: Perform optical flow field calculation on the video frame image units of the amusement park area in the continuous time series to calculate the pixel grayscale change between adjacent frames, and obtain a two-dimensional optical flow vector distribution map representing the instantaneous motion velocity of the pixel.

[0063] To obtain motion information for each pixel in the scene, the optical flow field between adjacent frames is calculated.

[0064] Step S1251: Select the current video frame image unit and the previous historical video frame image unit adjacent to the current video frame image unit from the amusement park area video frame image units in the continuous time series.

[0065] Two consecutive frames are selected from the video stream, for example, frame I_t at the current time t and frame I_{t-1} at the previous time t-1, as the input pair for optical flow calculation. Both frames are grayscale images. If the original images are color images, they are converted to grayscale using the weighted conversion formula Gray=0.299*R+0.587*G+0.114*B to ensure that the optical flow calculation is based on luminance information.

[0066] Step S1252: Perform multi-scale pyramid hierarchical downsampling processing on the current video frame image unit and the previous historical video frame image unit respectively, and construct an image pyramid hierarchical structure from the original resolution to the low resolution to obtain the current frame image pyramid and the historical frame image pyramid.

[0067] Image pyramids are constructed for I_t and I_{t-1} respectively. Using the original image as layer 0, lower-resolution images are generated layer by layer through Gaussian downsampling. The downsampling process uses a 5*5 Gaussian kernel for convolution, and then samples every other pixel, halving the image's width and height. A total of L pyramid layers are constructed, with L being 4. The current frame image pyramid is Pyr_t = [I_t_0, I_t_1, I_t_2, I_t_3], and the historical frame image pyramid is Pyr_{t-1} = [I_{t-1}_0, I_{t-1}_1, I_{t-1}_2, I_{t-1}_3], where index 0 represents the original resolution and index 3 represents the lowest resolution.

[0068] Step S1253: At the highest level of the current frame image pyramid and the historical frame image pyramid, initialize the optical flow vector value of the pixel at that level to zero.

[0069] The calculation begins at the highest level of the pyramid (the lowest resolution level), i.e., level 3. At the highest level, L_top=3, the image size is the smallest, and the motion amplitude is correspondingly reduced. The optical flow vector of each pixel is initialized to a zero vector (0, 0), indicating that there is an initial assumption of no motion at this level. The optical flow field data structure is a two-dimensional matrix with the same size as the current layer image, and each element is a vector (u, v) containing two floating-point numbers.

[0070] Step S1254: Starting from the highest level of the image pyramid, perform iterative calculation of optical flow vectors layer by layer down to the original resolution level. During the calculation process at each level, the optical flow vector result calculated by the previous layer is upsampled and amplified to serve as the initial estimate for the optical flow calculation of the current layer.

[0071] Starting from the highest layer L=3, the calculated optical flow vector field u_L3 is upsampled, doubling its size, and used as the initial optical flow estimate u_init_L2 for the next layer L=2. Upsampling uses bilinear interpolation to map the vector value of each point in u_L3 to a 2x2 region of the target image. Then, based on u_init_L2, the optical flow vector increment Δu_L2 is calculated at layer L=2, resulting in the final optical flow field u_L2 = u_init_L2 + Δu_L2. This process is repeated layer by layer, each time upsampling the calculation result u_{L+1} from the previous layer as the initial value u_init_L for the next layer, calculating the increment Δu_L, and obtaining the final optical flow field u_L for that layer. This process continues until the lowest layer L=0, yielding the final dense optical flow field u_L0.

[0072] Step S1255: At the current image resolution, construct constant brightness constraint equations and constant gradient constraint equations for the current layer image of the current frame image pyramid and the current layer image of the historical frame image pyramid based on the gray-level invariance assumption and the local neighborhood constant optical flow assumption.

[0073] In a certain layer of image L, for each pixel (x, y), a constant brightness constraint equation is established based on the assumption of gray-level invariance: I_t(x, y) - I_{t-1}(x+u_x, y+u_y)=0, where (u_x, u_y) is the optical flow vector at that point. Performing a first-order Taylor expansion of the equation yields I_x*Δu_x + I_y*Δu_y + I_t=0, where I_x and I_y are the image spatial gradients, I_t is the temporal gradient, and Δu is the optical flow increment. At the same time, the gradient constant constraint equation is introduced: I_{x_t}(x,y)-I_{x_{t-1}}(x+u_x,y+u_y)=0. Expanding the gradients in the x and y directions respectively, we get I_{xx}*Δu_x+I_{xy}*Δu_y+I_{xt}=0 and I_{yx}*Δu_x+I_{yy}*Δu_y+I_{yt}=0.

[0074] Step S1256: Solve the constant brightness constraint equation and the constant gradient constraint equation together to obtain the optical flow vector increment value containing the horizontal optical flow component and the vertical optical flow component.

[0075] The constant brightness constraint equation and the two constant gradient constraint equations are combined to form an overdetermined linear system of equations. Assume that within a small local neighborhood Ω, all pixels have the same optical flow vector increment (Δu_x, Δu_y). For each pixel within the neighborhood Ω, based on its grayscale value and gradient value, establish a linear equation for Δu_x and Δu_y: A_i*[Δu_x; Δu_y]=b_i, where A_i is a 2x2 matrix and b_i is a 2x1 vector. Combining all equations forms a matrix form A*[Δu_x; Δu_y]=b. By solving the normal equation (A^T*A)*[Δu_x; Δu_y]=(A^T*b), the optical flow vector increment (Δu_x, Δu_y) of the center pixel in this neighborhood is obtained. The size of the neighborhood Ω is typically 5x5 or 7x7 pixels.

[0076] Step S1257: The incremental value of the optical flow vector is accumulated and fused with the initial estimate passed down from the previous layer to update the refined optical flow vector field of the current layer.

[0077] The obtained optical flow vector increments (Δu_x, Δu_y) are added to the initial estimate u_init passed down from the previous layer to obtain the final optical flow vectors for this pixel in this layer: u_current_x = u_init_x + Δu_x, u_current_y = u_init_y + Δu_y. This process is repeated for all pixels in the image to obtain the refined optical flow vector field u_L for the current layer.

[0078] Step S1258: Repeat the step of iteratively calculating the optical flow vector layer by layer until the optical flow vector field calculation of the original resolution level is completed, and obtain the original resolution dense optical flow vector field map of the current video frame image unit relative to the previous historical video frame image unit.

[0079] Following steps S1254 to S1257, the process proceeds sequentially from the highest layer (L=3) to the lowest layer (L=0). At each layer, the calculation result from the previous layer is used as the initial value, and the optical flow increment is calculated for correction. Finally, at the lowest layer (L=0), a dense optical flow vector field map u_final with the same resolution as the original image is obtained. Each pixel in the map contains a two-dimensional optical flow vector (u_x, u_y).

[0080] Step S1259: Perform vector smoothing processing based on median filtering on the original resolution dense optical flow vector field map to eliminate isolated abnormal optical flow vectors caused by image noise or texture loss in the dense optical flow vector field map, and obtain a smoothed dense optical flow vector field map.

[0081] Median filtering is applied to the dense optical flow vector field map u_final at the original resolution. For each pixel (x, y), a filtering window is defined, for example, a 3x3 neighborhood. Optical flow vectors from all 9 pixels within the window are collected, and their horizontal component u_x and vertical component u_y sets are extracted. The u_x set is sorted, and the median value is taken as the smoothed horizontal component u_x_med for that point. The u_y set is sorted, and the median value is taken as the smoothed vertical component u_y_med. The original optical flow vector at that point is replaced with (u_x_med, u_y_med). Median filtering effectively removes abrupt and abnormal vectors while preserving edge information relatively well.

[0082] Step S12510: The smoothed dense optical flow vector field map is used as a two-dimensional optical flow vector distribution map to characterize the instantaneous motion velocity of the pixel. Each pixel position in the two-dimensional optical flow vector distribution map contains a two-dimensional optical flow vector composed of a horizontal motion velocity component and a vertical motion velocity component.

[0083] The dense optical flow vector field map after median filtering and smoothing is the final two-dimensional optical flow vector distribution map F_flow. The data structure of this two-dimensional optical flow vector distribution map is a two-dimensional matrix with the same size as the original image. Each element of the matrix is ​​a vector (u_x, u_y) containing two floating-point numbers, representing the instantaneous motion velocity of that pixel in the horizontal and vertical directions, respectively, in pixels per second. This two-dimensional optical flow vector distribution map fully describes the motion information of every point in the scene.

[0084] Step S126: Perform vector filtering on the two-dimensional optical flow vector distribution map based on the foreground pixel region mask image, retain the optical flow vector information in the foreground region corresponding to the tourist candidate object, and obtain the tourist individual motion optical flow vector map.

[0085] The final foreground pixel region mask image M_final generated in step S122 is aligned pixel-by-pixel with the two-dimensional optical flow vector distribution map F_flow generated in step S125. For each pixel (x, y) in the image, if M_final(x, y) == 1, it indicates that the point belongs to the foreground tourist candidate, and the optical flow vector (u_x, u_y) in F_flow(x, y) is retained; if M_final(x, y) == 0, it indicates that the point belongs to the background, and the optical flow vector of that point is set to the zero vector (0, 0). The processed result is a new optical flow vector map F_fg_flow, in which only the foreground region has a non-zero optical flow vector, while the optical flow vector of the background region is set to zero. This image is the optical flow vector map of individual tourist movement, with the same size as the original image, but the non-zero vectors only appear in the area where the tourist is located.

[0086] Step S127: Perform vector field segmentation processing based on connected region labeling on the motion optical flow vector map of the individual tourist, and cluster and merge all optical flow vectors belonging to the connected region of the same tourist candidate object to generate the average motion vector of each individual tourist.

[0087] The connected region label information generated in step S122 is utilized again. For each connected region R_k with a unique identifier ID_k, all pixels corresponding to it in the tourist individual motion optical flow vector map F_fg_flow are traversed, and the optical flow vectors at these points are collected. The arithmetic mean of these optical flow vectors is calculated. Specifically, the horizontal components u_x of all pixels in the region are summed to obtain sum_u_x, and the vertical components u_y are summed to obtain sum_u_y. These are divided by the total number of pixels in the region N_region_k to obtain the average optical flow vectors u_avg_x_k = sum_u_x / N_region_k and u_avg_y_k = sum_u_y / N_region_k for the tourist candidate. The average optical flow vectors (u_avg_x_k, u_avg_y_k) represent the overall instantaneous motion trend of the tourist individual in the current frame, in pixels per second.

[0088] Step S128: Perform motion trajectory association matching processing based on the average motion vector of each tourist individual and the spatial position of the tourist individual in adjacent video frame image units to obtain the displacement vector sequence of the tourist individual between consecutive video frame image units.

[0089] The tourist individuals detected in the current frame t are associated and matched with those detected in the previous frame t-1 to form continuous motion trajectories. The association matching is based on the principles of location proximity and motion vector consistency. For a tourist individual i in the current frame, its foot contact point coordinates are P_t_i, and its average optical flow vector is V_t_i. Its possible position in the previous frame is predicted as P_pred = P_t_i - V_t_i * Δt, where Δt is the frame interval. In the previous frame, all tourist individuals j within a radius R_search around this predicted position are searched. A cost matrix C is constructed, where C_ij = α * ||P_t_i - P_pred_j|| + β * (1 - cosine_similarity(V_t_i, V_{t-1}_j)), where ||·|| is the Euclidean distance, cosine_similarity is the cosine of the vector angle, and α and β are weighting coefficients. The Hungarian algorithm is used to solve this cost matrix to obtain the optimal matching pair. After a successful match, the difference between the current frame position P_t_i and the position P_{t-1}_matched of the corresponding individual in the previous frame is calculated to obtain the displacement vector D_t_i = P_t_i - P_{t-1}_matched. This displacement vector is then stored in the trajectory list of the tourist individual in chronological order to form the displacement vector sequence [D_t1, D_t2, D_t3, ...].

[0090] Step S129: Perform feature optimization processing on the displacement vector sequence in the time dimension to generate a set of tourist individual motion vector field features that characterizes the actual motion trend of individual tourists. The set of tourist individual motion vector field features includes the displacement vector direction and displacement vector magnitude of each individual tourist.

[0091] The original displacement vector sequence may contain noise and jitter, and needs to be optimized to reflect the true motion trend.

[0092] For example, step S1291: Obtain a sequence of displacement vectors associated with a specific tourist individual identifier, the sequence of displacement vectors containing multiple displacement vectors arranged in chronological order, each displacement vector consisting of a horizontal displacement component and a vertical displacement component.

[0093] From the trajectory database, extract the corresponding displacement vector sequence based on the visitor identifier ID_k. This displacement vector sequence is a list, such as [D_t1, D_t2, D_t3, ..., D_tn], where each displacement vector D_ti consists of a horizontal component dx_ti and a vertical component dy_ti, in pixels per frame.

[0094] Step S1292: Assign a time weight coefficient to each displacement vector in the displacement vector sequence. The time weight coefficient is negatively correlated with the time interval between the timestamp corresponding to the displacement vector and the current time.

[0095] Each displacement vector in the sequence is assigned a weight, with vectors closer to the current time having a larger weight. An exponential decay function is used to calculate the weights. Let the current time be t_now, the timestamp of the displacement vector D_ti be t_i, and the time interval Δt_i = t_now - t_i. Then the weight w_i = exp(-Δt_i / τ), where τ is the decay time constant, for example, 5 seconds. All weights form a weight sequence [w_1, w_2, w_3, ..., w_n], and the weight values ​​are between 0 and 1.

[0096] Step S1293: Construct a sliding time window of fixed length, extract multiple consecutive displacement vectors falling within the sliding time window from the displacement vector sequence and their corresponding time weight coefficients, and generate a vector queue to be smoothed.

[0097] Define a sliding window length N_window, set to 10. Starting from the beginning of the displacement vector sequence, extract the first N_window displacement vectors and their corresponding weight coefficients to form a vector queue to be smoothed Q=[(D_t1, w_1), (D_t2, w_2), ..., (D_tN, w_N)].

[0098] Step S1294: Perform a weighted average calculation on the horizontal displacement components in the vector queue to be smoothed. Multiply each horizontal displacement component by the corresponding time weight coefficient, sum the results, and then divide by the sum of the time weight coefficients to obtain the weighted average horizontal displacement components.

[0099] Calculate the weighted average of the horizontal components of all displacement vectors in queue Q. The weighted average horizontal component is dx_avg = (w_1*dx_t1 + w_2*dx_t2 + ... + w_N*dx_tN) / (w_1 + w_2 + ... + w_N).

[0100] Step S1295: Perform a weighted average calculation on the vertical displacement components in the vector queue to be smoothed. Multiply each vertical displacement component by the corresponding time weight coefficient, sum the results, and then divide by the sum of the time weight coefficients to obtain the weighted average vertical displacement components.

[0101] Similarly, calculate the weighted average of the vertical components. dy_avg=(w_1*dy_t1+w_2*dy_t2+...+w_N*dy_tN) / (w_1+w_2+...+w_N).

[0102] Step S1296: Combine the weighted average horizontal displacement component and the weighted average vertical displacement component into a new displacement vector, which serves as the smoothed instantaneous motion vector of the tourist individual corresponding to the center moment of the sliding time window.

[0103] Combine (dx_avg, dy_avg) into a new displacement vector D_smooth=(dx_avg, dy_avg), which is the smoothed instantaneous motion vector corresponding to the current sliding window center time t_center. t_center is the average of the earliest and latest timestamps within the window.

[0104] Step S1297: Slide the sliding time window backward along the time axis by one time unit, and repeat the step of extracting the vector queue to be smoothed and combining it into a new displacement vector until all displacement vectors in the displacement vector sequence have been traversed, resulting in a smoothed instantaneous motion vector sequence composed of multiple smoothed instantaneous motion vectors.

[0105] Move the sliding window one step backward, i.e., increment the window's starting index by one, and reconstruct a new queue Q_new=[(D_t2, w_2), (D_t3, w_3), ..., (D_t(N+1), w_{N+1})], and repeat steps S1294 to S1296 to calculate the next smoothed vector. Repeat this process until the window slides to the end of the sequence, obtaining a complete smoothed instantaneous motion vector sequence [D_smooth_1, D_smooth_2, ..., D_smooth_m].

[0106] Step S1298: Perform Kalman filter state estimation processing on the smoothed instantaneous motion vector sequence to establish a motion state space model of the tourist individual. Use the smoothed instantaneous motion vector sequence as the input of the observation value, and estimate the tourist individual's real motion speed and real motion direction at the current moment through the prediction update step.

[0107] A Kalman filter model is established to describe the motion of the tourists. The state vector is defined as X_k=[x_k, y_k, vx_k, vy_k]^T, representing the position and velocity of the tourist in the pixel coordinate system, respectively. The state transition matrix F is set as a constant velocity model: F=[[1, 0, Δt, 0]; [0, 1, 0, Δt]; [0, 0, 1, 0]; [0, 0, 0, 1]]. The observation matrix H maps the state space to the observation space: H=[[1, 0, 0, 0]; [0, 1, 0, 0]], because the observation value is the position. The observation noise covariance matrix R and the process noise covariance matrix Q are determined through offline calibration. The smoothed instantaneous motion vector sequence obtained in step S1297 is used as the position observation value Z_k and input into the Kalman filter. The filter performs the prediction step: X_pred=F*X_{k-1}, P_pred=F*P_{k-1}*F^T+Q. Then, the update step is executed: K = P_pred * H^T * (H * P_pred * H^T + R)^(-1), X_k = X_pred + K * (Z_k - H * X_pred), P_k = (IK * H) * P_pred. After multiple iterations, the vx_k and vy_k components are extracted from the final state vector X_k.

[0108] Step S1299: The latest real motion speed and real motion direction of the tourist individual obtained through the Kalman filter state estimation process are output as the tourist individual motion vector representing the real motion trend of the tourist individual.

[0109] The latest real velocity (vx_true, vy_true) estimated by Kalman filtering is used as the final motion vector of the tourist individual. The direction of this final motion vector is calculated by arctan(vy_true / vx_true), and the magnitude is calculated by √(vx_true^2+vy_true^2).

[0110] Step S12910: Organize and summarize the individual motion vectors of all tourists according to the individual tourist identifier to generate the individual tourist motion vector field feature set. Each element in the individual tourist motion vector field feature set includes the individual tourist identifier, the horizontal real motion velocity component, and the vertical real motion velocity component.

[0111] The IDs of all individual tourists and their corresponding real motion vectors (vx_true, vy_true) optimized by Kalman filtering are aggregated into a set M_motion. The data structure of this set is a list, and the list elements are tuples (ID_k, vx_true_k, vy_true_k).

[0112] Step S1210: Bind the tourist spatial location distribution feature set and the tourist individual motion vector field feature set based on the one-to-one correspondence of tourist individual identifiers to generate a tourist individual spatiotemporal state descriptor set with dual attributes of spatial location and motion state.

[0113] The tourist spatial location distribution feature set generated in step S124 and the tourist individual motion vector field feature set generated in step S129 are associated and bound through a common tourist identifier ID. For each ID_k, its foot contact point coordinates (x_feet_k, y_feet_k) at time t and its motion vector (vx_true_k, vy_true_k) are merged to form a spatiotemporal state descriptor S_k(t) = (ID_k, x_feet_k, y_feet_k, vx_true_k, vy_true_k). All tourist descriptors are merged to generate a set of tourist individual spatiotemporal state descriptors for the current monitoring period.

[0114] Step S130: Call the pre-configured amusement park area spatial semantic map data, perform regional spatial coordinate mapping and alignment processing on the tourist spatial location distribution feature set and the tourist individual motion vector field feature set, and generate a real-time tourist density distribution grid data matrix of each amusement park area based on the geographic coordinate system during the current monitoring period.

[0115] Tourist information in pixel coordinates is mapped onto a spatial semantic map with actual geographical significance.

[0116] Step S131: Obtain pre-configured spatial semantic map data of the amusement park area. The spatial semantic map data of the amusement park area includes a set of boundary geographic coordinate polygons of each area of ​​the amusement park, the geographic center point coordinates of each area, the functional attribute labels of each area, and the homography transformation matrix parameters that map pixel coordinate system points to geographic coordinate system points.

[0117] Acquire pre-constructed spatial semantic map data for the amusement park. This spatial semantic map data includes: the boundary of the "Adventure Island" area is a polygon P_adventure composed of a set of geographic latitude and longitude coordinates; the latitude and longitude coordinates of the center point of the area (C_lon_adventure, C_lat_adventure); the functional attribute label of the area is "Themed Amusement Area"; and a 3*3 homography transformation matrix H_pixel_to_geo. The parameters of this matrix are calculated through pre-calibration of cameras and alignment of ground control points, and are used to map the pixel coordinates (u, v) of points on the ground in the image (such as the foot contact point) to geographic coordinates (lon, lat).

[0118] Step S132: Extract the pixel coordinates of the foot contact point of each tourist individual from the tourist spatial location distribution feature set, and use the homography transformation matrix parameters to perform perspective projection transformation on the pixel coordinates of the foot contact point to convert the pixel coordinates of the foot contact point into a geographic spatial location point represented by geographic latitude and longitude coordinates.

[0119] From the tourist spatial location distribution feature set in step S124, extract the pixel coordinates (x_feet, y_feet) of the foot contact point of each tourist individual. Convert them into homogeneous coordinate form (x_feet, y_feet, 1). Then, transform them using the homography transformation matrix H_pixel_to_geo: calculate the temporary vector T=H_pixel_to_geo*[x_feet, y_feet, 1]^T. Finally, normalize T to obtain the geographic coordinates: geographic longitude lon=T[0] / T[2], geographic latitude lat=T[1] / T[2]. Thus, the geographic spatial location point (lon, lat) of the tourist is obtained.

[0120] Step S133: Perform spatial inclusion relationship judgment processing on the geospatial location points of all individual tourists obtained by conversion and the geographical coordinate polygons of each area boundary in the spatial semantic map data of the amusement park area to determine the specific amusement park area identifier where each individual tourist is currently located.

[0121] Iterate through the geographic locations (lon_k, lat_k) of each visitor. Using either the ray casting method or the classic PNPoly algorithm, determine whether the point is located inside a region boundary polygon P_region. For example, if the point (lon_k, lat_k) falls inside the polygon P_adventure, then the visitor's current region identifier is "Adventure Island". This process assigns a region label, Region_ID_k, to each visitor.

[0122] Step S134: Based on the boundary geographic coordinate polygons of each amusement park area, each amusement park area is spatially divided into a regional grid composed of multiple regular geographic grid units. Each geographic grid unit has a fixed geographic latitude and longitude span and a unique grid index code.

[0123] Taking the "Adventure Island" area as an example, the area is divided into a regular grid based on the smallest bounding rectangle of its boundary polygon. Each grid cell has a longitude span of Δlon and a latitude span of Δlat, for example, both Δlon and Δlat are 0.0001 degrees. Starting from the minimum longitude lon_min, a series of longitude boundary points are generated with a step size of Δlon; similarly, starting from the latitude lat_min, a series of latitude boundary points are generated with a step size of Δlat. These boundary points divide the area into a grid of I rows and J columns. Each grid cell has a unique index code (i, j), where i is the row index and j is the column index.

[0124] Step S135: Based on the specific geographic grid cell into which each tourist's geographic location falls, count the number of tourist individuals included in the current monitoring period within each geographic grid cell, and use this count as the initial tourist count for that geographic grid cell.

[0125] Iterate through all the geographical locations of tourists (lon_k, lat_k). For each point, calculate its grid cell index: i = floor((lat_k - lat_min) / Δlat), j = floor((lon_k - lon_min) / Δlon). Accumulate the tourist counts into the corresponding grid cells to obtain the initial tourist count matrix C_init, where C_init[i][j] represents the number of tourists in the i-th row and j-th column grid cell.

[0126] Step S136: Perform spatial diffusion smoothing processing on the initial tourist count based on neighboring grid cells, and allocate a portion of the tourist count of each geographic grid cell to the eight surrounding neighboring geographic grid cells according to distance weights to obtain the spatially smoothed tourist count distribution.

[0127] The initial counting matrix C_init is Gaussian smoothed. A 3*3 Gaussian kernel G is defined, for example, G=[[1 / 16, 2 / 16, 1 / 16]; [2 / 16, 4 / 16, 2 / 16]; [1 / 16, 2 / 16, 1 / 16]]. For each grid cell (i, j), its smoothed count value C_smooth[i][j]=Σ_{m=-1}^{1}Σ_{n=-1}^{1}(C_init[i+m][j+n]*G[m+1][n+1]). For cells at the boundary, only the effective neighborhood is weighted and averaged. This operation can eliminate the count abrupt changes caused by discretization, making the density distribution more continuous.

[0128] Step S137: Extract the horizontal and vertical true motion velocity components and corresponding time interval information of each tourist individual from the tourist individual motion vector field feature set. Combine this with the pixel coordinates of the tourist individual's foot contact point to calculate the tourist individual's displacement vector in the pixel coordinate system. Use the homography transformation matrix parameters to map the displacement vector to the geographic coordinate system to obtain the displacement vector in the geographic coordinate system. Based on the displacement vector in the geographic coordinate system and the time interval information, calculate the tourist motion velocity vector based on the geographic coordinate system.

[0129] Extract the motion vector (vx_true_k, vy_true_k) for each tourist individual ID_k from M_motion, in pixels per second. Multiply this motion vector by the time interval Δt (i.e., the time difference between two adjacent frames) to obtain the displacement vector D_pixel_k = (vx_true_k * Δt, vy_true_k * Δt) in the pixel coordinate system. Convert D_pixel_k to homogeneous coordinates (dx_pixel, dy_pixel, 0). Calculate the displacement vector in the geographic coordinate system using the homography transformation matrix H_pixel_to_geo: First, calculate two points: P1 = H_pixel_to_geo * [x_feet_k, y_feet_k, 1]^T, P2 = H_pixel_to_geo * [x_feet_k + dx_pixel, y_feet_k + dy_pixel, 1]^T. After normalization, we obtain the geographic coordinates P1_geo=(lon1, lat1) and P2_geo=(lon2, lat2). Then, the displacement vector in the geographic coordinate system is D_geo_k=(lon2-lon1, lat2-lat1). Finally, the velocity vector in the geographic coordinate system is V_geo_k=D_geo_k / Δt, in degrees per second.

[0130] Step S138: Perform vector summation and numerical averaging on the tourist movement speed vectors of all individual tourists falling into each geographic grid cell to calculate the average movement speed vector of each geographic grid cell. The average movement speed vector includes the average movement direction and the average movement speed.

[0131] For each grid cell (i, j), collect the geographic velocity vectors V_geo_k of all individual tourists falling within this cell. Calculate the vector sum of these vectors: sum_vx = Σvx_geo_k, sum_vy = Σvy_geo_k. The number of tourists within the cell is N_cell. Then the average velocity vector V_avg_geo[i][j] = (sum_vx / N_cell, sum_vy / N_cell). The direction of this average velocity vector is given by arctan(avg_vy / avg_vx), and its magnitude is given by √(avg_vx^2 + avg_vy^2).

[0132] Step S139: The spatially smoothed visitor count and the average movement speed vector are used as two attribute values ​​for each geographic grid unit, and filled into the corresponding positions of a two-dimensional matrix according to the grid index encoding, to generate a real-time visitor density distribution grid data matrix of each amusement park area based on the geographic coordinate system during the current monitoring period.

[0133] The smoothed visitor count C_smooth[i][j] obtained in step S136 and the average velocity vector V_avg_geo[i][j] obtained in step S138 are used as the two attribute values ​​for each grid cell. These are then filled into a two-dimensional matrix Density_map according to the grid index (i, j). The rows of the matrix correspond to the grid indices in the latitude direction, and the columns correspond to the grid indices in the longitude direction. Each matrix element is a data structure containing two fields: {"count":C_smooth[i][j], "velocity":(avg_vx, avg_vy)}. This matrix is ​​the real-time visitor density distribution grid data matrix, which fully describes the density and movement status of each area during the current monitoring period.

[0134] Step S140: Perform tourist flow trend analysis on the real-time tourist density distribution grid data matrix to obtain the tourist density change rate characteristic field and tourist flow main direction characteristic vector set for each amusement park area.

[0135] Based on the density distribution grid data matrix, the dynamic trend of tourist flow is analyzed.

[0136] Step S141: Extract the real-time density distribution grid data matrix of tourists from the real-time density distribution grid data matrix of tourists for multiple consecutive historical monitoring periods to form a three-dimensional density grid data tensor with a time dimension. The three dimensions of the three-dimensional density grid data tensor are the grid row index, the grid column index and the time index.

[0137] Extract a grid data matrix of real-time visitor density distribution for T consecutive time periods, including the current monitoring period, from the historical database. For example, T = 10 time periods, with each time period spaced 1 minute apart. Stack the above matrix in chronological order to form a three-dimensional data tensor D_tensor, with dimensions I (number of rows) * J (number of columns) * T (time index). The element D_tensor[i][j][t] in the tensor represents the visitor count in grid cell (i, j) at time period t.

[0138] Step S142: Perform adjacent time period difference calculation processing on the three-dimensional density grid data tensor in the time dimension. For each fixed grid cell position, calculate the difference between the tourist count in the current monitoring period and the tourist count in the previous historical monitoring period to obtain the tourist count time change rate of the grid cell.

[0139] For each grid cell (i, j), calculate the count difference between adjacent time periods. Take the count of the current time period t_cur as C_cur = D_tensor[i][j][t_cur], and the count of the previous time period t_prev as C_prev = D_tensor[i][j][t_prev]. Then, the tourist count change rate over time R_change[i][j] = C_cur - C_prev. This tourist count change rate over time reflects the increase or decrease in the number of tourists in this grid cell during the most recent time period; a positive value indicates an increase, and a negative value indicates a decrease.

[0140] Step S143: Fill the calculated tourist count change rate of all grid cells into a two-dimensional matrix with the same size as the real-time tourist density distribution grid data matrix according to the corresponding grid index, and generate a tourist density change rate feature field. The value of each grid cell in the tourist density change rate feature field represents the increase or decrease in the number of tourists at the corresponding location.

[0141] The R_change[i][j] values ​​of all grid cells calculated in step S142 are filled into a new two-dimensional matrix R_field according to their indices (i, j). The size of R_field is the same as that of Density_map, and each element is a floating-point number representing the rate of change of the number of visitors in that grid cell. R_field is the visitor density change rate feature field, which depicts the instantaneous change trend of density at each location during the current time period.

[0142] Step S144: Perform autoregressive moving average model fitting on the time series of visitor counts at each grid cell position in the three-dimensional density grid data tensor, extract the trend term and seasonal term of the visitor count time series, and predict the direction of change in visitor counts in the short term based on the slope of the trend term.

[0143] For each grid cell (i, j), extract its time series data C_t = D_tensor[i][j][:], which is a sequence of length T. Apply an autoregressive integrated moving average (ARIMA) model to fit this sequence. First, make the sequence stationary through differencing operations. Then, determine the orders p and q of the model based on the autocorrelation function and partial autocorrelation function. Establish an ARIMA(p, d, q) model, where d is the differencing order. Solve for the model parameters using maximum likelihood estimation. From the fitted model, decompose the time series into a trend term Trend_t, a seasonal term Season_t, and a residual term Residual_t. Calculate the slope k_trend of the trend term at the most recent few time points. If k_trend > 0, predict that the future count will increase; if k_trend < 0, predict that the future count will decrease.

[0144] Step S145: Identify grid cells in the aggregation hot spot areas where the temporal change rate of tourist count is positive and grid cells in the evacuation cold spot areas where the temporal change rate of tourist count is negative, based on the tourist density change rate feature field and the predicted change direction of the tourist count time series.

[0145] Traverse all grid cells in R_field. For each cell (i, j), if R_change[i][j] > Th_hot, where Th_hot is a preset hot spot threshold, and the predicted future change direction in Step S144 is an increase, then mark this grid cell as an aggregation hot spot area and add it to the hot spot set H_set. If R_change[i][j] < Th_cold, where Th_cold is a preset cold spot threshold (negative value), and the predicted future change direction in Step S144 is a decrease, then mark this grid cell as an evacuation cold spot area and add it to the cold spot set C_set.

[0146] Step S146: Extract the average motion speed vectors of all grid cells from the tourist real-time density distribution grid data matrix to construct a two-dimensional vector field, where each grid cell position in the two-dimensional vector field contains a motion vector composed of the average motion direction angle and the magnitude of the average motion speed.

[0147] From Density_map, extract the average motion speed vector V_avg_geo[i][j] = (avg_vx, avg_vy) of each grid cell (i, j). Construct a two-dimensional vector field V_field with these vectors of all grid cells. This two-dimensional vector field has the same number of rows and columns as the grid, and each position (i, j) corresponds to a two-dimensional vector (avg_vx, avg_vy), which describes the average motion state of tourists at this position.

[0148] Step S147: Perform flow field feature analysis processing on the two-dimensional vector field based on curl calculation and divergence calculation, calculate the curl distribution map of the two-dimensional vector field to identify the rotation center region of tourist flow, and calculate the divergence distribution map of the two-dimensional vector field to identify the convergence region and divergence region of tourist flow.

[0149] Curl and divergence are calculated for the two-dimensional vector field V_field. For each grid cell (i, j), the spatial derivative is calculated using the central difference method: ∂vx / ∂x≈(avg_vx[i][j+1]-avg_vx[i][j-1]) / (2*Δx), where Δx is the actual distance corresponding to the grid longitude span; ∂vy / ∂y≈(avg_vy[i+1][j]-avg_vy[i-1][j]) / (2*Δy). Curl[i][j]=∂vy / ∂x-∂vx / ∂y, the magnitude of curl represents the local rotation intensity, and positive and negative indicate the rotation direction. Divergence Div[i][j]=∂vx / ∂x+∂vy / ∂y, positive divergence indicates divergence (people flow outward), and negative divergence indicates convergence (people flow inward). Fill the two-dimensional matrix with Curl and Div respectively to obtain the curl distribution map Curl_map and the divergence distribution map Div_map.

[0150] Step S148: Based on the curl distribution map and the divergence distribution map, the two-dimensional vector field is divided into multiple sub-regions with similar flow patterns, and the motion vector directions within each sub-region are consistent.

[0151] Based on curl and divergence maps, and combined with vector directions, the V_field is segmented using a region growing algorithm or a watershed algorithm. Seed points are initialized: grid cells with an absolute curl value greater than the threshold Th_curl or an absolute divergence value greater than the threshold Th_div are selected as seeds. Starting from each seed, expansion is made to adjacent grid cells, merging neighborhoods that meet the following conditions: the angle between the neighborhood vector direction and the average vector direction of the seed region is less than the angle threshold Th_angle (e.g., 30°), and the curl or divergence has the same sign. Through iterative expansion, the entire vector field is finally segmented into multiple sub-regions R_flow_k, where the visitor flow pattern within each sub-region is consistent; for example, the vector directions within the same sub-region are roughly the same, or they exhibit a common rotation / convergence pattern.

[0152] Step S149: For each sub-region, perform vector superposition and averaging on the motion vectors of all grid cells within the sub-region to calculate a composite vector representing the overall tourist flow trend of the sub-region, which serves as the main direction feature vector of tourist flow in the sub-region.

[0153] For each segmented sub-region R_flow_k, collect the average velocity vector V_avg_geo[i][j] of all grid cells within that region. These vectors are then superimposed: sum_vx_k = Σavg_vx, sum_vy_k = Σavg_vy. The number of grid cells within the region is N_k. The resulting vector V_main_k = (sum_vx_k / N_k, sum_vy_k / N_k). This vector represents the main direction feature vector of visitor flow in sub-region R_flow_k; its direction represents the overall flow direction within the region, and its magnitude represents the average flow intensity.

[0154] Step S1410: Associate and store the boundary geographic coordinate range of each sub-region with the corresponding tourist flow main direction feature vector to generate a set of tourist flow main direction feature vectors covering the entire amusement park area.

[0155] For each sub-region R_flow_k, record its boundary geographic coordinate range, which is the minimum and maximum latitude and longitude of all grid cells constituting the sub-region: lon_min_k, lon_max_k, lat_min_k, lat_max_k. Associate the boundary range with the main direction feature vector V_main_k calculated in step S149 to form a data structure Region_flow_k={"bbox":[lon_min_k, lon_max_k, lat_min_k, lat_max_k], "main_direction":(vx_main_k, vy_main_k)}. Organize this data from all sub-regions into a list Flow_region_set, which is the set of main direction feature vectors for visitor flow covering the entire amusement park area.

[0156] Step S150: Based on the real-time density distribution grid data matrix of tourists, the tourist density change rate feature field, and the tourist flow main direction feature vector set, generate a region-by-region flow guidance signal encoding sequence for driving the indicator light arrays distributed in various areas of the amusement park to provide differentiated dynamic indication.

[0157] Based on density distribution, changing trends, and flow direction, signal codes are generated to control the indicator light array.

[0158] Step S151: Based on the tourist count scalar value of each grid cell in the real-time tourist density distribution grid data matrix, perform density level division processing on the grid cell, divide the tourist count scalar value into multiple consecutive density level intervals according to the numerical value, and assign a unique corresponding color code index to each density level interval.

[0159] Traverse all grid cells in the Density_map to obtain the count value of each cell. Set density level thresholds, such as Th_dense1, Th_dense2, and Th_dense3. Divide the levels according to the interval where the count value is located: If count < Th_dense1, the level is L1; if Th_dense1 ≤ count < Th_dense2, the level is L2; if Th_dense2 ≤ count < Th_dense3, the level is L3; if count ≥ Th_dense3, the level is L4. Assign color coding indices to each level: L1 corresponds to the green index C_green, L2 corresponds to the yellow index C_yellow, L3 corresponds to the orange index C_orange, and L4 corresponds to the red index C_red.

[0160] Step S152: According to the change rate of tourist count over time of each grid cell in the tourist density change rate feature field, perform trend level division processing on the change trend of the grid cell, divide the change rate of tourist count over time into multiple continuous change trend level intervals according to the numerical size, and assign a unique corresponding flashing frequency coding index to each change trend level interval.

[0161] Obtain the change rate R_change of each grid cell from the density change rate feature field R_field. Set change rate thresholds, such as Th_inc1, Th_inc2, and Th_dec1. Divide the trend levels according to R_change: If R_change < Th_dec1, it is the fast evacuation trend F1; if Th_dec1 ≤ R_change < 0, it is the slow evacuation trend F2; if 0 ≤ R_change < Th_inc1, it is the slow aggregation trend F3; if R_change ≥ Th_inc1, it is the fast aggregation trend F4. Assign flashing frequency coding indices to each trend level: F1 corresponds to the low-frequency flashing index B_slow, F2 corresponds to the medium-low-frequency flashing index B_medium_slow, F3 corresponds to the medium-high-frequency flashing index B_medium_fast, and F4 corresponds to the high-frequency flashing index B_fast.

[0162] Step S153: According to the tourist flow main direction feature vector of each sub-region in the tourist flow main direction feature vector set, quantify the tourist flow main direction feature vector into multiple preset standard direction codings, and each standard direction coding corresponds to a specific azimuth direction.

[0163] The main direction vector V_main = (vx_main, vy_main) for each sub-region is obtained from the feature vector set Flow_region_set. Its direction angle θ_main = arctan(vy_main / vx_main) is calculated. This direction angle is quantized into eight standard directions: North, Northeast, East, Southeast, South, Southwest, West, and Northwest. For example, if θ_main is between -22.5° and 22.5°, it is encoded as the East direction index D_east; if it is between 22.5° and 67.5°, it is encoded as the Northeast direction index D_northeast, and so on. Each standard direction code corresponds to a unique binary code.

[0164] Step S154: Perform statistical mode analysis on the color coding index of all grid cells in each amusement park area to determine the regional master color coding index that represents the overall density of the amusement park area.

[0165] For each amusement park region (Region_r), collect the color-coded indices of all grid cells within that region. Count the frequency of each color code and select the color code index with the highest frequency as the primary color-coded index (C_region_r) for that region. If multiple modes exist, choose the one with the higher density level.

[0166] Step S155: Perform a weighted average of the flicker frequency coding indices of all grid cells within each amusement park area, using the visitor count scalar value of each grid cell as the weight, to calculate the regional main flicker frequency coding index representing the overall density change trend of the amusement park area.

[0167] For each region_r, collect the flicker frequency encoding index B_i and the corresponding visitor count count_i for all grid cells within that region. Map the flicker frequency encoding indices to numerical values, for example, B_slow maps to 1, B_medium_slow to 2, B_medium_fast to 3, and B_fast to 4. Calculate the weighted average B_avg = Σ(B_val_i * count_i) / Σcount_i. Round B_avg to the nearest integer and then map it back to the flicker frequency encoding index to obtain the main flicker frequency encoding index B_region_r for the region.

[0168] Step S156: Combine the region main color encoding index and the region main flashing frequency encoding index for encoding to generate a first-level flow guidance signal code element. The first-level flow guidance signal code element is used to control the basic color and basic flashing mode of the corresponding region indicator light array.

[0169] The C_region_r and B_region_r are combined for encoding. For example, an 8-bit binary symbol is defined, with the high 4 bits storing the color encoding index and the low 4 bits storing the flashing frequency encoding index. Specifically, the first-level symbol Symbol1_r = (C_region_r << 4) | B_region_r, where << represents a left shift operation and | represents a bitwise OR operation. This symbol is directly mapped to the control instructions of the indicator light array, determining the color of the light and the basic flashing frequency.

[0170] Step S157: Obtain the main direction feature vectors of tourist flow in the surrounding areas that are spatially adjacent to the amusement park area from the set of main direction feature vectors of tourist flow. Analyze whether the main direction feature vectors of tourist flow in the surrounding areas point to the interior of the amusement park area. If they point to the interior, generate a directional warning identifier.

[0171] For each region_r, obtain the set of all its spatially adjacent regions, Neighbor_r. For each adjacent region N, obtain its main direction feature vector of visitor flow, V_main_N. Determine whether the direction of V_main_N points into the region_r. Specifically, calculate the vector V_N_to_R from the center of region N to the center of region R, and calculate the angle θ between V_main_N and V_N_to_R = arccos((V_main_N·V_N_to_R) / (|V_main_N|*|V_N_to_R|)). If θ < 45°, it is considered that the flow of people is from region N to region R. If multiple adjacent regions point to region_r, a directional warning identifier Alert_r is generated, encoded as 1; otherwise, it is 0.

[0172] Step S158: Based on the directional warning identifier, retrieve the enhanced indicator signal code that matches the directional warning identifier from the preset enhanced indicator mode library. The enhanced indicator signal code is used to drive the corresponding area indicator light array to superimpose enhanced indicator light effects on the basic color and basic flashing mode.

[0173] If Alert_r == 1, it indicates an influx of people from an adjacent area, requiring enhanced indication. The enhanced indication signal code Enhance_r is retrieved from a pre-defined enhanced indication mode library. This library predefines various enhanced lighting effects, such as the "breathing light effect" code E_breath, the "flowing light effect" code E_flow, and the "strobe effect" code E_flash. The appropriate enhancement code is selected based on the severity and direction of the warning. For example, if there are three or more areas pointing in different directions, E_flash is selected.

[0174] Step S159: The region main color encoding index, the region main flashing frequency encoding index, the standard direction encoding, and the enhanced indicator signal encoding are encapsulated and combined according to a predefined region-by-region flow guidance indicator signal frame structure to generate a binary signal encoding sequence containing multiple fields.

[0175] Define the signal frame structure, such as frame header + region ID + color code + blink code + direction code + enhancement code + checksum. The frame header is a fixed 8-bit binary sequence used to identify the start of the signal. The region ID is an 8-bit code. The color code and blink code constitute the first-level code. The direction code is 8 bits, with each bit representing whether a standard direction is the primary flow direction. The enhancement code is 8 bits. The checksum uses the CRC8 algorithm to calculate the check value of all preceding fields. All fields are concatenated sequentially to generate the complete binary signal encoding sequence Signal_r.

[0176] Step S1510: Associate and bind the binary signal encoding sequence with the network address identifiers of the indicator light arrays distributed in various areas of the amusement park to generate a region-by-region flow guidance signal encoding sequence for driving the indicator light arrays to perform differentiated dynamic indications, and send the region-by-region flow guidance signal encoding sequence to the corresponding indicator light array controller through the communication network.

[0177] The signal encoding sequence Signal_r generated for each area is associated and bound to the network address identifier MAC_r of the indicator light array in that area. A control command list is generated, with each element being (MAC_r, Signal_r). This list is encapsulated into a network data packet using the TCP / IP protocol stack and sent to the indicator light array controller of the corresponding area via the amusement park's internal wireless network or industrial Ethernet. After receiving the command, the controller parses the Signal_r field and drives the LED light array to emit light according to the specified color, flashing frequency, and enhancement effect, thereby achieving dynamic guidance of pedestrian flow.

[0178] For example, the step of generating a region-by-region flow guidance signal encoding sequence for driving the indicator light array distributed in various areas of the amusement park to provide differentiated dynamic indication based on the real-time density distribution grid data matrix of tourists, the tourist density change rate feature field, and the tourist flow main direction feature vector set, further includes: step S160: the step of dynamically optimizing and adjusting the region-by-region flow guidance signal encoding sequence.

[0179] Step S161: Extract the real-time density distribution grid data matrix of tourists from the real-time density distribution grid data matrix of tourists for multiple consecutive historical monitoring periods to form a time series dataset for predicting future density distribution.

[0180] From the historical database, extract a grid data matrix of real-time tourist density distribution for M consecutive historical time periods, including the current time, where M is 20 and each time period is 1 minute apart. Arrange the above matrix in chronological order to form a time series dataset TS_data, which will be used for subsequent predictive modeling.

[0181] Step S162: Perform seasonal decomposition on the time series of tourist counts for each grid cell in the time series dataset to separate the long-term trend component, periodic fluctuation component, and irregular random component of tourist counts.

[0182] For each grid cell (i, j), its counting time series C_t = TS_data[i][j][:] is extracted. The STL decomposition algorithm is used to perform seasonal decomposition on this series. The STL algorithm decomposes the series into a trend term (Trend_t), a season term (Season_t), and a residual term (Residual_t) through iterations of the inner and outer loops. The trend term reflects the long-term direction of change in tourist numbers, the season term reflects regular fluctuations with a period of days or hours, and the residual term represents random noise.

[0183] Step S163: Based on the long-term trend component and periodic fluctuation component, a prediction model based on a time convolutional network is used to predict the number of tourists in each grid cell within a preset future time period, thereby obtaining a future density distribution prediction grid data matrix.

[0184] A temporal convolutional network (TCNN) is constructed as the prediction model. This TCNN consists of multiple residual blocks, each containing two layers of dilated causal convolutions, with the dilation factor increasing layer by layer to expand the receptive field. ReLU activation and Dropout layers follow the convolutional layers. The network input is the count sequence of the past L time steps, and the output is the prediction sequence of the next K time steps. During training, supervised learning is performed using input-output pairs from historical data, with mean squared error as the loss function. The Adam optimizer is used, with an initial learning rate of 0.001. After training, for the current time step, the count sequence of each grid cell from the past L time steps is input into the network to obtain the predicted count for that cell in the next K time steps. The prediction results for all cells are then filled according to the grid index to obtain the future density distribution prediction grid data matrix Pred_density.

[0185] Step S164: Extract the motion vector of each tourist individual from the set of tourist individual motion vector field features, and combine it with the set of tourist flow main direction feature vectors to construct a tourist flow simulation model based on fluid dynamics.

[0186] The motion vector of each individual tourist is considered as the particle velocity in the flow field. A fluid simulation model based on the Navier-Stokes equations is constructed, using a two-dimensional vector field V_field as the initial velocity field and a density distribution Density_map as the initial density field. The model is discretized on a mesh using the finite difference method. The pressure field P is defined, and mass conservation is satisfied by solving the pressure Poisson equation ∇^2P=∇·V. Then, the velocity field is updated using the momentum equation: ∂V / ∂t=-(V·∇)V-(1 / ρ)∇P+ν∇^2V, where ρ is the density and ν is the viscosity coefficient.

[0187] Step S165: Input the predicted future density distribution grid data matrix as the initial condition into the tourist flow simulation model, and simulate the flow process of tourists between different areas in the future time period through iterative calculation to obtain the simulated density distribution grid data matrix sequence for each monitoring time in the future time period.

[0188] The future density distribution prediction grid data matrix Pred_density obtained in step S163 is used as the initial density field for simulation, and the current vector field V_field is used as the initial velocity field. The simulation time step is set to Δt_sim, with the total simulation duration corresponding to the next K time steps. At each simulation time step, the velocity field and density field are updated according to the fluid dynamics equations, and the density distribution at future times is iteratively calculated. After the simulation, the simulated density distribution grid data matrix sequence Sim_density_seq for each monitoring time within the future time period is obtained, with a sequence length of K.

[0189] Step S166: Perform bottleneck area identification processing on the simulated density distribution grid data matrix sequence, analyze the tourist count change curve of each grid cell during the simulated period, and identify grid cells whose tourist count continuously exceeds the preset capacity threshold as potential congestion bottleneck areas.

[0190] Iterate through the density matrix at each time step in Sim_density_seq. For each grid cell (i, j), extract its count sequence C_future[t] for the next K time steps. Set a capacity threshold Th_capacity, which is determined based on the physical space area and maximum number of people the grid cell can accommodate. If C_future[t] > Th_capacity for T_cont consecutive time steps, then mark the grid cell as a potential congestion bottleneck area and add it to the bottleneck set Bottleneck_set.

[0191] Step S167: Based on the geographical location and occurrence time of the potential congestion bottleneck area, and combined with the set of feature vectors of the main direction of tourist flow, trace back the tourist originating area that caused the congestion, and generate an advance diversion demand identifier for the originating area.

[0192] For each bottleneck region B in Bottleneck_set, obtain its occurrence time t_b and geographical location. Query the Flow_region_set of the main flow direction feature vectors at time t_b, and find all regions whose main directions point to B; these regions are the visitor inflow regions Source_set. For each region S in Source_set, generate a pre-direction demand identifier Demand_S, indicating that the flow of people from S to B needs to be pre-directed before congestion occurs.

[0193] Step S168: Perform association mapping processing between the advance diversion demand identifier and the area main color code index and area main flashing frequency code index in the area diversion indication signal encoding sequence of the current monitoring period to determine the amusement park area that needs to adjust the diversion indication signal in advance and its adjustment range parameters.

[0194] For each region S with a pre-emptive diversion demand indicator (Demand_S), its main color code index (C_S) and main flashing frequency code index (B_S) are extracted from the region-by-region diversion indication signal encoding sequence during the current monitoring period. The adjustment magnitude parameters are determined based on the severity of congestion and the expected occurrence time. For example, if congestion is expected to occur in 10 minutes and is severe, the color is shifted one level to a higher warning level (e.g., from yellow to orange), and the flashing frequency is increased one level (e.g., from medium frequency to high frequency).

[0195] Step S169: Based on the adjustment amplitude parameter, the first-level flow guidance signal code element of the amusement park area that needs to adjust the flow guidance signal in advance is corrected to enhance the basic color saturation and basic flashing frequency of the indicator light array in the area and generate the corrected first-level flow guidance signal code element.

[0196] For the region S that needs adjustment, the first-level code is corrected according to the adjustment range determined in step S168. The color encoding index C_S is adjusted to C_S', for example, from C_yellow to C_orange. The flashing frequency encoding index B_S is adjusted to B_S', for example, from B_medium_fast to B_fast. The corrected first-level code Symbol1_S'=(C_S'<<4)|B_S' is generated.

[0197] Step S1610: The modified first-level diversion indicator signal code and the enhanced indicator signal code are re-encapsulated and combined to generate an updated region-by-region diversion indicator signal encoding sequence containing advance diversion information.

[0198] The corrected first-level symbol Symbol1_S' is re-encapsulated with the original standard direction code and enhanced indicator signal code, and an updated signal code sequence Signal_S' is generated according to the signal frame structure of step S159. For other regions that do not require adjustment, the original signal remains unchanged. The updated signals of all regions are combined to form the updated region-by-region guiding indicator signal code sequence.

[0199] Step S1611: The updated zone-by-zone diversion indication signal encoding sequence is sent to the corresponding indicator light array controller through the communication network, driving the indicator light array to perform diversion indication actions in advance before the arrival of the future preset time period, so as to guide tourists to avoid the potential congestion bottleneck area.

[0200] The updated signal encoding sequence is associated and bound with the network address identifiers of the indicator light arrays in each area to generate the final control command list. The commands are then sent to the corresponding indicator light array controllers via the communication network. After parsing the commands, the controllers drive the indicator light arrays to illuminate in advance according to the corrected color and flashing frequency. For example, areas that originally displayed yellow are changed to orange and flash at a high frequency to warn visitors of potential congestion ahead and guide them to choose alternative routes, thus proactively managing pedestrian flow before congestion occurs.

[0201] In one exemplary embodiment, a vision-based real-time crowd density monitoring system for amusement parks is provided. This system can be a terminal, server, etc., and its internal structure diagram can be as follows: Figure 2As shown, this vision-based amusement park real-time crowd density monitoring system includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, near-field communication, or other technologies. When the computer program is executed by the processor, it implements a vision-based real-time crowd density monitoring method for amusement parks. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device can be a touch layer covering the display screen, or a button, trackball, or touchpad set on the shell of a vision recognition-based amusement park real-time crowd density monitoring system, or an external keyboard, touchpad, or mouse, etc.

[0202] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. A method for real-time crowd density monitoring in amusement parks based on visual recognition, characterized in that, The method includes: Acquire a collection of multi-view video stream data of the amusement park area during the current monitoring period, which is synchronously collected by the video acquisition equipment array deployed in various areas of the amusement park; Based on the video stream data set, stereoscopic visual features of the amusement park scene are extracted to obtain a set of spatial location distribution features of tourists based on pixel coordinates and a set of individual tourist motion vector field features. The pre-configured spatial semantic map data of the amusement park area is called, and the spatial coordinate mapping and alignment processing of the tourist spatial location distribution feature set and the tourist individual motion vector field feature set is performed to generate a real-time tourist density distribution grid data matrix of each amusement park area based on the geographic coordinate system during the current monitoring period. The real-time density distribution grid data matrix of tourists is processed by tourist flow trend analysis to obtain the characteristic field of tourist density change rate and the set of characteristic vectors of tourist flow main direction in each amusement park area. Based on the real-time density distribution grid data matrix of tourists, the tourist density change rate feature field, and the tourist flow main direction feature vector set, a region-by-region guidance indicator signal encoding sequence is generated to drive the indicator light arrays distributed in various areas of the amusement park to provide differentiated dynamic indications.

2. The method for real-time crowd density monitoring in amusement parks based on visual recognition according to claim 1, characterized in that, The step of extracting stereoscopic visual features of the amusement park scene based on the video stream data set to obtain a set of spatial location distribution features of tourists based on pixel coordinates and a set of individual tourist motion vector field features includes: Each video stream in the video stream data set is demultiplexed to obtain multiple amusement park area video frame image units with synchronization timestamps. The video frame image units of the amusement park area are subjected to initial segmentation processing to separate the foreground and background, resulting in a foreground pixel region mask image containing the outline of a suspected tourist. Each connected region in the foreground pixel region mask image corresponds to a two-dimensional projection region of a tourist candidate object in the pixel coordinate system. For each connected region in the foreground pixel region mask image, a fine-fitting process for contour edges is performed. The pixel coordinates of the head vertex and the pixel coordinates of the foot contact point of the tourist candidate object corresponding to the connected region are extracted to generate a pair of two-dimensional spatial positioning points of the tourist individual in the pixel coordinate system. Based on the pixel coordinates of the foot contact points in the two-dimensional spatial location points of the tourist individuals in the pixel coordinate system, a tourist spatial location distribution feature set marked by pixel coordinate values ​​is generated. The tourist spatial location distribution feature set contains the sequence of foot contact point pixel coordinates of all detected tourist individuals during the current monitoring period. The optical flow field of the video frame image unit of the amusement park area in the continuous time series is calculated by the pixel gray level change between adjacent frames to obtain a two-dimensional optical flow vector distribution map representing the instantaneous motion velocity of the pixel. The two-dimensional optical flow vector distribution map is subjected to vector filtering based on the foreground pixel region mask image to retain the optical flow vector information in the foreground region corresponding to the tourist candidate object, thereby obtaining the tourist individual motion optical flow vector map. The motion optical flow vector diagram of the individual tourists is subjected to vector field segmentation processing based on connected component labeling. All optical flow vectors belonging to the same candidate tourist within the connected component are clustered and merged to generate the average motion vector of each individual tourist. The motion trajectory association matching process is performed based on the average motion vector of each tourist individual and the spatial position of the tourist individual in adjacent video frame image units to obtain the displacement vector sequence of the tourist individual between consecutive video frame image units; The displacement vector sequence is subjected to feature optimization processing in the time dimension to generate a set of tourist individual motion vector field features that characterizes the actual movement trend of individual tourists. The set of tourist individual motion vector field features includes the displacement vector direction and displacement vector magnitude of each individual tourist. The set of tourist spatial location distribution features and the set of tourist individual motion vector field features are bound together based on a one-to-one correspondence with tourist individual identifiers to generate a set of tourist individual spatiotemporal state descriptors with dual attributes of spatial location and motion state.

3. The method for real-time crowd density monitoring in amusement parks based on visual recognition according to claim 2, characterized in that, The initial segmentation process of separating the foreground and background of the video frame image units of the amusement park area to obtain a foreground pixel region mask image containing the outline of a suspected visitor includes: The video frame image units of the amusement park area are processed by color space model conversion, converting the original red-green-blue color space pixel values ​​to hue-saturation-brightness color space to obtain hue-saturation-brightness color space image units. Based on the background modeling reference frame image corresponding to the video frame image unit of the amusement park area, calculate the difference in hue component, saturation component, and lightness component between the video frame image unit of the amusement park area and the background modeling reference frame image at the corresponding pixel positions in the hue-saturation-lightness color space. The hue component difference, saturation component difference, and brightness component difference are weighted and fused to generate a pixel difference metric map that characterizes the degree to which a pixel deviates from the background model. The pixel difference measurement map is subjected to preliminary binarization processing based on global adaptive threshold segmentation to obtain an initial foreground pixel label map. In the initial foreground pixel label map, the region with a pixel value of the first value is labeled as a suspected foreground region, and the region with a pixel value of the second value is labeled as a background region. Morphological opening operations are performed on the initial foreground pixel marker image to eliminate isolated noise pixel blocks in the suspected foreground region and connect the broken parts in the suspected foreground region caused by discontinuous segmentation, resulting in a morphologically optimized foreground marker image. The morphologically optimized foreground marker map is subjected to connected component analysis to extract all connected regions composed of first-value pixels, and the area of ​​the bounding rectangle and the perimeter feature parameters of the outline of each connected region are calculated. Based on the preset threshold values ​​for the minimum bounding rectangle area and the perimeter range of the tourist target, the connected regions are initially screened and filtered to remove connected regions whose area is smaller than the minimum bounding rectangle area threshold or whose perimeter exceeds the perimeter range threshold, thus obtaining a set of candidate tourist connected regions. For each connected region in the candidate tourist connected region set, contour internal hole filling processing is performed to ensure that each connected region is composed of continuous pixel regions, generating a complete connected region mask without internal holes; After the hole filling process, the edges of each connected region mask are smoothed. A polygon approximation algorithm is used to simplify the jagged edges of the connected region contour, resulting in a smooth connected region contour line. A foreground pixel region mask image containing the outline of a suspected tourist is generated based on the contour lines of the smooth connected regions. Each connected region in the foreground pixel region mask image represents a two-dimensional pixel projection of a tourist candidate object.

4. The method for real-time crowd density monitoring in amusement parks based on visual recognition according to claim 2, characterized in that, The step of performing contour edge refinement fitting on each connected region in the foreground pixel region mask image and extracting the head vertex pixel coordinates and foot contact point pixel coordinates of the tourist candidate object corresponding to the connected region includes: The target connected region in the foreground pixel region mask image is subjected to minimum bounding rectangle orientation calibration processing. The principal axis direction angle of the target connected region is calculated. The target connected region is rotated according to the principal axis direction angle so that the principal axis of the target connected region is parallel to the vertical axis of the image coordinate system, and the pose-normalized connected region mask is obtained. The pose-normalized connected region mask is subjected to vertical direction pixel projection distribution statistical processing. The number of foreground pixels contained in each row of pixels on the vertical axis of the connected region mask is calculated to obtain the vertical projection histogram. Peak detection processing is performed on the vertical projection histogram to identify the inflection point where the number of foreground pixels in the vertical projection histogram changes. The inflection point includes the head starting point, the shoulder peak point, and the foot starting point. Based on the vertical coordinates corresponding to the head starting point, the geometric center horizontal coordinates of the foreground pixels in the vertical coordinate row are located on the connected region mask. The combination of the geometric center horizontal coordinates and the vertical coordinates of the head starting point is determined as the first coordinates of the head vertex of the tourist candidate in the pose normalized image. Based on the vertical coordinates corresponding to the foot starting point, the geometric center horizontal coordinates of the foreground pixel in the vertical coordinate row are located on the connected region mask. The combination of the geometric center horizontal coordinates and the vertical coordinates of the foot starting point is determined as the second coordinates of the foot contact point of the tourist candidate in the pose normalized image. Based on the rotation angle and rotation center used in the posture normalization process, the first coordinate and the second coordinate are subjected to reverse rotation transformation to map the first coordinate and the second coordinate back to the pixel coordinate system of the original amusement park area video frame image unit. Sub-pixel precision edge search processing is performed on the head vertex coordinates after mapping back to the original pixel coordinate system within a local neighborhood window. Within a preset window around the head vertex coordinates, the image grayscale gradient information is used to adjust the head vertex coordinates to the position of the local maximum curvature point of the head contour. The coordinates of the foot contact point after mapping back to the original pixel coordinate system are optimized based on the geometric constraints of the ground contact point. Combined with the pixel distribution of the bottom edge of the connected region of the tourist candidate object, the coordinates of the foot contact point are adjusted to the midpoint of the bottom edge of the connected region and the coincidence point of the bottom edge and the ground. The optimized head vertex coordinates and optimized foot contact point coordinates in the original pixel coordinate system will be output as the head vertex pixel coordinates and foot contact point pixel coordinates of the tourist candidate object corresponding to the connected region. The extracted pixel coordinates of the head vertex and the pixel coordinates of the foot contact point are associated and stored with the corresponding tourist candidate object identifier to generate a two-dimensional spatial location point pair data structure of the tourist individual in the pixel coordinate system.

5. The method for real-time crowd density monitoring in amusement parks based on visual recognition according to claim 2, characterized in that, The optical flow field calculation process, which involves performing inter-frame pixel grayscale variation calculations on video frame image units of the amusement park area in a continuous time series, yields a two-dimensional optical flow vector distribution map characterizing the instantaneous motion velocity of each pixel. This includes: Select the current video frame image unit and the previous historical video frame image unit adjacent to the current video frame image unit from the video frame image units of the amusement park area in the continuous time series. Multi-scale pyramid hierarchical downsampling processing is performed on the current video frame image unit and the previous historical video frame image unit respectively to construct an image pyramid hierarchical structure from the original resolution to the low resolution, thus obtaining the current frame image pyramid and the historical frame image pyramid. At the highest level of the current frame image pyramid and the historical frame image pyramid, the optical flow vector value of the pixels at that level is initialized to zero. Starting from the highest level of the image pyramid, the optical flow vector is iteratively calculated layer by layer down to the original resolution level. During the calculation process at each level, the optical flow vector result calculated in the previous layer is upsampled and amplified to serve as the initial estimate for the optical flow calculation at the current layer. At the current image resolution, for the current layer image of the current frame image pyramid and the current layer image of the historical frame image pyramid, brightness constant constraint equations and gradient constant constraint equations are constructed based on the gray-level invariance assumption and the local neighborhood constant optical flow assumption. By jointly solving the constant brightness constraint equation and the constant gradient constraint equation, the incremental values ​​of the optical flow vector containing the horizontal optical flow component and the vertical optical flow component are obtained. The incremental value of the optical flow vector is accumulated and fused with the initial estimate passed down from the previous layer to update the refined optical flow vector field of the current layer. Repeat the step of iteratively calculating the optical flow vector layer by layer until the optical flow vector field calculation of the original resolution level is completed, and obtain the original resolution dense optical flow vector field map of the current video frame image unit relative to the previous historical video frame image unit. The original resolution dense optical flow vector field map is subjected to vector smoothing processing based on median filtering to eliminate isolated abnormal optical flow vectors caused by image noise or texture loss in the dense optical flow vector field map, and a smoothed dense optical flow vector field map is obtained. The smoothed dense optical flow vector field map is used as a two-dimensional optical flow vector distribution map to characterize the instantaneous motion velocity of the pixel. Each pixel position in the two-dimensional optical flow vector distribution map contains a two-dimensional optical flow vector composed of a horizontal motion velocity component and a vertical motion velocity component.

6. The method for real-time crowd density monitoring in amusement parks based on visual recognition according to claim 1, characterized in that, The process involves calling pre-configured spatial semantic map data of the amusement park area, performing regional spatial coordinate mapping and alignment processing on the set of tourist spatial location distribution features and the set of tourist individual motion vector field features, and generating a real-time tourist density distribution grid data matrix for each amusement park area within the current monitoring period, based on a geographic coordinate system. This includes: Obtain pre-configured spatial semantic map data of the amusement park area. The spatial semantic map data of the amusement park area includes a set of boundary geographic coordinate polygons of each area of ​​the amusement park, the geographic center point coordinates of each area, the functional attribute labels of each area, and the homography transformation matrix parameters that map pixel coordinate system points to geographic coordinate system points. Extract the pixel coordinates of the foot contact point of each tourist from the set of tourist spatial location distribution features, and use the homography transformation matrix parameters to perform perspective projection transformation on the pixel coordinates of the foot contact point to convert the pixel coordinates of the foot contact point into a geographic spatial location point represented by geographic latitude and longitude coordinates. The spatial inclusion relationship between the geospatial location points of all individual tourists obtained by conversion and the geographical coordinate polygons of each area boundary in the spatial semantic map data of the amusement park area is determined to identify the specific amusement park area where each individual tourist is currently located. Based on the boundary geographic coordinate polygons of each amusement park area, each amusement park area is spatially divided into a regional grid composed of multiple regular geographic grid units. Each geographic grid unit has a fixed geographic latitude and longitude span and a unique grid index code. Based on the specific geographic grid cell into which each tourist's geographic location falls, the number of tourist individuals included in each geographic grid cell during the current monitoring period is counted, which serves as the initial tourist count for that geographic grid cell. The initial visitor count is subjected to spatial diffusion smoothing based on neighboring grid cells. A portion of the visitor count of each geographic grid cell is distributed to the eight surrounding geographic grid cells according to distance weights to obtain the spatially smoothed visitor count distribution. The horizontal and vertical true motion velocity components and corresponding time interval information of each tourist individual are extracted from the tourist individual motion vector field feature set. Combined with the pixel coordinates of the tourist individual's foot contact point, the displacement vector of the tourist individual in the pixel coordinate system is calculated. The displacement vector is mapped to the geographic coordinate system using the homography transformation matrix parameters to obtain the displacement vector in the geographic coordinate system. Based on the displacement vector in the geographic coordinate system and the time interval information, the tourist motion velocity vector based on the geographic coordinate system is calculated. The average motion velocity vector of each tourist individual falling into each geographic grid cell is calculated by vector summation and numerical averaging, and the average motion velocity vector of each geographic grid cell is obtained. The average motion velocity vector includes the average motion direction and the average motion speed. The spatially smoothed visitor count and the average movement speed vector are used as two attribute values ​​for each geographic grid unit. They are then filled into the corresponding positions of a two-dimensional matrix according to the grid index encoding to generate a real-time visitor density distribution grid data matrix for each amusement park area based on the geographic coordinate system during the current monitoring period. In the real-time density distribution grid data matrix of tourists, the row index of the matrix corresponds to the grid number in the geographical latitude direction, the column index of the matrix corresponds to the grid number in the geographical longitude direction, and the value of the matrix element includes the tourist count scalar value and the average movement speed vector value of the grid cell.

7. The method for real-time crowd density monitoring in amusement parks based on visual recognition according to claim 1, characterized in that, The process of analyzing tourist flow trends in the real-time tourist density distribution grid data matrix yields a set of tourist density change rate characteristic fields and tourist flow main direction characteristic vectors for each amusement park area, including: Extract the real-time density distribution grid data matrix of tourists from multiple consecutive historical monitoring periods from the real-time density distribution grid data matrix of tourists to form a three-dimensional density grid data tensor with a time dimension. The three dimensions of the three-dimensional density grid data tensor are the grid row index, the grid column index and the time index, respectively. The three-dimensional density grid data tensor is processed by inter-time period difference calculation in the time dimension. For each fixed grid cell position, the difference between the tourist count in the current monitoring period and the tourist count in the previous historical monitoring period is calculated to obtain the tourist count time change rate of that grid cell. The calculated tourist count change rate of all grid cells is filled into a two-dimensional matrix with the same size as the real-time tourist density distribution grid data matrix according to the corresponding grid index, generating a tourist density change rate feature field. The value of each grid cell in the tourist density change rate feature field represents the increase or decrease in the number of tourists at the corresponding location. An autoregressive moving average model is used to fit the time series of visitor counts at each grid cell position in the three-dimensional density grid data tensor. The trend term and seasonal term of the visitor count time series are extracted, and the slope of the trend term is used to predict the direction of change in visitor counts in the short term. Based on the tourist density change rate characteristic field and the predicted change direction of the tourist count time series, grid cells of clustered hot spots with positive tourist count time change rate and grid cells of evacuation cold spots with negative tourist count time change rate are identified. Extract the average motion velocity vector of all grid cells from the real-time density distribution grid data matrix of tourists, and construct a two-dimensional vector field. Each grid cell position in the two-dimensional vector field contains a motion vector composed of the average motion direction angle and the average motion speed magnitude. The two-dimensional vector field is subjected to flow field feature analysis based on curl and divergence calculation. The curl distribution map of the two-dimensional vector field is calculated to identify the rotation center region of the tourist flow, and the divergence distribution map of the two-dimensional vector field is calculated to identify the convergence region and divergence region of the tourist flow. Based on the curl distribution map and the divergence distribution map, the two-dimensional vector field is divided into multiple sub-regions with similar flow patterns, and the motion vector directions within each sub-region are consistent. For each sub-region, the motion vectors of all grid cells within the sub-region are averaged by vector superposition to calculate a composite vector representing the overall tourist flow trend of the sub-region, which serves as the main direction feature vector of tourist flow in the sub-region. The boundary geographic coordinates of each sub-region are associated with and stored with the corresponding main direction feature vector of tourist flow, generating a set of main direction feature vectors of tourist flow covering the entire amusement park area.

8. The method for real-time crowd density monitoring in amusement parks based on visual recognition according to claim 1, characterized in that, The method, based on the real-time visitor density distribution grid data matrix, the visitor density change rate feature field, and the visitor flow main direction feature vector set, generates a region-by-region guidance indication signal encoding sequence for driving the indicator light arrays distributed in various areas of the amusement park to provide differentiated dynamic indications. This sequence includes: Based on the tourist count scalar value of each grid cell in the real-time tourist density distribution grid data matrix, the grid cell is divided into density levels. The tourist count scalar value is divided into multiple consecutive density level intervals according to the numerical value, and a unique color-coded index is assigned to each density level interval. Based on the tourist count time change rate of each grid cell in the tourist density change rate feature field, the change trend of the grid cell is divided into trend levels. The tourist count time change rate is divided into multiple continuous change trend level intervals according to the numerical value, and a unique corresponding flashing frequency code index is assigned to each change trend level interval. Based on the tourist flow main direction feature vector of each sub-region in the tourist flow main direction feature vector set, the tourist flow main direction feature vector is quantized into multiple preset standard direction codes, and each standard direction code corresponds to a specific orientation. Statistical mode analysis is performed on the color coding index of all grid cells in each amusement park area to determine the regional primary color coding index that represents the overall density of the amusement park area. The flicker frequency coding index of all grid cells in each amusement park area is weighted and averaged, and the scalar value of the visitor count of each grid cell is used as the weight to calculate the regional main flicker frequency coding index that represents the overall density change trend of the amusement park area. The region's main color encoding index and the region's main flashing frequency encoding index are combined and encoded to generate a first-level guiding indicator signal code element. The first-level guiding indicator signal code element is used to control the basic color and basic flashing mode of the corresponding region's indicator light array. Obtain the main direction feature vectors of tourist flow from the set of main direction feature vectors of tourist flow, and analyze whether the main direction feature vectors of tourist flow in the surrounding areas point to the interior of the amusement park area. If they point to the interior, generate a directional warning identifier. According to the directional warning identifier, an enhanced indicator signal code matching the directional warning identifier is retrieved from a preset enhanced indicator mode library. The enhanced indicator signal code is used to drive the corresponding area indicator light array to superimpose enhanced indicator light effects on the basic color and basic flashing mode. The region main color encoding index, the region main flashing frequency encoding index, the standard direction encoding, and the enhanced indicator signal encoding are encapsulated and combined according to a predefined region-by-region flow guidance indicator signal frame structure to generate a binary signal encoding sequence containing multiple fields. The binary signal encoding sequence is associated and bound with the network address identifiers of the indicator light arrays distributed in various areas of the amusement park to generate a region-by-region guiding indicator signal encoding sequence for driving the indicator light arrays to perform differentiated dynamic indications. The region-by-region guiding indicator signal encoding sequence is then sent to the corresponding indicator light array controller through the communication network.

9. A real-time crowd density monitoring system for amusement parks based on visual recognition, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the visual recognition-based real-time crowd density monitoring method for amusement parks as described in any one of claims 1 to 8 by executing the machine-executable instructions.

10. A computer program product, characterized in that, The computer program product includes machine-executable instructions stored in a computer-readable storage medium. The processor of the vision-based amusement park real-time crowd density monitoring system reads the machine-executable instructions from the computer-readable storage medium and executes the machine-executable instructions, causing the vision-based amusement park real-time crowd density monitoring system to perform the vision-based amusement park real-time crowd density monitoring method as described in any one of claims 1 to 8.