Airport flight video slice generation method and device
By using the methods of correlation between flight information and video source address, real-time video data feature extraction and clustering model optimization in the airport flight video surveillance system, video slices are generated, and the problems of low information storage and retrieval efficiency and slow manual monitoring response in traditional systems are solved, and efficient video management and analysis are achieved.
Patent Information
- Application Number
- CN202510480144.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Traditional airport flight video surveillance systems have problems such as information storage, low retrieval efficiency, slow manual monitoring response, and inability to effectively associate real-time data and video.
A method for generating video slicing at airport flights is proposed. By obtaining flight information and associated with the video source address, real-time video data feature extraction is carried out, clustering model is constructed, and dynamic division is optimized and matched with reference slice parameters through pulsed neural network to generate video slices.
It realizes accurate slicing and efficient storage of flight videos, supports multi-grained analysis, improves information retrieval efficiency and manual response speed, and enhances airport flight safety management capabilities.
Smart Images

Figure CN120017925A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video slicing, and in particular to a method and device for generating airport flight video slices. Background Art
[0002] In modern airport management, video surveillance and recording of airport flights have become increasingly important and have become an indispensable tool for ensuring flight safety and improving service quality.
[0003] Traditional flight video surveillance systems usually use fixed cameras for global 7*24 hours monitoring. Although they can provide continuous video streams, they have many shortcomings in information storage, search, processing and utilization.
[0004] Specifically, traditional methods usually use fixed resolution and time windows for video storage, resulting in a waste of storage resources and low retrieval efficiency. For example, for video clips without abnormal events for a long time, using high-resolution storage will take up a lot of storage space; and when a specific event needs to be retrieved quickly, the specific time period when the event occurred may not be accurately located due to the fixed time window.
[0005] Traditional airport management relies on manual monitoring of video data, which is passive and inefficient. Manual monitoring is difficult to capture all abnormal events in real time, and long-term monitoring can easily lead to fatigue and negligence. In addition, manual response is slow and cannot respond to emergencies in a timely manner, affecting the safety and operational efficiency of the airport.
[0006] Traditional systems cannot effectively associate real-time flight data with video, resulting in information islands. This makes it difficult for airport operators to quickly obtain relevant video information for specific flights when dealing with flight anomalies, settlement and billing, passenger complaints, etc. Summary of the invention
[0007] Based on this, the purpose of the present invention is to propose a method and device for generating airport flight video slices to solve the above-mentioned problems.
[0008] A method for generating airport flight video slices according to the present invention comprises: Get current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate, and baggage carousel number; Obtaining the video source address of the key area, wherein the key area includes the flight area, boarding gate, baggage carousel and waiting area; Associate flight information with the video source address and call real-time video data; Extract features from the real-time video data of the flight to obtain the target feature vector; Construct a clustering model, add the current target feature vector to the clustering model, and construct an event cognitive hypersurface through a pulse neural network to optimize the clustering model's dynamic partitioning of real-time video data and match the benchmark slice parameters; The real-time video data is sliced according to the matched reference slice parameters to generate video slices.
[0009] Furthermore, the event recognition hypersurface is constructed by using a pulse neural network to optimize the dynamic partitioning of the clustering model for real-time video data and the matching of the benchmark slice parameters, including: Calculate the pulse distance from the current target feature vector to each known event cluster hypersurface; If the pulse distance is less than the distance threshold, the current target feature vector is assigned to the nearest cluster, and the pulse plasticity update of the cluster center is triggered; If the pulse distance is greater than the distance threshold, a new event cluster is created and the reference slice parameters are initialized; Event boundaries are identified based on the newly added pulse patterns to adjust the shape of the event cognition hypersurface.
[0010] Furthermore, the calculation of the pulse distance from the current target feature vector to each known event cluster hypersurface includes: Calculate the current target feature vector x through the SNN impulse response kernel function t The pulse distance d to the hypersurface of the ith event cluster i , the formula is: , Among them, d i is the pulse distance of the ith event cluster, K is the number of neurons, and w ij is the synaptic weight of the jth neuron in the ith event cluster, δ is the impulse response function, σ j is the receptive field width of the jth neuron, is the end time t of the current time window and the pulse emission time t of neuron j j The absolute difference of .
[0011] Furthermore, the triggering cluster center pulse plasticity updating includes: For the matching event clusters, the synaptic weights are adjusted according to the pulse timing-dependent plasticity rule, and the adjustment formula is: , in, is the synaptic weight w ij The adjustment amount, η is the learning rate, τ is the time constant, Δt is the time difference between the postsynaptic neuron and the presynaptic neuron pulse, sign(Δt) is the sign function, when Δt>0, it is 1, the weight is strengthened, when Δt<0, it is -1, the weight is weakened.
[0012] Furthermore, the identifying event boundaries according to the newly added pulse pattern to adjust the shape of the event cognitive hypersurface includes: The objective function is to minimize the intra-cluster pulse distance variance and model complexity. The objective function is: , Among them, L is the objective function used to measure the clustering quality, λ is the regularization coefficient, is the Frobenius norm of the synaptic weight matrix.
[0013] By pulse gradient descent, the synaptic weights are iteratively updated, and the update formula is: , in, is the updated synaptic weight matrix of the ith event cluster, is the synaptic weight matrix of the ith event cluster before updating, The objective function for the synaptic weight matrix W i The gradient of , γ is the step size, and β is the momentum coefficient is the momentum term of the previous weight update.
[0014] Furthermore, the construction of the clustering model includes: Collect historical video streams of airport scenes and associated flight status data to generate historical target feature vectors; The historical target feature vector is converted into a pulse sequence, input into the pulse neural network, and the cluster center is initialized by the self-organizing map algorithm; The cluster centers are optimized based on the pulse gradient descent algorithm, and the hypersurface parameters of the event cluster are defined to minimize the pulse distance variance of samples within the cluster. According to the optimized hypersurface parameters, the reference slice parameters of each event cluster are calculated and stored, including the time window length, resolution threshold and coordinates of the region of interest.
[0015] Furthermore, slicing the real-time video data according to the reference slicing parameters of the matching event cluster includes: Read the predefined slice time window length T, resolution threshold R and region of interest ROI from the benchmark slice parameter library of the matching event cluster; Generate slicing instructions according to the slice time window length T, resolution threshold R and region of interest ROI; Slice the video stream according to the slicing instruction to generate segments matching the event type; The stability of the slice content is monitored through the pulse frequency of the SNN. If an abnormality is detected, the parameter re-evaluation mechanism is triggered.
[0016] Furthermore, the flight information is associated with the video source address and the real-time video data is called, including: Create a data structure to establish and store the mapping relationship between flight information and the corresponding video data source address; A timer is set to call the flight information API at a preset time to obtain the latest flight data, and the flight data returned by the flight information API is parsed to update the flight information status and the mapping relationship in the data structure in real time; According to the real-time flight information, the corresponding video data source is dynamically matched and switched from the updated mapping relationship of the data structure to obtain the real-time video stream data.
[0017] Furthermore, dynamically matching and switching to a corresponding video data source from the updated mapping relationship of the data structure according to the real-time flight information to obtain real-time video stream data includes: When the flight information is updated, the corresponding new video data source in the mapping relationship of the data structure is immediately searched, and the new video data source is switched to obtain real-time video stream data from the new video data source.
[0018] The present invention also proposes a device for generating airport flight video slices, which is used to implement the above-mentioned airport flight video slice generation method, and the device comprises: The first data module: used to obtain the current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate and baggage carousel number; The second data module is used to obtain the video source address of the key area, which includes the flight area, boarding gate, baggage carousel and waiting area; Association module: used to associate flight information with the video source address and call real-time video data; Feature extraction module: used to extract features from real-time video data of flights to obtain target feature vectors; Classification module: used to build a clustering model, add the current target feature vector to the clustering model, and build an event cognitive hypersurface through a pulse neural network to optimize the clustering model's dynamic division of real-time video data and match the benchmark slice parameters; Slicing module: used to slice real-time video data according to matching benchmark slicing parameters to generate video slices.
[0019] In summary, the airport flight video slice generation method of the present invention obtains the current flight information and accurately associates it with the video source address of the key area to ensure that the corresponding video data can be quickly called in real time according to the flight information, and the called real-time video data is feature extracted to obtain the target feature vector. Subsequently, the event cognitive hypersurface is constructed through the pulse neural network, and the clustering model is dynamically optimized to divide the real-time video data, so as to achieve accurate identification of event clusters and automatic matching of benchmark slice parameters, so as to adapt to the rapid changes of airport scenes, such as sudden passenger flow, flight delays, etc.; finally, the video is accurately cut according to the slice parameters matched by the event cluster, and structured video slices that are highly matched with the event type are generated, such as flight take-off and landing clips, baggage carousel abnormal clips, to support multi-granularity analysis and efficient storage.
[0020] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description or will be learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which: Figure 1 This is a flow chart of a method for generating airport flight video slices according to the first embodiment of the present invention; Figure 2 This is a system block diagram of a device for generating airport flight video slices according to a second embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0023] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0025] Example 1: Please refer to Figure 1 The present invention proposes a method for generating airport flight video slices, which includes steps S101 to S106: S101, obtaining current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate, and baggage carousel number.
[0026] Collect basic information related to the current flight, such as flight number (used to uniquely identify the flight), estimated departure / arrival time (used to determine the time range of flight activities), boarding gate (the specific location where passengers board the plane) and baggage carousel number (the location where passengers collect their luggage). This information is the basis for associating and retrieving video data in subsequent steps.
[0027] Specifically, you can connect to the information system of airlines or airports to obtain real-time flight information API. This API (the interface used to obtain flight-related information) can provide key information such as flight number, boarding gate, baggage carousel, estimated departure / arrival time, etc. Obtaining flight information through the API can ensure the real-time, comprehensive and accurate data, and avoid problems caused by manual input or outdated data.
[0028] S102, obtaining the video source address of the key area, wherein the key area includes the flight area, the boarding gate, the baggage carousel and the waiting area.
[0029] By obtaining the video source addresses of key areas, the dynamic conditions of key areas in the airport can be captured, including the flight area (where planes take off and land), boarding gates (where passengers get on and off the plane), baggage carousels (where passengers collect their luggage), and waiting areas (where passengers wait to board the plane).
[0030] Since there are many cameras installed in the airport, especially in key areas, each camera has an address (such as a website), through which the video stream captured by the camera can be viewed and called at any time. Through the video stream address (such as RTSP URL) corresponding to each camera in each key area, the system can obtain video data of these areas in real time.
[0031] S103, associate the flight information with the video source address, and call the real-time video data.
[0032] The specific information of the flight (such as flight number, boarding gate, baggage carousel number, etc.) is linked to the video source address of the corresponding area in the airport to ensure that the relevant video data can be quickly located based on the flight information. Specifically, this can be done by establishing a data structure (such as a mapping table or database) that stores the correspondence between flight information and video source addresses. When the flight information changes (such as a change in the boarding gate), the data structure is updated in a timely manner. In this way, the flight information is accurately associated with the video source address to ensure that the correct video data can be called later.
[0033] According to the latest flight information, the video data of the corresponding area is obtained in real time to provide real-time and accurate video data for video slice generation. Specifically, the video source address corresponding to the current flight information can be found by querying the data structure (such as a mapping table or database). Video data is obtained from the video source address in real time using video stream acquisition technology (such as RTSP, HLS and other protocols). Since the acquisition of video data is real-time and closely related to the flight status information, it can accurately reflect the current operation status of the flight in real time.
[0034] Further optionally, in step S103, the flight information is associated with the video source address and the real-time video data is called, including: Create a data structure to establish and store the mapping relationship between flight information and the corresponding video data source address (i.e. the corresponding camera address); A timer is set to call the flight information API at a preset time to obtain the latest flight data, and the flight data returned by the flight information API is parsed to update the flight information status and the mapping relationship in the data structure in real time; According to the real-time flight information, the corresponding video data source is dynamically matched and switched from the updated mapping relationship of the data structure to obtain the real-time video stream data.
[0035] Further optionally, dynamically matching and switching to a corresponding video data source from the updated mapping relationship of the data structure according to the real-time flight information to obtain real-time video stream data includes: When the flight information is updated, the corresponding new video data source (i.e., the new camera address) in the mapping relationship of the data structure is immediately searched, and the new video data source is switched to obtain real-time video stream data from the new video data source to ensure that the video image is consistent with the updated flight information.
[0036] Understandably, in order to effectively associate current flight information with real-time video data, it is necessary to obtain accurate current flight information, such as flight number, estimated departure time, estimated arrival time, boarding gate, baggage carousel number and other key data. By connecting with the information system of the airline or airport, you can obtain real-time flight information API, which can provide all the required flight information, ensuring the real-time and accuracy of the data.
[0037] The video stream address of the airport monitoring system can be used to obtain real-time video data of key areas (such as the flight area, boarding gate, baggage carousel and waiting area, which are important monitoring points in the flight operation process). Since the airport is full of cameras, especially in key areas, each camera has a unique address, through which the camera can be called and viewed at any time.
[0038] In order to closely associate flight information with real-time video data and call real-time video data based on flight information, a data structure (such as a mapping table or database) is created to establish and store the mapping relationship between flight information and the corresponding video data source address (i.e., camera address). The specific mapping example is as follows: Flight number: FL123 Area 1 (flying area): Camera address A1 Area 2 (boarding gate): Camera address A2 Area 3 (baggage carousel): Camera address A3 Area 4 (Waiting Area): Camera address A4 When flight FL123 is in the airfield, the system obtains video data from camera address A1; when the flight arrives at the gate, the system switches to camera address A2; and so on.
[0039] However, when the flight information changes, the mapping relationship will also be updated. Flight information, such as boarding gates, baggage carousels, estimated departure / arrival times, etc., may change due to various reasons (such as weather, mechanical failure, scheduling adjustments, etc.). In this case, the mapping relationship between flight information and video data sources needs to be updated in real time.
[0040] You can obtain real-time notifications of flight information changes through the information system of the airline or airport, and update the mapping relationship in the previously created data structure based on the change notification. Then, based on the updated mapping relationship, dynamically switch to the new video data source to obtain real-time video data corresponding to the changed flight information. However, although the status of each flight will change (such as changing boarding gate B1 to B2), the mapping relationship between each area and its corresponding video source address remains unchanged (such as the camera address corresponding to gate B2 has always been A2). Specific examples are as follows: The initial mapping relationship is: Flight number: FL123 Gate: B1 Baggage carousel: C1 Camera address: A1 (corresponding to gate B1) Flight information changed: boarding gate changed from B1 to B2 Updated mapping relationship: Flight number: FL123 Gate: B2 Baggage carousel: C1 Camera address: A2 (corresponding to Gate B2) Since flight information changes dynamically, in order to ensure that the obtained flight information is always the latest, you can set a timer to call the flight information API every preset time (such as every minute) to obtain the latest flight data. Then parse the returned data to update the status and mapping relationship of the flight information in real time.
[0041] Finally, according to the real-time updated flight information, the corresponding video data source is dynamically matched and switched from the updated data structure to obtain the real-time video stream data closely related to the current flight. In this way, the efficient association and processing of flight information and video data can be achieved, and the various situations in the flight operation process can be grasped in real time through the video monitoring system, providing strong support for flight safety and management.
[0042] S104, extracting features from the real-time video data of the flight to obtain a target feature vector.
[0043] Based on the synchronously collected airport scene monitoring video stream and flight status data, the spatiotemporal characteristics and structural information of dynamic targets are extracted to generate high-dimensional target feature vectors. The specific implementation process is as follows: First, feature extraction of the video stream is performed. The YOLOv8 target detection algorithm can be used to process the video stream in real time to identify dynamic targets such as pedestrians, vehicles, and luggage carts, and extract their spatial coordinates (x position, y position), motion parameters (speed, acceleration) and density (density, such as the number of targets per unit area). For example, in the boarding gate area, the pedestrian movement trajectory can be detected as [(100, 200), (105, 202), (110, 205)], and the speed is calculated to be 2.5 pixels / frame and the density is 5 people / square meter. In order to eliminate video noise (such as jitter and occlusion), the trajectory data is smoothed by the Kalman filter algorithm, and aligned with the flight status data based on the timestamp to solve the data asynchrony problem.
[0044] Secondly, the flight status data is parsed to obtain structured information such as flight number, scheduled take-off and landing time, actual take-off and landing time, boarding gate, baggage carousel number, and delay status (delay flag, such as 0 for normal and 1 for delayed) from the airport operation system. For example, flight CA1234 is scheduled to take off at 10:00, but actually takes off at 10:30, the delay flag is 1, and the boarding gate is B12. These data provide key contextual associations for video stream features, such as accurately identifying the relationship between flight delays and changes in crowd density in the waiting area.
[0045] Next, feature fusion is performed to align the video stream features with the flight status data by timestamp, and a target feature vector containing spatiotemporal coordinates, motion parameters, crowd density, and flight status is generated. Its structure can be: [x position, y position, speed, density, delay sign, timestamp, flight number, boarding gate, baggage carousel number].
[0046] For example, the feature vector [120, 300, 1.8, 4.5, 1, "2024-05-01 10:30:00", "CA1234","B12", "D05"] means: at 2024-05-01 10:30:00, near gate B12 of flight CA1234 (delayed), a pedestrian was detected at coordinates (120, 300), with a speed of 1.8 m / s, a density of 4.5 people / m2, and the associated baggage carousel is D05.
[0047] Finally, to improve feature stability, the feature vectors of consecutive frames are aggregated in sliding windows (such as 5-second windows) to calculate the average or maximum value. At the same time, high-dimensional feature vectors are reduced in dimension through principal component analysis to retain the main information and reduce computational complexity. For example, the original 128-dimensional video features can be reduced to 32 dimensions while retaining 95% of the variance information, significantly improving the efficiency of subsequent clustering models.
[0048] S105, constructing a clustering model, adding the current target feature vector to the clustering model, and constructing an event cognitive hypersurface through a pulse neural network to optimize the clustering model for dynamic partitioning of real-time video data and matching of benchmark slice parameters.
[0049] The construction of a dynamic clustering model based on a pulse neural network breaks through the static limitations of traditional methods. Since airport scenes are highly dynamic, such as sudden passenger flow and flight changes, traditional clustering models (such as K-means) cannot adjust the boundaries of event clusters in real time. The present invention introduces a pulse neural network, which simulates neuron pulse signals to model the spatiotemporal sequence of real-time video feature vectors (such as passenger density and luggage movement speed), and can dynamically capture the laws of event evolution. For example, when the waiting area changes from "normal" to "congested", the pulse neural network can adaptively adjust the event cluster division and match the benchmark slice parameters, thereby greatly improving the accuracy of event recognition and reducing the matching error of slice parameters, which can effectively cope with the complex scene changes at the airport.
[0050] Further optionally, the constructing of the clustering model includes: Collect historical video streams of airport scenes and associated flight status data to generate historical target feature vectors; The historical target feature vector is converted into a pulse sequence, input into the pulse neural network, and the cluster center is initialized by the self-organizing map algorithm; The cluster centers are optimized based on the pulse gradient descent algorithm, and the hypersurface parameters of the event cluster are defined to minimize the pulse distance variance of samples within the cluster. According to the optimized hypersurface parameters, the reference slice parameters of each event cluster are calculated and stored, including the time window length, resolution threshold and coordinates of the region of interest.
[0051] Specifically, for the synchronously collected surveillance video streams and flight status data of the airport scene, the temporal and spatial characteristics and structural information of dynamic targets are extracted. The video stream is processed by the target detection algorithm to obtain the trajectory, speed and density of pedestrians and vehicles; the flight status data provides key information such as flight number, take-off and landing time and delay status. After the two types of data are aligned by timestamp, the target feature vector is generated.
[0052] The historical target feature vector needs to be converted into the spatiotemporal pulse pattern of the spiking neural network. Pulse coding can be performed by combining frequency coding (the larger the eigenvalue, the higher the pulse frequency) and time coding (key events trigger early pulses) to simulate the information transmission mechanism of biological neurons. In the initialization stage, the self-organizing map algorithm is used to map the high-dimensional pulse sequence to the low-dimensional grid nodes. Through competitive learning, the grid nodes adaptively adjust the weights to approximate the data distribution and form the initial clustering center. There is no need to preset the number of clusters, and the intrinsic topological structure of the data can be automatically captured, providing a robust starting point for subsequent optimization.
[0053] In order to improve clustering accuracy, the pulse gradient descent algorithm is introduced to optimize the hypersurface parameters. The optimization goal is to minimize the pulse distance variance of samples within the cluster by defining the distance metric between pulse sequences (such as dynamic time warping) and adjusting parameters such as the hyperellipsoid equation coefficient. Pulse gradient descent iteratively updates the cluster center by back-propagating the gradient of the pulse activity, so that the hypersurface fits the data distribution better.
[0054] According to the optimized hypersurface parameters, the baseline slice parameters of each event cluster are calculated and stored, including the time window length, resolution threshold and region of interest coordinates. Among them, the time window length is inversely proportional to the average pulse frequency in the cluster. High-frequency events (such as dense flow of people) can use short windows to capture transients, and low-frequency events (such as flight takeoffs and landings) can use long windows to analyze trends. The resolution threshold is dynamically adjusted according to the variance of the feature vector. When the variance is large, the resolution is increased to capture details. The region of interest coordinates are generated based on the spatial distribution statistics of the samples in the cluster, such as the centroid or bounding box, which can locate the key monitoring areas. These parameters convert the clustering results into an operational slice configuration, ensuring that the slice content is highly matched with the event type, and improving the efficiency and accuracy of subsequent video analysis.
[0055] Event clusters can include typical or special event clusters such as flight takeoff and landing event clusters, boarding gate activity event clusters, baggage carousel event clusters, waiting area passenger flow event clusters, security check channel event clusters, abnormal behavior event clusters, equipment failure event clusters, and environmental interference event clusters.
[0056] The clustering model of the present invention realizes efficient recognition and slicing of airport scene events through the dynamic clustering capability and adaptive parameter generation mechanism of the pulse neural network.
[0057] Further optionally, constructing an event cognitive hypersurface through a spiking neural network and optimizing the clustering model for dynamic partitioning of real-time video data and matching of reference slice parameters include: Calculate the pulse distance from the current target feature vector to each known event cluster hypersurface; If the pulse distance is less than the distance threshold, the current target feature vector is assigned to the nearest cluster, and the pulse plasticity update of the cluster center is triggered; If the pulse distance is greater than the distance threshold, a new event cluster is created and the reference slice parameters are initialized; Event boundaries are identified based on the newly added pulse patterns to adjust the shape of the event cognition hypersurface.
[0058] Specifically, an event cognitive hypersurface is constructed through a spiking neural network to achieve dynamic adaptability and real-time event segmentation in airport scenarios. Specifically, the clustering model dynamically allocates data points by calculating the pulse distance from the current target feature vector to each known event cluster hypersurface, combined with a distance threshold mechanism. If the pulse distance is less than the threshold, the target feature vector is assigned to the nearest cluster, and the pulse plasticity update of the cluster center is triggered to adjust the synaptic weights to adapt to new data, such as changes in passenger flow caused by flight delays; if the pulse distance is greater than the threshold, a new event cluster is created, such as adding temporary flights or abnormal events, and the baseline slice parameters (such as time window length, resolution threshold, etc.) are initialized. This process simulates the mechanism of the human brain to recognize familiar and unfamiliar things through neuronal discharge patterns, enabling the clustering model to capture changes in event patterns in the airport in real time (such as fluctuations in flight take-off and landing times, abnormal baggage carousel flow), reduce reliance on manual annotation, and significantly improve the adaptability of dynamic scenarios, such as dynamically adjusting the number of open security check channels to cope with passenger flow peaks.
[0059] The clustering model further uses new pulse patterns (such as sudden changes in pulse frequency and changes in spatial distribution) to detect event boundaries and dynamically adjust synaptic connection strength and hypersurface parameters (such as curvature and direction). For example, when there is crowd gathering or abnormal evacuation in the boarding gate area, the sudden change in pulse frequency can trigger boundary detection, and the model fits the actual event characteristics more accurately by adjusting the hypersurface curvature. This process is similar to the airport navigation software dynamically adjusting the boarding gate path planning according to real-time flight conditions, so that the hypersurface can distinguish similar events, such as boarding gate activities of different flights or baggage carousel allocation conflicts, to reduce misjudgment. Through the boundary evolution mechanism, the clustering model can effectively identify subtle differences in airport operations, such as differences in passenger flow patterns between international flights and domestic flights, which can greatly improve the accuracy and reliability of event recognition.
[0060] Through dynamic clustering, event boundary detection and bio-inspired learning mechanisms, the present invention can capture complex event patterns such as flight dynamics (such as take-off and landing time fluctuations, taxi path changes), passenger flow distribution (such as boarding gate aggregation, baggage carousel flow abnormalities) in airport scenes in real time. And through the dynamic adjustment of the hypersurface shape, the model can accurately distinguish similar events (such as differences in passenger activities at different boarding gates, baggage sorting conflicts at adjacent baggage carousels), significantly reducing the event misjudgment rate. Finally, by optimizing the matching degree of benchmark slicing parameters (such as time window length, resolution threshold, etc.) and event clusters in the video stream, high robustness support is provided for efficient slicing and intelligent analysis of airport flight videos, helping airport operations to achieve refined resource scheduling and rapid response to abnormal events.
[0061] Further optionally, the calculating the pulse distance from the current target feature vector to each known event cluster hypersurface includes: Calculate the current target feature vector x through the SNN impulse response kernel function t The pulse distance d to the hypersurface of the ith event cluster i , the formula is: , Among them, d i is the pulse distance of the ith event cluster, K is the number of neurons, and w ij is the synaptic weight of the jth neuron in the ith event cluster, δ is the impulse response function, σ j is the receptive field width of the jth neuron, which is used to control the spatial decay speed of the impulse response. is the end time t of the current time window and the pulse emission time t of neuron j j The absolute difference of .
[0062] The present invention uses the impulse response kernel function of SNN to transform the spatiotemporal features (time difference , receptive field width σ j ) and the synaptic weight w ij Combined with the above, the dynamic similarity measurement of the event cluster hypersurface is realized, thus realizing dynamic clustering and event partitioning. The pulse distance reflects the similarity between the feature vector and the event cluster hypersurface, which is the key basis for judging the attribution of data points (assigning to existing clusters or creating new clusters).
[0063] Further optionally, before dynamically calculating the pulse distance from the current target feature vector to each known event cluster hypersurface, the method includes: Dynamically calculate the distance threshold θ i =μ i +ασ i , where μ i , σ i are the mean and standard deviation of the distance within the i-th event cluster, and α is a control parameter (the default value is 2).
[0064] By dynamically calculating the distance threshold and pulse distance, adaptive clustering and precise division of complex event patterns such as flight dynamics and passenger flow changes in airport scenes can be achieved. Specifically, by statistically analyzing the mean and standard deviation of the pulse distance in each event cluster, a personalized threshold is dynamically generated, and the similarity between the target feature vector and the event cluster is judged in real time in combination with the spatiotemporal response characteristics of the pulse neural network, so as to accurately distinguish similar events (such as activities at different boarding gates or baggage carousels), reduce the misjudgment rate, and provide highly matched benchmark slice parameters for airport flight video slices, ultimately improving resource scheduling efficiency and abnormal event response capabilities.
[0065] Further optionally, the updating of the impulse plasticity of the trigger cluster center includes: For the matching event clusters, the synaptic weights are adjusted according to the pulse timing-dependent plasticity rule, and the adjustment formula is: , in, is the synaptic weight w ij The adjustment amount, η is the learning rate, which is used to control the step size of weight update, τ is the time constant, which determines the influence range of pulse time difference on weight adjustment, Δt is the time difference between the postsynaptic neuron and the presynaptic neuron pulse emission, sign(Δt) is the sign function, when Δt>0, it is 1, the weight is enhanced, when Δt<0, it is -1, the weight is weakened.
[0066] By triggering the pulse plasticity update of the cluster center, the synaptic weights are dynamically adjusted based on the pulse timing-dependent plasticity rule to optimize the event clustering ability of the airport flight video slice generation method. Specifically, for the matched event clusters, the pulse timing-dependent plasticity rule is used to quantify the time difference between the pre-synaptic and post-synaptic neuron pulses. Through the sign function and directional weight adjustment mechanism (Δt>0 to enhance the weight, Δt<0 to weaken the weight), the clustering model can adaptively learn the intrinsic correlation of spatiotemporal features such as flight dynamics and passenger flow changes, thereby enhancing the representation accuracy and stability of event clusters.
[0067] Further optionally, identifying the event boundary according to the newly added pulse pattern to adjust the shape of the event cognitive hypersurface includes: The objective function is to minimize the intra-cluster pulse distance variance and model complexity. The objective function is: , Among them, L is the objective function used to measure the clustering quality, λ is the regularization coefficient, which is used to balance the intra-cluster variance and model complexity. is the Frobenius norm of the synaptic weight matrix, used for regularization.
[0068] By pulse gradient descent, the synaptic weights are iteratively updated, and the update formula is: , in, is the updated synaptic weight matrix of the ith event cluster, is the synaptic weight matrix of the ith event cluster before updating, The objective function for the synaptic weight matrix W i The gradient of , γ is the step size, which is used to control the update amplitude of each iteration, and β is the momentum coefficient, which is used to accelerate convergence and reduce oscillation. is the momentum term of the previous weight update.
[0069] Through the above steps, the hypersurface parameters can be adaptively adjusted in complex dynamic scenarios (such as flight takeoffs and landings, passenger flow fluctuations), thereby improving the robustness and efficiency of event boundary recognition.
[0070] S106, slicing the real-time video data according to the matched reference slicing parameters to generate video slices.
[0071] Traditional video storage usually uses fixed resolution and time window, resulting in storage waste and inefficient retrieval. The slice parameters of the present invention are matched according to the event cluster, that is, dynamically adjusted according to the event type. For example, high resolution (such as 1080p) is used in key areas (such as boarding gates), and low resolution (such as 720p) is used in non-key areas (such as channel backgrounds); the time window is adaptively adjusted according to the duration of the event (such as flight takeoff and landing slices covering the entire taxiing process). This greatly reduces the cost of video storage, and the retrieval time of abnormal events can be shortened from hours to minutes, supporting the collaborative work of global monitoring and local detection.
[0072] Further optionally, slicing the real-time video data according to the reference slicing parameters of the matching event cluster includes: Read the predefined slice time window length T, resolution threshold R and region of interest ROI from the benchmark slice parameter library of the matching event cluster; Generate slicing instructions according to the slice time window length T, resolution threshold R and region of interest ROI; Slice the video stream according to the slicing instruction to generate segments matching the event type; The stability of the slice content is monitored through the pulse frequency of the SNN. If an abnormality is detected, the parameter re-evaluation mechanism is triggered.
[0073] Specifically, we first extract predefined slice parameters from the benchmark slice parameter library of the matching event cluster, including time window length T, resolution threshold R, and region of interest ROI. These parameters provide spatiotemporal constraints for video slices, where T controls the temporal granularity, R determines the spatial clarity, and ROI locates the key area to ensure that the slice content is highly relevant to the event type (such as boarding, baggage claim).
[0074] Based on the extracted parameters, slice instructions are generated to drive video stream segmentation. This process converts the continuous video stream into segments that match the event semantics, such as cutting the video of the boarding gate area according to the flight take-off and landing time window T, or extracting the video of a specific period based on the ROI of the baggage claim carousel.
[0075] The pulse frequency of the spiking neural network is used to monitor the stability of the slice content in real time. If an abnormality is detected (such as image jitter or blur), the parameter re-evaluation mechanism is triggered to dynamically adjust T, R or ROI to maintain the slice quality. For example, when the baggage carousel video is blurred due to camera jitter, T can be shortened or R can be increased to enhance the key frame clarity.
[0076] In summary, the airport flight video slice generation method of the present invention obtains the current flight information and accurately associates it with the video source address of the key area to ensure that the corresponding video data can be quickly called in real time according to the flight information, and the called real-time video data is feature extracted to obtain the target feature vector. Subsequently, the event cognitive hypersurface is constructed through the pulse neural network, and the clustering model is dynamically optimized to divide the real-time video data, so as to achieve accurate identification of event clusters and automatic matching of benchmark slice parameters, so as to adapt to the rapid changes of airport scenes, such as sudden passenger flow, flight delays, etc.; finally, the video is accurately cut according to the slice parameters matched by the event cluster, and structured video slices that are highly matched with the event type are generated, such as flight take-off and landing clips, baggage carousel abnormal clips, to support multi-granularity analysis and efficient storage.
[0077] Example 2: Please refer to Figure 2 The present invention proposes a device for generating airport flight video slices, the device comprising: The first data module: used to obtain the current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate and baggage carousel number; The second data module is used to obtain the video source address of the key area, which includes the flight area, boarding gate, baggage carousel and waiting area; Association module: used to associate flight information with the video source address and call real-time video data; Feature extraction module: used to extract features from real-time video data of flights to obtain target feature vectors; Classification module: used to build a clustering model, add the current target feature vector to the clustering model, and build an event cognitive hypersurface through a pulse neural network to optimize the clustering model's dynamic division of real-time video data and match the benchmark slice parameters; Slicing module: used to slice real-time video data according to matching benchmark slicing parameters to generate video slices.
[0078] Further optionally, the classification module is also used for: Calculate the pulse distance from the current target feature vector to each known event cluster hypersurface; If the pulse distance is less than the distance threshold, the current target feature vector is assigned to the nearest cluster, and the pulse plasticity update of the cluster center is triggered; If the pulse distance is greater than the distance threshold, a new event cluster is created and the reference slice parameters are initialized; Event boundaries are identified based on the newly added pulse patterns to adjust the shape of the event cognition hypersurface.
[0079] Further optionally, the classification module is also used for: Calculate the current target feature vector x through the SNN impulse response kernel function t The pulse distance d to the hypersurface of the ith event cluster i , the formula is: , Among them, d i is the pulse distance of the ith event cluster, K is the number of neurons, and w ij is the synaptic weight of the jth neuron in the ith event cluster, δ is the impulse response function, σ j is the receptive field width of the jth neuron, is the end time t of the current time window and the pulse emission time t of neuron j j The absolute difference of .
[0080] Further optionally, the classification module is also used for: For the matching event clusters, the synaptic weights are adjusted according to the pulse timing-dependent plasticity rule, and the adjustment formula is: , in, is the synaptic weight w ij The adjustment amount, η is the learning rate, τ is the time constant, Δt is the time difference between the postsynaptic neuron and the presynaptic neuron pulse, sign(Δt) is the sign function, when Δt>0, it is 1, the weight is strengthened, when Δt<0, it is -1, the weight is weakened.
[0081] Further optionally, the classification module is also used for: The objective function is to minimize the intra-cluster pulse distance variance and model complexity. The objective function is: , Among them, L is the objective function used to measure the clustering quality, λ is the regularization coefficient, is the Frobenius norm of the synaptic weight matrix.
[0082] By pulse gradient descent, the synaptic weights are iteratively updated, and the update formula is: , in, is the updated synaptic weight matrix of the ith event cluster, is the synaptic weight matrix of the ith event cluster before updating, The objective function for the synaptic weight matrix W i The gradient of , γ is the step size, and β is the momentum coefficient is the momentum term of the previous weight update.
[0083] Further optionally, the classification module is also used for: Collect historical video streams of airport scenes and associated flight status data to generate historical target feature vectors; The historical target feature vector is converted into a pulse sequence, input into the pulse neural network, and the cluster center is initialized by the self-organizing map algorithm; The cluster centers are optimized based on the pulse gradient descent algorithm, and the hypersurface parameters of the event cluster are defined to minimize the pulse distance variance of samples within the cluster. According to the optimized hypersurface parameters, the reference slice parameters of each event cluster are calculated and stored, including the time window length, resolution threshold and coordinates of the region of interest.
[0084] Further optionally, the slicing module is also used for: Read the predefined slice time window length T, resolution threshold R and region of interest ROI from the benchmark slice parameter library of the matching event cluster; Generate slicing instructions according to the slice time window length T, resolution threshold R and region of interest ROI; Slice the video stream according to the slicing instruction to generate segments matching the event type; The stability of the slice content is monitored through the pulse frequency of the SNN. If an abnormality is detected, the parameter re-evaluation mechanism is triggered.
[0085] Further optionally, the association module is also used for: Create a data structure to establish and store the mapping relationship between flight information and the corresponding video data source address; A timer is set to call the flight information API at a preset time to obtain the latest flight data, and the flight data returned by the flight information API is parsed to update the flight information status and the mapping relationship in the data structure in real time; According to the real-time flight information, the corresponding video data source is dynamically matched and switched from the updated mapping relationship of the data structure to obtain the real-time video stream data.
[0086] Further optionally, the association module is also used for: When the flight information is updated, the corresponding new video data source in the mapping relationship of the data structure is immediately searched, and the new video data source is switched to obtain real-time video stream data from the new video data source.
[0087] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A method for generating airport flight video slices, characterized in that: The method comprises: Get current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate, and baggage carousel number; Obtaining the video source address of the key area, wherein the key area includes the flight area, boarding gate, baggage carousel and waiting area; Associate flight information with the video source address and call real-time video data; Extract features from the real-time video data of the flight to obtain the target feature vector; Construct a clustering model, add the current target feature vector to the clustering model, and construct an event cognitive hypersurface through a pulse neural network to optimize the clustering model's dynamic partitioning of real-time video data and match the benchmark slice parameters; The real-time video data is sliced according to the matched reference slice parameters to generate video slices.
2. The method for generating airport flight video slices according to claim 1, characterized in that: The method of constructing an event cognitive hypersurface through a pulse neural network and optimizing the dynamic partitioning of real-time video data by a clustering model and matching the parameters of the benchmark slices includes: Calculate the pulse distance from the current target feature vector to each known event cluster hypersurface; If the pulse distance is less than the distance threshold, the current target feature vector is assigned to the nearest cluster, and the pulse plasticity update of the cluster center is triggered; If the pulse distance is greater than the distance threshold, a new event cluster is created and the reference slice parameters are initialized; Event boundaries are identified based on the newly added pulse patterns to adjust the shape of the event cognition hypersurface.
3. The method for generating airport flight video slices according to claim 2, characterized in that: The step of calculating the pulse distance from the current target feature vector to each known event cluster hypersurface includes: Calculate the current target feature vector x through the SNN impulse response kernel function t The pulse distance d to the hypersurface of the ith event cluster i , the formula is: , Among them, d i is the pulse distance of the ith event cluster, K is the number of neurons, and w ij is the synaptic weight of the jth neuron in the ith event cluster, δ is the impulse response function, σ j is the receptive field width of the jth neuron, is the end time t of the current time window and the pulse emission time t of neuron j j The absolute difference of .
4. The method for generating airport flight video slices according to claim 3, characterized in that: The trigger cluster center pulse plasticity update includes: For the matching event clusters, the synaptic weights are adjusted according to the pulse timing-dependent plasticity rule, and the adjustment formula is: , in, is the synaptic weight w ij The adjustment amount, η is the learning rate, τ is the time constant, Δt is the time difference between the postsynaptic neuron and the presynaptic neuron pulse, sign(Δt) is the sign function, when Δt>0, it is 1, the weight is strengthened, when Δt<0, it is -1, the weight is weakened.
5. The method for generating airport flight video slices according to claim 3, characterized in that: The identifying event boundaries according to the newly added pulse pattern to adjust the shape of the event cognitive hypersurface includes: The objective function is to minimize the intra-cluster pulse distance variance and model complexity. The objective function is: , Among them, L is the objective function used to measure the clustering quality, λ is the regularization coefficient, is the Frobenius norm of the synaptic weight matrix; By pulse gradient descent, the synaptic weights are iteratively updated, and the update formula is: , in, is the updated synaptic weight matrix of the ith event cluster, is the synaptic weight matrix of the ith event cluster before updating, The objective function for the synaptic weight matrix W i The gradient of , γ is the step size, and β is the momentum coefficient is the momentum term of the previous weight update.
6. The method for generating airport flight video slices according to claim 1, characterized in that: The clustering model is constructed, comprising: Collect historical video streams of airport scenes and associated flight status data to generate historical target feature vectors; The historical target feature vector is converted into a pulse sequence, input into the pulse neural network, and the cluster center is initialized by the self-organizing map algorithm; The cluster centers are optimized based on the pulse gradient descent algorithm, and the hypersurface parameters of the event cluster are defined to minimize the pulse distance variance of samples within the cluster. According to the optimized hypersurface parameters, the reference slice parameters of each event cluster are calculated and stored, including the time window length, resolution threshold and coordinates of the region of interest.
7. The method for generating airport flight video slices according to claim 6, characterized in that: The step of slicing the real-time video data according to the reference slicing parameters of the matching event cluster includes: Read the predefined slice time window length T, resolution threshold R and region of interest ROI from the benchmark slice parameter library of the matching event cluster; Generate slicing instructions according to the slice time window length T, resolution threshold R and region of interest ROI; Slice the video stream according to the slicing instruction to generate segments matching the event type; The stability of the slice content is monitored through the pulse frequency of the SNN. If an abnormality is detected, the parameter re-evaluation mechanism is triggered.
8. The method for generating airport flight video slices according to claim 1, characterized in that: The step of associating the flight information with the video source address and calling the real-time video data includes: Create a data structure to establish and store the mapping relationship between flight information and the corresponding video data source address; A timer is set to call the flight information API at a preset time to obtain the latest flight data, and the flight data returned by the flight information API is parsed to update the flight information status and the mapping relationship in the data structure in real time; According to the real-time flight information, the corresponding video data source is dynamically matched and switched from the updated mapping relationship of the data structure to obtain the real-time video stream data.
9. The method for generating airport flight video slices according to claim 8, characterized in that: The method of dynamically matching and switching to a corresponding video data source from the updated mapping relationship of the data structure according to the real-time flight information to obtain real-time video stream data includes: When the flight information is updated, the corresponding new video data source in the mapping relationship of the data structure is immediately searched, and the new video data source is switched to obtain real-time video stream data from the new video data source.
10. An airport flight video slice generation device, used to implement the airport flight video slice generation method according to any one of claims 1 to 9, characterized in that: The device comprises: The first data module: used to obtain the current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate and baggage carousel number; The second data module is used to obtain the video source address of the key area, which includes the flight area, boarding gate, baggage carousel and waiting area; Association module: used to associate flight information with the video source address and call real-time video data; Feature extraction module: used to extract features from real-time video data of flights to obtain target feature vectors; Classification module: used to build a clustering model, add the current target feature vector to the clustering model, and build an event cognitive hypersurface through a pulse neural network to optimize the clustering model's dynamic division of real-time video data and match the benchmark slice parameters; Slicing module: used to slice real-time video data according to matching benchmark slicing parameters to generate video slices.
Citation Information
Patent Citations
Target monitoring method for airport field monitoring system on basis of video recognition
CN104243935A
Target tracking method, device and system, electronic device and storage medium
CN113536993A
Airport arrival and departure flight information accurate display system for serving passengers
CN119323905A
System for real time monitoring video of subway in relation to fire detection
KR101758008B1