An airport flight video slicing generation method and device
By obtaining the association between flight information and video source address in the airport flight video surveillance system, and using pulsed neural network to optimize the clustering model, dynamically divide the video data, and generate structured slices matching the event type, the problems of storage waste and inefficient retrieval in traditional systems are solved, and efficient video data management and rapid response are achieved.
Patent Information
- Application Number
- CN202510480144.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Traditional airport flight video surveillance systems have problems such as wasting storage resources, inefficient retrieval, inefficient manual monitoring and information islands. They cannot effectively associate real-time flight data with video, resulting in passive response and difficulty in quickly obtaining relevant video information.
By obtaining the precise association between flight information and video source address, using pulsed neural networks to build event cognitive hypersurfaces, optimize the clustering model to dynamically divide the real-time video data and match the reference slice parameters, and generate structured video slices that are highly matched with the event type.
It realizes the real-time and rapid call of video data, dynamically adapts to changes in airport scenes, generates video slices that are highly matched with event types, supports multi-grained analysis and efficient storage, and improves video retrieval efficiency and airport operation efficiency.
Smart Images

Figure CN120017925B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video slicing, and particularly to a method and device for generating airport flight video slices. Background Art
[0002] In modern airport management, video monitoring and recording of airport flights have become increasingly important and have become an indispensable tool for ensuring flight safety and improving service quality.
[0003] Traditional flight video monitoring systems usually use fixed cameras for 24 / 7 global monitoring. Although they can provide continuous video streams, there are many deficiencies in aspects such as information storage, retrieval, processing, and utilization.
[0004] Specifically, traditional methods usually store videos with fixed resolution and time windows, resulting in waste of storage resources and low retrieval efficiency. For example, for video segments without abnormal events for a long time, storing them at high resolution will occupy a large amount of storage space; while when quickly retrieving specific events, due to the fixed time window, it may not be possible to accurately locate the specific time period when the event occurred.
[0005] Traditional airport management relies on manual monitoring of video data, with passive response and low efficiency. Manual monitoring is difficult to capture all abnormal events in real time, and long-term monitoring is prone to fatigue and negligence. In addition, the manual response speed is slow, and it is impossible to respond to emergencies in a timely manner, affecting the safety and operational efficiency of the airport.
[0006] Traditional systems cannot effectively associate real-time data of flights with videos, resulting in the emergence of the information island problem. This makes it difficult for airport operators to quickly obtain relevant video information for specific flights when dealing with flight anomalies, settlement and billing, passenger complaints, etc. Summary of the Invention
[0007] Based on this, the objective of the present invention is to propose a method and device for generating airport flight video slices to solve the above-mentioned problems.
[0008] According to a method for generating airport flight video slices proposed by the present invention, the method includes:
[0009] Obtain current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate, and baggage carousel number;
[0010] Obtain the video source address of key areas, where the key areas include the flight area, boarding gate, baggage carousel, and waiting area;
[0011] Associate the flight information with the video source address and call real-time video data;
[0012] Extract features from the real-time video data of the flight to obtain the target feature vector;
[0013] Construct a clustering model, add the current target feature vector to the clustering model, and construct an event recognition hypersurface through a spiking neural network to optimize the dynamic division of the real-time video data by the clustering model and the matching of the reference slice parameters;
[0014] Slice the real-time video data according to the matched reference slice parameters to generate video slices.
[0015] Furthermore, the construction of the event recognition hypersurface through the spiking neural network to optimize the dynamic division of the real-time video data by the clustering model and the matching of the reference slice parameters includes:
[0016] Calculate the spiking distance from the current target feature vector to the hypersurface of each known event cluster;
[0017] If the spiking distance is less than the distance threshold, assign the current target feature vector to the nearest cluster and trigger the spiking plasticity update of the cluster center;
[0018] If the spiking distance is greater than the distance threshold, create a new event cluster and initialize the reference slice parameters;
[0019] Identify the event boundary based on the newly added spiking pattern to adjust the shape of the event recognition hypersurface.
[0020] Furthermore, the calculation of the spiking distance from the current target feature vector to the hypersurface of each known event cluster includes:
[0021] Calculate the spiking distance d t from the current target feature vector x i to the hypersurface of the i-th event cluster through the SNN spiking response kernel function, and the formula is:
[0022] ,
[0023] where d i is the spiking distance of the i-th event cluster, K is the number of neurons, w ij is the synaptic weight of the j-th neuron in the i-th event cluster, δ is the spiking response function, σ j is the receptive field width of the j-th neuron, is the absolute difference between the end time t of the current time window and the spiking time t j of neuron j.
[0024] Furthermore, the triggering of the spiking plasticity update of the cluster center includes:
[0025] For the matched event clusters, adjust the synaptic weights according to the pulse-timing-dependent plasticity rule. The adjustment formula is as follows:
[0026] ,
[0027] where, is the adjustment amount of the synaptic weight w ij , η is the learning rate, τ is the time constant, Δt is the time difference between the spike firings of the postsynaptic neuron and the presynaptic neuron, and sign(Δt) is the sign function, which is 1 when Δt > 0, enhancing the weight, and -1 when Δt < 0, weakening the weight.
[0028] Furthermore, recognize the event boundary according to the newly added pulse pattern to adjust the shape of the event recognition hypersurface, including:
[0029] Take minimizing the variance of the intra-cluster pulse distance and the model complexity as the objective function. The objective function is:
[0030] ,
[0031] where, L is the objective function used to measure the clustering quality, λ is the regularization coefficient, is the Frobenius norm of the synaptic weight matrix.
[0032] Update the synaptic weights iteratively through pulse gradient descent. The update formula is:
[0033] ,
[0034] where, is the synaptic weight matrix of the i-th event cluster after update, is the synaptic weight matrix of the i-th event cluster before update, is the gradient of the objective function with respect to the synaptic weight matrix W i , γ is the step size, and β is the momentum coefficient is the momentum term of the previous weight update.
[0035] Furthermore, the construction of the clustering model includes:
[0036] Collect the historical video stream of the airport scene and the associated flight status data to generate historical target feature vectors;
[0037] Convert the historical target feature vectors into pulse sequences, input them into the spiking neural network, and initialize the clustering centers through the self-organizing mapping algorithm;
[0038] Optimize the clustering centers based on the pulse gradient descent algorithm, define the hypersurface parameters of the event clusters, and minimize the variance of the pulse distance of the samples within the clusters;
[0039] Calculate and store the reference slice parameters for each event cluster according to the optimized hypersurface parameters, including the time window length, resolution threshold, and coordinates of the region of interest.
[0040] Furthermore, slicing the real-time video data according to the reference slice parameters of the matching event cluster includes:
[0041] Read the predefined time window length T, resolution threshold R, and region of interest ROI of the slice from the reference slice parameter library of the matching event cluster;
[0042] Generate a slice instruction according to the time window length T, resolution threshold R, and region of interest ROI of the slice;
[0043] Slice the video stream according to the slice instruction to generate a segment matching the event type;
[0044] Monitor the stability of the slice content through the pulse frequency of the SNN. If an anomaly is detected, trigger the parameter re-evaluation mechanism.
[0045] Furthermore, associating the flight information with the video source address and calling the real-time video data includes:
[0046] Create a data structure to establish and store the mapping relationship between the flight information and the corresponding video data source address;
[0047] Set a timer to call the flight information API every preset time to obtain the latest flight data, and parse the flight data returned by the flight information API to update the flight information status and the mapping relationship in the data structure in real time;
[0048] According to the real-time flight information, dynamically match and switch to the corresponding video data source from the updated mapping relationship of the data structure to obtain the real-time video stream data.
[0049] Furthermore, according to the real-time flight information, dynamically matching and switching to the corresponding video data source from the updated mapping relationship of the data structure to obtain the real-time video stream data includes:
[0050] When the flight information is updated, immediately search for the corresponding new video data source in the mapping relationship of the data structure, switch to the new video data source, and obtain the real-time video stream data from the new video data source.
[0051] The present invention also proposes an airport flight video slice generation device for implementing the above-mentioned airport flight video slice generation method. The device includes:
[0052] The first data module: used to obtain the current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate, and baggage carousel number;
[0053] The second data module: used to obtain the video source addresses of key areas, where the key areas include the flight area, boarding gate, baggage carousel, and waiting area;
[0054] The association module: used to associate the flight information with the video source addresses and call the real-time video data;
[0055] The feature extraction module: used to extract features from the real-time video data of the flight to obtain the target feature vector;
[0056] The classification module: used to construct a clustering model, add the current target feature vector to the clustering model, and construct an event recognition hypersurface through a spiking neural network to optimize the dynamic partitioning of the clustering model for the real-time video data and the matching of the benchmark slice parameters;
[0057] The slicing module: used to slice the real-time video data according to the matched benchmark slice parameters to generate video slices.
[0058] In summary, for the method for generating airport flight video slices of the present invention, the current flight information is obtained and accurately associated with the video source addresses of key areas to ensure that the corresponding video data can be called in real time according to the flight information. The features of the called real-time video data are extracted to obtain the target feature vector. Subsequently, an event recognition hypersurface is constructed through a spiking neural network to dynamically optimize the partitioning ability of the clustering model for the real-time video data, realize the accurate recognition of event clusters and the automatic matching of benchmark slice parameters, so as to adapt to the rapid changes in the airport scenario, such as sudden passenger flow, flight delays, etc.; finally, the video is accurately cut according to the slice parameters matched by the event clusters to generate structured video slices highly matching the event types, such as flight takeoff and landing segments, abnormal baggage carousel segments, to support multi-granularity analysis and efficient storage.
[0059] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the embodiments of the present invention. Description of the Drawings
[0060] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:
[0061] Figure 1 is a flowchart of a method for generating airport flight video slices according to Embodiment 1 of the present invention;
[0062] Figure 2It is a system block diagram of an airport flight video slicing generation device according to Embodiment 2 of the present invention. Detailed implementation manners
[0063] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0064] It should be noted that when an element is referred to as being "fixedly provided on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are only for the purpose of illustration.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0066] Embodiment 1: Please refer to Figure 1 , the present invention proposes an airport flight video slicing generation method, and the method includes steps S101 to S106:
[0067] S101, obtain the current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate and baggage carousel number.
[0068] Collect the basic information related to the current flight, such as flight number (used to uniquely identify the flight), estimated departure / arrival time (used to determine the time range of flight activities), boarding gate (the specific location where passengers board the plane), and baggage carousel number (the location where passengers pick up their luggage), etc. These information are the basis for associating and retrieving video data in the subsequent steps.
[0069] Specifically, it can be docked with the information system of the airline or airport to obtain the real-time flight information API. This API (i.e., the interface for obtaining flight-related information) can provide key information such as flight number, boarding gate, baggage carousel, estimated departure / arrival time, etc. Obtaining flight information through the API can ensure the real-time, comprehensive and accurate nature of the data, and avoid problems caused by manual input or outdated data.
[0070] S102, Obtain the video source addresses of key areas, where the key areas include the flight area, boarding gates, baggage carousels, and waiting areas.
[0071] By obtaining the video source addresses of key areas, the dynamic situations in key areas within the airport can be captured. The key areas include the flight area (where airplanes take off and land), boarding gates (where passengers board and alight from the plane), baggage carousels (where passengers claim their luggage), and waiting areas (where passengers wait for boarding), etc.
[0072] Since there are many cameras installed within the airport, especially in key areas, each camera has an address (such as a web address). Through this address, the video stream footage captured by the camera can be viewed and called at any time. Through the video stream address corresponding to each camera in the key area (such as RTSP URL), the system can obtain the video data of these areas in real time.
[0073] S103, Associate the flight information with the video source addresses and call the real-time video data.
[0074] Establish a connection between the specific flight information (such as flight number, boarding gate, baggage carousel number, etc.) and the video source addresses of the corresponding areas within the airport to ensure that relevant video data can be quickly located based on the flight information. Specifically, a data structure (such as a mapping table or database) can be established, which stores the correspondence between flight information and video source addresses. When the flight information changes (such as a change in the boarding gate), this data structure is updated in a timely manner. Thus, the flight information and video source addresses are accurately associated to ensure that the correct video data can be called subsequently.
[0075] According to the latest flight information, obtain the video data of the corresponding area in real time to provide real-time and accurate video data for video slicing generation. Specifically, the data structure (such as a mapping table or database) can be queried to find the video source address corresponding to the current flight information. Use video stream acquisition technologies (such as RTSP, HLS, etc. protocols) to obtain the video data from the video source address in real time. Since the acquisition of video data is real-time and closely associated with the flight status information, it can accurately reflect the current operating status of the flight in real time.
[0076] Further optionally, in step S103, the associating the flight information with the video source addresses and calling the real-time video data includes:
[0077] Create a data structure to establish and store the mapping relationship between the flight information and the corresponding video data source addresses (i.e., the corresponding camera addresses);
[0078] A timer is set to call the flight information API at a preset time to obtain the latest flight data, and the flight data returned by the flight information API is parsed to update the flight information status and the mapping relationship in the data structure in real time;
[0079] According to the real-time flight information, the corresponding video data source is dynamically matched and switched from the updated mapping relationship of the data structure to obtain the real-time video stream data.
[0080] Further optionally, dynamically matching and switching to a corresponding video data source from the updated mapping relationship of the data structure according to the real-time flight information to obtain real-time video stream data includes:
[0081] When the flight information is updated, the corresponding new video data source (i.e., the new camera address) in the mapping relationship of the data structure is immediately searched, and the new video data source is switched to obtain real-time video stream data from the new video data source to ensure that the video image is consistent with the updated flight information.
[0082] Understandably, in order to effectively associate current flight information with real-time video data, it is necessary to obtain accurate current flight information, such as flight number, estimated departure time, estimated arrival time, boarding gate, baggage carousel number and other key data. By connecting with the information system of the airline or airport, you can obtain real-time flight information API, which can provide all the required flight information, ensuring the real-time and accuracy of the data.
[0083] The video stream address of the airport monitoring system can be used to obtain real-time video data of key areas (such as the flight area, boarding gate, baggage carousel and waiting area, which are important monitoring points in the flight operation process). Since the airport is full of cameras, especially in key areas, each camera has a unique address, through which the camera can be called and viewed at any time.
[0084] In order to closely associate flight information with real-time video data and call real-time video data based on flight information, a data structure (such as a mapping table or database) is created to establish and store the mapping relationship between flight information and the corresponding video data source address (i.e., camera address). The specific mapping example is as follows:
[0085] Flight number: FL123
[0086] Area 1 (flying area): Camera address A1
[0087] Area 2 (boarding gate): Camera address A2
[0088] Area 3 (baggage carousel): Camera address A3
[0089] Area 4 (Waiting Area): Camera Address A4
[0090] When flight FL123 is in the flight area, the system will obtain video data from camera address A1; when the flight arrives at the boarding gate, the system will switch to camera address A2; and so on.
[0091] However, when the flight information changes, the mapping relationship will also be updated. Flight information, such as boarding gate, baggage carousel, estimated departure / arrival time, etc., may change due to various reasons (such as weather, mechanical failure, scheduling adjustment, etc.). In this case, it is necessary to update the mapping relationship between the flight information and the video data source in real time.
[0092] The change notice of flight information can be obtained in real time through the information system of the airline or the airport, and according to the change notice, the mapping relationship in the previously created data structure can be updated immediately. Then, according to the updated mapping relationship, dynamically switch to the new video data source to obtain the real-time video data corresponding to the changed flight information. However, although the status of each flight will change (such as changing the boarding gate from B1 to B2), the mapping relationship between each area and its corresponding video source address remains unchanged (for example, the camera address corresponding to boarding gate B2 is always A2). The specific example is as follows:
[0093] Initial mapping relationship: Flight number: FL123
[0094] Boarding gate: B1
[0095] Baggage carousel: C1
[0096] Camera address: A1 (corresponding to boarding gate B1)
[0097] Flight information changes, and the change content is: The boarding gate is changed from B1 to B2
[0098] Updated mapping relationship: Flight number: FL123
[0099] Boarding gate: B2
[0100] Baggage carousel: C1
[0101] Camera address: A2 (corresponding to boarding gate B2)
[0102] Since flight information is dynamically changing, in order to ensure that the obtained flight information is always the latest. A timer can be set to call the flight information API every preset time (such as every minute) to obtain the latest flight data. Then parse these returned data to update the status and mapping relationship of the flight information in real time.
[0103] Finally, according to the real-time updated flight information, dynamically match and switch to the corresponding video data source from the updated data structure, so as to obtain the real-time video stream data closely related to the current flight. In this way, the efficient association and processing of flight information and video data can be realized, and various situations during the flight operation can be grasped in real time through the video monitoring system, providing strong support for flight safety and management.
[0104] S104, extract the features of the real-time video data of the flight to obtain the target feature vector.
[0105] For the synchronously collected airport scene monitoring video stream and flight status data, extract the spatio-temporal features and structured information of dynamic targets to generate a high-dimensional target feature vector. The specific implementation process is as follows:
[0106] First, perform feature extraction on the video stream. The YOLOv8 object detection algorithm can be used to process the video stream in real time to identify dynamic targets such as pedestrians, vehicles, and luggage carts, and extract their spatial coordinates (x position, y position), motion parameters (speed, acceleration), and density (density, such as the number of targets per unit area). For example, in the boarding gate area, the pedestrian movement trajectory can be detected as [(100, 200), (105, 202), (110, 205)], and the speed is calculated to be 2.5 pixels / frame, and the density is 5 people / square meter. To eliminate video noise (such as jitter, occlusion), the trajectory data is smoothed through the Kalman filter algorithm and aligned with the flight status data based on the timestamp to solve the data asynchrony problem.
[0107] Secondly, parse the flight status data to obtain structured information such as flight number, scheduled departure and arrival times, actual departure and arrival times, boarding gate, baggage carousel number, and delay status (delay flag, such as 0 for normal and 1 for delay) from the airport operation system. For example, flight CA1234 is scheduled to depart at 10:00, actually departs at 10:30, the delay flag is 1, and the boarding gate is B12. These data provide key context associations for the video stream features, such as accurately identifying the relationship between flight delays and changes in the crowd density in the waiting area.
[0108] Next, perform feature fusion. After aligning the video stream features and flight status data by timestamp, generate a target feature vector containing spatio-temporal coordinates, motion parameters, crowd density, and flight status. Its structure can be: [x position, y position, speed, density, delay flag, timestamp, flight number, boarding gate, baggage carousel number].
[0109] For example, the feature vector [120, 300, 1.8, 4.5, 1, "2024-05-01 10:30:00", "CA1234","B12", "D05"] indicates that at 10:30:00 on May 1, 2024, near boarding gate B12 of flight CA1234 (delayed), a pedestrian was detected at coordinates (120, 300), with a speed of 1.8 m / s, a density of 4.5 people per square meter, and the associated baggage carousel is D05.
[0110] Finally, to improve the feature stability, a sliding window aggregation (such as a 5-second window) is performed on the feature vectors of consecutive frames, and the average value or maximum value is calculated. At the same time, the high-dimensional feature vectors are reduced in dimension through principal component analysis, retaining the main information and reducing the computational complexity. For example, the original 128-dimensional video features can be reduced to 32 dimensions while retaining 95% of the variance information, significantly improving the efficiency of the subsequent clustering model.
[0111] S105, construct a clustering model, add the current target feature vector to the clustering model, and construct an event recognition hypersurface through a spiking neural network to optimize the dynamic partitioning of the clustering model for real-time video data and the matching of benchmark slice parameters.
[0112] Constructing a dynamic clustering model based on a spiking neural network breaks through the static limitations of traditional methods. Due to the highly dynamic nature of airport scenarios, such as sudden passenger flows and flight schedule changes, traditional clustering models (such as K-means) cannot adjust the event cluster boundaries in real time. The present invention introduces a spiking neural network to perform spatio-temporal sequence modeling on real-time video feature vectors (such as passenger density, baggage movement speed) by simulating neuron pulse signals, and can dynamically capture the evolution law of events. For example, when the waiting area changes from "normal" to "congested", the spiking neural network can adaptively adjust the event cluster partitioning and match the benchmark slice parameters, thus greatly improving the accuracy of event recognition, reducing the matching error of slice parameters, and being able to effectively handle the complex scenario changes in the airport.
[0113] Further optionally, the constructing the clustering model includes:
[0114] Collect the historical video stream of the airport scenario and the associated flight status data to generate historical target feature vectors;
[0115] Convert the historical target feature vectors into pulse sequences, input them into the spiking neural network, and initialize the clustering centers through the self-organizing mapping algorithm;
[0116] Optimize the clustering centers based on the pulse gradient descent algorithm, define the hypersurface parameters of the event clusters, and minimize the pulse distance variance of the samples within the clusters;
[0117] Calculate and store the reference slice parameters for each event cluster according to the optimized hypersurface parameters, including the time window length, resolution threshold, and region of interest coordinates.
[0118] Specifically, for the synchronized surveillance video stream and flight status data of the airport scenario, extract the spatio-temporal features and structured information of dynamic targets. The video stream is processed by an object detection algorithm to obtain the trajectories, speeds, and densities of pedestrians and vehicles; the flight status data provides key information such as flight numbers, takeoff and landing times, and delay status. After aligning the two types of data by timestamp, a target feature vector is generated.
[0119] The historical target feature vectors need to be converted into spatio-temporal spike patterns of a spiking neural network. Pulse coding can be performed by combining frequency coding (the larger the eigenvalue, the higher the pulse frequency) and time coding (key events trigger early pulses) to simulate the information transmission mechanism of biological neurons. In the initialization stage, the self-organizing mapping algorithm is used to map the high-dimensional pulse sequence to low-dimensional grid nodes. Through competitive learning, the grid nodes adaptively adjust their weights to approximate the data distribution and form initial clustering centers. Thus, there is no need to preset the number of clusters, and the internal topological structure of the data can be automatically captured, providing a robust starting point for subsequent optimization.
[0120] To improve the clustering accuracy, a pulse gradient descent algorithm is introduced to optimize the hypersurface parameters. The optimization goal is to minimize the variance of the pulse distance within the cluster. By defining the distance metric between pulse sequences (such as dynamic time warping), parameters such as the coefficients of the hyper-ellipsoid equation are adjusted. Pulse gradient descent iteratively updates the clustering centers by backpropagating the gradient of the pulse activity, making the hypersurface better fit the data distribution.
[0121] Calculate and store the reference slice parameters for each event cluster according to the optimized hypersurface parameters, including the time window length, resolution threshold, and region of interest coordinates. Among them, the time window length is inversely proportional to the average pulse frequency within the cluster. High-frequency events (such as dense crowds) can use short windows to capture transients, and low-frequency events (such as flight takeoffs and landings) can use long windows to analyze trends. The resolution threshold is dynamically adjusted according to the variance of the feature vector, and the resolution is increased when the variance is large to capture details. The region of interest coordinates are generated based on the spatial distribution statistics of the samples within the cluster, such as the centroid or bounding box, which can locate the key monitoring areas. These parameters convert the clustering results into operable slice configurations, ensuring that the slice content highly matches the event type and improving the efficiency and accuracy of subsequent video analysis.
[0122] Event clusters can include various typical or special event clusters such as flight takeoff and landing event clusters, boarding gate activity event clusters, baggage carousel event clusters, passenger flow event clusters in the waiting area, security checkpoint event clusters, abnormal behavior event clusters, equipment failure event clusters, and environmental interference event clusters.
[0123] The clustering model of the present invention realizes the efficient recognition and slicing of airport scene events through the dynamic clustering ability and adaptive parameter generation mechanism of the spiking neural network.
[0124] Further optionally, constructing an event cognitive hypersurface through the spiking neural network to optimize the dynamic partitioning of the clustering model for real-time video data and the matching of benchmark slicing parameters, including:
[0125] Calculating the spiking distance from the current target feature vector to the hypersurface of each known event cluster;
[0126] If the spiking distance is less than the distance threshold, the current target feature vector is assigned to the nearest cluster, and the spiking plasticity update of the cluster center is triggered;
[0127] If the spiking distance is greater than the distance threshold, a new event cluster is created and the benchmark slicing parameters are initialized;
[0128] Identifying event boundaries based on the newly added spiking pattern to adjust the shape of the event cognitive hypersurface.
[0129] Specifically, an event cognitive hypersurface is constructed through the spiking neural network to achieve dynamic adaptability and real-time event partitioning in the airport scene. Specifically, the clustering model dynamically assigns data points by calculating the spiking distance from the current target feature vector to the hypersurface of each known event cluster and combining the distance threshold mechanism. If the spiking distance is less than the threshold, the target feature vector is assigned to the nearest cluster, and the spiking plasticity update of the cluster center is triggered to adjust the synaptic weights to adapt to the new data, such as the change in passenger flow caused by flight delays. If the spiking distance is greater than the threshold, a new event cluster is created, such as a newly added temporary flight or an abnormal event, and the benchmark slicing parameters (such as the time window length, resolution threshold, etc.) are initialized. This process simulates the mechanism of the human brain to recognize familiar and unfamiliar things through the neuron firing pattern, enabling the clustering model to capture the changes in event patterns in the airport in real time (such as fluctuations in flight arrival and departure times, abnormal luggage carousel traffic), reducing the dependence on manual annotation, and significantly improving the adaptability to dynamic scenarios, such as dynamically adjusting the number of open security check channels to cope with the passenger flow peak.
[0130] The clustering model further utilizes new pulse patterns (such as sudden changes in pulse frequency and changes in spatial distribution) to detect event boundaries and dynamically adjusts synaptic connection strengths and hypersurface parameters (such as curvature and direction). For example, when there are abnormal crowd aggregations or evacuations in the boarding gate area, sudden changes in pulse frequency can trigger boundary detection, and the model can more accurately fit the actual event characteristics by adjusting the hypersurface curvature. This process is similar to an airport navigation software adjusting the boarding gate path planning according to real-time flight dynamics, enabling the hypersurface to distinguish between similar events, such as boarding gate activities of different flights or conflicts in baggage carousel assignments, to reduce misjudgments. Through the boundary evolution mechanism, the clustering model can effectively identify subtle differences in airport operations, such as differences in passenger flow patterns between international and domestic flights, which can greatly improve the accuracy and reliability of event recognition.
[0131] Through dynamic clustering, event boundary detection, and biologically inspired learning mechanisms, the present invention can capture complex event patterns in airport scenarios in real time, such as fluctuations in takeoff and landing times and changes in taxiing paths of flights, and passenger flow distributions (such as aggregations at boarding gates and abnormal flows at baggage carousels). By dynamically adjusting the shape of the hypersurface, the model can accurately distinguish between similar events (such as differences in passenger activities at different boarding gates and conflicts in baggage sorting at adjacent baggage carousels), significantly reducing the event misjudgment rate. Finally, by optimizing the matching degree between benchmark slice parameters (such as time window length, resolution threshold, etc.) and event clusters in the video stream, it provides highly robust support for efficient slicing and intelligent analysis of airport flight videos, helping airport operations achieve refined resource scheduling and rapid response to abnormal events.
[0132] Further optionally, calculating the pulse distance from the current target feature vector to each known event cluster hypersurface includes:
[0133] Calculating the pulse distance d from the current target feature vector x t to the i-th event cluster hypersurface, the formula is: i
[0134]
[0135] where d i is the pulse distance of the i-th event cluster, K is the number of neurons, w ij is the synaptic weight of the j-th neuron in the i-th event cluster, δ is the pulse response function, σ j is the receptive field width of the j-th neuron, used to control the spatial attenuation rate of the pulse response, is the absolute difference between the end time t of the current time window and the pulse firing time t j of neuron j.
[0136] Through the pulse response kernel function of the SNN, the present invention combines spatio-temporal features (time difference 、Receptive field width σ j ), combined with synaptic weight w ij ), realizes the dynamic similarity measurement of the event cluster hypersurface, thus realizing dynamic clustering and event partitioning. The impulse distance reflects the similarity between the feature vector and the event cluster hypersurface, and is the key basis for judging the attribution of data points (assigned to existing clusters or creating new clusters).
[0137] Further optionally, before dynamically calculating the impulse distance from the current target feature vector to each known event cluster hypersurface, it includes:
[0138] Dynamically calculate the distance threshold θ i = μ i + ασ i , where μ i and σ i are the mean distance and standard deviation within the i-th event cluster respectively, and α is a control parameter (default value is 2).
[0139] By dynamically calculating the distance threshold and the impulse distance, it realizes the adaptive clustering and precise partitioning of complex event patterns such as flight dynamics and passenger flow changes in the airport scenario. Specifically, by statistically calculating the mean and standard deviation of the impulse distances within each event cluster, a personalized threshold is dynamically generated. Combining with the spatio-temporal response characteristics of the spiking neural network, the similarity between the target feature vector and the event cluster is judged in real time, so as to accurately distinguish similar events (such as activities at different boarding gates or baggage carousels), reduce the misjudgment rate, and provide benchmark slice parameters with high matching degree for airport flight video slices, ultimately improving the resource scheduling efficiency and the ability to respond to abnormal events.
[0140] Further optionally, the impulse plasticity update of the trigger cluster center includes:
[0141] For the matching event cluster, adjust the synaptic weight according to the impulse timing-dependent plasticity rule. The adjustment formula is:
[0142] ,
[0143] where is the adjustment amount of the synaptic weight w ij . η is the learning rate, which is used to control the step size of weight update. τ is the time constant, which determines the influence range of the impulse time difference on weight adjustment. Δt is the time difference between the spike firing of the postsynaptic neuron and the presynaptic neuron. sign(Δt) is the sign function, which is 1 when Δt > 0, enhancing the weight, and -1 when Δt < 0, weakening the weight.
[0144] By triggering the pulse plasticity update at the cluster center, the synaptic weights are dynamically adjusted based on the spike-timing-dependent plasticity rule to optimize the event clustering ability of the airport flight video slicing method. Specifically, for the matched event clusters, the spike-timing-dependent plasticity rule is used to quantify the time difference between the pre- and post-synaptic neuron spikes. Through the sign function and the directional weight adjustment mechanism (enhancing the weight when Δt>0 and weakening the weight when Δt<0), the clustering model can adaptively learn the internal correlations of spatio-temporal features such as flight dynamics and passenger flow changes, thereby enhancing the representation accuracy and stability of the event clusters.
[0145] Further optionally, the method of recognizing the event boundary according to the newly added pulse pattern to adjust the shape of the event recognition hypersurface includes:
[0146] Taking the minimization of the within-cluster pulse distance variance and the model complexity as the objective function, the objective function is:
[0147] ,
[0148] where L is the objective function for measuring the clustering quality, λ is the regularization coefficient for balancing the within-cluster variance and the model complexity, is the Frobenius norm of the synaptic weight matrix for regularization.
[0149] Through pulse gradient descent, the synaptic weights are iteratively updated, and the update formula is:
[0150] ,
[0151] where, is the synaptic weight matrix of the i-th event cluster after update, is the synaptic weight matrix of the i-th event cluster before update, is the gradient of the objective function with respect to the synaptic weight matrix W i γ is the step size for controlling the update amplitude of each iteration, and β is the momentum coefficient for accelerating convergence and reducing oscillations, is the momentum term of the previous weight update.
[0152] Through the above steps, the hypersurface parameters can be adaptively adjusted in complex dynamic scenarios (such as flight takeoffs and landings, passenger flow fluctuations), thereby improving the robustness and efficiency of event boundary recognition.
[0153] S106, slice the real-time video data according to the matched reference slice parameters to generate video slices.
[0154] Traditional video storage usually adopts fixed resolution and time window, resulting in storage waste and inefficient retrieval. The slicing parameters of the present invention are matched according to event clusters, that is, dynamically adjusted according to event types. For example, high resolution (such as 1080p) is adopted for key areas (such as boarding gates), and low resolution (such as 720p) is adopted for non-key areas (such as corridor backgrounds); the time window is adaptively adjusted according to the duration of the event (such as the slicing of flight takeoff and landing covers the whole taxiing process). Thus, the video storage cost is greatly reduced, and at the same time, the abnormal event retrieval time can be shortened from the hour level to the minute level, supporting the collaborative work of global monitoring and local detection.
[0155] Further optionally, slicing the real-time video data according to the reference slicing parameters of the matching event cluster includes:
[0156] Reading the predefined time window length T, resolution threshold R, and region of interest ROI of the slice from the reference slicing parameter library of the matching event cluster;
[0157] Generating a slicing instruction according to the time window length T, resolution threshold R, and region of interest ROI of the slice;
[0158] Slicing the video stream according to the slicing instruction to generate a segment matching the event type;
[0159] Monitoring the stability of the slice content through the pulse frequency of the SNN, and triggering a parameter re-evaluation mechanism if an abnormality is detected.
[0160] Specifically, first extract the predefined slicing parameters from the reference slicing parameter library of the matching event cluster, including the time window length T, resolution threshold R, and region of interest ROI. These parameters provide spatio-temporal constraints for video slicing. Among them, T controls the time granularity, R determines the spatial clarity, and ROI locates the key area to ensure that the slice content is highly relevant to the event type (such as boarding, baggage claim).
[0161] Generating a slicing instruction based on the extracted parameters to drive video stream segmentation. This process converts the continuous video stream into segments matching the event semantics. For example, the video of the boarding gate area is cut according to the flight takeoff and landing time window T, or the video of a specific period is extracted according to the ROI of the baggage claim carousel.
[0162] Utilize the pulse frequency of the spiking neural network to monitor the stability of the slice content in real time. If an abnormality (such as jitter or blur in the picture) is detected, the parameter re-evaluation mechanism is triggered to dynamically adjust T, R, or ROI to maintain the slice quality. For example, when the video of the baggage carousel is blurred due to camera jitter, T can be shortened or R can be increased to enhance the clarity of the key frames.
[0163] In summary, for the method of generating airport flight video slices of the present invention, the current flight information is obtained and accurately associated with the video source address of the key area to ensure that the corresponding video data can be called in real time and quickly according to the flight information. Feature extraction is performed on the called real-time video data to obtain the target feature vector. Subsequently, an event recognition hypersurface is constructed through a spiking neural network, and the partitioning ability of the clustering model for real-time video data is dynamically optimized to achieve accurate identification of event clusters and automatic matching of benchmark slice parameters, so as to adapt to the rapid changes in the airport scenario, such as sudden passenger flow, flight delays, etc.; finally, the video is accurately cut according to the slice parameters matched by the event clusters to generate structured video slices that highly match the event types, such as flight takeoff and landing segments, abnormal baggage carousel segments, to support multi-granularity analysis and efficient storage.
[0164] Embodiment 2: Please refer to Figure 2 , an airport flight video slice generating device proposed by the present invention, the device includes:
[0165] The first data module: used to obtain the current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate, and baggage carousel number;
[0166] The second data module: used to obtain the video source address of the key area, and the key area includes the flight area, boarding gate, baggage carousel, and waiting area;
[0167] The association module: used to associate the flight information with the video source address and call the real-time video data;
[0168] The feature extraction module: used to perform feature extraction on the real-time video data of the flight to obtain the target feature vector;
[0169] The classification module: used to construct a clustering model, add the current target feature vector to the clustering model, and construct an event recognition hypersurface through a spiking neural network to optimize the dynamic partitioning of the real-time video data by the clustering model and the matching of benchmark slice parameters;
[0170] The slicing module: used to slice the real-time video data according to the matched benchmark slice parameters to generate video slices.
[0171] Further optionally, the classification module is further used for:
[0172] Calculating the spiking distance from the current target feature vector to the hypersurface of each known event cluster;
[0173] If the spiking distance is less than the distance threshold, then assign the current target feature vector to the nearest cluster and trigger the spiking plasticity update of the cluster center;
[0174] If the spiking distance is greater than the distance threshold, then create a new event cluster and initialize the benchmark slice parameters;
[0175] Adjust the shape of the event perception hypersurface according to the newly added pulse pattern recognition event boundary.
[0176] Further optionally, the classification module is further configured to:
[0177] Calculate the current target feature vector x through the SNN pulse response kernel function t The pulse distance d to the i-th event cluster hypersurface i , and the formula is:
[0178] ,
[0179] where d i is the pulse distance of the i-th event cluster, K is the number of neurons, w ij is the synaptic weight of the j-th neuron in the i-th event cluster, δ is the pulse response function, σ j is the receptive field width of the j-th neuron, is the absolute difference between the end time t of the current time window and the pulse firing time t j of neuron j.
[0180] Further optionally, the classification module is further configured to:
[0181] For the matched event clusters, adjust the synaptic weights according to the pulse-timing-dependent plasticity rule, and the adjustment formula is:
[0182] ,
[0183] where, is the adjustment amount of the synaptic weight w ij , η is the learning rate, τ is the time constant, Δt is the time difference between the post-synaptic neuron and the pre-synaptic neuron pulse firing, and sign(Δt) is the sign function, which is 1 when Δt>0, enhancing the weight, and -1 when Δt<0, weakening the weight.
[0184] Further optionally, the classification module is further configured to:
[0185] Take minimizing the variance of the intra-cluster pulse distance and the model complexity as the objective function, and the objective function is:
[0186] ,
[0187] where L is the objective function, used to measure the clustering quality, λ is the regularization coefficient, is the Frobenius norm of the synaptic weight matrix.
[0188] Update the synaptic weights iteratively through pulse gradient descent, and the update formula is:
[0189] ,
[0190] wherein, is the synaptic weight matrix of the updated i-th event cluster, is the synaptic weight matrix of the i-th event cluster before update, is the gradient of the objective function with respect to the synaptic weight matrix W i , γ is the step size, and β is the momentum coefficient is the momentum term of the previous weight update.
[0191] Further optionally, the classification module is further configured to:
[0192] Collect the historical video stream of the airport scene and associated flight status data, and generate historical target feature vectors;
[0193] Convert the historical target feature vectors into pulse sequences, input them into the spiking neural network, and initialize the clustering centers through the self-organizing mapping algorithm;
[0194] Optimize the clustering centers based on the spiking gradient descent algorithm, define the hypersurface parameters of the event clusters, and minimize the pulse distance variance of the samples within the clusters;
[0195] According to the optimized hypersurface parameters, calculate and store the benchmark slice parameters of each event cluster, including the time window length, resolution threshold, and region of interest coordinates.
[0196] Further optionally, the slicing module is further configured to:
[0197] Read the predefined time window length T, resolution threshold R, and region of interest ROI of the slice from the benchmark slice parameter library of the matching event cluster;
[0198] Generate a slicing instruction according to the time window length T, resolution threshold R, and region of interest ROI of the slice;
[0199] Slice the video stream according to the slicing instruction to generate segments matching the event type;
[0200] Monitor the stability of the slice content through the pulse frequency of the SNN, and trigger the parameter re-evaluation mechanism if an anomaly is detected.
[0201] Further optionally, the association module is further configured to:
[0202] Create a data structure to establish and store the mapping relationship between flight information and the corresponding video data source address;
[0203] Set a timer to call the flight information API every preset time to obtain the latest flight data, and parse the flight data returned by the flight information API to update the flight information status and the mapping relationship in the data structure in real time;
[0204] According to the real-time flight information, dynamically match and switch to the corresponding video data source from the updated mapping relationship of the data structure, and obtain the real-time video stream data.
[0205] Further optionally, the association module is further configured to:
[0206] When the flight information is updated, immediately search for the corresponding new video data source in the mapping relationship of the data structure, switch to the new video data source, and obtain the real-time video stream data from the new video data source.
[0207] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A method for generating airport flight video slices, characterized in that, The method includes: Obtain the current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate, and baggage carousel number; Obtain the video source addresses of key areas, where the key areas include the flight area, boarding gate, baggage carousel, and waiting area; Associate the flight information with the video source addresses and call the real-time video data; Extract features from the real-time video data of the flight to obtain the target feature vector; Construct a clustering model, add the current target feature vector to the clustering model, and construct an event cognition hypersurface through a spiking neural network to optimize the dynamic partitioning of the real-time video data by the clustering model and the matching of the reference slice parameters; Slice the real-time video data according to the matched reference slice parameters to generate video slices; Among them, the constructing of the clustering model includes: Collect the historical video streams of the airport scene and the associated flight status data to generate historical target feature vectors; Convert the historical target feature vectors into spike trains, input them into the spiking neural network, and initialize the clustering centers through the self-organizing mapping algorithm; Optimize the clustering centers based on the spike gradient descent algorithm, define the hypersurface parameters of the event clusters, and minimize the variance of the spike distances of the samples within the clusters; According to the optimized hypersurface parameters, calculate and store the reference slice parameters of each event cluster, including the time window length, resolution threshold, and region of interest coordinates.
2. The method for generating airport flight video slices according to claim 1, wherein The constructing of the event cognition hypersurface through the spiking neural network to optimize the dynamic partitioning of the real-time video data by the clustering model and the matching of the reference slice parameters includes: Calculate the spike distance from the current target feature vector to the hypersurfaces of each known event cluster; If the spike distance is less than the distance threshold, assign the current target feature vector to the nearest cluster and trigger the spike plasticity update of the cluster center; If the spike distance is greater than the distance threshold, create a new event cluster and initialize the reference slice parameters; Identify the event boundaries according to the newly added spike pattern to adjust the shape of the event cognition hypersurface.
3. The method for generating airport flight video slices according to claim 2, wherein The calculating of the spike distance from the current target feature vector to the hypersurfaces of each known event cluster includes: Calculate the current target feature vector x through the SNN impulse response kernel function t The impulse distance d to the hypersurface of the i-th event cluster i , and the formula is: , where d i is the pulse distance of the i-th event cluster, K is the number of neurons, w ij is the synaptic weight of the j-th neuron in the i-th event cluster, δ is the impulse response function, σ j is the receptive field width of the j-th neuron, is the absolute difference between the current time window end time t and the spike emission time t j of neuron j.
4. The method for generating airport flight video slices according to claim 3, wherein, The triggering of the spike plasticity update of the cluster center includes: For the matched event cluster, adjust the synaptic weights according to the spike timing-dependent plasticity rule, and the adjustment formula is: , wherein, is the adjustment amount of the synaptic weight w ij , η is the learning rate, τ is the time constant, Δt is the time difference between the spike firings of the postsynaptic neuron and the presynaptic neuron, sign(Δt) is the sign function, which is 1 when Δt > 0 to enhance the weight, and -1 when Δt < 0 to weaken the weight.
5. The method for generating airport flight video slices according to claim 3, wherein The identifying of the event boundaries according to the newly added spike pattern to adjust the shape of the event cognition hypersurface includes: Take minimizing the variance of the spike distances within the cluster and the model complexity as the objective function; Through spike gradient descent, iteratively update the synaptic weights, and the update formula is: , Among them, is the synaptic weight matrix of the i-th event cluster after update, is the synaptic weight matrix of the i-th event cluster before update, is the gradient of the objective function with respect to the synaptic weight matrix W i , γ is the step size, and β is the momentum coefficient is the momentum term of the previous weight update.
6. The method for generating airport flight video slices according to claim 1, wherein The slicing of the real-time video data according to the matched reference slice parameters includes: Read the predefined time window length T, resolution threshold R, and region of interest ROI of the slice from the reference slice parameter library of the matched event cluster; Generate a slice instruction according to the time window length T, resolution threshold R, and region of interest ROI of the slice; Slice the video stream according to the slice instruction to generate segments matching the event types; Monitor the stability of the slice content through the spike frequency of the SNN. If an anomaly is detected, trigger the parameter re-evaluation mechanism.
7. The method for generating airport flight video slices according to claim 1, characterized in that, The associating of the flight information with the video source addresses and the calling of the real-time video data includes: Create a data structure to establish and store the mapping relationship between flight information and the corresponding video data source address; Set a timer to call the flight information API every preset time to obtain the latest flight data, and parse the flight data returned by the flight information API to update the flight information status and the mapping relationship in the data structure in real time; According to the real-time flight information, dynamically match and switch to the corresponding video data source from the updated mapping relationship of the data structure, and obtain real-time video stream data.
8. The method for generating airport flight video slices according to claim 7, wherein, The step of, according to the real-time flight information, dynamically matching and switching to the corresponding video data source from the updated mapping relationship of the data structure, and obtaining real-time video stream data, includes: When the flight information is updated, immediately search for the corresponding new video data source in the mapping relationship of the data structure, switch to the new video data source, and obtain real-time video stream data from the new video data source.
9. An airport flight video slice generation device for implementing the airport flight video slice generation method described in any one of claims 1 to 8, characterized in that, The device includes: The first data module: used to obtain the current flight information, including flight number, estimated departure time, estimated arrival time, boarding gate, and baggage carousel number; The second data module: used to obtain the video source addresses of key areas, where the key areas include the flight area, boarding gate, baggage carousel, and waiting area; The association module: used to associate flight information with video source addresses and call real-time video data; The feature extraction module: used to extract features from the real-time video data of the flight to obtain the target feature vector; The classification module: used to build a clustering model, add the current target feature vector to the clustering model, and build an event recognition hypersurface through a spiking neural network to optimize the dynamic partitioning of the clustering model for real-time video data and the matching of benchmark slice parameters; The slicing module: used to slice the real-time video data according to the matched benchmark slice parameters to generate video slices; Among them, the classification module is also used for: Collect the historical video stream of the airport scene and associated flight status data to generate historical target feature vectors; Convert the historical target feature vectors into pulse sequences, input them into the spiking neural network, and initialize the clustering centers through the self-organizing mapping algorithm; Optimize the clustering centers based on the pulse gradient descent algorithm, define the hypersurface parameters of the event clusters, and minimize the pulse distance variance of the samples within the clusters; According to the optimized hypersurface parameters, calculate and store the benchmark slice parameters of each event cluster, including the time window length, resolution threshold, and coordinates of the region of interest.
Citation Information
Patent Citations
Target monitoring method for airport field monitoring system on basis of video recognition
CN104243935A
Airport arrival and departure flight information accurate display system for serving passengers
CN119323905A