Efficient airport scene real-time comprehensive risk prevention and control method and system

By applying deep learning models and dynamic density peak clustering technology in airport scenes, real-time risk prevention and control of airport scenes is achieved, the shortcomings in efficiency and response speed of traditional monitoring methods are solved, and the overall safety level is improved.

CN120071241APending Publication Date: 2025-05-30CHENGDU CIVIL AVIATION AIR TRAFFIC CONTROL SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510063205.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The traditional comprehensive monitoring method for airport scenes is not sufficient in efficiency, response speed and data analysis capabilities, and there are obvious shortcomings, so it is impossible to effectively manage airport scenes in real time.

Method used

Deep learning model, especially the YOLO model, is adopted to conduct multi-objective real-time and efficient detection, and combine dynamic density peak clustering technology and ReID processing to achieve real-time risk prevention and control of airport scenes.

Benefits of technology

It improves detection accuracy and speed, achieves efficient and real-time risk management of airport scenes, enhances the overall safety level, and can promptly detect and deal with potential safety hazards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071241A_ABST
    Figure CN120071241A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of airport risk prevention and control, and particularly relates to an efficient airport scene real-time comprehensive risk prevention and control method and system, and the method comprises the steps: transmitting a to-be-detected image of an airport scene to a preset deep learning model, and carrying out the multi-target real-time efficient detection, so as to obtain the types, unique IDs and physical coordinates of different objects; performing risk prevention and control according to the category, the unique ID and the physical coordinate in combination with a set judgment rule; the judgment rule comprises class-based unauthorized animal intrusion judgment, unique ID-based safe distance judgment and high-risk area judgment based on physical coordinates and combined with a dynamic density peak value clustering technology; when the risk exists in the judgment result, an alarm notice is quickly sent to specified safety management personnel, and associated processing measures are implemented; the beneficial effects of the invention are that the system can achieve the efficient and real-time risk management of the airport scene, and improves the overall safety level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of airport risk prevention and control, and particularly to an efficient real-time comprehensive risk prevention and control method and system for airport surface Background Art

[0002] With the rapid development of the air transportation industry, as an important transportation hub, the comprehensive risk management of the airport surface has become particularly important. The airport surface refers to various activities and scenarios occurring within the airport, including the processes of aircraft takeoff and landing, the facilities and personnel activities within the airport, etc.

[0003] The safety of the airport surface area is directly related to the safety of flight takeoff and landing and the life and property safety of ground personnel. Therefore, how to effectively prevent and control risks and manage airport surface hazards, such as preventing unauthorized animals (animals other than patrol police dogs, such as cats, birds, or other animals, etc.) from entering, preventing high-risk areas, preventing safety distances, etc., and timely detecting and handling potential safety hazards has become a key issue in airport safety management.

[0004] However, the traditional comprehensive monitoring of the airport surface mainly relies on security and manual patrols and video monitoring by fixed cameras. This method has obvious deficiencies in terms of efficiency, response speed, data analysis ability, and the balance of accuracy and real-time performance. Therefore, it is urgent to introduce more intelligent and efficient technical solutions to improve the overall safety level. Summary of the Invention

[0005] Aiming at the technical defects mentioned in the background art, the purpose of the embodiments of the present invention is to provide an efficient real-time comprehensive risk prevention and control method and system for airport surface, so as to achieve efficient and real-time risk management of the airport surface and improve the overall safety level.

[0006] To achieve the above purpose, in the first aspect, the embodiments of the present invention provide an efficient real-time comprehensive risk prevention and control method for airport surface, and the method includes:

[0007] Transmitting the to-be-detected image of the airport surface to a preset deep learning model for multi-object real-time and efficient detection to obtain the categories, unique IDs, and physical coordinates of different objects;

[0008] Performing risk prevention and control according to the categories, unique IDs, and physical coordinates in combination with the set determination rules; wherein, the determination rules include unauthorized animal intrusion judgment based on categories, safety distance judgment based on unique IDs, and high-risk area judgment based on physical coordinates in combination with the dynamic density peak clustering technology;

[0009] When a risk exists in the judgment result, an alarm notification will be quickly sent to the designated safety management personnel, and the associated treatment measures will be implemented.

[0010] As a specific preferred embodiment of the present application, the multi-target real-time and efficient detection includes the following steps:

[0011] Preprocess the image to be detected;

[0012] Transmit the preprocessed image to a preset deep learning model to complete multi-target real-time detection, and output the category and pixel coordinates of each detection box; wherein, the deep learning model uses a pre-trained YOLO model;

[0013] Post-process the YOLO model to perform bounding box filtering and bounding box decoding;

[0014] Finally, perform coordinate transformation and ReID processing to obtain the category, unique ID, and physical coordinates.

[0015] As a specific implementation manner of the present application, when the YOLO model is trained, it includes the following data augmentation training operations:

[0016] Random erasing, rotation, flipping, histogram equalization, Gamma correction, and Gaussian blur; to enhance the model's adaptability to complex environments.

[0017] As a specific implementation manner of the present application, the coordinate transformation specifically includes converting the image pixel coordinates into actual physical coordinates by using four-point perspective transformation.

[0018] As a specific implementation manner of the present application, the ReID processing specifically includes:

[0019] Based on the pixel coordinates output by the YOLO model, extract each object from the image to be detected as an object to be encoded;

[0020] After preprocessing the object to be encoded, encode it with a ReID model and convert it into a 1x1024-dimensional encoding vector; wherein, the encoding vector after ReID encoding is the appearance feature of the object;

[0021] Combine the encoding vector and physical coordinates to measure similarity, and perform ID assignment based on the calculation result.

[0022] As a specific implementation manner of the present application, the safety distance judgment includes the judgment of the distance between static objects and the safety distance judgment of moving objects;

[0023] The determination of the distance between stationary objects includes calculating the distance between two objects using the Euclidean distance formula based on the coordinates of the detected object center points or bounding boxes, and comparing it with a preset safety distance threshold; if the calculated distance is less than the threshold, it is considered that the distance between the two objects is too close and there is a safety hazard;

[0024] The determination of the safety distance for moving objects includes obtaining the displacement change of the same object in two consecutive frames through ReID to obtain the corresponding speed, and analyzing the dynamic changes between moving objects for judgment; among them, the greater the speed in this judgment, the smaller the safety distance.

[0025] As a specific implementation manner of the present application, the determination of high-risk areas is completed through density peak clustering technology based on a dynamic threshold, and the specific steps include:

[0026] Calculate the distance matrix;

[0027] Determine the cut-off distance for calculating the local density;

[0028] Calculate the local density;

[0029] Calculate the relative distance of each point to a point with a higher density;

[0030] By drawing a decision graph, where the horizontal axis represents the local density value and the vertical axis represents the relative distance value, select points with high local density and relative distance values as clustering centers;

[0031] Assign other points to the nearest clustering center;

[0032] According to the changes in the density of people and objects in different scenarios, dynamically adjust the threshold for aggregation judgment; among them, the dynamic threshold adjustment model is as follows:

[0033]

[0034] d c1 is the cut-off distance after dynamic adjustment, d c0 is the initial cut-off distance, k is the adjustment coefficient, ρ m is the current highest density;

[0035] Then, based on the clustering results and the dynamically adjusted threshold, determine whether it is a high-risk area.

[0036] In a second aspect, an embodiment of the present invention further provides an efficient real-time comprehensive risk prevention and control system for airport surface, and the system includes:

[0037] An efficient detection module for transmitting the to-be-detected image of the airport surface to a preset deep learning model for multi-object real-time and efficient detection to obtain the categories, unique IDs, and physical coordinates of different objects;

[0038] A risk prevention and control module for performing risk prevention and control according to the category, unique ID, and physical coordinates in combination with set determination rules; wherein, the determination rules include unauthorized animal intrusion judgment based on the category, safety distance judgment based on the unique ID, and high-risk area judgment based on the physical coordinates in combination with the dynamic density peak clustering technology;

[0039] A processing module for, when a risk exists in the judgment result, quickly sending an alarm notification to designated security management personnel and implementing associated processing measures.

[0040] The technical solution provided by the embodiments of the present invention has the following advantages:

[0041] 1. Improve detection accuracy and speed: The present invention adopts the advanced YOLO model. By utilizing its significant advantages in detection accuracy and speed, it can effectively balance the requirements of detection accuracy and real-time performance; The YOLO model demonstrates excellent accuracy-latency balance ability under different model scales, and is particularly suitable for the efficient and real-time requirements of the present invention;

[0042] 2. The present invention collects diverse airport scene data containing these factors for retraining the YOLO model to enhance the model's adaptability to complex environments; In addition, the application of data augmentation technology further improves the performance of the model in actual deployment;

[0043] 3. This solution improves accuracy by converting the output coordinate points to be close to the actual physical distance ratio rather than simply the pixel distance ratio, especially in the detection of individuals at different distances and angles;

[0044] 4. Dynamically adjust the aggregation judgment threshold: In view of the characteristics that the density of people and objects on the airport surface changes over time, the present invention realizes the function of dynamically adjusting the high-risk area judgment threshold according to the current scene characteristics; This enables the algorithm to flexibly respond to the specific requirements at different times and locations, improving the adaptability and practicality of the system;

[0045] 5. The present invention innovatively combines the high-speed detection ability of the YOLO model and the accurate positioning ability of people and objects by dynamic DPC clustering, not only solving the coordinate error problem existing in traditional technologies, but also significantly improving the overall performance of high-risk area detection; This integrated method greatly enhances the accuracy and robustness of high-risk area detection on the airport surface, providing strong technical support for the safety management of the airport;

[0046] 6. Multiple real-time risk prevention and control: The present invention simultaneously realizes real-time unauthorized animal intrusion, high-risk area monitoring, and safety distance monitoring; When potential risks are detected, the system can issue warnings or take preventive measures in a timely manner, effectively avoiding the occurrence of accidents;

[0047] 7. Safety distance prevention and control for balancing speed and position: In addition to the safety distance prevention and control between static objects, the present invention also introduces ReID to monitor the speed between moving objects, achieving the prevention and control of speed and distance balance; ultimately realizing the efficient and real-time risk management of the airport scene and ensuring the safe operation of the airport. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art.

[0049] Figure 1 is a flowchart of an efficient real-time comprehensive risk prevention and control method for airport scenes provided by an embodiment of the present invention;

[0050] Figure 2 is a schematic diagram of the process of multi-target real-time and efficient detection provided by an embodiment of the present invention;

[0051] Figure 3 is a schematic diagram of visual deviation provided by an embodiment of the present invention;

[0052] Figure 4 is a schematic diagram of the process of ReID processing provided by an embodiment of the present invention;

[0053] Figure 5 is a schematic diagram of the processing of risk prevention and control provided by an embodiment of the present invention;

[0054] Figure 6 is a schematic block diagram of the principle of an efficient real-time comprehensive risk prevention and control system for airport scenes provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0056] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0057] Throughout the specification, references to "an embodiment", "embodiments", "an example", or "examples" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present invention. Thus, the phrases "in an embodiment", "in embodiments", "an example", or "examples" that appear throughout the specification do not necessarily all refer to the same embodiment or example. Furthermore, the particular features, structures, or characteristics may be combined in any suitable combination and / or sub-combination in one or more embodiments or examples.

[0058] It should be noted that, unless otherwise specified, the technical terms in this embodiment have the ordinary meanings understood in the technical field to which they belong.

[0059] DPC: Densitypeak clustering, a dynamic density peak clustering technique, is an efficient density-based clustering algorithm. It does not require specifying the number of clusters, can automatically discover the clustering centers, is not affected by outliers, and is particularly suitable for non-clustered data, which conforms to the data characteristics of the present invention.

[0060] ReID: The full name is Re-identification, that is, re-identification, which is a model for feature extraction of input data.

[0061] YOLO model: You Only Look Once is a very efficient object detection model.

[0062] Please refer to Figure 1 and Figure 5 , an efficient real-time comprehensive risk prevention and control method for airport surface provided by an embodiment of the present invention, the method includes:

[0063] S101, transmitting the image to be detected of the airport surface to a preset deep learning model for multi-target real-time and efficient detection to obtain the categories, unique IDs, and physical coordinates of different objects;

[0064] S102, performing risk prevention and control according to the categories, unique IDs, and physical coordinates and in combination with set determination rules; wherein, the determination rules include unauthorized animal intrusion judgment based on categories, safety distance judgment based on unique IDs, and high-risk area judgment based on physical coordinates and in combination with the dynamic density peak clustering technique;

[0065] S103, when a risk exists in the judgment result, quickly send an alarm notification to the designated safety management personnel and implement the associated processing measures.

[0066] In this embodiment, referring to Figure 2 , the multi-target real-time and efficient detection includes the following steps:

[0067] Preprocess the image to be detected;

[0068] Transmit the preprocessed image to a preset deep learning model to complete multi-object real-time detection, and output the category and pixel coordinates of each detection box; among them, the deep learning model uses the pre-trained YOLO model;

[0069] Post-process the YOLO model to perform bounding box filtering and bounding box decoding;

[0070] Finally, perform coordinate conversion and ReID processing to obtain the category, unique ID, and physical coordinates.

[0071] Specifically, the image to be detected needs to be preprocessed first to convert the input into a format that can be detected by YOLO, including the following two steps:

[0072] 1) Resize the image to 640x640 pixels, which is actually a scaling operation.

[0073] Assume the original image size is (W×H), and the adjusted image size is (640×640), then the scaling ratio of width and height is

[0074]

[0075] 2) Normalization: Normalize the pixel values to the range [0, 1].

[0076]

[0077] In the above formula, x is the pixel value, min(x) is the minimum pixel value of the current image, max(x) is the maximum pixel value of the current image, and y is the pixel value after normalization, with a value range of [0, 1].

[0078] It should be noted that if operations involving cross-shot fusion of images are involved, different algorithms can be used for self-fusion to ensure that images from different perspectives can be accurately aligned and integrated. The present invention will not elaborate too much on this.

[0079] In this embodiment, although YOLO has been updated to version YOLOv11 at present, the present invention still uses YOLOv10 instead of YOLOv11 due to the following experimental results:

[0080] 1). Precision comparison: In the actual experiment of the present invention, it is found that the detection precisions of YOLOv10 and YOLOv11 both meet the expected requirements of this experiment and meet the requirements.

[0081] 2). Inference speed comparison:

[0082] a. Pretreatment: The pretreatment steps of YOLOv10 and YOLOv11 are the same, and the time consumption here is the same.

[0083] b. Inference: The model inference time of the two models is positively correlated with the type and quantity of model parameters. For models of the same order of magnitude, the inference speeds of the two are quite comparable, with a very small gap.

[0084] c. Postprocessing:

[0085] The output dimension of YOLOv10 is (300, 6), that is, the total number of prediction boxes is 300, and the dimension of each prediction box is 6, representing the upper-left and lower-right coordinates (left, top, right, bottom) of the detection box, the class ID (classId), and the class confidence (p) (2 + 1 + 1 + 1 + 1 = 6). The postprocessing steps are bounding box filtering and bounding box decoding.

[0086] The output dimension of YOLOv11 is (8400, 84), that is, the total number of prediction boxes is 8400, and the dimension of each prediction box is 84, representing the center coordinates (cx, cy), width (w), height (h) of each detection box, and the confidence of 80 classes (p) (2 + 1 + 1 + 80 = 84). The postprocessing steps of YOLOv11 include bounding box filtering, non-maximum suppression, and bounding box decoding.

[0087] In contrast, the output dimension of YOLOv10 is smaller, the postprocessing steps are one less, the inference speed is faster, and it is more friendly to the requirement of real-time performance. Therefore, the present invention selects YOLOv10 to complete the function of object detection, and outputs the class and coordinates of each detection box;

[0088] The postprocessing of the YOLO model is the postprocessing of YOLOv10, that is, the bounding box filtering and bounding box decoding mentioned above.

[0089] 1). Bounding box filtering: If the confidence of a certain class is lower than the set threshold, then discard the prediction box.

[0090] 2). Bounding box decoding:

[0091] This step involves mapping the position of the prediction box from the output space of the model back to the space of the original image. This usually requires using an inverse affine transformation matrix IM to adjust the position of the bounding box. Here, IM is a matrix containing scaling and translation information. Assuming the size of the original image is (W×H), the upper-left coordinates (left', top') and lower-right coordinates (right', bottom') after decoding are as follows:

[0092]

[0093] (left', top') = (left, top) * IM

[0094] (right', bottom') = (right, bottom) * IM

[0095] Furthermore, to improve the detection effect of YOLOv10, the present invention retrains YOLOv10 based on a self-built airport scene dataset. This self-built dataset introduces diverse scene samples in the real environment of the airport scene, including different occlusion situations, lighting changes, perspective transformations, etc. In addition, for specific airport environments, such as areas like runways and aprons, scene samples related to people, aircraft, and unauthorized animals are particularly added. This is used to enrich the training dataset of the YOLOv10 model. To a certain extent, the higher the quality of the training dataset, the better the detection effect in the actual scene.

[0096] Meanwhile, when the YOLO model is being trained, it includes the following data augmentation training operations:

[0097] Random erasing, rotation, flipping, histogram equalization, Gamma correction, and Gaussian blur.

[0098] Specifically, the present invention introduces some data augmentation methods to enhance the self-built dataset. The dataset augmentation not only enhances the model's adaptability to complex environments but also improves the detection accuracy:

[0099] 1). Random erasing: Erase a part of the image randomly: Randomly select the upper-left coordinate (x1, y1) and the lower-right coordinate (x2, y2) in the image, and set the pixel values in this area to 0. The mathematical expression is as follows:

[0100] set I(x, y) = 0, y1 ≤ y ≤ y2;

[0101] 2). Rotation: Rotate the image randomly by a certain angle θ. Assume the coordinates of the image are ((x, y)), and the rotated coordinates ((x’, y’)) are calculated by the following formula:

[0102]

[0103] 3). Flipping: Randomly flip the image horizontally or vertically.

[0104] When flipping horizontally, the pixel coordinates ((x, y)) become ((W - x, y)); when flipping vertically, the pixel coordinates ((x, y)) become ((x, H - y)). Where W and H are the width and height of the data.

[0105] 4). Histogram equalization

[0106] Histogram equalization enhances contrast by adjusting the gray value distribution of the image. The formula is as follows:

[0107]

[0108] where I(x, y) is the pixel value of the original image, I’(x, y) is the enhanced pixel value, and L is the number of gray levels.

[0109] 5). Gamma correction: used to adjust the brightness of the image. The formula is as follows:

[0110] I’(x, y) = I(x, y) r ; where r is the Gamma value, usually between 0.5 and 2.5.

[0111] 6). Gaussian blur: smooths the image through convolution operations to reduce noise. The formula for the Gaussian kernel is as follows:

[0112]

[0113] where δ is the standard deviation, which determines the degree of blur.

[0114] In this embodiment, the coordinate transformation specifically includes converting the image pixel coordinates into actual physical coordinates by using four-point perspective transformation.

[0115] Specifically, for the output (category and pixel coordinates) of the multi-object detection in the previous step, due to the different installation angles and positions of the camera, there are significant differences between the pixel coordinates in the image and the actual physical coordinates. For example, as Figure 3 shown, the pixel coordinates of Person 1 and Person 2 in the image are far apart, while in reality, the distance between them may be very close; on the contrary, Person 3 and Person 4 appear very close in the image, but in fact, the distance between them may be large.

[0116] To effectively correct this visual deviation and make the pixel coordinates as accurately reflect the actual distance between objects as possible.

[0117] The present invention uses four-point perspective transformation to convert the image pixel coordinates into actual coordinates. The calculation process of the perspective transformation matrix is as follows:

[0118] (1). Select four origin points (x i , y i ) on the original plane, i = 1, 2, 3, 4.

[0119] Select four target points (x' i , y' i ) corresponding to the four origin points on the target plane, i = 1, 2, 3, 4.

[0120]

[0121] (2). Establish equations:

[0122] For each pair of origin points and corresponding target points, the following system of linear equations can be established:

[0123] x′ i (h 31 x i +h 32 y i +h 33 )-(h 11 x i +h 12 y i +h 13 ) = 0

[0124] y′ i (h 31 x i +h 32 y i +h 33 )-(h 21 x i +h 22 y i +h 23 ) = 0

[0125] (3). Solve the system of equations:

[0126] Combine the equations of all points to form a matrix equation:

[0127] Ah = b

[0128] Among them, the h matrix is the perspective transformation matrix, and the A, h, and b matrices are as follows:

[0129]

[0130] h = [h 11 , h 12 , h 13 , h 21 , h 22 , h 23 , h 31 , h 32 , h 33 T

[0131] b = [0, 0, 0, 0, 0, 0, 0, 0] T

[0132] ​By solving the above linear equations, the perspective transformation matrix h can be obtained. With the perspective transformation matrix h, the fitting relationship of any coordinate pair between two planes can be obtained. The mapping relationship from the original plane coordinates (x, y) to the three-dimensional plane coordinates (X, Y, Z) is as follows:

[0133]

[0134] Then, the mapping relationship from the original plane pixel coordinates (x, y) to the three-dimensional plane coordinates (X, Y, Z) satisfies the following formula:

[0135] X = h 11 *x + h 12 *y + h 13

[0136] Y = h 21 *x + h 22 *y + h 23

[0137] Z = h 31 *x + h 32 *y + h 13

[0138] The mapping relationship from the three-dimensional plane coordinates (X, Y, Z) to the coordinates (x', y') on the target plane satisfies the following formula:

[0139]

[0140] Thus, the target plane coordinates (x', y') can be fitted from the original plane coordinates (x, y);

[0141] After perspective transformation, the relative coordinates obtained can be used to obtain the actual physical coordinates through the following formula, as follows:

[0142]

[0143] Among them, A and B represent that the boundary lengths and widths are A meters and B meters, the relative coordinates after coordinate fitting are (x', y'), the true coordinates are (x, y), and the pixel sizes are (W, H);

[0144] From the above, the calibrated physical coordinates are obtained.

[0145] Furthermore, in order to monitor the position changes and speeds of each detected object, the present invention also needs to continuously identify the positions of the same object in different frames, that is, the re-identification algorithm. The present invention uses ReID to implement the re-identification function, and the specific ReID processing includes:

[0146] Based on the pixel coordinates output by the YOLO model, each object is extracted from the image to be detected as the object to be encoded;

[0147] After preprocessing the encoding object, it is encoded using the ReID model and converted into a 1x1024-dimensional encoding vector; among them, the encoding vector after ReID encoding is the appearance feature of the object.

[0148] The similarity is measured by combining the encoding vector and the physical coordinates, and ID assignment is performed based on the calculation results.

[0149] Specifically, referring to Figure 4 , the preprocessing method of ReID is as follows:

[0150] 1) Resize the image to 256x128 pixels, which is actually a scaling operation:

[0151] Assume that the size of the original image is (W×H), and the size of the resized image is (256×128), then the scaling ratios of the width and height are

[0152]

[0153] 2). Normalization: Normalize the pixel values to the range [0,1].

[0154]

[0155] In the above formula, x is the pixel value, min(x) is the minimum pixel value of the current image, max(x) is the maximum pixel value of the current image, and y is the pixel value after normalization, with a value range of [0,1].

[0156] 3). Standardization:

[0157]

[0158]

[0159] The significance of the feature vector after ReID encoding is the appearance feature. For different frames of the same object, the appearance is similar, and this feature vector after ReID encoding should be similar.

[0160] When performing re-identification, the present invention takes the following two points into consideration:

[0161] 1). The result of the encoding vector of Reid is equivalent to the appearance feature of the object. The more similar two vectors a and b are, the larger the value of their cosine similarity cos(a,b), and the closer it is to 1. (The range of cosine similarity is from -1 to 1, and it measures the similarity in direction by calculating the cosine value of the angle between two non-zero vectors.)

[0162] 2). The physical coordinates are equivalent to the position characteristics of the object. The time interval between video frames is very small (usually 30 FPS for video frames, and with such a small time gap, the displacement cannot be large). To a certain extent, the smaller the normalized distance, the greater the possibility that the two objects are the same object.

[0163] If the similarity is measured by combining the encoding vector and the physical coordinates, better results will be obtained: The present invention proposes to calculate the similarity by combining the cosine similarity cos(a, b) and the normalized distance d of the physical coordinates to re-identify each object:

[0164] L = 0.2 * (1 - d) + 0.8 * cos(a, b)

[0165]

[0166] Among them, the physical coordinates and the encoding vector of object a are (x a , y a ), a i ; the physical coordinates and the encoding vector of object b are (x b , y b ), b i ; the width and height of the site are W and H.

[0167] Next is to assign IDs: When the similarity between the newly observed object and the object corresponding to the existing ID is higher than a certain preset threshold, it is considered the same object and the same and unique ID is assigned. If the similarity is much lower than the threshold, it is considered a new object and a new ID is assigned. For the object that has been assigned an ID, if it cannot be recognized again within a continuous number of frames (for example, 5 frames), it can be considered that the pedestrian has left the current monitoring area, and the corresponding ID can be marked as "lost" status.

[0168] In this embodiment, unauthorized animal intrusion: Through the output class ID of YOLOv10, it can be determined whether the detected object is an unauthorized animal.

[0169] Safety distance: The physical coordinate conversion has been completed in the foregoing steps. For the scenarios in the present invention, it can be divided into the following two situations, that is, the safety distance judgment includes the distance judgment between static objects and the safety distance judgment of moving objects.

[0170] The distance judgment between static objects includes calculating the distance between two objects using the Euclidean distance formula according to the coordinates of the center point or the bounding box of the detected object, and comparing it with the preset safety distance threshold; if the calculated distance is less than the threshold, it is considered that the distance between the two objects is too close and there is a safety hazard;

[0171] That is, the distance between static objects (static people, people, static airplanes, airplanes, static people, airplanes): In this case, only according to the coordinates of the detected object center point or bounding box, use the Euclidean distance formula to calculate the distance between pairwise objects, and compare it with the preset safety distance threshold. If the calculated distance is less than the threshold, it is considered that the distance between these two objects is too close and there may be potential safety hazards.

[0172] The determination of the safety distance of the moving object includes obtaining the displacement change of the same object in the previous and subsequent frames through ReID, obtaining the corresponding speed, and analyzing the dynamic changes between moving objects for judgment; among them, the greater the speed in this judgment, the smaller the safety distance.

[0173] In this embodiment, the safety distance of the moving object: For an object in a moving state, not only the safety distance between object bodies needs to be considered, but also the speed and direction of other objects need to be considered. Between different moving objects, the requirements for speed and safety distance are different. For example, the speed and safety distance between moving people and people are much smaller than the safety distance between moving people and airplanes. In this case, it is necessary to calculate the speed according to the ID of each object (that is, the output of the re-identification algorithm), and then make a comprehensive judgment; since when determining the safety distance between moving objects, the speed difference, reaction time, and braking performance are mainly considered. For objects with extremely different speeds, such as pedestrians and airplanes, a larger safety distance is required to ensure safety, while a smaller safety distance is required between people; at the same time, environmental conditions and regulatory restrictions will also affect the specific safety distance requirements.

[0174] By real-time monitoring and analyzing these factors, the appropriate safety distance can be dynamically adjusted and maintained.

[0175] Through ReID, the displacement change of the same object in the previous and subsequent frames can be obtained, and thus the speed can be obtained. Then, the dynamic changes between moving objects can be effectively monitored and analyzed, providing reliable technical support for the real-time monitoring of the safety distance. The safety distance risk prevention and control rule is that the greater the speed, the smaller the safety distance. Set a threshold (not set a specific threshold according to the specific situation) to balance the speed and distance of different objects to achieve the detection of the safety distance between moving objects.

[0176] Furthermore, in places where there are more objects such as people, airplanes, and animals, unexpected situations are more likely to occur, which is a high-risk area and requires key prevention and control. In order to detect this high-risk area, the present invention is completed through density peak clustering based on a dynamic threshold (DPC clustering).

[0177] The determination of the high-risk area is completed through density peak clustering technology based on a dynamic threshold. The specific steps include:

[0178] Calculate the distance matrix; that is, calculate the Euclidean distance between the physical coordinate points of each person or object in the airport surface area.

[0179]

[0180] d ij represents the Euclidean distance, where (x i , y i ) and (x j , y j ) represent the physical coordinates of objects i and j respectively.

[0181] Determine the cut-off distance for calculating the local density; where the cut-off distance is represented by d c , and the local density is represented by ρ i .

[0182] Calculate the local density; for each point, calculate its local density ρ i , that is, the number of points within the distance d c .

[0183]

[0184] ε(*) represents a step function:

[0185]

[0186] Calculate the relative distance of each point from the point with higher density; that is, calculate the relative distance δ i of each point from the point with higher density: for point i, calculate its minimum distance δ i from the point with higher density, that is, satisfy the following formula:

[0187]

[0188] If point i is the point with the highest density, then define:

[0189]

[0190] By plotting a decision graph, where the horizontal axis represents the local density value and the vertical axis represents the relative distance value, select the points with high local density and relative distance values as the clustering centers; that is, the horizontal axis represents the ρ i value,

[0191] the vertical axis represents the δ i value, and select the points with high ρ i and δ i values as the clustering centers.

[0192] Assign other points to the nearest clustering center;

[0193] Dynamically adjust the threshold for aggregation judgment according to the changes in the density of people and objects in different scenarios. The dynamic threshold adjustment model is as follows:

[0194]

[0195] d c1 is the truncated distance after dynamic adjustment, d c0 is the initial truncated distance, k is the adjustment coefficient, ρ m is the current highest density;

[0196] The dynamic threshold adjustment is to better adapt to the clustering requirements in different environments. For example, in scenarios where the density of people and objects changes greatly, a fixed truncated distance may not be suitable for all clustering tasks. The present invention provides a method for dynamically adjusting the density distribution of people and objects based on the current scenario: the truncated distance can be adjusted according to the current density, achieving the effect that the greater the current maximum density, the smaller the dynamic adjustment of the truncated distance.

[0197] Then, based on the clustering results and the dynamically adjusted threshold, determine whether it is a high-risk area.

[0198] That is, evaluate the currently detected situation. If it is found that the density of people and objects in any area exceeds the standard, it is marked as "high risk". This method can not only ensure the consistent performance of the algorithm in different environments and time periods, but also automatically adapt to the changes in the airport passenger flow, effectively preventing various safety problems that may be caused in high-risk areas.

[0199] In this embodiment, for animals that enter the monitoring area without permission, the system has an instant detection and response function. When such a situation is detected, the system will quickly send an alarm notification to the designated security management personnel, and can selectively play a repelling sound on the spot according to the actual situation, which not only effectively avoids damage to the facilities caused by animals, but also ensures the safety of personnel. In addition, this mechanism can also promote the staff to quickly reach the scene and implement further treatment measures.

[0200] To ensure the safety of all activity participants and equipment, the present invention continuously monitors the relative positions and movement speeds of various objects. Once it is found that someone or something approaches a critical point that may cause potential safety hazards, an emergency communication link is immediately activated to send a warning message to the relevant responsible persons. At the same time, by activating the on-site audio equipment to emit a clear alarm sound, this is to remind the people on the scene to pay attention to the changes in the surrounding environment and adjust their own behaviors in time to avoid risks.

[0201] Once the density of people and objects is detected to be higher than the preset threshold, the area will be automatically marked as "high risk" by the system. The system will remind the security officer to pay key attention to and prevent it through means such as alarm bells and text messages to ensure timely response to potential threats. Conduct a detailed risk assessment of the monitored area according to the set safety standards. (For example: the density of people and objects per square meter in the area should not exceed 5 people or objects; there should be a safety distance of at least 1 meter between the stationary aircraft and other objects or personnel; there should be a distance of at least 10 meters between the moving people and the aircraft; and there should be a distance of at least 20 meters between two moving aircraft; for stationary aircraft, the distance between them should not be less than the specified minimum value.)

[0202] The above solution has the following advantages:

[0203] 1. Improve detection accuracy and speed: The present invention adopts an advanced YOLO model. By utilizing its significant advantages in detection accuracy and speed, it can effectively balance the requirements of detection accuracy and real-time performance; the YOLO model demonstrates excellent accuracy-latency balance capabilities under different model scales, which is particularly suitable for the efficient and real-time requirements of the present invention;

[0204] 2. The present invention collects diverse airport scene data containing these factors for retraining the YOLO model to enhance the model's adaptability to complex environments; in addition, the application of data augmentation technology further improves the performance of the model in actual deployment;

[0205] 3. The solution improves accuracy by converting the output coordinate points to be close to the actual physical distance ratio rather than simply the pixel distance ratio, especially in the detection of individuals at different distances and angles;

[0206] 4. Dynamically adjust the aggregation judgment threshold: In view of the characteristics that the density of people and objects on the airport surface changes over time, the present invention realizes the function of dynamically adjusting the judgment threshold of high-risk areas according to the current scene characteristics; this enables the algorithm to flexibly meet the specific requirements at different times and locations, improving the adaptability and practicality of the system;

[0207] 5. The present invention innovatively integrates the high-speed detection ability of the YOLO model and the precise positioning ability of people and objects by dynamic DPC clustering. It not only solves the coordinate error problem existing in traditional technologies but also significantly improves the overall performance of high-risk area detection; this integrated method greatly enhances the accuracy and robustness of high-risk area detection on the airport surface, providing strong technical support for the safety management of the airport;

[0208] 6. Multiple real-time risk prevention and control: The present invention simultaneously realizes real-time unauthorized animal intrusion, high-risk area monitoring, and safety distance monitoring; when potential risks are detected, the system can promptly issue warnings or take preventive measures to effectively avoid accidents.

[0209] 7. Safety distance prevention and control that balances speed and position: In addition to the safety distance prevention and control between static objects, the present invention also introduces ReID to monitor the speed between moving objects, achieving balanced prevention and control of speed and distance; ultimately realizing efficient and real-time risk management of the airport scene and ensuring the safe operation of the airport.

[0210] Refer to Figure 6 , based on the same inventive concept, the embodiment of the present invention also provides an efficient real-time integrated risk prevention and control system for the airport scene, and the system includes:

[0211] An efficient detection module for transmitting the image to be detected in the airport scene to a preset deep learning model for multi-object real-time and efficient detection to obtain the category, unique ID, and physical coordinates of different objects.

[0212] A risk prevention and control module for performing risk prevention and control according to the category, unique ID, and physical coordinates in combination with the set determination rules; wherein, the determination rules include unauthorized animal intrusion judgment based on the category, safety distance judgment based on the unique ID, and high-risk area judgment based on the physical coordinates in combination with the dynamic density peak clustering technology.

[0213] A processing module for promptly sending an alarm notification to the designated safety management personnel and implementing the associated processing measures when a risk exists in the judgment result.

[0214] Furthermore, the multi-object real-time and efficient detection includes the following steps:

[0215] Preprocess the image to be detected.

[0216] Transmit the preprocessed image to a preset deep learning model to complete multi-object real-time detection and output the category and pixel coordinates of each detection box; wherein, the deep learning model uses the pre-trained YOLO model.

[0217] Post-process the YOLO model for bounding box filtering and bounding box decoding.

[0218] Finally, perform coordinate conversion and ReID processing to obtain the category, unique ID, and physical coordinates.

[0219] Furthermore, when the YOLO model is trained, it includes the following data augmentation training operations:

[0220] Random erasing, rotation, flipping, histogram equalization, Gamma correction, and Gaussian blur.

[0221] In this embodiment, the coordinate transformation specifically includes converting the image pixel coordinates into actual physical coordinates by using four-point perspective transformation;

[0222] The ReID processing specifically includes:

[0223] Based on the pixel coordinates output by the YOLO model, each object is extracted from the image to be detected as an object to be encoded;

[0224] After preprocessing the object to be encoded, it is encoded with the ReID model and converted into a 1x1024-dimensional encoding vector; among them, the encoding vector after ReID encoding is the appearance feature of the object;

[0225] The similarity is measured by combining the encoding vector and the physical coordinates, and ID assignment is performed based on the calculation result;

[0226] The safety distance judgment includes the distance judgment between static objects and the safety distance judgment of moving objects;

[0227] The high-risk area judgment is completed by the density peak clustering technology based on a dynamic threshold. The specific steps include:

[0228] Calculate the distance matrix;

[0229] Determine the cut-off distance for calculating the local density;

[0230] Calculate the local density;

[0231] Calculate the relative distance of each point to the point with higher density;

[0232] By drawing a decision diagram, where the horizontal axis represents the local density value and the vertical axis represents the relative distance value, select the points with high local density and relative distance values as the clustering centers;

[0233] Assign other points to the nearest clustering center;

[0234] According to the changes in the density of people and objects in different scenarios, dynamically adjust the threshold for aggregation judgment; among them, the dynamic threshold adjustment model is as follows:

[0235]

[0236] d c1 is the cut-off distance after dynamic adjustment, d c0 is the initial cut-off distance, k is the adjustment coefficient, ρ m is the current highest density;

[0237] Then, based on the clustering results and the dynamically adjusted threshold, determine whether it is a high-risk area.

[0238] It should be noted that for a more specific description of the working process of the system embodiment, please refer to the foregoing method embodiment section and will not be elaborated herein.

[0239] The entire solution is based on deep learning and dynamic density peak clustering (DPC) technology, aiming to ensure the safety of the airport surface. Specifically, the present invention can achieve real-time and effective prevention and control of unauthorized animals (except patrol police dogs, such as other animals like cats, birds, etc.) and unauthorized person intrusions, while effectively monitoring high-risk areas and ensuring that the safety distance is maintained; and separately considering according to the object type and speed, detecting the speeds and safety distances of different objects, achieving prevention and control of the safety distances between different moving objects and different speeds; and having the ability of risk assessment and response.

[0240] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An efficient real-time comprehensive risk prevention and control method for airport scenes, characterized in that: The method comprises: The images to be detected at the airport scene are transmitted to the preset deep learning model for real-time and efficient multi-target detection to obtain the categories, unique IDs and physical coordinates of different objects; Risk prevention and control is performed according to the category, unique ID and physical coordinates in combination with the set judgment rules; wherein the judgment rules include unauthorized animal intrusion judgment based on category, safe distance judgment based on unique ID and high-risk area judgment based on physical coordinates combined with dynamic density peak clustering technology; When a risk is identified, an alert notification will be promptly sent to the designated security manager and the associated processing measures will be implemented.

2. An efficient real-time comprehensive risk prevention and control method for airport scenes as claimed in claim 1, characterized in that: The multi-target real-time and efficient detection comprises the following steps: Preprocess the image to be detected; The pre-processed image is transmitted to a preset deep learning model to complete multi-target real-time detection, and the category and pixel coordinates of each detection frame are output; wherein the deep learning model adopts a pre-trained YOLO model; Post-process the YOLO model to perform bounding box filtering and bounding box decoding; Finally, coordinate transformation and ReID processing are performed to obtain the category, unique ID and physical coordinates.

3. An efficient real-time comprehensive risk prevention and control method for airport scenes as claimed in claim 2, characterized in that: The YOLO model includes the following data enhancement training operations during training: Random erase, rotation, flip, histogram equalization, gamma correction and Gaussian blur.

4. An efficient real-time comprehensive risk prevention and control method for airport scenes as claimed in claim 3, characterized in that: The coordinate conversion specifically includes converting the image pixel coordinates into actual physical coordinates by using a four-point perspective transformation.

5. An efficient real-time comprehensive risk prevention and control method for airport scenes as claimed in claim 4, characterized in that: The ReID processing specifically includes: Based on the pixel coordinates output by the YOLO model, each object is extracted from the image to be detected as an object to be encoded; After preprocessing the object to be encoded, it is encoded using the ReID model and converted into a 1x1024-dimensional encoding vector; the encoding vector after ReID encoding is the appearance feature of the object; The encoding vector and the physical coordinates are combined to measure similarity, and an ID is assigned based on the calculation result.

6. An efficient real-time comprehensive risk prevention and control method for airport scenes as claimed in claim 5, characterized in that: The safety distance judgment includes the distance judgment between stationary objects and the safety distance judgment of objects in motion; The distance determination between the stationary objects includes calculating the distance between the two objects using the Euclidean distance formula according to the coordinates of the detected center point or bounding box of the object, and comparing it with a preset safety distance threshold; if the calculated distance is less than the threshold, it is considered that the distance between the two objects is too close and there is a safety hazard; The safe distance judgment of the moving object includes obtaining the displacement change of the same object in the previous and next frames through ReID, obtaining the corresponding speed, and analyzing the dynamic changes between the moving objects for judgment; wherein, in this judgment, the greater the speed, the smaller the safe distance.

7. An efficient real-time comprehensive risk prevention and control method for airport scenes as claimed in claim 6, characterized in that: The high-risk area determination is completed by using a density peak clustering technique based on a dynamic threshold, and the specific steps include: Calculate the distance matrix; Determine the cutoff distance for calculating local density; Calculate local density; Calculate the relative distance of each point to the points with higher density; By drawing a decision graph, the horizontal axis represents the local density value and the vertical axis represents the relative distance value, and the points with high local density and relative distance value are selected as cluster centers; Assign other points to the nearest cluster center; According to the changes in different scenes and the density of people and objects, the threshold of aggregation judgment is dynamically adjusted; the dynamic threshold adjustment model is as follows: d c1 is the dynamically adjusted cutoff distance, d c0 is the initial cutoff distance, k is the adjustment coefficient, ρ m is the current highest density; Then, based on the clustering results and the dynamically adjusted threshold, determine whether it is a high-risk area.

8. An efficient real-time comprehensive risk prevention and control system for airport scenes, characterized by: The system comprises: An efficient detection module is used to transmit the images to be detected at the airport scene to the preset deep learning model for real-time and efficient detection of multiple targets to obtain the categories, unique IDs and physical coordinates of different objects; A risk prevention and control module, for performing risk prevention and control according to the category, unique ID and physical coordinates in combination with set judgment rules; wherein the judgment rules include unauthorized animal intrusion judgment based on category, safe distance judgment based on unique ID and high-risk area judgment based on physical coordinates in combination with dynamic density peak clustering technology; The processing module is used to quickly send an alarm notification to the designated security manager and implement the associated processing measures when the judgment result indicates that there is a risk.

9. An efficient airport scene real-time comprehensive risk prevention and control system as claimed in claim 8, characterized in that: The multi-target real-time and efficient detection comprises the following steps: Preprocess the image to be detected; The pre-processed image is transmitted to a preset deep learning model to complete multi-target real-time detection, and the category and pixel coordinates of each detection frame are output; wherein the deep learning model adopts a pre-trained YOLO model; Post-process the YOLO model to perform bounding box filtering and bounding box decoding; Finally, coordinate transformation and ReID processing are performed to obtain the category, unique ID and physical coordinates.

10. An efficient airport scene real-time comprehensive risk prevention and control system as claimed in claim 8, characterized in that: The high-risk area determination is completed by using a density peak clustering technique based on a dynamic threshold, and the specific steps include: Calculate the distance matrix; Determine the cutoff distance for calculating local density; Calculate local density; Calculate the relative distance of each point to the points with higher density; By drawing a decision graph, the horizontal axis represents the local density value and the vertical axis represents the relative distance value, and the points with high local density and relative distance value are selected as cluster centers; Assign other points to the nearest cluster center; According to the changes in different scenes and the density of people and objects, the threshold of aggregation judgment is dynamically adjusted; the dynamic threshold adjustment model is as follows: d c1 is the dynamically adjusted cutoff distance, d c0 is the initial cutoff distance, k is the adjustment coefficient, ρ m is the current highest density; Then, based on the clustering results and the dynamically adjusted threshold, determine whether it is a high-risk area.